Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Coronavirus IBV: partial amino terminal sequencing of spike polypeptide S2 identifies the sequence Arg-Arg-Phe-Arg-Arg at the cleavage site of the spike precursor propolypeptide of IBV strains Beaudette and M41.

The spike protein of avian infectious bronchitis coronavirus comprises two glycopolypeptides S1 and S2 derived by cleavage of a proglycopolypeptide So, the nucleotide sequence of which has recently been determined for the Beaudette strain (Binns, M.M. et al., 1985, J. Gen. Virol. 66, 719-726). The order of the two glycopolypeptides within So is aminoterminus(N)-S1-S2-carboxyterminus(C). To locate the N-terminus of S2 we have performed partial amino acid sequencing on S2 from IBV-Beaudette labelled with [3H]serine and from the related strain labelled with [3H]valine, leucine and isoleucine. The residues identified and their positions relative to the N-terminus of S2 were: serine, 13; valine, 6, 12; leucine, none in the first 20 residues; isoleucine, 2, 19. These results identified the N-terminus of S2 of IBV-Beaudette as serine, 520 residues from the N-terminus of S1, excluding the signal sequence. Immediately to the N-terminal side of residue 520 So has the sequence Arg-Arg-Phe-Arg-Arg; similar basic connecting peptides are a feature of several other virus spike glycoproteins. It was deduced that for IBV-Beaudette S1 comprises 519 residues (Mr 57.0K) or 514 residues (56.2K) if the connecting peptide was to be removed by carboxypeptidase-like activity in vivo while S2 has 625 residues (69.2K). Nucleotide sequencing of the cleavage region of the So gene of IBV-M41 revealed the same connecting peptide as IBV-Beaudette and that the first 20 N-terminal residues of S2 of IBV-M41 were identical to those of the Beaudette strain. IBV-Beaudette grown in Vero cells had some uncleaved So; this was cleavable by 10 micrograms/ml of trypsin and of chymotrypsin. Partial N-terminal analysis of S1 from IBV-M41 identified leucine and valine residues at positions 2 and 9 respectively from the N-terminus. This confirms the identification, made by Binns et al. (1985), of the N-terminus of S1 and the end of the signal sequence of the IBV-Beaudette spike propolypeptide. N-terminal sequencing of [3H]leucine-labelled IBV-Beaudette membrane (M) polypeptide showed leucine residues at positions 8, 16 and 22 from the N-terminus; these results confirm the open reading frame identified by M.E.G. Boursnell et al. (1984, Virus Res. 1, 303-313) in the nucleotide sequence of M. The N-terminus of the nucleocapsid (N) polypeptide appeared to be blocked.

Amino Acid Sequence↗

Insertion of a short Alu sequence into the hMSH2 gene following a double cross over next to sequences with chi homology.

Alu repeat sequences and other multiple copy repetitive elements are present throughout the human genome and are active in promoting recombination. It is believed that reverse transcription of transcribed Alu repeats followed by chromosomal integration has been responsible for the wide dispersion and high copy number of these sequences. During studies on the hMSH2 gene we have used RT-PCR to amplify from peripheral blood lymphocytes a cDNA species in which 553 base pairs of hMSH2 cDNA have been deleted to be replaced by a short 36 base pair Alu sequence as a result of a genomic insertion/deletion event. The 36 base pair Alu insert is homologous to a 26 base pair Alu sequence previously implicated in the promotion of recombination and contains the GCTGG motif which is part of the prokaryotic chi sequence. A second chi-like sequence is also located within the deleted hMSH2 region. Both chi-like sequences are located within 4 bp of the two 4-bp regions of cross over containing the insertion/deletion breakpoints. This suggest that a double recombination event has occurred, providing direct evidence for the recombinogenic activity of this Alu element. Furthermore, it suggests that chi-like sequences may define recombination hotspots as in prokaryotes.

DNA, Complementary↗

Relation between mRNA expression and sequence information in Desulfovibrio vulgaris: combinatorial contributions of upstream regulatory motifs and coding sequence features to variations in mRNA abundance.

The context-dependent expression of genes is the core for biological activities, and significant attention has been given to identification of various factors contributing to gene expression at genomic scale. However, so far this type of analysis has been focused either on relation between mRNA expression and non-coding sequence features such as upstream regulatory motifs or on correlation between mRNA abundance and non-random features in coding sequences (e.g., codon usage and amino acid usage). In this study multiple regression analyses of the mRNA abundance and all sequence information in Desulfovibrio vulgaris were performed, with the goal to investigate how much coding and non-coding sequence features contribute to the variations in mRNA expression, and in what manner they act together. Using the AlignACE program, 442 over-represented motifs were identified from the upstream 100bp region of 293 genes located in the known regulons. Regression of mRNA expression data against the measures of coding and non-coding sequence features indicated that 54.1% of the variations in mRNA abundance can be explained by the presence of upstream motifs, while coding sequences alone contribute to 29.7% of the variations in mRNA abundance. Interestingly, most of contribution from coding sequences is overlapping with that from upstream motifs; thereby a total of 60.3% of the variations in mRNA abundance can be explained when coding and non-coding information was included. This result demonstrates that upstream regulatory motifs and coding sequence information contribute to the overall mRNA expression in a combinatorial rather than an additive manner.

Base Sequence↗

Sequence-based typing of the complete coding sequence of DQA1 and phenotype frequencies in the Dutch Caucasian population.

Typing of DQA1 by sequencing has been a challenge because of a 3-nucleotide deletion in exon 2 in half of the alleles. Furthermore, 19 of the 28 alleles cannot be identified on basis of exon 2 alone, but need additional exon information. With the sequencing strategy presented here the complete exons 1-4 are sequenced heterozygously, enabling identification of all DQA1 alleles by sequence-based typing (SBT). Exons 1-4 were amplified and sequenced separately, the combined sequences were used for automated allele assignment. The method was validated by typing 21 individuals with all possible different allele group combinations. In addition 26 quality control samples were correctly typed by this method. To determine the phenotype frequencies 155 unrelated Dutch Caucasian individuals were DQA1 typed. In total 15 known and two new DQA1 alleles were identified. DQA1*0103 and *0505 were the most frequent alleles with phenotype frequencies of 30% and 29%, respectively. The SBT method presented here is an improvement compared to already existing protocols in that the complete exon sequence is obtained for all coding exons, using identical polymerase chain reaction conditions. Furthermore, all exons are sequenced heterozygously, facilitating allele assignment and reducing the number of amplification reactions.

Alleles↗

Initiation of a Sarcocystis neurona expressed sequence tag (EST) sequencing project: a preliminary report.

To accelerate genetic and molecular characterization of Sarcocystis neurona, the primary causative agent of equine protozoal myeloencephalitis (EPM), a sequencing project has been initiated that will generate approximately 7000-8000 expressed sequence tags (ESTs) from this apicomplexan parasite. Poly(A)(+) RNA was isolated from culture-derived S. neurona merozoites, and a cDNA library was constructed in a unidirectional lambda phage cloning vector. Sixty phage clones were randomly picked from the library, and the cDNA inserts were amplified from these clones using the T3 and T7 primers that flank the multi-cloning site of the lambda vector. This analysis demonstrated that 100% (60/60) of the clones selected from this library contained recombinant cDNA inserts ranging in size from 0.4 to 4.0 kilobases (kb) with an average size of 1.23kb. Single-pass sequencing from the 5' end of the 60 amplified cDNAs produced high-quality nucleotide sequence from 53 of the clones. Comparison of these ESTs to the current gene databases revealed significant matches for 10 of the ESTs, six of which are similar to sequences from other Apicomplexa (i.e., Toxoplasma gondii). Importantly, none of the ESTs were of obvious mammalian origin, thus indicating that the cDNAs in this library were derived primarily from parasite mRNA and not from mRNA of the bovine turbinate host cells. Collectively, these data indicate that the described cDNA library will provide an excellent substrate for generating a portion of the ESTs that are planned from S. neurona. This sequencing project will greatly hasten gene discovery for this protozoan pathogen thereby enhancing efforts towards the development of improved diagnostics, treatments, and preventatives for EPM. In addition, the S. neurona ESTs will represent a significant contribution to the extensive database of sequences from the Apicomplexa. Comparative analyses of these apicomplexan sequences will likely offer a multitude of important information about the biology and evolutionary history of this phylogenetic grouping of parasites.

Animals↗

The accuracy of DNA sequences: estimating sequence quality.

In this paper we describe a method for the statistical reconstruction of a large DNA sequence from a set of sequenced fragments. We assume that the fragments have been assembled and address the problem of determining the degree to which the reconstructed sequence is free from errors, i.e., its accuracy. A consensus distribution is derived from the assembled fragment configuration based upon the rates of sequencing errors in the individual fragments. The consensus distribution can be used to find a minimally redundant consensus sequence that meets a prespecified confidence level, either base by base or across any region of the sequence. A likelihood-based procedure for the estimation of the sequencing error rates, which utilizes an iterative EM algorithm, is described. Prior knowledge of the error rates is easily incorporated into the estimation procedure. The methods are applied to a set of assembled sequence fragments from the human G6PD locus. We close the paper with a brief discussion of the relevance and practical implications of this work.

Algorithms↗

Characteristic sequences for DNA primary sequence.

A DNA sequence can be identified with a word over an alphabet N = [A, C, G, T]. Characteristic sequences of a DNA sequence are given in term of classifications of bases of nucleic acids. Using the characteristic sequences, we construct a set of 2 x 2 matrices to represent DNA primary sequences, which are based on counting of the frequency of occurrence of all (0,1) triplets of characteristic sequences. Furthermore, the leading eigenvalues of these matrices are computed and considered as invariants for the DNA primary sequences. Similarity and dissimilarity analysis based on the characteristic sequences are given for eight exon-1 genes of beta-globin about eight species.

Animals↗

Sequence of terminal regions of cowpox virus DNA: arrangement of repeated and unique sequence elements.

One terminal EcoRI fragment of the genome of cowpox virus (CPV) strain Brighton red has been cloned in plasmid pBR325, and the nucleotide sequence of the 2,725-base-pair Sal I fragment corresponding to that at the end of the viral genome has been determined. The fragment consists of three unique sequence regions flanking two sets of repeated sequence. The repeated sequence sets are composed of four types of subunits, the majority of which are arranged in higher-order repeat units. The subunits are themselves closely related; two are subsets of a third, whereas the fourth is a recombinant of the first two. The fragment possesses no long open reading frames (maximal coding potential, 65 amino acids). The sequence of this CPV DNA Sal I fragment is compared with that of the corresponding fragment of vaccinia virus WR DNA [Baroudy, B. M., Venkatesan, S. & Moss, B. (1982) Cell 28, 315-324; Venkatesan, S., Baroudy, B. M. & Moss, B. (1981) Cell 25, 805-813]. Two of the unique sequence regions of the two viruses are related to the extent of 96%, and the third contains at least one sequence of 112 residues that is 98% homologous. As for the repeated sequence sets, those of vaccinia virus are composed of only two, rather than four, types of subunit, one of which is identical to one of the CPV subunits, whereas the other differs from another CPV subunit by only three mismatches and one deletion. However, the arrangement of subunits in the two viruses is different, that in vaccinia virus DNA being simpler. Both subunits as well as repeat units probably arose as a result of unequal crossover.

Base Sequence↗

Cloning and sequence analysis of rat bone sialoprotein (osteopontin) cDNA reveals an Arg-Gly-Asp cell-binding sequence.

The primary structure of a bone-specific sialoprotein was deduced from cloned cDNA. One of the cDNA clones isolated from a rat osteosarcoma (ROS 17/2.8) phage lambda gt11 library had a 1473-base-pair-long insert that encoded a protein with 317 amino acid residues. This cDNA clone appears to represent the complete coding region of sialoprotein mRNA, including a putative AUG initiation codon and a signal peptide sequence. The amino acid sequence deduced from the cDNA contains several Ser-Xaa-Glu sequences, possibly representing attachment points for O-glycosidically linked oligosaccharides and one Asn-Xaa-Ser sequence representing a likely site for the N-glycosidically linked oligosaccharide. An interesting observation is the Gly-Arg-Gly-Asp-Ser sequence, which is identical to the cell-binding sequence identified in fibronectin. The presence of this sequence prompted us to investigate the cell-binding properties of sialoprotein. The ROS 17/2.8 cells attached and attained a spread morphology on surfaces coated with sialoprotein. We could demonstrate that synthetic Arg-Gly-Asp-containing peptides efficiently inhibited the attachment of cells to sialoprotein-coated substrates. The results show that the Arg-Gly-Asp sequence also confers cell-binding properties on bone-specific sialoprotein. To better reflect the potential function of bone sialoprotein--we propose the name "osteopontin" for this protein.

Amino Acid Sequence↗

Tn10 insertion specificity is strongly dependent upon sequences immediately adjacent to the target-site consensus sequence.

Transposon Tn10 inserts preferentially into particular "hotspots" that have been shown by sequence analysis to contain the symmetrical consensus sequence 5'-GCTNAGC-3'. This consensus is necessary but not sufficient to determine insertion specificity. We have mutagenized a known hotspot to identify other determinants for insertion into this site. This genetic dissection of the sequence context of a protein binding site shows that a second major determinant for Tn10 insertion specificity is contributed by the 6-9 base pairs that flank each end of the consensus sequence. Variations in these context base pairs can confer variations of at least 1000-fold in insertion frequency. There is no discernible consensus sequence for the context determinant, suggesting that sequence-specific protein-DNA contacts are not playing a major role. Taken together with previous work, the observations presented suggest a model for the interaction of transposase with the insertion site: symmetrically disposed subunits bind with specific contacts to the major groove of consensus-sequence base pairs, while flanking sequences influence the interaction through effects on DNA helix structure. We also show that the determinants important for insertion into a site are not important for transposition out of that site.

Base Sequence↗

Defining target sequences of DNA-binding proteins by random selection and PCR: determination of the GCN4 binding sequence repertoire.

We developed a simple and accurate method to define the sequence recognition properties of DNA-binding proteins. The method employs polymerase chain reaction (PCR) amplification of sequences selected from a mixture of random oligonucleotides by the gel mobility-shift assay. We used this method to define the sequence requirement of the binding domain of the yeast transcriptional activator GCN4. Using a total of 200 ng of purified protein and four cycles of binding and subsequent amplification, we identified the TGA-(C/G)TCA sequence as the binding consensus of GCN4, which is consistent with the previously reported recognition sequence. In addition, our data indicate that GCN4 can bind with lower affinity to sequences that differ from the optimal sequence in one or even two positions. The most common variation was the C to A at position +2. The majority of the substitutions that still allowed binding were 3' to the central C residue indicating that the two sides of the palindromic recognition sequence are not equivalent.

Base Sequence↗

Extraordinarily stable mini-hairpins: electrophoretical and thermal properties of the various sequence variants of d(GCGAAAGC) and their effect on DNA sequencing.

A small DNA fragment having a characteristic sequence d(GCGAAAGC) has been shown to form an extraordinarily stable mini-hairpin structure and to have an unusually rapid mobility in polyacrylamide gel electrophoresis, even when containing 7M urea. Here, we have studied the stability of the various sequence variants of d(GCGAAAGC) and the corresponding RNA fragments. Many such sequence variants form stable mini-hairpins in a similar manner to the d(GCGAAAGC) sequence. The RNA fragment, r(GCGAAAGC) also forms a mini-hairpin structure with less stability. The DNA mini-hairpins with GAAA or GAA loop are much more stable than DNA and RNA mini-hairpins with other loop sequence so far as has been examined. The stability difference between DNA and RNA mini-hairpins may be deduced to the stem structures formed by DNA (B form) and RNA (A form). The stable hairpins consisting of the GCGAAAGC sequence cause strong band compression on the sequencing gel. This phenomenon should be carefully considered in DNA sequencing.

Base Sequence↗

A comparison of expressed sequence tags (ESTs) to human genomic sequences.

The Expressed Sequence Tag (EST) division of GenBank, dbEST, is a large repository of the data being generated by human genome sequencing centers. ESTs are short, single pass cDNA sequences generated from randomly selected library clones. The approximately 415 000 human ESTs represent a valuable, low priced, and easily accessible biological reagent. As many ESTs are derived from yet uncharacterized genes, dbEST is a prime starting point for the identification of novel mRNAs. Conversely, other genes are represented by hundreds of ESTs, a redundancy which may provide data about rare mRNA isoforms. Here we present an analysis of >1000 ESTs generated by the WashU-Merck EST project. These ESTs were collected by querying dbEST with the genomic sequences of 15 human genes. When we aligned the matching ESTs to the genomic sequences, we found that in one gene, 73% of the ESTs which derive from spliced or partially spliced transcripts either contain intron sequences or are spliced at previously unreported sites; other genes have lower percentages of such ESTs, and some have none. This finding suggests that ESTs could provide researchers with novel information about alternative splicing in certain genes. In a related analysis of pairs of ESTs which are reported to derive from a single gene, we found that as many as 26% of the pairs do not BOTH align with the sequence of the same gene. We suspect that some of these unusual ESTs result from artifacts in EST generation, and caution researchers that they may find such clones while analyzing sequences in dbEST.

Alternative Splicing↗

Characterisation of polyoma late mRNA leader sequences by molecular cloning and DNA sequence analysis.

The leader sequences of two of the three polyoma virus late mRNAs were characterised by molecular cloning and DNA sequence analysis. A short single-stranded DNA fragment complementary to the 5' end of the body of mVP1 was used to prime cDNA synthesis, and double-stranded cDNA was inserted into a derivative of pAT 153. Analysis of fourteen mVP1 cDNAs and two mVP3 cDNAs allowed the precise determination of the leader-body joints, and demonstrated that the majority of polyoma late leader sequences consist of exact tandem repeats of a 57 nucleotide sequence present only once in the genomic DNA at 66-67 m.u. The sequences in the genomic DNA borderline this unit are typical of those found at RNA splice points. Leader sequences contained on average three to four repeat units. Nuclease S1 mapping of total late mRNA demonstrated that most mRNA 5' ends map heterogeneously in the 50 nucleotides 5' to the repeated sequence unit. The structures of the leader sequences strongly suggest that they are generated by appropriate splicing events from a tandemly repeated transcript of the entire circular viral genome.

Base Sequence↗

Nucleotide sequence of complementary DNA and derived amino acid sequence of murine complement protein C3.

The nucleotide sequences coding for murine complement component C3 have been determined from a cloned genomic DNA fragment and several overlapping cloned complementary DNA fragments. The amino acid sequence of the protein was deduced. The mature beta and alpha subunits contain 642 and 993 amino acids respectively. Including a 24 amino acid signal peptide and four arginines in the beta-alpha transition region, which are probably not contained in the mature protein, the unglycosylated single chain precursor protein preproC3 would have a molecular mass of 186 484 Da and consist of 1663 amino acid residues. The C3 messenger RNA would be composed of a 56 +/- 2 nucleotide long 5' non-translated region, 4992 nucleotides of coding sequence, and a 3' non-translated region of 39 nucleotides, excluding the poly A tail. The beta chain contains only three cysteine residues, the alpha chain 24, ten of which are clustered in the carboxy terminal stretch of 175 amino acids. Two potential carbohydrate attachment sites are predicted for the alpha chain, none for the beta chain. From a comparison with human C3 cDNA sequence (of which over 80% has been determined) an extensive overall sequence homology was observed. Human and murine preproC3 would be of very similar length and share several noteworthy properties: the same order of the subunits in the precursor, the same basic residue multiplet in the beta-alpha transition region, and a glutamine residue in the thioester region. The equivalent position of the known factor I cleavage sites in human C3 alpha could be located in the murine C3 alpha chain and the size and sequence of the resulting peptide were deduced. A comparison of the amino acid sequences of murine C3 and human alpha 2-macroglobulin is given. Several areas of strong sequence homology are observed, and we conclude that the two genes must have evolved from a common ancestor.

Amino Acid Sequence↗

Sequence haplotypes revealed by sequence-tagged site fine mapping of the Ror1 gene in the centromeric region of barley chromosome 1H.

We describe the development of polymerase chain reaction-based, sequence-tagged site (STS) markers for fine mapping of the barley (Hordeum vulgare) Ror1 gene required for broad-spectrum resistance to powdery mildew (Blumeria graminis f. sp. hordei). After locating Ror1 to the centromeric region of barley chromosome 1H using a combined amplified fragment length polymorphism/restriction fragment-length polymorphism (RFLP) approach, sequences of RFLP probes from this chromosome region of barley and corresponding genome regions from the related grass species oat (Avena spp.), wheat, and Triticum monococcum were used to develop STS markers. Primers based on the RFLP probe sequences were used to polymerase chain reaction-amplify and directly sequence homologous DNA stretches from each of four parents that were used for mapping. Over 28,000 bp from 22 markers were compared. In addition to one insertion/deletion of at least 2.0 kb, 79 small unique sequence polymorphisms were observed, including 65 single nucleotide substitutions, two dinucleotide substitutions, 11 insertion/deletions, and one 5-bp/10-bp exchange. The frequency of polymorphism between any two barley lines ranged from 0.9 to 3.0 kb, and was greatest for comparisons involving an Ethiopian landrace. Haplotype structure was observed in the marker sequences over distances of several hundred basepairs. Polymorphisms in 16 STSs were used to generate genetic markers, scored by restriction enzyme digestion or by direct sequencing. Over 2,300 segregants from three populations were used in Ror1 linkage analysis, mapping Ror1 to a 0.2- to 0.5-cM marker interval. We discuss the implications of sequence haplotypes and STS markers for the generation of high-density maps in cereals.

Base Sequence↗

Sequence of the 5'-end quarter of the human-thyroglobulin messenger ribonucleic acid and of its deduced amino-acid sequence.

Thyroglobulin, the dimeric glycoprotein (19 S, 2 X 330 kDa), specific to the thyroid gland, is the support for thyroid hormone synthesis. Elucidation of the mechanism for thyroid hormone synthesis requires the knowledge of the primary sequence of the protein. In this paper the sequence of the first coding 2190 nucleotides from the 5' end of the human mRNA is presented. This was obtained by sequencing two previously described overlapping clones and by construction and sequencing of a single-stranded cDNA corresponding to the 5' end of the mRNA. The nucleotide sequence represents a quarter of the human thyroglobulin mRNA, from which a polypeptide sequence of 730 amino acids at the NH2-terminal end of the monomer has been deduced. This sequence shows a repetition of five highly conserved motifs each of approximately 50 amino acids, the analysis of which allowed us to establish a consensus sequence. We have also demonstrated (a) the hormonogenic tyrosine residue recently described in the mature protein, which is located four amino acids after the NH2-terminal Asn; (b) a prepeptide signal of thyroglobulin secretion comprising 19 amino acids preceding the Asn residue, the NH2-terminal residue of the mature protein and (c) a six-signal tripeptide (Asn-Xaa-Thr or Ser) of N-glycosylation of the chain.

Amino Acid Sequence↗

Sequence-specific complex formation of DNA and a eukaryotic sequence-specific endonuclease, SceI.

Endo.SceI is a eukaryotic sequence-specific endonuclease of 120 kDa that causes sequence-specific double-stranded scission of DNA. Unlike results with restriction enzymes, we found a consensus sequence around the cleavage sites for Endo.SceI instead of a common sequence. We searched for conditions for studying the binding of Endo.SceI to DNA other than cutting. Under optimized conditions including gel mobility shift assay, Endo.SceI exhibited sequence-specific binding to a short double-stranded DNA (41 base pairs) containing a cleavage site and the DNA reisolated from the protein-DNA complex was not cleaved. The analysis of the complex of Endo.SceI and DNA isolated by the gel mobility shift experiments showed that the DNA-binding entity in the Endo.SceI preparation does have Endo.SceI activity and consists of an equal amount of 75-kDa and 50-kDa polypeptides. Based on this observation and those from previous studies, we conclude that Endo.SceI is a heterodimer of the 75-kDa and 50-kDa subunits. Under the present assay conditions, Endo.SceI did not show binding to single-stranded DNA having the same sequence of either plus or minus strand of the double-stranded DNA containing the cleavage site (the 41-bp DNA). Endo.SceI showed significantly higher affinity for the consensus sequence than the major cleavage site in pBR322 DNA. Unlike the cleavage of DNA by Endo.SceI which requires Mg2+, this sequence-specific binding is independent of but stimulated by Mg2+.

Base Sequence↗