Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Receptor sequence in the terminal protein of bacteriophage M2 that interacts with an RGD (Arg-Gly-Asp) sequence of the primer protein.

At the initiation of protein-primed DNA replication of bacteriophages M2 and phi 29, the Arg-Gly-Asp (RGD) sequence of primer protein participates in the recognition of terminal protein (TP), where the initiation site for protein-primed DNA replication of template DNA is located. We compared the sequences of M2 and phi 29 TP with those of the members of the integrin superfamily and found the highly homologous sequences Lys-Lys-Ile-Pro-Pro-Asp-Asp (KKIPPDD) in M2 and phi 29 TP and Lys-Lys-Gly-Cys-Pro-Pro-Asp-Asp (KKGCPPDD) in the beta-subunit of fibronectin receptor protein. A synthetic 20mer peptide that contained the KKIPPDD sequence interfered with the inhibitory effect of the RGD peptide on both transfection and the protein-priming reaction in vitro. We propose that the sequence KKIPPDD of M2 TP is the receptor sequence for RGD.

Amino Acid Sequence↗

The cDNA sequence of a human epidermal keratin: divergence of sequence but conservation of structure among intermediate filament proteins.

We have determined the DNA sequence of a cloned cDNA that is complementary to the mRNA for the 50 kilodalton (kd) human epidermal keratin. This provides the first amino acid sequence for a cytoskeletal keratin. Comparison of this sequence with those of other keratins reveals an evolutionary relationship between the cytoskeletal and the microfibrillar keratins, but shows no homology to matrix or feather keratins. The 50 kd keratin shares 28%-30% homology with partial sequences of other intermediate filament proteins, which suggests that keratins may be the most distantly related members of this class of fibrous proteins. Our computer analyses predict that the 50 kd keratin contains two long alpha-helical domains separated by a cluster of helix-inhibitory residues in the middle of the protein. These findings indicate that despite major sequence divergence among intermediate filament proteins, they retain sequences compatible with secondary structural features that appear to be common to all of them.

Amino Acid Sequence↗

Quantification analysis of 5'-splice signal sequences in mRNA precursors. Mutations in 5'-splice signal sequence of human beta-globin gene and beta-thalassemia.

Concerning the signals which direct excision of introns from mRNA precursors in higher eukaryotic genes, consensus 9-nucleotide sequence, (CA)AG/GT(AG)AGT, has been proposed with the 5'-splice site, but actual 5'-splice site sequences differ from it in a greater or lesser degree. We analyzed 5'-splice site sequence of human beta-globin gene by quantification method (categorical discriminant analysis) proposed previously. Analysis of 13-nucleotide sequences and deleted sequences showed that 9-nucleotide sequences in the consensus region are almost sufficient to define 5'-splice signal. To confirm this view, we examined a number of beta-globin mutant genes, where nucleotide changes occur at the authentic 5'-splice site of the first intron and cause beta-thalassemia phenotype. Our method could explain why such mutations abolish the 5'-splice site and cryptic 5'-splice sites are activated.

Base Sequence↗

mRNA sequence predictions from homologous protein sequences.

A model has been developed that permits the prediction of mRNA nucleic acid sequence from the sequences of the translated proteins. The model relies on the information obtained from the comparison of protein sequences in related species to reduce the number of possible codons for those amino acids where mutations are observed. The predictions so obtained have been tested by applying the model to proteins whose mRNA sequences are known. The model's predictions have been found to be 100% accurate if three or more different amino acids are known at a given position and if the protein sequences are restricted to relatively closely related species (within the same class). The use of this model may permit a reduction of the mRNA sequence degeneracy and therefore be helpful in the synthesis of cDNA probes or for the prediction of restriction endonuclease sites. Computer programs have been developed to ease the use of the model.

Amino Acid Sequence↗

Primary structure of two distinct rat pancreatic preproelastases determined by sequence analysis of the complete cloned messenger ribonucleic acid sequences.

The mRNA sequences for two rat pancreatic elastolytic enzymes have been cloned by recombinant DNA technology and their nucleotide sequences determined. Rat elastase I mRNA is 1113 nucleotides in length, plus a poly(A) tail, and encodes a preproelastase of 266 amino acids. The amino acid sequence of the predicted active form of rat elastase I is 84% homologous to porcine elastase 1. Key amino acid residues involved in determining substrate specificity of porcine elastase 1 are retained in the rat enzyme. The activation peptide of the zymogen does not appear related to that of other mammalian pancreatic serine proteases. The mRNA for elastase I is localized in the rough endoplasmic reticulum of acinar cells, as expected for the site of synthesis of an exocrine secretory enzyme. Rat elastase II mRNA is 910 nucleotides in length, plus a poly(A) tail, and encodes a preproenzyme of 271 amino acids. The amino acid sequence is more closely related to porcine elastase 1 (58% sequence identity) than to the other pancreatic serine proteases (33-39% sequence identity). Predictions of substrate preference based upon key amino acid residues that define the substrate binding cleft are consistent with the broad specificity observed for mammalian pancreatic elastase 2. The activation peptide is similar to that of the chymotrypsinogens and retains an N-terminal cysteine available to form a disulfide link to an internal conserved cysteine residue.

Amino Acid Sequence↗

Analysis of the P1 gene sequences and the 3'-terminal sequences and secondary structures of the single-stranded RNA genome of Potato virus V.

The immunocapture reverse transcriptase PCR method (IC-RT-PCR) was used to selectively amplify specific genome sequences of an isolate of Potato virus V (PVV, genus Potyvirus) from a potato plant infected by multiple viruses in the field in Finland. The sequences of the 5'- and 3'-non-translated regions (NTR) and the P1-and coat protein (CP)-encoding sequences were determined because they are the most variable genomic regions in potyviruses. The sequences of the new Finnish PVV isolate obtained in this study were compared to the sequences of eight PVV isolates characterized from other European countries. The results indicated little genetic variability among the European PVV isolates (nt identity values > 96% for all regions examined). Most PVV isolates were grouped according to their geographical origin using P1-sequences for a phylogenetic analysis. The nucleotide substitution patterns for P1 and the CP genes revealed that variability is due to a random genetic drift. Comparison of the secondary structure predictions for the unusually long 3'-non-translated region (3'-NTR) of PVV to the 3'-NTR of Tobacco etch virus (TEV) and other potyviruses revealed that some of the structures defining regions crucial for viral genome replication in TEV are present in PVV and other potyviruses.

Base Sequence↗

DNA sequence around the Escherichia coli unc operon. Completion of the sequence of a 17 kilobase segment containing asnA, oriC, unc, glmS and phoS.

The nucleotide sequence is described of a region of the Escherichia coli chromosome extending from oriC to phoS that also includes the loci gid, unc and glmS. Taken with known sequences for asnA and phoS this completes the sequence of a segment of about 17 kilobases or 0.4 min of the E. coli genome. Sequences that are probably transcriptional promoters for unc and phoS can be detected and the identity of the unc promoter has been confirmed by experiments in vitro with RNA polymerase. Upstream of the promoter sequence is an extensive region that appears to be non-coding. Conserved sequences are found that may serve to concentrate RNA polymerase in the vicinity of the unc promoter. Hairpin loop structures resembling known rho-independent transcription termination signals are evident following the unc operon and glmS. The glmS gene encoding the amidotransferase, glucosamine synthetase, has been identified by homology with glutamine 5-phosphoribosylpyrophosphate amidotransferase.

Amidophosphoribosyltransferase↗

Planarian mitochondria sequence heterogeneity: relationships between the type of cytochrome c oxidase subunit I gene sequence, karyotype and genital organ.

Freshwater planarians Dugesia japonica from three localities were examined for cytochrome c oxidase subunit I (COI) gene sequence, karyotype and the presence of genital organ. The planarians from Mt Fujiwara in Japan were composed of two different groups; one revealed inter- and intraindividual COI gene heterogeneity, while another revealed no sequence heterogeneity. The sequence in planarians from Mt Alishan in Taiwan was homogeneous, while that from the Kenting National Park in Taiwan revealed a considerable heterogeneity. All the planarians having the homogeneous gene sequences carry the 2X karyotype and many of them had genital organs. These are assumed to belong to the sexual lineage. In contrast, almost all planarians having heterogeneous sequences carry the karyotype of either 3X plus 2X (mixoploid) or 3X, and all of them lack genital organs. These lineages are assumed to be asexual. The heterogeneity of COI gene sequences in the presumed asexual lineages would have resulted from an accumulation of mutations by repeated asexual reproduction.

Animals↗

Complete sequence of a chicken lambda light chain immunoglobulin derived from the nucleotide sequence of its mRNA.

Recombinant cDNA plasmids have been constructed from chicken spleen poly(A)-containing RNA. Two clones have been selected and provide the sequence determination of a chicken lambda immunoglobulin light chain: they include the complete variable, constant, and 3' untranslated regions of the chicken lambda light chain mRNA and part of the leader sequence. Comparison of the chicken light chain constant region with both human and mouse lambda constant sequences indicates 61% homology at the amino acid level. Unexpectedly, the chicken variable sequence is 53-63% homologous to human variable sequences when it is compared to the various lambda subgroups and only 42% homologous to the mouse V lambda 1 sequence. The degree of homology between the variable regions of these three species does not easily correlate with their phylogenetic relationship.

Amino Acid Sequence↗

Nucleotide sequence of cDNA and derived amino acid sequence of human complement component C9.

The nucleotide sequence coding for the ninth component of human complement (C9) has been determined and the corresponding amino acid sequence has been derived. A human liver cDNA library was screened by the colony-hybridization technique using two radiolabeled oligonucleotide probes that correspond to known regions of the C9 amino acid sequence. Two recombinant plasmids were isolated and their cDNA inserts were sequenced. The derived protein sequence consists of 537 amino acids in a single polypeptide chain. A profile of the hydropathic index versus sequence number indicates that the amino-terminal half of C9 is predominantly hydrophilic in character whereas the carboxyl-terminal section of this protein is more hydrophobic. The amphipathic organization of the primary structure of C9 is consistent with the known potential of polymerized C9 to penetrate lipid bilayers, causing the formation of transmembrane channels.

Amino Acid Sequence↗

Complete nucleotide sequence of cDNA and deduced amino acid sequence of rat liver catalase.

We have isolated five cDNA clones for rat liver catalase (hydrogen peroxide:hydrogen peroxide oxidoreductase, EC 1.11.1.6). These clones overlapped with each other and covered the entire length of the mRNA, which had been estimated to be 2.4 kilobases long by blot hybridization analysis of electrophoretically fractionated RNA. Nucleotide sequencing was carried out on these five clones and the composite nucleotide sequence of catalase cDNA was determined. The 5' noncoding region contained 83 bases and was followed by 1581 bases of an open reading frame that encoded 527 amino acids. The 3' noncoding region was 831 bases long and contained long repeats of the unit AC. The amino acid sequence deduced from the nucleotide sequence of the cDNAs showed about 90% homology with the reported primary structure of bovine liver catalase. The molecular weight of rat liver catalase was calculated to be 59,758 from the predicted amino acid sequence. The amino acid residues in contact with the heme group are completely identical for bovine liver and rat liver catalases. The amino acid sequence at the COOH terminus was confirmed by the results of carboxypeptidase P treatment of the protein purified from rat liver in the presence of leupeptin. Rat liver catalase has no cleavable signal peptide for translocation of the enzyme into peroxisomes.

Amino Acid Sequence↗

Isolation and partial nucleotide sequence of the laccase gene from Neurospora crassa: amino acid sequence homology of the protein to human ceruloplasmin.

The laccase (benzenediol:oxygen oxidoreductase, EC 1.10.3.2) gene from Neurospora crassa was cloned and part of its nucleotide sequence corresponding to the carboxyl-terminal region of the protein has been determined. The gene was cloned by cDNA synthesis with a laccase-specific synthetic deoxyundecanucleotide as primer and poly(A) RNA isolated from cycloheximide-treated N. crassa cultures as template. Based on the nucleotide sequence of the cDNA obtained, a unique 21-mer was synthesized and used to screen a genomic DNA library from N. crassa. Five different positive clones were isolated and shown to share an overlapping DNA region with the same pattern of restriction sites. Sequence analysis of the common 1.36-kilobase Sal I fragment revealed an open reading frame of 726 nucleotides. The amino acid sequence deduced is in complete agreement with the primary structures of several tryptic peptides isolated previously from N. crassa laccase. The analyzed carboxyl-terminal region of laccase exhibits a striking sequence homology to the carboxyl-terminal part of the third homology unit of the multicopper oxidase ceruloplasmin and to a smaller extent, to the low molecular weight blue copper proteins plastocyanin and azurin. Based on amino acid sequence comparison between these proteins, putative copper ligands of N. crassa laccase are proposed. Moreover, these data further support the hypothesis that the small blue copper proteins and the multicopper oxidases have evolved from the same ancestral gene.

Amino Acid Sequence↗

Intron sequence directs RNA editing of the glutamate receptor subunit GluR2 coding sequence.

The Ca2+ permeability and the rectifying properties of the glutamate receptors assembled from the subunits GluR1-GluR4 depend upon a critical Arg in the GluR2 subunit located in a domain that has been proposed to span the membrane. The GluR2 subunit gene encodes a Gln (CAG) at this position, whereas the mRNA is edited so that it encodes an Arg (CGG) at this position [Sommer, B., Kohler, M., Sprengel, R. & Seeburg, P. H. (1991) Cell 67, 11-20]. The editing process is specific since only the GluR2 subunit RNA is edited even though the GluR1, GluR3, and GluR4 RNAs have a similar sequence. We show that this selective RNA editing depends upon a critical intron sequence in the GluR2 gene. This critical intron sequence is sufficient to cause editing of the GluR3 subunit exon in a chimera minigene constructed so that the GluR3 exon is placed upstream to the GluR2 intron sequence. Transfections of a neuronal cell line, N2a, with minigene constructs encoding different fragments of the GluR2 gene demonstrate that the 5' part of the 3' intron is essential for editing. Part of the exon and this critical intron sequence contains an inverted repeat that can fold into a structure consisting of three helical elements. Similar conclusions were reached by Higuchi, M., Single, F. n., Köhler, M., Sommer, B., Sprengel, R. & Seeburg, P. H. [(1993) Cell 75, 1361-1370]. These experiments demonstrate that the low Ca2+ permeability of the ionotropic non-N-methyl-D-aspartate glutamate receptors depends upon RNA editing, which requires a sequence in an intron 3' to the exon.

Algorithms↗

Sequence tag identification of intact proteins by matching tanden mass spectral data against sequence data bases.

Molecular and fragment ion data of intact 8- to 43-kDa proteins from electrospray Fourier-transform tandem mass spectrometry are matched against the corresponding data in sequence data bases. Extending the sequence tag concept of Mann and Wilm for matching peptides, a partial amino acid sequence in the unknown is first identified from the mass differences of a series of fragment ions, and the mass position of this sequence is defined from molecular weight and the fragment ion masses. For three studied proteins, a single sequence tag retrieved only the correct protein from the data base; a fourth protein required the input of two sequence tags. However, three of the data base proteins differed by having an extra methionine or by missing an acetyl or heme substitution. The positions of these modifications in the protein examined were greatly restricted by the mass differences of its molecular and fragment ions versus those of the data base. To characterize the primary structure of an unknown represented in the data base, this method is fast and specific and does not require prior enzymatic or chemical degradation.

Amino Acid Sequence↗

Heteroduplex mobility assay-guided sequence discovery: elucidation of the small subunit (18S) rDNA sequences of Pfiesteria piscicida and related dinoflagellates from complex algal culture and environmental sample DNA pools.

The newly described heterotrophic estuarine dinoflagellate Pfiesteria piscicida has been linked with fish kills in field and laboratory settings, and with a novel clinical syndrome of impaired cognition and memory disturbance among humans after presumptive toxin exposure. As a result, there is a pressing need to better characterize the organism and these associations. Advances in Pfiesteria research have been hampered, however, by the absence of genomic sequence data. We employed a sequencing strategy directed by heteroduplex mobility assay to detect Pfiesteria piscicida 18S rDNA "signature" sequences in complex pools of DNA and used those data as the basis for determination of the complete P. piscicida 18S rDNA sequence. Specific PCR assays for P. piscicida and other estuarine heterotrophic dinoflagellates were developed, permitting their detection in algal cultures and in estuarine water samples collected during fish kill and fish lesion events. These tools should enhance efforts to characterize these organisms and their ecological relationships. Heteroduplex mobility assay-directed sequence discovery is broadly applicable, and may be adapted for the detection of genomic sequence data of other novel or nonculturable organisms in complex assemblages.

Animals↗

Complete nucleotide sequence of a cDNA derived from calf lens gamma-crystallin mRNA: presence of Alu I-like DNA sequences.

The nucleotide sequence of a cloned cDNA derived from gamma-crystallin mRNA of calf lens was determined. The cloned cDNA contains the entire coding region 522 bp long, 30 nucleotides of the 5' noncoding region, and 67 residues in the 3' noncoding region followed by a poly(A) tail of 25 nucleotides. The deduced amino acid sequence directly demonstrates for the first time that the calf gamma-crystallin contains 174 residues. The nucleotide sequence contains a number of interesting features including a 32-bp sequence in the 3' region with 70% complementarity to the 3' end of the first monomer unit of the consensus Alu I DNA. Within this region, a 32-bp sequence shows about 80% homology with a segment of hamster 4.5S RNA. The possible evolutionary and regulatory significance of these sequences is discussed.

Amino Acid Sequence↗

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

DNA sequence analysis: a general, simple and rapid method for sequencing large oligodeoxyribonucleotide fragments by mapping.

Several electrophoretic and chromatographic systems have been investigated and compared for sequence analysis of oligodeoxyribonucleotides. Three systems were found to be useful for the separation of a series of sequential degradation products resulting from a labeled oligonucleotide: (I) 2-D electrophoresisdagger; (II) 2-D PEI-cellulose; and (III) 2-D homochromatography. System (III) proved generally most informative regardless of base composition and sequence. Furthermore, only in this system will the omission of an oligonucleotide in a series of oligonucleotides be self-evident from the two-dimensional map. The sequence of up to fifteen nucleotides can be determined solely by the characteristic mobility shifts of its sequential degradation products distributed on the two-dimensional map. With this method, ten nucleotides from the double-stranded region adjacent to the left-hand 3'-terminus and seven from the right-hand 3'-terminus of bacteriophage lambda DNA have been sequenced. Similarly, nine nucleotides from the double-stranded region adjacent to the left-hand 3'-terminus and five nucleotides from the right-hand terminus of bacteriophage phi80 DNA have also been sequenced. The advantages and disadvantages of each separation system with respect to sequence analysis are discussed.

Base Sequence↗