Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequences”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Molecular cloning and DNA sequence of rat amelogenin and a comparative analysis of mammalian amelogenin protein sequence divergence.

The developing rat incisor is a common model used in the study of enamel development. It has been impossible to study correlation between rat enamel structure and the sequence of the major developing enamel protein in this species as to date a DNA sequence for rat amelogenin has not been reported. This study presents the first cloning of a full-length cDNA copy of rat amelogenin and its deduced primary sequence. Detailed analysis of this sequence provides evidence that the gene has evolved by internal sequence duplication. Comparison of the rat amelogenin primary sequence with those published for other species provides evidence that this protein, while exhibiting extreme levels of sequence conservation, has been subject to significant structural changes that may be related to alterations in enamel structure in different mammalian groups.

Amelogenin↗

Sequence analysis of group B rotavirus gene 1 and definition of a rotavirus-specific sequence motif within the RNA polymerase gene.

The complete nucleic acid sequence was determined for the largest genomic segment of the IDIR strain of group B rotavirus and compared with RNA polymerase genes of rotavirus groups A and C as well as other RNA viruses. IDIR gene 1 contained 3509 bases with a single, long open reading frame which encoded a deduced polypeptide of 1159 amino acids (MW = 131.6 kDa; pl 8.851). The deduced amino acid sequence of IDIR gene 1 shared 50% similar sequences and 27.6% identical sequences with VP1 of the RF strain of group A rotavirus. IDIR gene 1 also contained 45.4% similar and 26.5% identical amino acid sequences in comparison with gene 1 of the Cowden strain of group C rotavirus. Amino acids 643-689 of IDIR gene 1 corresponded to the conserved viral RNA polymerase domains, "SG . . . T . . . NS . . N" and "GDD." Within these domains, group A, B, and C rotaviruses displayed substantial homologies that were not shared with other RNA viruses. These sequences indicated the presence of highly conserved structural or functional components among groups of rotaviruses which were otherwise quite heterogeneous. The identification of rotavirus-specific residues within RNA polymerase sequence may prove valuable in devising strategies aimed at the control of rotavirus replication and infection.

Amino Acid Sequence↗

Purification, amino acid sequence, and cDNA sequence of a novel calcium-precipitating proteolipid involved in calcification of corynebacterium matruchotii.

Corynebacterium matruchotii is a microbial inhabitant of the oral cavity associated with dental calculus formation. It produces membrane-associated proteolipid capable of inducing hydroxyapatite formation in vitro. This proteolipid was purified from chloroform:methanol extracts by chromatography on Sephadex LH-20 and migrated on SDS-polyacrylamide gel electrophoresis at 6-9 kDa. Removal of covalently attached acyl moieties by methanolic KOH decreased its molecular mass to approximately 5.5 kDa. The amino acid sequence of the apoproteolipid indicated a peptide of 50 amino acids, a calculated molecular weight of 5354 Da, and an isoelectric point of 4.28. Sequence analysis revealed an 8 amino acid sequence with homology to human phosphoprotein phosphatase 2A as well as several potential acylation sites and one phosphorylation site. The purified proteolipid induced calcium precipitation in vitro. Deacylation of the proteolipid by hydroxylamine treatment resulted in >50% loss of calcium-precipitating activity, suggesting that covalently attached lipids are required. Degenerate oligonucleotide primers, based on the amino acid sequence, were used to amplify the gene for the 5.5 kDa proteolipid from total chromosomal DNA of C. matruchotii by PCR. A 166 bp cDNA was isolated and sequenced, confirming the amino acid sequence of the proteolipid. Thus, we have sequenced a unique bacterial proteolipid that is involved in the formation of dental calculus by precipitating Ca2+ and possibly in transport of inorganic phosphate, necessary for hydroxyapatite formation.

Amino Acid Sequence↗

Isolation and sequencing of a gene coding for glyoxalase I activity from Salmonella typhimurium and comparison with other glyoxalase I sequences.

The glyoxalase I gene (gloA) from Salmonella typhimurium has been isolated in Escherichia coli on a multi-copy pBR322-derived plasmid, selecting for resistance to 3 mM methylglyoxal on Luria-Bertani agar. The region of the plasmid which confers the methylglyoxal resistance in E. coli was sequenced. The deduced protein sequence was compared to the known sequences of the Homo sapiens and Pseudomonas putida glyoxalase I (GlxI) enzymes, and regions of strong homology were used to probe the National Center for Biotechnology Information protein database. This search identified several previously known glyoxalase I sequences and other open reading frames with unassigned function. The clustal alignments of the sequences are presented, indicating possible Zn2+ ligands and active site regions. In addition, the S. typhimurium sequence aligns with both the N-terminal half and the C-terminal half of the proposed GlxI sequences from Saccharomyces cerevisiae and Schizosaccharomyces pombe, suggesting that the structures of the yeast enzymes are those of fused dimers.

Amino Acid Sequence↗

NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Amino Acid Sequence↗

Sequences annotated by structure: a tool to facilitate the use of structural information in sequence analysis.

With the aim of bridging the gap between protein sequence and structural analyses, we have developed a tool to aid the identification of new protein sequences by recognizing distant homologues using structural information. The tool generates sequence annotated by structure (SAS) files, applying structural information derived from structural analyses to a given protein sequence. A World Wide Web interface allows a given sequence to be submitted either for structural annotation or, where its structure is unknown, for search and alignment against sequences of known structure. In both cases, SAS will colour residues in the sequence of known structure according to a selection of properties, including secondary structure, interatomic contacts and active site information. SAS can also be used to inspect properties of a single structure.

Amino Acid Sequence↗

IS1631 occurrence in Bradyrhizobium japonicum highly reiterated sequence-possessing strains with high copy numbers of repeated sequences RSalpha and RSbeta.

From Bradyrhizobium japonicum highly reiterated sequence-possessing (HRS) strains indigenous to Niigata and Tokachi in Japan with high copy numbers of the repeated sequences RSalpha and RSbeta (K. Minamisawa, T. Isawa, Y. Nakatsuka, and N. Ichikawa, Appl. Environ. Microbiol. 64:1845-1851, 1998), several insertion sequence (IS)-like elements were isolated by using the formation of DNA duplexes by denaturation and renaturation of total DNA, followed by treatment with S1 nuclease. Most of these sequences showed structural features of bacterial IS elements, terminal inverted repeats, and homology with known IS elements and transposase genes. HRS and non-HRS strains of B. japonicum differed markedly in the profiles obtained after hybridization with all the elements tested. In particular, HRS strains of B. japonicum contained many copies of IS1631, whereas non-HRS strains completely lacked this element. This association remained true even when many field isolates of B. japonicum were examined. Consequently, IS1631 occurrence was well correlated with B. japonicum HRS strains possessing high copy numbers of the repeated sequence RSalpha or RSbeta. DNA sequence analysis indicated that IS1631 is 2,712 bp long. In addition, IS1631 belongs to the IS21 family, as evidenced by its two open reading frames, which encode putative proteins homologous to IstA and IstB of IS21, and its terminal inverted repeat sequences with multiple short repeats.

Amino Acid Sequence↗

Saccharomyces cerevisiae RAD5-encoded DNA repair protein contains DNA helicase and zinc-binding sequence motifs and affects the stability of simple repetitive sequences in the genome.

rad5 (rev2) mutants of Saccharomyces cerevisiae are sensitive to UV light and other DNA-damaging agents, and RAD5 is in the RAD6 epistasis group of DNA repair genes. To unambiguously define the function of RAD5, we have cloned the RAD5 gene, determined the effects of the rad5 deletion mutation on DNA repair, DNA damage-induced mutagenesis, and other cellular processes, and analyzed the sequence of RAD5-encoded protein. Our genetic studies indicate that RAD5 functions primarily with RAD18 in error-free postreplication repair. We also show that RAD5 affects the rate of instability of poly(GT) repeat sequences. Genomic poly(GT) sequences normally change length at a rate of about 10(-4); this rate is approximately 10-fold lower in the rad5 deletion mutant than in the corresponding isogenic wild-type strain. RAD5 encodes a protein of 1,169 amino acids of M(r) 134,000, and it contains several interesting sequence motifs. All seven conserved domains found associated with DNA helicases are present in RAD5. RAD5 also contains a cysteine-rich sequence motif that resembles the corresponding sequences found in 11 other proteins, including those encoded by the DNA repair gene RAD18 and the RAG1 gene required for immunoglobin gene arrangement. A leucine zipper motif preceded by a basic region is also present in RAD5. The cysteine-rich region may coordinate the binding of zinc; this region and the basic segment might constitute distinct DNA-binding domains in RAD5. Possible roles of RAD5 putative ATPase/DNA helicase activity in DNA repair and in the maintenance of wild-type rates of instability of simple repetitive sequences are discussed.

Adenosine Triphosphatases↗

Molecular cloning, sequencing, and chromosome mapping of a 1A-encoded omega-type prolamin sequence from wheat.

Gliadins are the most abundant component of the seed storage proteins in cereals and, in combination with glutenins, are important for the bread-making quality of wheat. They are divided into four subfamilies, the alpha-, beta-, gamma-, and omega-gliadins, depending on their electrophoresis pattern, chromosomal location, and DNA and protein structures. Using a PCR-based strategy we isolated and sequenced an omega-gliadin sequence. We also determined the chromosomal subarm location of this sequence using wheat aneuploids and deletion lines. The gene is 1858 bp long and contains a coding sequence 1248 bp in length. Like all other gliadin gene families characterized in cereals, the omega-gliadin gene described here had characteristic features including two repeated sequences 300 bp upstream of the start codon. At the DNA level, the gene had a high degree of similarity to the omega-secalin and C-hordein genes of rye and barley, but exhibited much less homology to the alpha- and beta-gliadin gene families. In terms of the deduced amino acid sequence, this gene has about 80 and 70% similarity to the omega-secalin and C-hordein genes, respectively, and possesses all the features reported for other gliadin gene families. The omega-gliadin gene has about 30 repeats of the core consensus sequences PQQPX and XQQPQQX, twice as many as other gliadin gene families. Southern blotting and PCR analysis with aneuploid and deletion lines for the short arm of chromosome 1A showed that the omega-gliadin was located on the distal 25% of the short arm of chromosome 1A. By comparison of PCR and A-PAGE profiles for deletion stocks, its genomic location must be at a different locus from gli-Ala in 'Chinese Spring'.

Amino Acid Sequence↗

An in silico mining for simple sequence repeats from expressed sequence tags of zebrafish, medaka, Fundulus, and Xiphophorus.

Teleost fish genome projects involving model species are resulting in a rapid accumulation of genomic and expressed DNA sequences in public databases. The expressed sequence tags (ESTs) collected in the databases can be mined for the analysis of both structural and functional genomics. In this study, we in silico analyzed 49,430 unigenes representing a total of 692,654 ESTs from four model fish for their potential use in developing simple sequence repeats (SSRs), or microsatellites. After bioinformatical mining, a total of 3,018 EST derived SSRs (EST-SSRs) were identified for 2,335 SSR containing ESTs (SSR-ESTs). The frequency of identified SSR-ESTs ranged from 1.5% for Xiphophorus to 7.3% for zebrafish. The dinucleotide repeat motif is the most abundant SSR, accounting for 47%, 52%, 64%, and 78% for medaka, Fundulus, zebrafish, and Xiphophorus, respectively. Simulation analysis suggests that a majority of these EST-SSRs have sufficient flanking sequences for polymerase chain reaction (PCR) primer design. Comparative DNA sequence analyses of SSR-ESTs identified several cross-species SSRs and sequences that may be used as cross-reference genes in comparative studies. For example, the flanking sequences of one SSR (CTG)n within the pituitary tumor-transforming gene (PTTG) 1 interacting protein (PTTGIP), showed conservation spanning the medaka, Fundulus, human, and mouse genomes. This study provides a large body of information on EST-SSRs that can be useful for the development of polymorphic markers, gene mapping, and comparative genome analysis. Functional analysis of these SSR-ESTs may reveal their role in metabolism and gene evolution of these model species.

Amino Acid Sequence↗

Complete genomic sequence of the murine low affinity Fc receptor for IgE. Demonstration of alternative transcripts and conserved sequence elements.

The complete sequence of the murine low affinity Fc receptor for IgE (Fc epsilon RII), including the 5' and 3' flanking sequences, is reported. The murine Fc epsilon RII gene spans 12.9 kb and includes 12 exons surrounding 11 introns. The composite exon sequence is virtually identical to previously reported murine Fc epsilon RII cDNA sequences. Much of the proximal promoter regions of the mouse and human homologues of Fc epsilon RII show remarkable homology to each other, including three promoter elements previously identified for MHC class II genes. The reported exon/intron structure of the human FC epsilon RII is similar to the murine homologue, except that the latter has an additional exon coding for a fourth amino acid repetitive sequence (vs three in the human gene). RNase protection studies have identified an additional transcript within intron 2 of murine Fc epsilon RIIa, similar to the human Fc epsilon RIIb form but with a different predicted sequence of the first six amino acids. This transcript is present in the mRNA of purified splenic B cells, but not in the mRNA of the Fc epsilon RII+ B lymphoma cell line M12.4.5. The murine Fc epsilon RII gene contains a large intron (4.2 kb) separating the lectin and nonlectin coding regions, and several repetitive sequences are found clustered within this intron. These results emphasize the importance of the demarcation between these domains and allude to their evolutionary and functional significance.

Amino Acid Sequence↗

Sequence motifs and free energies of selected natural and non-natural nucleosome positioning DNA sequences.

Our laboratories recently completed SELEX experiments to isolate DNA sequences that most-strongly favor or disfavor nucleosome formation and positioning, from the entire mouse genome or from even more diverse pools of chemically synthetic random sequence DNA. Here we directly compare these selected natural and non-natural sequences. We find that the strongest natural positioning sequences have affinities for histone binding and nucleosome formation that are sixfold or more lower than those possessed by many of the selected non-natural sequences. We conclude that even the highest-affinity sequence regions of eukaryotic genomes are not evolved for the highest affinity or nucleosome positioning power. Fourier transform calculations on the selected natural sequences reveal a special significance for nucleosome positioning of a motif consisting of approximately 10 bp periodic placement of TA dinucleotide steps. Contributions to histone binding and nucleosome formation from periodic TA steps are more significant than those from other periodic steps such as AA (=TT), CC (=GG) and more important than those from the other YR steps (CA (=TG) and CG), which are reported to have greater conformational flexibility in protein-DNA complexes even than TA. We report the development of improved procedures for measuring the free energies of even stronger positioning sequences that may be isolated in the future, and show that when the favorable free energy of histone-DNA interactions becomes sufficiently large, measurements based on the widely used exchange method become unreliable.

Animals↗

Synergistic effect of upstream sequences, CCAAT box elements, and HSE sequences for enhanced expression of chimaeric heat shock genes in transgenic tobacco.

The thermoregulated expression of the soybean heat shock (hs) gene Gmhsp17.3-B is regulated via the heat shock promoter elements (HSEs), but full promoter activity requires additional sequences located upstream of the HSE-containing region. Structural features within this putative enhancer region include a run of simple sequences which are also present upstream of HSE-like sequences of other soybean hs genes, and three perfect CCAAT box sequences located immediately upstream from the most distal HSE of the promoter. A series of heterologous and homologous promoter fusions linked to the chloramphenicol acetyl transferase (CAT) gene was constructed and examined in transgenic tobacco plants. The region containing the AT-rich domain of the 5' flanking region was unable to direct transcription from the TATA box of a truncated delta CaMV35S promoter. Heat-inducible CAT activity was detectable when additional sequences from the native promoter containing three CCAAT boxes and a single HSE were present in the constructions. Complete reconstitution of the native hs promoter/enhancer region increased hs specific CAT activities only very little, but deletion of CCAAT box sequences reduced CAT expression five-fold. Our results suggest that AT-rich sequences have a moderate effect on thermoinducible expression levels of the soybean heat shock gene and that CCAAT box sequences act cooperatively with HSEs to increase the hs promoter activity.

Base Sequence↗

Intraspecies analysis: comparison of ITS sequence data and gene intron sequence data with breeding data for a worldwide collection of Gonium pectorale.

The morphologically uniform species Gonium pectorale is a colonial green flagellate of worldwide distribution. The affinities of 25 isolates from 18 sites on five continents were assessed by both DNA sequence comparisons and sexual compatibility. Complete sequences were obtained (i) for the internal transcribed spacer ITS-1 and ITS-2 regions of ribosomal DNA and (ii) for each of three single-copy spliceosomal introns, two in a small G protein and one in the actin gene. ITS sequences appeared to homogenize sufficiently rapidly to behave as a single copy gene. Intron sequence differences between isolates in this species reached nucleotide substitution saturation, while ITS sequences did not. Parsimony and evolutionary distance analysis of the two types of DNA data gave essentially the same tree conformation. By all these criteria, the group of G. pectorale isolates fell into two main clades, A and B. Clade A, with isolates from four continents, was comprised of four subclades of quite closely related isolates, plus one strain of ambiguous affinity. Clade B was comprised of two subclades represented by South African and South American isolates, respectively; thus, only subclades of clade B showed geographical localization. With respect to mating, all isolates except one homothallic strain and one apparently sterile strain fell into either one or the other of two mating types. Pairings in all possible combinations revealed that isolates from the same site formed abundant zygotes, which germinated to produce new, sexually active organisms. Zygotes were also formed in many pairings of other combinations, including crosses of clade A with clade B organisms, but none of the latter produced viable germlings. The ability to mate and produce viable progeny that were themselves capable of sexual reproduction was restricted to members of subclades established on the basis of DNA sequence similarities. Thus, the grades of difference in both nuclear intron sequences and rDNA ITS sequences paralleled those observed in the sexual analysis.

Actins↗

Nucleotide sequence of melon yellow spot virus M RNA segment and characterization of non-viral sequences in subgenomic RNA.

The nucleotide sequence of melon yellow spot virus (MYSV) M RNA segment was determined. The M RNA segment contains one open reading frame (ORF) encoding 308 amino acids (aa) in the sense orientation and another ORF encoding 1,127 aa in the complementary orientation, which were homologous to the NSm protein and G1/G2 glycoprotein precursor (Gp) protein, respectively. Amino acid sequences identities with the other tospovirus suggested that MYSV is closely related to groundnut bud necrosis virus and watermelon silver mottle virus. To analyze subgenomic RNA of the M RNA segment, RNA transcripts corresponding to the NSm and Gp genes were specifically amplified, and the nucleotide sequence of the 5' terminal region was determined. Sequence analysis of the NSm and Gp transcripts showed that they had a non-viral sequence 12-18 and 10-18 nucleotides long, respectively. Although these sequences varied considerably, in more than half of the cases, a cytosine residue was observed at the 3' end of the non-viral leader sequence, which suggests that the viral transcriptase prefers certain cap-donor sequences harboring a 3'CA dinucleotide.

Base Sequence↗

Structure and function of adenovirus DNA binding protein: comparison of the amino acid sequences of the Ad5 and Ad12 proteins derived from the nucleotide sequence of the corresponding genes.

The adenoviral DNA binding protein (DBP) is a multifunctional protein involved in DNA replication and gene expression. In order to investigate the relation between structure and function of DBP, the amino acid sequences of the serotypes 5 and 12 (Ad5 and Ad12) have been compared. The amino acid sequence of Ad5 DBP was previously established by nucleotide sequence analysis of the Ad5 DBP gene (W. Kruijer, F. M. A. Van Schaik, and J. S. Sussenbach, Nucl. Acids Res. 9, 4439-4457, 1981). In this study the analysis of the Ad5 DBP gene and adjacent regions by determination of the sequence of the first leader in late DBP mRNA's and the splice point between the tripartite leader and the main body of the mRNA encoding the 100-kDa protein has been extended. The nucleotide sequence of the Ad12 DBP gene is also described. From the nucleotide sequence and RNA mapping data of Ad12 DBP mRNA's (I. Saito, J. Sato H. Handa, K. Shiraki, and H. Shimojo, Virology 114, 379-398, 1981) the complete Ad12 DBP amino acid sequence could be deduced. Ad12 DBP contains 484 amino acids and has an actual Mr of 54,992. It is 45 amino acids shorter than Ad5 DBP. Comparison of the Ad12 and Ad5 DBP amino acid sequences shows that several longer deletions are present in the N-terminal 125 amino acid residues of Ad12 DBP. In contrast, only a single amino acid deletion and insertion is found in the C-terminal 359 amino acids of Ad12 DBP. The N- and C-terminal domains of Ad12 and Ad5 DBP are 45 and 80% homologous, respectively. This suggests that both domains of DBP are subjected to different evolutionary pressures. Analysis of various Ad5 mutants with an altered DBP gene, has indicated that the C-terminal domain is involved in DNA replication and early gene expression, while the N-terminal domain has a role in late gene expression in monkey cells. These results are discussed in relation to the structure and function of adenovirus DBP.

Adenoviruses, Human↗

Nucleotide sequence of the marmoset herpesvirus thymidine kinase gene and predicted amino acid sequence of thymidine kinase polypeptide.

The nucleotide sequence of a 2549-bp DNA fragment containing the entire coding region of the marmoset herpesvirus (MarHV) thymidine kinase gene (tk) and the flanking sequences was determined by the dideoxynucleotide chain termination method. The MarHV thymidine kinase polypeptide predicted from the nucleotide sequence contained 376 amino acids and had a molecular weight of 41,281. The sequencing data also reveal that the coding portion of another MarHV gene probably begins only 292 nucleotides downstream from the stop codon of the MarHV tk gene. There was relatively little nucleotide sequence homology between the MarHV tk gene and that of the herpes simplex virus (HSV) types 1 and 2 tk genes. Comparisons of the predicted amino acid sequences of the MarHV thymidine kinase polypeptide with that of the HSV-1 and HSV-2 thymidine kinase polypeptides, however, revealed clear, but interrupted, homology within several regions of the polypeptide chains. Amino acid sequence homology was particularly striking at residues 10 to 27 of the MarHV thymidine kinase polypeptide and residues 49 to 66 of the HSV-1 and HSV-2 thymidine kinase polypeptides. These same amino acid residues exhibit noticeable sequence homology to the mitochondrial beta subunit ATPase, oncogene p21 protein, adenylate kinase, and to other nucleotide-binding proteins. It has been proposed that the indicated regions of homology are elements of a nucleotide-binding pocket in ATPase, p21, and adenylate kinase, raising the possibility that amino acid residues 15 to 25 of the MarHV thymidine kinase and 54 to 64 of the HSV-1 and HSV-2 enzymes are likewise parts of nucleotide-binding sites.

Amino Acid Sequence↗

Complete sequence of constant and 3' noncoding regions of an immunoglobulin mRNA using the dideoxynucleotide method of RNA sequencing.

Three synthetic oligonucleotides were prepared to be complementary to known regions of the mouse immunoglublin light chain mRNA, and their ability to prime the transcription of complementary DNA (cDNA) was studied. The sequence of the cDNA was determined by adapting for mRNA the DNA sequencing method of Sanger, Nicklen and Coulson (1977) which uses 2'3' dideoxy ribonucleotides. A continuous sequence of 532 nucleotides was obtained, 321 corresponding to the whole of the constant region of the mRNA and the remaining 211 being the complete 3' noncoding region of the mRNA. The termination codon U-A-G occurs at the expected position in the mRNA corresponding to the triplet following the C terminal cystine. The nucleotide sequence is partially corroborated by the sequence of fragments obtained previously from 32P-mRNA fingerprints and endonuclease IV digests of 32P-cDNA, and is in agreement with the amino acid sequence of the constant region, except for a rearrangement of four amino acids (between amino acid positions 163 and 166). A revision of the amino acid sequence confirms the nucleic acid sequence.

Amino Acid Sequence↗