Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

Spiroplasma virus 4: nucleotide sequence of the viral DNA, regulatory signals, and proposed genome organization.

The replicative form (RF) of spiroplasma virus 4 (SpV4) has been cloned in Escherichia coli, and the cloned RF has been shown to be infectious by transfection (M. C. Pascarel-Devilder, J. Renaudin, and J.-M. Bové, Virology 151:390-393, 1986). The cloned SpV4 RF was randomly subcloned and was fully sequenced by the dideoxy chain termination technique, using the M13 cloning and sequencing system. The nucleotide sequence of the SpV4 genome contains 4,421 nucleotides with a G+C content of 32 mol%. The triplet TGA is not a termination codon but, as in Mycoplasma capricolum (F. Yamao, A. Muto, Y. Kawauchi, M. Iwami, S. Iwagani, Y. Azumi, and S. Osawa, Proc. Natl. Acad. Sci. USA 82:2306-2309, 1985), probably codes for tryptophan. With these assumptions, nine open reading frames (ORFs) were identified. All nine are characterized by an ATG or GTG initiation codon, one or several termination codons, and a Shine-Dalgarno sequence upstream of the initiation codon. The nine ORFs are distributed in all three reading frames. One of the ORFs (ORF1) corresponds to the 60,000-dalton capsid protein gene. Analysis of codon usage showed that T- and A-terminated codons are preferably used, reflecting the low G+C content (32 mol%) of the SpV4 genome. The viral DNA contains two G+C-rich inverted repeat sequences. One could be involved in transcription termination and the other in initiation of cDNA strand synthesis. The SpV4 genome was found to contain at least three promoterlike sequences quasi-identical to those of eubacteria. These results fully support the bacterial origin of spiroplasmas.

Amino Acid Sequence↗

Cloning of two glutamate dehydrogenase cDNAs from Asparagus officinalis: sequence analysis and evolutionary implications.

Two different amplification products, termed c1 and c2, showing a high similarity to glutamate dehydrogenase sequences from plants, were obtained from Asparagus officinalis using two degenerated primers and RT-PCR (reverse transcriptase polymerase chain reaction). The genes corresponding to these cDNA clones were designated aspGDHA and aspGDHB. Screening of a cDNA library resulted in the isolation of cDNA clones for aspGDHB only. Analysis of the deduced amino acid (aa) sequence from the full-length cDNA suggests that the gene product contains all regions associated with metabolic function of NAD glutamate dehydrogenase (NAD-GDH). A first phylogenetic analysis including only GDHs from plants suggested that the two GDH genes of A. officinalis arose by an ancient duplication event, pre-dating the divergence of monocots and dicots. Codon usage analysis showed a bias towards A/T ending codons. This tendency is likely due to the biased nucleotide composition of the asparagus genome, rather than to the translational selection for specific codons. Using principal coordinate analysis, the evolutionary relatedness of plant GDHs with homologous sequences from a large spectrum of organisms was investigated. The results showed a closer affinity of plant GDHs to GDHs of thermophilic archaebacterial and eubacterial species, when compared to those of unicellular eukaryotic fungi. Sequence analysis at specific amino acid signatures, known to affect the thermal stability of GDH, and assays of enzyme activity at non-physiological temperatures, showed a greater adaptation to heat-stress conditions for the asparagus and tobacco enzymes compared with the Saccharomyces cerevisiae enzyme.

Amino Acid Sequence↗

Spiroplasmas: gene structure and expression.

Upon sequencing of the SpV4 genome, eight putative open reading frames (ORFs) including that for the 65-kilodalton (kDa) capsid protein were detected. They involve all three reading frames. Three promoter sequences were found, as well as a transcription terminator and the initiation site for complementary strand synthesis. Ribosome binding sites and regulatory sequences are closely related to those of Eubacteria. Codon usage analysis showed that A and T terminated codons are preferably used. UAA is the major termination codon. Upon cloning of the full-size SpV4 replicative form, the capsid protein gene could not be expressed in Escherichia coli, whereas the spiralin gene cloned in the same bacterium is expressed. These results suggest that in spiroplasmas, as in Mycoplasma capricolum, UGA is not a termination codon, but very probably codes for tryptophan. Spiralin contains no tryptophan. Hence, its gene contains no UGA codons and can thus be expressed in E. coli. On the other hand, the gene for capsid protein has nine UGA codons and cannot be fully expressed in the bacterium. Our results fully support the bacterial origin of spiroplasmas.

Bacteria↗

The araBAD operon of Salmonella typhimurium LT2. I. Nucleotide sequence of araB and primary structure of its product, ribulokinase.

Hybrid plasmids containing the araBAD operon of Salmonella typhimurium LT2 were characterized by Southern blot and genetic analyses. The nucleotide sequence of araB was determined. The araB gene product, ribulokinase (EC 2.7.1.16), was purified and the results of amino acid composition analysis and partial amino acid sequence are in agreement with predictions from the DNA sequence. Ribulokinase is 569 amino acid residues long and has a calculated Mr of 61 793. Ribulokinase shares significant homology with xylulose kinase from Escherichia coli. Codon usage in the araB gene does not favor those codons which have intermediate codon-anticodon binding energy.

Amino Acid Sequence↗

Analysis of nucleotide sequences of two ligninase cDNAs from a white-rot filamentous fungus, Phanerochaete chrysosporium.

An analysis of nucleotide sequences of two types of ligninase cDNAs isolated from the basidiomycete Phanerochaete chrysosporium, designated CLG4 and CLG5, are presented here. The amino acid sequences of the corresponding ligninase proteins, designated LG4 and LG5, respectively, have been deduced from the cDNA sequences. Mature ligninases LG4 and LG5 are preceded by leader sequences containing 28 and 27 amino acids (aa), respectively, and each contains 344 aa residues. The estimated Mrs of mature LG4 and LG5 are 36,540 and 36,607, respectively. Potential N-glycosylation site(s) with the general sequence Asn-X-Thr/Ser are found in both LG4 and LG5. Nucleotide sequence homology between the coding region of CLG4 and CLG5 is 71.5%, whereas the amino acid sequence homology between the two ligninases is 68.5%. The codon usage of ligninases is extremely biased in favor of codons rich in cytosine and guanine. Amino acid sequences of two tryptic peptides of ligninase H8 have exactly matching sequences in ligninase LG5. Also, the sequences of the oligodeoxynucleotide probes, which correspond to the sequences in the tryptic peptides of ligninase H8 and which were used in isolating the ligninase clones from the cDNA library, have exactly matching sequences in CLG5. The experimentally determined N-terminal sequence of purified ligninase H8 is found in the deduced N-terminal amino acid sequence of LG5. These results suggest that CLG5 encodes ligninase H8 and that CLG4 represents a related but different ligninase gene.

Amino Acid Sequence↗

Cloning and molecular characterization of the acetamidase-encoding gene (amdS) from Aspergillus oryzae.

We have isolated an acetamidase-encoding gene (amdS) from Aspergillus oryzae by heterologous hybridization using the corresponding Aspergillus nidulans gene as a probe. The gene is located on a 3.5-kb SacI fragment and its nucleotide (nt) sequence was determined. Compared with the A. nidulans amdS gene, the coding region of A. oryzae gene consists of seven exons interrupted by six introns and encodes 545 amino acid (aa) residues. The deduced aa sequence has a high degree of homology with that of the A. nidulans acetamidase protein. Three introns (IVS-1, IVS-2, and IVS-4) exist at the same positions as those of A. nidulans amdS, whilst three additional introns (IVS-3, IVS-5, and IVS-6) are also present. There is no preference in its codon usage (G + C content in the third position of codons is 51%). Gene disruption experiments demonstrate that the resulting mutants show significantly reduced growth on acetamide-containing medium, indicating that the A. oryzae amdS gene encodes a functional acetamidase that is required for acetamide utilization. Transcriptional analysis by Northern blot reveals a 1.8-kb transcript in RNA extracted from mycelium grown in medium containing acetamide or acetate plus beta-alanine as the sole carbon and nitrogen sources.

Amidohydrolases↗

Gene synthesis, expression in Escherichia coli, purification and characterization of the recombinant bovine acyl-CoA-binding protein.

A synthetic gene encoding the 86 amino acid residues of mature acyl-CoA-binding protein (ACBP), and the initiating methionine was constructed. The synthetic gene was assembled from eight partially overlapping oligonucleotides. Codon usage and nucleotides surrounding the ATG translation-initiation codon were chosen to allow efficient expression in Escherichia coli as well as in yeast. The synthetic gene was inserted into the expression vector pKK223-3 and expressed in E. coli. In maximally induced cultures, recombinant ACBP constitutes 12-15% of total cellular protein. A fraction highly enriched for recombinant ACBP was obtained by extracting induced E. coli cells with 1 M-acetic acid. Recombinant ACBP was purified to homogeneity by successive use of gel-filtration chromatography, ion-exchange chromatography and reverse-phase h.p.l.c. Recombinant ACBP differed from native ACBP by lacking the N-terminal acetyl group. The acyl-CoA-binding characteristics of recombinant ACBP did not differ from those of native ACBP, and the two proteins showed the same ability to induce medium-chain acyl-CoA synthesis by goat mammary-gland fatty acid synthetase. It was concluded that the N-terminal acetyl group is not important for acyl-CoA binding.

Acyl Coenzyme A↗

Syk mutation in Jurkat E6-derived clones results in lack of p72syk expression.

The human leukemic Jurkat cell line is commonly used as a model cellular system to study T lymphocyte signal transduction. Various clonal derivatives of Jurkat T cells exist which display different characteristics with regard to responses to external stimuli. Among these, the E6-1 clone of Jurkat T cells has been used as a parental line from which numerous important somatic mutant clones have been generated. During the course of experiments examining signals initiated by the T cell antigen receptor in an E6-1-derived Jurkat cell clone J.CaM1, we observed that the 72-kilodalton Syk protein tyrosine kinase previously found in other Jurkat cells was not detected. Upon further analysis it was determined that Syk transcripts from the J.CaM1 cells as well as the parental E6-1 cells contain a single guanine nucleotide insertion at position 92. This nucleotide insertion results in a shift in the Syk open reading frame leading to alternate codon usage as well as the generation of a termination codon at position 109. Thus, Syk transcripts in E6-1 cells and E6-1-derived clones are predicted to be capable of encoding only the first 33 amino acids of the 630-amino acid wild type Syk. These findings are incompatible with a recently proposed model of T cell antigen receptor signal transduction based, in part, on experiments conducted using E6-1-derived cells, suggesting that Syk might play a role upstream of Lck and Zap70.

Amino Acid Sequence↗

Primary structure of the ompF gene that codes for a major outer membrane protein of Escherichia coli K-12.

The nucleotide sequence of the ompF gene coding for a major outer membrane protein of Escherichia coli K-12 has been determined and the amino acid sequence of the OmpF protein was deduced from it. The OmpF protein contains 340 amino acid residues, and is produced from a precursor having 22 extra amino acid residues, the signal peptide, at the amino terminus. The expected secondary structure of the OmpF protein had a high beta-sheet content with a low alpha-helix content. The promoter region and the transcription termination region of the ompF gene had a significantly high AT content, while the AT content of the coding region was about the same as the average AT content of the E. coli chromosome. Following the termination codon, a typical rho-independent transcription termination signal was observed. The codon usage in the ompF gene was highly nonrandom; the codons preferably utilized are those recognized by the most abundant species of isoaccepting tRNAs or those, among synonymous codons recognized by the same tRNA, that can interact more properly with the anticodon.

Amino Acid Sequence↗

Sequences of the E. coli uvrC gene and protein.

We have determined the sequence of a 2400 bp region of E. coli chromosomal DNA containing the uvrC gene. The coding region of uvrc is 2267 bp in length, encodes a polypeptide with a calculated molecular weight of 66,038 daltons, and is preceded by a typical E. coli ribosome binding site. By constructing deletion derivatives we have established that a uvrC promoter lies within the 113 bp region preceding the translational start of uvrC. The codon usage in uvrC is strongly biased in favor of codons used infrequently in E. coli, which may contribute to the relatively low intracellular concentration of uvrC protein.

Amino Acid Sequence↗

Nucleotide sequence of the LuxC gene and the upstream DNA from the bioluminescent system of Vibrio harveyi.

The nucleotide sequence of the luxC gene (1431 bp) and the upstream DNA (1049 bp) of the luminescent bacterium Vibrio harveyi has been determined. The luxC gene can be translated into a polypeptide of 55 kDa in excellent agreement with the molecular mass of the reductase polypeptide required for synthesis of the aldehyde substrate for the bioluminescent reaction. Analyses of codon usage showed a high frequency (1.9%) of the isoleucine codon, AUA, in the luxC gene compared to that found in Escherichia coli genes (0.2%) and its absence in the luxA, B and D genes. The low G/C content of the luxC gene and upstream DNA (38-39%) compared to that found in the other lux genes of V. harveyi (45%) was primarily due to a stretch of 500 nucleotides with only a 24% G/C content, extending from 200 bp inside lux C to 300 bp upstream. Moreover, an open reading frame did not extend for more than 48 codons between the luxC gene and 600 bp upstream at which point a gene transcribed in the opposite direction started. As the lux system in the luminescent bacterium, V. fischeri, contains a regulatory gene immediately upstream of luxC transcribed in the same direction, these results show that the organization and regulation of the lux genes have diverged in different luminescent bacteria.

Amino Acid Sequence↗

TransTerm, the translational signal database, extended to include full coding sequences and untranslated regions.

TransTerm is a database of mRNA sequences and parameters useful for detecting translational control signals in general. TransTerm-98 has been expanded beyond previous years to include full coding sequences and UTRs, while retaining the original small contexts about the coding sequence start- and stop-codons. The database contains more than 130 000 non-redundant coding sequences with associated untranslated regions (UTRs) from over 450 species. This includes the complete genomes of 12 prokaryotic and one eukaryotic organism. Several coding sequence parameters are available: coding sequence length, Nc, GC3 and, when it is computable, Codon Adaptation Index (CAI). Codon usage tables and summaries of start- and stop-codon contexts are also included. TransTerm-98 has both a relational database form with a WWW interface and a flatfile format, also available by Internet browser. TransTerm is available at: http://biochem.otago.ac.nz:800/Transterm/homepage.h tml

Codon↗

The complete mitochondrial sequence of Tarsius bancanus: evidence for an extensive nucleotide compositional plasticity of primate mitochondrial DNA.

Inconsistencies between phylogenetic interpretations obtained from independent sources of molecular data occasionally hamper the recovery of the true evolutionary history of certain taxa. One prominent example concerns the primate infraordinal relationships. Phylogenetic analyses based on nuclear DNA sequences traditionally represent Tarsius as a sister group to anthropoids. In contrast, mitochondrial DNA (mtDNA) data only marginally support this affiliation or even exclude Tarsius from primates. Two possible scenarios might cause this conflict: a period of adaptive molecular evolution or a shift in the nucleotide composition of higher primate mtDNAs through directional mutation pressure. To test these options, the entire mt genome of Tarsius bancanus was sequenced and compared with mtDNA of representatives of all major primate groups and mammals. Phylogenetic reconstructions at both the amino acid (AA) and DNA level of the protein-coding genes led to faulty tree topologies depending on the algorithms used for reconstruction. We propose that these artifactual affiliations rather reflect the nucleotide compositional similarity than phylogenetic relatedness and favor the directional mutation pressure hypothesis because: (1) the overall nucleotide composition changes dramatically on the lineage leading to higher primates at both silent and nonsilent sites, and (2) a highly significant correlation exists between codon usage and the nucleotide composition at the third, silent codon position. Comparisons of mt genes with mt pseudogenes that presumably transferred to the nucleus before the directional mutation pressure took place indicate that the ancestral DNA composition is retained in the relatively fossilized mtDNA-like sequences, and that the directed acceleration of the substitution rate in higher primates is restricted to mtDNA.

Animals↗

Nucleotide sequence and characteristics of the gene for L-lactate dehydrogenase of Thermus caldophilus GK24 and the deduced amino-acid sequence of the enzyme.

The gene for L-lactate dehydrogenase (LDH) (EC 1.1.1.27) of Thermus caldophilus GK24 was cloned in Escherichia coli using synthetic oligonucleotides as hybridization probes. The nucleotide sequence of the cloned DNA was determined. The primary structure of the LDH was deduced from the nucleotide sequence. The deduced amino acid sequence agreed with the NH2-terminal and COOH-terminal sequences previously reported and the determined amino acid sequences of the peptides obtained from trypsin-digested T. caldophilus LDH. The LDH comprised 310 amino acid residues and its molecular mass was determined to be 32,808. On alignment of the whole amino acid sequences, the T. caldophilus LDH showed about 40% identity with the Bacillus stearothermophilus, Lactobacillus casei and dogfish muscle LDHs. The T. caldophilus LDH gene was expressed with the E. coli lac promoter in E. coli, which resulted in the production of the thermophilic LDH. The gene for the T. caldophilus LDH showed more than 40% identity with those for the human and mouse muscle LDHs on alignment of the whole nucleotide sequences. The G + C content of the coding region for the T. caldophilus LDH was 74.1%, which was higher than that of the chromosomal DNA (67.2%). The G + C contents in the first, second and third positions of the codons used were 77.7%, 48.1% and 95.5% respectively. The high G + C content in the third base caused extremely non-random codon usage in the LDH gene. About half (48.7%) the codons in the LDH gene started with G, and hence there were relatively high contents of Val, Ala, Glu and Gly in the LDH. The contents of Pro, Arg, Ala and Gly, which have high G + C contents in their codons, were also high. Rare codons with U or A as the third base were sometimes used to avoid the TCGA sequence, the recognition site for the restriction endonuclease, TaqI. Two TCGA sequences were found only in the sequence of CTCGAG (XhoI site) in the sequenced region of the T. caldophilus DNA. There were three segments with similar sequences in the two 5' non-coding regions, probably the promoter and ribosome-binding regions, of the genes for the T. caldophilus LDH and the Thermus thermophilus 3-isopropylmalate dehydrogenase.

Amino Acid Sequence↗

Expression of the strA-strB streptomycin resistance genes in Pseudomonas syringae and Xanthomonas campestris and characterization of IS6100 in X. campestris.

Expression of the strA-strB streptomycin resistance (SMr) genes was examined in Pseudomonas syringae pv. syringae and Xanthomonas campestris pv. vesicatoria. The strA-strB genes in P. syringae and X. campestris were encoded on elements closely related to Tn5393 from Erwinia amylovora and designated Tn5393a and Tn5393b, respectively. The putative recombination site (res) and resolvase-repressor (tnpR) genes of Tn5393 from E. amylovora, P syringae, and X. campestris were identical; however, IS6100 mapped within tnpR in X. campestris, and IS1133 was previously located downstream of tnpR in E. amylovora (C.-S Chiou and A. L. Jones, J. Bacteriol. 175:732-740, 1993). Transcriptional fusions (strA-strB::uidA) indicated that a strong promoter sequence was located within res in Tn5393a. Expression from this promoter sequence was reduced when the tnpR gene was present in cis position relative to the promoter. In X. campestris pv. vesicatoria, analysis of promoter activity with transcriptional fusions indicated that IS6100 increased the expression of strA-strB. Analysis of codon usage patterns and percent G+C in the third codon position indicated that IS6100 could have originated in a gram-negative bacterium. The data obtained in the present study help explain differences observed in the levels of SMr expressed by three genera which share common genes for resistance. Furthermore, the widespread dissemination of Tn5393 and derivatives in phytopathogenic prokaryotes confirms the importance of these bacteria as reservoirs of antibiotic resistance in the environment.

Base Sequence↗

Characterization of the rfc region of Shigella flexneri.

The O antigen of the Shigella flexneri lipopolysaccharide (LPS) is an important virulence determinant and immunogen. We have isolated S. flexneri mutants which produce a semi-rough LPS by using an O-antigen-specific phage, Sf6c. Western immunoblotting was used to show that the LPS produced by the semi-rough mutants contained only one O-antigen repeat unit. Thus, the mutants are deficient in production of the O-antigen polymerase and were termed rfc mutants. Complementation experiments were used to locate the rfc adjacent to the rfb genes on plasmid clones previously isolated and containing this region (D. F. Macpherson, R. Morona, D. W. Beger, K.-C. Cheah, and P. A. Manning, Mol. Microbiol 5:1491-1499, 1991). A combination of deletions and subcloning analysis located the rfc gene as spanning a 2-kb region. Insertion of a kanamycin resistance cartridge into a SalI site in this region inactivated the rfc gene. The DNA sequence of the rfc region was determined. An open reading frame spanning the SalI site was identified and encodes a protein with a predicted molecular mass of 43.7 kDa. The predicted protein is highly hydrophobic and showed little sequence homology with any other protein. Comparison of its hydropathy plot with that of other Rfc proteins from Salmonella enterica (typhimurium) and Salmonella enterica (muenchen) revealed that the profiles were similar and that the proteins have 12 or more potential membrane-spanning segments. A comparison of the S. flexneri rfc gene and protein product with other rfc and rfc-like proteins revealed that they have a similarly low percentage of G + C content and have similar codon usage, and all have a high percentage of rare codons. An attempt to identify the S. flexneri Rfc protein was unsuccessful, although proteins encoded upstream and downstream of the rfc gene could be identified. Examination of the distribution of rare or minor codons in the rfc gene revealed that it has several minor codons within the first 25 amino acids. This is in contrast to the upstream gene rfbG, which also has a high percentage of rare codons but whose gene product could be detected. The positioning of the rare codons in the rfc gene may restrict translation and suggests that minor isoaccepting tRNA species may be involved in translational regulation of rfc expression. The low percentage of G + C content of rfc genes may be a consequence of the selection pressure to maintain this form of control.

Amino Acid Sequence↗

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗