Search PubMedSearch

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Codon bias evolution in Drosophila. Population genetics of mutation-selection drift.

Although non-random patterns of synonymous codon usage are a prominent feature in the genomes of many organisms, the relatives roles of mutational biases and natural selection in maintaining codon bias remain a contentious issue. In some species, patterns of codon bias and empirical findings on the biology of translation suggest 'major codon preference', a balance among mutation pressure, genetic drift, and weak selection in favor of translationally superior codons. Population genetics theory makes testable predictions to distinguish such a model from a strictly mutational model of codon bias. Major codon preference predicts two fitness classes of synonymous DNA changes: 'preferred' mutations from non-major to major codons and 'unpreferred' changes in the opposite direction. An extension of current statistical methods is employed to reveal differences in the within and between species dynamics of preferred and unpreferred silent mutations in Drosophila simulans. In this lineage, codon bias appears to be maintained under roughly equal magnitudes of natural selection and genetic drift. In the sibling species, D. melanogaster, however, a reduction in N(e)s, the product of effective population size and selection coefficient, appears to have allowed a genome-wide reduction in codon bias.

Animals

Codon usage in the Mycobacterium tuberculosis complex.

The usage of alternative synonymous codons in Mycobacterium tuberculosis (and M. bovis) genes has been investigated. This species is a member of the high-G+C Gram-positive bacteria, with a genomic G+C content around 65 mol%. This G+C-richness is reflected in a strong bias towards C- and G-ending codons for every amino acid: overall, the G+C content at the third positions of codons is 83%. However, there is significant variation in codon usage patterns among genes, which appears to be associated with gene expression level. From the variation among genes, putative optimal codons were identified for 15 amino acids. The degree of bias towards optimal codons in an M. tuberculosis gene is correlated with that in homologues from Escherichia coli and Bacillus subtilis. The set of selectively favoured codons seems to be quite highly conserved between M. tuberculosis and another high-G+C Gram-positive bacterium, Corynebacterium glutamicum, even though the genome and overall codon usage of the latter are much less G+C-rich.

Bacillus subtilis

Codon usage in Caenorhabditis elegans: delineation of translational selection and mutational biases.

Synonymous codon usage varies considerably among Caenorhabditis elegans genes. Multivariate statistical analyses reveal a single major trend among genes. At one end of the trend lie genes with relatively unbiased codon usage. These genes appear to be lowly expressed, and their patterns of codon usage are consistent with mutational biases influenced by the neighbouring nucleotide. At the other extreme lie genes with extremely biased codon usage. These genes appear to be highly expressed, and their codon usage seems to have been shaped by selection favouring a limited number of translationally optimal codons. Thus, the frequency of these optimal codons in a gene appears to be correlated with the level of gene expression, and may be a useful indicator in the case of genes (or open reading frames) whose expression levels (or even function) are unknown. A second, relatively minor trend among genes is correlated with the frequency of G at synonymously variable sites. It is not yet clear whether this trend reflects variation in base composition (or mutational biases) among regions of the C.elegans genome, or some other factor. Sequence divergence between C.elegans and C.briggsae has also been studied.

Animals

Hydropathic characteristics of adenovirus hexons.

The complete nucleotide sequence and the predicted amino acid sequence of the adenovirus type 7 hexon gene were determined. The hydropathy of the hexon proteins from human adenovirus types 2, 3, 4, 5, 7, 12, 16, 40, 41, and 48, bovine adenovirus type 3, murine adenovirus type 1, and avian adenovirus types 1 and 10 was analysed. The presence of purines and pyrimidines in the second position of the codons was correlated to hydrophilicity and hydrophobicity, respectively. Comparison of the hydrophilicity plots of eight hexons showed seven hypervariable regions to be distributed on the surface. A large portion of the hypervariable regions manifests hydrophilicity. The strength of the surface charge accumulated on the hydrophilic and hydrophobic regions correlated to the tissue tropism of the different adenovirus types. Analysis of codon usage for adenovirus hexons showed that among synonymous codons those with cytidine in the third position were preferably used to a great extent. Analysis of the nucleotide and amino acid sequence pair distances and the phylogenetic tree of 14 hexon proteins showed members of subgenera B, D and E to be closely related, especially Ad4 and Ad16, and subgenus A to be closely related to subgenus F.

Adenoviruses, Human

Codon usage in Entamoeba histolytica.

The codon usage of 10 E. histolytica genes comprising 4455 codons was analysed. The codon usage revealed an extremely biased use of synonymous codons with a preference for NNU (44%) and NNA (41.4%) codons. Codons CGG (arg), AGG (arg) and CCG (pro) were absent in the E. histolytica genes examined. The codon usage of E. histolytica resembled that of Plasmodium falciparum.

Animals

[Analysis of the primary structure of mRNA from Escherichia coli: occurrence of nucleotides on the 3'-side of the codon].

The occurrence of nucleotides of the 3' side of codons has been determined in highly and weakly expressed genes from Escherichia coli. It was found that the usage of some amino acid codons in highly expressed genes was site specific, depending on the base 3' to the codon. The role of the 3' nucleotide as a modulator of codon translation effectiveness is discussed. The rules of synonymous codon usage in relation to the 3' flanking nucleotide have been established for highly expressed genes. For example, if a triplet next to the lysine codon starts with guanosine, lysine is preferably encoded by AAA and not by AAG (P less than 10(-8), while of cytidine is 3' to the lysine codon, AAG is preferred over AAA (P less than 0.001). These rules are observed in highly and absent in weakly expressed mRNAs and can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence

DNA Translator and Aligner: HyperCard utilities to aid phylogenetic analysis of molecules.

DNA Translator and Aligner are molecular phylogenetics HyperCard stacks for Macintosh computers. They manipulate sequence data to provide graphical gene mapping, conversions, translations and manual multiple-sequence alignment editing. DNA Translator is able to convert documented GenBank or EMBL documented sequences into linearized, rescalable gene maps whose gene sequences are extractable by clicking on the corresponding map button or by selection from a scrolling list. Provided gene maps, complete with extractable sequences, consist of nine metazoan, one yeast, and one ciliate mitochondrial DNAs and three green plant chloroplast DNAs. Single or multiple sequences can be manipulated to aid in phylogenetic analysis. Sequences can be translated between nucleic acids and proteins in either direction with flexible support of alternate genetic codes and ambiguous nucleotide symbols. Multiple aligned sequence output from diverse sources can be converted to Nexus, Hennig86 or PHYLIP format for subsequent phylogenetic analysis. Input or output alignments can be examined with Aligner, a convenient accessory stack included in the DNA Translator package. Aligner is an editor for the manual alignment of up to 100 sequences that toggles between display of matched characters and normal unmatched sequences. DNA Translator also generates graphic displays of amino acid coding and codon usage frequency relative to all other, or only synonymous, codons for approximately 70 select organism-organelle combinations. Codon usage data is compatible with spreadsheet or UWGCG formats for incorporation of additional molecules of interest. The complete package is available via anonymous ftp and is free for non-commercial uses.

Amino Acid Sequence

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing

Codon bias and mutability in HIV sequences.

A survey of the patterns of synonymous codon preference in the HIV env gene reveals a correlation between the codon bias and the mutability requirements of different regions of the protein. At hypervariable regions in gp120 one finds a greater proportion of codons that tend to mutate nonsynonymously, but to a target that is similar in hydrophobicity and volume. We argue that this strategy results from a compromise between the selective pressure placed on the virus by the induced immune response, which favors amino acid substitutions in the complementarity determining regions, and the negative selection against missense mutations that violate structural constraints of the env protein.

Codon

Ribosome-mediated translational pause and protein domain organization.

Because regions on the messenger ribonucleic acid differ in the rate at which they are translated by the ribosome and because proteins can fold cotranslationally on the ribosome, a question arises as to whether the kinetics of translation influence the folding events in the growing nascent polypeptide chain. Translationally slow regions were identified on mRNAs for a set of 37 multidomain proteins from Escherichia coli with known three-dimensional structures. The frequencies of individual codons in mRNAs of highly expressed genes from E. coli were taken as a measure of codon translation speed. Analysis of codon usage in slow regions showed a consistency with the experimentally determined translation rates of codons; abundant codons that are translated with faster speeds compared with their synonymous codons were found to be avoided; rare codons that are translated at an unexpectedly higher rate were also found to be avoided in slow regions. The statistical significance of the occurrence of such slow regions on mRNA spans corresponding to the oligopeptide domain termini and linking regions on the encoded proteins was assessed. The amino acid type and the solvent accessibility of the residues coded by such slow regions were also examined. The results indicated that protein domain boundaries that mark higher-order structural organization are largely coded by translationally slow regions on the RNA and are composed of such amino acids that are stickier to the ribosome channel through which the synthesized polypeptide chain emerges into the cytoplasm. The translationally slow nucleotide regions on mRNA possess the potential to form hairpin secondary structures and such structures could further slow the movement of ribosome. The results point to an intriguing correlation between protein synthesis machinery and in vivo protein folding. Examination of available mutagenic data indicated that the effects of some of the reported mutations were consistent with our hypothesis.

Bacterial Proteins

The evolution of codon preferences in Drosophila: a maximum-likelihood approach to parameter estimation and hypothesis testing.

Synonymous codon usage in related species may differ as a result of variation in mutation biases, differences in the overall strength and efficiency of selection, and shifts in codon preference-the selective hierarchy of codons within and between amino acids. We have developed a maximum-likelihood method to employ explicit population genetic models to analyze the evolution of parameters determining codon usage. The method is applied to twofold degenerate amino acids in 50 orthologous genes from D. melanogaster and D. virilis. We find that D. virilis has significantly reduced selection on codon usage for all amino acids, but the data are incompatible with a simple model in which there is a single difference in the long-term Ne, or overall strength of selection, between the two species, indicating shifts in codon preference. The strength of selection acting on codon usage in D. melanogaster is estimated to be |Nes| approximately 0.4 for most CT-ending twofold degenerate amino acids, but 1.7 times greater for cysteine and 1.4 times greater for AG-ending codons. In D. virilis, the strength of selection acting on codon usage for most amino acids is only half that acting in D. melanogaster but is considerably greater than half for cysteine, perhaps indicating the dual selection pressures of translational efficiency and accuracy. Selection coefficients in orthologues are highly correlated (rho = 0.46), but a number of genes deviate significantly from this relationship.

Amino Acids

Gene length and codon usage bias in Drosophila melanogaster, Saccharomyces cerevisiae and Escherichia coli.

The relationship between gene length and synonymous codon usage bias was investigated in Drosophila melanogaster, Escherichia coli and Saccharomyces cerevisiae. Simulation studies indicate that the correlations observed in the three organisms are unlikely to be due to sampling errors or any potential bias in the methods used to measure codon usage bias. The correlation was significantly positive in E.coli genes, whereas negative correlations were obtained for D. melanogaster and S.cerevisiae genes. When only ribosomal protein genes were used, whose expression levels are assumed to be similar, E.coli and S.cerevisiae showed significantly positive correlations. For the two eukaryotes, the distribution of effective number of codons was different in short genes (300-500 bp) compared with longer genes; this was not observed in E.coli. Both positive and negative correlations can be explained by translational selection. Energetically costly longer genes have higher codon usage bias to maximize translational efficiency. Selection may also be acting to reduce the size of highly expressed proteins, and the effect is particularly pronounced in eukaryotes. The different relationships between codon usage bias and gene length observed in prokaryotes and eukaryotes may be the consequence of these different types of selection.

Animals

Adjustment of the tRNA population to the codon usage in chloroplasts.

In chloroplasts there is a correlation between the amounts of tRNAs specific for a given amino acid and the codons specifying this amino acid. Furthermore, for the amino acids coded for by more than one codon, the population of isoaccepting tRNAs is adjusted to the frequency of synonymous codons used in chloroplast protein genes. A comparison by two-dimensional gel electrophoresis of the tRNA populations extracted from chloroplasts and from chloroplast polysomes shows that all chloroplast tRNAs are involved in protein biosynthesis.

Chloroplasts

Nucleotide sequence of the Escherichia coli recJ chromosomal region and construction of recJ-overexpression plasmids.

The nucleotide sequence of the recJ gene of Escherichia coli K-12 and two upstream coding regions was determined. Three regions were identified within these two upstream genes that exhibited weak to moderate promoter activity in fusions to the galK gene and are candidates for the recJ promoter. recJ appeared to be poorly translated: the recJ nucleotide sequence revealed a suboptimal initiation codon GUG, no discernible ribosome-binding consensus sequence, and relatively nonbiased synonymous codon usage. Comparison of the sequence of this region of the chromosome with DNA data bases identified the gene immediately downstream of recJ as prfB, which encodes translational release factor 2 and has been mapped near recJ at 62 min. No significant homology between recJ and other previously sequenced regions of DNA was detected. However, protein sequence comparisons with a gene upstream of recJ, denoted xprB, revealed significant homology with several site-specific recombination proteins. Its genetic function is presently unknown. Knowledge of the nucleotide sequence of recJ allowed the construction of a plasmid from which overexpression of RecJ protein could be induced. Supporting the notion that translation of recJ is limiting, a strong T7 bacteriophage promoter upstream of recJ did not, by itself, allow high-level expression of RecJ protein. The addition of a ribosome-binding sequence fused to the initiator GTG of recJ in this construction was necessary to promote expression of high levels of RecJ protein.

Amino Acid Sequence

Structure of the Escherichia coli K12 regulatory gene tyrR. Nucleotide sequence and sites of initiation of transcription and translation.

The nucleotide sequence of 1964 base pairs of the Escherichia coli K12 chromosome containing the autogenously regulated regulatory gene tyrR has been determined. The site of initiation of transcription of tyrR has been mapped by primer-extension analysis, and the initiation codon has been identified by site-specific deletion mutagenesis. The nucleotide sequence predicts a subunit molecular weight of 53,099 for the TyrR protein. Codon usage in the tyrR structural gene shows a bias toward those synonymic codons which are used rarely in efficiently expressed E. coli genes. The nucleotide sequence of a 22-base pair region adjacent to the promoter and distal to the structural gene exhibits considerable identity with corresponding regions of other genes regulated by tyrR. It is proposed that this is a site for repression by the TyrR protein.

Amino Acid Sequence

DNA sequence variability at the rplX locus of Bacillus subtilis.

The pattern and extent of DNA sequence variability at the rplX locus (encoding ribosomal protein L24) has been investigated in nine strains of Bacillus subtilis. Overall, there is a very low level of nucleotide diversity, even at silent sites, which is probably due to selection among synonymous codons. By analogy with Escherichia coli, there may also be some effect of the relative proximity of rplX to the chromosomal origin of replication. The small number of nucleotide substitutions are non-randomly distributed: all of the synonymous changes are in valine codons. From the sequence differences the strains can be divided into two groups, which are not coincident with their previous classification; this observation is consistent with recombination among strains.

Amino Acid Sequence

Codon equilibrium II: Its use in estimating silent-substitution rates.

We study the equilibrium in the use of synonymous codons by eukaryotic organisms and find five equations involving substitution rates that we believe embody the important implications of equilibrium for the process of silent substitution. We then combine these five equations with additional criteria to determine sets of substitution rates applicable to eukaryotic organisms. One method employs the equilibrium equations and a principle of maximum entropy to find the most uniform set of rates consistent with equilibrium. In a second method we combine the equilibrium equations with data on the man-mouse divergence to determine that set of rates that is most neutral yet consistent with both types of data (i.e., equilibrium and divergence data). Simulations show this second method to be quite reliable in spite of significant saturation in the substitution process. We find that when divergence data are included in the calculation of rates, even though these rates are chosen to be as neutral as possible, the strength of selection inferred from the nonuniformity of the rates is approximately doubled. Both sets of rates are applied to estimate the human-mouse divergence time based on several independent subsets of the divergence data consisting of the quartet, C- or T-ending duet, and A- or G-ending duet codon sets. Both rate sets produce patterns of divergence times that are shortest for the quartet data, intermediate for the CT-ending duets, and longest for the AG-ending duets. This indicates that rates of transitions in the duet-codon sets are significantly higher than those in the quartet-codon sets; this effect is especially marked for A----G, the rate of which in duets must be about double that in quartets.

Animals

Rare codons are not sufficient to destabilize a reporter gene transcript in tobacco.

In plants, as in other eukaryotes, most synonymous codons of the genetic-code are not used with equal frequency, but instead some codons are preferred, whereas others are rare. Circumstantial evidence led to the suggestion that rare codons have a negative influence on mRNA stability. To address this question experimentally, rare codons encoded by a Bacillus thuringiensis (B.t.) toxin gene (cryIA(c)) or a synthetic sequence were introduced into a phytohemagglutinin (PHA) reporter gene. In neither case was the mRNA stability appreciably diminished in stably transformed tobacco cell cultures nor was the accumulation of mRNA in transgenic plants affected. Thus rare codons do not appear to be sufficient to cause rapid degradation of the PHA mRNA and potentially other mRNAs in plants.

Cell Line