Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Proteome composition and codon usage in spirochaetes: species-specific and DNA strand-specific mutational biases.

The genomes of the spirochaetes Borrelia burgdorferi and Treponema pallidum show strong strand-specific skews in nucleotide composition, with the leading strand in replication being richer in G and T than the lagging strand in both species. This mutation bias results in codon usage and amino acid composition patterns that are significantly different between genes encoded on the two strands, in both species. There are also substantial differences between the species, with T.pallidum having a much higher G+C content than B. burgdorferi. These changes in amino acid and codon compositions represent neutral sequence change that has been caused by strong strand- and species-specific mutation pressures. Genes that have been relocated between the leading and lagging strands since B. burgdorferi and T.pallidum diverged from a common ancestor now show codon and amino acid compositions typical of their current locations. There is no evidence that translational selection operates on codon usage in highly expressed genes in these species, and the primary influence on codon usage is whether a gene is transcribed in the same direction as replication, or opposite to it. The dnaA gene in both species has codon usage patterns distinctive of a lagging strand gene, indicating that the origin of replication lies downstream of this gene, possibly within dnaN. Our findings strongly suggest that gene-finding algorithms that ignore variability within the genome may be flawed.

Amino Acids↗

Translational selection is operative for synonymous codon usage in Clostridium perfringens and Clostridium acetobutylicum.

Here, the codon usage patterns of two Clostridium species (Clostridium perfringens and Clostridium acetobutylicum) are reported. These prokaryotes are characterized by a strong mutational bias towards A+T, a striking excess of coding sequences and purine-rich leading strands of replication, strong GC-skews and a high frequency of genomic rearrangements. As expected, it was found that the mutational bias dominates codon usage but there is some variation of synonymous codon choices among genes in the two species. This variation was investigated using a multivariate statistical approach. In the two species, two major trends were detected. One was related to the location of the sequences in the leading or lagging strand of replication, and the other was associated with the preferential use of putatively translational optimal codons in heavily expressed genes. Analyses of the estimated number of synonymous and non-synonymous substitutions among orthologous genes permit us to postulate that optimal codons might be selected not only for speed but also for accuracy during translation.

Amino Acids↗

Doublet preference and gene evolution.

Doublet preference analysis was carried out on coding and noncoding regions of Escherichia coli, Saccharomyces cerevisiae, and human mitochondrial and nuclear DNA. The preference pattern in 1-2 and 2-3 doublets in E. coli and S. cerevisiae correlated with that in noncoding regions. The 3-1 doublet preference in E. coli genes with low optimal codon frequency and in S. cerevisiae genes also showed a correlation with each of their noncoding doublet preference. A mechanism to explain these double preference correlations in doublet preference is presented: mutational biases, the origin of the noncoding region doublet preference, evolved so as to maintain the 1-2 and 2-3 doublet preference, which is determined by codon usage. These biases then acted on the 3-1 doublet, which was almost free of coding constraints, resulting in a similar preference in this doublet.

Base Composition↗

[Molecular evolution of MHC DQA genes. II. Phylogenetic analysis based on nucleotide substitution and SCU bias].

Phylogenetics of 23 alleles at MHC DQA loci in 7 mammalian species was studied based on their nucleotide (NT) substitution and synonymous codon usage (SCU) bias. (1) It was demonstrated that the NT substitution rates are 1.0 x 10(-9) NT/site/yr for exon2 and 1.3 x 10(-9) NT/site/yr for exon2-4 in a large time scale, which is similar to other nuclear genes, while for mouse and rat the rates are nearly twice as high as above mentioned. (2) The DQA locus diversity and their interallelic diversity developed long after the radiation of mammalian 80Mya (million years ago). The bovine counterpart, of, and with the same recent ancestor of ovine DQA2, remains to be discovered. HLA-DQA2 locus split from HLA-DQA1 ancestor at the time between 12 approximately 20 Mya while allele diversity of HLA-DQA1 emerged and developed from 24 Mya to less than 1 Mya. (3) The phylogenetic trees based on SCU divergence reflect the phylogenetics of MHC DQA genes quite well generally in a new respect and reveal that HLA-DQA2 has a distinctive SCU bias different from all other MHC DQA locianalyzed. It indicates that SCU statistics plays an important and unique role in phylogenetic analysis of orthologous genes. The method to estimate the SCU divergence and SCU similarity was improved in this research.

Animals↗

Relationships between transcriptional and translational control of gene expression in Saccharomyces cerevisiae: a multiple regression analysis.

Natural selection for an increased translation efficiency has been proposed as the main determinant for the bias in codon usage observed in many genes of Saccharomyces cerevisiae. Recently, the efficiency of transcription of a large number of yeast genes has been determined, based on the cellular content of the respective mRNAs: this provides an additional dimension to the study of the multisep process of gene expression. Using a representative set of yeast genes with a known level of transcription, the relationship between transcriptional and translational steps was evaluated by a multiple linear regression model. This analysis demonstrated a positive correlation between the amount of transcript, given as the number of mRNA copies per cell for each individual gene, and indices evaluating the effects of translational selection on the corresponding codon usage pattern. This finding suggests a close association of the cellular mRNA content, regulated also at the transcriptional level, to its efficiency of translation, mediated by a fine-tuning of codon usage strategy. Moreover, multiple regression analysis demonstrated that the transcription level of a gene can be approximately predicted using indices of bias deriving from its nucleotide sequence. This allowed for an extensive investigation of uncharacterized regions of the complete genome sequence of S. cerevisiae, to detect new potential short protein coding genes that were not considered by previous searching procedures. Several small open reading frames exhibiting a statistically significant coding potential were thus identified as good candidates for functional analysis.

Amino Acid Sequence↗

The mitochondrial genome of the mosquito Anopheles gambiae: DNA sequence, genome organization, and comparisons with mitochondrial sequences of other insects.

The entire 15,363 bp mitochondrial genome was cloned and sequenced from the mosquito Anopheles gambiae. With respect to the protein-coding genes, rRNA genes and the control region, the gene order was identical to that reported for other insects. There were significant differences, however, in the position and orientation of specific tRNA loci. The overall nucleotide composition was heavily biased towards adenine and thymine, which accounted for 77.6% of all nucleotides. Comparisons were made with the mitochondrial genomes of other insects on the basis genome size and organization, DNA and putative amino acid sequence data, nucleotide substitutions, codon usage and bias, and patterns of AT enrichment.

Amino Acid Sequence↗

Intron length and codon usage.

The correlation was shown between the length of introns and the codon usage of the coding sequences of the corresponding genes, which in some cases can be related to the level of gene expression. The link is positive in the unicellular organisms, i.e., genes with the longer introns show the higher bias of codon usage. It is most pronounced in baker's yeast, where it is definitely related to the level of gene expression--genes with the higher level of expression have the longer introns. The correlation is inverted in multicellular organisms as compared to unicellular ones. Some organisms, however, do not show the link. The presence or absence of the link does not seem to be related to the GC percent of the coding sequences.

Animals↗

In vivo introduction of unpreferred synonymous codons into the Drosophila Adh gene results in reduced levels of ADH protein.

The evolution of codon bias, the unequal usage of synonymous codons, is thought to be due to natural selection for the use of preferred codons that match the most abundant species of isoaccepting tRNA, resulting in increased translational efficiency and accuracy. We examined this hypothesis by introducing 1, 6, and 10 unpreferred codons into the Drosophila alcohol dehydrogenase gene (Adh). We observed a significant decrease in ADH protein production with number of unpreferred codons, confirming the importance of natural selection as a mechanism leading to codon bias. We then used this empirical relationship to estimate the selection coefficient (s) against unpreferred synonymous mutations and found the value (s >or= 10(-5)) to be approximately one order of magnitude greater than previous estimates from population genetics theory. The observed differences in protein production appear to be too large to be consistent with current estimates of the strength of selection on synonymous sites in D. melanogaster.

Alcohol Dehydrogenase↗

Correlation of codon bias measures with mRNA levels: analysis of transcriptome data from Escherichia coli.

Although codon usage is often represented by a 61-dimensional vector, the ability of determining the codon bias in a gene relies on a uni-dimensional vector which measures the total bias in usage of synonymous codons. Codon usage is receiving more and more focus because codon biases might be valuable tools to predict and optimize gene/protein expression. How good any of these measures is for correlating codon usage with gene and protein expression has yet to be investigated. In this study, we correlated gene transcript levels in Escherichia coli with codon usage, using a number of different codon bias measures. We found that there is a significant correlation between transcript levels and codon bias measures, suggesting that these measures can be used to assess or predict gene expression. The codon bias measure performing best in this context was the codon adaptation index.

Codon↗

Large-scale analyses of synonymous substitution rates can be sensitive to assumptions about the process of mutation.

A popular approach to examine the roles of mutation and selection in the evolution of genomes has been to consider the relationship between codon bias and synonymous rates of molecular evolution. A significant relationship between these two quantities is taken to indicate the action of weak selection on substitutions among synonymous codons. The neutral theory predicts that the rate of evolution is inversely related to the level of functional constraint. Therefore, selection against the use of non-preferred codons among those coding for the same amino acid should result in lower rates of synonymous substitution as compared with sites not subject to such selection pressures. However, reliably measuring the extent of such a relationship is problematic, as estimates of synonymous rates are sensitive to our assumptions about the process of molecular evolution. Previous studies showed the importance of accounting for unequal codon frequencies, in particular when synonymous codon usage is highly biased. Yet, unequal codon frequencies can be modeled in different ways, making different assumptions about the mutation process. Here we conduct a simulation study to evaluate two different ways of modeling uneven codon frequencies and show that both model parameterizations can have a dramatic impact on rate estimates and affect biological conclusions about genome evolution. We reanalyze three large data sets to demonstrate the relevance of our results to empirical data analysis.

Amino Acid Substitution↗

Codon usage decreases the error minimization within the genetic code.

The genetic code is not random but instead is organized in such a way that single nucleotide substitutions are more likely to result in changes between similar amino acids. This fidelity, or error minimization, has been proposed to be an adaptation within the genetic code. Many models have been proposed to measure this adaptation within the genetic code. However, we find that none of these consider codon usage differences between species. Furthermore, use of different indices of amino acid physicochemical characteristics leads to different estimations of this adaptation within the code. In this study, we try to establish a more accurate model to address this problem. In our model, a weighting scheme is established for mistranslation biases of the three different codon positions, transition/transversion biases, and codon usage. Different indices of amino acids' physicochemical characteristics are also considered. In contrast to pervious work, our results show that the natural genetic code is not fully optimized for error minimization. The genetic code, therefore, is not the most optimized one for error minimization, but one that balances between flexibility and fidelity for different species.

Amino Acid Substitution↗

Codon context.

The analysis of coding sequences reveals nonrandomness in the context of both sense and stop codons. Part of this is related to nucleotide doublet preference, seen also in non-coding sequences and thought to arise from the dependence of mutational events on surrounding sequence. Another nonrandom context element, relating the wobble nucleotides of successive codons, is observed even when doublet preference, codon usage and bias in amino acid doublets are all allowed for. Several phenomena related to protein synthesis have been shown in vivo to be affected by the nucleotide sequence around codons. Thus, nonsense and missense suppression, elongation rate, precision of tRNA selection and polypeptide chain termination are all affected by codon context. At present, it remains unclear how these phenomena may influence the evolution of nonrandomness in the context of codons in natural sequences.

Codon↗

Distribution and evolution of sequence characteristics in the E. coli genome.

The mean (G + C) composition (51.0%) and standard deviation (+/- 3.8%) of published DNA sequences accounting for 10% of the E. coli genome is in excellent agreement with the principal overall distribution determined by high resolution melting. While differences in base and neighbor characteristics are small and uniform throughout all regions of the genome, it is found that the (G + C) content of sequences varies in segmented fashion within boundaries corresponding to coding (53% G + C) and noncoding (46% G + C) regions; with variances in the latter being six-fold greater than in coding regions. The variance in different regions shows a strong negative dependence on (G + C) content of the region, reflecting the condition that A-T and G-C base pairs are preferred neighbors of A-T and C-G pairs, respectively; with the bias increasing with decreasing (G + C) content. Neighbor analysis indicates the most extreme positive biases occur in AA, TT, GC and CG throughout all regions, but particularly in noncoding regions. Extraordinary numbers of oligomeric strings of (A)n, etc., are the further consequence of this bias. These and other characteristics point to the existence of inherent biases in neighbor frequencies levied during replication or repair, and which reflect, in turn, neighbor influences during mutation. The bias in codon usage noted by Grantham and others is seen here as due, in part, to the adaptation of coding sequences to this microenvironment through selection among synonymous codons so as to preserve inherent neighbor biases.

Base Composition↗

Codon usage and tRNA genes in eukaryotes: correlation of codon usage diversity with translation efficiency and with CG-dinucleotide usage as assessed by multivariate analysis.

The species-specific diversity of codon usage in five eukaryotes (Schizosaccharomyces pombe, Caenorhabditis elegans, Drosophila melanogaster, Xenopus laevis, and Homo sapiens) was investigated with principal component analysis. Optimal codons for translation were predicted on the basis of tRNA-gene copy numbers. Highly expressed genes, such as those encoding ribosomal proteins and histones in S. pombe, C. elegans, and D. melanogaster, have biased patterns of codon usage which have been observed in a wide range of unicellular organisms. In S. pombe and C. elegans, codons contributing positively to the principal component with the largest variance (Z1-parameter) corresponded to the optimal codons which were predicted on the basis of tRNA gene numbers. In D. melanogaster, this correlation was less evident, and the codons contributing positively to the Z1-parameter corresponded primarily to codons with a C or G in the codon third position. In X. laevis and H. sapiens, codon usage in the genes encoding ribosomal proteins and histones was not significantly biased, suggesting that the primary factor influencing codon-usage diversity in these species is not translation efficiency. Codon-usage diversity in these species is known to reflect primarily isochore structures. In the present study, the second additional factor was explained by the level of use of codons containing CG-dinucleotides, and this is discussed with respect to transcription regulation via methylation of CG-dinucleotides, which is observed in mammalian genomes.

Animals↗

Sequence analysis of the cDNA encoding human liver glycogen phosphorylase reveals tissue-specific codon usage.

We have cloned the cDNA encoding glycogen phosphorylase (1,4-alpha-D-glucan:orthophosphate alpha-D-glucosyl-transferase, EC 2.4.1.1) from human liver. Blot-hybridization analysis using a large fragment of the cDNA to probe mRNA from rabbit brain, muscle, and liver tissues shows preferential hybridization to liver RNA. Determination of the entire nucleotide sequence of the liver message has allowed a comparison with the previously determined rabbit muscle phosphorylase sequence. Despite an amino acid identity of 80%, the two cDNAs exhibit a remarkable divergence in G+C content. In the muscle phosphorylase sequence, 86% of the nucleotides at the third codon position are either deoxyguanosine or deoxycytidine residues, while in the liver homolog the figure is only 60%, resulting in a strikingly different pattern of codon usage throughout most of the sequence. The liver phosphorylase cDNA appears to represent an evolutionary mosaic; the segment encoding the N-terminal 80 amino acids contains greater than 90% G+C at the third codon position. A survey of other published mammalian cDNA sequences reveals that the data for liver and muscle phosphorylases reflects a bias in codon usage patterns in liver and muscle coding sequences in general.

Animals↗

Codon-anticodon assignment and detection of codon usage trends in seven microbial genomes.

We have assigned codon-anticodon recognition patterns for the whole set of transfer RNAs of Haemophilus influenzae Rd, Methanococcus jannaschii, and Synechocystis sp. PCC6803 using sequence information derived from the complete genome sequence of these organisms and have tabulated them along with those previously reported for Escherichia coli, Mycoplasma genitalium, Mycoplasma pneumoniae, and Saccharomyces cerevisiae. Using the resulting codon-anticodon tables, the bias in codon usage of genes encoding the entire protein and ribosomal protein complement of each of the seven microbial genomes was analyzed. Then, the codon adaptation index (CAIrp) for each protein gene was calculated using the codon usage preference of the ribosomal protein genes of the corresponding organism. Of the seven genomes examined, six showed CAIrp scores that roughly coincided with the expected level of gene expression. The result demonstrates that CAIrp analysis may be useful for prediction of the expression level of unknown genes when all or at least considerable portions of the genome sequence are available.

Codon↗

Mitochondrial cytochrome C oxidase subunit I of Manduca sexta and a comparison with other invertebrate genes.

A cDNA encoding mitochondrial cytochrome c oxidase subunit I (mt COI) from Manduca sexta (Lepidoptera: Sphingidae) was cloned and sequenced. AT (adenine-thymine) content is high and codon usage is biased and likely reflects the role of mt COI in electron transport. The encoded protein is 514 amino acids long, contains seven invariant His residues observed in COIs in all organisms and would be predicted to be composed of 12 transmembrane regions.

Amino Acid Sequence↗

Three yeast genes, PIR1, PIR2 and PIR3, containing internal tandem repeats, are related to each other, and PIR1 and PIR2 are required for tolerance to heat shock.

We isolated three highly homologous genes, PIR1, PIR2 and PIR3, collectively called the PIR genes. The remarkable feature of their putative amino acid sequence is that they contain a sequence consisting of 18-19 amino acid residues repeated tandemly seven to ten times. Genes homologous to PIR were found in Kluyveromyces lactis and Zygosaccharomyces rouxii but not in Schizosaccharomyces pombe, suggesting that a set of PIR genes plays some role in budding yeast. Bias of codon usage seen in each of the PIR translation products suggests that they are expressed abundantly. The fact that disruption of each gene is viable indicates that none of them is essential. The double disruptants, pir1 pir2, were viable under various conditions, such as higher temperature (37 degrees C) or high salt concentration, but showed a slow-growing phenotype on an agar slab. Furthermore, they were sensitive to heat shock. Addition of a pir3 disruption to the pir1 pir2 double disruptant brought about no phenotypic difference from the original double mutant. PIR1 and PIR3 are closely linked to each other and are on chromosome XI.

Amino Acid Sequence↗