Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Unusual codon bias occurring within insertion sequences in Escherichia coli.

The large open reading frames of insertion sequences from Escherichia coli were examined for their spatial pattern of codon usage bias and distribution of rarely used codons. There is a bias in codon usage that is generally lower toward the terminal ends of the coding regions, which is reflected in the occurrence of an excess of nonpreferred codons in the 3' portions of the coding regions as compared with the 5' portions. In contrast, typical chromosomal genes have a lower codon usage bias toward the 5' ends of the coding regions. These results imply that the selective forces reflected in codon usage bias may differ according to position within the coding sequence. In addition, these constraints apparently differ in important ways between genes contained in insertion sequences and those in the chromosome.

Chromosomes, Bacterial↗

Comparison of the patterns of codon usage and bias between Brugia, Echinococcus, Onchocerca and Schistosoma species.

Patterns of codon usage and bias were compared among taxa of the genera Brugia, Echinococcus, Onchocerca and Schistosoma by metric multidimensional scaling and three commonly used indices of bias: Nc, GC3S and B. The overall codon usage for each taxon was compared, as was the codon usage for each individual gene within the taxa. Differences in the patterns of codon usage observed between taxa were dependent on the overall base composition of the genes analysed. The codon usage of Echinococcus was distinct from that of the other taxa. Furthermore, the pattern of codon usage detected by the average codon usage summed across all genes for each taxon was not shown by all genes from that taxon.

Animals↗

On the rate of DNA sequence evolution in Drosophila.

Analysis of the rate of nucleotide substitution at silent sites in Drosophila genes reveals three main points. First, the silent rate varies (by a factor of two) among nuclear genes; it is inversely related to the degree of codon usage bias, and so selection among synonymous codons appears to constrain the rate of silent substitution in some genes. Second, mitochondrial genes may have evolved only as fast as nuclear genes with weak codon usage bias (and two times faster than nuclear genes with high codon usage bias); this is quite different from the situation in mammals where mitochondrial genes evolve approximately 5-10 times faster than nuclear genes. Third, the absolute rate of substitution at silent sites in nuclear genes in Drosophila is about three times higher than the average silent rate in mammals.

Animals↗

[Expression and secretion of human bone morphogenetic protein-7 in Pichia pastoris].

The synonymous codons are used in a highly non-random manner in hosts of widely divergent species, which is termed "codon usage bias". Several reports suggest that codon usage bias sometimes frustrate attempts to express high levels of exogenous genes. In this study, we attempted to express mature peptide of human bone morphogenetic protein-7(hBMP7), with optimized codons in P. pastoris expression system. Three low-usage ARG codons (CGG or CGA) of gene fragment coding the mature peptide of hBMP7 have been successfully converted into P. pastoris-preferred ARG codons (AGA) by overlap extension PCR-based multiple-site-directed mutagenesis for a high level expression of hBMP7 mature peptide. The present results showed that the production level (25.45 mg/L) of codon-optimized hbmp7 had a remarkably improvement of 4.6-fold relative to that (5.5 mg/L) of non-codon-optimized hbmp7. Furthermore, a strain haboring multi-copy of codon-optimized hbmp7 expression cassette was screened, and showed a increased level of expression with 2-fold more potent than the single-copy one. The recombinant hBMP7 mature peptide were produced as a 18 kD monomer proteins, and were easily purified from culture supernatants by using ion-exchange chromatography. Functional assay demonstrated that rhBMP7 could induce ectopic cartilage formation, although its inductive ability was much less active than CHO cell-derived hBMP7.

Animals↗

Relationship of codon bias to mRNA concentration and protein length in Saccharomyces cerevisiae.

In 1982, Ikemura reported a strikingly unequal usage of different synonymous codons, in five Saccharomyces cerevisiae nuclear genes having high protein levels. To study this trend in detail, we examined data from three independent studies that used oligonucleotide arrays or SAGE to estimate mRNA concentrations for nearly all genes in the genome. Correlation coefficients were calculated for the relationship of mRNA concentration to four commonly used measures of synonymous codon usage bias: the codon adaptation index (CAI), the codon bias index (CBI), the frequency of optimal codons (F(op)), and the effective number of codons (N(c)). mRNA concentration was best approximated as an exponential function of each of these four measures. Of the four, the CAI was the most strongly correlated with mRNA concentration (r(s)=0.62+/-0.01, n=2525, p<10(-17)). When we controlled for CAI, mRNA concentration and protein length were negatively correlated (partial r(s)=-0.23+/-0.01, n=4765, p<10(-17)). This may result from selection to reduce the size of abundant proteins to minimize transcriptional and translational costs. When we controlled for mRNA concentration, protein length and CAI were positively correlated (partial r(s)=0.16+/-0.01, n=4765, p<10(-17)). This may reflect more effective selection in longer genes against missense errors during translation. The correlation coefficients between the mRNA levels of individual genes, as measured by different investigators and methods, were low, in the range r(s)=0.39-0.68.

Codon↗

Analysis of synonymous codon usage in SARS Coronavirus and other viruses in the Nidovirales.

In this study, we calculated the codon usage bias in severe acute respiratory syndrome Coronavirus (SARSCoV) and performed a comparative analysis of synonymous codon usage patterns in SARSCoV and 10 other evolutionary related viruses in the Nidovirales. Although there is a significant variation in codon usage bias among different SARSCoV genes, codon usage bias in SARSCoV is a little slight, which is mainly determined by the base compositions on the third codon position. By comparing synonymous codon usage patterns in different viruses, we observed that synonymous codon usage pattern in these virus genes was virus specific and phylogenetically conserved, but it was not host specific. Phylogenetic analysis based on codon usage pattern suggested that SARSCoV was diverged far from all three known groups of Coronavirus. Compositional constraints could explain most of the variation of synonymous codon usage among these virus genes, while gene function is also correlated to synonymous codon usages to a certain extent. However, translational selection and gene length have no effect on the variations of synonymous codon usage in these virus genes.

Base Composition↗

Codon bias in actin multigene families and effects on the reconstruction of phylogenetic relationships.

Codon usage patterns and phylogenetic relationships in the actin multigene family have been analyzed for three dipteran species--Drosophila melanogaster, Bactrocera dorsalis, and Ceratitis capitata. In certain phylogenetic tree reconstructions, using synonymous distances, some gene relationships are altered due to a homogenization phenomenon. We present evidence to show that this homogenization phenomenon is due to codon usage bias. A survey of the pattern of synonymous codon preferences for 11 actin genes from these three species reveals that five out of the six Drosophila actin genes show high degrees of codon bias as indicated by scaled chi 2 values. In contrast to this, four out of the five actin genes from the other species have low codon bias values. A Monte Carlo contingency test indicates that for those Drosophila actin genes which exhibit codon bias, the patterns of codon usage are different compared to actin genes from the other species. In addition, the genes exhibiting codon bias also appear to have reduced rates of synonymous substitution. The homogenization phenomenon seen in terms of synonymous substitutions is not observed for nonsynonymous changes. Because of this homogenization phenomenon, "trees" constructed based on synonymous substitutions will be affected. These effects can be overt in the case of multigene families, but similar distortions may underlie reconstructions based on single-copy genes which exhibit codon usage bias.

Actins↗

Evolutionary patterns of codon usage in the chloroplast gene rbcL.

In this study we reconstruct the evolution of codon usage bias in the chloroplast gene rbcL using a phylogeny of 92 green-plant taxa. We employ a measure of codon usage bias that accounts for chloroplast genomic nucleotide content, as an attempt to limit plausible explanations for patterns of codon bias evolution to selection- or drift-based processes. This measure uses maximum likelihood-ratio tests to compare the performance of two models, one in which a single codon is overrepresented and one in which two codons are overrepresented. The measure allowed us to analyze both the extent of bias in each lineage and the evolution of codon choice across the phylogeny. Despite predictions based primarily on the low G + C content of the chloroplast and the high functional importance of rbcL, we found large differences in the extent of bias, suggesting differential molecular selection that is clade specific. The seed plants and simple leafy liverworts each independently derived a low level of bias in rbcL, perhaps indicating relaxed selectional constraint on molecular changes in the gene. Overrepresentation of a single codon was typically plesiomorphic, and transitions to overrepresentation of two codons occurred commonly across the phylogeny, possibly indicating biochemical selection. The total codon bias in each taxon, when regressed against the total bias of each amino acid, suggested that twofold amino acids play a strong role in inflating the level of codon usage bias in rbcL, despite the fact that twofolds compose a minority of residues in this gene. Those amino acids that contributed most to the total codon usage bias of each taxon are known through amino acid knockout and replacement to be of high functional importance. This suggests that codon usage bias may be constrained by particular amino acids and, thus, may serve as a good predictor of what residues are most important for protein fitness.

Amino Acids↗

The primary structure of the alcohol dehydrogenase gene from the fission yeast Schizosaccharomyces pombe.

We have cloned and sequenced the alcohol dehydrogenase gene of the fission yeast Schizosaccharomyces pombe. The gene was isolated by transformation and complementation of a Saccharomyces cerevisiae strain which lacked functional alcohol dehydrogenase with an S. pombe gene bank constructed in the autonomously replicating yeast plasmid YEp13. Southern hybridization analysis indicates that S. pombe contains only one alcohol dehydrogenase gene. The structural region of the gene is 50% homologous to the alcohol dehydrogenase encoding genes of the budding yeast S. cerevisiae. The gene exhibits a very strong codon usage bias; with the set of predominantly used codons generally resembling that which S. cerevisiae employs preferentially. All of the differences in codon usage bias between S. pombe and S. cerevisiae are in the direction of greater G + C content in S. pombe codons. It is argued that this observation supports the hypothesis that selection toward uniform codon-anticodon binding energies contributes to codon usage bias and that the optimum binding energy is, on the average, higher in S. pombe than S. cerevisiae.

Alcohol Dehydrogenase↗

Relating physicochemical properties of amino acids to variable nucleotide substitution patterns among sites.

Markov-process models of codon substitution were implemented that account for features of DNA sequence evolution (such as transition/transversion bias and codon usage bias) as well as heterogeneity of amino acid substitution pattern over sites. The codon (amino acid) sites are assumed to come from several classes (such as secondary structure categories), among which the rate of amino acid substitution and the effect of amino acid chemical properties vary. Parameters are estimated by the maximum likelihood method, which accounts for the phylogenetic relationship among species and corrects for multiple hits at the same site. The likelihood ratio test is used to compare models. Mitochondrial cytochrome b genes of 28 primate species are analyzed. The site-heterogeneity models provide much better fit to previous homogeneous models.

Amino Acid Substitution↗

Selection on codon usage for error minimization at the protein level.

Given the structure of the genetic code, synonymous codons differ in their capacity to minimize the effects of errors due to mutation or mistranslation. I suggest that this may lead, in protein-coding genes, to a preference for codons that minimize the impact of errors at the protein level. I develop a theoretical measure of error minimization for each codon, based on amino acid similarity. This measure is used to calculate the degree of error minimization for 82 genes of Drosophila melanogaster and 432 rodent genes and to study its relationship with CG content, the degree of codon usage bias, and the rate of nucleotide substitution. I show that (i) Drosophila and rodent genes tend to prefer codons that minimize errors; (ii) this cannot be merely the effect of mutation bias; (iii) the degree of error minimization is correlated with the degree of codon usage bias; (iv) the amino acids that contribute more to codon usage bias are the ones for which synonymous codons differ more in the capacity to minimize errors; and (v) the degree of error minimization is correlated with the rate of nonsynonymous substitution. These results suggest that natural selection for error minimization at the protein level plays a role in the evolution of coding sequences in Drosophila and rodents.

Amino Acids↗

Mutation and selection on the anticodon of tRNA genes in vertebrate mitochondrial genomes.

The H-strand of vertebrate mitochondrial DNA is left single-stranded for hours during the slow DNA replication. This facilitates C-->U mutations on the H-strand (and consequently G-->A mutations on the L-strand) via spontaneous deamination which occurs much more frequently on single-stranded than on double-stranded DNA. For the 12 coding sequences (CDS) collinear with the L-strand, NNY synonymous codon families (where N stands for any of the four nucleotides and Y stands for either C or U) end mostly with C, and NNR and NNN codon families (where R stands for either A or G) end mostly with A. For the lone ND6 gene on the other strand, the codon bias is the opposite, with NNY codon families ending mostly with U and NNR and NNN codon families ending mostly with G. These patterns are consistent with the strand-specific mutation bias. The codon usage biased towards C-ending and A-ending in the 12 CDS sequences affects the codon-anticodon adaptation. The wobble site of the anticodon is always G for NNY codon families dominated by C-ending codons and U for NNR and NNN codon families dominated by A-ending codons. The only, but consistent, exception is the anticodon of tRNA-Met which consistently has a 5'-CAU-3' anticodon base-pairing with the AUG codon (the translation initiation codon) instead of the more frequent AUA. The observed CAU anticodon (matching AUG) would increase the rate of translation initiation but would reduce the rate of peptide elongation because most methionine codons are AUA, whereas the unobserved UAU anticodon (matching AUA) would increase the elongation rate at the cost of translation initiation rate. The consistent CAU anticodon in tRNA-Met suggests the importance of maximizing the rate of translation initiation.

Animals↗

Analysis of codon usage pattern in the radioresistant bacterium Deinococcus radiodurans.

The main factors shaping codon usage bias in the Deinococcus radiodurans genome were reported. Correspondence analysis (COA) was carried out to analyze synonymous codon usage bias. The results showed that the main trend was strongly correlated with gene expression level assessed by the "Codon Adaptation Index" (CAI) values, a result that was confirmed by the distribution of genes along the first axis. The results of correlation analysis, variance analysis and neutrality plot indicated that gene nucleotide composition was clearly contributed to codon bias. CDS length was also key factor in dictating codon usage variation. A general tendency of more biased codon usage of genes with longer CDS length to higher expression level was found. Further, the hydrophobicity of each protein also played a role in shaping codon usage in this organism, which could be confirmed by the significant correlation between the positions of genes placed on the first axis and the hydrophobicity values (r=-0.100, P<0.01). In summary, gene expression level played a crucial role, nucleotide mutational bias, CDS length and the hydrophobicity of each protein just in a minor way in shaping the codon usage pattern of D. radiodurans. Notably, 19 codons firstly defined as "optimal codons" may provide useful clues for molecular genetic engineering and evolutionary studying.

Codon↗

An environmental signature for 323 microbial genomes based on codon adaptation indices.

BACKGROUND: Codon adaptation indices (CAIs) represent an evolutionary strategy to modulate gene expression and have widely been used to predict potentially highly expressed genes within microbial genomes. Here, we evaluate and compare two very different methods for estimating CAI values, one corresponding to translational codon usage bias and the second obtained mathematically by searching for the most dominant codon bias. RESULTS: The level of correlation between these two CAI methods is a simple and intuitive measure of the degree of translational bias in an organism, and from this we confirm that fast replicating bacteria are more likely to have a dominant translational codon usage bias than are slow replicating bacteria, and that this translational codon usage bias may be used for prediction of highly expressed genes. By analyzing more than 300 bacterial genomes, as well as five fungal genomes, we show that codon usage preference provides an environmental signature by which it is possible to group bacteria according to their lifestyle, for instance soil bacteria and soil symbionts, spore formers, enteric bacteria, aquatic bacteria, and intercellular and extracellular pathogens. CONCLUSION: The results and the approach described here may be used to acquire new knowledge regarding species lifestyle and to elucidate relationships between organisms that are far apart evolutionarily.

Bacteria↗

Comparative analysis of expressed sequences reveals a conserved pattern of optimal codon usage in plants.

Codon usage bias is a ubiquitous phenomenon, which may be caused by mutational bias, selection, or both. The patterns of codon usage in plants are not well understood. Datasets of expressed sequence tags (ESTs) available for many plant species provide the resources for large-scale comparative analysis of codon usage patterns. We developed a computational approach to translate EST or assembled contig sequences, and then used the coding information for comparative analysis of codon usage in 12 plant species, including 6 eudicots, 5 monocots and the green alga Chlamydomonas reinhardtii. While codon nucleotide composition is highly conserved within eudicots or monocots, there is a significant difference between these two major taxonomic groups of higher plants. The third nucleotide position of codons is AU-rich in the eudicot genomes (35-42% of G+C content), but GC-rich in the monocot genomes (59-61% of G+C content). To identify optimal codons in these species, we used EST counts to estimate gene transcript levels. It was demonstrated that codon usage bias is correlated positively with gene transcript levels. Interestingly, the use of optimal codons appears to be well conserved between eudicots and monocots, and to a lesser degree between the higher plants and C. reinhardtii. Most of the optimal codons end with a C or G base, regardless of the different nucleotide composition in these genomes. The results suggest that plant codon usage is affected by translational selection, and the selective pressure appears to be conserved in the plant kingdom.

Animals↗

A codon-based model of nucleotide substitution for protein-coding DNA sequences.

A codon-based model for the evolution of protein-coding DNA sequences is presented for use in phylogenetic estimation. A Markov process is used to describe substitutions between codons. Transition/transversion rate bias and codon usage bias are allowed in the model, and selective restraints at the protein level are accommodated using physicochemical distances between the amino acids coded for by the codons. Analyses of two data sets suggest that the new codon-based model can provide a better fit to data than can nucleotide-based models and can produce more reliable estimates of certain biologically important measures such as the transition/transversion rate ratio and the synonymous/nonsynonymous substitution rate ratio.

Animals↗

Correlation between Shine--Dalgarno sequence conservation and codon usage of bacterial genes.

In this study, we analyzed the correlation between codon usage bias and Shine--Dalgarno (SD) sequence conservation, using complete genome sequences of nine prokaryotes. For codon usage bias, we adopted the codon adaptation index (CAI), which is based on the codon usage preference of genes encoding ribosomal proteins, elongation factors, heat shock proteins, outer membrane proteins, and RNA polymerase subunit proteins. To compute SD sequence conservation, we used SD motif sequences predicted by Tompa and systematically aligned them with 5'UTR sequences. We found that there exists a clear correlation between the CAI values and SD sequence conservation in the genomes of Escherichia coli, Bacillus subtilis, Haemophilus influenzae, Archaeoglobus fulgidus, Methanobacterium thermoautotrophicum, and Methanococcus jannaschii, and no relationship is found in M. genitalium, M. pneumoniae, and Synechocystis. That is, genes with higher CAI values tend to have more conserved SD sequences than do genes with lower CAI values in these organisms. Some organisms, such as M. thermoautotrophicum, do not clearly show the correlation. The biological significance of these results is discussed in the context of the translation initiation process and translation efficiency.

Base Sequence↗

Evolutionary basis of codon usage and nucleotide composition bias in vertebrate DNA viruses.

Understanding the extent and causes of biases in codon usage and nucleotide composition is essential to the study of viral evolution, particularly the interplay between viruses and host cells or immune responses. To understand the common features and differences among viruses we analyzed the genomic characteristics of a representative collection of all sequenced vertebrate-infecting DNA viruses. This revealed that patterns of codon usage bias are strongly correlated with overall genomic GC content, suggesting that genome-wide mutational pressure, rather than natural selection for specific coding triplets, is the main determinant of codon usage. Further, we observed a striking difference in CpG content between DNA viruses with large and small genomes. While the majority of large genome viruses show the expected frequency of CpG, most small genome viruses had CpG contents far below expected values. The exceptions to this generalization, the large gammaherpesviruses and iridoviruses and the small dependoviruses, have sufficiently different life-cycle characteristics that they may help reveal some of the factors shaping the evolution of CpG usage in viruses.

Animals↗