Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Synonymous codon usage in environmental chlamydia UWE25 reflects an evolutional divergence from pathogenic chlamydiae.

Publication of the complete genome sequence for the Acanthamoeba sp. endosymbiont UWE25 has illuminated the evolution history of chlamydiae. In this study, the codon usage bias in UWE25 and five other species of pathogenic chlamydiae was calculated. It was found that genomic composition constraints are the major source of codon usage variation in UWE25. This result is different from the former observation in pathogenic chlamydiae, whose genomic base composition is more unbiased. Four other factors, such as strand-specific mutational bias, natural selection acting at the level of translation, hydropathy level of each protein and the conservation level of amino acids also have influence in shaping the codon usage in these six species to some extent. Further analysis suggests that the high stability of the UWE25 genome partially account for the difference in codon usage pattern between environmental and pathogenic chlamydiae. Moreover, our results imply that the replicational selection pressure in pathogenic chlamydiae is stronger than that in UWE25. Analyzing the codon usage pattern in the environmental chlamydia and comparing it with that of the pathogenic chlamydiae may provide clues how the chlamydiae have evolved from their common ancestor.

Amino Acids↗

HIV-1 gag expression is quantitatively dependent on the ratio of native and optimized codons.

There is a significant variation of codon usage bias among different species and even among genes within the same organisms. Codon optimization, this is, gene redesigning with the use of codons preferred for the specific expression system, results in improved expression of heterologous genes in bacteria, plants, yeast, mammalian cells, and transgenic animals. The mechanisms preventing expression of genes with rare or low-usage codons at adequate levels are not completely elucidated. Human immunodeficiency virus (HIV) represents an interesting model for studying how differences in codon usage affect gene expression in heterologous systems. Construction of synthetic genes with optimized codons demonstrated that the codon-usage effects might be a major impediment to the efficient expression of HIV gag/pol and env gene products in mammalian cells. According to another hypothesis, the poor expression of HIV structural proteins even without HIV context is attributed to the so-called cis-acting inhibitory elements (INS), which are located within the protein-coding region. They consist of AU-rich sequences and may be inactivated through the introduction of multiple mutations over the large regions of gag gene. In our work, we evaluated expression of hybrid HIV-1 gag mRNAs where wild-type (A-rich) gag sequences were combined with artificial sequences. In such "humanized" gag fragments with adapted codon usage, AT-content was significantly reduced in favor of G and C nucleotides without any changes in protein sequence. We show that wild-type gag sequences negatively influence expression of gag-reporter, and the addition of fragments with optimized codons to gag mRNA partially rescues its expression. The results demonstrate that the expression of HIV-1 gag is determined by the ratio of optimized and rare codons within mRNA. Our data also indicates that some wtgag fragments counteract the influence of the other wtgag sequences, which cause the inhibition of gag expression. The presented data do not contradict the concept of INS; yet, it makes the definition of INS more complex. This supports the idea of a broader role of the selected codon usage in influencing the expression of HIV proteins in mammalian cells.

Codon↗

Complete mitochondrial genome sequence of Urechis caupo, a representative of the phylum Echiura.

BACKGROUND: Mitochondria contain small genomes that are physically separate from those of nuclei. Their comparison serves as a model system for understanding the processes of genome evolution. Although hundreds of these genome sequences have been reported, the taxonomic sampling is highly biased toward vertebrates and arthropods, with many whole phyla remaining unstudied. This is the first description of a complete mitochondrial genome sequence of a representative of the phylum Echiura, that of the fat innkeeper worm, Urechis caupo. RESULTS: This mtDNA is 15,113 nts in length and 62% A+T. It contains the 37 genes that are typical for animal mtDNAs in an arrangement somewhat similar to that of annelid worms. All genes are encoded by the same DNA strand which is rich in A and C relative to the opposite strand. Codons ending with the dinucleotide GG are more frequent than would be expected from apparent mutational biases. The largest non-coding region is only 282 nts long, is 71% A+T, and has potential for secondary structures. CONCLUSIONS: Urechis caupo mtDNA shares many features with those of the few studied annelids, including the common usage of ATG start codons, unusual among animal mtDNAs, as well as gene arrangements, tRNA structures, and codon usage biases.

Amino Acid Sequence↗

On the evolution of codon volatility.

Volatility of a codon is defined as the probability that a random point mutation in the codon generates a nonsynonymous change. It has been proposed that higher-than-expected mean codon volatility of a gene indicates that positive selection for nonsynonymous changes has acted on the gene in the recent past. I show that strong frequency-dependent selection (minority advantage) in large populations can increase codon volatility slightly, whereas directional positive selection has no effect on volatility. Factors unrelated to positive selection, such as expression-related or GC-content-related codon usage bias, also affect volatility. These and other considerations suggest that codon volatility has only limited utility for detecting positive selection at the DNA sequence level.

Animals↗

Use and misuse of correspondence analysis in codon usage studies.

Correspondence analysis has frequently been used for codon usage studies but this method is often misused. Because amino acid composition exerts constraints on codon usage, it is common to use tables containing relative codon frequencies (or ratios of frequencies) instead of simple codon counts to get rid of these amino acid biases. The problem is that some important properties of correspondence analysis, such as rows weighting, are lost in the process. Moreover, the use of relative measures sometimes introduces other biases and often diminishes the quantity of information to analyse, occasionally resulting in interpretation errors. For instance, in the case of an organism such as Borrelia burgdorferi, the use of relative measures led to the conclusion that there was no translational selection, while analyses based on codon counts show that there is a possibility of a selective effect at that level. In this paper, we expose these problems and we propose alternative strategies to correspondence analysis for studying codon usage biases when amino acid composition effects must be removed.

Bacillus subtilis↗

Estimating the "effective number of codons": the Wright way of determining codon homozygosity leads to superior estimates.

In 1990, Frank Wright introduced a method for measuring synonymous codon usage bias in a gene by estimation of the "effective number of codons," N(c). Several attempts have been made recently to improve Wright's estimate of N(c), but the methods that work in cases where a gene encodes a protein not containing all amino acids with degenerate codons have not been tested against each other. In this article I derive five new estimators of N(c) and test them together with the two published estimators, using resampling under rigorous testing conditions. Estimation of codon homozygosity, F, turns out to be a key to the estimation of N(c). F can be estimated in two closely related ways, corresponding to sampling with or without replacement, the latter being what Wright used. The N(c) methods that are based on sampling without replacement showed much better accuracy at short gene lengths than those based on sampling with replacement, indicating that Wright's homozygosity method is superior. Surprisingly, the methods based on sampling with replacement displayed a superior correlation with mRNA levels in Escherichia coli.

Codon↗

Comparison and evolutionary analysis of the glycosomal glyceraldehyde-3-phosphate dehydrogenase from different Kinetoplastida.

In this work, we present the sequences and a comparison of the glycosomal GAPDHs from a number of Kinetoplastida. The complete gene sequences have been determined for some species (Crithidia fasciculata, Herpetomonas samuelpessoai, Leptomonas seymouri, and Phytomonas sp), whereas for other species (Trypanosoma brucei gambiense, Trypanosoma congolense, Trypanosoma vivax, and Leishmania major), only partial sequences have been obtained by PCR amplification. The structure of all available glycosomal GAPDH genes was analyzed in detail. Considerable variations were observed in both their nucleotide composition and their codon usage. The GC content varies between 64.4% in L. seymouri and 49.5% in the previously sequenced GAPDH gene from Trypanoplasma borreli. A highly biased codon usage was found in C. fasciculata, with only 34 triplets used, whereas in T. borreli 57 codons were employed. No obvious correlation could be observed between the codon usage and either the nucleotide composition or the level of gene expression. The glycosomal GAPDH is a very well-conserved enzyme. The maximal overall difference observed in the amino acid sequences is only 25%. Specific insertions and extensions are retained in all sequences. The residues involved in catalysis, substrate, and inorganic phosphate binding are fully conserved, whereas some variability is observed in the cofactor-binding pocket. The implications of these data for the design of new trypanocidal drugs targeted against GAPDH are discussed. All available gene and amino acid sequences of glycosomal GAPDHs were used for a phylogenetic analysis. The division of the Kinetoplastida into two suborders, Bodonina and Trypanosomatina, was well supported. Within the letter group, the Trypanosoma species appeared to be monophyletic, whereas the other trypanosomatids form a second clade.

Amino Acid Sequence↗

Synonymous substitution rates in enterobacteria.

It has been shown previously that the synonymous substitution rate between Escherichia coli and Salmonella typhimurium is lower in highly than in weakly expressed genes, and it has been suggested that this is due to stronger selection for translational efficiency in highly expressed genes as reflected in their greater codon usage bias. This hypothesis is tested here by comparing the substitution rate in codon families with different patterns of synonymous codon use. It is shown that the decline in the substitution rate across expression levels is as great for codon families that do not appear to be subject to selection for translational efficiency as for those that are. This implies that selection on translational efficiency is not responsible for the decline in the substitution rate across genes. It is argued that the most likely explanation for this decline is a decrease in the mutation rate. It is also shown that a simple evolutionary model in which synonymous codon use is determined by a balance between mutation, selection for an optimal codon, and genetic drift predicts that selection should have little effect on the substitution rate in the present case.

Codon↗

Background selection in single genes may explain patterns of codon bias.

Background selection involves the reduction in effective population size caused by the removal of recurrent deleterious mutations from a population. Previous work has examined this process for large genomic regions. Here we focus on the level of a single gene or small group of genes and investigate how the effects of background selection caused by nonsynonymous mutations are influenced by the lengths of coding sequences, the number and length of introns, intergenic distances, neighboring genes, mutation rate, and recombination rate. We generate our predictions from estimates of the distribution of the fitness effects of nonsynonymous mutations, obtained from DNA sequence diversity data in Drosophila. Results for genes in regions with typical frequencies of crossing over in Drosophila melanogaster suggest that background selection may influence the effective population sizes of different regions of the same gene, consistent with observed differences in codon usage bias along genes. It may also help to cause the observed effects of gene length and introns on codon usage. Gene conversion plays a crucial role in determining the sizes of these effects. The model overpredicts the effects of background selection with large groups of nonrecombining genes, because it ignores Hill-Robertson interference among the mutations involved.

Animals↗

Large-scale, multi-genome analysis of alternate open reading frames in bacteria and archaea.

Analysis of over 300,000 annotated genes in 105 bacterial and archaeal genomes reveals an unexpectedly high frequency of large (>300 nucleotides) alternate open reading frames (ORFs). Especially notable is the very high frequency of alternate ORFs in frames +3 and -1 (where the annotated gene is defined as frame +1). The occurrence of alternate ORFs is correlated with genomic G+C content and is strongly influenced by synonymous codon usage bias. The frequency of alternate ORFs in frame -1 is also influenced by the occurrence of codons encoding leucine and serine in frame +1. Although some alternate ORFs have been shown to encode proteins, many others are probably not expressed because they lack appropriate signals for transcription and translation. These latter can be mis-annotated by automatic gene finding programs leading to errors in public databases. Especially prone to mis-annotation is frame -1, because it exhibits a potential codon usage and theoretical capacity to encode proteins with an amino acid composition most similar to real genes. Some alternate ORFs are conserved across bacterial or archaeal species, and can give rise to misannotated "conserved hypothetical" genes, while others are unique to a genome and are misidentified as "hypothetical orphan" genes, contributing significantly to the orphan gene paradox.

Algorithms↗

DNAskew: statistical analysis of base compositional asymmetry and prediction of replication boundaries in the genome sequences.

Sueoka and Lobry declared respectively that, in the absence of bias between the two DNA strands for mutation and selection, the base composition within each strand should be A=T and C=G (this state is called Parity Rule type 2, PR2). However, the genome sequences of many bacteria, vertebrates and viruses showed asymmetries in base composition and gene direction. To determine the relationship of base composition skews with replication orientation, gene function, codon usage biases and phylogenetic evolution, in this paper a program called DNAskew was developed for the statistical analysis of strand asymmetry and codon composition bias in the DNA sequence. In addition, the program can also be used to predict the replication boundaries of genome sequences. The method builds on the fact that there are compositional asymmetries between the leading and the lagging strand for replication. DNAskew was written in Perl script language and implemented on the LINUX operating system. It works quickly with annotated or unannotated sequences in GBFF (GenBank flatfile) or fasta format. The source code is freely available for academic use at http://www.epizooty.com/pub/stat/DNAskew.

Algorithms↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

Codon discrimination due to presence of abundant non-cognate competitive tRNA.

It has been thought that preferential use of synonymous codons provides high efficiency and fidelity of protein synthesis through specific codon-anticodon interactions. In yeast genes, some codon boxes seem to prefer a codon which is unsuited for its cognate anticodon. Now, we propose that codon usage biases may arise due to presence of abundant non-cognate competitive tRNA capable of misreading a codon by C-U or G-U pairing in the middle position.

Codon↗

Codon usage in regulatory genes in Escherichia coli does not reflect selection for 'rare' codons.

It has often been suggested that differential usage of codons recognized by rare tRNA species, i.e. "rare codons", represents an evolutionary strategy to modulate gene expression. In particular, regulatory genes are reported to have an extraordinarily high frequency of rare codons. From E. coli we have compiled codon usage data for highly expressed genes, moderately/lowly expressed genes, and regulatory genes. We have identified a clear and general trend in codon usage bias, from the very high bias seen in very highly expressed genes and attributed to selection, to a rather low bias in other genes which seems to be more influenced by mutation than by selection. There is no clear tendency for an increased frequency of rare codons in the regulatory genes, compared to a large group of other moderately/lowly expressed genes with low codon bias. From this, as well as a consideration of evolutionary rates of regulatory genes, and of experimental data on translation rates, we conclude that the pattern of synonymous codon usage in regulatory genes reflects primarily the relaxation of natural selection.

Base Sequence↗

Bradyrhizobium japonicum does not require alpha-ketoglutarate dehydrogenase for growth on succinate or malate.

The sucA gene, encoding the E1 component of alpha-ketoglutarate dehydrogenase, was cloned from Bradyrhizobium japonicum USDA110, and its nucleotide sequence was determined. The gene shows a codon usage bias typical of non-nif and non-fix genes from this bacterium, with 89.1% of the codons being G or C in the third position. A mutant strain of B. japonicum, LSG184, was constructed with the sucA gene interrupted by a kanamycin resistance marker. LSG184 is devoid of alpha-ketoglutarate dehydrogenase activity, indicating that there is only one copy of sucA in B. japonicum and that it is completely inactivated in the mutant. Batch culture experiments on minimal medium revealed that LSG184 grows well on a variety of carbon substrates, including arabinose, malate, succinate, beta-hydroxybutyrate, glycerol, formate, and galactose. The sucA mutant is not a succinate auxotroph but has a reduced ability to use glutamate as a carbon or nitrogen source and an increased sensitivity to growth inhibition by acetate, relative to the parental strain. Because LSG184 grows well on malate or succinate as its sole carbon source, we conclude that B. japonicum, unlike most other bacteria, does not require an intact tricarboxylic acid (TCA) cycle to meet its energy needs when growing on the four-carbon TCA cycle intermediates. Our data support the idea that B. japonicum has alternate energy-yielding pathways that could potentially compensate for inhibition of alpha-ketoglutarate dehydrogenase during symbiotic nitrogen fixation under oxygen-limiting conditions.

Amino Acid Sequence↗

Molecular considerations in the evolution of bacterial genes.

Synonymous and nonsynonymous substitution rates at the loci encoding glyceraldehyde-3-phosphate dehydrogenase (gap) and outer membrane protein 3A (ompA) were examined in 12 species of enteric bacteria. By examining homologous sequences in species of varying degrees of relatedness and of known phylogenetic relationships, we analyzed the patterns of synonymous and nonsynonymous substitutions within and among these genes. Although both loci accumulate synonymous substitutions at reduced rates due to codon usage bias, portions of the gap and ompA reading frames show significant deviation in synonymous substitution rates not attributable to local codon bias. A paucity of synonymous substitutions in portions of the ompA gene may reflect selection for a novel mRNA secondary structure. In addition, these studies allow comparisons of homologous protein-coding sequences (gap) in plants, animals, and bacteria, revealing differences in evolutionary constraints on this glycolytic enzyme in these lineages.

Amino Acid Sequence↗

Gene expressivity is the main factor in dictating the codon usage variation among the genes in Pseudomonas aeruginosa.

Codon usage biases of all DNA sequences (length greater than or equal to 300 bp) from the complete genome of Pseudomonas aeruginosa have been analyzed. As P. aeruginosa is a GC-rich organism, G and/or C are expected to predominate in their codons. Overall codon usage data analysis indicates that indeed codons ending in G and/or C are predominant in this organism. But multivariate statistical analysis indicates that there is a single major trend in the codon usage variation among the genes in this organism, which has a strong negative correlation with the expressivities of the genes. The majority of the lowly expressed genes are scattered towards the positive end of the major axis whereas the highly expressed genes are clustered towards the negative end. This is the first report where the prokaryotic organism having highly skewed base composition is dictated mainly by translational selection, though some other factors such as the lengths of the genes as well as the hydrophobicity of genes also influence the codon usage variation among the genes in this organism in a minor way.

Codon↗

Translation in Bacillus subtilis: roles and trends of initiation and termination, insights from a genome analysis.

We analysed the Bacillus subtilis protein coding sequences termini, and compared it to other genomes. The analysis focused on signals, com-positional biases of nucleotides, oligonucleotides, codons and amino acids and mRNA secondary structure. AUG is the preferred start codon in all genomes, independent of their G+C content, and seems to induce less stable mRNA structures. However, it is not conserved between homologous genes neither is it preferred in highly expressed genes. In B.subtilis the ribosome binding site is very strong. We found that downstream boxes do not seem to exist either in Escherichia coli or in B.subtilis. UAA stop codon usage is correlated with the G+C content and is strongly selected in highly expressed genes. We found less stable mRNA structures at both termini, which we related to mRNA-ribosome and mRNA-release-factor interactions. This pattern seems to impose a peculiar A-rich nucleotide and codon usage bias in these regions. Finally the analysis of all proteins from B.subtilis revealed a similar amino acid bias near both termini of proteins consisting of over-representation of hydrophilic residues. This bias near the stop codon is partially release-factor specific.

Algorithms↗