Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

An evolutionary perspective on synonymous codon usage in unicellular organisms.

Observed patterns of synonymous codon usage are explained in terms of the joint effects of mutation, selection, and random drift. Examination of the codon usage in 165 Escherichia coli genes reveals a consistent trend of increasing bias with increasing gene expression level. Selection on codon usage appears to be unidirectional, so that the pattern seen in lowly expressed genes is best explained in terms of an absence of strong selection. A measure of directional synonymous-codon usage bias, the Codon Adaptation Index, has been developed. In enterobacteria, rates of synonymous substitution are seen to vary greatly among genes, and genes with a high codon bias evolve more slowly. A theoretical study shows that the patterns of extreme codon bias observed for some E. coli (and yeast) genes can be generated by rather small selective differences. The relative plausibilities of various theoretical models for explaining nonrandom codon usage are discussed.

Amino Acid Sequence↗

Revisiting the codon adaptation index from a whole-genome perspective: analyzing the relationship between gene expression and codon occurrence in yeast using a variety of models.

Highly expressed genes in many bacteria and small eukaryotes often have a strong compositional bias, in terms of codon usage. Two widely used numerical indices, the codon adaptation index (CAI) and the codon usage, use this bias to predict the expression level of genes. When these indices were first introduced, they were based on fairly simple assumptions about which genes are most highly expressed: the CAI was originally based on the codon composition of a set of only 24 highly expressed genes, and the codon usage on assumptions about which functional classes of genes are highly expressed in fast-growing bacteria. Given the recent advent of genome-wide expression data, we should be able to improve on these assumptions. Here, we measure, in yeast, the degree to which consideration of the current genome-wide expression data sets improves the performance of both numerical indices. Indeed, we find that by changing the parameterization of each model its correlation with actual expression levels can be somewhat improved, although both indices are fairly insensitive to the exact way they are parameterized. This insensitivity indicates a consistent codon bias amongst highly expressed genes. We also attempt direct linear regression of codon composition against genome-wide expression levels (and protein abundance data). This has some similarity with the CAI formalism and yields an alternative model for the prediction of expression levels based on the coding sequences of genes. More information is available at http://bioinfo.mbb.yale.edu/expression/codons.

Codon↗

Clonorchis sinensis: codon usage in nuclear genes.

Codon usage in Clonorchis sinensis was analyzed using 12,515 codons from 38 coding sequences. Total GC content was 49.83%, and GC1, GC2 and GC3 contents were 56.32%, 43.15% and 50.00%, respectively. The effective number of codons converged at 51-53 codons. When plotted against total GC content or GC3, codon usage was distributed in relation to GC3 biases. Relative synonymous codon usage for each codon revealed a single major trend, which was highly correlated with GC content at the third position when codons began with A or U at the first two positions. In codons beginning with G or C base at the first two positions, the G or C base rarely occurred at the third position. These results suggest that codon usage is shaped by a bias towards G or C at the third base, and that this is affected by the first and second bases.

Amino Acids↗

Small regions of preferential codon usage and their effect on overall codon bias--the case of the plp gene.

Preferential codon usage within synonymous groups (codon bias), is hypothesised to be a consequence of either mutational pressure or translational selection. This paper examines PLP and DM-20, two transcripts differentially spliced from the plp gene, using a variety of previously described indices measuring codon bias as well as a novel index for quantifying the response to directional mutation pressure. The results demonstrate that small regions of extreme codon bias may have appreciable effects on the quantification of total bias, and that this effect may be either an increase or a decrease depending which index is used.

Animals↗

Codon usage analysis of Ascaris species influence of base and intercodon frequencies on the synonymous codon usage.

Patterns of codon usage and bias were characterized in the genus Ascaris. Furthermore, the influence of base composition and intercodon frequencies on codon usage was investigated using freely available analytical software. Results showed that A and T were present in the genome at a higher frequency than G and C. As well, codons with AT base pairs at the wobble position were used more often than those with GC at this site, suggesting that the bias extends to the codon level. The presence of T at the intercodon position inhibited the use of codon with A at the wobble site. With respect to amino acid frequency, Serine, Leucine, and Arginine accounted for a disproportionately large fraction of the overall amino acid distribution per gene, implying that the bias was also conserved at the protein level.

Amino Acids↗

Selective constraints on codon usage of nuclear genes from Arabidopsis thaliana.

Highly expressed nuclear genes from Arabidopsis thaliana show an increased frequency of codons that match abundant tRNAs, and it has been suggested that this reflects a selective pressure to increase translation efficiency. Here we explore the possibility that the difference in codon usage between highly expressed genes and other Arabidopsis genes is not the result of selection but, rather, arises from mutation biases. Specifically, we explore the possibility that an influence of transcription level on mutational properties coupled with a context dependency of mutations, both of which have been observed in various organisms, contribute to variation in codon-usage bias across genes. Using noncoding sites immediately flanking both high- and low-expression-coding sequences to infer context-dependent composition biases, we analyze codon-usage bias across genes. The data show that mutation bias cannot explain codon usage of high-expression genes in Arabidopsis and, surprisingly, also indicate that even low-expression genes are under selective constraints. In addition, the data indicate that the general preference for certain codons is context dependent; the composition of the 3' nucleotide, that is, the first position of the next codon, is correlated with what codon is found at an increased frequency in highly expressed genes. This context dependency indicates that selective pressure on codon usage is more complex than previously thought. Overall, the study supports previous suggestions that selection plays a significant role in determining codon usage of nuclear genes in A. thaliana.

Arabidopsis↗

Nucleotide sequence of the mitochondrial structural gene for subunit 9 of yeast ATPase complex.

We have determined the nucleotide sequence of a segment of Saccharomyces mtDNA that contains the structural gene for one of the subunits (the dicyclohexylcarbodiimide-binding protein) of the mitochondrial ATPase complex. The sequence fits the known amino acid sequence of this protein with the exception of one amino acid. Codon usage is biased in favor of A + T-rich codons. On both sides of the gene, the nucleotide sequence contains less than 4% (mol/mol) G + C for at least 180 nucleotides; these A + T sequences show no evidence of internal repetition. The gene and all the A + T-rich sequence preceding the gene are present in a 12S RNA that is the major transcript of this segment of mtDNA. The nature of the sequences responsible for binding ribosomes to mitochondrial mRNA and for termination of RNA synthesis is considered.

Adenosine Triphosphatases↗

Patterns of synonymous codon usage in Drosophila melanogaster genes with sex-biased expression.

The nonrandom use of synonymous codons (codon bias) is a well-established phenomenon in Drosophila. Recent reports suggest that levels of codon bias differ among genes that are differentially expressed between the sexes, with male-expressed genes showing less codon bias than female-expressed genes. To examine the relationship between sex-biased gene expression and level of codon bias on a genomic scale, we surveyed synonymous codon usage in 7276 D. melanogaster genes that were classified as male-, female-, or non-sex-biased in their expression in microarray experiments. We found that male-biased genes have significantly less codon bias than both female- and non-sex-biased genes. This pattern holds for both germline and somatically expressed genes. Furthermore, we find a significantly negative correlation between level of codon bias and degree of sex-biased expression for male-biased genes. In contrast, female-biased genes do not differ from non-sex-biased genes in their level of codon bias and show a significantly positive correlation between codon bias and degree of sex-biased expression. These observations cannot be explained by differences in chromosomal distribution, mutational processes, recombinational environment, gene length, or absolute expression level among genes of the different expression classes. We propose that the observed codon bias differences result from differences in selection at synonymous and/or linked nonsynonymous sites between genes with male- and female-biased expression.

Animals↗

Rates of nucleotide substitution and mammalian nuclear gene evolution. Approximate and maximum-likelihood methods lead to different conclusions.

Rates and patterns of synonymous and nonsynonymous substitutions have important implications for the origin and maintenance of mammalian isochores and the effectiveness of selection at synonymous sites. Previous studies of mammalian nuclear genes largely employed approximate methods to estimate rates of nonsynonymous and synonymous substitutions. Because these methods did not account for major features of DNA sequence evolution such as transition/transversion rate bias and unequal codon usage, they might not have produced reliable results. To evaluate the impact of the estimation method, we analyzed a sample of 82 nuclear genes from the mammalian orders Artiodactyla, Primates, and Rodentia using both approximate and maximum-likelihood methods. Maximum-likelihood analysis indicated that synonymous substitution rates were positively correlated with GC content at the third codon positions, but independent of nonsynonymous substitution rates. Approximate methods, however, indicated that synonymous substitution rates were independent of GC content at the third codon positions, but were positively correlated with nonsynonymous rates. Failure to properly account for transition/transversion rate bias and unequal codon usage appears to have caused substantial biases in approximate estimates of substitution rates.

Animals↗

Codon usage in Aspergillus nidulans.

Synonymous codon usage in genes from the ascomycete (filamentous) fungus Aspergillus nidulans has been investigated. A total of 45 gene sequences has been analysed. Multivariate statistical analysis has been used to identify a single major trend among genes. At one end of this trend are lowly expressed genes, whereas at the other extreme lie genes known or expected to be highly expressed. The major trend is from nearly random codon usage (in the lowly expressed genes) to codon usage that is highly biased towards a set of 19-20 "optimal" codons. The G + C content of the A. nidulans genome is close to 50%, indicating little overall mutational bias, and so the codon usage of lowly expressed genes is as expected in the absence of selection pressure at silent sites. Most of the optimal codons are C- or G- ending, making highly expressed genes more G + C-rich at silent sites.

Aspergillus nidulans↗

Characterization of two divergent beta-tubulin genes from Colletotrichum graminicola.

We have cloned and sequenced two beta-tubulin genes, TUB1 and TUB2, from the phytopathogenic fungus, Colletotrichum graminicola. The nucleotide sequences of the coding regions of the two genes are only 72.8% homologous. This divergence is reflected in the deduced amino acid (aa) sequences which differ at 94 aa residues. Comparison with the aa sequences of other fungal beta-tubulins indicates that the C. graminicola TUB2 gene encodes a conserved isotype, whereas the C. graminicola TUB1 product is highly divergent. Both genes contain six identically placed introns and the position of each intron is conserved in other fungal beta-tubulin genes. Also typical of other fungal beta-tubulin genes, there is a pronounced bias in codon usage in the C. graminicola TUB2 gene; there is a lesser codon bias in TUB1 from C. graminicola. Both C. graminicola beta-tubulin genes are transcribed and yield similar sized messages.

Amino Acid Sequence↗

Codon usage in Giardia lamblia.

A codon usage table for the intestinal parasite Giardia lamblia was generated by analysis of the nucleotide sequences of eight genes comprising 3,135 codons. Codon usage revealed a biased use of synonymous codons with a preference for NNC codons (42.1%). The codon usage of G. lamblia more closely resembles that of the prokaryote Halobacterium halobium (correlation coefficient r = 0.73) rather than that of other eukaryotic protozoans, i.e. Trypanosoma brucei (r = 0.434) and Plasmodium falciparum (r = -0.31). These observations are consistent with the view that G. lamblia represents the first line of descent from the ancestral cells that first took on eukaryotic features.

Animals↗

Analysis of messages expressed by Echinostoma paraensei miracidia and sporocysts, obtained by random EST sequencing.

A lambdaZAP Express cDNA library was constructed with mRNA obtained from immature miracidia within eggs, hatched miracidia, and sporocysts of Echinostoma paraensei. This cDNA library was amplified and 213 expressed sequence tag (EST) sequences (averaging 466 nucleotides in length) were obtained. The mean percentage of unresolved bases within the EST sequences was 0.4%, ranging from 0 to 4.6%. The 213 ESTs represent 151 unique messages. BLAST (version 2.0.8) analysis disclosed that 64 unique E. paraensei messages (42.4%) had significant similarities (BLAST score < or =e-5), at deduced amino acid or nucleotide levels, with known sequences in the nonredundant GenBank databases or the dbEST database (NCBI). The remainder, 57.6% of the unique EST-encoded messages, scored nonsignificant hits. Most of the E. paraensei messages that could be assigned a cellular role based on sequence similarities were involved in gene/protein expression. Several ESTs scored highest similarities with sequences obtained from trematode species. A total of 22,560 nucleotides present in open reading frames from ESTs that aligned with known sequences was used to determine codon usage for E. paraensei. Analysis of a subset of eight ESTs that contained full-length open reading frames did not reveal a bias in codon usage. Also, EST sequences were found to contain 3' untranslated regions with an average length of 69.9 +/- 88.4 nucleotides (n = 46). The EST sequences were submitted to GenBank/dbEST, adding to the 51 available Echinostoma-derived sequences, to provide reference information for both phylogenetic analysis and study of general trematode biology.

Animals↗

Primary structure of the reaction center from Rhodopseudomonas sphaeroides.

The reaction center is a pigment-protein complex that mediates the initial photochemical steps of photosynthesis. The amino-terminal sequences of the L, M, and H subunits and the nucleotide and derived amino acid sequences of the L and M structural genes from Rhodopseudomonas sphaeroides have previously been determined. We report here the sequence of the H subunit, completing the primary structure determination of the reaction center from R. sphaeroides. The nucleotide sequence of the gene encoding the H subunit was determined by the dideoxy method after subcloning fragments into single-stranded M13 phage vectors. This information was used to derive the amino acid sequence of the corresponding polypeptide. The termini of the primary structure of the H subunit were established by means of the amino and carboxy terminal sequences of the polypeptide. The data showed that the H subunit is composed of 260 residues, corresponding to a molecular weight of 28,003. A molecular weight of 100,858 for the reaction center was calculated from the primary structures of the subunits and the cofactors. Examination of the genes encoding the reaction center shows that the codon usage is strongly biased towards codons ending in G and C. Hydropathy analysis of the H subunit sequence reveals one stretch of hydrophobic residues near the amino terminus; the L and M subunits contain five such stretches. From a comparison of the sequences of homologous proteins found in bacterial reaction centers and photosystem II of plants, an evolutionary tree was constructed. The analysis of evolutionary relationships showed that the L and M subunits of reaction centers and the D1 and D2 proteins of photosystem II are descended from a common ancestor, and that the rate of change in these proteins was much higher in the first billion years after the divergence of the reaction center and photosystem II than in the subsequent billion years represented by the divergence of the species containing these proteins.

Amino Acid Sequence↗

Isolation and characterization of a Ustilago maydis glyceraldehyde-3-phosphate dehydrogenase-encoding gene.

The complete nucleotide sequence of the glyceraldehyde-3-phosphate dehydrogenase gene from the corn smut fungus Ustilago maydis is reported. The gene encodes a 337-amino acid protein, parts of which show sequence identity to corresponding regions of GAPDH-encoding genes from other organisms. A single, putative 407-bp intron interrupts the tenth codon. Codon usage is highly biased for codons ending in cytosine.

Amino Acid Sequence↗

Comparison and cross-species expression of the acetyl-CoA synthetase genes of the Ascomycete fungi, Aspergillus nidulans and Neurospora crassa.

The genes encoding the acetate-inducible enzyme acetyl-coenzyme A synthetase from Neurospora crassa and Aspergillus nidulans (acu-5 and facA, respectively) have been cloned and their sequences compared. The predicted amino acid sequence of the Aspergillus enzyme has 670 amino acid residues and that of the Neurospora enzyme either 626 or 606 residues, depending upon which of the two possible initiation codons is used. The amino acid sequences following the second alternative AUG show 86% homology between the two species; the extended N-terminal sequences show no homology. The Neurospora protein is characterized by the appearance of the S(T)PXX sequence motif where the amino acid homologies break down. The codon usage is biased in both genes, with a marked deficiency, especially in Neurospora, of codons with A in the third position. The facA transcribed sequence contains six introns: one in the long leader sequence, one in the 5' coding sequence not homologous with acu-5, and four within the sequence that is largely similar to that of acu-5. Only one intron, corresponding in size and position to the furthest downstream of the facA introns, is found in acu-5. The evolution of introns during the divergence of these two Ascomycete fungi is discussed. Each of the two genes has been transferred by transformation into the other species. Each species is evidently able to splice out the other's introns. Most transformants have normal acetate-induction of acetyl-CoA synthetase, implying that the two genes respond to transcriptional control signals common to both species, in spite of the striking divergence of their 5' ends.

Acetate-CoA Ligase↗

Codon usage in Tetrahymena and other ciliates.

Codon usage in ciliates was examined by analyzing the coding regions of 22 ciliate genes corresponding to a total of 26,142 nucleotides (8,714 codons). It was found that Tetrahymena, Paramecium and the hypotrichs (Oxytricha and Stylonychia) differed in which synonymous codons were used most frequently by their genes. In fact, the codon choices in highly expressed Tetrahymena genes were more similar to those of yeast genes than those of Paramecium genes. The ciliates do not appear to have unusually strong biases in codon usage frequency when compared to other protists such as yeast. The analysis of the Tetrahymena genes indicated that genes which are highly expressed during normal cell growth have a stronger bias towards using the "preferred" codons than those expressed at lower levels during growth or for brief periods during processes such as conjugation. This conforms to what is found in other protists.

Amino Acids↗

Organization and structure of Volvox alpha-tubulin genes.

Southern analysis of Volvox genomic DNA revealed two genes homologous to Chlamydomonas reinhardtii alpha-tubulin cDNA. Restriction fragment length polymorphism analysis indicated that the two genes are not genetically linked. Clones representing one of the alpha-tubulin genes have been isolated from a genomic library of Volvox carteri f. nagariensis. A 3153 bp BamHI fragment containing the entire alpha-tubulin gene (1802 bp) plus 707 bp of the 5'- and 644 bp of the 3'-untranslated regions has been sequenced, revealing the following features: (1) the derived alpha-tubulin primary structure of 451 amino acids is highly conserved, differing in two residues from the alpha 1- and in two additional residues from the alpha 2-tubulin of C. reinhardtii; (2) in comparison to the C. reinhardtii genes, the Volvox alpha-tubulin gene contains a third intron; positions of the other two introns are precisely conserved; (3) codon usages are biased towards G or C, and against A, in the third position; 19 codons are absent from the alpha-tubulin coding sequence, and 5 of these are not used in any of 7 compiled Volvox genes; (4) transcription begins with an A, 30 bp downstream of the putative TATA box; upstream of the TATA box is a 14 bp sequence similar to consensus sequences found in all 4 C. reinhardtii tubulin genes and believed to regulate promoter function.

Amino Acid Sequence↗