Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Structural comparison of two nontandemly repeated yeast glyceraldehyde-3-phosphate dehydrogenase genes.

A hybrid plasmid (pgap63) was isolated which contains a second yeast glyceraldehyde-3-phosphate dehydrogenase structural gene. The complete nucleotide sequence of this gene was determined and compared with the primary structure of a yeast glyceraldehyde-3-phosphate dehydrogenase gene (pgap49) which was reported previously (Holland, J.P., and Holland, M.J. (1979) J. Biol. Chem. 254, 9839-9845). Based on the restriction endonuclease cleavage maps of the isolated segments of yeast DNA which contain these genes, the two genes are nontandemly duplicated. Greater than 94% of the nucleotides within the coding regions of these genes are homologous and the polypeptides encoded by the two structural genes differ by only 15 amino acid residues. Both genes have the same, highly biased, codon usage pattern and neither contains intervening sequences. Approximately 100 nucleotides adjacent to the ATG initiation codons and 130 nucleotides beyond the TAA termination codons are greater than 70% homologous. Structures within the flanking sequences of the genes which are potentially relevant to transcriptional and translational control are described. Several sequences (8 to 15 nucleotides in length) are repeated in both the 5' and 3' flanking sequences of the genes in a noninverted fashion. Finally, a rapid procedure for the isolation of spontaneous deletions within hybrid plasmid DNAs is described, as is the isolation of a structural gene deletion in pgap49.

Amino Acid Sequence↗

Factors influencing the synonymous codon and amino acid usage bias in AT-rich Pseudomonas aeruginosa phage PhiKZ.

To reveal how the AT-rich genome of bacteriophage PhiKZ has been shaped in order to carry out its growth in the GC-rich host Pseudomonas aeruginosa, synonymous codon and amino acid usage bias of PhiKZ was investigated and the data were compared with that of P. aeruginosa. It was found that synonymous codon and amino acid usage of PhiKZ was distinct from that of P. aeruginosa. In contrast to P. aeruginosa, the third codon position of the synonymous codons of PhiKZ carries mostly A or T base; codon usage bias in PhiKZ is dictated mainly by mutational bias and, to a lesser extent, by translational selection. A cluster analysis of the relative synonymous codon usage values of 16 myoviruses including PhiKZ shows that PhiKZ is evolutionary much closer to Escherichia coli phage T4. Further analysis reveals that the three factors of mean molecular weight, aromaticity and cysteine content are mostly responsible for the variation of amino acid usage in PhiKZ proteins, whereas amino acid usage of P. aeruginosa proteins is mainly governed by grand average of hydropathicity, aromaticity and cysteine content. Based on these observations, we suggest that codons of the phage-like PhiKZ have evolved to preferentially incorporate the smaller amino acid residues into their proteins during translation, thereby economizing the cost of its development in GC-rich P. aeruginosa.

Amino Acids↗

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena↗

Models of nearly neutral mutations with particular implications for nonrandom usage of synonymous codons.

The population dynamics of nearly neutral mutations are studied using a single-site and a multisite model. In the latter model, the nucleotides in a sequence are completely linked and the selection schemes employed are additive, multiplicative, and additive with a threshold. Although the third selection scheme is very different from the first two, the three schemes produce identical results for a wide range of parameter values. Thus the present study provides a general theory for the population dynamics of nearly neutral mutations because the results can also be used to draw inferences about other selection schemes such as stabilizing selection and synergistic selection. It is shown that the number of slightly deleterious mutations accumulated in a sequence can be considerably larger under the multisite model than under the single-site model, particularly if the sequence is long or if the mutation rate per site is high. The results show that even a very slight selective difference between synonymous codons can produce a strong bias in codon usage. Three alternative explanations for the strong bias in codon usage in bacterial and yeast genes are considered. The implications of the present results for molecular evolution are discussed.

Biological Evolution↗

Selection intensity for codon bias.

The patterns of nonrandom usage of synonymous codons (codon bias) in enteric bacteria were analyzed. Poisson random field (PRF) theory was used to derive the expected distribution of frequencies of nucleotides differing from the ancestral state at aligned sites in a set of DNA sequences. This distribution was applied to synonymous nucleotide polymorphisms and amino acid polymorphisms in the gnd and putP genes of Escherichia coli. For the gnd gene, the average intensity of selection against disfavored synonymous codons was estimated as approximately 7.3 x 10(-9); this value is significantly smaller than the estimated selection intensity against selectively disfavored amino acids in observed polymorphisms (2.0 x 10(-8)), but it is approximately of the same order of magnitude. The selection coefficients for optimal synonymous codons estimated from PRF theory were consistent with independent estimates based on codon usage for threonine and glycine. Across 118 genes in E. coli and Salmonella typhimurium, the distribution of estimated selection coefficients, expressed as multiples of the effective population size, has a mean and standard deviation of 0.5 +/- 0.4. No significant differences were found in the degree of codon bias between conserved positions and replacement positions, suggesting that translational misincorporation is not an important selective constraint among synonymous polymorphic codons in enteric bacteria. However, across the first 100 codons of the genes, conserved amino acids with identical codons have significantly greater codon bias than that of either synonymous or nonidentical codons, suggesting that there are unique selective constraints, perhaps including mRNA secondary structures, in this part of the coding region.

Codon↗

Cloning and characterization of the gene encoding the highly expressed ribosomal protein l3 of the ciliated protozoan Tetrahymena thermophila. Evidence for differential codon usage in highly expressed genes.

We have cloned and characterized the cDNA and the macronuclear genomic copy of the highly conserved ribosomal protein (r-protein) L3 of Tetrahymena thermophila. The r-protein L3 is encoded by a single copy gene interrupted by one intron. The organization of the promoter region exhibits features characteristic of ribosomal protein genes in Tetrahymena. The codon usage of the L3 gene is highly biased. A thorough analysis of codon usage in Tetrahymena genes revealed that genes could be categorized into two classes according to codon usage bias. Class A comprises r-protein genes and a number of other highly expressed genes. Class B comprises weakly expressed genes such as the conjugation induced CnjB and CnjC genes, but surprisingly, this class also contains abundantly expressed genes such as the genes encoding the surface antigens SerH3 and SerH1. Codon usage is slightly more restricted in class A than in class B, but both classes exhibit distinct and different codon usage biases. Class A genes preferentially use C and U in the silent third codon positions, whereas class B genes preferentially use A and U in the silent third codon positions. The analysis suggests that two different strategies have been employed for optimization of codon usage in the A+T-rich genome of Tetrahymena.

Animals↗

Synonymous codon usage in Zea mays L. nuclear genes is varied by levels of C and G-ending codons.

A multivariate statistical method called correspondence analysis was used to examine the codon usage of one-hundred-and-one nuclear genes of maize (Zea mays L.). Forty percent of the variation in codon usage was due to bias toward G or C-ending versus A or U-ending codons. Differences in levels of G-ending codons showed the weakest correlation with the major codon usage bias. The bias toward C or U versus A or G in the silent third nucleotide position of synonymous codons accounted for approximately 10% of the variation in codon usage. The G+C content of the silent third nucleotide position of coding regions was not strongly correlated with G+C content of introns. Codon usage was strongly biased toward codons ending in G or C for a number of highly expressed genes including most light-regulated chloroplast proteins, ABA-induced proteins, histones, and anthocyanin biosynthetic enzymes. Codon usage of genes encoding storage proteins and regulatory proteins, such as transposases, kinases, phosphatases and transcription factors, was more random than that of genes encoding cytosolic enzymes with similar bias toward G or C-ending codons. Codon usage in maize may reflect both regional bias on nucleotide composition and selection on the silent third nucleotide position.

Base Composition↗

[Bias of base composition and codon usage in pseudorabies virus genes].

The complete sequence of the Pseudorabies Virus (PRV) genomic DNA has not yet been determined, primarily because of the high content of G + C nucleotides of about 74%. We examined the base composition and codon usage of the 68 known PRV genes. As a result, we found a strong bias towards GC-rich codons especially NNC or NNG (N represents any one of four nucleotides) in PRV genes. This demonstrated that the usage bias of synonymous codon and amino acid is the main cause of the high G + C content of PRV. The results showed that the genome regions adjacent UL48, UL40, UL14, IE180 genes where the G + C content occurs as pronounced waves are corresponding to the replication origins. It was also found that the codon usage patterns of regulatory genes are apparently different from other PRV genes. A corresponding analysis of amino acid compositions indicated that the bias of codon usage could be related to the differences of gene function.

Amino Acids↗

The targeting of somatic hypermutation.

Somatic hypermutation does not occur randomly within immunoglobulin V genes but, rather, is preferentially targeted to certain nucleotide positions (hot spots) and away from others (cold spots). Cold spots often coincide with residues essential for V gene folding. Hotspots, which appear to be strategically located to favour affinity maturation, are most frequently located in the CDRs (particularly CDR1) though conserved hotspots are also found at the base of FR3. Hotspots are in part created by local DNA sequence and the strong biases of codon usage in V genes indicate that the genes have evolved such that somatic hypermutation is targeted to those parts of the V where it is likely to prove most useful. These features of mutational hotspots and biased codon usage are also evident in V genes of lower animals suggesting that diversification by strategic targeting of non-templated mutation may have evolved early in antigen receptor evolution.

Animals↗

The selection-mutation-drift theory of synonymous codon usage.

It is argued that the bias in synonymous codon usage observed in unicellular organisms is due to a balance between the forces of selection and mutation in a finite population, with greater bias in highly expressed genes reflecting stronger selection for efficiency of translation. A population genetic model is developed taking into account population size and selective differences between synonymous codons. A biochemical model is then developed to predict the magnitude of selective differences between synonymous codons in unicellular organisms in which growth rate (or possibly growth yield) can be equated with fitness. Selection can arise from differences in either the speed or the accuracy of translation. A model for the effect of speed of translation on fitness is considered in detail, a similar model for accuracy more briefly. The model is successful in predicting a difference in the degree of bias at the beginning than in the rest of the gene under some circumstances, as observed in Escherichia coli, but grossly overestimates the amount of bias expected. Possible reasons for this discrepancy are discussed.

Amino Acyl-tRNA Synthetases↗

An analysis of codon usage in mammals: selection or mutation bias?

A new statistical test has been developed to detect selection on silent sites. This test compares the codon usage within a gene and thus does not require knowledge of which genes are under the greatest selection, that there exist common trends in codon usage across genes, or that genes have the same mutation pattern. It also controls for mutational biases that might be introduced by the adjacent bases. The test was applied to 62 mammalian sequences, and significant codon usage biases were detected in all three species examined (humans, rats, and mice). However, these biases appear not to be the consequence of selection, but of the first base pair in the codon influencing the mutation pattern at the third position.

Animals↗

Codon usage in highly expressed genes of Haemophillus influenzae and Mycobacterium tuberculosis: translational selection versus mutational bias.

Biases in the codon usage and base compositions at three codon sites in different genes of A+T-rich Gram-negative bacterium Haemophillus influenzae and G+C-rich Gram-positive bacterium Mycobacterium tuberculosis have been examined to address the following questions: (1) whether the synonymous codon usage in organisms having highly skewed base compositions is totally dictated by the mutational bias as reported previously (Sharp, P.M., Devine, K.M., 1989. Codon usage and gene expression level in Dictyostelium discoideum: highly expressed genes do 'prefer' optimal codons. Nucleic Acids Res. 17, 5029-5039), or is also controlled by translational selection; (2) whether preference of G in the first codon positions by highly expressed genes, as reported in Escherichia coli (Gutierrez, G., Marquez, L., Marin, A., 1996. Preference for guanosine at first codon position in highly expressed Escherichia coli genes. A relationship with translational efficiency. Nucleic Acids Res. 24, 2525-2527), is true in other bacteria; and (3) whether the usage of bases in three codon positions is species-specific. Result presented here show that even in organisms with high mutational bias, translational selection plays an important role in dictating the synonymous codon usage, though the set of optimal codons is chosen in accordance with the mutational pressure. The frequencies of G-starting codons are positively correlated to the level of expression of genes, as estimated by their Codon Adaptation Index (CAI) values, in M. tuberculosis as well as in H. influenzae in spite of having an A+T-rich genome. The present study on the codon preferences of two organisms with oppositely skewed base compositions thus suggests that the preference of G-starting codons by highly expressed genes might be a general feature of bacteria, irrespective of their overall G+C contents. The ranges of variations in the frequencies of individual bases at the first and second codon positions of genes of both H. influenzae and M. tuberculosis are similar to those of E. coli, implying that though the composition of all three codon positions is governed by a selection-mutation balance, the mutational pressure has little influence in the choice of bases at the first two codon positions, even in organisms with highly biased base compositions.

Animals↗

Codon usage in mammalian genes is biased by sequence slippage mechanisms.

The codons for some conserved amino acids are found to be the same between homologous genes from different species when the statistics of codon usage would suggest that they should be different. I examine whether this 'coincidence' of codon usage could be due to genetic mechanisms homogenising the DNA around specific sites. This paper describes the further analysis of the coincident codons in 19 genes (a total of 96 homologues) for slippage. Coincident codons arise in contexts of increased sequence simplicity, and have a high chance of occurring within sequences similar to the recombination-prone minisatellite 'core' sequence. This suggests a role of genetic homogenisation in their generation.

Amino Acid Sequence↗

Synonymous codon usage in Drosophila melanogaster: natural selection and translational accuracy.

I present evidence that natural selection biases synonymous codon usage to enhance the accuracy of protein synthesis in Drosophila melanogaster. Since the fitness cost of a translational misincorporation will depend on how the amino acid substitution affects protein function, selection for translational accuracy predicts an association between codon usage in DNA and functional constraint at the protein level. The frequency of preferred codons is significantly higher at codons conserved for amino acids than at nonconserved codons in 38 genes compared between D. melanogaster and Drosophila virilis or Drosophila pseudoobscura (Z = 5.93, P < 10(-6)). Preferred codon usage is also significantly higher in putative zinc-finger and homeodomain regions than in the rest of 28 D. melanogaster transcription factor encoding genes (Z = 8.38, P < 10(-6)). Mutational alternatives (within-gene differences in mutation rates, amino acid changes altering codon preference states, and doublet mutations at adjacent bases) do not appear to explain this association between synonymous codon usage and amino acid constraint.

Animals↗

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis↗

Codon optimization for high-level expression of human erythropoietin (EPO) in mammalian cells.

Codon bias has been observed in many species. The usage of selective codons in a given gene is positively correlated with its expression efficiency. As an experimental approach to study codon-usage effects on heterologous gene expression in mammalian cells, we designed two human erythropoietin (EPO) genes, one in which native codons were systematically substituted with codons frequently found in highly expressed human genes and the other with codons prevalent in yeast genes. Relative performances of the re-engineered EPO genes were evaluated with various combinations of promoters and signal leader sequences. Under the comparable set of combinations, mature EPO gene with human high-frequency codons gave a considerably higher level of expression than that with yeast high-frequency codons. However, the levels of EPO expression varied, depending on the alternate combinations. Since the promoters and the signal leader sequences that we used are known to be equally efficient in gene expression, we hypothesized that the varied expression levels were due to the linear sequence between the promoter and the coding gene sequence. To test this possibility, we designed the EPO gene with hybrid codon usage in which the 5'-proximal region of the EPO gene was synthesized with yeast-biased codons and the rest with human-biased codons. This codon-usage hybrid EPO gene substantially enhanced the level of EPO transcripts and proteins up to 2.9-fold and 13.8-fold, respectively, when compared to the level reached by the original counterpart. Our results suggest that the linear sequence between the promoter and the 5'-proximal region of a gene plays an important role in achieving high-level expression in mammalian cells.

Amino Acid Sequence↗

Natural selection versus primitive gene structure as determinant of codon usage.

Different codons are not utilized equally in known gene sequences. One of the important biases of codon usage is observed in the form of an enrichment of RNY codons, especially within RNN codon families. Such biases could represent the residue of a primitive repeating-RNY gene structure, or the outcome of natural selection, or both. Analyses based on the rates of silent substitutions, the frequencies of base doublets, and synonymous codon ratios for Escherichia coli, yeast, Drosophila and Xenopus proteins have been performed. The results rule out any significant support for a primitive repeating-RNY or repeating-RRY gene structure, and establish the important role of natural selection in determining the choice of codons. With strong intervention by natural selection, the relationship between primitive gene structure and codon usage necessarily becomes minimal.

Animals↗

Codon usage in bony fishes.

Bony fishes are excellent experimental models that have been used extensively in biochemical and molecular genetic studies. As proteins are isolated and characterized from these organisms, information on codon usage by bony fishes can be used for subsequent recombinant DNA studies. Codon usage and nucleotide bias within codons from three species of bony fishes and two composites of 14 and 15 bony-fish species were analyzed. Although differences in codon usage increased from seven amino acids between fish species to eight amino acids between fish genera, the small number of differences (3 amino acids) between a single species and a fish composite minus that species suggests that codon usage tables constructed from large numbers of fish species are representative of bony fish in general. Furthermore, we found few differences in codon usage between two vertebrate phyla (fish and rat). Codons in fish DNA sequences end predominantly in G or C, even though the coding sequences are not enriched in these nucleotides. This positional base bias can be used to locate putative protein coding regions in fish DNA sequences.

Amino Acids↗