Search PubMedSearch

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals

Codon usage divergence of homologous vertebrate genes and codon usage clock.

This paper is concerned with the divergence of synonymous codon usage and its bias in three homologous genes within vertebrate species. Genetic distances among species are described in terms of synonymous codon usage divergence and the correlation is found between the genetic distances and taxonomic distances among species under study. A codon usage clock is reported in alpha-globin and beta-globin. A method is developed to define the synonymous codon preference bias and it is observed that the bias changes considerably among species.

Animals

Two regions in human DNA polymerase beta mRNA suppress translation in Escherichia coli.

Although human DNA polymerase beta (DNA pol beta) shows 96% identity with rat DNA pol beta at the amino acid level, it is weakly expressed in Escherichia (E.) coli relative to the rat enzyme. The mechanism of this suppression was investigated. Pulse-chase protein labeling and steady state mRNA analysis showed that mature human DNA pol beta protein is relatively stable in E. coli and the levels of human and rat DNA pol beta mRNA were comparable indicating that the human DNA pol beta expression is suppressed at the translational level. By systematic expression analysis of a number of chimeric genes composed of human and rat cDNAs, two strong translational suppression regions were mapped in the human DNA pol beta mRNA; one was named TSR-1, corresponding to CGG encoding arginine (arg) at position 4 and the other, termed TSR-2, is located between codons 153 and 199. Since substitution of the rat Arg-4 codon with synonymous codons showed strong effects upon the expression level, we propose that the arg codon at the N-terminal coding region plays a role in modulating expression.

Amino Acid Sequence

[Constraints on base sequences in a polynucleotide: I. Significance of the degeneration of the code].

The statistical study of polynucleotide sequences constituting the genes of E. coli, bacteriophages lambda and T7 reveals that constraints act upon nucleic acids (DNA or RNA) and contribute to determine the choice between the synonymous codons. The existence of synonymous codons seems to be the way of satisfying these constraints, keeping the possibility of specifying a large variety of polypeptides. At least in the case of amino acids with a small number of codons, these constraints are strong enough to influence the primary structure of proteins.

Bacteriophages

Frequencies of codons in histones, tubulins and fibrinogen: bias due to interference between transcription signals and protein function.

The distribution of codons was studied in 65 proteins: 48 histones, 14 tubulins, and three fibrinogens, With the methodology used, (1) we confirmed that the preterminator state of a codon has no detectable effect on codon bias. (2) The well-known effect of CG suppression was visible. We also found that (3) some codons which are very rare, are equal to parts of known transcription signals. Thus, we advanced that to avoid signal interference, the use of these codons is suppressed when a synonymous codon is available. In addition we found that in the whole series of codons, transcription signals are less frequent than in a random sequence of equal composition. Finally we observed (4) that tryptophan is absent in histones. This absence was related not to the TGG codon itself, but to characteristics of the amino acid. We conclude that the functional constraints of a protein can influence, at least for synonymous codon usage, the evolution of its own coding sequence.

Animals

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena

Constraints on codon context in Escherichia coli genes. Their possible role in modulating the efficiency of translation.

The constraints on nucleotide sequences of highly and weakly expressed genes from Escherichia coli have been analysed and compared. Differences in synonymous codon spectra in highly and weakly expressed genes lead to different frequencies of nucleotides (in the first and third codon positions) and dinucleotides in the two groups of genes. It has been found that the choice of synonymous codons in highly expressed genes depends on the nucleotides adjacent to the codon. For example, lysine is preferably encoded by the AAA codon if guanosine is 3' to the lysine codon (AAA-G, P less than 10(-9)). And, on the contrary, AAG is used more often than AAA (P less than 0.001) if cytidine is 3' adjacent to lysine. Guanosine occurs more frequently than adenosine 5' to all the lysine codons (AAR, P less than 10(-5), i.e. NNG codons are preferred over the synonymous NNA codons 5' to the positions of lysine in the genes. The context effect was observed in nonsense and missense suppression experiments. Therefore, a hypothesis has been suggested that the efficiency of translation of some codons (for which the constraints on the adjacent nucleotides were found) can be modulated by the codon context. The rules for preferable synonymous codon choice in highly expressed genes depending on the nucleotides surrounding the codon are presented. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Amino Acid Sequence

Codon usage in muscle genes and liver genes.

Synonymous codon usage frequencies, derived from cDNA clone sequences, were compared for several sets of vertebrate genes. Gene sets as diverse as those expressed in avian skeletal muscle and in mammalian liver showed similar patterns of synonymous codon usage. There were no significant differences suggesting tissue-specific co-adaptation of codon usage patterns and tRNA anticodon profiles. The results indicate a consensus codon usage pattern for vertebrate genes which is largely independent of taxonomic class, tissue of expression, and the cellular fate and rate of evolution of the encoded proteins. Certain elements of the consensus codon usage pattern indicate that it is the product of natural selection and not simply a mutational equilibrium among phenotypically equivalent synonyms.

Animals

Selection on silent sites in the rodent H3 histone gene family.

Selection promoting differential use of synonymous codons has been shown for several unicellular organisms and for Drosophila, but not for mammals. Selection coefficients operating on synonymous codons are likely to be extremely small, so that a very large effective population size is required for selection to overcome the effects of drift. In mammals, codon-usage bias is believed to be determined exclusively by mutation pressure, with differences between genes due to large-scale variation in base composition around the genome. The replication-dependent histone genes are expressed at extremely high levels during periods of DNA synthesis, and thus are among the most likely mammalian genes to be affected by selection on synonymous codon usage. We suggest that the extremely biased pattern of codon usage in the H3 genes is determined in part by selection. Silent site G + C content is much higher than expected based on flanking sequence G + C content, compared to other rodent genes with similar silent site base composition but lower levels of expression. Dinucleotide-mediated mutation bias does affect codon usage, but the affect is limited to the choice between G and C in some fourfold degenerate codons. Gene conversion between the two clusters of histone genes has not been an important force in the evolution of the H3 genes, but gene conversion appears to have had some effect within the cluster on chromosome 13.

Animals

Codon usage and tRNA content in unicellular and multicellular organisms.

Choices of synonymous codons in unicellular organisms are here reviewed, and differences in synonymous codon usages between Escherichia coli and the yeast Saccharomyces cerevisiae are attributed to differences in the actual populations of isoaccepting tRNAs. There exists a strong positive correlation between codon usage and tRNA content in both organisms, and the extent of this correlation relates to the protein production levels of individual genes. Codon-choice patterns are believed to have been well conserved during the course of evolution. Examination of silent substitutions and tRNA populations in Enterobacteriaceae revealed that the evolutionary constraint imposed by tRNA content on codon usage decelerated rather than accelerated the silent-substitution rate, at least insofar as pairs of taxonomically related organisms were examined. Codon-choice patterns of multicellular organisms are briefly reviewed, and diversity in G+C percentage at the third position of codons in vertebrate genes--as well as a possible causative factor in the production of this diversity--is discussed.

Animals

Synonymous substitution-rate constants in Escherichia coli and Salmonella typhimurium and their relationship to gene expression and selection pressure.

Based on the differences in synonymous codon use between E. coli and S. typhimurium, the synonymous substitution rates can be estimated. In contrast to previous studies on the substitution rates in these two organisms, we use a kinetic model that explicitly takes the selection bias into account. The selection pressure on synonymous codons for a particular amino acid can be calculated from the observed codon bias. This offers a unique opportunity to study systematically the relationship between substitution-rate constants and selection pressure. The results indicate that the codon bias in these organisms is determined by a mutation-selection balance rather than by stabilizing selection. A best fit to the data implies that the mutation rate constant increases about threefold in genes at low expression levels relative to those that are highly expressed.

Base Sequence

[Regularities of the nucleotide sequence at the 5'-end of the codon in Escherichia coli genes].

The frequencies of occurrence of nucleotides at the 5' side of codons have been determined in highly and weakly expressed genes from E. coli. Significant constraints on the nucleotide 5' to some codons were found in highly expressed genes. Certain rules of synonymous codon usage depending on the amino acid 3' of the codon were established. E. g., codon possessing quanosine in the third position (NNG) are preferred over NNA if the next amino acid is lysine (P less than 10(-5)). On the other hand, rules of synonymous codon usage in relation to 5' flanking nucleotide were found. For example, when coding for aspartic acid, GAC codon is preferred over GAU (P less than 0.001) if uridine is 5' to codon and on the contrary GAU is favoured (P less than 0.0001) if quanosine is at the 5' side of aspartic acid codon. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.

Base Sequence

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence

Essential factors determining codon usage in ubiquitin genes.

Ubiquitin is ubiquitous in all eukaryotes and its amino acid sequence shows extreme conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences coding for 13 ubiquitin genes from 11 species reported so far have been compiled and analyzed. The G + C content of codon third base reveals a positive linear correlation with the genome G + C content of the corresponding species. The slope strongly suggests that the overall G + C content of codons of polyubiquitin genes clearly reflects the genome G + C content by AT/GC substitutions at the codon third position. The G + C content of ubiquitin codon third base also shows a positive linear correlation with the overall G + C content of coding regions of compiled genes, indicating the codon choices among synonymous codons reflect the average codon usage pattern of corresponding species. On the other hand, the monoubiquitin gene, which is different from the polyubiquitin gene in gene organization, gene expression, and function of the encoding protein, shows a different codon usage pattern compared with that of the polyubiquitin gene. From comparisons of the levels of synonymous substitutions among ubiquitin repeats and the homology of the amino acid sequence of the tail of monomeric ubiquitin genes, we propose that the molecular evolution of ubiquitin genes occurred as follows: Plural primitive ubiquitin sequences were dispersed on genome in ancestral eukaryotes. Some of them situated in a particular environment fused with the tail sequence to produce monomeric ubiquitin genes that were maintained across species. After divergence of species, polyubiquitin genes were formed by duplication of the other primitive ubiquitin sequences on different chromosomes. Differences in the environments in which ubiquitin genes are embedded reflect the differences in codon choice and in gene expression pattern between poly- and monomeric ubiquitin genes.

Amino Acid Sequence

Codon usage in yeast: cluster analysis clearly differentiates highly and lowly expressed genes.

Codon usage data has been compiled for 110 yeast genes. Cluster analysis on relative synonymous codon usage revealed two distinct groups of genes. One group corresponds to highly expressed genes, and has much more extreme synonymous codon preference. The pattern of codon usage observed is consistent with that expected if a need to match abundant tRNAs, and intermediacy of tRNA-mRNA interaction energies are important selective constraints. Thus codon usage in the highly expressed group shows a higher correlation with tRNA abundance, a greater degree of third base pyrimidine bias, and a lesser tendency to the A+T richness which is characteristic of the yeast genome. The cluster analysis can be used to predict the likely level of gene expression of any gene, and identifies the pattern of codon usage likely to yield optimal gene expression in yeast.

Base Composition

Codon usage in Pseudomonas aeruginosa.

We have generated a codon usage table for Pseudomonas aeruginosa. Codon usage in P. aeruginosa is extremely biased. In contrast to E. coli and yeast, P. aeruginosa preferentially uses those codons within a synonymous codon group with the strongest predicted codon-anticodon interaction. We were unable to correlate a particular codon usage pattern with predicted levels of mRNA expressivity. The choice of a third base reflects the high guanine plus cytosine content of the P. aeruginosa genome (67.2%) and cytosine is the preferred nucleotide for the third codon position.

Bacteriophages

Rapidly evolving mouse alpha-globin-related pseudo gene and its evolutionary history.

The nucleotide sequences of a mouse pseudo alpha-globin gene and two adult alpha-globin genes from mouse and rabbit were compared. A close examination of sequence differences among the three genes revealed that the mouse pseudo alpha-globin gene was derived from one of the mouse alpha-globin genes about 24 million years ago by gene duplication and lost its original function, presumably due to loss of both the intervening sequences or frameshift mutations that prevented production of a functional globin polypeptide, and eventually became an inactive gene about 17 million years ago. After this event, this gene evolved at a very high rate, approximately 1.9 times the rate of synonymous codon change in productive genes. This suggests the presence of functional constraints against synonymous codon changes in normally functioning genes and also suggests that the "nucleotide arrangements" serve almost no important functions beside protein coding ability in the greater part of this gene, except for a very limited number of nucleotides. A continuous stretch of the pseudo alpha-globin gene consisting of a third exon and a 5' half of the 3' noncoding region shows marked sequence homology to the alpha-globin gene, suggesting transfer from one of the globin or globin-like genes by recombination very recently.

Animals

Codon usage patterns in Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster and Homo sapiens; a review of the considerable within-species diversity.

The genetic code is degenerate, but alternative synonymous codons are generally not used with equal frequency. Since the pioneering work of Grantham's group it has been apparent that genes from one species often share similarities in codon frequency; under the "genome hypothesis" there is a species-specific pattern to codon usage. However, it has become clear that in most species there are also considerable differences among genes. Multivariate analyses have revealed that in each species so far examined there is a single major trend in codon usage among genes, usually from highly biased to more nearly even usage of synonymous codons. Thus, to represent the codon usage pattern of an organism it is not sufficient to sum over all genes as this conceals the underlying heterogeneity. Rather, it is necessary to describe the trend among genes seen in that species. We illustrate these trends for six species where codon usage has been examined in detail, by presenting the pooled codon usage for the 10% of genes at either end of the major trend. Closely-related organisms have similar patterns of codon usage, and so the six species in Table 1 are representative of wider groups. For example, with respect to codon usage, Salmonella typhimurium closely resembles E. coli, while all mammalian species so far examined (principally mouse, rat and cow) largely resemble humans.

Amino Acids