Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals

Stable structure of thermophilic proton ATPase beta subunit.

F1-ATPase is the major enzyme for ATP synthesis in mitochondria, chloroplasts, and bacterial plasma membranes. F1-ATPase obtained from thermophilic bacterium PS3 (TF1) is the only ATPase which can be reconstituted from its primary structure. Its beta subunit constitutes the catalytic site, and is capable of forming hybrid F1's with E. coli alpha and gamma subunits. Since the stability of TF1 resides in its primary structure, we cloned a gene coding for TF1, and the primary structure of the beta subunit was deduced from the nucleotide sequence of the gene to compare the sequence with those of beta's of three major categories of F1's; prokaryotic membranes, chloroplasts, and mitochondria. The following results were obtained. Homology: The primary structure of the TF1 beta subunit (473 residues, Mr = 51,995.6) showed 89.3% homology with 270 residues which are identical in the beta subunits from human mitochondria, spinach chloroplasts, and E. coli. It contained regions homologous to several nucleotide-binding proteins. Secondary structure: The deduced alpha-helical (30.1%) and beta-sheet (22.3%) contents were consistent with those determined from the circular dichroism spectra. Residues forming reverse turns (Gly and Pro) were highly conserved among the F1 beta subunits. Substituted residues and stability of TF1: We compared the amino acid sequence of the TF1 beta subunit with those of the other F1 beta subunits mentioned above. The observed substitutions in the thermophilic subunit increased its propensities to form secondary structures, and its external polarity to form tertiary structure. Codon usage: The codon usage of the TF1 beta gene was found to be unique. The changes in codons that achieved these amino acid substitutions were much larger than those caused by minimal mutations, and the third letters of the optimal codons were either guanine or cytosine, except in codons for Gln, Lys, and Glu.

Amino Acid Sequence

The translational termination signal database (TransTerm) now also includes initiation contexts.

The TransTerm database of termination codon contexts has been extended to include sense codon usage, and initiation codon contexts. The database was constructed from 23,721 coding sequences from 93 organisms. The database contains: a) the sequence around the termination codon (-10, +10); b) the sequence around the initiation codon (-20, +10); c) the length, 'G+C%' of the third position of codons (GC3), the 'codon adaptation index' (CAI) and the 'effective number of codons' statistic (Nc); d) summary tables for each organism including total codon usage, stop codon and tetranucleotide stop-signal usage, and matrices tallying base frequencies at each position around the initiation and termination codons. The data are arranged to facilitate investigation of the relationships between the three phases of protein synthesis. The database is available electronically from EMBL.

Animals

Intrastrand parity rules of DNA base composition and usage biases of synonymous codons.

When there are no biases in mutation and selection between the two strands of DNA, the 12 possible substitution rates of the four nucleotides reduces to six (type 1 parity rule or PR1), and the intrastrand average base composition is expected to be A = T and G = C at equilibrium without regard to the G + C content of DNA (type 2 parity rule or PR2). Significant deviations from the parity rules in the third codon letters of the four-codon amino acids result mostly from selective biases rather than mutational biases between the two strands of DNA during evolution. The parity rules lay the foundation for evaluating the biases in synonymous codon usage in terms of (1) directional mutation pressure for variation of the DNA G + C content due to mutational biases between alpha-bases (A or T) and gamma-bases (G or C), (2) strand-bias mutation, for example, by DNA repair during transcription, and (3) functional selection in evolution, for example, due to tRNA abundance. The present analysis shows that, although the PR2 violation is common in the third codon letters of four-codon amino acids, the contribution of PR2 violation to the DNA G + C content of the third codon position is small and, in majority of cases, mildly counteracts the effect of the directional mutation pressure on the G + C content.

Animals

Consecutive low-usage leucine codons block translation only when near the 5' end of a message in Escherichia coli.

Insertion of nine consecutive low-usage CUA leucine codons after codon 13 of a 313-codon test mRNA strongly inhibited its translation without apparent effect on translation of other mRNAs containing CUA codons. In contrast, nine consecutive high-usage CUG leucine codons at the same position had no apparent effect, and neither low- nor high-usage codons affected translation when inserted after codon 223 or 307. Additional experiments indicated that the strong positional effect of the low-usage codons could not be accounted for by differences in stability of the mRNAs or in stringency of selection of the correct tRNA. The positional effect could be explained if translation complexes are less stable near the beginning of a message: slow translation through low-usage codons early in the message may allow most translation complexes to dissociate before they read through.

Blotting, Northern

Unusual codon bias occurring within insertion sequences in Escherichia coli.

The large open reading frames of insertion sequences from Escherichia coli were examined for their spatial pattern of codon usage bias and distribution of rarely used codons. There is a bias in codon usage that is generally lower toward the terminal ends of the coding regions, which is reflected in the occurrence of an excess of nonpreferred codons in the 3' portions of the coding regions as compared with the 5' portions. In contrast, typical chromosomal genes have a lower codon usage bias toward the 5' ends of the coding regions. These results imply that the selective forces reflected in codon usage bias may differ according to position within the coding sequence. In addition, these constraints apparently differ in important ways between genes contained in insertion sequences and those in the chromosome.

Chromosomes, Bacterial

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus

Usage of the three termination codons: compilation and analysis of the known eukaryotic and prokaryotic translation termination sequences.

The published translation termination sequences have been compiled and analysed to aid the interpretation of experiments on termination codon usage in the Xenopus oocyte (Bienz et al. 1981). There are significant differences between prokaryotes and eukaryotes concerning the usage of the three termination codons and of tandem stops. In addition viruses show termination strategies that differ from those of their hosts. Preferred context sequences flanking termination codons are described. Contexts vary within the last codon according to the nature of the termination codon, but are uniform within the first triplet following the terminators.

Animals

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals

On the rate of DNA sequence evolution in Drosophila.

Analysis of the rate of nucleotide substitution at silent sites in Drosophila genes reveals three main points. First, the silent rate varies (by a factor of two) among nuclear genes; it is inversely related to the degree of codon usage bias, and so selection among synonymous codons appears to constrain the rate of silent substitution in some genes. Second, mitochondrial genes may have evolved only as fast as nuclear genes with weak codon usage bias (and two times faster than nuclear genes with high codon usage bias); this is quite different from the situation in mammals where mitochondrial genes evolve approximately 5-10 times faster than nuclear genes. Third, the absolute rate of substitution at silent sites in nuclear genes in Drosophila is about three times higher than the average silent rate in mammals.

Animals

Beta tubulin gene of the parasitic protozoan Leishmania mexicana.

A genomic DNA library was generated with Sau3A cut DNA derived from promastigotes of Leishmania mexicana amazonensis and the lambda vector EMBL3. The library was screened for beta tubulin clones using 32P-labeled heterologous probe of chicken beta tubulin cDNA. From the various genomic clones the one designated 23.1, which gave the simplest hybridization banding pattern, was further characterized. The leishmanial insert DNA was subcloned into plasmid vectors and the resulting clones were designated as T11, T28 and T50. Using these clones leishmanial beta tubulin coding region was sequenced by the dideoxy method. The result shows that the beta tubulin has 445 amino acids, a carboxyl terminal tyrosine, and no intron. Leishmanial beta tubulin has 93% amino acid sequence similarity with that of trypanosome and 82% with that of man: and there is a strong bias in codon usage for codons possessing guanine or cytosine in the third base.

Amino Acid Sequence

DNA sequences from the str operon of Escherichia coli.

The str operon at 72 min on the Escherichia coli chromosome contains genes for ribosomal proteins (r-proteins) S12 (str or rpsL) and S7 (rpsG) and elongation factors G (fus) and Tu (tufA). The sequence of the entire S12 gene, the S12-S7 intercistronic region, and the beginning of the S7 gene is reported. Also, the sequence of the end of the S7 gene, the S7-G intercistronic region, and the beginning of the elongation fractor G gene is reported. The S12-S7 intercistronic region is 96 base pairs long, in contrast to other intercistronic regions in r-protein operons which have been found to vary from 3 to 66 base pairs. The S7-G intercistronic region is only 27 bases long, supporting the previous conclusion that r-protein and elongation factor genes are co-transcribed. A comparison of translation initiation sites of the S12 and S7 genes, and other examples of co-transcribed r-protein genes, reveals no obvious features that could account for equimolar synthesis of all r-proteins. The codon usage in the S12 and S7 genes follows the pattern observed in other r-protein genes; that is, there is a highly preferential usage of codons recognized by the most abundant of isoaccepting tRNA species. This pattern could reflect the cell's need for efficient translation or minimal errors, or both, in r-protein synthesis.

Amino Acid Sequence

The primary structure of the alcohol dehydrogenase gene from the fission yeast Schizosaccharomyces pombe.

We have cloned and sequenced the alcohol dehydrogenase gene of the fission yeast Schizosaccharomyces pombe. The gene was isolated by transformation and complementation of a Saccharomyces cerevisiae strain which lacked functional alcohol dehydrogenase with an S. pombe gene bank constructed in the autonomously replicating yeast plasmid YEp13. Southern hybridization analysis indicates that S. pombe contains only one alcohol dehydrogenase gene. The structural region of the gene is 50% homologous to the alcohol dehydrogenase encoding genes of the budding yeast S. cerevisiae. The gene exhibits a very strong codon usage bias; with the set of predominantly used codons generally resembling that which S. cerevisiae employs preferentially. All of the differences in codon usage bias between S. pombe and S. cerevisiae are in the direction of greater G + C content in S. pombe codons. It is argued that this observation supports the hypothesis that selection toward uniform codon-anticodon binding energies contributes to codon usage bias and that the optimum binding energy is, on the average, higher in S. pombe than S. cerevisiae.

Alcohol Dehydrogenase

Complementary DNA and amino acid sequence of rat liver microsomal, xenobiotic epoxide hydrolase.

The coding nucleotide sequence for rat liver microsomal, xenobiotic epoxide hydrolase was determined from two overlapping cDNA clones, which together contain 1750 nucleotides complementary to epoxide hydrolase mRNA. The single open reading frame of 1365 nucleotides codes for a 455 amino acid polypeptide with a molecular weight of 52,581. The deduced amino acid composition agrees well with those determined by direct amino acid analysis of the rat protein, and the amino acid sequence is 81% identical to that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and that of rabbit epoxide hydrolase. Analysis of codon usage for epoxide hydrolase, and comparison to codon usage for NADPH-cytochrome P-450 oxidoreductase and cytochromes P-450b, P-450d, and P-450PCN, suggest that epoxide hydrolase is more conserved than cytochromes P-450b and P-450PCN; comparison of the extent of sequence conservation for 12 homologous proteins between the rat and rabbit, including cytochrome P-450b, supports this hypothesis, and indicates that much of epoxide hydrolase is constrained to maintain its hydrophobic character, consistent with its intramembranous location. The predicted membrane topology of epoxide hydrolase delineates 6 membrane-spanning segments, less than the 8 or 10 predicted for two cytochrome P-450 isozymes; the lower number of membrane-spanning segments predicted for epoxide hydrolase correlates with its lesser dependence on the membrane for maintenance of its tertiary structure and catalytic activity.

Amino Acid Sequence

Codon bias in actin multigene families and effects on the reconstruction of phylogenetic relationships.

Codon usage patterns and phylogenetic relationships in the actin multigene family have been analyzed for three dipteran species--Drosophila melanogaster, Bactrocera dorsalis, and Ceratitis capitata. In certain phylogenetic tree reconstructions, using synonymous distances, some gene relationships are altered due to a homogenization phenomenon. We present evidence to show that this homogenization phenomenon is due to codon usage bias. A survey of the pattern of synonymous codon preferences for 11 actin genes from these three species reveals that five out of the six Drosophila actin genes show high degrees of codon bias as indicated by scaled chi 2 values. In contrast to this, four out of the five actin genes from the other species have low codon bias values. A Monte Carlo contingency test indicates that for those Drosophila actin genes which exhibit codon bias, the patterns of codon usage are different compared to actin genes from the other species. In addition, the genes exhibiting codon bias also appear to have reduced rates of synonymous substitution. The homogenization phenomenon seen in terms of synonymous substitutions is not observed for nonsynonymous changes. Because of this homogenization phenomenon, "trees" constructed based on synonymous substitutions will be affected. These effects can be overt in the case of multigene families, but similar distortions may underlie reconstructions based on single-copy genes which exhibit codon usage bias.

Actins

Selection on silent sites in the rodent H3 histone gene family.

Selection promoting differential use of synonymous codons has been shown for several unicellular organisms and for Drosophila, but not for mammals. Selection coefficients operating on synonymous codons are likely to be extremely small, so that a very large effective population size is required for selection to overcome the effects of drift. In mammals, codon-usage bias is believed to be determined exclusively by mutation pressure, with differences between genes due to large-scale variation in base composition around the genome. The replication-dependent histone genes are expressed at extremely high levels during periods of DNA synthesis, and thus are among the most likely mammalian genes to be affected by selection on synonymous codon usage. We suggest that the extremely biased pattern of codon usage in the H3 genes is determined in part by selection. Silent site G + C content is much higher than expected based on flanking sequence G + C content, compared to other rodent genes with similar silent site base composition but lower levels of expression. Dinucleotide-mediated mutation bias does affect codon usage, but the affect is limited to the choice between G and C in some fourfold degenerate codons. Gene conversion between the two clusters of histone genes has not been an important force in the evolution of the H3 genes, but gene conversion appears to have had some effect within the cluster on chromosome 13.

Animals

Influence of the codon following the initiation codon on the expression of the lacZ gene in Saccharomyces cerevisiae.

A set of 32 different codons were introduced in a lacZ expression vector (pPTK400) immediately 3' from the AUG initiation codon. Expression of the lacZ gene was determined in Saccharomyces cerevisiae by measuring the amount of beta-galactosidase fusion protein using immuno-gel electrophoresis. A 5.3-fold difference in expression was found among the various constructs. It was found that there was no preference for a certain nucleotide in any position of the second codon and there was no distinct correlation between the level of tRNA corresponding to any particular second codon and expression. No correlation could be found between the local secondary structure and expression. When the overall codon usage in yeast and the codon usage in the second position of the mRNA is compared, there is no obvious significant difference in preference. This indicates that in yeast, in contrast to Escherichia coli, the codon choice at the beginning of the mRNA does not deviate from the one further downstream and is determined by the requirements for optimal translation elongation. Important determinants of the optimal context for an initiation codon in yeast therefore must be located mainly 5' from this codon.

Amino Acid Sequence

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena