Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Codon usage, genetic code and phylogeny of Dictyostelium discoideum mitochondrial DNA as deduced from a 7.3-kb region.

We have sequenced a region (7,376-bp) of the mitochondrial (mt) DNA (54 kb) of the cellular slime mold, Dictyostelium discoideum. From the DNA and amino-acid sequence comparisons with known sequences, genes for ATPase subunit 9 (ATP9), cytochrome b (CYTB), NADH dehydrogenase subunits 1, 3 and 6 (ND1, ND3 and ND6), small subunit rRNA (SSU rRNA) and seven tRNAs (Arg, Asn, Cys, Lys, f-Met, Met and Pro) have been identified. The sequenced region of the mtDNA has a high average A + T-content (70.8%). The A + T-content of protein-genes (73.6%) is considerably higher than that of RNA genes (61.3%). Even with the strong AT-bias, the genetic code employed is most probably the universal one. All seven tRNAs are able to form typical clover leaf structures. The molecular phylogenetic trees of CYTB and SSU rRNA suggest that D. discoideum is closer to green plants than to animals and fungi.

Amino Acid Sequence↗

Models of nearly neutral mutations with particular implications for nonrandom usage of synonymous codons.

The population dynamics of nearly neutral mutations are studied using a single-site and a multisite model. In the latter model, the nucleotides in a sequence are completely linked and the selection schemes employed are additive, multiplicative, and additive with a threshold. Although the third selection scheme is very different from the first two, the three schemes produce identical results for a wide range of parameter values. Thus the present study provides a general theory for the population dynamics of nearly neutral mutations because the results can also be used to draw inferences about other selection schemes such as stabilizing selection and synergistic selection. It is shown that the number of slightly deleterious mutations accumulated in a sequence can be considerably larger under the multisite model than under the single-site model, particularly if the sequence is long or if the mutation rate per site is high. The results show that even a very slight selective difference between synonymous codons can produce a strong bias in codon usage. Three alternative explanations for the strong bias in codon usage in bacterial and yeast genes are considered. The implications of the present results for molecular evolution are discussed.

Biological Evolution↗

In vivo evidence for non-universal usage of the codon CUG in Candida maltosa.

An alkane-assimilating yeast Candida maltosa had been studied in order to establish systems suitable for biotransformation of hydrophobic compounds. However, functional expression of heterologous genes tested for this purpose had not been successful in several cases. On the other hand, it had been reported that the codon CUG, a universal leucine codon, is read as serine in C. cylindracea. The same altered codon usage had also been suggested by in vitro experiments in some Candida yeasts which are phylogenetically closely related to C. maltosa. In this study we have shown that the failure in functional expression of a heterologous gene is due to the fact that the codon CUG is read as serine in C. maltosa. This conclusion was drawn from the following experimental results: (1) when a cytochrome P450 gene of C. maltosa containing a CTG codon was expressed in C. maltosa, the corresponding amino acid was found to be serine, and not leucine; (2) a tRNA gene with an almost identical structure to that of the tRNASerCAG gene of C. albicans could be isolated from the genome of C. maltosa; (3) the Saccharomyces cerevisiae URA3 gene, which has one CTG codon, could not complement the ura3 mutation of C. maltosa as itself, but when the CTG codon was changed to another leucine codon, CTC, the mutated gene could complement the ura3 mutation. The last result is the first example of succeeding in functional expression of a heterologous gene in Candida species having an altered codon usage by changing the CTG codon in the gene to another codon.

Amino Acid Sequence↗

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals↗

Stable structure of thermophilic proton ATPase beta subunit.

F1-ATPase is the major enzyme for ATP synthesis in mitochondria, chloroplasts, and bacterial plasma membranes. F1-ATPase obtained from thermophilic bacterium PS3 (TF1) is the only ATPase which can be reconstituted from its primary structure. Its beta subunit constitutes the catalytic site, and is capable of forming hybrid F1's with E. coli alpha and gamma subunits. Since the stability of TF1 resides in its primary structure, we cloned a gene coding for TF1, and the primary structure of the beta subunit was deduced from the nucleotide sequence of the gene to compare the sequence with those of beta's of three major categories of F1's; prokaryotic membranes, chloroplasts, and mitochondria. The following results were obtained. Homology: The primary structure of the TF1 beta subunit (473 residues, Mr = 51,995.6) showed 89.3% homology with 270 residues which are identical in the beta subunits from human mitochondria, spinach chloroplasts, and E. coli. It contained regions homologous to several nucleotide-binding proteins. Secondary structure: The deduced alpha-helical (30.1%) and beta-sheet (22.3%) contents were consistent with those determined from the circular dichroism spectra. Residues forming reverse turns (Gly and Pro) were highly conserved among the F1 beta subunits. Substituted residues and stability of TF1: We compared the amino acid sequence of the TF1 beta subunit with those of the other F1 beta subunits mentioned above. The observed substitutions in the thermophilic subunit increased its propensities to form secondary structures, and its external polarity to form tertiary structure. Codon usage: The codon usage of the TF1 beta gene was found to be unique. The changes in codons that achieved these amino acid substitutions were much larger than those caused by minimal mutations, and the third letters of the optimal codons were either guanine or cytosine, except in codons for Gln, Lys, and Glu.

Amino Acid Sequence↗

The translational termination signal database (TransTerm) now also includes initiation contexts.

The TransTerm database of termination codon contexts has been extended to include sense codon usage, and initiation codon contexts. The database was constructed from 23,721 coding sequences from 93 organisms. The database contains: a) the sequence around the termination codon (-10, +10); b) the sequence around the initiation codon (-20, +10); c) the length, 'G+C%' of the third position of codons (GC3), the 'codon adaptation index' (CAI) and the 'effective number of codons' statistic (Nc); d) summary tables for each organism including total codon usage, stop codon and tetranucleotide stop-signal usage, and matrices tallying base frequencies at each position around the initiation and termination codons. The data are arranged to facilitate investigation of the relationships between the three phases of protein synthesis. The database is available electronically from EMBL.

Animals↗

Cervical lesions are associated with human papillomavirus type 16 intratypic variants that have high transcriptional activity and increased usage of common mammalian codons.

Human papillomavirus type 16 (HPV-16) is a major cause of cervical neoplasia, but only a minority of HPV-16 infections result in cancer. Whether particular HPV-16 variants are associated with cervical disease has not yet been clearly established. An investigation of whether cervical neoplasia is associated with infection with HPV-16 intratypic variants was undertaken by using RFLP analyses in a study of 100 HPV-16 DNA-positive women with or without neoplasia. RFLP variant 2 was positively associated [odds ratio (OR)=2.57] and variant 5 was negatively associated with disease (OR=0.2). Variant 1, which resembles the reference isolate of HPV-16, was found at a similar prevalence among those with and without neoplasia. Variants 1 and 2 were also more likely to be associated with detectable viral mRNA than variant 5 (respectively P=0.03 and P=0.00). When HPV-16 E5 ORFs in 50 clones from 36 clinical samples were sequenced, 19 variant HPV-16 E5 DNA sequences were identified. Twelve of these DNA sequences encoded variant E5 amino acid sequences, 10 of which were novel. Whilst the associations between HPV-16 E5 RFLP variants and neoplasia could not be attributed to differences in amino acid sequences, correlation was observed in codon usage. DNA sequences of RFLP variant 2 (associated with greatest OR for neoplasia) had a significantly greater usage of common mammalian codons compared with RFLP pattern 1 variants.

Amino Acid Sequence↗

Mutation pressure, natural selection, and the evolution of base composition in Drosophila.

Genome sequencing in a number of taxa has revealed variation in nucleotide composition both among regions of the genome and among functional classes of sites in DNA. Mutational biases, biased gene conversion, and natural selection have been proposed as causes of this variation. Here, we review patterns of base composition in Drosophila DNA. Nucleotide composition in Drosophila melanogaster varys regionally, and base composition is correlated between introns and exons. Drosophila species also show striking patterns of non-random codon usage. Patterns of synonymous codon usage and the biochemistry of translation suggest that natural selection may act at 'silent' sites. A relationship between recombination rates and codon usage and comparisons of the evolutionary dynamics of silent mutations within and between species support natural selection discriminating among synonymous codons. The causes of regional base composition variation are less clear. Progress in functional studies of non-coding DNA, further investigations of genome patterns, and statistical tests based on evolutionary theory will lead to a greater understanding of the contributions of mutational processes and natural selection in patterning genome-wide nucleotide composition.

Animals↗

Intrastrand parity rules of DNA base composition and usage biases of synonymous codons.

When there are no biases in mutation and selection between the two strands of DNA, the 12 possible substitution rates of the four nucleotides reduces to six (type 1 parity rule or PR1), and the intrastrand average base composition is expected to be A = T and G = C at equilibrium without regard to the G + C content of DNA (type 2 parity rule or PR2). Significant deviations from the parity rules in the third codon letters of the four-codon amino acids result mostly from selective biases rather than mutational biases between the two strands of DNA during evolution. The parity rules lay the foundation for evaluating the biases in synonymous codon usage in terms of (1) directional mutation pressure for variation of the DNA G + C content due to mutational biases between alpha-bases (A or T) and gamma-bases (G or C), (2) strand-bias mutation, for example, by DNA repair during transcription, and (3) functional selection in evolution, for example, due to tRNA abundance. The present analysis shows that, although the PR2 violation is common in the third codon letters of four-codon amino acids, the contribution of PR2 violation to the DNA G + C content of the third codon position is small and, in majority of cases, mildly counteracts the effect of the directional mutation pressure on the G + C content.

Animals↗

Consecutive low-usage leucine codons block translation only when near the 5' end of a message in Escherichia coli.

Insertion of nine consecutive low-usage CUA leucine codons after codon 13 of a 313-codon test mRNA strongly inhibited its translation without apparent effect on translation of other mRNAs containing CUA codons. In contrast, nine consecutive high-usage CUG leucine codons at the same position had no apparent effect, and neither low- nor high-usage codons affected translation when inserted after codon 223 or 307. Additional experiments indicated that the strong positional effect of the low-usage codons could not be accounted for by differences in stability of the mRNAs or in stringency of selection of the correct tRNA. The positional effect could be explained if translation complexes are less stable near the beginning of a message: slow translation through low-usage codons early in the message may allow most translation complexes to dissociate before they read through.

Blotting, Northern↗

Protein-encoding genes in the sulfothermophilic archaea Sulfolobus and Pyrococcus.

A number of unrelated protein-encoding genes from sulfothermophilic archaea, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Pyrococcus furiosus and Pyrococcus woesei, has been analyzed. In the Sulfolobus genus, the content of A + T is significantly higher than that of C + G and the base usage follows the order, A > T > G > C. In Pyrococcus, the A + T content is also higher than that of C + G, but with lower values; in the order of base usage, G precedes T. The codon usage of these sulfothermophiles has been determined; alternative start codons are frequently used in both genera; codon preferences reflect the rich A + T composition of the corresponding genomes; for both genera the codon bias is particularly evident within the different arginine triplets, where AGA and AGG are predominant. From the similarities in the codon usage, close taxonomic relationships become evident within the Sulfolobus or the Pyrococcus genus; a lower, but significant similarity is also clear between these genera. The synonymous codon usage of these sulfothermophiles shows similarities with that of Saccharomyces cerevisiae and bovine mitochondria, whereas clear divergences are observed with the halophilic archaeal genus, Halobacterium, or the eubacterium, Escherichia coli. The unrelated proteins of the considered sulfothermophiles have been analyzed for the content of hydrophobic residues; the comparison with mesophiles reveals a significant increase in the average hydrophobicity of amino acid residues. This finding could indicate a mechanism of adaptation of proteins in organisms living under extreme environments. It is noteworthy that an opposite trend, i.e. a decreased average hydrophobicity, occurs in unrelated halophilic proteins.

Animals↗

Unusual codon bias occurring within insertion sequences in Escherichia coli.

The large open reading frames of insertion sequences from Escherichia coli were examined for their spatial pattern of codon usage bias and distribution of rarely used codons. There is a bias in codon usage that is generally lower toward the terminal ends of the coding regions, which is reflected in the occurrence of an excess of nonpreferred codons in the 3' portions of the coding regions as compared with the 5' portions. In contrast, typical chromosomal genes have a lower codon usage bias toward the 5' ends of the coding regions. These results imply that the selective forces reflected in codon usage bias may differ according to position within the coding sequence. In addition, these constraints apparently differ in important ways between genes contained in insertion sequences and those in the chromosome.

Chromosomes, Bacterial↗

What drives codon choices in human genes?

Synonymous codon usage is based and the bias seems to be different in different organisms. Factors with proposed roles in causing codon bias include degree and timing of gene expression, codon-anticodon interactions, transcription and translation rate and fidelity, codon context, and global and local G + C content. We offer a new perspective and new methods for elucidating codon choices applied especially to the human genome. We present data supporting the thesis that codon choices for human genes are largely a consequence of two factors: (1) amino acid constraints, (2) maintaining DNA structures dependent on base-step conformational tendencies consistent with the organism's genome signature that is determined by genome-wide processes of DNA modification, replication and repair. The related codon signature defined as the dinucleotide relative abundances at the distinct codon positions (1,2), (2,3), and (3,4) (4 = 1 of the next codon) accommodates both the global genome signature and amino acid constraints. In human genes, codon positions (2,3) and (3,4) containing the silent site have similar codon signatures reflecting DNA symmetry. Strong CG and TA dinucleotide underrepresentation is observed at all codon positions as well as in non-coding regions. Estimates of synonymous codon usage based on codon signatures are in excellent agreement with the actual codon usage in human and general vertebrate genes. These properties are largely independent of the isochore compartment (G + C content), gene size, and transcriptional and translational constraints. We hypothesize that major influences on codon usage in human genes result from residue preferences and diresidue associations in proteins coupled to biases on the DNA level, related to replication and repair processes and/or DNA structural requirements.

Codon↗

DNA sequence evolution: the sounds of silence.

Silent sites (positions that can undergo synonymous substitutions) in protein-coding genes can illuminate two evolutionary processes. First, despite being silent, they may be subject to natural selection. Among eukaryotes this is exemplified by yeast, where synonymous codon usage patterns are shaped by selection for particular codons that are more efficiently and/or accurately translated by the most abundant tRNAs; codon usage across the genome, and the abundance of different tRNA species, are highly co-adapted. Second, in the absence of selection, silent sites reveal underlying mutational patterns. Codon usage varies enormously among human genes, and yet silent sites do not appear to be influenced by natural selection, suggesting that mutation patterns vary among regions of the genome. At first, the yeast and human genomes were thought to reflect a dichotomy between unicellular and multicellular organisms. However, it now appears that natural selection shapes codon usage in some multicellular species (e.g. Drosophila and Caenorhabditis), and that regional variations in mutation biases occur in yeast. Silent sites (in serine codons) also provide evidence for mutational events changing adjacent nucleotides simultaneously.

Animals↗

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus↗

Horizontal gene transfer contributes to the wide distribution and evolution of type II restriction-modification systems.

Restriction modification (RM) systems serve to protect bacteria against bacteriophages. They comprise a restriction endonuclease activity that specifically cleaves DNA and a corresponding methyltransferase activity that specifically methylates the DNA, thereby protecting it from cleavage. Such systems are very common in bacteria. To find out whether the widespread distribution of RM systems is due to horizontal gene transfer, we have compared the codon usages of 29 type II RM systems with the average codon usage of their respective bacterial hosts. Pronounced deviations in codon usage were found in six cases: EcoRI, EcoRV, KpnI, SinI, SmaI, and TthHB81. They are interpreted as evidence for horizontal gene transfer in these cases. As the methodology is expected to detect only one-fourth to one-third of all horizontal gene transfer events, this result implies that horizontal gene transfer had a considerable influence on the distribution and evolution of RM systems. In all of these six cases the codon usage deviations of the restriction enzyme genes are much more pronounced than those of the methyltransferase genes. This result suggests that in these cases horizontal gene transfer had occurred sequentially with the gene for the methyltransferase being first acquired by the cell. This can be explained by the fact that an active restriction endonuclease is highly toxic in cells whose DNA is not protected from cleavage by a corresponding methyltransferase.

Bacteria↗