Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Efficient synthesis of secreted murine interleukin-2 by Saccharomyces cerevisiae: influence of 3'-untranslated regions and codon usage.

Several expression vectors were compared which directed the synthesis of secreted murine interleukin-2 (mIL2) in the culture medium of Saccharomyces cerevisiae. We used the prepro-sequence of the alpha 1 mating-factor precursor as a secretion signal in S. cerevisiae in combination with different promoters. The yield of mature mIL2 was significantly improved by deleting the major part of the 3'-untranslated region (UTR). In Northern-blotting experiments we showed that a destabilizing sequence present in the 3' UTR might be responsible for rapid degradation of the mIL2 mRNA. The highest expression (about 10 micrograms/ml) was obtained under control of the GAL1 promoter in an S. cerevisiae strain where the regulatory GAL4 gene was overexpressed. No difference in expression level was observed in a construct wherein twelve consecutive codons were replaced by optimal codons for S. cerevisiae.

Animals

Codon usage and mistranslation. In vivo basal level misreading of the MS2 coat protein message.

The coat protein of the small RNA virus MS2 shows charge heterogeneity in vivo. In most strains there is a basic satellite of the native protein. We have shown that this basic satellite is greatly diminished or absent in strains with the streptomycin-resistant allele, rpsL, a mutation which leads to increased translational accuracy. Further, the satellite is present in cells where the coat protein is encoded by duplex DNA. Tryptic digests of the satellite show that it contains new lysine-containing peptides which appear to be the same as those found in derivatives of coat protein which have a lysine for asparagine substitution. Sequencing of the NH2-terminal 19 amino acids of the satellite protein shows that the asparagine codon AAU at amino acid 12 is misread approximately 8 times more frequently than the AAC at amino acid 3. We conclude that the satellite species is the result of basal level lysine for asparagine substitution. These substitutions are most likely caused by preferential misreading of AAU codons at a frequency of approximately 5 X 10(-3), 10-fold higher than the average error frequency.

Amino Acid Sequence

Co-evolution of base composition and codon usage in Xenopus laevis and human globin genes with long-range DNA organization of their genome.

Eucaryotic DNA is punctuated by many A+T-rich segments that we named A+T-rich linkers. Two types of these A+T-rich linkers can be distinguished: (i) isolated A+T-rich linkers, and (ii) A+T-rich linkers crowded in clusters. We have analysed the distribution of A+T-rich linker across the alpha- and beta-globin gene domain in Xenopus laevis and human genomes using isodenaturation and electron microscopy. Comparison of our data with those previously obtained for the avian globin genes leads us to conclude that genes can be harboured indifferently in either domain. A correlation is established between the presence of A+T-rich linker inside introns and flanking regions and the A+T content of the coding sequence. For the coding sequence, a high A+T content is strongly correlated with high A+T content in the codon's third position and weakly in the first position.

Animals

Codon usage and secondary structure of MS2 phage RNA.

MS2 is an RNA bacteriophage (3569 bases). The secondary structure of the RNA has been determined, and is known to play an important role in regulating translation. Paired regions of the genome have a higher G+C content than unpaired regions. It has been suggested that this reflects selection for high G+C content to encourage pairing, but a re-analysis of the data together with computer simulation suggest that it is an automatic consequence in any RNA sequence of the way it folds up to minimise its free energy. It has also been suggested that the three registers in which pairing can occur in a coding region are used differentially to optimise the use of the redundancy of the genetic code, but re-analysis of the data shows only weak statistical support for this hypothesis.

Base Composition

Function of 3' non-coding sequences and stop codon usage in expression of the chloroplast psaB gene in Chlamydomonas reinhardtii.

The rate of mRNA decay is an important step in the control of gene expression in prokaryotes, eukaryotes and cellular organelles. Factors that determine the rate of mRNA decay in chloroplasts are not well understood. Chloroplast mRNAs typically contain an inverted repeat sequence within the 3' untranslated region that can potentially fold into a stem-loop structure. These stem-loop structures have been suggested to stabilize the mRNA by preventing degradation by exonuclease activity, although such a function in vivo has not been clearly established. Secondary structures within the translation reading frame may also determine the inherent stability of an mRNA. To test the function of the inverted repeat structures in chloroplast mRNA stability mutants were constructed in the psaB gene that eliminated the 3' flanking sequences of psaB or extended the open reading frame into the 3' inverted repeat. The mutant psaB genes were introduced into the chloroplast genome of Chlamydomonas reinhardtii. Mutants lacking the 3' stem-loop exhibited a 75% reduction in the level of psaB mRNA. The accumulation of photosystem I complexes was also decreased by a corresponding amount indicating that the mRNA level is limiting to PsaB protein synthesis. Pulse-chase labeling of the mRNA showed that the decay rate of the psaB mRNA was significantly increased demonstrating that the stem-loop structure is required for psaB mRNA stability. When the translation reading frame was extended into the 3' inverted repeat the mRNA level was reduced to only 2% of wild-type indicating that ribosome interaction with stem-loop structures destabilizes chloroplast mRNAs. The non-photosynthetic phenotype of the mutant with an extended reading frame allowed us to test whether infrequently used stop codons (UAG and UGA) can terminate translation in vivo. Both UAG and UGA are able to effectively terminate PsaB synthesis although UGA is never used in any of the Chlamydomonas chloroplast genes that have been sequenced.

Amino Acid Sequence

Codon usage, genetic code and phylogeny of Dictyostelium discoideum mitochondrial DNA as deduced from a 7.3-kb region.

We have sequenced a region (7,376-bp) of the mitochondrial (mt) DNA (54 kb) of the cellular slime mold, Dictyostelium discoideum. From the DNA and amino-acid sequence comparisons with known sequences, genes for ATPase subunit 9 (ATP9), cytochrome b (CYTB), NADH dehydrogenase subunits 1, 3 and 6 (ND1, ND3 and ND6), small subunit rRNA (SSU rRNA) and seven tRNAs (Arg, Asn, Cys, Lys, f-Met, Met and Pro) have been identified. The sequenced region of the mtDNA has a high average A + T-content (70.8%). The A + T-content of protein-genes (73.6%) is considerably higher than that of RNA genes (61.3%). Even with the strong AT-bias, the genetic code employed is most probably the universal one. All seven tRNAs are able to form typical clover leaf structures. The molecular phylogenetic trees of CYTB and SSU rRNA suggest that D. discoideum is closer to green plants than to animals and fungi.

Amino Acid Sequence

Models of nearly neutral mutations with particular implications for nonrandom usage of synonymous codons.

The population dynamics of nearly neutral mutations are studied using a single-site and a multisite model. In the latter model, the nucleotides in a sequence are completely linked and the selection schemes employed are additive, multiplicative, and additive with a threshold. Although the third selection scheme is very different from the first two, the three schemes produce identical results for a wide range of parameter values. Thus the present study provides a general theory for the population dynamics of nearly neutral mutations because the results can also be used to draw inferences about other selection schemes such as stabilizing selection and synergistic selection. It is shown that the number of slightly deleterious mutations accumulated in a sequence can be considerably larger under the multisite model than under the single-site model, particularly if the sequence is long or if the mutation rate per site is high. The results show that even a very slight selective difference between synonymous codons can produce a strong bias in codon usage. Three alternative explanations for the strong bias in codon usage in bacterial and yeast genes are considered. The implications of the present results for molecular evolution are discussed.

Biological Evolution

In vivo evidence for non-universal usage of the codon CUG in Candida maltosa.

An alkane-assimilating yeast Candida maltosa had been studied in order to establish systems suitable for biotransformation of hydrophobic compounds. However, functional expression of heterologous genes tested for this purpose had not been successful in several cases. On the other hand, it had been reported that the codon CUG, a universal leucine codon, is read as serine in C. cylindracea. The same altered codon usage had also been suggested by in vitro experiments in some Candida yeasts which are phylogenetically closely related to C. maltosa. In this study we have shown that the failure in functional expression of a heterologous gene is due to the fact that the codon CUG is read as serine in C. maltosa. This conclusion was drawn from the following experimental results: (1) when a cytochrome P450 gene of C. maltosa containing a CTG codon was expressed in C. maltosa, the corresponding amino acid was found to be serine, and not leucine; (2) a tRNA gene with an almost identical structure to that of the tRNASerCAG gene of C. albicans could be isolated from the genome of C. maltosa; (3) the Saccharomyces cerevisiae URA3 gene, which has one CTG codon, could not complement the ura3 mutation of C. maltosa as itself, but when the CTG codon was changed to another leucine codon, CTC, the mutated gene could complement the ura3 mutation. The last result is the first example of succeeding in functional expression of a heterologous gene in Candida species having an altered codon usage by changing the CTG codon in the gene to another codon.

Amino Acid Sequence

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals

Stable structure of thermophilic proton ATPase beta subunit.

F1-ATPase is the major enzyme for ATP synthesis in mitochondria, chloroplasts, and bacterial plasma membranes. F1-ATPase obtained from thermophilic bacterium PS3 (TF1) is the only ATPase which can be reconstituted from its primary structure. Its beta subunit constitutes the catalytic site, and is capable of forming hybrid F1's with E. coli alpha and gamma subunits. Since the stability of TF1 resides in its primary structure, we cloned a gene coding for TF1, and the primary structure of the beta subunit was deduced from the nucleotide sequence of the gene to compare the sequence with those of beta's of three major categories of F1's; prokaryotic membranes, chloroplasts, and mitochondria. The following results were obtained. Homology: The primary structure of the TF1 beta subunit (473 residues, Mr = 51,995.6) showed 89.3% homology with 270 residues which are identical in the beta subunits from human mitochondria, spinach chloroplasts, and E. coli. It contained regions homologous to several nucleotide-binding proteins. Secondary structure: The deduced alpha-helical (30.1%) and beta-sheet (22.3%) contents were consistent with those determined from the circular dichroism spectra. Residues forming reverse turns (Gly and Pro) were highly conserved among the F1 beta subunits. Substituted residues and stability of TF1: We compared the amino acid sequence of the TF1 beta subunit with those of the other F1 beta subunits mentioned above. The observed substitutions in the thermophilic subunit increased its propensities to form secondary structures, and its external polarity to form tertiary structure. Codon usage: The codon usage of the TF1 beta gene was found to be unique. The changes in codons that achieved these amino acid substitutions were much larger than those caused by minimal mutations, and the third letters of the optimal codons were either guanine or cytosine, except in codons for Gln, Lys, and Glu.

Amino Acid Sequence

The translational termination signal database (TransTerm) now also includes initiation contexts.

The TransTerm database of termination codon contexts has been extended to include sense codon usage, and initiation codon contexts. The database was constructed from 23,721 coding sequences from 93 organisms. The database contains: a) the sequence around the termination codon (-10, +10); b) the sequence around the initiation codon (-20, +10); c) the length, 'G+C%' of the third position of codons (GC3), the 'codon adaptation index' (CAI) and the 'effective number of codons' statistic (Nc); d) summary tables for each organism including total codon usage, stop codon and tetranucleotide stop-signal usage, and matrices tallying base frequencies at each position around the initiation and termination codons. The data are arranged to facilitate investigation of the relationships between the three phases of protein synthesis. The database is available electronically from EMBL.

Animals

Mutation pressure, natural selection, and the evolution of base composition in Drosophila.

Genome sequencing in a number of taxa has revealed variation in nucleotide composition both among regions of the genome and among functional classes of sites in DNA. Mutational biases, biased gene conversion, and natural selection have been proposed as causes of this variation. Here, we review patterns of base composition in Drosophila DNA. Nucleotide composition in Drosophila melanogaster varys regionally, and base composition is correlated between introns and exons. Drosophila species also show striking patterns of non-random codon usage. Patterns of synonymous codon usage and the biochemistry of translation suggest that natural selection may act at 'silent' sites. A relationship between recombination rates and codon usage and comparisons of the evolutionary dynamics of silent mutations within and between species support natural selection discriminating among synonymous codons. The causes of regional base composition variation are less clear. Progress in functional studies of non-coding DNA, further investigations of genome patterns, and statistical tests based on evolutionary theory will lead to a greater understanding of the contributions of mutational processes and natural selection in patterning genome-wide nucleotide composition.

Animals

Intrastrand parity rules of DNA base composition and usage biases of synonymous codons.

When there are no biases in mutation and selection between the two strands of DNA, the 12 possible substitution rates of the four nucleotides reduces to six (type 1 parity rule or PR1), and the intrastrand average base composition is expected to be A = T and G = C at equilibrium without regard to the G + C content of DNA (type 2 parity rule or PR2). Significant deviations from the parity rules in the third codon letters of the four-codon amino acids result mostly from selective biases rather than mutational biases between the two strands of DNA during evolution. The parity rules lay the foundation for evaluating the biases in synonymous codon usage in terms of (1) directional mutation pressure for variation of the DNA G + C content due to mutational biases between alpha-bases (A or T) and gamma-bases (G or C), (2) strand-bias mutation, for example, by DNA repair during transcription, and (3) functional selection in evolution, for example, due to tRNA abundance. The present analysis shows that, although the PR2 violation is common in the third codon letters of four-codon amino acids, the contribution of PR2 violation to the DNA G + C content of the third codon position is small and, in majority of cases, mildly counteracts the effect of the directional mutation pressure on the G + C content.

Animals

Consecutive low-usage leucine codons block translation only when near the 5' end of a message in Escherichia coli.

Insertion of nine consecutive low-usage CUA leucine codons after codon 13 of a 313-codon test mRNA strongly inhibited its translation without apparent effect on translation of other mRNAs containing CUA codons. In contrast, nine consecutive high-usage CUG leucine codons at the same position had no apparent effect, and neither low- nor high-usage codons affected translation when inserted after codon 223 or 307. Additional experiments indicated that the strong positional effect of the low-usage codons could not be accounted for by differences in stability of the mRNAs or in stringency of selection of the correct tRNA. The positional effect could be explained if translation complexes are less stable near the beginning of a message: slow translation through low-usage codons early in the message may allow most translation complexes to dissociate before they read through.

Blotting, Northern

Protein-encoding genes in the sulfothermophilic archaea Sulfolobus and Pyrococcus.

A number of unrelated protein-encoding genes from sulfothermophilic archaea, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Pyrococcus furiosus and Pyrococcus woesei, has been analyzed. In the Sulfolobus genus, the content of A + T is significantly higher than that of C + G and the base usage follows the order, A > T > G > C. In Pyrococcus, the A + T content is also higher than that of C + G, but with lower values; in the order of base usage, G precedes T. The codon usage of these sulfothermophiles has been determined; alternative start codons are frequently used in both genera; codon preferences reflect the rich A + T composition of the corresponding genomes; for both genera the codon bias is particularly evident within the different arginine triplets, where AGA and AGG are predominant. From the similarities in the codon usage, close taxonomic relationships become evident within the Sulfolobus or the Pyrococcus genus; a lower, but significant similarity is also clear between these genera. The synonymous codon usage of these sulfothermophiles shows similarities with that of Saccharomyces cerevisiae and bovine mitochondria, whereas clear divergences are observed with the halophilic archaeal genus, Halobacterium, or the eubacterium, Escherichia coli. The unrelated proteins of the considered sulfothermophiles have been analyzed for the content of hydrophobic residues; the comparison with mesophiles reveals a significant increase in the average hydrophobicity of amino acid residues. This finding could indicate a mechanism of adaptation of proteins in organisms living under extreme environments. It is noteworthy that an opposite trend, i.e. a decreased average hydrophobicity, occurs in unrelated halophilic proteins.

Animals

Unusual codon bias occurring within insertion sequences in Escherichia coli.

The large open reading frames of insertion sequences from Escherichia coli were examined for their spatial pattern of codon usage bias and distribution of rarely used codons. There is a bias in codon usage that is generally lower toward the terminal ends of the coding regions, which is reflected in the occurrence of an excess of nonpreferred codons in the 3' portions of the coding regions as compared with the 5' portions. In contrast, typical chromosomal genes have a lower codon usage bias toward the 5' ends of the coding regions. These results imply that the selective forces reflected in codon usage bias may differ according to position within the coding sequence. In addition, these constraints apparently differ in important ways between genes contained in insertion sequences and those in the chromosome.

Chromosomes, Bacterial