RNA codons and protein synthesis. IX. Synonym codon recognition by multiple species of valine-, alanine-, and methionine-sRNA.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We have placed aminoacyl-tRNA selection at individual codons in competition with a frameshift that is assumed to have a uniform rate. By assaying a reporter in the shifted frame, relative rates for association of the 29 YNN codons and their cognate aminoacyl-tRNAs were obtained during logarithmic growth in Escherichia coli. For five codons, three beginning with C and two with U, these relative rates agree with relative in vitro rates for elongation factor Tu-mediated aminoacyl-tRNA binding to ribosomes and subsequent GTP hydrolysis. Therefore, the frameshift assay probably measures this process in vivo. Observed rates for aminoacyl-tRNA selection span a 25-fold range. Therefore, the time required to transit different codons in vivo probably differs substantially. Codons very frequently used in highly expressed genes generally select aminoacyl-tRNAs more quickly than do rarely used codons. This suggests that speed of aminoacyl-tRNA selection is a significant factor determining biased use of synonymous codons. However, the preferential use of codons appears to be marked only for codons with the highest rates of aminoacyl-tRNA selection. Rapid selection in vivo is usually effected by elevation of the tRNA concentration for codons with moderate intrinsic speed (rate constant), not by choosing intrinsically fast codons. Despite a preference for high rate, there are quickly translated codons that are not commonly used, and common codons that are translated relatively slowly. Other factors are therefore more important than speed for some codons. Strong preference for rapid aminoacyl-tRNA selection is not observed in weakly expressed genes. Instead, there is a slight preference for slower aminoacyl-tRNA selection. The rate of aminoacyl-tRNA selection by a YNC codon is always greater than the rate of the corresponding YNU codon even though in many YNC/U pairs both codons react with the same elongation factor Tu/GTP/aminoacyl-tRNA complex. Thus, for these tRNAs, the differences between in vivo rate constants of tRNAs are dependent on the nature of anticodon base-pairing. However, no more general relationship is evident between codon/anticodon composition and rate of aminoacyl-tRNA selection. The frameshift method can be extended to all codons.
Codon usage is compared between four classes of species, with an emphasis on characterization of low-usage codons. The classes of species analyzed include the bacterium Escherichia coli (ECO), the yeast Saccharomyces cerevisiae (YSC), the fruit fly Drosophila melanogaster (DRO), and several species of primates (PRI) (taken as a group; includes eleven species for which nucleotide sequence data have been reported to GenBank, however, greater than 90% of the sequences were from Homo sapiens). The number of protein-coding sequences analyzed were 968 for ECO, 484 for YSC, 244 for DRO, and 1518 for PRI. Three methods have been used to determine low-usage codons in these species. The first and most common way of assessing codon usage is by summing the number of time codons appear in reading frames of the genome in question. The second way is to examine the distribution of usage in different genes by scoring the number of protein reading frames in which a particular codon does not appear. The third way starts with a similar notion, but instead considers combinations of codons that are missing from the maximum number of genes. These three methods give very similar results. Each species has a unique combination of eight least-used codons, but all species contain the arginine codons, CGA and CGG. The agreement between YSC and PRI is particularly striking as they share six low-usage codons. All six carry the dinucleotide sequence, CG. The eight least-used codons in PRI include all codons that contain the CG dinucleotide sequence. Low-usage codons are clearly avoided in genes encoding abundant proteins for ECO, YSC DRO. In all species, proteins containing a high percentage of low-usage codons could be characterized as cases where an excess of the protein could be detrimental. Low codon usage is relatively insensitive to gross base composition. However, dinucleotide usage can sometimes influence codon usage. This is particularly notable in the case of CG dinucleotides in PRI.
This investigation of the codon context of enterobacteria, plasmid, and phage protein genes was based on a search for correlations between the presence of one base type at codon position III and the presence of another base type at some other position in adjacent codons. Enterobacterial genes were compared with eukaryotic sequences for codon context effects. In enterobacterial genes, base usage at codon position III is correlated with the third position of the upstream adjacent codon and with all three positions of the downstream codon. Plasmid genes are free of context biases. Phage genes are heterogeneous: MS2 codons have no biased context, whereas lambda genes partly follow the trends of the host bacterium, and T7 genes have biased codon contexts that differ from those of the host. It has been reported that two successive third-codon positions tend to be occupied by two purines or two pyrimidines in Escherichia coli genes of low expression level. Here, the extent to which highly expressed protein genes can modulate base usage at two successive codon positions III, given the constraints on codon usage and protein sequence that act on them, was quantified. This demonstrates that the above-mentioned favored patterns are not a characteristic of weakly expressed genes but occur in all genes in which codon context can vary appreciably. The correlation between successive third-codon positions is a distinct feature of enterobacteria and of some phages, one that may result from adaptation of gene structure to translational efficiency. Conversely, codon context in yeast and human genes is biased--but for reasons unrelated to translation.
A series of Saccharomyces cerevisiae plasmids and mutant derivatives containing fusions of the Escherichia coli galactokinase gene, galK, to the yeast iso-1-cytochrome c CYC1 transcription unit were used to study the sequences affecting the initiation of translation in S. cerevisiae. When the CYC1 AUG initiation codon preceded the galK AUG codon and coding sequence and either the two AUGs were out of frame with each other or a nonsense codon was located between them, the expression of the galK gene was extremely low. Deletion of the CYC1 AUG and its surrounding sequences resulted in a 100-fold increase in galK expression. This dependence of galK expression on the elimination of the CYC1 AUG codon was used to select mutations in that codon. Then the ability of these altered initiation codons to serve in translational initiation was determined by reconstruction of the CYC1 gene 3' to and in frame with them. Initiation was found to occur at the codons UUG and AUA, but not at the codons AAA and AUC. Furthermore the codon UUG, when preceded by an A three nucleotides upstream, served as a better initiation codon than when a U was substituted for the A. The efficiency of translation from these non-AUG codons was quantitated by using a CYC1/galK protein-coding fusion and measuring cellular galactokinase levels. Initiation at the UUG codon was 6.9% as efficient as initiation at the wild-type AUG codon when preceded by an A three nucleotides upstream, but was over 10-fold less efficient when a U was substituted for that A. Initiation at AUA was 0.5% as efficient as at AUG. The effects of the sequences preceding the initiation codon are discussed in light of these results.
The constraints on nucleotide sequences of highly and weakly expressed genes from Escherichia coli have been analysed and compared. Differences in synonymous codon spectra in highly and weakly expressed genes lead to different frequencies of nucleotides (in the first and third codon positions) and dinucleotides in the two groups of genes. It has been found that the choice of synonymous codons in highly expressed genes depends on the nucleotides adjacent to the codon. For example, lysine is preferably encoded by the AAA codon if guanosine is 3' to the lysine codon (AAA-G, P less than 10(-9)). And, on the contrary, AAG is used more often than AAA (P less than 0.001) if cytidine is 3' adjacent to lysine. Guanosine occurs more frequently than adenosine 5' to all the lysine codons (AAR, P less than 10(-5), i.e. NNG codons are preferred over the synonymous NNA codons 5' to the positions of lysine in the genes. The context effect was observed in nonsense and missense suppression experiments. Therefore, a hypothesis has been suggested that the efficiency of translation of some codons (for which the constraints on the adjacent nucleotides were found) can be modulated by the codon context. The rules for preferable synonymous codon choice in highly expressed genes depending on the nucleotides surrounding the codon are presented. These rules can be used in the chemical synthesis of genes designed for expression in E. coli.
We report the sequences of Neurospora crassa mitochondrial alanine, leucine(1), leucine(2), threonine, tryptophan, and valine tRNAs. On the basis of the anticodon sequences of these tRNAs and of a glutamine tRNA, whose sequence analysis is nearly complete, we infer the following: (i) The N. crassa mitochondrial tRNA species for alanine, leucine(2), threonine, and valine, amino acids that belong to four-codon families (GCN, CUN, ACN, and GUN, respectively; N = U, C, A, or G) all contain an unmodified U in the first position of the anticodon. In contrast, tRNA species for glutamine, leucine(1), and tryptophan, amino acids that use codons ending in purines (CA(G) (A), UU(G) (A), and UG(G) (A), respectively) contain a modified U derivative in the same position. These findings and the fact that we have not detected any other isoacceptor tRNAs for these amino acids suggest that N. crassa mitochondrial tRNAs containing U in the first position of the anticodon are capable of reading all four codons of a four-codon family whereas those containing a modified U are restricted to reading codons ending in A or G. Such an expanded codon-reading ability of certain mitochondrial tRNAs will explain how the mitochondrial protein-synthesizing system operates with a much lower number of tRNA species than do systems present in prokaryotes or in eukaryotic cytoplasm. (ii) The anticodon sequence of the N. crassa mitochondrial tryptophan tRNA is U(*)CA and not CCA or CmCA as is the case with tryptophan tRNAs from prokaryotes or from eukaryotic cytoplasm. Because a tRNA with U(*)CA in the anti-codon would be expected to read the codon UGA, as well as the normal tryptophan codon UGG, this suggests that in N. crassa mitochondria, as in yeast and in human mitochondria, UGA is a codon for tryptophan and not a signal for chain termination. (iii) The anticodon sequences of the two leucine tRNAs indicate that N. crassa mitochondria use both families of leucine codons (UU(A) (G) and CUN; N = U, C, A, or G) for leucine, in contrast to yeast mitochondria [Li, M. & Tzagoloff, A. (1979) Cell 18, 47-53] in which the CUA leucine codon and possibly the entire CUN family of leucine codons may be translated as threonine.
The cat-86 gene specifies chloramphenicol acetyltransferase (CAT). The cat-86 start codon is UUG, although related genes have AUG as the start codon. Changing the start codon to AUG increased expression of cat-86 by 36% in Bacillus subtilis. Changing the start codon to GUG and CUG decreased expression to 65% and 30%, respectively, of the level obtained when AUG was the start codon. CUG has not been previously shown to function as a start codon in B. subtilis. N-terminal sequencing of purified CAT protein specified by the CUG mutant, revealed that CUG was indeed the start codon and specified methionine. The gene xylE, which specifies catechol 2,3-dioxygenase, has AUG as its start codon. Changing the start codon for xylE to CUG decreased expression by 98%. However, when the ribosome-binding site sequence for xylE was optimized and the spacing between it and the start codon was increased to 8 nucleotides, xylE activity increased to 13% of the activity observed for AUG. CUG did not function efficiently as a start codon for cat-86 in Escherichia coli. These data suggest conditions under which CUG can function, with modest efficiency, as a start codon in B. subtilis.
This study reports the analysis of codon usage in 35 complete Homo sapiens genes. Both codon frequency and inter-codon interference exhibit patterns of evolutionary interest. There is a significant positive correlation between the frequency with which a given codon is used and the frequency with which its complement is used. Since the frequency of appearance of the complementary codon on the coding strand is equal to the frequency of appearance of the original codon on the non-coding strand, in the same phase, the non-coding strand is found to resemble the coding strand in triplet composition. The same effect has been observed in Escherichia coli. This preference for the use of certain complementary triplets as codons suggests that the evolution of the use of the genetic code depended to some extent upon the double-stranded nature of the coding material. In addition, the effect of discrimination against the use of two dinucleotides, CpG and UpA, is observed in codon usage and also in adjacent codon interference. Codons beginning with G, or A, are unlikely to be preceded by codons ending in C, or U, respectively. Consideration of codon assignment in the genetic code together with the observed CpG infrequency suggests that the evolution of the code may have been influenced by conditions in which the use of CpG dinucleotides was unfavorable. The infrequent use of UpA dinucleotides can be explained as the result of frameshift mutation during gene evolution.
Extreme codon bias is seen for the Saccharomyces cerevisiae genes for the fermentative alcohol dehydrogenase isozyme I (ADH-I) and glyceraldehyde-3-phosphate dehydrogenase. Over 98% of the 1004 amino acid residues analyzed by DNA sequencing are coded for by a select 25 of the 61 possible coding triplets. These preferred codons tend to be highly homologous to the anticodons of the major yeast isoacceptor tRNA species. Codons which necessitate site by side GC base pairs between the codons and the tRNA anticodons are always avoided whenever possible. Codons containing 100% G, C, A, U, GC, or AU are also avoided. This provides for approximately equivalent codon-anticodon binding energies for all preferred triplets. All sequenced yeast genes show a distinct preference for these same 25 codons. The degree of preference varies from greater than 90% for glyceraldehyde-3-phosphate dehydrogenase and ADH-I to less than 20% for iso-2 cytochrome c. The degree of bias for these 25 preferred triplets in each gene is correlated with the level of its mRNA in the cytoplasm. Genes which are strongly expressed are more biased than genes with a lower level of expression. A similar phenomenon is observed in the codon preferences of highly expressed genes in Escherichia coli. High levels of gene expression are well correlated with high levels of codon bias toward 22 of the 61 coding triplets. As in yeast, these preferred codons are highly complementary to the major cellular isoacceptor tRNA species. In at least four cases (Ala, Arg, Leu, and Val), these preferred E. coli codons are incompatible with the preferred yeast codons.
To test features of the current model of transcription attenuation in amino acid biosynthetic operons, alterations were introduced into the trp operon leader region and expression of the mutated operons was examined in miaA and miaA+ Escherichia coli strains that lacked the trp repressor. The miaA mutation prevents modification of the adenosine residue immediately 3' of the anticodon of tRNAs that interact with codons beginning with uridine. The undermodified tRNA(Trp) in miaA strains is thought to increase readthrough at the trp attenuator by slowing ribosome movement over two tandem Trp codons in the 14-codon leader peptide coding region. The rate of translation of these two "control codons" is thought to be the key step in determining the extent of transcription attenuation in the trp leader region. Sequential deletion of trpL DNA specifying the leader peptide initiation region, RNA segment 1, RNA segment 2 and RNA segment 3 alternately decreased and increased trp operon expression, a result consistent with previous findings in another bacterium and the generally accepted model for transcription attenuation. Replacement of the tandem Trp control codons by AGG-UGC (Arg-Cys) codons eliminated the miaA-dependent increase in transcription readthrough. Replacement of the Trp control codons by AGG-UGA (Arg-stop) codons caused complete readthrough at the trp attenuator as well as abolishing the miaA effect. Presumably, the ribosome terminating translation at the new UGA codon mimics the effect of a stalled ribosome at the Trp control codons. This finding suggests that ribosome dissociation at some stop codons is slow relative to the time required for transcription of the trp leader region. Thus, most ribosomes translating the trp leader peptide coding region may remain attached to the natural UGA stop codon until after the attenuation decision is made. The interpretation supports models for trp operon attenuation in which the elevated basal level readthrough is determined by occasional ribosome release prior to synthesis of the 3:4 terminator hairpin.
We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.
I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.
CGG is an arginine codon in the universal genetic code. We previously reported that in Mycoplasma capricolum, a relative of Gram-positive eubacteria, codon CGG did not appear in coding frames, including termination sites, and tRNA(ArgCCG) pairing with codon CGG, was not detected. These facts suggest that CGG is a nonsense (unassigned and untranslatable) codon--i.e., not assigned to arginine or to any other amino acid. We have investigated whether CGG is really an unassigned codon by using a cell-free translation system prepared from M. capricolum. Translation of synthetic mRNA containing in-frame CGG codons does not result in "read-through" to codons beyond the CGG codons--i.e., translation ceases just before CGG. Sucrose-gradient centrifugation profiles of the reaction mixture have shown that the bulk of peptide that has been synthesized is attached to 70S ribosomes and is released upon further incubation with puromycin. The result suggests that the peptide is in the P site of ribosome in the form of peptidyl-tRNA, leaving the A site empty. When in-frame CGG codons are replaced by UAA codons in mRNA, no read-through occurs beyond UAA, just as in the case of CGG. However, the synthesized peptide is released from 70S ribosomes, presumably by release factor 1. These data suggest strongly that CGG is an unassigned codon and differs from UAA in that CGG is not used for termination.
To study translation initiation in Chlamydomonas chloroplasts, we mutated the initiation codon AUG to AUU, ACG, ACC, ACU, and UUC in the chloroplast petA gene, which encodes cytochrome f of the cytochrome b6/f complex. Cytochrome f accumulated to detectable levels in all mutant strains except the one with a UUC codon, but only the mutant with an AUU codon grew well at 24 degrees C under conditions that require photosynthesis. Because no cytochrome f was detectable in the UUC mutant and because each mutant that accumulated cytochrome f did so at a different level, we concluded that any residual translation probably initiates at the mutant codon. As a further demonstration that alternative initiation sites are not used in vivo, we introduced in-frame UAA stop codons immediately downstream or upstream or in place of the initiation codon. Stop codons at or downstream of the initiation codon prevented accumulation of cytochrome f, whereas the one immediately upstream of the initiation codon had no effect on the accumulation of cytochrome f. These results suggest that an AUG codon is not required to specify the site of translation initiation in chloroplasts but that the efficiency of translation initiation depends on the identity of the initiation codon.
Leucine participates in multivalent repression of the Serratia marcescens ilvGMEDA operon by attenuation (J.-H. Hsu, E. Harms, and H.E. Umbarger, J. Bacteriol. 164:217-222, 1985), although there is only one single leucine codon that could be involved in this type of control. This leucine codon is the rarely used CUA. The contribution of this leucine codon to the control of transcription by attenuation was examined by replacing it with the commonly used leucine codon CUG and with a nonregulatory proline codon, CCG. These changes left intact the proposed secondary structure of the leader. The effects of the codon changes were assessed by placing the mutant leader regions upstream of the ilvGME structural genes or the cat gene and measuring acetohydroxy acid synthase II, transaminase B, or chloramphenicol acetyltransferase activities in cells grown under limiting and repressing conditions. The presence of the common leucine codon in place of the rare leucine codon reduced derepression by about 70%. Eliminating the leucine codon by converting it to proline abolished leucine control. Furthermore, a possible context effect of the adjacent upstream serine codon on leucine control was examined by changing it into a glycine codon.
The context requirements for recognition of an initiator codon were evaluated in vitro by monitoring the relative use of two AUG codons that were strategically positioned to produce long (pre-chloramphenicol acetyl transferase [CAT]) and short versions of CAT protein. The yield of pre-CAT initiated from the 5'-proximal AUG codon increased, and synthesis of CAT from the second AUG codon decreased, as sequences flanking the first AUG codon increasingly resembled the eucaryotic consensus sequence. Thus, under prescribed conditions, the fidelity of initiation in extracts from animal as well as plant cells closely mimics what has been observed in vivo. Unexpectedly, recognition of an AUG codon in a suboptimal context was higher when the adjacent downstream sequence was capable of assuming a hairpin structure than when the downstream region was unstructured. This finding adds a new, positive dimension to regulation by mRNA secondary structure, which has been recognized previously as a negative regulator of initiation. Translation of pre-CAT from an AUG codon in a weak context was not preferentially inhibited under conditions of mRNA competition. That result is consistent with the scanning model, which predicts that recognition of the AUG codon is a late event that occurs after the competition-sensitive binding of a 40S ribosome-factor complex to the 5' end of mRNA. Initiation at non-AUG codons was evaluated in vitro and in vivo by introducing appropriate mutations in the CAT and preproinsulin genes. GUG was the most efficient of the six alternative initiator codons tested, but GUG in the optimal context for initiation functioned only 3 to 5% as efficiently as AUG. Initiation at non-AUG codons was artifactually enhanced in vitro at supraoptimal concentrations of magnesium.