Search PubMedSearch

SEARCH · Search PubMed

Results for “Codons”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Clustering of low usage codons and ribosome movement.

A model is presented in which the distribution of low-usage codons in a message is a major factor in determining the impact that they will have on the translation rate and distribution of ribosomes on that message. This model is based on the assumption that low-usage codons are translated more slowly than normal codons, an assumption supported by various lines of published experimental evidence. Although the parameters used to develop this model are somewhat arbitrary, the main conclusions of this paper are consistent with a wide variation in the values of those parameters. In the model, low-usage codons arranged in clusters are much more effective in blocking ribosome movement on the message than ones that are dispersed. The effective size of the cluster is limited to the dimensions of the ribosome. It has been estimated that ribosomes on a message are spaced at least 27 nucleotides or nine codons apart. A ribosome translating a cluster of nine codons in which some or all of the codons are low-usage will move more slowly than over a comparable stretch of message containing no low-usage codons. Owing to ribosome size, the ribosome immediately behind the stalled ribosome will move as slowly; it must wait for the stalled ribosome to move on before it can even begin to translate the difficult region containing the low-usage codons. When the low-usage codon cluster is at the 3' end, the message will eventually be occupied by a ribosome jam that will transmit back to the 5' end of the message. In the steady state, the slowing effect imposed by a cluster of nine low-usage codons at the 3' end of a message would be just as great as if the entire message was composed of them. If the cluster is situated in the middle of a message, the ribosomes will form a jam upstream of the cluster. The ribosome density downstream of the cluster will be considerably reduced from what it would be for the same message with no cluster. If the cluster is at the 5' end of the message, the density of ribosomes will be reduced over the entire length of the message but the overall translation rate per ribosome will be only slightly reduced. However, owing to the reduced number of ribosomes initiating, the efficiency of the message in protein synthesis will be considerably reduced.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals

Contextual constraints on synonymous codon choice.

We have studied the statistical constraints on synonymous codon choice to evaluate various proposals regarding the origin of the bias in synonymous codon usage observed by Fiers et al. (1975), Air et al. (1976), Grantham et al. (1980) and others. We have determined the statistical dependence of the degenerate third base on either of its nearest neighbors in mitochondrial, prokaryotic, and eukaryotic coding sequences. We noted an increasing dependence of the third base on its nearest neighbors in moving from mitochondria to prokaryotes to eukaryotes. A statistical model assuming random equiprobable selection of synonymous codons was found grossly adequate for the mitochondria, but totally inadequate for prokaryotes and eukaryotes. A model assuming selection of synonymous codons reflecting a genomic strategy, i.e. the genome hypothesis of Grantham et al. (1980), gave a good approximation of the mitochondrial sequences. A statistical model which exactly maintains codon frequency, but allows the position of corresponding synonymous codons to vary was only grossly adequate for prokaryotes and totally inadequate for eukaryotes. The results of these simulations are consistent with the measures on experimental sequences and suggest that a "frequency constraint" model such as that of Grantham et al. (1980) may be an adequate explanation of the codon usage in mitochondria. However, in addition to this frequency constraint, there may be constraints on synonymous codon choice in prokaryotes due to codon context. Furthermore, any proposal to explain codon usage in eukaryotes must involve a constraint on the context of a codon in the sequence.

Amino Acid Sequence

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins

Codon usage in Kluyveromyces lactis and in yeast cytochrome c-encoding genes.

Codon usage (CU) in Kluyveromyces lactis has been studied. Comparison of CU in highly and lowly expressed genes reveals the existence of 21 optimal codons; 18 of them are also optimal in other yeasts like Saccharomyces cerevisiae or Candida albicans. Codon bias index (CBI) values have been recalculated with reference to the assignment of optimal codons in K. lactis and compared to those previously reported in the literature taking as reference the optimal codons from S. cerevisiae. A new index, the intrinsic codon deviation index (ICDI), is proposed to estimate codon bias of genes from species in which optimal codons are not known; its correlation with other index values, like CBI or effective number of codons (Nc), is high. A comparative analysis of CU in six cytochrome-c-encoding genes (CYC) from five yeasts is also presented and the differences found in the codon bias of these genes are discussed in relation to the metabolic type to which the corresponding yeasts belong. Codon bias in the CYC from K. lactis and S. cerevisiae is correlated to mRNA levels.

Amino Acids

CGG: an unassigned or nonsense codon in Mycoplasma capricolum.

CGG is an arginine codon in the universal genetic code. We previously reported that in Mycoplasma capricolum, a relative of Gram-positive eubacteria, codon CGG did not appear in coding frames, including termination sites, and tRNA(ArgCCG) pairing with codon CGG, was not detected. These facts suggest that CGG is a nonsense (unassigned and untranslatable) codon--i.e., not assigned to arginine or to any other amino acid. We have investigated whether CGG is really an unassigned codon by using a cell-free translation system prepared from M. capricolum. Translation of synthetic mRNA containing in-frame CGG codons does not result in "read-through" to codons beyond the CGG codons--i.e., translation ceases just before CGG. Sucrose-gradient centrifugation profiles of the reaction mixture have shown that the bulk of peptide that has been synthesized is attached to 70S ribosomes and is released upon further incubation with puromycin. The result suggests that the peptide is in the P site of ribosome in the form of peptidyl-tRNA, leaving the A site empty. When in-frame CGG codons are replaced by UAA codons in mRNA, no read-through occurs beyond UAA, just as in the case of CGG. However, the synthesized peptide is released from 70S ribosomes, presumably by release factor 1. These data suggest strongly that CGG is an unassigned codon and differs from UAA in that CGG is not used for termination.

Amino Acid Sequence

Selection intensity for codon bias.

The patterns of nonrandom usage of synonymous codons (codon bias) in enteric bacteria were analyzed. Poisson random field (PRF) theory was used to derive the expected distribution of frequencies of nucleotides differing from the ancestral state at aligned sites in a set of DNA sequences. This distribution was applied to synonymous nucleotide polymorphisms and amino acid polymorphisms in the gnd and putP genes of Escherichia coli. For the gnd gene, the average intensity of selection against disfavored synonymous codons was estimated as approximately 7.3 x 10(-9); this value is significantly smaller than the estimated selection intensity against selectively disfavored amino acids in observed polymorphisms (2.0 x 10(-8)), but it is approximately of the same order of magnitude. The selection coefficients for optimal synonymous codons estimated from PRF theory were consistent with independent estimates based on codon usage for threonine and glycine. Across 118 genes in E. coli and Salmonella typhimurium, the distribution of estimated selection coefficients, expressed as multiples of the effective population size, has a mean and standard deviation of 0.5 +/- 0.4. No significant differences were found in the degree of codon bias between conserved positions and replacement positions, suggesting that translational misincorporation is not an important selective constraint among synonymous polymorphic codons in enteric bacteria. However, across the first 100 codons of the genes, conserved amino acids with identical codons have significantly greater codon bias than that of either synonymous or nonidentical codons, suggesting that there are unique selective constraints, perhaps including mRNA secondary structures, in this part of the coding region.

Codon

The initiation codon determines the efficiency but not the site of translation initiation in Chlamydomonas chloroplasts.

To study translation initiation in Chlamydomonas chloroplasts, we mutated the initiation codon AUG to AUU, ACG, ACC, ACU, and UUC in the chloroplast petA gene, which encodes cytochrome f of the cytochrome b6/f complex. Cytochrome f accumulated to detectable levels in all mutant strains except the one with a UUC codon, but only the mutant with an AUU codon grew well at 24 degrees C under conditions that require photosynthesis. Because no cytochrome f was detectable in the UUC mutant and because each mutant that accumulated cytochrome f did so at a different level, we concluded that any residual translation probably initiates at the mutant codon. As a further demonstration that alternative initiation sites are not used in vivo, we introduced in-frame UAA stop codons immediately downstream or upstream or in place of the initiation codon. Stop codons at or downstream of the initiation codon prevented accumulation of cytochrome f, whereas the one immediately upstream of the initiation codon had no effect on the accumulation of cytochrome f. These results suggest that an AUG codon is not required to specify the site of translation initiation in chloroplasts but that the efficiency of translation initiation depends on the identity of the initiation codon.

Animals

Role of codon choice in the leader region of the ilvGMEDA operon of Serratia marcescens.

Leucine participates in multivalent repression of the Serratia marcescens ilvGMEDA operon by attenuation (J.-H. Hsu, E. Harms, and H.E. Umbarger, J. Bacteriol. 164:217-222, 1985), although there is only one single leucine codon that could be involved in this type of control. This leucine codon is the rarely used CUA. The contribution of this leucine codon to the control of transcription by attenuation was examined by replacing it with the commonly used leucine codon CUG and with a nonregulatory proline codon, CCG. These changes left intact the proposed secondary structure of the leader. The effects of the codon changes were assessed by placing the mutant leader regions upstream of the ilvGME structural genes or the cat gene and measuring acetohydroxy acid synthase II, transaminase B, or chloramphenicol acetyltransferase activities in cells grown under limiting and repressing conditions. The presence of the common leucine codon in place of the rare leucine codon reduced derepression by about 70%. Eliminating the leucine codon by converting it to proline abolished leucine control. Furthermore, a possible context effect of the adjacent upstream serine codon on leucine control was examined by changing it into a glycine codon.

Base Sequence

Context effects and inefficient initiation at non-AUG codons in eucaryotic cell-free translation systems.

The context requirements for recognition of an initiator codon were evaluated in vitro by monitoring the relative use of two AUG codons that were strategically positioned to produce long (pre-chloramphenicol acetyl transferase [CAT]) and short versions of CAT protein. The yield of pre-CAT initiated from the 5'-proximal AUG codon increased, and synthesis of CAT from the second AUG codon decreased, as sequences flanking the first AUG codon increasingly resembled the eucaryotic consensus sequence. Thus, under prescribed conditions, the fidelity of initiation in extracts from animal as well as plant cells closely mimics what has been observed in vivo. Unexpectedly, recognition of an AUG codon in a suboptimal context was higher when the adjacent downstream sequence was capable of assuming a hairpin structure than when the downstream region was unstructured. This finding adds a new, positive dimension to regulation by mRNA secondary structure, which has been recognized previously as a negative regulator of initiation. Translation of pre-CAT from an AUG codon in a weak context was not preferentially inhibited under conditions of mRNA competition. That result is consistent with the scanning model, which predicts that recognition of the AUG codon is a late event that occurs after the competition-sensitive binding of a 40S ribosome-factor complex to the 5' end of mRNA. Initiation at non-AUG codons was evaluated in vitro and in vivo by introducing appropriate mutations in the CAT and preproinsulin genes. GUG was the most efficient of the six alternative initiator codons tested, but GUG in the optimal context for initiation functioned only 3 to 5% as efficiently as AUG. Initiation at non-AUG codons was artifactually enhanced in vitro at supraoptimal concentrations of magnesium.

Animals

The use of logistic models for the analysis of codon frequencies of DNA sequences in terms of explanatory variables.

The development of the regressive logistic model applicable to the analysis of codon frequencies of DNA sequences in terms of explanatory variables is presented. A codon is a triplet of nucleotides that code for an amino acid, and may be considered as a trivariate response (B1, B2, B3), where Bi (i = 1, 2, 3) is a categorical random variable with values A, C, G, T. The linear order of bases in the DNA and possible statistical dependence of the bases in a given codon make the regressive logistic model a suitable tool for the analysis of codon frequencies. A problem of structural zeros arises from the fact that the stopping codons (terminators) do not code for amino acids; this is solved by normalizing the likelihood function. Codon frequencies may also depend on the function of the gene and they are known to differ between genes of the same genome. Differences also occur between synonymous codons for the same amino acid. Thus, the use of covariates that differ between synonymous codons as well as covariates that are constant within codons of the same amino acid may be useful in explaining the frequencies. As an illustration, the method is applied to the human mitochondrial genome using the following as explanatory variables: (1) TSCORE, a measure of the number of single base mutations required for a given codon to become a terminator; (2) AARISK, an indicator of a codon's ability of changing by a single base substitution to triplets coding for amino acids with very different characteristics; (3) AVDIST, a measure of the typicality of the amino acid coded for by the triplets. The results indicate that models that incorporate dependency structure and covariates are to be preferred to either the models comprising covariates alone or dependency structure alone.

Amino Acid Sequence

Codon usage and bias among individual genes of the coccidia and piroplasms.

Codon usage has been analysed in individual gene sequences, derived from a variety of parasitic protozoa in the class Sporozoa of the phylum Apicomplexa using metric multidimensional scaling. The two groups of codon usage patterns detected reflect the two main subgroups of organisms studied (the coccidia and the piroplasms), and it is the pattern of usage of synonymous codons that has the largest influence on overall codon usage in the individual genes, rather than being the pattern of amino acid composition of the gene product. The magnitude of the codon usage bias in the sequences was determined using three commonly used indices-NC, GC3S and B. In general, although relatively low levels of codon usage bias were detected in these gene sequences, codon usage bias does explain at least some of the codon usage patterns observed. Codon usage bias was observed to be dependent on the overall base composition of the genes analysed, which in turn was reflected in the types of codons that were either over- or under-represented in the nucleotide sequences. In keeping with observations on prokaryotic organisms, it is speculated that the codon usage patterns detected in these parasitic protozoa are the result of directional mutation pressure on the base composition of the genomic DNA.

Animals

The second to last amino acid in the nascent peptide as a codon context determinant.

Forty-two different sense codons, coding for all 20 amino acids, were placed at the ribosomal E site location, two codons upstream of a UGA or UAG codon. The influence of these variable codons on readthrough of the stop codons was measured in Escherichia coli. A 30-fold difference in readthrough of the UGA codon was observed. Readthrough is not related to any property of the upstream codon, its cognate tRNA or the nature of its codon-anticodon interaction. Instead, it is the amino acid corresponding to the second upstream codon, in particular the acidic/basic property of this amino acid, which seems to be a major determinant. This amino acid effect is influenced by the identity of the A site stop codon and the efficiency of its decoding tRNA, which suggests a correlation with ribosomal pausing. The magnitude of the amino acid effect is in some cases different when UGA is decoded by a wildtype form of tRNA(Trp) as compared with a suppressor form of the same tRNA. This indicates that the structure of the A site decoding tRNA is also a determinant for the amino acid effect.

Amino Acids

Chloroplast DNA codon use: evidence for selection at the psb A locus based on tRNA availability.

Codon use in the three sequenced chloroplast genomes (Marchantia, Oryza, and Nicotiana) is examined. The chloroplast has a bias in that codons NNA and NNT are favored over synonymous NNC and NNG codons. This appears to be a consequence of an overall high A + T content of the genome. This pattern of codon use is not followed by the psb A gene of all three genomes and other psb A sequences examined. In this gene, the codon use favors NNC over NNT for twofold degenerate amino acids. In each case the only tRNA coded by the genome is complementary to the NNC codon. This codon use is similar to the codon use by chloroplast genes examined from Chlamydomonas reinhardtii. Since psb A is the major translation product of the chloroplast, this suggests that selection is acting on the codon use of this gene to adapt codons to tRNA availability, as previously suggested for unicellular organisms.

Animals

Evolution of the mitochondrial genetic code. I. Origin of AGR serine and stop codons in metazoan mitochondria.

AGA and AGG (AGR) are arginine codons in the universal genetic code. These codons are read as serine or are used as stop codons in metazoan mitochondria. The arginine residues coded by AGR in yeast or Trypanosoma are coded by arginine CGN throughout metazoan mitochondria. AGR serine sites in metazoan mitochondria are occupied mainly in corresponding sites in yeast or Trypanosoma mitochondria by UCN serine, AGY serine, or codons for amino acids other than serine or arginine. Based on these observations, we propose the following evolutionary events. AGR codons became unassigned because of deletion of tRNA Arg (UCU) and elimination of AGR codons by conversion to CGN arginine codons. Upon acquisition by serine tRNA of pairing ability with AGR codons, some codons for amino acids other than arginine mutated to AGR, and were captured by anticodon GCU in serine tRNA. During vertebrate mitochondrial evolution, AGR stop codons presumably were created from UAG stop by deletion of the first nucleotide U and by use of R as the third nucleotide that had existed next to the ancestral UAG stop.

Animals

Analysis of the stop codon context in plant nuclear genes.

A region of 18 nucleotides surrounding the stop codon (the stop codon context) in 748 plant nuclear genes was analyzed. Non-randomness was found both upstream and downstream from the stop codon, suggesting that these sequences may help in ensuring efficient termination of translation. The UAG amber codon is the least-used stop codon and the bias in the nucleotide distribution 5' and 3' to the stop codon was more pronounced for the amber codon than for the other stop codons. This might indicate that the codon context affects termination more at UAG than at UGA or UAA stop codons.

Base Composition

A study of the purine/pyrimidine codon occurrence with a reduced centered variable and an evaluation compared to the frequency statistic.

With the three-letter alphabet [R,Y,N] (R = purine, Y = pyrimidine, N = R or Y), there are 26 codons (NNN being excluded): RNN,...,NNY (six codons at two unspecified bases N), RRN,...,NYY (12 codons at one unspecified base N), RRR,...,YYY (eight specified codons). A statistical methodology that uses the codon frequency and a reduced centered variable leads to similar results for a codon occurrence study, regardless of gene function and regardless of a particular protein coding gene taxonomic population. Therefore, this variable can be considered a new codon usage index, whose use removes certain nonsignificant results found with the frequency statistic. This methodology identifies the common and rare codons (i.e., the codons having the highest and lowest occurrence) and leads to a model of codon evolution at three successive states: RNN, then RNY, and finally RYY. Some biological relations between this model and the YRY(N)6YRY preferential occurrence are also presented.

Base Sequence

Correlation between codon usage, regional genomic nucleotide composition, and amino acid composition in the cytochrome P-450 gene superfamily.

The codon usage bias of 110 mammalian cytochrome P-450 genes has been determined and analyzed in relation to a variety of genetic, biochemical, and physiological parameters. In those P-450 genes exhibiting biased usage the preferred codons generally do not differ among the four species examined (rat, rabbit, man, and mouse) or from the predominantly used codons identified for all sequenced genes in a recent data base analysis (Wada et al. (1992) Nucleic Acids Res. 20 (Suppl.), 2111-2118). Codon usage bias does not correlate with evolutionary relationships, evolutionary age, or with the extent of evolutionary conservation of orthologous proteins; there is no obvious correlation with the level of expression of a given P-450, with its inducibility, nor with its physiologic role; and neither the preferred codons nor the degree of bias differ for P-450s expressed in different tissues. Codon usage bias does correlate with the C+G content at the codon third position, and thus preferred codons usually end in C or G; for those P-450s for which gene sequences are available this bias also correlates with the C + G content of the intronic and flanking regions of these genes. Moreover, a lesser increase in the C + G content at the codon first and second positions is also evident in genes located in regions of high C + G content; this leads to predictable differences in the amino acid compositions of P-450 enzymes that correlate with genomic nucleotide composition and the degree of bias in codon usage.

Amino Acids

Translational efficiency of the Escherichia coli adenylate cyclase gene: mutating the UUG initiation codon to GUG or AUG results in increased gene expression.

Roy et al. [Roy, A., Haziza, C. & Danchin, A. (1983) EMBO J. 2, 791-797] established that translation of Escherichia coli adenylate cyclase initiates at a UUG codon, and they suggested this might decrease the efficiency of translation. We investigated the effect of varying the initiation codon on the expression of the adenylate cyclase (cya) gene. Using oligonucleotide-directed mutagenesis, we changed the UUG initiation codon to GUG and the more common initiator AUG and assayed for cya gene expression in a number of ways. First, the GUG initiation codon, in place of UUG, doubled cya expression when cya was expressed from the dual cya P1/P2 promoters. The corresponding AUG codon construct was nonviable. Second, when the cya gene was placed under the transcriptional control of the thermoinducible phage lambda PL promoter, the relative amounts of cya gene product were 1:2:6 for the UUG, GUG, and AUG initiation codons, respectively. Finally, the cya P2 promoter, Shine-Dalgarno sequence, and the DNA corresponding to the first 86 codons of cya were fused to DNA encoding the E. coli galactokinase gene beginning at the second codon. The relative amounts of the fusion polypeptides, which had galactokinase activity, were 1:2:3 for the UUG, GUG, and AUG initiation codons, respectively. These results demonstrate that the cya UUG initiation codon limits cya expression at the level of translation.

Adenylyl Cyclases