Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Translation rate modification by preferential codon usage: intragenic position effects.

We present a model for calculating the protein production rate as a function of the translation rate. The model takes into account that the elongation rate along an mRNA molecule is non-uniform as a result of different tRNA availabilities for different codons. Initiation of ribosomes on an mRNA is normally the rate-limiting step in the translation process, and blocking of the initiation site can be avoided if the codons closest to this site allow fast translation by the ribosome. Hence, different selective forces may act on the choice of synonymous codons in the initiation region than elsewhere on a given mRNA. We show that the elongation rate along the whole mRNA influences the production rate of abundant proteins, whereas only the elongation rate in the initiation region is of importance for the production rate of rare proteins. We also present an analysis of the codon distribution along known mRNAs coding for abundant and rare proteins.

Bacterial Proteins↗

A cluster of vitellogenin genes in the Mediterranean fruit fly Ceratitis capitata: sequence and structural conservation in dipteran yolk proteins and their genes.

Four genes encoding the major egg yolk polypeptides of the Mediterranean fruit fly Ceratitis capitata, vitellogenins 1 and 2 (VG1 and VG2), were cloned, characterized and partially sequenced. The genes are located on the same region of chromosome 5 and are organized in pairs, each encoding the two polypeptides on opposite DNA strands. Restriction and nucleotide sequence analysis indicate that the gene pairs have arisen from an ancestral pair by a relatively recent duplication event. The transcribed part is very similar to that of the Drosophila melanogaster yolk protein genes Yp1, Yp2 and Yp3. The Vg1 genes have two introns at the same positions as those in D. melanogaster Yp3; the Vg2 genes have only one of the introns, as do D. melanogaster Yp1 and Yp2. Comparison of the five polypeptide sequences shows extensive homology, with 27% of the residues being invariable. The sequence similarity of the processed proteins extends in two regions separated by a nonconserved region of varying size. Secondary structure predictions suggest a highly conserved secondary structure pattern in the two regions, which probably correspond to structural and functional domains. The carboxy-end domain of the C. capitata proteins shows the same sequence similarities with triacyglycerol lipases that have been reported previously for the D. melanogaster yolk proteins. Analysis of codon usage shows significant differences between D. melanogaster and C. capitata vitellogenins with the latter exhibiting a less biased representation of synonymous codons.

Alleles↗

Codon usage in Cryptosporidium parvum differs from that in other Eimeriorina.

Codon usage of Crytosporidium parvum was compared with those of other Eimeriorina Toxoplasma gondii and Eimeria tenella and revealed a biased use of synonymous codons with a preference for NNU (40.0%) and NNA (33.4%). There was no close resemblance of the codon usage of C. parvum to T. gondii (correlation coefficient, r = 0.14) or E. tenella (r = 0.14) but it was similar to Entamoeba histolytica (r = 0.75) and Plasmodium falciparum (r = 0.5). Analysis of the codon usage in homologous gene sequences (actin, beta-tubulin) also failed to reveal a close relationship between C. parvum and T. gondii or E. tenella. The low usage codons in C. parvum were most frequently used codons in T. gondii and E. tenella. These observations are consistent with 18S rRNA sequence analysis which shows no close relationship of Cryptosporidium with other Eimeriorina (Sarcocystis, Toxoplasma and Eimeria) and questions the validity of the current classification of C. parvum.

Actins↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

Studies of codon usage and tRNA genes of 18 unicellular organisms and quantification of Bacillus subtilis tRNAs: gene expression level and species-specific diversity of codon usage based on multivariate analysis.

We examined codon usage in Bacillus subtilis genes by multivariate analysis, quantified its cellular levels of individual tRNAs, and found a clear constraint of tRNA contents on synonymous codon choice. Individual tRNA levels were proportional to the copy number of the respective tRNA genes. This indicates that the tRNA gene copy number is an important factor to determine in cellular tRNA levels, which is common with Escherichia coli and yeast Saccharomyces cerevisiae. Codon usage in 18 unicellular organisms whose genomes have been sequenced completely was analyzed and compared with the composition of tRNA genes. The 18 organisms are as follows: yeast S. cerevisiae, Aquifex aeolicus, Archaeoglobus fulgidus, B. subtilis, Borrelia burgdorferi, Chlamydia trachomatis, E. coli, Haemophilus influenzae, Helicobacterpylori, Methanococcusjannaschii, Methanobacterium thermoautotrophicum, Mycobacterium tuberculosis, Mycoplasma genitalium, Mycoplasma pneumoniae, Pyrococcus horikoshii, Rickettsia prowazekii, Synechocystis sp., and Treponema pallidum. Codons preferred in highly expressed genes were related to the codons optimal for the translation process, which were predicted by the composition of isoaccepting tRNA genes. Genes with specific codon usage are discussed in connection with their evolutionary origins and functions. The origin and terminus of replication could be predicted on the basis of codon usage when the usage was analyzed relative to the transcription direction of individual genes.

Archaea↗

Codon usage and bias among individual genes of the coccidia and piroplasms.

Codon usage has been analysed in individual gene sequences, derived from a variety of parasitic protozoa in the class Sporozoa of the phylum Apicomplexa using metric multidimensional scaling. The two groups of codon usage patterns detected reflect the two main subgroups of organisms studied (the coccidia and the piroplasms), and it is the pattern of usage of synonymous codons that has the largest influence on overall codon usage in the individual genes, rather than being the pattern of amino acid composition of the gene product. The magnitude of the codon usage bias in the sequences was determined using three commonly used indices-NC, GC3S and B. In general, although relatively low levels of codon usage bias were detected in these gene sequences, codon usage bias does explain at least some of the codon usage patterns observed. Codon usage bias was observed to be dependent on the overall base composition of the genes analysed, which in turn was reflected in the types of codons that were either over- or under-represented in the nucleotide sequences. In keeping with observations on prokaryotic organisms, it is speculated that the codon usage patterns detected in these parasitic protozoa are the result of directional mutation pressure on the base composition of the genomic DNA.

Animals↗

Codon usage and base composition in Rickettsia prowazekii.

Codon usage and base composition in sequences from the A + T-rich genome of Rickettsia prowazekii, a member of the alpha Proteobacteria, have been investigated. Synonymous codon usage patterns are roughly similar among genes, even though the data set includes genes expected to be expressed at very different levels, indicating that translational selection has been ineffective in this species. However, multivariate statistical analysis differentiates genes according to their G + C contents at the first two codon positions. To study this variation, we have compared the amino acid composition patterns of 21 R. prowazekii proteins with that of a homologous set of proteins from Escherichia coli. The analysis shows that individual genes have been affected by biased mutation rates to very different extents: genes encoding proteins highly conserved among other species being the least affected. Overall, protein coding and intergenic spacer regions have G + C content values of 32.5% and 21.4%, respectively. Extrapolation from these values suggests that R. prowazekii has around 800 genes and that 60-70% of the genome may be coding.

Base Composition↗

Protein-encoding genes in the sulfothermophilic archaea Sulfolobus and Pyrococcus.

A number of unrelated protein-encoding genes from sulfothermophilic archaea, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Pyrococcus furiosus and Pyrococcus woesei, has been analyzed. In the Sulfolobus genus, the content of A + T is significantly higher than that of C + G and the base usage follows the order, A > T > G > C. In Pyrococcus, the A + T content is also higher than that of C + G, but with lower values; in the order of base usage, G precedes T. The codon usage of these sulfothermophiles has been determined; alternative start codons are frequently used in both genera; codon preferences reflect the rich A + T composition of the corresponding genomes; for both genera the codon bias is particularly evident within the different arginine triplets, where AGA and AGG are predominant. From the similarities in the codon usage, close taxonomic relationships become evident within the Sulfolobus or the Pyrococcus genus; a lower, but significant similarity is also clear between these genera. The synonymous codon usage of these sulfothermophiles shows similarities with that of Saccharomyces cerevisiae and bovine mitochondria, whereas clear divergences are observed with the halophilic archaeal genus, Halobacterium, or the eubacterium, Escherichia coli. The unrelated proteins of the considered sulfothermophiles have been analyzed for the content of hydrophobic residues; the comparison with mesophiles reveals a significant increase in the average hydrophobicity of amino acid residues. This finding could indicate a mechanism of adaptation of proteins in organisms living under extreme environments. It is noteworthy that an opposite trend, i.e. a decreased average hydrophobicity, occurs in unrelated halophilic proteins.

Animals↗

Distribution and evolution of sequence characteristics in the E. coli genome.

The mean (G + C) composition (51.0%) and standard deviation (+/- 3.8%) of published DNA sequences accounting for 10% of the E. coli genome is in excellent agreement with the principal overall distribution determined by high resolution melting. While differences in base and neighbor characteristics are small and uniform throughout all regions of the genome, it is found that the (G + C) content of sequences varies in segmented fashion within boundaries corresponding to coding (53% G + C) and noncoding (46% G + C) regions; with variances in the latter being six-fold greater than in coding regions. The variance in different regions shows a strong negative dependence on (G + C) content of the region, reflecting the condition that A-T and G-C base pairs are preferred neighbors of A-T and C-G pairs, respectively; with the bias increasing with decreasing (G + C) content. Neighbor analysis indicates the most extreme positive biases occur in AA, TT, GC and CG throughout all regions, but particularly in noncoding regions. Extraordinary numbers of oligomeric strings of (A)n, etc., are the further consequence of this bias. These and other characteristics point to the existence of inherent biases in neighbor frequencies levied during replication or repair, and which reflect, in turn, neighbor influences during mutation. The bias in codon usage noted by Grantham and others is seen here as due, in part, to the adaptation of coding sequences to this microenvironment through selection among synonymous codons so as to preserve inherent neighbor biases.

Base Composition↗

Does the 'non-coding' strand code?

The hypothesis that DNA strands complementary to the coding strand contain in phase coding sequences has been investigated. Statistical analysis of the 50 genes of bacteriophage T7 shows no significant correlation between patterns of codon usage on the coding and non-coding strands. In Bacillus and yeast genes the correlation observed is not different from that expected with random synonymous codon usage, while a high correlation seen in 52 E. coli genes can be explained in terms of an excess of RNY codons. A deficiency of UUA, CUA and UCA codons (complementary to termination) seems to be restricted to the E. coli genes, and may be due to low abundance of the relevant cognate tRNA species. Thus the analysis shows that the non-coding strand has the properties expected of a sequence complementary to a coding strand, with no indications that it encodes, or may have encoded, proteins.

Bacillus↗

Transgene sequence codon optimization and composition determines replication competence of self-amplifying RNA.

Self-amplifying RNA (saRNA) is an emerging RNA therapeutic modality that can facilitate higher magnitude and more durable protein expression at substantially lower doses than nonreplicating mRNA. Unlike conventional messenger RNA (mRNA), alphavirus-derived saRNA must support a replicase-driven RNA amplification step in addition to translation, raising the possibility that transgene coding sequences impose sequence-level constraints on replication. Here, saRNA replication was found to be dependent on the codon composition of the transgene; multiple therapeutic transgenes were replication defective despite an intact Venezuelan Equine Encephalitis Virus (VEEV)-derived saRNA backbone. Replication defects were rescued by synonymous codon re-optimization of the same transgenes, indicating that nucleotide-level features of the coding sequence, rather than the encoded protein, govern replication competence. Comparative compositional analyses identified a distinct signature associated with productive replication, characterized by elevated GC (>53%) and GC3 (>63%) content, higher codon adaptation to human (>0.75), and reduced UpA (<43/kb) and UpU (<41/kb) dinucleotide density. Moreover, deliberate compositional perturbation of an otherwise replication-competent transgene shifted these features and abolished replication, supporting a causal and combinatorial role for sequence composition in defining saRNA replication outcome. These findings define an underappreciated constraint in saRNA therapeutics and motivate saRNA-specific payload design frameworks that incorporate alphavirus-associated compositional biases during transgene sequence optimization.

Codon↗

Construction of genetic code from evolutionary stability.

The construction of the genetic code is investigated based on a stability principle. The concept and formulation of mutational deterioration (MD) of the genetic code is proposed. It is proved that the degeneracies of codon multiplets obey the rule to best resist MD. The MD for each ideal multiplet of codons is expressed by four parameters and it takes on a minimum value for real distributions of codons in the multiplet. Then the global mutational deterioration (GMD) of code table is calculated and the minimal code is deduced. The domain-like distribution of hydrophobic and hydrophilic amino acids on the genetic code is explained from the minimization of GMD. It is demonstrated that the standard code is approximately GMD-minimal. By introducing some constraints that are related to the initial condition of the system, we have deduced the standard genetic code from the minimization of GMD. The minimization shows the general trend of evolutionary process to some stable state while the constraints reflect a 'frozen accident.' Many deviant codon assignments are also explained through MD minimization assuming the changeable degrees of degeneracies for some multiplets. So, a possible answer to the question of "Why are synonymous codons and amino acids distributed in the code table just as they are?" is given.

Biological Evolution↗

Codon frequencies in 119 individual genes confirm consistent choices of degenerate bases according to genome type.

The poor printing of our previous Figure 2 (1) is corrected. Codon usage in mRNA sequences just published is also given. A new correspondence analysis is done, based on simultaneous comparison in all mRNA of use of the 61 codons. This analysis reinforces our claim that most genes in a genome, or genome type, have the same coding strategy; that is, they show similar choices among synonymous codons, or among degenerate bases (2). Like analysis on frequency variation in the amino acids coded reveals an entirely different pattern.

Animals↗

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence↗

Preferential codon usage in prokaryotic genes: the optimal codon-anticodon interaction energy and the selective codon usage in efficiently expressed genes.

By considering the nucleotide sequence of several highly expressed coding regions in bacteriophage MS2 and mRNAs from Escherichia coli, it is possible to deduce some rules which govern the selection of the most appropriate synonymous codons NNU or NNC read by tRNAs having GNN, QNN or INN as anticodon. The rules fit with the general hypothesis that an efficient in-phase translation is facilitated by proper choice of degenerate codewords promoting a codon-anticodon interaction with intermediate strength (optimal energy) over those with very strong or very weak interaction energy. Moreover, codons corresponding to minor tRNAs are clearly avoided in these efficiently expressed genes. These correlations are clearcut in the normal reading frame but not in the corresponding frameshift sequences +1 and +2. We hypothesize that both the optimization of codon-anticodon interaction energy and the adaptation of the population to codon frequency or vice versa in highly expressed mRNAs of E. coli are part of a strategy that optimizes the efficiency of translation. Conversely, codon usage in weakly expressed genes such as repressor genes follows exactly the opposite rules. It may be concluded that, in addition to the need for coding an amino acid sequence, the energetic consideration for codon-anticodon pairing, as well as the adaptation of codons to the tRNA population, may have been important evolutionary constraints on the selection of the optimal nucleotide sequence.

Anticodon↗

Codon usage in regulatory genes in Escherichia coli does not reflect selection for 'rare' codons.

It has often been suggested that differential usage of codons recognized by rare tRNA species, i.e. "rare codons", represents an evolutionary strategy to modulate gene expression. In particular, regulatory genes are reported to have an extraordinarily high frequency of rare codons. From E. coli we have compiled codon usage data for highly expressed genes, moderately/lowly expressed genes, and regulatory genes. We have identified a clear and general trend in codon usage bias, from the very high bias seen in very highly expressed genes and attributed to selection, to a rather low bias in other genes which seems to be more influenced by mutation than by selection. There is no clear tendency for an increased frequency of rare codons in the regulatory genes, compared to a large group of other moderately/lowly expressed genes with low codon bias. From this, as well as a consideration of evolutionary rates of regulatory genes, and of experimental data on translation rates, we conclude that the pattern of synonymous codon usage in regulatory genes reflects primarily the relaxation of natural selection.

Base Sequence↗

Intercodon dinucleotides affect codon choice in plant genes.

In this work, 710 CDSs corresponding to over 290 000 codons equally distributed between Brassica napus, Arabidopsis thaliana, Lycopersicon esculentum, Nicotiana tabacum, Pisum sativum, Glycine max, Oryza sativa, Triticum aestivum, Hordeum vulgare and Zea mays were considered. For each amino acid, synonymous codon choice was determined in the presence of A, G, C or T as the initial nucleotide of the subsequent triplet; data were statistically analysed under the hypothesis of an independent assortment of codons. In 33.4% of cases, a frequency significantly (P: = 0.01) different from that expected was recorded. This was mainly due to a pervasive intercodon TpA and CpG deficiency. As a general rule, intercodon TpAs and CpGs were preferably replaced by CpAs and TpGs, respectively. In several instances, codon frequencies were also modified to avoid homotetramer and homotrimer formation, to reduce intercodon ApCs downstream (1,2) GG or AG dinucleotides, as well as to increase GpA or ApG intercodons under certain contexts. Since TpA, CpG and homotetra(tri)mer deficiency directly or indirectly accounted for 77% of significant variation in the codon frequency, it can be concluded that codon usage mirrors precise needs at the DNA structure level. Plant species exhibited a phylogenetically-related adaptation to structural constraints. Codon usage flexibility was reflected in strikingly different arrays of optimum codons for probe design.

Base Composition↗

Codon usage in plant genes.

We have examined codon bias in 207 plant gene sequences collected from Genbank and the literature. When this sample was further divided into 53 monocot and 154 dicot genes, the pattern of relative use of synonymous codons was shown to differ between these taxonomic groups, primarily in the use of G + C in the degenerate third base. Maize and soybean codon bias were examined separately and followed the monocot and dicot codon usage patterns respectively. Codon preference in ribulose 1,5 bisphosphate and chlorophyll a/b binding protein, two of the most abundant proteins in leaves was investigated. These highly expressed are more restricted in their codon usage than plant genes in general.

Amino Acid Sequence↗