Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Codon usage in plant peroxidase genes.

Codon preference and asymmetry in usage in the DNA sequences encoding the mature enzyme protein of 24 plant peroxidases from 12 different species were examined. Codon usage in highly conserved/non-conserved areas of the sequences was analysed, as well as possible deficiency/excess in CpG dinucleotides in the pairs of codon positions. Sequence relationships displayed by overall codon usage, dinucleotide frequencies within codons, and amino acid sequences were also studied. The main findings were: (1) Monocots clustered separately from dicots for overall codon usage and dinucleotide frequencies in codon positions 2 and 3, with six and seven clusters respectively discernible among these 24 peroxidase sequences. The monocot/dicot distinction disappeared in the four clusters among the mature protein amino acid sequences. Overall codon usage in sequences from monocotyledon and dicotyledon species differed, the monocots favouring codons with C or G in the third position. (2) Codon usage was biassed in many sequences, asymmetry was particularly noticeable in the monocots. (3) For repeated amino acids within conserved areas, codon preference appeared dependent on the order in which the repeated amino acid occurred, so that its usage of synonymous codons frequently balanced out.

Amino Acid Sequence↗

How optimized is the translational machinery in Escherichia coli, Salmonella typhimurium and Saccharomyces cerevisiae?

The optimization of the translational machinery in cells requires the mutual adaptation of codon usage and tRNA concentration, and the adaptation of tRNA concentration to amino acid usage. Two predictions were derived based on a simple deterministic model of translation which assumes that elongation of the peptide chain is rate-limiting. The highest translational efficiency is achieved when the codon recognized by the most abundant tRNA reaches the maximum frequency. For each codon family, the tRNA concentration is optimally adapted to codon usage when the concentration of different tRNA species matches the square-root of the frequency of their corresponding synonymous codons. When tRNA concentration and codon usage are well adapted to each other, the optimal content of all tRNA species carrying the same amino acid should match the square-root of the frequency of the amino acid. These predictions are examined against empirical data from Escherichia coli, Salmonella typhimurium, and Saccharomyces cerevisiae.

Codon↗

Codon usage in nucleopolyhedroviruses.

Phylogenetic analyses based on baculovirus polyhedrin nucleotide and amino acid sequences revealed two major nucleopolyhedrovirus (NPV) clades, designated Group I and Group II. Subsequent phylogenetic analyses have revealed three Group II subclades, designated A, B and C. Variations in amino acid frequencies determine the extent of dissimilarity for divergent but structurally and functionally conserved genes and therefore significantly influence the analysis of phylogenetic relationships. Hence, it is important to consider variations in amino acid codon usage. The Genome Hypothesis postulates that genes in any given genome use the same coding pattern with respect to synonymous codons and that genes in phylogenetically related species generally show the same pattern of codon usage. We have examined codon usage in six genes from six NPVs and found that: (1) there is significant variation in codon use by genes within the same virus genome; (2) there is significant variation in the codon usage of homologous genes encoded by different NPVs; (3) there is no correlation between the level of gene expression and codon bias in NPVs; (4) there is no correlation between gene length and codon bias in NPVs; and (5) that while codon use bias appears to be conserved between viruses that are closely related phylogenetically, the patterns of codon usage also appear to be a direct function of the GC-content of the virus-encoded genes.

Animals↗

Genetic code and optimal resistance to the effects of mutations.

This paper deals with the notion of resistance of the genetic code to the effects of mutations. We measure the resistance of a group of t codons as the number of pairs of those which differ from each other in only one of their three bases. We find for each value of t the maximum possible value of the resistance and we describe some groups of codons giving this value. Important examples of such configurations are found in the genetic code, among these are the groups of synonymous codons, as observed elsewhere, and the cluster of codons which have an hydrophobic amino acid for translation.

Base Sequence↗

Codon usage bias and base composition in MHC genes in humans and common chimpanzees.

Codon bias and base composition in major histocompatibility complex (MHC) sequences have been studied for both class I and II loci in Homo sapiens and Pan troglodytes. There is low to moderate codon bias for the MHC of humans and chimpanzees. In the class I loci, the same level of moderate codon bias is seen for HLA-B, HLA-C, Patr-A, Patr-B, and Patr-C, while at HLA-A the level of codon bias is lower. There is a correlation between codon usage bias and G+C content in the A and B loci in humans and chimps, but not at the C locus. To examine the effect of diversifying selection on codon bias, we subdivided class I alleles into antigen recognition site (ARS) and non-ARS codons. ARS codons had lower bias than non-ARS codons. This may indicate that the constraint of codon bias on nucleotide substitution may be selected against in ARS codons. At the class II loci, there are distinct differences between alpha and beta chain genes with respect to codon usage, with the beta chain genes being much more biased. Species-specific differences in base composition were seen in exon 2 at the DRB1 locus, with lower GC content in chimpanzees. Considering the complex evolutionary history of MHC genes, the study of codon usage patterns provides us with a better understanding of both the evolutionary history of these genes and the evolution of synonymous codon usage in genes under natural selection.

Animals↗

Translation rate modification by preferential codon usage: intragenic position effects.

We present a model for calculating the protein production rate as a function of the translation rate. The model takes into account that the elongation rate along an mRNA molecule is non-uniform as a result of different tRNA availabilities for different codons. Initiation of ribosomes on an mRNA is normally the rate-limiting step in the translation process, and blocking of the initiation site can be avoided if the codons closest to this site allow fast translation by the ribosome. Hence, different selective forces may act on the choice of synonymous codons in the initiation region than elsewhere on a given mRNA. We show that the elongation rate along the whole mRNA influences the production rate of abundant proteins, whereas only the elongation rate in the initiation region is of importance for the production rate of rare proteins. We also present an analysis of the codon distribution along known mRNAs coding for abundant and rare proteins.

Bacterial Proteins↗

A cluster of vitellogenin genes in the Mediterranean fruit fly Ceratitis capitata: sequence and structural conservation in dipteran yolk proteins and their genes.

Four genes encoding the major egg yolk polypeptides of the Mediterranean fruit fly Ceratitis capitata, vitellogenins 1 and 2 (VG1 and VG2), were cloned, characterized and partially sequenced. The genes are located on the same region of chromosome 5 and are organized in pairs, each encoding the two polypeptides on opposite DNA strands. Restriction and nucleotide sequence analysis indicate that the gene pairs have arisen from an ancestral pair by a relatively recent duplication event. The transcribed part is very similar to that of the Drosophila melanogaster yolk protein genes Yp1, Yp2 and Yp3. The Vg1 genes have two introns at the same positions as those in D. melanogaster Yp3; the Vg2 genes have only one of the introns, as do D. melanogaster Yp1 and Yp2. Comparison of the five polypeptide sequences shows extensive homology, with 27% of the residues being invariable. The sequence similarity of the processed proteins extends in two regions separated by a nonconserved region of varying size. Secondary structure predictions suggest a highly conserved secondary structure pattern in the two regions, which probably correspond to structural and functional domains. The carboxy-end domain of the C. capitata proteins shows the same sequence similarities with triacyglycerol lipases that have been reported previously for the D. melanogaster yolk proteins. Analysis of codon usage shows significant differences between D. melanogaster and C. capitata vitellogenins with the latter exhibiting a less biased representation of synonymous codons.

Alleles↗

Codon usage in Cryptosporidium parvum differs from that in other Eimeriorina.

Codon usage of Crytosporidium parvum was compared with those of other Eimeriorina Toxoplasma gondii and Eimeria tenella and revealed a biased use of synonymous codons with a preference for NNU (40.0%) and NNA (33.4%). There was no close resemblance of the codon usage of C. parvum to T. gondii (correlation coefficient, r = 0.14) or E. tenella (r = 0.14) but it was similar to Entamoeba histolytica (r = 0.75) and Plasmodium falciparum (r = 0.5). Analysis of the codon usage in homologous gene sequences (actin, beta-tubulin) also failed to reveal a close relationship between C. parvum and T. gondii or E. tenella. The low usage codons in C. parvum were most frequently used codons in T. gondii and E. tenella. These observations are consistent with 18S rRNA sequence analysis which shows no close relationship of Cryptosporidium with other Eimeriorina (Sarcocystis, Toxoplasma and Eimeria) and questions the validity of the current classification of C. parvum.

Actins↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

Studies of codon usage and tRNA genes of 18 unicellular organisms and quantification of Bacillus subtilis tRNAs: gene expression level and species-specific diversity of codon usage based on multivariate analysis.

We examined codon usage in Bacillus subtilis genes by multivariate analysis, quantified its cellular levels of individual tRNAs, and found a clear constraint of tRNA contents on synonymous codon choice. Individual tRNA levels were proportional to the copy number of the respective tRNA genes. This indicates that the tRNA gene copy number is an important factor to determine in cellular tRNA levels, which is common with Escherichia coli and yeast Saccharomyces cerevisiae. Codon usage in 18 unicellular organisms whose genomes have been sequenced completely was analyzed and compared with the composition of tRNA genes. The 18 organisms are as follows: yeast S. cerevisiae, Aquifex aeolicus, Archaeoglobus fulgidus, B. subtilis, Borrelia burgdorferi, Chlamydia trachomatis, E. coli, Haemophilus influenzae, Helicobacterpylori, Methanococcusjannaschii, Methanobacterium thermoautotrophicum, Mycobacterium tuberculosis, Mycoplasma genitalium, Mycoplasma pneumoniae, Pyrococcus horikoshii, Rickettsia prowazekii, Synechocystis sp., and Treponema pallidum. Codons preferred in highly expressed genes were related to the codons optimal for the translation process, which were predicted by the composition of isoaccepting tRNA genes. Genes with specific codon usage are discussed in connection with their evolutionary origins and functions. The origin and terminus of replication could be predicted on the basis of codon usage when the usage was analyzed relative to the transcription direction of individual genes.

Archaea↗

Codon usage and bias among individual genes of the coccidia and piroplasms.

Codon usage has been analysed in individual gene sequences, derived from a variety of parasitic protozoa in the class Sporozoa of the phylum Apicomplexa using metric multidimensional scaling. The two groups of codon usage patterns detected reflect the two main subgroups of organisms studied (the coccidia and the piroplasms), and it is the pattern of usage of synonymous codons that has the largest influence on overall codon usage in the individual genes, rather than being the pattern of amino acid composition of the gene product. The magnitude of the codon usage bias in the sequences was determined using three commonly used indices-NC, GC3S and B. In general, although relatively low levels of codon usage bias were detected in these gene sequences, codon usage bias does explain at least some of the codon usage patterns observed. Codon usage bias was observed to be dependent on the overall base composition of the genes analysed, which in turn was reflected in the types of codons that were either over- or under-represented in the nucleotide sequences. In keeping with observations on prokaryotic organisms, it is speculated that the codon usage patterns detected in these parasitic protozoa are the result of directional mutation pressure on the base composition of the genomic DNA.

Animals↗

Codon usage and base composition in Rickettsia prowazekii.

Codon usage and base composition in sequences from the A + T-rich genome of Rickettsia prowazekii, a member of the alpha Proteobacteria, have been investigated. Synonymous codon usage patterns are roughly similar among genes, even though the data set includes genes expected to be expressed at very different levels, indicating that translational selection has been ineffective in this species. However, multivariate statistical analysis differentiates genes according to their G + C contents at the first two codon positions. To study this variation, we have compared the amino acid composition patterns of 21 R. prowazekii proteins with that of a homologous set of proteins from Escherichia coli. The analysis shows that individual genes have been affected by biased mutation rates to very different extents: genes encoding proteins highly conserved among other species being the least affected. Overall, protein coding and intergenic spacer regions have G + C content values of 32.5% and 21.4%, respectively. Extrapolation from these values suggests that R. prowazekii has around 800 genes and that 60-70% of the genome may be coding.

Base Composition↗

Protein-encoding genes in the sulfothermophilic archaea Sulfolobus and Pyrococcus.

A number of unrelated protein-encoding genes from sulfothermophilic archaea, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Pyrococcus furiosus and Pyrococcus woesei, has been analyzed. In the Sulfolobus genus, the content of A + T is significantly higher than that of C + G and the base usage follows the order, A > T > G > C. In Pyrococcus, the A + T content is also higher than that of C + G, but with lower values; in the order of base usage, G precedes T. The codon usage of these sulfothermophiles has been determined; alternative start codons are frequently used in both genera; codon preferences reflect the rich A + T composition of the corresponding genomes; for both genera the codon bias is particularly evident within the different arginine triplets, where AGA and AGG are predominant. From the similarities in the codon usage, close taxonomic relationships become evident within the Sulfolobus or the Pyrococcus genus; a lower, but significant similarity is also clear between these genera. The synonymous codon usage of these sulfothermophiles shows similarities with that of Saccharomyces cerevisiae and bovine mitochondria, whereas clear divergences are observed with the halophilic archaeal genus, Halobacterium, or the eubacterium, Escherichia coli. The unrelated proteins of the considered sulfothermophiles have been analyzed for the content of hydrophobic residues; the comparison with mesophiles reveals a significant increase in the average hydrophobicity of amino acid residues. This finding could indicate a mechanism of adaptation of proteins in organisms living under extreme environments. It is noteworthy that an opposite trend, i.e. a decreased average hydrophobicity, occurs in unrelated halophilic proteins.

Animals↗

Distribution and evolution of sequence characteristics in the E. coli genome.

The mean (G + C) composition (51.0%) and standard deviation (+/- 3.8%) of published DNA sequences accounting for 10% of the E. coli genome is in excellent agreement with the principal overall distribution determined by high resolution melting. While differences in base and neighbor characteristics are small and uniform throughout all regions of the genome, it is found that the (G + C) content of sequences varies in segmented fashion within boundaries corresponding to coding (53% G + C) and noncoding (46% G + C) regions; with variances in the latter being six-fold greater than in coding regions. The variance in different regions shows a strong negative dependence on (G + C) content of the region, reflecting the condition that A-T and G-C base pairs are preferred neighbors of A-T and C-G pairs, respectively; with the bias increasing with decreasing (G + C) content. Neighbor analysis indicates the most extreme positive biases occur in AA, TT, GC and CG throughout all regions, but particularly in noncoding regions. Extraordinary numbers of oligomeric strings of (A)n, etc., are the further consequence of this bias. These and other characteristics point to the existence of inherent biases in neighbor frequencies levied during replication or repair, and which reflect, in turn, neighbor influences during mutation. The bias in codon usage noted by Grantham and others is seen here as due, in part, to the adaptation of coding sequences to this microenvironment through selection among synonymous codons so as to preserve inherent neighbor biases.

Base Composition↗

Does the 'non-coding' strand code?

The hypothesis that DNA strands complementary to the coding strand contain in phase coding sequences has been investigated. Statistical analysis of the 50 genes of bacteriophage T7 shows no significant correlation between patterns of codon usage on the coding and non-coding strands. In Bacillus and yeast genes the correlation observed is not different from that expected with random synonymous codon usage, while a high correlation seen in 52 E. coli genes can be explained in terms of an excess of RNY codons. A deficiency of UUA, CUA and UCA codons (complementary to termination) seems to be restricted to the E. coli genes, and may be due to low abundance of the relevant cognate tRNA species. Thus the analysis shows that the non-coding strand has the properties expected of a sequence complementary to a coding strand, with no indications that it encodes, or may have encoded, proteins.

Bacillus↗

Transgene sequence codon optimization and composition determines replication competence of self-amplifying RNA.

Self-amplifying RNA (saRNA) is an emerging RNA therapeutic modality that can facilitate higher magnitude and more durable protein expression at substantially lower doses than nonreplicating mRNA. Unlike conventional messenger RNA (mRNA), alphavirus-derived saRNA must support a replicase-driven RNA amplification step in addition to translation, raising the possibility that transgene coding sequences impose sequence-level constraints on replication. Here, saRNA replication was found to be dependent on the codon composition of the transgene; multiple therapeutic transgenes were replication defective despite an intact Venezuelan Equine Encephalitis Virus (VEEV)-derived saRNA backbone. Replication defects were rescued by synonymous codon re-optimization of the same transgenes, indicating that nucleotide-level features of the coding sequence, rather than the encoded protein, govern replication competence. Comparative compositional analyses identified a distinct signature associated with productive replication, characterized by elevated GC (>53%) and GC3 (>63%) content, higher codon adaptation to human (>0.75), and reduced UpA (<43/kb) and UpU (<41/kb) dinucleotide density. Moreover, deliberate compositional perturbation of an otherwise replication-competent transgene shifted these features and abolished replication, supporting a causal and combinatorial role for sequence composition in defining saRNA replication outcome. These findings define an underappreciated constraint in saRNA therapeutics and motivate saRNA-specific payload design frameworks that incorporate alphavirus-associated compositional biases during transgene sequence optimization.

Codon↗

Construction of genetic code from evolutionary stability.

The construction of the genetic code is investigated based on a stability principle. The concept and formulation of mutational deterioration (MD) of the genetic code is proposed. It is proved that the degeneracies of codon multiplets obey the rule to best resist MD. The MD for each ideal multiplet of codons is expressed by four parameters and it takes on a minimum value for real distributions of codons in the multiplet. Then the global mutational deterioration (GMD) of code table is calculated and the minimal code is deduced. The domain-like distribution of hydrophobic and hydrophilic amino acids on the genetic code is explained from the minimization of GMD. It is demonstrated that the standard code is approximately GMD-minimal. By introducing some constraints that are related to the initial condition of the system, we have deduced the standard genetic code from the minimization of GMD. The minimization shows the general trend of evolutionary process to some stable state while the constraints reflect a 'frozen accident.' Many deviant codon assignments are also explained through MD minimization assuming the changeable degrees of degeneracies for some multiplets. So, a possible answer to the question of "Why are synonymous codons and amino acids distributed in the code table just as they are?" is given.

Biological Evolution↗

Codon frequencies in 119 individual genes confirm consistent choices of degenerate bases according to genome type.

The poor printing of our previous Figure 2 (1) is corrected. Codon usage in mRNA sequences just published is also given. A new correspondence analysis is done, based on simultaneous comparison in all mRNA of use of the 61 codons. This analysis reinforces our claim that most genes in a genome, or genome type, have the same coding strategy; that is, they show similar choices among synonymous codons, or among degenerate bases (2). Like analysis on frequency variation in the amino acids coded reveals an entirely different pattern.

Animals↗