Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Natural selection versus primitive gene structure as determinant of codon usage.

Different codons are not utilized equally in known gene sequences. One of the important biases of codon usage is observed in the form of an enrichment of RNY codons, especially within RNN codon families. Such biases could represent the residue of a primitive repeating-RNY gene structure, or the outcome of natural selection, or both. Analyses based on the rates of silent substitutions, the frequencies of base doublets, and synonymous codon ratios for Escherichia coli, yeast, Drosophila and Xenopus proteins have been performed. The results rule out any significant support for a primitive repeating-RNY or repeating-RRY gene structure, and establish the important role of natural selection in determining the choice of codons. With strong intervention by natural selection, the relationship between primitive gene structure and codon usage necessarily becomes minimal.

Animals↗

Heterogeneity in codon usages of sobemovirus genes.

When conventional phylogenetic trees were built using 14 genome sequences of 9 sobemoviruses, two main lineages were apparent: monocot-infecting viruses and dicot-infecting viruses. To investigate whether members of the genus Sobemovirus originated from monocot hosts or from dicot hosts, we constructed relationship trees based on Relative Synonymous Codon Usage (RSCU) of the viruses. The RSCU relationship trees grouped the monocot-infecting and dicot-infecting viruses even better than the genome phylogenetic trees. The RSCU approach also enabled direct comparisons among viral and host species. When host species were added into the RSCU tree, the viral species clustered with the monocot hosts, indicating codon usage homologies to monocots. The stability of the RSCU tree was improved when RSCU values were calculated for individual viral open reading frames (ORFs). Most interestingly, the codon usages of the viral ORF-2 that encodes the replicase showed affinity to that of the plants whereas codon usages of the other viral ORFs were not relevant to the host species. All ORF-2s from 3 monocot viruses and 4 out of 6 dicot viruses had greater RSCU affinities to sequences of ORFs in monocot than to dicot hosts, possibly indicating that ORF-2, and therefore the replicase module of sobemovirus has a monocot origin.

Arabidopsis↗

Divergence in codon usage of Lactobacillus species.

We have analyzed codon usage patterns of 70 sequenced genes from different Lactobacillus species. Codon usage in lactobacilli is highly biased. Both inter-species and intra-species heterogeneity of codon usage bias was observed. Codon usage in L. acidophilus is similar to that in L. helveticus, but dissimilar to that in L. bulgaricus, L. casei, L. pentosus and L. plantarum. Codon usage in the latter three organisms is not significantly different, but is different from that in L. bulgaricus. Inter-species differences in codon usage can, at least in part, be explained by differences in mutational drift. L. bulgaricus shows GC drift, whereas all other species show AT drift. L. acidophilus and L. helveticus rarely use NNG in family-box (a set of synonymous) codons, in contrast to all other species. This result may be explained by assuming that L. acidophilus and L. helveticus, but not other species examined, use a single tRNA species for translation of family-box codons. Differences in expression level of genes are positively correlated with codon usage bias. Highly expressed genes show highly biased codon usage, whereas weakly expressed genes show much less biased codon usage. Codon usage patterns at the 5'-end of Lactobacillus genes is not significantly different from that of entire genes. The GC content of codons 2-6 is significantly reduced compared with that of the remainder of the gene. The possible implications of a reduced GC content for the control of translation efficiency are discussed.

Base Sequence↗

Evidence for selection on synonymous mutations affecting stability of mRNA secondary structure in mammals.

BACKGROUND: In mammals, contrary to what is usually assumed, recent evidence suggests that synonymous mutations may not be selectively neutral. This position has proven contentious, not least because of the absence of a viable mechanism. Here we test whether synonymous mutations might be under selection owing to their effects on the thermodynamic stability of mRNA, mediated by changes in secondary structure. RESULTS: We provide numerous lines of evidence that are all consistent with the above hypothesis. Most notably, by simulating evolution and reallocating the substitutions observed in the mouse lineage, we show that the location of synonymous mutations is non-random with respect to stability. Importantly, the preference for cytosine at 4-fold degenerate sites, diagnostic of selection, can be explained by its effect on mRNA stability. Likewise, by interchanging synonymous codons, we find naturally occurring mRNAs to be more stable than simulant transcripts. Housekeeping genes, whose proteins are under strong purifying selection, are also under the greatest pressure to maintain stability. CONCLUSION: Taken together, our results provide evidence that, in mammals, synonymous sites do not evolve neutrally, at least in part owing to selection on mRNA stability. This has implications for the application of synonymous divergence in estimating the mutation rate.

Amino Acid Substitution↗

The problem of counting sites in the estimation of the synonymous and nonsynonymous substitution rates: implications for the correlation between the synonymous substitution rate and codon usage bias.

Most methods for estimating the rate of synonymous and nonsynonymous substitution per site define a site as a mutational opportunity: the proportion of sites that are synonymous is equal to the proportion of mutations that would be synonymous under the model of evolution being considered. Here we demonstrate that this definition of a site can give misleading results and that a physical definition of site should be used in some circumstances. We illustrate our point by reexamining the relationship between codon usage bias and the synonymous substitution rate. It has recently been shown that the rate of synonymous substitution, calculated using the Goldman-Yang method, which encapsulates the mutational-opportunity definition of a site at a high level of sophistication, is either positively correlated or uncorrelated to synonymous codon bias in Drosophila. Using other methods, which account for synonymous codon bias but define a site physically, we show that there is a negative correlation between the synonymous substitution rate and codon bias and that the lack of a negative correlation using the Goldman-Yang method is due to the way in which the number of synonymous sites is counted. We also show that there is a positive correlation between the synonymous substitution rate and third position GC content in mammals, but that the relationship is considerably weaker than that obtained using the Goldman-Yang method. We argue that the Goldman-Yang method is misleading in this context and conclude that methods that rely on a mutational-opportunity definition of a site should be used with caution.

Animals↗

An analysis of the codon usage of Pasteurella haemolytica A1.

Analysis of approximately 17 kbp of nucleotide sequences from three different regions of the genome of Pasteurella haemolytica A1 showed that the mol% G+C of P. haemolytica A1 DNA is 38.5%. When only the coding sequences (approx. 10 kbp) were analysed, a similar value of 38.8% was obtained. A comparison of the relative synonymous codon usage values of the cloned genes showed that P. haemolytica A1 has a very different codon usage pattern from that of Escherichia coli.

Base Composition↗

Differences in codon bias cannot explain differences in translational power among microbes.

BACKGROUND: Translational power is the cellular rate of protein synthesis normalized to the biomass invested in translational machinery. Published data suggest a previously unrecognized pattern: translational power is higher among rapidly growing microbes, and lower among slowly growing microbes. One factor known to affect translational power is biased use of synonymous codons. The correlation within an organism between expression level and degree of codon bias among genes of Escherichia coli and other bacteria capable of rapid growth is commonly attributed to selection for high translational power. Conversely, the absence of such a correlation in some slowly growing microbes has been interpreted as the absence of selection for translational power. Because codon bias caused by translational selection varies between rapidly growing and slowly growing microbes, we investigated whether observed differences in translational power among microbes could be explained entirely by differences in the degree of codon bias. Although the data are not available to estimate the effect of codon bias in other species, we developed an empirically-based mathematical model to compare the translation rate of E. coli to the translation rate of a hypothetical strain which differs from E. coli only by lacking codon bias. RESULTS: Our reanalysis of data from the scientific literature suggests that translational power can differ by a factor of 5 or more between E. coli and slowly growing microbial species. Using empirical codon-specific in vivo translation rates for 29 codons, and several scenarios for extrapolating from these data to estimates over all codons, we find that codon bias cannot account for more than a doubling of the translation rate in E. coli, even with unrealistic simplifying assumptions that exaggerate the effect of codon bias. With more realistic assumptions, our best estimate is that codon bias accelerates translation in E. coli by no more than 60% in comparison to microbes with very little codon bias. CONCLUSIONS: While codon bias confers a substantial benefit of faster translation and hence greater translational power, the magnitude of this effect is insufficient to explain observed differences in translational power among bacterial and archaeal species, particularly the differences between slowly growing and rapidly growing species. Hence, large differences in translational power suggest that the translational apparatus itself differs among microbes in ways that influence translational performance.

Bacterial Physiological Phenomena↗

A novel intra-molecular protein-protein interaction code based on partial complementary coding of co-locating amino acids.

Proteins are assumed to contain all the information necessary for unambiguous folding and specific interaction with each other. However, ab initio structure prediction is often not successful because the amino acid sequence itself is simply not sufficient to guide between endless folding possibilities. It seems to be logical to try to find the "missing" information in nucleic acids, in the redundant codon. Statistical analyses of approximately 35K amino acid co-locations in 80 different protein structures indicate the existence of a weak intra-molecular protein-protein interaction code. Co-locating amino acids are preferentially coded by codons which are complementary in reverse orientation to each other at the 1st and 3rd codon positions, but not necessarily at the 2nd. This code, called D-1 X 3/RC-3 X 1, limits the number of preferred amino acid pairs from 20 to 10.3+/-0.8 (SEM, n=20) and emphasizes the importance of "strictly" defined amino acids (those having less synonymous codons). The existence of this code does not by any means violate the known physicochemical rules of protein folding or interaction. It is suggested that the biological source of preferential (specific) amino acid co-locations is the partial complementarity of their codons. This special coding of co-locating amino acids is important to better understanding of some fundamental biochemical processes and observations such as: (a) protein folding; (b) specific and high affinity protein-protein interactions; (c) the role of the wobble bases; (d) the significance of the redundant genetic code; (e) the origin of specific protein-protein interactions. Furthermore it might be useful even in protein design.

Amino Acid Sequence↗

Compositional pressure and translational selection determine codon usage in the extremely GC-poor unicellular eukaryote Entamoeba histolytica.

It is widely accepted that the compositional pressure is the only factor shaping codon usage in unicellular species displaying extremely biased genomic compositions. This seems to be the case in the prokaryotes Mycoplasma capricolum, Rickettsia prowasekii and Borrelia burgdorferi (GC-poor), and in Micrococcus luteus (GC-rich). However, in the GC-poor unicellular eukaryotes Dictyostelium discoideum and Plasmodium falciparum, there is evidence that selection, acting at the level of translation, influences codon choices. This is a twofold intriguing finding, since (1) the genomic GC levels of the above mentioned eukaryotes are lower than the GC% of any studied bacteria, and (2) bacteria usually have larger effective population sizes than eukaryotes, and hence natural selection is expected to overcome more efficiently the randomizing effects of genetic drift among prokaryotes than among eukaryotes. In order to gain a new insight about this problem, we analysed the patterns of codon preferences of the nuclear genes of Entamoeba histolytica, a unicellular eukaryote characterised by an extremely AT-rich genome (GC = 25%). The overall codon usage is strongly biased towards A and T in the third codon positions, and among the presumed highly expressed sequences, there is an increased relative usage of a subset of codons, many of which are C-ending. Since an increase in C in third codon positions is 'against' the compositional bias, we conclude that codon usage in E. histolytica, as happens in D. discoideum and P. falciparum, is the result of an equilibrium between compositional pressure and selection. These findings raise the question of why strongly compositionally biased eukaryotic cells may be more sensitive to the (presumed) slight differences among synonymous codons than compositionally biased bacteria.

Animals↗

Codon adaptation and synonymous substitution rate in diatom plastid genes.

Diatom plastid genes are examined with respect to codon adaptation and rates of silent substitution (Ks). It is shown that diatom genes follow the same pattern of codon usage as other plastid genes studied previously. Highly expressed diatom genes display codon adaptation, or a bias toward specific major codons, and these major codons are the same as those in red algae, green algae, and land plants. It is also found that there is a strong correlation between Ks and variation in codon adaptation across diatom genes, providing the first evidence for such a relationship in the algae. It is argued that this finding supports the notion that the correlation arises from selective constraints, not from variation in mutation rate among genes. Finally, the diatom genes are examined with respect to variation in Ks among different synonymous groups. Diatom genes with strong codon adaptation do not show the same variation in synonymous substitution rate among codon groups as the flowering plant psbA gene which, previous studies have shown, has strong codon adaptation but unusually high rates of silent change in certain synonymous groups. The lack of a similar finding in diatoms supports the suggestion that the feature is unique to the flowering plant psbA due to recent relaxations in selective pressure in that lineage.

Adaptation, Physiological↗

Analysis of interactions between the codon-anticodon duplexes within the ribosome: their role in translation.

Computer graphics simulation of interactions between the codon-anticodon duplexes formed by normal elongator tRNAs at the ribosomal A, P and E-sites (the AP and PE interduplex interactions) was made. This demonstrated that only the correct duplexes at the A-site are compatible with the AP interduplex interaction. The selection of synonymous codons and anticodon wobble bases, together with the AP interduplex interaction, prevents frameshifting. In the absence of this interaction the efficiency of the selection falls off sharply. This suggests that the AP interduplex interaction should be retained during translocation and in the post-translocation state, i.e. the PE interduplex interaction that is identical with that of AP should exist to avoid frameshifting. In such a model the P-site duplex provides an indirect linkage between the A and E-site duplexes. The indirect linkage prohibits the simultaneous existence of the A and E-site duplexes. The wobble pairs of the P and E-site duplexes can affect the rate of the A-site occupation via the AP interduplex interaction and the AE interduplex indirect linkage. It is demonstrated that frameshifting can occur from the AP or PE codon-anticodon complex destabilization caused, for example, by small mobility of the wobble pairs, misreading of the codon, unmodified adenine and guanine at tRNA positions 34 (wobble) and 37, respectively. The results obtained can be subjected to direct experimental tests.

Anticodon↗

Amino acid composition is correlated with protein abundance in Escherichia coli: can this be due to optimization of translational efficiency?

Amino acid occurrence frequencies were found for four groups of Escherichia coli proteins with different abundance levels in the cell. These frequencies decrease with increasing protein abundance for amino acids whose codons are translated by tRNAs present at low concentrations (e.g., Cys, Trp, Ser, etc.); the opposite tendency was observed for amino acids translated by abundant tRNAs (Lys, Val, etc.). The efficiency (rate and accuracy) of codon translation is expected to be proportional to the concentration of the cognate tRNA. Therefore, the observed constraints on amino acid composition may be explained as resulting from evolutionary pressure optimizing the translational efficiency of a gene (the same pressure is responsible for the nonrandom choice of synonymous codons).

Amino Acids↗

Genetic code redundancy and the evolutionary stability of protein secondary structure.

The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. The extent to which these biases are expressed in protein secondary structure is described through the analysis of the overall amino acid compositions of the alpha-helix, beta-sheet, beta-turn and random coil segments elucidated by X-ray crystallography. Given the concept of neutral mutation in proteins, the allocation of synonyms in the genetic code appears to protect secondary structures from amino acid changes and discourages the appearance of chemically complex residues. The level of protection is similar for each structural form, despite their clear preferences for certain amino acids. The organization of the code is therefore relevant to the preservation of conformation seen in the evolution of many protein families.

Amino Acids↗

High-level expression of codon optimized foot-and-mouth disease virus complex epitopes and cholera toxin B subunit chimera in Hansenula polymorpha.

A codon optimized DNA sequence coding for foot-and-mouth disease virus (FMDV) capsid protein complex epitopes of VP1 amino acid residues 21-40, 135-160, and 200-213 was genetically fused to the N-terminal end of a 6x His-tagged cholera toxin B subunit (CTB) gene with the similar synonymous codons preferred by the methylotropic yeast Hansenula polymorpha. The fusion gene was synthesized based on a polymerase chain reaction (PCR) and subsequently overexpressed in H. polymorpha. The chimeric protein was successfully secreted into the culture medium (up to 100mg/L) and retained the antigenicity associated with CTB and FMDV antibodies by Western blot analysis. The chimera after purification through Co(2+)-charged resin column bound specifically to GM1 ganglioside receptor and thus retained the biological activity of CTB. This study has important implications in the construction of CTB chimera for mucosal vaccines against FMDV.

Base Sequence↗

Detection of foot-and-mouth virus antibodies using a purified protein from the high-level expression of codon-optimized, foot-and-mouth disease virus complex epitopes in Escherichia coli.

A codon optimized DNA sequence coding for foot-and-mouth disease virus (FMDV) capsid protein complex epitopes of VP1 amino acid residues 21-40, 135-160, and 200-213 was genetically fused to the C-terminal end of a glutathione-S-transferase (GST) gene in pGEX-6P-1 vector with the synonymous codons preferred by Escherichia coli . The gene was synthesized using PCR and subsequently expressed in E. coli producing an intracellular, soluble fusion protein that retained antigenicity associated with FMDV antibodies by western blot analysis. The chimera was purified from bacterial lysates by affinity chromatography and could be used in ELISA tests for antibodies against FMDV.

Animals↗

The quality of merC, a module of the mer mosaic.

We examined a region of high variability in the mosaic mercury resistance (mer) operon of natural bacterial isolates from the primate intestinal microbiota. The region between the merP and merA genes of nine mer loci was sequenced and either the merC, the merF, or no gene was present. Two novel merC genes were identified. Overall nucleotide diversity, pi (per 100 sites), of the merC gene was greater (49.63) than adjacent merP (35.82) and merA (32.58) genes. However, the consequences of this variability for the predicted structure of the MerC protein are limited and putative functional elements (metal-binding ligands and transmembrane domains) are strongly conserved. Comparison of codon usage of the merTP, merC, and merA genes suggests that several merC genes are not coeval with their flanking sequences. Although evidence of homologous recombination within the very variable merC genes is not apparent, the flanking regions have higher homologies than merC, and recombination appears to be driving their overall sequence identities higher. The synonymous codon usage bias (EN(C)) values suggest greater variability in expression of the merC gene than in flanking genes in six different bacterial hosts. We propose a model for the evolution of MerC as a host-dependent, adventitious module of the mer operon.

Amino Acid Sequence↗

Molecular evolution in the Drosophila melanogaster species subgroup: frequent parameter fluctuations on the timescale of molecular divergence.

Although mutation, genetic drift, and natural selection are well established as determinants of genome evolution, the importance (frequency and magnitude) of parameter fluctuations in molecular evolution is less understood. DNA sequence comparisons among closely related species allow specific substitutions to be assigned to lineages on a phylogenetic tree. In this study, we compare patterns of codon usage and protein evolution in 22 genes (>11,000 codons) among Drosophila melanogaster and five relatives within the D. melanogaster subgroup. We assign changes to eight lineages using a maximum-likelihood approach to infer ancestral states. Uncertainty in ancestral reconstructions is taken into account, at least to some extent, by weighting reconstructions by their posterior probabilities. Four of the eight lineages show potentially genomewide departures from equilibrium synonymous codon usage; three are decreasing and one is increasing in major codon usage. Several of these departures are consistent with lineage-specific changes in selection intensity (selection coefficients scaled to effective population size) at silent sites. Intron base composition and rates and patterns of protein evolution are also heterogeneous among these lineages. The magnitude of forces governing silent, intron, and protein evolution appears to have varied frequently, and in a lineage-specific manner, within the D. melanogaster subgroup.

Amino Acid Substitution↗

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae↗