Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

The relationship between synonymous codon usage and protein structure.

The hypothesis that synonymous codon usage is related to protein three-dimensional structure is examined by investigating the correlation between synonymous codon usage and protein secondary structure. All except two codons in E. coli show the same secondary structural preference for alpha-helix, beta-strand or coil as that of amino acids to be encoded by the respective codons, while 17 codons show secondary structural bias in mammalian proteins. The results indicate that there is no significant correlation between synonymous codon usage and protein secondary structure in E. coli, but there is a correlation in mammals. It could be deduced that synonymous codons carry much less structural information in prokaryotes than in eukaryotes due to their divergent evolutionary mechanism.

Animals↗

Codon usage in selected AT-rich bacteria.

The relationship between DNA base composition and codon bias in very AT-rich bacteria was analyzed. Five clostridial genes, five mycoplasmal genes and three rickettsial genes constituted the data base. In the genes of these three organisms, the rule for codon bias was very simple: use U or A in the first and third positions of the codon when possible. This was contrasted with the bias found in Bacillus subtilis and Escherichia coli. The rule for Bacillus subtilis was equally straightforward: use all codons without bias. Only in E. coli, amongst the species examined, did the codon bias appear to be a complicated codon 'choice'.

Adenine↗

Molecular evolution of an imprinted gene: repeatability of patterns of evolution within the mammalian insulin-like growth factor type II receptor.

The repeatability of patterns of variation in Ka/Ks and Ks is expected if such patterns are the result of deterministic forces. We have contrasted the molecular evolution of the mammalian insulin-like growth factor type II receptor (Igf2r) in the mouse-rat comparison with that in the human-cow comparison. In so doing, we investigate explanations for both the evolution of genomic imprinting and for Ks variation (and hence putatively for mutation rate evolution). Previous analysis of Igf2r, in the mouse-rat comparison, found Ka/Ks patterns that were suggested to be contrary to those expected under the conflict theory of imprinting. We find that Ka/Ks variation is repeatable and hence confirm these patterns. However, we also find that the molecular evolution of Igf2r signal sequences suggests that positive selection, and hence conflict, may be affecting this region. The variation in Ks across Igf2r is also repeatable. To the best of our knowledge this is the first demonstration of such repeatability. We consider three explanations for the variation in Ks across the gene: (1) that it is the result of mutational biases, (2) that it is the result of selection on the mutation rate, and (3) that it is the product of selection on codon usage. Explanations 2 and 3 predict a Ka-Ks correlation, which is not found. Explanation 3 also predicts a negative correlation between codon bias and Ks, which is also not found. However, in support of explanation 1 we do find that in rodents the rate of silent C --> T mutations at CpG sites does covary with Ks, suggesting that methylation-induced mutational patterns can explain some of the variation in Ks. We find evidence to suggest that this CpG effect is due to both variation in CpG density, and to variation in the frequency with which CpGs mutate. Interestingly, however, a GC4 analysis shows no covariance with Ks, suggesting that to eliminate methyl-associated effects CpG rates themselves must be analyzed. These results suggest that, in contrast to previous studies of intragenic variation, Ks patterns are not simply caused by the same forces responsible for Ka/Ks correlations.

Animals↗

Molecular evolution of duplicated ray finned fish HoxA clusters: increased synonymous substitution rate and asymmetrical co-divergence of coding and non-coding sequences.

In this study the molecular evolution of duplicated HoxA genes in zebrafish and fugu has been investigated. All 18 duplicated HoxA genes studied have a higher non-synonymous substitution rate than the corresponding genes in either bichir or paddlefish, where these genes are not duplicated. The higher rate of evolution is not due solely to a higher non-synonymous-to-synonymous rate ratio but to an increase in both the non-synonymous as well as the synonymous substitution rate. The synonymous rate increase can be explained by a change in base composition, codon usage, or mutation rate. We found no changes in nucleotide composition or codon bias. Thus, we suggest that the HoxA genes may experience an increased mutation rate following cluster duplication. In the non-Hox nuclear gene RAG1 only an increase in non-synonymous substitutions could be detected, suggesting that the increased mutation rate is specific to duplicated Hox clusters and might be related to the structural instability of Hox clusters following duplication. The divergence among paralog genes tends to be asymmetric, with one paralog diverging faster than the other. In fugu, all b-paralogs diverge faster than the a-paralogs, while in zebrafish Hoxa-13a diverges faster. This asymmetry corresponds to the asymmetry in the divergence rate of conserved non-coding sequences, i.e., putative cis-regulatory elements. These results suggest that the 5' HoxA genes in the same cluster belong to a co-evolutionary unit in which genes have a tendency to diverge together.

Animals↗

Protein-encoding genes in the sulfothermophilic archaea Sulfolobus and Pyrococcus.

A number of unrelated protein-encoding genes from sulfothermophilic archaea, Sulfolobus acidocaldarius, Sulfolobus solfataricus, Pyrococcus furiosus and Pyrococcus woesei, has been analyzed. In the Sulfolobus genus, the content of A + T is significantly higher than that of C + G and the base usage follows the order, A > T > G > C. In Pyrococcus, the A + T content is also higher than that of C + G, but with lower values; in the order of base usage, G precedes T. The codon usage of these sulfothermophiles has been determined; alternative start codons are frequently used in both genera; codon preferences reflect the rich A + T composition of the corresponding genomes; for both genera the codon bias is particularly evident within the different arginine triplets, where AGA and AGG are predominant. From the similarities in the codon usage, close taxonomic relationships become evident within the Sulfolobus or the Pyrococcus genus; a lower, but significant similarity is also clear between these genera. The synonymous codon usage of these sulfothermophiles shows similarities with that of Saccharomyces cerevisiae and bovine mitochondria, whereas clear divergences are observed with the halophilic archaeal genus, Halobacterium, or the eubacterium, Escherichia coli. The unrelated proteins of the considered sulfothermophiles have been analyzed for the content of hydrophobic residues; the comparison with mesophiles reveals a significant increase in the average hydrophobicity of amino acid residues. This finding could indicate a mechanism of adaptation of proteins in organisms living under extreme environments. It is noteworthy that an opposite trend, i.e. a decreased average hydrophobicity, occurs in unrelated halophilic proteins.

Animals↗

Codon usages in different gene classes of the Escherichia coli genome.

A new measure for assessing codon bias of one group of genes with respect to a second group of genes is introduced. In this formulation, codon bias correlations for Escherichia coli genes are evaluated for level of expression, for contrasts along genes, for genes in different 200 kb (or longer) contigs around the genome, for effects of gene size, for variation over different function classes, for codon bias in relation to possible lateral transfer and for dicodon bias for some gene classes. Among the function classes, codon biases of ribosomal proteins are the most deviant from the codon frequencies of the average E. coli gene. Other classes of 'highly expressed genes' (e.g. amino acyl tRNA synthetases, chaperonins, modification genes essential to translation activities) show less extreme codon biases. Consistently for genes with experimentally determined expression rates in the exponential growth phase, those of highest molar abundances are more deviant from the average gene codon frequencies and are more similar in codon frequencies to the average ribosomal protein gene. Independent of gene size, the codon biases in the 5' third of genes deviate by more than a factor of two from those in the middle and 3' thirds. In this context, there appear to be conflicting selection pressures imposed by the constraints of ribosomal binding, or more generally the early phase of protein synthesis (about the first 50 codons) may be more biased than the complete nascent polypeptide. In partitioning the E. coli genome into 10 equal lengths, pronounced differences in codon site 3 G+C frequencies accumulate. Genes near to oriC have 5% greater codon site 3 G+C frequencies than do genes from the ter region. This difference also is observed between small (100-300 codons) and large (>800 codons) genes. This result contrasts with that for eukaryotic genomes (including human, Caenorhabditis elegans and yeast) where long genes tend to have site 3 more AT rich than short genes. Many of the above results are special for E. coli genes and do not apply to genes of most bacterial genomes. A gene is defined as alien (possibly horizontally transferred) if its codon bias relative to the average gene exceeds a high threshold and the codon bias relative to ribosomal proteins is also appropriately high. These are identified, including four clusters (operons). The bulk of these genes have no known function.

Amino Acyl-tRNA Synthetases↗

Cloning and sequence analysis of putative glyceraldehyde-3-phosphate dehydrogenase gene from Monascus purpureus KCCM11832.

Using a synthetic oligonucleotide probe, glyceraldehyde-3-phosphate dehydrogenase gene (gpd1) was cloned from Monascus purpureus KCCM11832. The 2834 bp EcoRV-HindIII region harbored 1183 bp 5'-UTR containing such regulatory elements as CT box, common in fungal gpd's, and gpd box previously found exclusively in Aspergillus gpd's. Full-length cDNA was cloned by PCR, and its sequence was determined. Transcription starting point was located 88 bp upstream from start codon. Polyadenylation signal sequence occurred 201 bp downstream from stop codon. Region from start codon ATG to stop codon TAA including introns showed 62 approximately 69% nucleotide sequence identity to those of Aspergillus gpd's. Significant bias in third position, with pyrimidines favored over purines, was observed in codon usage. The deduced amino acid sequence had 81 approximately 85% identity to Aspergillus gpd's. Monascus purpureus GPD was located at the same clade with Aspergillus GPD's.

Amino Acid Sequence↗

Proteome analysis of the plant pathogen Xylella fastidiosa reveals major cellular and extracellular proteins and a peculiar codon bias distribution.

The bacteria Xylella fastidiosa is the causative agent of a number of economically important crop diseases, including citrus variegated chlorosis. Although its complete genome is already sequenced, X. fastidiosa is very poorly characterized by biochemical approaches at the protein level. In an initial effort to characterize protein expression in X. fastidiosa we used one- and two-dimensional gel electrophoresis and mass spectrometry to identify the products of 142 genes present in a whole cell extract and in an extracellular fraction of the citrus isolated strain 9a5c. Of particular interest for the study of pathogenesis are adhesion and secreted proteins. Homologs to proteins from three different adhesion systems (type IV fimbriae, mrk pili and hsf surface fibrils) were found to be coexpressed, the last two being detected only as multimeric complexes in the high molecular weight region of one-dimensional electrophoresis gels. Using a procedure to extract secreted proteins as well as proteins weakly attached to the cell surface we identified 30 different proteins including toxins, adhesion related proteins, antioxidant enzymes, different types of proteases and 16 hypothetical proteins. These data suggest that the intercellular space of X. fastidiosa colonies is a multifunctional microenvironment containing proteins related to in vivo bacterial survival and pathogenesis. A codon usage analysis of the most expressed proteins from the whole cell extract revealed a low biased distribution, which we propose is related to the slow growing nature of X. fastidiosa. A database of the X. fastidiosa proteome was developed and can be accessed via the internet (URL: www.proteome.ibi.unicamp.br).

Antioxidants↗

Codon usage in higher plants, green algae, and cyanobacteria.

Codon usage is the selective and nonrandom use of synonymous codons by an organism to encode the amino acids in the genes for its proteins. During the last few years, a large number of plant genes have been cloned and sequenced, which now permits a meaningful comparison of codon usage in higher plants, algae, and cyanobacteria. For the nuclear and organellar genes of these organisms, a small set of preferred codons are used for encoding proteins. Codon usage is different for each genome type with the variation mainly occurring in choices between codons ending in cytidine (C) or guanosine (G) versus those ending in adenosine (A) or uridine (U). For organellar genomes, chloroplastic and mitochrondrial proteins are encoded mainly with codons ending in A or U. In most cyanobacteria and the nuclei of green algae, proteins are encoded preferentially with codons ending in C or G. Although only a few nuclear genes of higher plants have been sequenced, a clear distinction between Magnoliopsida (dicot) and Liliopsida (monocot) codon usage is evident. Dicot genes use a set of 44 preferred codons with a slight preference for codons ending in A or U. Monocot codon usage is more restricted with an average of 38 codons preferred, which are predominantly those ending in C or G. But two classes of genes can be recognized in monocots. One set of monocot genes uses codons similar to those in dicots, while the other genes are highly biased toward codons ending in C or G with a pattern similar to nuclear genes of green algae. Codon usage is discussed in relation to evolution of plants and prospects for intergenic transfer of particular genes.

Journal Article↗

Key for protein coding sequences identification: computer analysis of codon strategy.

The signal qualifying an AUG or GUG as an initiator in mRNAs processed by E. coli ribosomes is not found to be a systematic, literal homology sequence. In contrast, stability analysis reveals that initiators always occur within nucleic acid domains of low stability, for which a high A/U content is observed. Since no aminoacid selection pressure can be detected at N-termini of the proteins, the A/U enrichment results from a biased usage of the code degeneracy. A computer analysis is presented which allows easy detection of the codon strategy. N-terminal codons carry rather systematically A or U in third position, which suggests a mechanism for translation initiation and helps to detect protein coding sequences in sequenced DNA.

Amino Acid Sequence↗

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code↗

Codon usage between genomes is constrained by genome-wide mutational processes.

Analysis of genome-wide codon bias shows that only two parameters effectively differentiate the genome-wide codon bias of 100 eubacterial and archaeal organisms. The first parameter correlates with genome GC content, and the second parameter correlates with context-dependent nucleotide bias. Both of these parameters may be calculated from intergenic sequences. Therefore, genome-wide codon bias in eubacteria and archaea may be predicted from intergenic sequences that are not translated. When these two parameters are calculated for genes from nonmammalian eukaryotic organisms, genes from the same organism again have similar values, and genome-wide codon bias may also be predicted from intergenic sequences. In mammals, genes from the same organism are similar only in the second parameter, because GC content varies widely among isochores. Our results suggest that, in general, genome-wide codon bias is determined primarily by mutational processes that act throughout the genome, and only secondarily by selective forces acting on translated sequences.

Animals↗

Comparative complete genome sequence analysis of the amino acid replacements responsible for the thermostability of Corynebacterium efficiens.

Corynebacterium efficiens is the closest relative of Corynebacterium glutamicum, a species widely used for the industrial production of amino acids. C. efficiens but not C. glutamicum can grow above 40 degrees C. We sequenced the complete C. efficiens genome to investigate the basis of its thermostability by comparing its genome with that of C. glutamicum. The difference in GC content between the species was reflected in codon usage and nucleotide substitutions. Our comparative genomic study clearly showed that there was tremendous bias in amino acid substitutions in all orthologous ORFs. Analysis of the direction of the amino acid substitutions suggested that three substitutions are important for the stability of the C. efficiens proteins: from lysine to arginine, serine to alanine, and serine to threonine. Our results strongly suggest that the accumulation of these three types of amino acid substitutions correlates with the acquisition of thermostability and is responsible for the greater GC content of C. efficiens.

Amino Acid Sequence↗

Molecular characterization of the principal symbiotic bacteria of the weevil Sitophilus oryzae: a peculiar G + C content of an endocytobiotic DNA.

The principal intracellular symbiotic bacteria of the cereal weevil Sitophilus oryzae were characterized using the sequence of the 16S rDNA gene (rrs gene) and G + C content analysis. Polymerase chain reaction amplification with universal eubacterial primers of the rrs gene showed a single expected sequence of 1,501 bp. Comparison of this sequence with the available database sequences placed the intracellular bacteria of S. oryzae as members of the Enterobacteriaceae family, closely related to the free-living bacteria, Erwinia herbicola and Escherichia coli, and the endocytobiotic bacteria of the tsetse fly and aphids. Moreover, by high-performance liquid chromatography, we measured the genomic G + C content of the S. oryzae principal endocytobiotes (SOPE) as 54%, while the known genomic G + C content of most intracellular bacteria is about 39.5%. Furthermore, based on the third codon position G + C content and the rrs gene G + C content, we demonstrated that most intracellular bacteria except SOPE are A + T biased irrespective of their phylogenetic position. Finally, using the hsp60 gene sequence, the codon usage of SOPE was compared with that of two phylogenetically closely related bacteria: E. coli, a free-living bacterium, and Buchnera aphidicola, the intracellular symbiotic bacteria of aphids. Taken together, these results show a peculiar and distinctly different DNA composition of SOPE with respect to the other obligate intracellular bacteria, and, combined with biological and biochemical data, they elucidate the evolution of symbiosis in S. oryzae.

Animals↗

A Rev-independent human immunodeficiency virus type 1 (HIV-1)-based vector that exploits a codon-optimized HIV-1 gag-pol gene.

The human immunodeficiency virus (HIV) genome is AU rich, and this imparts a codon bias that is quite different from the one used by human genes. The codon usage is particularly marked for the gag, pol, and env genes. Interestingly, the expression of these genes is dependent on the presence of the Rev/Rev-responsive element (RRE) regulatory system, even in contexts other than the HIV genome. The Rev dependency has been explained in part by the presence of RNA instability sequences residing in these coding regions. The requirement for Rev also places a limitation on the development of HIV-based vectors, because of the requirement to provide an accessory factor. We have now synthesized a complete codon-optimized HIV-1 gag-pol gene. We show that expression levels are high and that expression is Rev independent. This effect is due to an increase in the amount of gag-pol mRNA. Provision of the RRE in cis did not lower protein or RNA levels or stimulate a Rev response. Furthermore we have used this synthetic gag-pol gene to produce HIV vectors that now lack all of the accessory proteins. These vectors should now be safer than murine leukemia virus-based vectors.

Base Sequence↗

Guanine-adenine bias: a general property of retroid viruses that is unrelated to host-induced hypermutation.

The recently discovered mammalian enzymes, APOBEC3G and 3F, induce guanine-to-adenine hypermutation in retroviruses. However, the preference of adenine over guanine in retroviral codon usage is not correlated with the presence or absence of APOBEC3G or its viral inhibitor (Vif), and its pattern does not reflect the biochemical properties of APOBEC3G action. The guanine-adenine bias of retroviruses is thus probably not a result of host-induced mutational pressure, but rather reflects a general predisposition associated with reverse transcription.

APOBEC-3G Deaminase↗

Widespread selection for local RNA secondary structure in coding regions of bacterial genes.

Redundancy of the genetic code dictates that a given protein can be encoded by a large collection of distinct mRNA species, potentially allowing mRNAs to simultaneously optimize desirable RNA structural features in addition to their protein-coding function. To determine whether natural mRNAs exhibit biases related to local RNA secondary structure, a new randomization procedure was developed, DicodonShuffle, which randomizes mRNA sequences while preserving the same encoded protein sequence, the same codon usage, and the same dinucleotide composition as the native message. Genes from 10 of 14 eubacterial species studied and one eukaryote, the yeast Saccharomyces cerevisiae, exhibited statistically significant biases in favor of local RNA structure as measured by folding free energy. Several significant associations suggest functional roles for mRNA structure, including stronger secondary structure bias in the coding regions of intron-containing yeast genes than in intronless genes, and significantly higher folding potential in polycistronic messages than in monocistronic messages in Escherichia coli. Potential secondary structure generally increased in genes from the 5' to the 3' end of E. coli operons, and secondary structure potential was conserved in homologous Salmonella typhi operons. These results are interpreted in terms of possible roles of RNA structures in RNA processing, regulation of mRNA stability, and translational control.

Computational Biology↗

Mitochondrial genes collectively suggest the paraphyly of Crustacea with respect to Insecta.

Complete sequences of seven protein coding genes from Penaeus notialis mitochondrial DNA were compared in base composition and codon usage with homologous genes from Artemia franciscana and four insects. The crustacean genes are significantly less A + T-rich than their counterpart in insects and the pattern of codon usage (ratio of G + C-rich versus A + T-rich codon) is less biased. A phylogenetic analysis using amino acid sequences of the seven corresponding polypeptides supports a sister-taxon status for mollusks-annelid and arthropods. Furthermore, a distance matrix-based tree and two most-parsimonious trees both suggest that crustaceans are paraphyletic with respect to insects. This is also supported by the inclusion of Panulirus argus COII (complete) and COI and COIII (partial) sequence data. From analysis of single and combined genes to infer phylogenies, it is observed that obtained from single genes are not well supported in most topologies cases and notably differ from that of the tree based on all seven genes.

Animals↗