Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Shannon information theoretic computation of synonymous codon usage biases in coding regions of human and mouse genomes.

Exonic GC of human mRNA reference sequences (RefSeqs), as well as A, C, G, and T in codon position 3 are linearly correlated with genomic GC. These observations utilize information from the completed human genome sequence and a large, high-quality set of human and mouse coding sequences, and are in accord with similar determinations published by others. A Shannon Information Theoretic measure of bias in synonymous codon usage was developed. When applied to either human or mouse RefSeqs, this measure is nonlinearly correlated with genomic, exonic, and third codon position A, C, G, and T. Information values between orthologous mouse and human RefSeqs are linearly correlated: mouse = 0.092 + 0.55 human. Mouse genes were consistently placed in genomic regions whose GC content was closer to 50% than was the GC content of the human ortholog. Since the (nonlinear) information versus percent GC curve has a minimum at 50% GC and monotonically increases with increasing distance from 50% GC, this phenomenon directly results in the low slope of 0.55. This appears to be a manifestation of an evolutionary strategy for placement of genes in regions of the genome with a GC content that relates synonymous codon bias and protein folding.

Animals↗

Design, synthesis and expression of a human interleukin-2 gene incorporating the codon usage bias found in highly expressed Escherichia coli genes.

A synthetic gene encoding human interleukin-2 (IL-2) was designed such that the codon usage bias resembled that found in highly expressed Escherichia coli genes. The percentage of preferred codons was increased from 43% in the native cDNA sequence to 85% in the synthetic sequence. The cDNA and synthetic IL-2 genes were placed under the control of the trc promoter and expressed in E. coli JM101. While Northern blot analysis of IL-2 mRNA from each genetic construct demonstrated equivalent message half-lives, immunoblot and bioactivity analyses showed the synthetic gene to direct the synthesis of up to 16 times more IL-2 than the native cDNA sequence.

Amino Acid Sequence↗

Codon usage bias and mutation constraints reduce the level of error minimization of the genetic code.

Studies on the origin of the genetic code compare measures of the degree of error minimization of the standard code with measures produced by random variant codes but do not take into account codon usage, which was probably highly biased during the origin of the code. Codon usage bias could play an important role in the minimization of the chemical distances between amino acids because the importance of errors depends also on the frequency of the different codons. Here I show that when codon usage is taken into account, the degree of error minimization of the standard code may be dramatically reduced, and shifting to alternative codes often increases the degree of error minimization. This is especially true with a high CG content, which was probably the case during the origin of the code. I also show that the frequency of codes that perform better than the standard code, in terms of relative efficiency, is much higher in the neighborhood of the standard code itself, even when not considering codon usage bias; therefore alternative codes that differ only slightly from the standard code are more likely to evolve than some previous analyses suggested. My conclusions are that the standard genetic code is far from being an optimum with respect to error minimization and must have arisen for reasons other than error minimization.

Base Composition↗

Patterns of codon usage bias in three dicot and four monocot plant species.

Codon usage in nuclear genes of four monocot and three dicot species was analyzed to find general patterns in codon choice of plant species. Codon bias was correlated with GC content at the third codon position. GC contents were higher in monocot species than in dicot species at all codon positions. The high GC contents of monocot species might be the result of relatively strong mutational bias that occurred in the lineage of the Poaceae species. In both dicot and monocot species, the effective number of codons (ENCs) for most genes was similar to that for the expected ENCs based on the GC content at the third codon positions. G and C ending codons were detected as the "preferred" codons in monocot species, as in Drosophila. Also, many "preferred" codons are the same in dicot species. Pyrimidine (C and T) is used more frequently than purine (G and A) in four-fold degenerate codon groups.

Amino Acid Sequence↗

The role of context-dependent mutations in generating compositional and codon usage bias in grass chloroplast DNA.

The influence of local base composition on mutations in chloroplast DNA (cpDNA) is studied in detail and the resulting, empirically derived, mutation dynamics are used to analyze both base composition and codon usage bias. A 4 x 4 substitution matrix is generated for each of the 16 possible flanking base combinations (contexts) using 17,253 noncoding sites, 1309 of which are variable, from an alignment of three complete grass chloroplast genome sequences. It is shown that substitution bias at these sites is correlated with flanking base composition and that the A+T content of these flanking sites as well as the number of flanking pyrimidines on the same strand appears to have general influences on substitution properties. The context-dependent equilibrium base frequencies predicted from these matrices are then applied to two analyses. The first examines whether or not context dependency of mutations is sufficient to generate average compositional differences between noncoding cpDNA and silent sites of coding sequences. It is found that these two classes of sites exist, on average, in very different contexts and that the observed mutation dynamics are expected to generate significant differences in overall composition bias that are similar to the differences observed in cpDNA. Context dependency, however, cannot account for all of the observed differences: although silent sites in coding regions appear to be at the equilibrium predicted, noncoding cpDNA has a significantly lower A+T content than expected from its own substitution dynamics, possibly due to the influence of indels. The second study examines the codon usage of low-expression chloroplast genes. When context is accounted for, codon usage is very similar to what is predicted by the substitution dynamics of noncoding cpDNA. However, certain codon groups show significant deviation when followed by a purine in a manner suggesting some form of weak selection other than translation efficiency. Overall, the findings indicate that a full understanding of mutational dynamics is critical to understanding the role selection plays in generating composition bias and sequence structure.

Base Composition↗

[Synonymous codon usage bias in the rice cultivar 93-11 (Oryza sativa L. ssp. indica)].

By using the whole genome sequences and EST data from the indica rice cultivar 93-11, a detailed relative analysis is made of the effect of some impact factors on synonymous codon usage. The results showed that the gene expression level assessed by mRNA abundance is positive relative to the "codon adaptation index" (CAI, 0.227**), and "codon preference parameter" (CPP, 0.145**), but negative relative to "effective number of codons" (ENC, -0.147**), indicating that genes with higher expression showed more significant variation in codon usage. There are significant negative correlations between gene length and CAI, CPP (r = -0.413** and -0.480** respectively), but a positive correlation between gene length and ENC(r = 0.210**), which suggested a tendency of shorter genes to higher expression of the transcriptional activity in 93-11. From the results that a higher negative correlation between GC content and ENC(r = -0.740**), but higher positive correlations between GC content and CAI, CPP (r = 0.877** and 0.832**, respectively), we can concluded that the GC content in coding region gave far more contribution to codon usage bias than that mRNA abundance and gene length. Four kinds of bases showed a three-period distribution in the translation initiation region, the bias at the first codon sites, which located +4, and +6, in the downstream of ATG being the largest. That suggested that there was a strong action of natural selection on these specific positions in the 93-11 genome. In this paper twenty-five codons defined firstly as "optimal codons" in 93-11 may provide some more useful information for rice gene-transformation.

Base Composition↗

Synonymous codon usage bias in 16 Staphylococcus aureus phages: implication in phage therapy.

To reveal the factors influencing architecture of protein-coding genes in staphylococcal phages, relative synonymous codon usage variation has been investigated in 920 protein-coding genes of 16 staphylococcal phages. As expected for AT rich genomes, there are predominantly A and T ending codons in all 16 phages. Both Nc plot and correspondence analysis on relative synonymous codon usage indicates that mutation bias influences codon usage variation in the 16 phages. Correspondence analysis also suggests that translational selection and gene length also influence the codon usage variation in the phages to some extent and codon usage in staphylococcal phages is phage-specific but not S. aureus-specific. Further analysis indicates that among 16 staphylococcal phages, 44AHJD, P68 and K may be extremely virulent in nature as most of their genes have high translation efficiency. If this is true, then above three phages may be useful for curing staphylococcal infections.

Animals↗

Synonymous codon usage bias and the expression of human glucocerebrosidase in the methylotrophic yeast, Pichia pastoris.

The lysosomal hydrolase glucocerebrosidase catalyzes the penultimate step in the breakdown of membrane glycosphingolipids. An inherited deficiency in this enzyme leads to the onset of Gaucher disease, the most common lysosomal storage disorder. Exogenous sources of this protein are required for biochemical and biophysical investigations and enzyme replacement therapy of Gaucher disease. Heterologous expression of glucocerebrosidase has been successful in mammalian and insect cell lines and although its use in enzyme replacement therapy of Gaucher disease has proven efficacious, current production levels limit the availability of the enzyme. Initial attempts to express human glucocerebrosidase using the methylotrophic yeast Pichia pastoris had limited success, despite significant levels of transcription. Using fragments of the glucocerebrosidase cDNA fused to the luciferase cDNA as a translational read-through reporter, the impact of synonymous codon usage bias on protein expression in P. pastoris was examined. A table of preferred codons was determined for P. pastoris and the codon usage of a 186-bp fragment of the glucocerebrosidase gene was optimized to that of the P. pastoris preferred set. A second construct with altered G+C content but no codon optimization was created for comparison. While the native glucocerebrosidase coding region limited luciferase activity to baseline levels, the codon optimized and G+C altered constructs increased luciferase activity 10.6- and 7.5-fold, respectively. Optimized G+C content, regardless of corresponding codon optimization, appears to be the major contributor to increased translational efficiency in this heterologous expression host.

Amino Acid Sequence↗

Unconventional codon usage bias mediates mRNA translational dynamics in macrophages.

Macrophages require rapid and tightly controlled regulatory mechanisms to respond to environmental disruptions. While transcriptional regulation has been well characterized, the mechanisms underlying translational control in macrophages remain poorly understood. Here, we investigated the dynamics of mRNA translation in mouse macrophages during acute, intermediate, and prolonged LPS exposure. Our results reveal clear phase-specific translational regulation during macrophage polarization, which initially increases the synthesis of inflammatory mediators and cytokines, while simultaneously suppressing the expression of cell cycle-related genes. Mechanistically, we observed pervasive upstream translation in the 5' UTRs of cell cycle-related mRNAs, which contributes to cell cycle arrest during the early phase of inflammatory response. Notably, we identified a unique codon preference toward A/U in the third position of codons in macrophages, which contrasts with the G/C preference commonly observed in other tissues. AU codon preference increases the stability and translation efficiency of cell cycle-related mRNAs, promoting cell cycle restoration after extended LPS exposure. These findings reveal that uORF translation and codon usage bias are critical components of translational regulation during macrophage polarization, highlighting a potential therapeutic intervention for modulating immune activation via macrophage-specific codon optimization.

Animals↗

The problem of counting sites in the estimation of the synonymous and nonsynonymous substitution rates: implications for the correlation between the synonymous substitution rate and codon usage bias.

Most methods for estimating the rate of synonymous and nonsynonymous substitution per site define a site as a mutational opportunity: the proportion of sites that are synonymous is equal to the proportion of mutations that would be synonymous under the model of evolution being considered. Here we demonstrate that this definition of a site can give misleading results and that a physical definition of site should be used in some circumstances. We illustrate our point by reexamining the relationship between codon usage bias and the synonymous substitution rate. It has recently been shown that the rate of synonymous substitution, calculated using the Goldman-Yang method, which encapsulates the mutational-opportunity definition of a site at a high level of sophistication, is either positively correlated or uncorrelated to synonymous codon bias in Drosophila. Using other methods, which account for synonymous codon bias but define a site physically, we show that there is a negative correlation between the synonymous substitution rate and codon bias and that the lack of a negative correlation using the Goldman-Yang method is due to the way in which the number of synonymous sites is counted. We also show that there is a positive correlation between the synonymous substitution rate and third position GC content in mammals, but that the relationship is considerably weaker than that obtained using the Goldman-Yang method. We argue that the Goldman-Yang method is misleading in this context and conclude that methods that rely on a mutational-opportunity definition of a site should be used with caution.

Animals↗

Amino acid cost and codon-usage biases in 6 prokaryotic genomes: a whole-genome analysis.

For most prokaryotic organisms, amino acid biosynthesis represents a significant portion of their overall energy budget. The difference in the cost of synthesis between amino acids can be striking, differing by as much as 7-fold. Two prokaryotic organisms, Escherichia coli and Bacillus subtilis, have been shown to preferentially utilize less costly amino acids in highly expressed genes, indicating that parsimony in amino acid selection may confer a selective advantage for prokaryotes. This study confirms those findings and extends them to 4 additional prokaryotic organisms: Chlamydia trachomatis, Chlamydophila pneumoniae AR39, Synechocystis sp. PCC 6803, and Thermus thermophilus HB27. Adherence to codon-usage biases for each of these 6 organisms is inversely correlated with a coding region's average amino acid biosynthetic cost in a fashion that is independent of chemoheterotrophic, photoautotrophic, or thermophilic lifestyle. The obligate parasites C. trachomatis and C. pneumoniae AR39 are incapable of synthesizing many of the 20 common amino acids. Removing auxotrophic amino acids from consideration in these organisms does not alter the overall trend of preferential use of energetically inexpensive amino acids in highly expressed genes.

Adaptation, Biological↗

Analysis of synonymous codon usage bias in Chlamydia.

Chlamydiae are obligate intracellular bacterial pathogens that cause ocular and sexually transmitted diseases, and are associated with cardiovascular diseases. The analysis of codon usage may improve our understanding of the evolution and pathogenesis of Chlamydia and allow reengineering of target genes to improve their expression for gene therapy. Here, we analyzed the codon usage of C. muridarum, C. trachomatis (here indicating biovar trachoma and LGV), C. pneumoniae, and C. psittaci using the codon usage database and the CUSP (Create a codon usage table) program of EMBOSS (The European Molecular Biology Open Software Suite). The results show that the four genomes have similar codon usage patterns, with a strong bias towards the codons with A and T at the third codon position. Compared with Homo sapiens, the four chlamydial species show discordant seven or eight preferred codons. The ENC (effective number of codons used in a gene)-plot reveals that the genetic heterogeneity in Chlamydia is constrained by the G+C content, while translational selection and gene length exert relatively weaker influences. Moreover, mutational pressure appears to be the major determinant of the codon usage variation among the chlamydial genes. In addition, we compared the codon preferences of C. trachomatis with those of E. coli, yeast, adenovirus and Homo sapiens. There are 23 codons showing distinct usage differences between C. trachomatis and E. coli, 24 between C. trachomatis and adenovirus, 21 between C. trachomatis and Homo sapiens, but only six codons between C. trachomatis and yeast. Therefore, the yeast system may be more suitable for the expression of chlamydial genes. Finally, we compared the codon preferences of C. trachomatis with those of six eukaryotes, eight prokaryotes and 23 viruses. There is a strong positive correlation between the differences in coding GC content and the variations in codon bias (r=0.905, P<0.001). We conclude that the variation of codon bias between C. trachomatis and other organisms is much less influenced by phylogenetic lineage and primarily determined by the extent of disparities in GC content.

Animals↗

Structural features of multiple nifH-like sequences and very biased codon usage in nitrogenase genes of Clostridium pasteurianum.

The structural gene (nifH1) encoding the nitrogenase iron protein of Clostridium pasteurianum has been cloned and sequenced. It is located on a 4-kilobase EcoRI fragment (cloned into pBR325) that also contains a portion of nifD and another nifH-like sequence (nifH2). C. pasteurianum nifH1 encodes a polypeptide (273 amino acids) identical to that of the isolated iron protein, indicating that the smaller size of the C. pasteurianum iron protein does not result from posttranslational processing. The 5' flanking region of nifH1 or nifH2 does not contain the nif promoter sequences found in several gram-negative bacteria. Instead, a sequence resembling the Escherichia coli consensus promoter (TTGACA-N17-TATAAT) is present before C. pasteurianum nifH2, and a TATAAT sequence is present before C pasteurianum nifH1. Codon usage in nifH1, nifH2, and nifD (partial) is very biased. A preference for A or U in the third position of the codons is seen. nifH2 could encode a protein of 272 amino acid residues, which differs from the iron protein (nifH1 product) in 23 amino acid residues (8%). Another nifH-like sequence (nifH3) is located on a nonadjacent EcoRI fragment and has been partially sequenced. C. pasteurianum nifH2 and nifH3 may encode proteins having several amino acids that are conserved in other proteins but not in C. pasteurianum iron protein, suggesting a possible role for the multiple nifH-like sequences of C. pasteurianum in the evolution of nifH. Among the nine sequenced iron proteins, only the C. pasteurianum protein lacks a conserved lysine residue which is near the extended C terminus of the other iron proteins. The absence of this positive charge in the C. pasteurianum iron protein might affect the cross-reactivity of the protein in heterologous systems.

Amino Acid Sequence↗

The 'weighted sum of relative entropy': a new index for synonymous codon usage bias.

Shannon entropy from information theory has been applied to estimate the degree of deviation from equal usage of synonymous codons; however, previous attempts have failed to take into account all three aspects of amino acid usage, i.e. (i) the number of distinct amino acids, (ii) their relative frequencies, and (iii) their degree of codon degeneracy. A new index taking into account all of these aspects is proposed. The index, designated as the 'weighted sum of relative entropy' (E(w)), is defined as the sum of the relative entropy of each amino acid weighted by its relative frequency in the sequence. In this paper, we demonstrate that E(w) allows us to avoid some amino acid usage biases and can yield results contradictory to those obtained by previous methods.

Algorithms↗

Lactococcus lactis glyceraldehyde-3-phosphate dehydrogenase gene, gap: further evidence for strongly biased codon usage in glycolytic pathway genes.

The gene gap, encoding glyceraldehyde-3-phosphate dehydrogenase (EC 1.2.1.12), was isolated from a genomic library of Lactococcus lactis LM0230 DNA. Plasmids containing the L. lactis gene were able to complement a gap mutant of Escherichia coli. The nucleotide sequence of gap predicted a polypeptide chain of 337 amino acids for the enzyme and a subunit molecular mass of 36,043. The codon usage in gap and four other glycolytic genes from L. lactis showed a high degree of bias, when compared with 84 other chromosomal genes. Northern blot analysis of total L. lactis RNA showed that gap hybridized strongly with a 1.3 kb transcript. The 5' end of the transcript was determined by primer extension analysis to be a C located 35 bp upstream from the gap start codon. These transcript analyses, and the orientation of the open reading frames in the DNA flanking gap, indicated that in L. lactis gap is expressed on a monocistronic transcript. Nucleotide sequencing indicated that the DNA adjacent to gap did not encode other glycolytic pathway enzymes. The DNA sequence flanking gap contained two open reading frames (ORF156 and ORF211) of unknown function. The 3' end of a clpA homologue was identified in the sequence upstream of ORF156. The location of gap on the L. lactis DL11 chromosome map was determined to be between map coordinates 0.530 and 0.660.

Amino Acid Sequence↗

The repertoire of transfer RNA genes is tuned to codon usage bias in the genomes of Phytophthora sojae and Phytophthora ramorum.

In all, 238 and 155 transfer (t)RNA genes were predicted from the genomes of Phytophthora sojae and P. ramorum, respectively. After omitting pseudogenes and undetermined types of tRNA genes, there remained 208 P. sojae tRNA genes and 140 P. ramorum tRNA genes. There were 45 types of tRNA genes, with distinct anticodons, in each species. Fourteen common anticodon types of tRNAs are missing altogether from the genome in the two species; however, these appear to be compensated by wobbling of other tRNA anticodons in a manner which is tied to the codon bias in Phytophthora genes. The most abundant tRNA class was arginine in both P. sojae and P. ramorum. A codon usage table was generated for these two organisms from a total of 9,803,525 codons in P. sojae and 7,496,598 codons in P. ramorum. The most abundant codon type detected from the codon usage tables was GAG (encoding glutamic acid), whereas the most numerous tRNA gene had a methionine anticodon (CAT). The correlation between the frequencies of tRNA genes and the codon frequencies in protein-coding genes was very low (0.12 in P. sojae and 0.19 in P. ramorum); however, the correlation between amino acid tRNA gene frequency and the corresponding amino acid codon frequency in P. sojae and P. ramorum was substantially higher (0.53 in P. sojae and 0.77 in P. ramorum). The codon usage frequencies of P. sojae and P ramorum were very strongly correlated (0.99), as were tRNA gene frequencies (0.77). Approximately 60% of orthologous tRNA gene pairs in P sojae and P. ramorum are located in regions that have conserved synteny in the two species.

Anticodon↗

Comprehensive analysis of synonymous codon usage bias and evolutionary dynamics in the chloroplast genomes of eight Coptis species.

Coptis is a medically important genus renowned for producing valuable isoquinoline alkaloids. Although its chloroplast genomes encode key components for photosynthesis and plastid gene expression, the evolutionary constraints acting on their coding sequences and synonymous codon usage remain poorly resolved. Here, we combined a transparent taxon-level sampling strategy with comparative analyses of chloroplast CDSs from eight Coptis taxa. We quantified nucleotide composition, relative synonymous codon usage, effective number of codons, neutrality and PR2 patterns, and correspondence analysis, and then integrated these results with a core-CDS distance analysis and gene-wise pairwise dN/dS estimates. The chloroplast genomes showed a conserved AT-rich composition, especially at the third codon position (GC3 approximately 30.3-30.8%), with a consistent GC1&#x2009;>&#x2009;GC2&#x2009;>&#x2009;GC3 trend. Thirty preferred codons were detected, 28 ending in A/T, and eleven optimal codons were shared across the genus. The core-CDS distance analysis recovered a close relationship between C. chinensis and C. chinensis var. brevisepala, whereas most coding genes showed dN/dS values below one, consistent with pervasive purifying constraint. Across 48 consistently filtered CDSs, GC3s was negatively associated with mean dN (Spearman rho = -0.404, P&#x2009;=&#x2009;0.00439) and CAI was positively associated with mean dN (rho&#x2009;=&#x2009;0.303, P&#x2009;=&#x2009;0.0361), whereas the remaining associations were not significant (all P&#x2009;>&#x2009;=&#x2009;0.0972). These results extend codon-usage analysis by linking synonymous-site composition to coding-sequence evolution within Coptis, while providing a hypothesis-generating resource for future plastid engineering studies.

Genome, Chloroplast↗