Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Correlation between codon usage and thermostability.

Variations of arginine codon usage between organisms may have important implications to thermostability. The preferential usage of AGR codons for arginine in thermophiles and hyperthermophiles implies positive error minimization, contributing to avoid mutations that could harm protein thermostability. This bias is not a mere consequence of increased G + C content, as it has been previously suggested, and may represent a new mechanism of adaptation to protein thermostability.

Arginine↗

A strong effect of AT mutational bias on amino acid usage in Buchnera is mitigated at high-expression genes.

The advent of full genome sequences provides exceptionally rich data sets to explore molecular and evolutionary mechanisms that shape divergence among and within genomes. In this study, we use multivariate analysis to determine the processes driving genome-wide patterns of amino usage in the obligate endosymbiont Buchnera and its close free-living relative Escherichia coli. In the AT-rich Buchnera genome, the primary source of variation in amino acid usage differentiates high- and low-expression genes. Amino acids of high-expression Buchnera genes are generally less aromatic and use relatively GC-rich codons, suggesting that selection against aromatic amino acids and against amino acids with AT-rich codons is stronger in high-expression genes. Selection to maintain hydrophobic amino acids in integral membrane proteins is a primary factor driving protein evolution in E. coli but is a secondary factor in Buchnera. In E. coli, gene expression is a secondary force driving amino acid usage, and a correlation with tRNA abundance suggests that translational selection contributes to this effect. Although this and previous studies demonstrate that AT mutational bias and genetic drift influence amino acid usage in Buchnera, this genome-wide analysis argues that selection is sufficient to affect the amino acid content of proteins with different expression and hydropathy levels.

Adenine↗

Secondary structure of MS2 phage RNA and bias in code word usage.

Based on the secondary structural model of MS2 RNA, it is shown that, in base-pairing regions of the RNA, there is a bias in the use of synonymous codons which favours C and/or G over U and/or A in the third codon positions, and that in non-pairing regions, there is an opposite bias which favours U and/or A over C and/or G. This nature is interpreted as a result of selective constraint which stabilises the secondary structure of the single-stranded RNA genome of the MS2 phage.

Base Sequence↗

Minor shift in background substitutional patterns in the Drosophila saltans and willistoni lineages is insufficient to explain GC content of coding sequences.

BACKGROUND: Several lines of evidence suggest that codon usage in the Drosophila saltans and D. willistoni lineages has shifted towards a less frequent use of GC-ending codons. Introns in these lineages show a parallel shift toward a lower GC content. These patterns have been alternatively ascribed to either a shift in mutational patterns or changes in the definition of preferred and unpreferred codons in these lineages. RESULTS AND DISCUSSION: To gain additional insight into this question, we quantified background substitutional patterns in the saltans/willistoni group using inactive copies of a novel, Q-like retrotransposable element. We demonstrate that the pattern of background substitutions in the saltans/willistoni lineage has shifted to a significant degree, primarily due to changes in mutational biases. These differences predict a lower equilibrium GC content in the genomes of the saltans/willistoni species compared with that in the D. melanogaster species group. The magnitude of the difference can readily account for changes in intronic GC content, but it appears insufficient to explain changes in codon usage within the saltans/willistoni lineage. CONCLUSION: We suggest that the observed changes in codon usage in the saltans/willistoni clade reflects either lineage-specific changes in the definitions of preferred and unpreferred codons, or a weaker selective pressure on codon bias in this lineage.

Animals↗

Development of a GFP reporter gene for Chlamydomonas reinhardtii chloroplast.

Reporter genes have been successfully used in chloroplasts of higher plants, and high levels of recombinant protein expression have been reported. Reporter genes have also been used in the chloroplast of Chlamydomonas reinhardtii, but in most cases the amounts of protein produced appeared to be very low. We hypothesized that the inability to achieve high levels of recombinant protein expression in the C. reinhardtii chloroplast was due to the codon bias seen in the C. reinhardtii chloroplast genome. To test this hypothesis, we synthesized a gene encoding green fluorescent protein (GFP) de novo, optimizing its codon usage to reflect that of major C. reinhardtii chloroplast-encoded proteins. We monitored the accumulation of GFP in C. reinhardtii chloroplasts transformed with the codon-optimized GFP cassette (GFPct), under the control of the C. reinhardtii rbcL 5'- and 3'-UTRs. We compared this expression with the accumulation of GFP in C. reinhardtii transformed with a non-optimized GFP cassette (GFPncb), also under the control of the rbcL 5'- and 3'-UTRs. We demonstrate that C. reinhardtii chloroplasts transformed with the GFPct cassette accumulate approximately 80-fold more GFP than GFPncb-transformed strains. We further demonstrate that expression from the GFPct cassette, under control of the rbcL 5'- and 3'-UTRs, is sufficiently robust to report differences in protein synthesis based on subtle changes in environmental conditions, showing the utility of the GFPct gene as a reporter of C. reinhardtii chloroplast gene expression.

Amino Acid Sequence↗

Codon usage and nucleotide composition in Coxiella burnetii.

Coxiella burnetii, the causative agent of Q fever, is an obligate intracellular bacterium. With the development of molecular biology techniques, there have been increasing efforts on gene cloning and other genetic analyses of this organism. In this report, we tabulate the codon usage (CU) and nucleotide (nt) co-occurrence in C. burnetii, based on available nt sequence data. The average G+C content of the C. burnetii genome is 42.4%, where the G+C content is 42.7% for the chromosome and 38.7% for the plasmid. In comparison to Escherichia coli, there is biased CU. Some codons are frequently used in C. burnetii, but rarely used in E. coli and vice versa. Plasmid genes prefer A or T at the first or third position of a codon. However, TAA remains the most used stop codon. In the AT-rich DNA of C. burnetii, A or T tend to occur together, forming A or T tracks.

Bacterial Proteins↗

Characterization and phylogenetic utility of the mammalian protamine p1 gene.

We sequenced the protamine P1 gene (ca. 450 bp) from 20 bats (order Chiroptera) and the flying lemur (order Dermoptera). We compared these sequences with published sequences from 19 other mammals representing seven orders (Artiodactyla, Carnivora, Cetacea, Perissodactyla, Primates, Proboscidea, and Rodentia) to assess structure, base compositional bias, and phylogenetic utility. Approximately 80% of second codon positions were guanine, resulting in protamine proteins containing a high frequency of arginine residues. Our data indicate that codon usage for arginine differs among higher mammalian taxa. Parsimony analysis of 40 species representing nine orders produced a well-resolved tree in which most nodes were supported strongly, except at the lowest taxonomic levels (e.g., within Artiodactyla and Vespertilionidae). These data support monophyly of several taxa proposed by morphologic and molecular studies (all nine orders: Laurasiatheria, Cetartiodactytla, Yangochiroptera, Noctilionoidea, Rhinolophoidea, Vespertilionoidea, Phyllostomidae, Natalidae, and Vespertilionidae) and, in agreement with recent molecular studies, reject monophyly of Archonta, Volitantia, and Microchiroptera. Bats were sister to a clade containing Perissodactyla, Carnivora, and Cetartiodactyla, and, although not unequivocally, rhinolophoid bats (traditional microchiropterans) were sister to megachiropterans. Sequences of the protamine P1 gene are useful for resolving relationships at and above the familial level in bats, and generally within and among mammalian orders, but with some drawbacks. The coding and intervening sequences are small, producing few phylogenetically informative characters, and aligning the intron is difficult, even among closely related families. Given these caveats, the protamine P1 gene may be important to future systematic studies because its functional and evolutionary constraints differ from other genes currently used in systematic studies.

Amino Acid Sequence↗

Selection on the codon bias of Chlamydomonas reinhardtii chloroplast genes and the plant psbA gene.

Plant chloroplast genes have a codon use that reflects the genome compositional bias of a high A+T content with the single exception of the highly translated psbA gene which codes for the photosystem II D1 protein. The codon usage of plant psbA corresponds more closely to the limited tRNA population of the chloroplast and is very similar to the codon use observed in the chloroplast genes of the green alga Chlamydomonas reinhardtii. This pattern of codon use may be an adaptation for increased translation efficiency. A correspondence between codon use of plant psbA and Chlamydomonas chloroplast genes and the tRNAs coded by the chloroplast genome, however, is not observed in all synonymous codon groups. It is shown here that the degree of correspondence between codon use and tRNA population in different synonymous groups is correlated with the second codon position composition. Synonymous groups with an A or T at the second codon position have a high representation of codons for which a complementary tRNA is coded by the chloroplast genome. Those with a G or C at the second position have an increased representation of codons that bind a chloroplast tRNA by wobble. It is proposed that the difference between synonymous groups in terms of codon adaptation to the tRNA population in plant psbA and Chlamydomonas chloroplast genes may be the result of differences in second position composition.

Animals↗

Analysis of a shift in codon usage in Drosophila.

In order to gain further insight into a shift in codon usage first observed in Drosophila willistoni we have analyzed seven genes in six species in the lineage leading to D. willistoni. This lineage contains the willistoni and saltans species groups. Sequences were obtained from GenBank or newly sequenced for this study. All species studied showed significant difference in codon usage compared to D. melanogaster for about one third of all amino acids. Within the willistoni/saltans lineage, codon usage is homogeneous, indicating that the shift in codon usage occurred prior to the diversification of extant species in this lineage which we estimate to date to about 20 million years ago. Thus the shift is old and has been stable. We also examined introns from these genes and the G/C composition at four-fold degenerate sites in an effort to detect a change in mutation bias. There is little or no evidence for a difference in mutation bias compared to D. melanogaster. We also considered whether relaxed selection (possibly due to reduced population sizes) or reduced recombination (due to numerous naturally occurring inversions) could account for the shift and concluded these factors alone are insufficient to explain the patterns observed. A change in the relative abundance of isoaccepting tRNAs is one of the few explanations that can account for the observations. Particularly intriguing is the fact that the greatest changes in codon usage have occurred for amino acids with two-fold C/T ending codons for which it is known that posttranscriptional modification occurs in tRNAs from a G in the wobble position to Queuosine that changes optimal binding from C to a slight preference for U. However, we do not argue that this shift was adaptive in nature, rather it may be an example of a "frozen accident."

Amino Acids↗

Relationship between "proto-splice sites" and intron phases: evidence from dicodon analysis.

The coding sequence at the boundaries of exons flanking nuclear introns shows some degree of conservation. To the extent that such sequences might be recognized by the splicing machinery, this conservation may be a derived result of evolution for efficient splicing. Alternatively, such conserved sequences might be remnants of proto-splice sites, which might have existed early in eukaryotic genes and served as the targets for the insertion of introns, as has been proposed by the introns-late theory. The distribution of intron phases, the position of the intron within a codon, is biased with an over-representation of phase 0 introns. Could any distribution of proto-splice sites account for today's intron phase distribution? Here, we examine the dicodon usage in six model organisms, based on current sequences in the GenBank database, and predict the phase distribution that would be expected if introns had been inserted into proto-splice sites. However, these predictions differ between the various model organisms and disagree with the observed intron phase distributions. Thus, we reject the hypothesis that introns are inserted into hypothetical proto-splice sites. Finally, we analyze the sequences around the splice sites of introns in all six of the species to show that the actual conservation of sequence in exon regions near introns is very small and differs considerably between these species, which is inconsistent with a general proto-splice sites model.

Animals↗

Effects of rare codon clusters on high-level expression of heterologous proteins in Escherichia coli.

Within Escherichia coli and other species, a clear codon bias exists among the 61 amino acid codons found within the population of mRNA molecules, and the level of cognate tRNA appears directly proportional to the frequency of codon usage. Given this situation, one would predict translational problems with an abundant mRNA species containing an excess of rare low tRNA codons. Such a situation might arise after the initiation of transcription of a cloned heterologous gene in the E. coli host. Recent studies suggest clusters of AGG/AGA, CUA, AUA, CGA or CCC codons can reduce both the quantity and quality of the synthesized protein. In addition, it is likely that an excess of any of these codons, even without clusters, could create translational problems.

Arginine↗

[Synonymous codon usage in Pichia pastoris].

According to the synonymous codons used in 28 open reading frames from Pichia pastoris, the codon usage in this species was calculated and 19 codons have been inferred to be its optimal codons. The results show that pattern of the codon usage in P. pastoris is similar to that in S. cerevisiae and in K. lactis except for the synonymous codon of glutamic acid, which may be the special bias of P. pastoris.

Codon↗

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals↗

Optimization of codon usage is required for effective genetic immunization against Art v 1, the major allergen of mugwort pollen.

BACKGROUND: As the major allergen of mugwort pollen, Art v 1 is an important target for specific immunotherapy. However, both recombinant protein as well as a gene vaccine for Art v 1 failed to be immunogenic in mice. In order to improve immunogenicity we focused on genetic immunization because interspecific differences of codon usage have been shown as an obstacle for effective induction of immune responses with gene vaccines encoding infectious pathogens. OBJECTIVE: In order to find out, whether codon usage might also be used to improve genetic immunization with allergen genes, the response against a gene vaccine expressing the wild-type gene of Art v 1 (pCMV-wtArt) was compared with a synthetic codon-optimized vector with human codon usage (pCMV-humArt). METHODS: Balb/c mice were injected intradermally with pCMV-wtArt or pCMV-humArt. In vitro expression levels of both constructs were compared in transfection experiments. Total immunoglobulin G (IgG), IgG1, IgG2a and IgE antibodies were analyzed by enzyme-linked immunosorbent assay and the anaphylactic activity of the sera was determined by allergen-specific degranulation of rat basophil leukemia-2H3 cells. RESULTS: No immune response was detectable with the gene vaccine expressing the wildtype Art v 1, but immunization with pCMV-humArt revealed a strong and allergen-specific induction of antibody responses. The antibodies recognized both the recombinant as well as the purified natural (glycosylated) Art v 1 molecule. The response type was Th1-biased, as indicated by high levels of IgG2a antibodies. Expression analysis with B16 mouse melanoma cells transfected with pCMV-humArt or pCMV-wtArt revealed an impaired expression of the wild-type vector but normal translation after recoding. CONCLUSION: The results demonstrate that optimization of codon usage offers a simple way to improve immunogenicity and therefore should be routinely considered in the development of gene vaccines for the treatment of allergy.

Allergens↗

Universal replication biases in bacteria.

Analysis of 15 complete bacterial chromosomes revealed important biases in gene organization. Strong compositional asymmetries between the genes lying on the leading versus lagging strands were observed at the level of nucleotides, codons and, surprisingly, amino acids. For some species, the bias is so high that the sole knowledge of a protein sequence allows one to predict with almost no errors whether the gene is transcribed from one strand or the other. Furthermore, we show that these biases are not species specific but appear to be universal. These findings may have important consequences in our understanding of fundamental biological processes in bacteria, such as replication fidelity, codon usage in genes and even amino acid usage in proteins.

Amino Acids↗

Mammalian housekeeping genes evolve more slowly than tissue-specific genes.

Do housekeeping genes, which are turned on most of the time in almost every tissue, evolve more slowly than genes that are turned on only at specific developmental times or tissues? Recent large-scale gene expression studies enable us to have a better definition of housekeeping genes and to address the above question in detail. In this study, we examined 1581 human-mouse orthologous gene pairs for their patterns of sequence evolution, contrasting housekeeping genes with tissue-specific genes. Our results show that, in comparison to tissue-specific genes, housekeeping genes on average evolve more slowly and are under stronger selective constraints as reflected by significantly smaller values of Ka/Ks. Besides stronger purifying selection, we explored several other factors that can possibly slow down nonsynonymous rates in housekeeping genes. Although mutational bias might slightly slow the nonsynonymous rates in housekeeping genes, it is unlikely to be the major cause of the rate difference between the two types of genes. The codon usage pattern of housekeeping genes does not seem to differ from that of tissue-specific genes. Moreover, contrary to the old textbook concept, we found that approximately 74% of the housekeeping genes in our study belong to multigene families, not significantly different from that of the tissue-specific genes ( approximately 70%). Therefore, the stronger selective constraints on housekeeping genes are not due to a lower degree of genetic redundancy.

Animals↗

Differential codon usage for conserved amino acids: evidence that the serine codons TCN were primordial.

The availability of specialized sequence databanks for Escherichia coli, Saccharomyces cerevisiae and Bacillus subtilis made it possible to build a set of 105 protein-coding genes that are homologous in these three species. An analysis of the triplets at both the nucleotide and amino acid level revealed that the codon bias of some amino acids are significantly higher at conserved rather than at non-conserved positions. Comparisons of homologous genes in E. coli and Salmonella typhimurium, and in S. cerevisiae and Drosophila melanogaster, led to the same conclusion. A special case was made for serine in E. coli, whose major codon is AGC for non-conserved and TCC for conserved residues. We interpret this observation as evidence that the primordial codons for serine were TCN, while codons AGY appeared later. This conclusion is substantiated by an analysis of the codon usage of catalytic serine residues in ancient, ubiquitous and essential proteins (ATP synthases and topoisomerases). It is shown that in these proteins the proportion of the catalytic serine residues coded by TCN is significantly higher than the one expected from the overall codon usage of serine residues.

Amino Acid Sequence↗

Quantitative assessment of peptide sequence diversity in M13 combinatorial peptide phage display libraries.

Novel statistical methods have been developed and used to quantitate and annotate the sequence diversity within combinatorial peptide libraries on the basis of small numbers (1-200) of sequences selected at random from commercially available M13 p3-based phage display libraries. These libraries behave statistically as though they correspond to populations containing roughly 4.0+/-1.6% of the random dodecapeptides and 7.9+/-2.6% of the random constrained heptapeptides that are theoretically possible within the phage populations. Analysis of amino acid residue occurrence patterns shows no demonstrable influence on sequence censorship by Escherichia coli tRNA isoacceptor profiles or either overall codon or Class II codon usage patterns, suggesting no metabolic constraints on recombinant p3 synthesis. There is an overall depression in the occurrence of cysteine, arginine and glycine residues and an overabundance of proline, threonine and histidine residues. The majority of position-dependent amino acid sequence bias is clustered at three positions within the inserted peptides of the dodecapeptide library, +1, +3 and +12 downstream from the signal peptidase cleavage site. Conformational tendency measures of the peptides indicate a significant preference for inserts favoring a beta-turn conformation. The observed protein sequence limitations can primarily be attributed to genetic codon degeneracy and signal peptidase cleavage preferences. These data suggest that for applications in which maximal sequence diversity is essential, such as epitope mapping or novel receptor identification, combinatorial peptide libraries should be constructed using codon-corrected trinucleotide cassettes within vector-host systems designed to minimize morphogenesis-related censorship.

Amino Acids↗