Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Evidence that natural selection acts on silent mutation.

Analysis of nucleic acid sequence data of mammalian hemoglobin, yeast cytochrome c, and human interferon reveals strong biases in favor of specific codons. These biases do not appear to dissipate over time, suggesting that an indirect form of selection acts on silent mutations. The data are compatible with the "bootstrapping" hypothesis that silent mutations which alter the rate of evolution can hitchhike with traits whose appearance they facilitate. Selection involving modulating effects of codon usage on gene expression may also be involved, but the data appear to exclude simple maximization of gene expression.

Animals↗

Translation and stability of an Escherichia coli beta-galactosidase mRNA expressed under the control of pyruvate kinase sequences in Saccharomyces cerevisiae.

Plasmids were assembled in which the coding region of the pyruvate kinase (PYK) gene of Saccharomyces cerevisiae was replaced by that of the B-galactosidase (LacZ) gene from Escherichia coli. Analysis of the resultant, chimaeric transcripts from low copy number, centromeric plasmids indicated that this substitution caused a dramatic reduction in the steady-state level of the messenger RNA (mRNA). This fluctuation cannot be wholly accounted for by the 2-fold decrease in mRNA stability observed. This is consistent with the existence of a transcriptional Downstream Activation Site (DAS) within the PYK coding region, analogous to the DAS reported within the yeast phosphoglycerate kinase gene (PGK; Kingsman, S M et al. (1985) Biotech. Gen. Eng. Rev. 3, 377). At these low levels of heterologous gene expression, comparison of the distribution of PYK and PYK/LacZ transcripts across polysome gradients revealed no significant effect mediated by their striking disparity in codon usage. Nevertheless, upon increasing B-galactosidase mRNA levels, via manipulation of plasmid copy number, a distinct decline in ribosome loading was observed for the heterologous PYK/LacZ transcript which was not mirrored by either endogenous PYK transcripts or other yeast mRNAs of high (Ribosomal protein 1) or moderate (Actin) codon bias. However, high levels of the PYK/LacZ mRNA did affect the translation of an endogenous mRNA with poor codon bias (TRP2). The possible basis for this phenomenon is discussed.

Bacterial Proteins↗

[Codon optimization and expression in Pichia pastoris of E2 gene of classical swine fever virus].

Codon bias was one of the important parameter which influence heterogenous gene expression, optimizing codon sequence could improve expression level of heterogenous gene. In the preview study, wildtype E2 gene was expressed poorly in Pichia pastoris, in order to improve the expression level of E2 gene in Pichia pastoris, the low usage codons of E2 gene were mutated into high usage codons in Pichia pastoris by directed-mutagenesis based on PCR. The result showed that, compared with the results reported in preview study, the expression level of E2 gene in Pichia pastoris was improved observably by substituting 24 low usage codons of E2 gene for the high usage synonymous codons. It suggested the stragety to improve the expression of E2 gene in Pichia pastoris by codon optimization was successful.

Base Sequence↗

Mutation pressure, natural selection, and the evolution of base composition in Drosophila.

Genome sequencing in a number of taxa has revealed variation in nucleotide composition both among regions of the genome and among functional classes of sites in DNA. Mutational biases, biased gene conversion, and natural selection have been proposed as causes of this variation. Here, we review patterns of base composition in Drosophila DNA. Nucleotide composition in Drosophila melanogaster varys regionally, and base composition is correlated between introns and exons. Drosophila species also show striking patterns of non-random codon usage. Patterns of synonymous codon usage and the biochemistry of translation suggest that natural selection may act at 'silent' sites. A relationship between recombination rates and codon usage and comparisons of the evolutionary dynamics of silent mutations within and between species support natural selection discriminating among synonymous codons. The causes of regional base composition variation are less clear. Progress in functional studies of non-coding DNA, further investigations of genome patterns, and statistical tests based on evolutionary theory will lead to a greater understanding of the contributions of mutational processes and natural selection in patterning genome-wide nucleotide composition.

Animals↗

Synonymous codon usage and gene function are strongly related in Oryza sativa.

The relationship between codon usage and gene function was investigated while considering a dataset of 2106 nuclear genes of Oryza sativa. The results of standard chi(2) test and F-statistic showed that for every 59 synonymous codons, a strongly significant association with gene functional categories existed in rice, indicating that codon usage was generally coordinated with gene function whether it was at the level of individual amino acids or at the level of nucleotides. However, it could not be directly said that the use of every codons differed significantly between any two functional categories. Notably, there existed large difference both in selection for biased codons or selection intensity among functional categories. Therefore, we identified at least two classes of genes: one group of genes, mainly belonging to the "METABOLISM" category, was tended to use G- and/or C-ending codons while the other was more biased to choose codons ending with A and/or U. The latter group contained genes of various functions, especially those genes classified into the "Nuclear Structure" category. These observations will be more important for molecular genetic engineering and genome functional annotation.

Chromosome Mapping↗

Synonymous codon usage in Cryptosporidium parvum: identification of two distinct trends among genes.

The usage of alternative synonymous codons in the apicomplexan Cryptosporidium parvum has been investigated. A data set of 54 genes was analysed. Overall, A- and U-ending codons predominate, as expected in an A+T-rich genome. Two trends of codon usage variation among genes were identified using correspondence analysis. The primary trend is in the extent of usage of a subset of presumably translationally optimal codons, that are used at significantly higher frequencies in genes expected to be expressed at high levels. Fifteen of the 18 codons identified as optimal are more G+C-rich than the otherwise common codons, so that codon selection associated with translation opposes the general mutation bias. Among 40 genes with lower frequencies of these optimal codons, a secondary trend in G+C content was identified. In these genes, G+C content at synonymously variable third positions of codons is correlated with that in 5' and 3' flanking sequences, indicative of regional variation in G+C content, perhaps reflecting regional variation in mutational biases.

Animals↗

Compositional bias and size of genomes of human DNA viruses.

Genomes of 144 human DNA viruses were analyzed in the aspect of their compositional asymmetry. DNA viruses were divided into two groups according to their genome sizes. The analysis revealed that the level of guanine and cytosine (GC content) in the coding sequences of small genome DNA viruses was significantly lower than that of large genome DNA viruses. Because small genome viruses replicate their genomes using cellular enzymes, while large genome viruses use their own enzymes for genome replication, the two groups of viruses may be under different mutational bias and/or selection pressure. In these viruses, GC content at the third codon position correlated with GC content at the first and second codon position. However, the relationship in small genome DNA viruses was weaker than that in large genome DNA viruses, suggesting that their genome composition may be more strongly influenced by codon usage preference or restriction on amino acid composition.

Base Composition↗

The unusual nucleotide content of the HIV RNA genome results in a biased amino acid composition of HIV proteins.

Extremely high frequencies of the A nucleotide are found in the RNA genomes of the lentivirus group of retroviruses. It is presently unknown what molecular force is responsible for this A-pressure. In this manuscript, we demonstrate a correlation between this 'A-pressure' and the amino acid-usage of the lentivirus family. We compared the amino acid composition of the Gag and Pol proteins of the human immunodeficiency viruses type 1 and 2 (HIV-1 and HIV-2) with that of the second group of human retroviruses; the human T-cell leukemia viruses type I and II (HTLV-I and HTLV-II). Differences in total amino acid content correlate with the preference for A-rich codons in the HIV genome. A pair-wise comparison of homologous amino acid positions in the Pol proteins indicates that both conservative and non-conservative changes can be accounted for by this A-bias. The putative molecular mechanism underlying this A-pressure and the evolutionary consequences are discussed.

Amino Acid Sequence↗

Optimizing doped libraries by using genetic algorithms.

The insertion of random sequences into protein-encoding genes in combination with biological selection techniques has become a valuable tool in the design of molecules that have useful and possibly novel properties. By employing highly effective screening protocols, a functional and unique structure that had not been anticipated can be distinguished among a huge collection of inactive molecules that together represent all possible amino acid combinations. This technique is severely limited by its restriction to a library of manageable size. One approach for limiting the size of a mutant library relies on 'doping schemes', where subsets of amino acids are generated that reveal only certain combinations of amino acids in a protein sequence. Three mononucleotide mixtures for each codon concerned must be designed, such that the resulting codons that are assembled during chemical gene synthesis represent the desired amino acid mixture on the level of the translated protein. In this paper we present a doping algorithm that "reverse translates' a desired mixture of certain amino acids into three mixtures of mononucleotides. The algorithm is designed to optimally bias these mixtures towards the codons of choice. This approach combines a genetic algorithm with local optimization strategies based on the downhill simplex method. Disparate relative representations of all amino acids (and stop codons) within a target set can be generated. Optional weighing factors are employed to emphasize the frequencies of certain amino acids and their codon usage, and to compensate for reaction rates of different mononucleotide building blocks (synthons) during chemical DNA synthesis. The effect of statistical errors that accompany an experimental realization of calculated nucleotide mixtures on the generated mixtures of amino acids is simulated. These simulations show that the robustness of different optima with respect to small deviations from calculated values depends on their concomitant fitness. Furthermore, the calculations probe the fitness landscape locally and allow a preliminary assessment of its structure.

Algorithms↗

eCodonOpt: a systematic computational framework for optimizing codon usage in directed evolution experiments.

We present a systematic computational framework, eCodonOpt, for designing parental DNA sequences for directed evolution experiments through codon usage optimization. Given a set of homologous parental proteins to be recombined at the DNA level, the optimal DNA sequences encoding these proteins are sought for a given diversity objective. We find that the free energy of annealing between the recombining DNA sequences is a much better descriptor of the extent of crossover formation than sequence identity. Three different diversity targets are investigated for the DNA shuffling protocol to showcase the utility of the eCodonOpt framework: (i) maximizing the average number of crossovers per recombined sequence; (ii) minimizing bias in family DNA shuffling so that each of the parental sequence pair contributes a similar number of crossovers to the library; and (iii) maximizing the relative frequency of crossovers in specific structural regions. Each one of these design challenges is formulated as a constrained optimization problem that utilizes 0-1 binary variables as on/off switches to model the selection of different codon choices for each residue position. Computational results suggest that many-fold improvements in the crossover frequency, location and specificity are possible, providing valuable insights for the engineering of directed evolution protocols.

Aldose-Ketose Isomerases↗

Predicted highly expressed genes in the genomes of Streptomyces coelicolor and Streptomyces avermitilis and the implications for their metabolism.

Highly expressed genes in bacteria often have a stronger codon bias than genes expressed at lower levels, due to translational selection. In this study, a comparative analysis of predicted highly expressed (PHX) genes in the Streptomyces coelicolor and Streptomyces avermitilis genomes was performed using the codon adaptation index (CAI) as a numerical estimator of gene expression level. Although it has been suggested that there is little heterogeneity in codon usage in G+C-rich bacteria, considerable heterogeneity was found among genes in these two G+C-rich Streptomyces genomes. Using ribosomal protein genes as references, approximately 10% of the genes were predicted to be PHX genes using a CAI cutoff value of greater than 0.78 and 0.75 in S. coelicolor and S. avermitilis, respectively. The PHX genes showed good agreement with the experimental data on expression levels obtained from proteomic analysis by previous workers. Among 724 and 730 PHX genes identified from S. coelicolor and S. avermitilis, 368 are orthologue genes present in both genomes, which were mostly 'housekeeping' genes involved in cell growth. In addition, 61 orthologous gene pairs with unknown functions were identified as PHX. Only one polyketide synthase gene from each Streptomyces genome was predicted as PHX. Nevertheless, several key genes responsible for producing precursors for secondary metabolites, such as crotonyl-CoA reductase and propionyl-CoA carboxylase, and genes necessary for initiation of secondary metabolism, such as adenosylmethionine synthetase, were among the PHX genes in the two Streptomyces species. The PHX genes exclusive to each genome, and what they imply regarding cellular metabolism, are also discussed.

Bacterial Proteins↗

DNA and protein sequence homologies between the adhesins of Mycoplasma genitalium and Mycoplasma pneumoniae.

Mycoplasma genitalium and Mycoplasma pneumoniae are morphologically and serologically related pathogens that colonize the human host. Their successful parasitism appears to be dependent on the product, an adhesin protein, of a gene that is carried by each of these mycoplasmas. Here we describe the cloning and determine the sequence of the structural gene for the putative adhesin of M. genitalium and compare its sequence to the counterpart P1 gene of M. pneumoniae. Regions of homology that were consistent with the observed serological cross-reactivity between these adhesins were detected at both DNA and protein levels. However, the degree of homology between these two genes and their products was much higher than anticipated. Interestingly, the A + T content of the M. genitalium adhesin gene was calculated as 60.1%, which is substantially higher tham that of the P1 gene (46.5%). Comparisons of codon usage between the two organisms revealed that M. genitalium preferentially used A- and T-rich codons. A total of 65% of positions 3 and 56% of positions 1 in M. genitalium codons were either A or T, whereas M. pneumoniae utilized A or T for positions 3 and 1 at a frequency of 40 and 47%, respectively. The biased choice of the A- and T-rich codons in M. genitalium could also account for the preferential use of A- and T-rich codons in conservative amino acid substitutions found in the M. genitalium adhesin. These facts suggest that M. genitalium might have evolved independently of other human mycoplasma species, including M. pneumoniae.

Amino Acid Sequence↗

Synonymous codon usage in Escherichia coli: selection for translational accuracy.

In many organisms, selection acts on synonymous codons to improve translation. However, the precise basis of this selection remains unclear in the majority of species. Selection could be acting to maximize the speed of elongation, to minimize the costs of proofreading, or to maximize the accuracy of translation. Using several data sets, we find evidence that codon use in Escherichia coli is biased to reduce the costs of both missense and nonsense translational errors. Highly conserved sites and genes have higher codon bias than less conserved ones, and codon bias is positively correlated to gene length and production costs, both indicating selection against missense errors. Additionally, codon bias increases along the length of genes, indicating selection against nonsense errors. Doublet mutations or replacement substitutions do not explain our observations. The correlations remain when we control for expression level and for conflicting selection pressures at the start and end of genes. Considering each amino acid by itself confirms our results. We conclude that selection on synonymous codon use in E. coli is largely due to selection for translational accuracy, to reduce the costs of both missense and nonsense errors.

Codon↗

The pattern and distribution of immunoglobulin VH gene mutations in chronic lymphocytic leukemia B cells are consistent with the canonical somatic hypermutation process.

The overexpanded clone in most B-cell-type chronic lymphocytic leukemia (BCLL) patients expresses an immunoglobulin (Ig) heavy chain variable (V(H)) region gene with some level of mutation. While it is presumed that these mutations were introduced in the progenitor cell of the leukemic clone by the canonical somatic hypermutation (SHM) process, direct evidence of such is lacking. Nucleotide sequences of the Ig V(H) genes from 172 B-CLL patients were analyzed. Previously described V(H) gene usage biases were noted. As with canonical SHM, mutations found in B-CLL were more frequent in RGYW hot spots (mutations in an RGYW motif = 44.1%; germ line frequency of RGYW motifs = 25.6%) and favored transitions over transversions (transition-transversion ratio = 1.29). Significantly, transition preference was also noted when only mutations in the wobble position of degenerate codons were considered. Wobble positions are inherently unselected since regardless of change an identical amino acid is encoded; therefore, they represent a window into the nucleotide bias of the mutational mechanism. B-CLL V(H) mutations concentrated in complementarity-determining region 1 (CDR1) and CDR2, which exhibited higher replacement-to-silent ratios (CDR R/S, 4.60; framework region [FR] R/S, 1.72). These results are consistent with the notion that V(H) mutations in B-CLL cells result from canonical SHM and select for altered, structurally sound antigen receptors.

Amino Acid Motifs↗

Codon usage, genetic code and phylogeny of Dictyostelium discoideum mitochondrial DNA as deduced from a 7.3-kb region.

We have sequenced a region (7,376-bp) of the mitochondrial (mt) DNA (54 kb) of the cellular slime mold, Dictyostelium discoideum. From the DNA and amino-acid sequence comparisons with known sequences, genes for ATPase subunit 9 (ATP9), cytochrome b (CYTB), NADH dehydrogenase subunits 1, 3 and 6 (ND1, ND3 and ND6), small subunit rRNA (SSU rRNA) and seven tRNAs (Arg, Asn, Cys, Lys, f-Met, Met and Pro) have been identified. The sequenced region of the mtDNA has a high average A + T-content (70.8%). The A + T-content of protein-genes (73.6%) is considerably higher than that of RNA genes (61.3%). Even with the strong AT-bias, the genetic code employed is most probably the universal one. All seven tRNAs are able to form typical clover leaf structures. The molecular phylogenetic trees of CYTB and SSU rRNA suggest that D. discoideum is closer to green plants than to animals and fungi.

Amino Acid Sequence↗

Selection, mutations and codon usage in a bacterial model.

We present a statistical model of bacterial evolution based on the coupling between codon usage and tRNA abundance. Such a model interprets this aspect of the evolutionary process as a balance between the codon homogenization effect due to mutation process and the improvement of the translation phase due to natural selection. We develop a thermodynamical description of the asymptotic state of the model. The analysis of naturally occurring sequences shows that the effect of natural selection on codon bias affects genes whose products are largely required at maximal growth rate conditions or undergo rapid transient increases.

Bacteria↗

Domain-specific bias in arginine/lysine usage by protein toxins.

The content of lysine and arginine residues in a number of A-B type protein toxins has been examined. It is found that the A subunit, or its equivalent, often shows a strong bias in the type of basic amino acid residue used tending towards nearly exclusive use of either arginine or lysine rather than use of both, whereas the B subunit or its equivalent shows no such bias. Although arginine codons are GC-rich and lysine codons are AT-rich, the content of GC and AT in the genes coding for the toxins does not adequately explain this bias. Other explanations are discussed, including the possibility that the bias is linked to catalytic function or membrane interaction. Understanding this bias may yield valuable insights into toxin structure and function. Furthermore, identification of bias in sequences may be a useful tool for identifying new toxins and their domains.

ADP Ribose Transferases↗

Codon usage in streptococci.

Codon usage was analysed for 14 streptococcal genes or significant open reading frames and found to be different from that in Escherichia coli and Bacillus subtilis. In particular, the preferred use of WWT codons over WWC was inconsistent with the rule of optimal codon-anticodon interaction energy. On the other hand, for SSTC codons, adherence to this rule was better in streptococci than in E. coli. A preliminary codon bias table generated with the Pustell computer program for the analysed streptococcal genes may prove useful for the detection of protein coding regions in newly sequenced DNAs from both streptococci and staphylococci.

Bacillus subtilis↗