Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Codon adaptation index as a measure of dominating codon bias.

UNLABELLED: We propose a simple algorithm to detect dominating synonymous codon usage bias in genomes. The algorithm is based on a precise mathematical formulation of the problem that lead us to use the Codon Adaptation Index (CAI) as a 'universal' measure of codon bias. This measure has been previously employed in the specific context of translational bias. With the set of coding sequences as a sole source of biological information, the algorithm provides a reference set of genes which is highly representative of the bias. This set can be used to compute the CAI of genes of prokaryotic and eukaryotic organisms, including those whose functional annotation is not yet available. An important application concerns the detection of a reference set characterizing translational bias which is known to correlate to expression levels; in this case, the algorithm becomes a key tool to predict gene expression levels, to guide regulatory circuit reconstruction, and to compare species. The algorithm detects also leading-lagging strands bias, GC-content bias, GC3 bias, and horizontal gene transfer. The approach is validated on 12 slow-growing and fast-growing bacteria, Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. AVAILABILITY: http://www.ihes.fr/~materials.

Adaptation, Physiological↗

Mitochondrial genomes of Dactylogyrus wunderi (Monopisthocotyla: Dactylogyridae): structural features, codon usage patterns, and phylogenetic implications.

BACKGROUND: Codon usage bias (CUB) is a common phenomenon reported among many species and genes, but its unique characteristics in the mitochondrial genome of class Monopisthocotyla remain unknown. METHODS: The complete mitochondrial genome of Dactylogyrus wunderi was sequenced and characterized, and the mitochondrial genome compositions and CUB of six Dactylogyrus species and 35 Monopisthocotyla species were analyzed using bioinformatics methods. RESULTS: The mitochondrial genome of D. wunderi is a typical circular structure in length of 14,920 bp. The A&#x2009;+&#x2009;T contents of the six Dactylogyrus species (58.4% &#xb1; 5.7%) were significantly lower than that of Monopisthocotyla species (71.0% &#xb1; 5.80%, p&#x2009;<&#x2009;0.01). Neutrality plot analysis showed slopes of 0.3136 and 0.389 in the six Dactylogyrus and the 35 Monopisthocotyla species, respectively. Furthermore, 98.3% and 77.4% of the genes in the six Dactylogyrus and the Monopisthocotyla species, respectively, had effective number of codons (ENC) higher than 35, but 23.3% and 0.5% genes of ENC ratio ranged from -&#x2009;0.05 to 0.05 in the six Dactylogyrus and Monopisthocotyla species. Phylogenetic analysis revealed that, within the context of the sampled taxa, the families of Monopisthocotyla were monophyletic groups, except for Ancyrocephalidae. CONCLUSIONS: The nucleotide composition had AT base bias in Monopisthocotyla, and natural selection was the main factor affecting CUB in the mitochondrial genomes of Monopisthocotyla species. These results provided insights into the factors affecting CUB in Monopisthocotyla species and deepened our insight of phylogeny, evolution, and codon usage of Monopisthocotyla.

Genome, Mitochondrial↗

Synonymous codon usage in Lactococcus lactis: mutational bias versus translational selection.

In this study codon usage bias of all experimentally known genes of Lactococcus lactis has been analyzed. Since Lactococcus lactis is an AT rich organism, it is expected to occur A and/or T at the third position of codons and detailed analysis of overall codon usage data indicates that A and/or T ending codons are predominant in this organism. However, multivariate statistical analyses based both on codon count and on relative synonymous codon usage (RSCU) detect a large number of genes, which are supposed to be highly expressed are clustered at one end of the first major axis, while majority of the putatively lowly expressed genes are clustered at the other end of the first major axis. It was observed that in the highly expressed genes C and T ending codons are significantly higher than the lowly expressed genes and also it was observed that C ending codons are predominant in the duets of highly expressed genes, whereas the T endings codons are abundant in the quartets. Abundance of C and T ending codons in the highly expressed genes suggest that, besides, compositional biases, translational selection are also operating in shaping the codon usage variation among the genes in this organism as observed in other compositionally skewed organisms. The second major axis generated by correspondence analysis on simple codon counts differentiates the genes into two distinct groups according to their hydrophobicity values, but the same analysis computed with relative synonymous codon usage values could not discriminate the genes according to the hydropathy values. This suggests that amino acid composition exerts constraints on codon usage in this organism. On the other hand the second major axis produced by correspondence analysis on RSCU values differentiates the genes into two groups according to the synonymous codon usage for cysteine residues (rarest amino acids in this organism), which is nothing but a artifactual effect induced by the RSCU values. Other factors such as length of the genes and the positions of the genes in the leading and lagging strand of replication have practically no influence in the codon usage variation among the genes in this organism.

Codon↗

Global mRNA stability is not associated with levels of gene expression in Drosophila melanogaster but shows a negative correlation with codon bias.

A multitude of factors contribute to the regulation of gene expression in living cells. The relationship between codon usage bias and gene expression has been extensively studied, and it has been shown that codon bias may have adaptive significance in many unicellular and multicellular organisms. Given the central role of mRNA in post-transcriptional regulation, we hypothesize that mRNA stability is another important factor associated either with positive or negative regulation of gene expression. We have conducted genome-wide studies of the association between gene expression (measured as transcript abundance in public EST databases), mRNA stability, codon bias, GC content, and gene length in Drosophila melanogaster. To remove potential bias of gene length inherently present in EST libraries, gene expression is measured as normalized transcript abundance. It is demonstrated that codon bias and GC content in second codon position are positively associated with transcript abundance. Gene length is negatively associated with transcript abundance. The stability of thermodynamically predicted mRNA secondary structures is not associated with transcript abundance, but there is a negative correlation between mRNA stability and codon bias. This finding does not support the hypothesis that codon bias has evolved as an indirect consequence of selection favoring thermodynamically stable mRNA molecules.

Animals↗

The cytochrome b region in the mitochondrial DNA of the ant Tetraponera rufoniger: sequence divergence in Hymenoptera may be associated with nucleotide content.

Polymerase chain reaction (PCR) followed by sequencing of single-stranded DNA yielded sequence information from the cytochrome b (cyt b) region in mitochondrial DNA from the ant Tetraponera rufoniger. Compared with the cyt b genes from Apis mellifera, Drosophila melanogaster, and D. yakuba, the overall A+T content (A+T%) of that of T. rufoniger is lower (69.9% vs 80.7%, 74.2%, and 73.9%, respectively) than those of the other three. The codon usage in the cyt b gene of T. rufoniger is biased although not as much as in A. mellifera, D. melanogaster, and D. yakuba; T. rufoniger has eight unused codons whereas D. melanogaster, D. yakuba, and A. mellifera have 21, 20, and 23, respectively. The inferred cyt b polypeptide chain (PPC) of T. rufoniger has diverged at least as much from a common ancestor with D. yakuba as has that of A. mellifera (approximately 3.5 vs approximately 2.9). Despite the lower A+T%, the relative frequencies of amino acids in the cyt b PPC of T. rufoniger are significantly (P < 0.05) associated with the content of adenine and thymine (A+T%) and size of codon families. The mitochondrially located cytochrome oxidase subunit II genes (CO-II) of endopterygote insects have significantly higher average A+T% (approximately 75%) than those of exopterygous (approximately 69%) and paleopterous (approximately 69%) insects. The increase in A+T% of endopterygote insects occurred in Upper Carboniferous and coincided with a significant acceleration of PPC divergence. However, acceleration of PPC divergence is not significantly correlated with the increase of the A+T% (P > 0.1). The high A+T%, the biased codon usage, and the increased PPC divergence of Hymenoptera can in that respect most easily be explained by directional mutation pressure which began in the Upper Carboniferous and still occurs in most members of the order. Given the roughly identical A+T% of the cyt b and CO-II genes from the other insects whose DNA sequences are known (A. mellifera, D. melanogaster, and D. yakuba), it seems most likely that the A+T% of T. rufoniger declined secondarily within the last 100 Myr as a result of a reduced directional mutation pressure.

Amino Acid Sequence↗

Correlated evolution of synonymous and nonsynonymous sites in Drosophila.

Recent work has shown that Drosophila melanogaster genes with fast-evolving nonsynonymous sites have lower codon usage bias. This pattern has been attributed to interference between positive selection at nonsynonymous sites and weak selection on codon usage. Here we have looked for this correlation in a much larger and less biased dataset, comprising 630 gene pairs from D. melanogaster and D. yakuba. We confirmed that there is a negative correlation between the rate of nonsynonymous substitutions (d(N)) and codon bias in D. melanogaster. We then tested the interference hypothesis and other alternative explanations, including one involving gene expression. We found that d(N) indeed correlates with the level of gene expression. Given that gene expression is a strong determinant of codon bias, the relationship between d(N) and codon bias might be a by-product of gene expression. However, our tests show that none of the hypotheses we consider seem to explain the data fully.

Amino Acid Substitution↗

Codon usage patterns in cytochrome oxidase I across multiple insect orders.

Synonymous codon usage bias is determined by a combination of mutational biases, selection at the level of translation, and genetic drift. In a study of mtDNA in insects, we analyzed patterns of codon usage across a phylogeny of 88 insect species spanning 12 orders. We employed a likelihood-based method for estimating levels of codon bias and determining major codon preference that removes the possible effects of genome nucleotide composition bias. Three questions are addressed: (1) How variable are codon bias levels across the phylogeny? (2) How variable are major codon preferences? and (3) Are there phylogenetic constraints on codon bias or preference? There is high variation in the level of codon bias values among the 88 taxa, but few readily apparent phylogenetic patterns. Bias level shifts within the lepidopteran genus Papilio are most likely a result of population size effects. Shifts in major codon preference occur across the tree in all of the amino acids in which there was bias of some level. The vast majority of changes involves double-preference models, however, and shifts between single preferred codons within orders occur only 11 times. These shifts among codons in double-preference models are phylogenetically conservative.

Animals↗

Comparative analysis of the base composition and codon usages in fourteen mycobacteriophage genomes.

To study the possible codon usage and base composition variation in the bacteriophages, fourteen mycobacteriophages were used as a model system here and both the parameters in all these phages and their plating bacteria, M. smegmatis had been determined and compared. As all the organisms are GC-rich, the GC contents at third codon positions were found in fact higher than the second codon positions as well as the first + second codon positions in all the organisms indicating that directional mutational pressure is strongly operative at the synonymous third codon positions. Nc plot indicates that codon usage variation in all these organisms are governed by the forces other than compositional constraints. Correspondence analysis suggests that: (i) there are codon usage variation among the genes and genomes of the fourteen mycobacteriophages and M. smegmatis, i.e., codon usage patterns in the mycobacteriophages is phage-specific but not the M. smegmatis-specific; (ii) synonymous codon usage patterns of Barnyard, Che8, Che9d, and Omega are more similar than the rest mycobacteriophages and M. smegmatis; (iii) codon usage bias in the mycobacteriophages are mainly determined by mutational pressure; and (iv) the genes of comparatively GC rich genomes are more biased than the GC poor genomes. Translational selection in determining the codon usage variation in highly expressed genes can be invoked from the predominant occurrences of C ending codons in the highly expressed genes. Cluster analysis based on codon usage data also shows that there are two distinct branches for the fourteen mycobacteriophages and there is codon usage variation even among the phages of each branch.

Bacteriophages↗

Codon usage and G + C content in Bradyrhizobium japonicum genes are not uniform.

To date, the sequences of 45 Bradyrhizobium japonicum genes are known. This provides sufficient information to determine their codon usage and G + C content. Surprisingly, B. japonicum nodulation and NifA-regulated genes were found to have a less biased codon usage and a lower G + C content than genes not belonging to these two groups. Thus, the coding regions of nodulation genes and NifA-regulated genes could hardly be identified in codon preference plots whereas this was not difficult with other genes. The codon frequency table of the highly biased genes was used in a codon preference plot to analyze the RSRj alpha 9 sequence which is an insertion sequence (IS)-like element. The plot helped identify a new open reading frame (ORF355) that escaped previous detection because of two sequencing errors. These were now corrected. The deduced gene product of ORF355 in RSRj alpha 9 showed extensive similarity to a putative protein encoded by an ORF in the T-DNA of Agrobacterium rhizogenes. The DNA sequences bordering both ORFs showed inverted repeats and potential target site duplications which supported the assumption that they were IS-like elements.

Amino Acid Sequence↗

Compositional heterogeneity and patterns of molecular evolution in the Drosophila genome.

The rates and patterns of molecular evolution in many eukaryotic organisms have been shown to be influenced by the compartmentalization of their genomes into fractions of distinct base composition and mutational properties. We have examined the Drosophila genome to explore relationships between the nucleotide content of large chromosomal segments and the base composition and rate of evolution of genes within those segments. Direct determination of the G + C contents of yeast artificial chromosome clones containing inserts of Drosophila melanogaster DNA ranging from 140-340 kb revealed significant heterogeneity in base composition. The G + C content of the large segments studied ranged from 36.9% G + C for a clone containing the hunchback locus in polytene region 85, to 50.9% G + C for a clone that includes the rosy region in polytene region 87. Unlike other organisms, however, there was no significant correlation between the base composition of large chromosomal regions and the base composition at fourfold degenerate nucleotide sites of genes encompassed within those regions. Despite the situation seen in mammals, there was also no significant association between base composition and rate of nucleotide substitution. These results suggest that nucleotide sequence evolution in Drosophila differs from that of many vertebrates and does not reflect distinct mutational biases, as a function of base composition, in different genomic regions. Significant negative correlations between codon-usage bias and rates of synonymous site divergence, however, provide strong support for an argument that selection among alternative codons may be a major contributor to variability in evolutionary rates within Drosophila genomes.

Animals↗

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus↗

Comparative study of translation termination sites and release factors (RF1 and RF2) in procaryotes.

Translation termination is catalyzed by release factors that recognize stop codons. However, previous works have shown that in some bacteria, the termination process also involves bases around stop codons. Recently, Ito et al. analyzed release factors and identified the amino acids therein that recognize stop codons. However, the amino acids that recognize bases around stop codons remain unclear. To identify the candidate amino acids that recognize the bases around stop codons, we aligned the protein sequences of the release factors of various bacteria and searched for amino acids that were conserved specifically in the sequence of bacteria that seemed to regulate translation termination by bases around stop codons. As a result, species having several highly conserved residues in RF1 and RF2 showed positive correlations between their codon usage bias and conservation of the bases around the stop codons. In addition, some of the residues were located very close to the SPF motif, which deciphers stop codons. These results suggest that these conserved amino acids enable the release factors to recognize the bases around the stop codons.

Amino Acid Sequence↗

Error minimization explains the codon usage of highly expressed genes in Escherichia coli.

Different organisms use synonymous codons with different preferences. Several measures have been introduced to compute the extent of codon usage bias within a gene or genome, among which the codon adaptation index (CAI) has been shown to be well correlated with mRNA levels of Escherichia coli. In this work an error adaptation index (eAI) is introduced, which estimates the level at which a gene can tolerate the effects of mistranslations. It is shown that the eAI has a strong correlation with CAI, as well as with mRNA levels, which suggests that the codons of highly expressed genes are selected so that mistranslation would have the minimum possible effect on the structure and function of the related proteins.

Base Composition↗

A periodic pattern of mRNA secondary structure created by the genetic code.

Single-stranded mRNA molecules form secondary structures through complementary self-interactions. Several hypotheses have been proposed on the relationship between the nucleotide sequence, encoded amino acid sequence and mRNA secondary structure. We performed the first transcriptome-wide in silico analysis of the human and mouse mRNA foldings and found a pronounced periodic pattern of nucleotide involvement in mRNA secondary structure. We show that this pattern is created by the structure of the genetic code, and the dinucleotide relative abundances are important for the maintenance of mRNA secondary structure. Although synonymous codon usage contributes to this pattern, it is intrinsic to the structure of the genetic code and manifests itself even in the absence of synonymous codon usage bias at the 4-fold degenerate sites. While all codon sites are important for the maintenance of mRNA secondary structure, degeneracy of the code allows regulation of stability and periodicity of mRNA secondary structure. We demonstrate that the third degenerate codon sites contribute most strongly to mRNA stability. These results convincingly support the hypothesis that redundancies in the genetic code allow transcripts to satisfy requirements for both protein structure and RNA structure. Our data show that selection may be operating on synonymous codons to maintain a more stable and ordered mRNA secondary structure, which is likely to be important for transcript stability and translation. We also demonstrate that functional domains of the mRNA [5'-untranslated region (5'-UTR), CDS and 3'-UTR] preferentially fold onto themselves, while the start codon and stop codon regions are characterized by relaxed secondary structures, which may facilitate initiation and termination of translation.

3' Untranslated Regions↗

Mutation and selection at silent and replacement sites in the evolution of animal mitochondrial DNA.

Two patterns are presented that illustrate the interaction of mutation and selection in the evolution of animal mtDNA: 1) variation among taxa in the ratio of polymorphism to divergence (rpd) at silent and replacement sites in protein-coding genes, and 2) strand-differences in polymorphism and divergence at 'silent' sites that suggest a mutation-selection balance in the evolution of codon usage. Cytochrome b data from GenBank show that about half of the species pairs tested have a significant excess of amino acid polymorphism, relative to divergence. The remaining half of species pairs do not depart from neutrality, but generally do show an excess of amino acid polymorphism. Sequences from Drosophila pseudoobscura displaying a signature of an expanding population show a slight, but non-significant, deficiency of amino acid polymorphism suggestive of recently intensified selection on mildly deleterious mutations. Genes whose reading frames lie on the major coding strand of Drosophila mtDNA show a preponderance of T- > C substitutions, while genes encoded on the minor strand experience more A- > G than T- > C substitutions between species at both silent and replacement sites. However, silent mutations at third codon positions are introduced into the population in proportions opposite to those observed as fixed differences between species (e.g., an excess of T- > C polymorphisms are found at the ND5 gene on the minor coding strand). The high A + T content of insect mtDNAs imposes strong codon usage bias favoring A-ending and T-ending codons resulting in a distinct mutation-selection balance for genes encoded on opposites strands. Thus, at both replacement and silent sites, mutations that appear to be constrained in terms of divergence between species are in excess within species. The data suggest that mildly deleterious mutations are common in mitochondrial genes. A test of this, and a competing, hypothesis is proposed that requires additional sequence surveys of polymorphism and divergence. An important challenge is to tease apart the impact of mutation and selection on levels of polymorphism versus divergence in a genome that does not generally recombine.

Animals↗

Codon usage of human DNA viruses and its similarity to certain host genes.

Codon usages of DNA viruses had previously been shown to associate with their genome size. Codon usage of various human DNA viruses was compared to those of human genes to further understand viral codon usage and its roles in viral-host interaction. Codon usage bias in both large and small genome human DNA viruses was dominantly driven by translation selection. Non-optimal codon usage in small DNA viruses showed similarity to cell cycle-related genes, whereas codon usage of large DNA viruses was more diverse, herpesviruses showed more heterogeneity than human adenoviruses, while poxviruses showed a clear bimodal pattern. Some of the large DNA viruses such as herpes simplex and molluscum contagiosum viruses showed more optimal codon usage. Enrichment analysis identified some groups of human genes with similar codon usage to each group of these viruses. These host genes with similarity in codon usages to those of viruses may be efficiently expressed in infected cells and involved in their life cycle, pathogenesis and/or immune evasion.

Humans↗

[Use of the hygromycin phosphotransferase gene as the dominant selective marker for Chlamydomonas reinhardtii transformation].

The hygromycin phosphotransferase gene (hpt) from E. coli under the control of the SV40 early promoter was used as a dominant selectable marker for transformation of Chlamydomonas reinhardtii. Cells were transformed by electroporation (pulse length, 2 ms, field strength, 1 kV/cm). The culture growth phase was a crucial parameter for transformation (optimal density approximately 10(6) cells/ml). It was possible to obtain approximately 10(3) Hyg-resistant colonies under these conditions. Foreign DNA integrated into the Chlamydomonas genome was maintained for at least 8 months but the Hyg-resistant phenotype of the transformed clones was unstable. The frequency of codon usage in the hpt gene was compared with the one in Chlamydomonas nuclear genes. It is supposed that highly biased codon usage in Chlamydomonas does not preclude expression. Advantages of this selection system for studying Chlamydomonas transformation by heterologous genes are discussed.

Animals↗

Synonymous codon usage in environmental chlamydia UWE25 reflects an evolutional divergence from pathogenic chlamydiae.

Publication of the complete genome sequence for the Acanthamoeba sp. endosymbiont UWE25 has illuminated the evolution history of chlamydiae. In this study, the codon usage bias in UWE25 and five other species of pathogenic chlamydiae was calculated. It was found that genomic composition constraints are the major source of codon usage variation in UWE25. This result is different from the former observation in pathogenic chlamydiae, whose genomic base composition is more unbiased. Four other factors, such as strand-specific mutational bias, natural selection acting at the level of translation, hydropathy level of each protein and the conservation level of amino acids also have influence in shaping the codon usage in these six species to some extent. Further analysis suggests that the high stability of the UWE25 genome partially account for the difference in codon usage pattern between environmental and pathogenic chlamydiae. Moreover, our results imply that the replicational selection pressure in pathogenic chlamydiae is stronger than that in UWE25. Analyzing the codon usage pattern in the environmental chlamydia and comparing it with that of the pathogenic chlamydiae may provide clues how the chlamydiae have evolved from their common ancestor.

Amino Acids↗