Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Horizontal transfer of accessory chromosomes in fungi - a regulated process for exchange of genetic material?

Horizontal transfer of entire chromosomes has been reported in several fungal pathogens, often significantly impacting the fitness of the recipient fungus. All documented instances of horizontal chromosome transfers (HCTs) showed a marked propensity for accessory chromosomes, consistently involving the transfer of an accessory chromosome while other chromosomes were seldom, if ever, co-transferred. The mechanisms underlying HCTs, as well as the factors regulating the specificity of HCTs for accessory chromosomes, remain unclear. In this perspective, we provide an overview of the observed propensity in reported cases of horizontal chromosome transfers. We hypothesize the existence of a signal that distinguishes mobile, i.e., horizontally transferred, accessory chromosomes from the rest of the donor genome. Recent findings in Metarhizium robertsii and Magnaporthe oryzae, suggest that a mobile accessory chromosome may contain putative histones and/or histone modifiers, which could generate such a signal. Based on this, we propose that mobile accessory chromosomes may encode the machinery required for their own horizontal transmission, implying that HCT could be a regulated process. Finally, we present evidence of substantial differences in codon usage bias between core and accessory chromosomes in 14 out of 19 analysed fungal species and strains. Such differences in codon usage bias could indicate past horizontal transfers of these accessory chromosomes. Interestingly, HCT was previously unknown for many of these species, suggesting that the horizontal transfer of accessory chromosomes may be more widespread than previously thought, and therefore an important factor in fungal genome evolution.

Gene Transfer, Horizontal↗

Comparative studies on codon usage pattern of chloroplasts and their host nuclear genes in four plant species.

A detailed comparison was made of codon usage of chloroplast genes with their host (nuclear) genes in the four angiosperm species Oryza sativa, Zea mays, Triticum aestivum and Arabidopsis thaliana. The average GC content of the entire genes, and at the three codon positions individually, was higher in nuclear than in chloroplast genes, suggesting different genomic organization and mutation pressures in nuclear and chloroplast genes. The results of Nc-plots and neutrality plots suggested that nucleotide compositional constraint had a large contribution to codon usage bias of nuclear genes in O. sativa, Z. mays, and T. aestivum, whereas natural selection was likely to be playing a large role in codon usage bias in chloroplast genomes. Correspondence analysis and chi-test showed that regardless of the genomic environment (species) of the host, the codon usage pattern of chloroplast genes differed from nuclear genes of their host species by their AU-richness. All the chloroplast genomes have predominantly A- and/or U-ending codons, whereas nuclear genomes have G-, C- or U-ending codons as their optimal codons. These findings suggest that the chloroplast genome might display particular characteristics of codon usage that are different from its host nuclear genome. However, one feature common to both chloroplast and nuclear genomes in this study was that pyrimidines were found more frequently than purines at the synonymous codon position of optimal codons.

Arabidopsis↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

Analysis of codon usage in beta-tubulin sequences of helminths.

Codon usage bias has been shown to be correlated with gene expression levels in many organisms, including the nematode Caenorhabditis elegans. Here, the codon usage (cu) characteristics for a set of currently available beta-tubulin coding sequences of helminths were assessed by calculating several indices, including the effective codon number (Nc), the intrinsic codon deviation index (ICDI), the P2 value and the mutational response index (MRI). The P2 value gives a measure of translational pressure, which has been shown to be correlated to high gene expression levels in some organisms, but it has not yet been analysed in that respect in helminths. For all but two of the C. elegans beta-tubulin coding sequences investigated, the P2 value was the only index that indicated the presence of codon usage bias. Therefore, we propose that in general the helminth beta-tubulin sequences investigated here are not expressed at high levels. Furthermore, we calculated the correlation coefficients for the cu patterns of the helminth beta-tubulin sequences compared with those of highly expressed genes in organisms such as Escherichia coli and C. elegans. It was found that beta-tubulin cu patterns for all sequences of members of the Strongylida were significantly correlated to those for highly expressed C. elegans genes. This approach provides a new measure for comparing the adaptation of cu of a particular coding sequence with that of highly expressed genes in possible expression systems.Finally, using the cu patterns of the sequences studied, a phylogenetic tree was constructed. The topology of this tree was very much in concordance with that of a phylogeny based on small subunit ribosomal DNA sequence alignments.

Animals↗

Microevolutionary divergence pattern of the segmentation gene hunchback in Drosophila.

To study the microevolutionary processes shaping the evolution of the segmentation gene hunchback (hb) from Drosophila melanogaster, we cloned and sequenced the gene from 12 isofemale lines representing wild-type populations of D. melanogaster, as well as from the closely related species Drosophila sechellia, Drosophila orena, and Drosophila yakuba. We find a relatively low degree of sequence variation in D. melanogaster (theta = 0.0017), which is, however, consistent with its chromosomal location in a region of low recombination. Tests of neutrality do not reject a neutral-evolution model for the whole region. However, pairwise tests with different subregions indicate that there is a relative excess of polymorphic sites in the leader and the intron. Codon usage pattern analysis shows a particularly biased codon usage in the highly conserved regions, which is in line with the hypothesis that selection on translational accuracy is the driving force behind such a bias. A comparison of the expression pattern of hb in different sibling species of D. melanogaster reveals some regulatory changes in D. yakuba, which could be interpreted as changes in the timing of secondary expression domains.

Animals↗

Adaptive basis of codon usage in the haploid moss Physcomitrella patens.

Patterns of codon usage bias were studied in the moss model species Physcomitrella patens. A total of 92 nuclear, protein coding genes were employed, and estimated levels of gene expression were tested for association with two measures of codon usage bias and other variables hypothesized to be associated with gene expression. Codon bias was found to be positively associated both with estimated levels of gene expression and GC content in the coding parts of studied genes. However, GC content in noncoding parts, that is, introns and 5' and 3' untranslated regions (UTRs), was not associated with estimated levels of gene expression. It is argued that codon bias is not shaped by mutational bias, but rather by weak natural selection for translational efficiency in P. patens. The possible role of life history characteristics in shaping patterns of codon usage in this species is discussed.

Bryopsida↗

Low codon bias and high rates of synonymous substitution in Drosophila hydei and D. melanogaster histone genes.

We have evaluated codon usage bias in Drosophila histone genes and have obtained the nucleotide sequence of a 5,161-bp D. hydei histone gene repeat unit. This repeat contains genes for all five histone proteins (H1, H2a, H2b, H3, and H4) and differs from the previously reported one by a second EcoRI site. These D. hydei repeats have been aligned to each other and to the 5.0-kb (i.e., long) and 4.8-kb (i.e., short) histone repeat types from D. melanogaster. In each species, base composition at synonymous sites is similar to the average genomic composition and approaches that in the small intergenic spacers of the histone gene repeats. Accumulation of synonymous changes at synonymous sites after the species diverged is quite high. Both of these features are consistent with the relatively low codon usage bias observed in these genes when compared with other Drosophila genes. Thus, the generalization that abundantly expressed genes in Drosophila have high codon bias and low rates of silent substitution does not hold for the histone genes.

Animals↗

Synonymous codon usage in Kluyveromyces lactis.

The nature and variation of synonymous codon usage in 47 open reading frames from Kluyveromyces lactis have been investigated. Using multivariate statistical analysis, a single major trend among K. lactis genes was identified that differentiates among genes by expression level: highly expressed genes have high codon usage bias, while genes of low expression level have low bias. A relatively minor secondary trend differentiates among genes according to G+C content at silent sites. In these respects, K. lactis is similar to both Saccharomyces cerevisiae and Candida albicans, and the same 'optimal' codons appear to be selected in highly expressed genes in all three species. In addition, silent sites in K. lactis and S. cerevisiae have similar G+C contents, but in C. albicans genes they are more A+T-rich. Thus, in all essential features, codon usage in K. lactis is very similar to that in S. cerevisiae, even though silent sites in genes compared between these two species have undergone sufficient mutation to be saturated with changes. We conclude that the factors influencing overall codon usage, namely mutational biases and the abundances of particular tRNAs, have not diverged between the two species. Nevertheless, in a few cases, codon usage differs between homologous genes from K. lactis and S. cerevisiae. The strength of codon usage bias in cytochrome c genes differs considerably, presumably because of different expression patterns in the two species. Two other, linked, genes have very different G+C content at silent sites in the two species, which may be a reflection of their chromosomal locations. Correspondence analysis was used to identify two open reading frames with highly atypical codon usage that are probably not genes.

Base Sequence↗

Support vector machine for classification of meiotic recombination hotspots and coldspots in Saccharomyces cerevisiae based on codon composition.

BACKGROUND: Meiotic double-strand breaks occur at relatively high frequencies in some genomic regions (hotspots) and relatively low frequencies in others (coldspots). Hotspots and coldspots are receiving increasing attention in research into the mechanism of meiotic recombination. However, predicting hotspots and coldspots from DNA sequence information is still a challenging task. RESULTS: We present a novel method for classification of hot and cold ORFs located in hotspots and coldspots respectively in Saccharomyces cerevisiae, using support vector machine (SVM), which relies on codon composition differences. This method has achieved a high classification accuracy of 85.0%. Since codon composition is a fusion of codon usage bias and amino acid composition signals, the ability of these two kinds of sequence attributes to discriminate hot ORFs from cold ORFs was also investigated separately. Our results indicate that neither codon usage bias nor amino acid composition taken separately performed as well as codon composition. Moreover, our SVM based method was applied to the full genome: We predicted the hot/cold ORFs from the yeast genome by using cutoffs of recombination rate. We found that the performance of our method for predicting cold ORFs is not as good as that for predicting hot ORFs. Besides, we also observed a considerable correlation between meiotic recombination rate and amino acid composition of certain residues, which probably reflects the structural and functional dissimilarity between the hot and cold groups. CONCLUSION: We have introduced a SVM-based novel method to discriminate hot ORFs from cold ones. Applying codon composition as sequence attributes, we have achieved a high classification accuracy, which suggests that codon composition has strong potential to be used as sequence attributes in the prediction of hot and cold ORFs.

Algorithms↗

A model of protein translation including codon bias, nonsense errors, and ribosome recycling.

We present and analyse a model of protein translation at the scale of an individual messenger RNA (mRNA) transcript. The model we develop is unique in that it incorporates the phenomena of ribosome recycling and nonsense errors. The model conceptualizes translation as a probabilistic wave of ribosome occupancy traveling down a heterogeneous medium, the mRNA transcript. Our results show that the heterogeneity of the codon translation rates along the mRNA results in short-scale spikes and dips in the wave. Nonsense errors attenuate this wave on a longer scale while ribosome recycling reinforces it. We find that the combination of nonsense errors and codon usage bias can have a large effect on the probability that a ribosome will completely translate a transcript. We also elucidate how these forces interact with ribosome recycling to determine the overall translation rate of an mRNA transcript. We derive a simple cost function for nonsense errors using our model and apply this function to the yeast (Saccharomyces cervisiae) genome. Using this function we are able to detect position dependent selection on codon bias which correlates with gene expression levels as predicted a priori. These results indirectly validate our underlying model assumptions and confirm that nonsense errors can play an important role in shaping codon usage bias.

Animals↗

Phylogenetic estimation under codon models can be biased by codon usage heterogeneity.

In theory, codon models that account for the dependence of nucleotide substitutions between codon positions as well as differences between synonymous and non-synonymous changes best describe the sequence evolution in protein coding genes. However, in practice we know little about the degree to which violations of the assumptions of codon model-based estimates occur, and how significant these artifacts may be. In nucleotide-based phylogenies from first and second codon positions in a concatenated plastid gene data set, two distantly related taxa--dinoflagellate and haptophyte plastids--were robustly grouped together. This artifactual grouping is attributed to the parallel heterogeneity in leucine (Leu) and serine (Ser) codon usages in the data set. Here, by using this data set, we demonstrated that codon-based phylogenetic estimations are seriously biased, robustly uniting the dinoflagellate and haptophyte plastids into a monophyletic clade, when the model assumption of homogeneity of codon composition was violated. Our results suggest that similar phylogenetic artifacts may occur via codon usage heterogeneity in any amino acids in codon model-based estimations. We advise that homogeneity in codon usage across taxa in a data set be confirmed before codon model-based phylogenetic estimation is attempted.

Codon↗

Genomic analysis of influenza A viruses, including avian flu (H5N1) strains.

This study was designed to conduct genomic analysis in two steps, such as the overall relative synonymous codon usage (RSCU) analysis of the five virus species in the orthomyxoviridae family, and more intensive pattern analysis of the four subtypes of influenza A virus (H1N1, H2N2, H3N2, and H5N1) which were isolated from human population. All the subtypes were categorized by their isolated regions, including Asia, Europe, and Africa, and most of the synonymous codon usage patterns were analyzed by correspondence analysis (CA). As a result, influenza A virus showed the lowest synonymous codon usage bias among the virus species of the orthomyxoviridae family, and influenza B and influenza C virus were followed, while suggesting that influenza A virus might have an advantage in transmitting across the species barrier due to their low codon usage bias. The ENC values of the host-specific HA and NA genes represented their different HA and NA types very well, and this reveals that each influenza A virus subtype uses different codon usage patterns as well as the amino acid compositions. In NP, PA and PB2 genes, most of the virus subtypes showed similar RSCU patterns except for H5N1 and H3N2 (A/HK/1774/1999) subtypes which were suspected to be transmitted across the species barrier, from avian and porcine species to human beings, respectively. This distinguishable synonymous codon usage patterns in non-human origin viruses might be useful in determining the origin of influenza A viruses in genomic levels as well as the serological tests. In this study, all the process, including extracting sequences from GenBank flat file and calculating codon usage values, was conducted by Java codes, and these bioinformatics-related methods may be useful in predicting the evolutionary patterns of pandemic viruses.

Africa↗

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals↗

Codon usage and intragenic position.

Data on codon usage bias in E. coli are re-examined with respect to intragenic position. The bias is less extreme near the beginning than in the rest of the gene, particularly in highly expressed genes. This is contrary to the previous finding that there is a linear decline in codon usage bias with position along weakly expressed genes but little or no change in bias along highly expressed genes. The effect is not confined to genes coding for proteins with leader peptides, as suggested earlier (Burns and Beacham, 1985). There is some evidence of a similar but smaller effect in yeast.

Codon↗

Genetic immunization with codon-optimized equine infectious anemia virus (EIAV) surface unit (SU) envelope protein gene sequences stimulates immune responses in ponies.

In the context of DNA vaccines the native equine infectious anemia virus (EIAV)-envelope gene has proven to be an extremely weak immunogen in horses probably because the RNA transcripts are poorly expressed owing to an unusual codon-usage bias, the possession of multiple RNA splice sites and potential adenosine-rich RNA instability elements. To overcome these problems a synthetic version of sequences encoding the EIAV surface unit (SU) envelope glycoprotein was produced (SYNSU) in which the codon-usage bias was modified to conform to that of highly expressed horse and human genes. In transfected COS-1 cell cultures, the steady state expression levels of SYNSU were at least 30-fold greater than equivalent native SU sequences. More importantly, EIAV-specific humoral and lymphocyte proliferation responses were induced in ponies immunized with a mammalian expression vector encoding SYNSU. However, these immunological responses were unable to confer protection against infection with a virulent EIAV strain.

Amino Acid Sequence↗

Molecular evolution between Drosophila melanogaster and D. simulans: reduced codon bias, faster rates of amino acid substitution, and larger proteins in D. melanogaster.

Both natural selection and mutational biases contribute to variation in codon usage bias within Drosophila species. This study addresses the cause of codon bias differences between the sibling species, Drosophila melanogaster and D. simulans. Under a model of mutation-selection-drift, variation in mutational processes between species predicts greater base composition differences in neutrally evolving regions than in highly biased genes. Variation in selection intensity, however, predicts larger base composition differences in highly biased loci. Greater differences in the G+C content of 34 coding regions than 46 intron sequences between D. melanogaster and D. simulans suggest that D. melanogaster has undergone a reduction in selection intensity for codon bias. Computer simulations suggest at least a fivefold reduction in Nes at silent sites in this lineage. Other classes of molecular change show lineage effects between these species. Rates of amino acid substitution are higher in the D. melanogaster lineage than in D. simulans in 14 genes for which outgroup sequences are available. Surprisingly, protein sizes are larger in D. melanogaster than in D. simulans in the 34 genes compared between the two species. A substantial fraction of silent, replacement, and insertion/deletion mutations in coding regions may be weakly selected in Drosophila.

Amino Acids↗

Mitochondrial genomes of Clymenella torquata (Maldanidae) and Riftia pachyptila (Siboglinidae): evidence for conserved gene order in annelida.

Mitochondrial genomes are useful tools for inferring evolutionary history. However, many taxa are poorly represented by available data. Thus, to further understand the phylogenetic potential of complete mitochondrial genome sequence data in Annelida (segmented worms), we examined the complete mitochondrial sequence for Clymenella torquata (Maldanidae) and an estimated 80% of the sequence of Riftia pachyptila (Siboglinidae). These genomes have remarkably similar gene orders to previously published annelid genomes, suggesting that gene order is conserved across annelids. This result is interesting, given the high variation seen in the closely related Mollusca and Brachiopoda. Phylogenetic analyses of DNA sequence, amino acid sequence, and gene order all support the recent hypothesis that Sipuncula and Annelida are closely related. Our findings suggest that gene order data is of limited utility in annelids but that sequence data holds promise. Additionally, these genomes show AT bias (approximately 66%) and codon usage biases but have a typical gene complement for bilaterian mitochondrial genomes.

Animals↗