Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Maximizing transcription efficiency causes codon usage bias.

The rate of protein synthesis depends on both the rate of initiation of translation and the rate of elongation of the peptide chain. The rate of initiation depends on the encountering rate between ribosomes and mRNA; this rate in turn depends on the concentration of ribosomes and mRNA. Thus, patterns of codon usage that increase transcriptional efficiency should increase mRNA concentration, which in turn would increase the initiation rate and the rate of protein synthesis. An optimality model of the transcriptional process is presented with the prediction that the most frequently used ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide in the cellular matrix where mRNA is transcribed. This prediction is supported by four kinds of evidence. First, A-ending codons are the most frequently used synonymous codons in mitochondria, where ATP is much more abundant than that of the three other ribonucleotides. Second, A-ending codons are more frequently used in mitochondrial genes than in nuclear genes. Third, protein genes from organisms with a high metabolic rate use more A-ending codons and have higher A content in their introns than those from organisms with a low metabolic rate.

Animals↗

Annotation pattern of ESTs from Spodoptera frugiperda Sf9 cells and analysis of the ribosomal protein genes reveal insect-specific features and unexpectedly low codon usage bias.

MOTIVATION: A whole set of Expressed Sequence Tags (ESTs) from the Sf9 cell line of Spodoptera frugiperda is presented here for the first time. By this way we want to identify both conserved and specific genes of this pest species. We also expect from this analysis to find a class of protein sequences providing a tool to explore genomic features and phylogeny of Lepidoptera. RESULTS: The ESTs display both housekeeping as well as developmentally regulated genes, and a high percentage of sequences with unknown function. Among the identified ORFs, almost all ribosomal proteins (RPs) were found with high EST redundancy and hence sequence accuracy. The codon usage found among RP genes is in average surprisingly much less biased in Lepidoptera than in other organisms. Other Spodoptera genes also displayed a low bias, suggesting a general genome expression feature in this Lepidoptera. We also found that the L35A and L36 RP sequences, respectively, display 40 and 10 amino-acid insertions, both being present only in insects. Sequence analysis suggests that they are probably not subjected to a strong selective pressure and may be good phylogenetic markers for Lepidoptera. Most interestingly, the Lepidoptera sequences of 9 RP genes displayed a specific signature different from the canonical one. We conclude that the RP family allows valuable comparative genomics and phylogeny of Lepidoptera. AVAILABILITY: All EST sequence data are available from the private 'Spodo-Base' upon request.

Abstracting and Indexing↗

Biased codon usage near intron-exon junctions: selection on splicing enhancers, splice-site recognition or something else?

Two groups recently argued that, in human genes, synonymous sites near intron-exon junctions undergo selection for correct splicing. However, neither study controlled for the possibility of an underlying nucleotide bias at the ends of exons. In this article, we show that generalized A and T enrichment exists, which could be independent of splicing regulation. Evidence for selection between synonymous codons that are associated with splicing enhancers remains after controlling for this bias, whereas support for cryptic splice-site avoidance is diminished.

Animals↗

Strand asymmetry and codon usage bias in the chloroplast genome of Euglena gracilis.

It is shown that the two strands of the chloroplast genome from Euglena gracilis are asymmetric with regards to nucleotide composition. This asymmetry switches at both the origin of replication and a location that is halfway around the circular genome from the origin. In both halves of the genome the leading strand is G+T-rich, having a bias toward G over C and T over A, and the lagging strand is A+C-rich. This asymmetry is probably the result of a difference in mutation dynamics between the leading and lagging strands. In addition to composition asymmetry, the two strands differ with regards to coding content. In both halves of the genome the vast majority of genes are coded by the leading strand. These two aspects of strand asymmetry are then applied to a statistical test for selection on codon usage. The results indicate that selection on codon usage is limited to genes on the leading strand; no gene on the A+C-rich lagging strand shows evidence for selection, suggesting that highly expressed genes are coded predominantly on the strand of DNA that is the leading strand during replication. On the basis of these observations it is proposed that the coding strand bias is generated by selection to code highly expressed genes on the leading strand to coordinate the direction of replication and transcription, thereby increasing the potential rate of both reactions.

Animals↗

DNA G+C content of the third codon position and codon usage biases of human genes.

The human genome, as in other eukaryotes, has a wide heterogeneity in the DNA base composition. The evolutionary basis for this heterogeneity has been unknown. A previous study of the human genome (846 genes analyzed) has shown that, in the major range of the G+C content in the third codon position (0.25-0.75), biases from the Parity Rule 2 (PR2) among the synonymous codons of the four-codon amino acids are similar except in the highest G+C range (Sueoka, N., 1999. Translation-coupled violation of Parity Rule 2 in human genes is not the cause of heterogeneity of the DNA G+C content of third codon position. Gene 238, 53-58.). PR2 is an intra-strand rule where A=T and G=C are expected when there are no biases between the two complementary strands of DNA in mutation and selection rates (substitution rates). In this study, 14,026 human genes were analyzed. In addition, the third codon positions of two-codon amino acids were analyzed. New results show the following: (a) The G+C contents of the third codon position of human genes are scattered in the G+C range of 0.22-0.96 in the third codon position. (b) The PR2 biases are similar in the range of 0.25-0.75, whereas, in the high G+C range (0.75-0.96; 13% of the genes), the PR2-bias fingerprints are different from those of the major range. (c) Unlike the PR2 biases, the G+C contents of the third codon position for both four-codon and two-codon amino acids are all correlated almost perfectly with the G+C content of the third codon position over the total G+C ranges. These results support the notion that the directional mutation pressure, rather than the directional selection pressure, is mainly responsible for the heterogeneity of the G+C content of the third codon position.

Amino Acids↗

Nucleotide substitution pattern in rice paralogues: implication for negative correlation between the synonymous substitution rate and codon usage bias.

Understanding the correlation between synonymous substitution rate and GC content is essential to decipher the gene evolution. However, it has been controversial on their relationship. We analyzed the GC content and synonymous substitution rate in 1092 paralogues produced by two large-scale duplication events in the rice genome. According to the GC content at the third codon sites (GC3), the paralogues were classified into GC3-rich and GC3-poor genes. By referring to their outgroup sequences, we inferred the last common ancestor of sister paralogues and, consequently, calculated the average synonymous substitution rate for two gene classes. The results suggest that average synonymous substitution rate is lower in GC3-rich genes than that in GC3-poor genes, indicating that the synonymous substitution rate is negatively correlated with GC content in the rice genome. Through characterizing the synonymous nucleotide substitution pattern, we found a strong synonymous nucleotide substitution frequency bias from AT to GC in GC3-rich genes. This indicates possible limitations of commonly used methods developed to estimate the synonymous substitution rate. Their estimates might produce misleading results on correlation between the synonymous substitution rate and GC content.

Base Composition↗

Protein evolution and codon usage bias on the neo-sex chromosomes of Drosophila miranda.

The neo-sex chromosomes of Drosophila miranda constitute an ideal system to study the effects of recombination on patterns of genome evolution. Due to a fusion of an autosome with the Y chromosome, one homolog is transmitted clonally. Here, I compare patterns of molecular evolution of 18 protein-coding genes located on the recombining neo-X and their homologs on the nonrecombining neo-Y chromosome. The rate of protein evolution has significantly increased on the neo-Y lineage since its formation. Amino acid substitutions are accumulating uniformly among neo-Y-linked genes, as expected if all loci on the neo-Y chromosome suffer from a reduced effectiveness of natural selection. In contrast, there is significant heterogeneity in the rate of protein evolution among neo-X-linked genes, with most loci being under strong purifying selection and two genes showing evidence for adaptive evolution. This observation agrees with theory predicting that linkage limits adaptive protein evolution. Both the neo-X and the neo-Y chromosome show an excess of unpreferred codon substitutions over preferred ones and no difference in this pattern was observed between the chromosomes. This suggests that there has been little or no selection maintaining codon bias in the D. miranda lineage. A change in mutational bias toward AT substitutions also contributes to the decline in codon bias. The contrast in patterns of molecular evolution between amino acid mutations and synonymous mutations on the neo-sex-linked genes can be understood in terms of chromosome-specific differences in effective population size and the distribution of selective effects of mutations.

Animals↗

Substitution rates in Drosophila nuclear genes: implications for translational selection.

The relationships between synonymous and nonsynonymous substitution rates and between synonymous rate and codon usage bias are important to our understanding of the roles of mutation and selection in the evolution of Drosophila genes. Previous studies used approximate estimation methods that ignore codon bias. In this study we reexamine those relationships using maximum-likelihood methods to estimate substitution rates, which accommodate the transition/transversion rate bias and codon usage bias. We compiled a sample of homologous DNA sequences at 83 nuclear loci from Drosophila melanogaster and at least one other species of Drosophila. Our analysis was consistent with previous studies in finding that synonymous rates were positively correlated with nonsynonymous rates. Our analysis differed from previous studies, however, in that synonymous rates were unrelated to codon bias. We therefore conducted a simulation study to investigate the differences between approaches. The results suggested that failure to properly account for multiple substitutions at the same site and for biased codon usage by approximate methods can lead to an artifactual correlation between synonymous rate and codon bias. Implications of the results for translational selection are discussed.

Animals↗

The Vitreoscilla hemoglobin gene: molecular cloning, nucleotide sequence and genetic expression in Escherichia coli.

Vitreoscilla hemoglobin is involved in oxygen metabolism of this bacterium, possibly in an unusual role for a microbe. We have isolated the Vitreoscilla hemoglobin structural gene from a pUC19 genomic library using mixed oligodeoxy-nucleotide probes based on the reported amino acid sequence of the protein. The gene is expressed in Escherichia coli from its natural promoter as a major cellular protein. The nucleotide sequence, which is in complete agreement with the known amino acid sequence of the protein, suggests the existence of promoter and ribosome binding sites with a high degree of homology to consensus E. coli upstream sequences. In the case of at least some amino acids, a codon usage bias can be detected which is different from the biased codon usage pattern in E. coli. The downstream sequence exhibits homology with the 3' end sequences of several plant leghemoglobin genes. E. coli cells expressing the gene contain greater than fivefold more heme than controls.

Amino Acid Sequence↗

Codon usage and bias among individual genes of the coccidia and piroplasms.

Codon usage has been analysed in individual gene sequences, derived from a variety of parasitic protozoa in the class Sporozoa of the phylum Apicomplexa using metric multidimensional scaling. The two groups of codon usage patterns detected reflect the two main subgroups of organisms studied (the coccidia and the piroplasms), and it is the pattern of usage of synonymous codons that has the largest influence on overall codon usage in the individual genes, rather than being the pattern of amino acid composition of the gene product. The magnitude of the codon usage bias in the sequences was determined using three commonly used indices-NC, GC3S and B. In general, although relatively low levels of codon usage bias were detected in these gene sequences, codon usage bias does explain at least some of the codon usage patterns observed. Codon usage bias was observed to be dependent on the overall base composition of the genes analysed, which in turn was reflected in the types of codons that were either over- or under-represented in the nucleotide sequences. In keeping with observations on prokaryotic organisms, it is speculated that the codon usage patterns detected in these parasitic protozoa are the result of directional mutation pressure on the base composition of the genomic DNA.

Animals↗

Directionality of point mutation and 5-methylcytosine deamination rates in the chimpanzee genome.

BACKGROUND: The pattern of point mutation is important for studying mutational mechanisms, genome evolution, and diseases. Previous studies of mutation direction were largely based on substitution data from a limited number of loci. To date, there is no genome-wide analysis of mutation direction or methylation-dependent transition rates in the chimpanzee or its categorized genomic regions. RESULTS: In this study, we performed a detailed examination of mutation direction in the chimpanzee genome and its categorized genomic regions using 588,918 SNPs whose ancestral alleles could be inferred by mapping them to human genome sequences. The C-->T (G-->A) changes occurred most frequently in the chimpanzee genome. Each type of transition occurred approximately four times more frequently than each type of transversion. Notably, the frequency of C-->T (G-->A) was the highest in exons among the genomic categories regardless of whether we calculated directly, normalized with the nucleotide content, or removed the SNPs involved in the CpG effect. Moreover, the directionality of the point mutation in exons and CpG islands were opposite relative to their corresponding intergenic regions, indicating that different forces govern the nucleotide changes. Our analysis suggests that the GC content is not in equilibrium in the chimpanzee genome. Further quantitative analysis revealed that the 5-methylcytosine deamination rates at CpG sites were highly dependent on the local GC content and the lengths of SNP flanking sequences and varied among categorized genomic regions. CONCLUSION: We present the first mutational spectrum, estimated by three different approaches, in the chimpanzee genome. Our results provide detailed information on recent nucleotide changes and methylation-dependent transition rates in the chimpanzee genome after its split from the human. These results have important implications for understanding genome composition evolution, mechanisms of point mutation, and other genetic factors such as selection, biased codon usage, biased gene conversion, and recombination.

5-Methylcytosine↗

Correlation between codon usage, regional genomic nucleotide composition, and amino acid composition in the cytochrome P-450 gene superfamily.

The codon usage bias of 110 mammalian cytochrome P-450 genes has been determined and analyzed in relation to a variety of genetic, biochemical, and physiological parameters. In those P-450 genes exhibiting biased usage the preferred codons generally do not differ among the four species examined (rat, rabbit, man, and mouse) or from the predominantly used codons identified for all sequenced genes in a recent data base analysis (Wada et al. (1992) Nucleic Acids Res. 20 (Suppl.), 2111-2118). Codon usage bias does not correlate with evolutionary relationships, evolutionary age, or with the extent of evolutionary conservation of orthologous proteins; there is no obvious correlation with the level of expression of a given P-450, with its inducibility, nor with its physiologic role; and neither the preferred codons nor the degree of bias differ for P-450s expressed in different tissues. Codon usage bias does correlate with the C+G content at the codon third position, and thus preferred codons usually end in C or G; for those P-450s for which gene sequences are available this bias also correlates with the C + G content of the intronic and flanking regions of these genes. Moreover, a lesser increase in the C + G content at the codon first and second positions is also evident in genes located in regions of high C + G content; this leads to predictable differences in the amino acid compositions of P-450 enzymes that correlate with genomic nucleotide composition and the degree of bias in codon usage.

Amino Acids↗

The correlation between synonymous and nonsynonymous substitutions in Drosophila: mutation, selection or relaxed constraints?

Codon usage bias, the preferential use of particular codons within each codon family, is characteristic of synonymous base composition in many species, including Drosophila, yeast, and many bacteria. Preferential usage of particular codons in these species is maintained by natural selection acting largely at the level of translation. In Drosophila, as in bacteria, the rate of synonymous substitution per site is negatively correlated with the degree of codon usage bias, indicating stronger selection on codon usage in genes with high codon bias than in genes with low codon bias. Surprisingly, in these organisms, as well as in mammals, the rate of synonymous substitution is also positively correlated with the rate of nonsynonymous substitution. To investigate this correlation, we carried out a phylogenetic analysis of substitutions in 22 genes between two species of Drosophila, Drosophila pseudoobscura and D. subobscura, in codons that differ by one replacement and one synonymous change. We provide evidence for a relative excess of double substitutions in the same species lineage that cannot be explained by the simultaneous mutation of two adjacent bases. The synonymous changes in these codons also cannot be explained by a shift to a more preferred codon following a replacement substitution. We, therefore, interpret the excess of double codon substitutions within a lineage as being the result of relaxed constraints on both kinds of substitutions in particular codons.

Animals↗

Association of the phi nucleotide with codon bias, amino acid usage and expressivity: differences between Bacillus subtilis and Escherichia coli.

By measuring the non-randomness in Shine-Dalgarno regions it was recently shown that the compositional non-randomness peaks approximately 10 nucleotides upstream of the start codons. This position, termed the phi position, was furthermore shown to be associated with certain characteristics of the gene/protein and start codon usage. This raises the question whether codon usage in general is associated with the phi position. In this study, the connection between the phi nucleotide and general codon usage, both gene-wide and at the level of individual amino acids, was studied in Eschericia coli and Bacillus subtilis. E. coli but not B. subtilis shows a strong general association between the phi position and codon usage bias. In both species, the genes with higher expressivity show stronger conservation in the Shine-Dalgarno region compared to the genes with lower expressivity.

Amino Acids↗

Selection on codon usage in Drosophila americana.

Synonymous codons are not used at random, significantly influencing the base composition of the genome. The selection-mutation-drift model proposes that this bias reflects natural selection in favor of a subset of preferred codons. Previous estimates in Drosophila of the intensity of selective forces involved seem too large to be reconciled with theoretical predictions of the level of codon bias. This probably results from confounding effects of the demographic histories of the species concerned. We have studied three species of the virilis group of Drosophila, which are more likely to satisfy the assumptions of the evolutionary models. We analyzed the patterns of polymorphism and divergence in a sample of 18 genes and applied a new method for estimating the intensity of selection on synonymous mutations based on the frequencies of unpreferred mutations among polymorphic sites. This yielded estimates of selection intensities (N(e)s) of the order of 0.65, which is more compatible with the observed levels of codon bias. Our results support the action of both selection and mutational bias on codon usage bias and suggest that codon usage and genome base composition in the D. americana lineage are in approximate equilibrium. Biased gene conversion may also contribute to the observed patterns.

Animals↗

Synonymous codon usage in Bacillus subtilis reflects both translational selection and mutational biases.

Codon usage data for 56 Bacillus subtilis genes show that synonymous codon usage in B. subtilis is less biased than in Escherichia coli, or in Saccharomyces cerevisiae. Nevertheless, certain genes with a high codon bias can be identified by correspondence analysis, and also by various indices of codon bias. These genes are very highly expressed, and a general trend (a decrease) in codon bias across genes seems to correspond to decreasing expression level. This, then, may be a general phenomenon in unicellular organisms. The unusually small effect of translational selection on the pattern of codon usage in lowly expressed genes in B. subtilis yields similar dinucleotide frequencies among different codon positions, and on complementary strands. These patterns could arise through selection on DNA structure, but more probably are largely determined by mutation. This prevalence of mutational bias could lead to difficulties in assessing whether open reading frames encode proteins.

Bacillus subtilis↗