Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Dinucleotide frequencies and codon usage in jawless and cartilaginous fishes.

Dinucleotide frequencies and codon usage in terms of strong-weak codon choices were examined in gene coding regions of 5 jawless and cartilaginous fish species. These dinucleotide frequencies were then compared to gene-coding regions from a vertebrate and an invertebrate species. These primitive vertebrate fishes exhibited species specificity in the hierarchy of dinucleotide frequencies. The most frequently occurring dinucleotide varied among coding regions of jawless and cartilaginous fishes, but it always contained G. Of the 16 dinucleotides, TA had the lowest frequency of occurrence in all species, and it had considerably lower frequencies in jawless fish genes than in cartilaginous fish genes. Dinucleotide frequency analysis suggested CpG conversion to TG in ray genes. Strong-weak codon usage analysis indicated that all 5 fish species used strong-weak-strong bonding codons most frequently; furthermore, each species used any-weak-strong codons in greater than expected levels. Gene coding regions from all 5 species exhibited a bias toward strong bonding nucleotides at codon position 3, with the greatest bias in sea lamprey genes. This bias may reflect the overall G+C content of localized regions of chromosomal DNA in which these genes reside.

Animals↗

RNA editing in Arabidopsis mitochondria effects 441 C to U changes in ORFs.

On the basis of the sequence of the mitochondrial genome in the flowering plant Arabidopsis thaliana, RNA editing events were systematically investigated in the respective RNA population. A total of 456 C to U, but no U to C, conversions were identified exclusively in mRNAs, 441 in ORFs, 8 in introns, and 7 in leader and trailer sequences. No RNA editing was seen in any of the rRNAs or in several tRNAs investigated for potential mismatch corrections. RNA editing affects individual coding regions with frequencies varying between 0 and 18.9% of the codons. The predominance of RNA editing events in the first two codon positions is not related to translational decoding, because it is not correlated with codon usage. As a general effect, RNA editing increases the hydrophobicity of the coded mitochondrial proteins. Concerning the selection of RNA editing sites, little significant nucleotide preference is observed in their vicinity in comparison to unedited C residues. This sequence bias is, per se, not sufficient to specify individual C nucleotides in the total RNA population in Arabidopsis mitochondria.

Arabidopsis↗

Effect of strong directional selection on weakly selected mutations at linked sites: implication for synonymous codon usage.

The fixation of weakly selected mutations can be greatly influenced by strong directional selection at linked loci. Here, I investigate a two-locus model in which weakly selected, reversible mutations occur at one locus and recurrent strong directional selection occurs at the other locus. This model is analogous to selection on codon usage at synonymous sites linked to nonsynonymous sites under strong directional selection. Two approximations obtained here describe the expected frequency of the weakly selected preferred alleles at equilibrium. These approximations, as well as simulation results, show that the level of codon bias declines with an increasing rate of substitution at the strongly selected locus, as expected from the well-understood theory that selection at one locus reduces the efficacy of selection at linked loci. These solutions are used to examine whether the negative correlation between codon bias and nonsynonymous substitution rates recently observed in Drosophila can be explained by this hitchhiking effect. It is shown that this observation can be reasonably well accounted for if a large fraction of the nonsynonymous substitutions on genes in the data set are driven by strong directional selection.

Alleles↗

Patterns of nucleotide substitution among simultaneously duplicated gene pairs in Arabidopsis thaliana.

We characterized rates and patterns of synonymous and nonsynonymous substitution in 242 duplicated gene pairs on chromosomes 2 and 4 of Arabidopsis thaliana. Based on their collinear order along the two chromosomes, the gene pairs were likely duplicated contemporaneously, and therefore comparison of genetic distances among gene pairs provides insights into the distribution of nucleotide substitution rates among plant nuclear genes. Rates of synonymous substitution varied 13.8-fold among the duplicated gene pairs, but 90% of gene pairs differed by less than 2.6-fold. Average nonsynonymous rates were approximately fivefold lower than average synonymous rates; this rate difference is lower than that of previously studied nonplant lineages. The coefficient of variation of rates among genes was 0.65 for nonsynonymous rates and 0.44 for synonymous rates, indicating that synonymous and nonsynonymous rates vary among genes to roughly the same extent. The causes underlying rate variation were explored. Our analyses tentatively suggest an effect of physical location on synonymous substitution rates but no similar effect on nonsynonymous rates. Nonsynonymous substitution rates were negatively correlated with GC content at synonymous third codon positions, and synonymous substitution rates were negatively correlated with codon bias, as observed in other systems. Finally, the 242 gene pairs permitted investigation of the processes underlying divergence between paralogs. We found no evidence of positive selection, little evidence that paralogs evolve at different rates, and no evidence of differential codon usage or third position GC content.

Arabidopsis↗

Strand-specific nucleotide composition bias in echinoderm and vertebrate mitochondrial genomes.

The gene organization of starfish mitochondrial DNA is identical with that of the sea urchin counterpart except for a reported inversion of an approximately 4.6-kb segment containing two structural genes for NADH dehydrogenase subunits 1 and 2 (ND 1 and ND 2). When the codon usage of each structural gene in starfish, sea urchin, and vertebrate mitochondrial DNAs is examined, it is striking that codons ending in T and G are preferentially used more for heavy strand-encoded genes, including starfish ND 1 and ND 2, than for light strand-encoded genes, including sea urchin ND 1 and ND 2. On the contrary, codons ending in A and C are preferentially used for the light strand-encoded genes rather than for the heavy strand-encoded ones. Moreover, G-U base pairs are more frequently found in the possible secondary structures of heavy strand-encoded tRNAs than in those of light strand-encoded tRNAs. These observations suggest the existence of a certain constraint operating on mitochondrial genomes from various animal phyla, which results in the accumulation of G and T on one strand, and A and C on the other.

Amino Acid Sequence↗

The complete mitochondrial genome of the articulate brachiopod Terebratalia transversa.

We sequenced the complete mitochondrial DNA (mtDNA) of the articulate brachiopod Terebratalia transversa. The circular genome is 14,291 bp in size, relatively small compared with other published metazoan mtDNAs. The 37 genes commonly found in animal mtDNA are present; the size decrease is due to the truncation of several tRNA, rRNA, and protein genes, to some nucleotide overlaps, and to a paucity of noncoding nucleotides. Although the gene arrangement differs radically from those reported for other metazoans, some gene junctions are shared with two other articulate brachiopods, Laqueus rubellus and Terebratulina retusa. All genes in the T. transversa mtDNA, unlike those in most metazoan mtDNAs reported, are encoded by the same strand. The A+T content (59.1%) is low for a metazoan mtDNA, and there is a high propensity for homopolymer runs and a strong base-compositional strand bias. The coding strand is quite G+T-rich, a skew that is shared by the confamilial (laqueid) species L. rubellus but is the opposite of that found in T. retusa, a cancellothyridid. These compositional skews are strongly reflected in the codon usage patterns and the amino acid compositions of the mitochondrial proteins, with markedly different usages being observed between T. retusa and the two laqueids. This observation, plus the similarity of the laqueid noncoding regions to the reverse complement of the noncoding region of the cancellothyridid, suggests that an inversion that resulted in a reversal in the direction of first-strand replication has occurred in one of the two lineages. In addition to the presence of one noncoding region in T. transversa that is comparable with those in the other brachiopod mtDNAs, there are two others with the potential to form secondary structures; one or both of these may be involved in the process of transcript cleavage.

Amino Acid Sequence↗

Codon usage and lateral gene transfer in Bacillus subtilis.

Bacillus subtilis possesses three classes of genes, differing by their codon preference. One class corresponds to prophages or prophage-like elements, indicative of the existence of systematic lateral gene transfer in this organism. The nature of the selection pressure that operates on codon bias is beginning to be understood.

Bacillus subtilis↗

The role of the codon first letter in the relationship between genomic GC content and protein amino acid composition.

Analysis of the statistical distribution of amino acid compositions within 22 protein families shows that a GC bias generally affects proteins with a variety of functions from the extreme thermophile Thermus. This results in evident enrichment in amino acids of the group L, V, A, P, R and G and underrepresentation of amino acids of the group I, M, F, S, T, C and W. The strong amino acid composition biases noted in Thermus proteins are not related to thermoadaptation; they were also found in mesophilic homologues encoded by GC-rich genes. The results of a comparative analysis on large samples of translated sequences from 30 organisms, representing the three major kingdoms of life and including extremophiles, indicate a universal correlation between the usage of particular amino acids and the genomic GC content. It is concluded that the codon first letter plays a dominant role in translating the genomic GC signature into protein amino acid composition and sequences.

Amino Acid Sequence↗

Gene organization features in A/T-rich organisms.

Several species have genomes in which the four nucleotides are not equally represented (Glöckner 2000). Interestingly, shifts to very high A/T or G/C levels can occur in several distinct branches of the tree of life. The underlying reasons for these shifts therefore may be of different origin. Now entire chromosome sequences from two different A/T-rich genomes, Dictyostelium discoideum and Plasmodium falciparum, are available (Bowman et al. 1999; Gardner et al. 2002; Glöckner et al. 2002). This gives us the opportunity to investigate how a high A/T content may influence the signals that are the landmarks for gene specification. We found that, in contrast with most known metazoan and plant genomes, splice signals contain, little information other than the canonical GT-AG dinucleotides. Intron lengths in A/T rich organisms, on the other hand, are comparable to those of other lower eukaryotes. Intergenic regions show, dependent on the orientation of adjacent genes, a size pattern with a ratio of 1 (3'-3') to 2 (3'-5') to 3 (5'-5'). Overall, gene organization patterns seem not to be influenced by the A/T bias. Surprisingly, the slightly higher A/T content of the P. falciparum genome compared to that of D. discoideum (80.1 versus 77.4%) is not achieved by increased A/T richness in intergenic regions. Instead both the shift of the nucleotide usage in coding regions to A/T-rich codons and the longer intergenic regions make an equal contribution to the higher A/T content in this organism.

AT Rich Sequence↗

Strong associations between gene function and codon usage.

The association between codon usage and gene function was analyzed in the complete genomes of Eschericia coli, Bacillus subtilis, Lactococcus lactis and Campylobacter jejuni, using the functional annotation provided by NCBI. Two distinctly different ways of quantifying codon usage were used in the analysis. By using contingency tables it was found that for most amino acids a highly significant association with gene function exists for all species, indicating that codon usage at the level of individual amino acids is generally closely coordinated with gene function. By computing the effective number of codons in the annotated genes and comparing the median values in groups of different gene functions it was shown for all species that codon bias gene by gene also differs.

Amino Acids↗

Codon usage in Plasmodium vivax nuclear genes.

Codon usage in Plasmodium vivax nuclear genes was analysed and compared with that in Plasmodium falciparum nuclear genes. Preferred codons were determined for P. vivax. Unlike P. falciparum, P. vivax genes are about 15% less A+T rich in the coding regions, with no obvious A+T bias at the third position of the codons. The amino-acid composition of P. vivax gene products is also different from that of P. falciparum. These results provide valuable information to facilitate gene cloning as well as expression and transfection studies for P. vivax.

Amino Acids↗

The guanine and cytosine content of genomic DNA and bacterial evolution.

The genomic guanine and cytosine (G + C) content of eubacteria is related to their phylogeny. The G + C content of various parts of the genome (protein genes, stable RNA genes, and spacers) reveals a positive linear correlation with the G + C content of their genomic DNA. However, the plotted correlation slopes differ among various parts of the genome or among the first, second, and third positions of the codons depending on their functional importance. Facts suggest that biased mutation pressure, called A X T/G X C pressure, has affected whole DNA during evolution so as to determine the genomic G + C content in a given bacterium. The role of A X T/G X C pressure in diversification of bacterial DNA sequences and codon usage patterns is discussed in the perspective of the neutral theory of molecular evolution.

Biological Evolution↗

Support vector machines for separation of mixed plant-pathogen EST collections based on codon usage.

MOTIVATION: Discovery of host and pathogen genes expressed at the plant-pathogen interface often requires the construction of mixed libraries that contain sequences from both genomes. Sequence identification requires high-throughput and reliable classification of genome origin. When using single-pass cDNA sequences difficulties arise from the short sequence length, the lack of sufficient taxonomically relevant sequence data in public databases and ambiguous sequence homology between plant and pathogen genes. RESULTS: A novel method is described, which is independent of the availability of homologous genes and relies on subtle differences in codon usage between plant and fungal genes. We used support vector machines (SVMs) to identify the probable origin of sequences. SVMs were compared to several other machine learning techniques and to a probabilistic algorithm (PF-IND) for expressed sequence tag (EST) classification also based on codon bias differences. Our software (Eclat) has achieved a classification accuracy of 93.1% on a test set of 3217 EST sequences from Hordeum vulgare and Blumeria graminis, which is a significant improvement compared to PF-IND (prediction accuracy of 81.2% on the same test set). EST sequences with at least 50 nt of coding sequence can be classified using Eclat with high confidence. Eclat allows training of classifiers for any host-pathogen combination for which there are sufficient classified training sequences. AVAILABILITY: Eclat is freely available on the Internet (http://mips.gsf.de/proj/est) or on request as a standalone version. CONTACT: friedel@informatik.uni-muenchen.de.

Algorithms↗

Expression of tetanus toxin fragment C in E. coli: high level expression by removing rare codons.

Tetanus toxin fragment C had been previously expressed in Escherichia coli at 3-4% cell protein. The codon bias for tetanus toxin in Clostridium tetani is very different from that of highly expressed homologous genes in E. coli, resulting in the presence of many rare E. coli codons in the sequence encoding fragment C. We have replaced the coding sequence by sequence optimized for codon usage in E. coli, and show that the expression of fragment C is increased. Although the level of mRNA also increased this appeared to be a secondary consequence of more efficient translation. Complete sequence replacement increased expression to approximately 11-14% cell protein but only after the promoter strength had been improved.

Amino Acid Sequence↗

Mammalian mutation pressure, synonymous codon choice, and mRNA degradation.

The usage of synonymous codons (SCs) in mammalian genes is highly correlated with local base composition and is therefore thought to be determined by mutation pressure. The usage is nonetheless structured. For instance, mammals share with Saccharomyces and Drosophila most preferences for the C-ending over the G-ending codon (or vice versa) within each fourfold-degenerate SC family and the fact that their SCs are placed along coding regions in ways that minimize the number of T|A and C|G dinucleotides ("|" being the codon boundary). TA and CG underrepresentations are observed everywhere in the mammalian genome affecting the SC usage, the amino acid composition of proteins, and the primary structure of introns and noncoding DNA. While the rarity of CG is ascribed to the high mutability of this dinucleotide, the rarity of TA in coding regions is considered adaptive because UA dinucleotides are cleaved by endoribonucleases. Here we present in vivo experimental evidence indicating that the number of T|A and/or C|G dinucleotides of a human gene can affect strongly the expression level and degradation of its mRNA. Our results are consistent with indirect evidence produced by other workers and with the detailed work that has been devoted to characterize UA cleavage in vitro and in vivo. We conclude that SC choice can influence strongly mRNA function and gene expression through effects not directly related to the codon-anticodon interaction. These effects should constrain heavily the nucleotide motif composition of the most abundant mRNAs in the transcriptome, in particular, their SC usage, a usage that must be reflected by cellular tRNA concentrations and thus defines for all other genes which SCs are translated fastest and most accurately. Furthermore, the need to avoid such effects genome-wide appears serious enough to have favored the evolution of biases in context-dependent mutation that reduce the occurrence of intrinsically unfavorable motifs, and/or, when possible, to have induced the molecular machinery mediating such effects to rely opportunistically on already existing motif rarities and abundances. This may explain why nucleotide motif preferences are very similar in transcribed and nontranscribed mammalian DNA even though the preferences appear to be adaptive only in transcribed DNA.

Animals↗

Limitations of codon adaptation index and other coding DNA-based features for prediction of protein expression in Saccharomyces cerevisiae.

The relationship between codon usage and protein/mRNA expression in S. cerevisiae has been extensively studied. Recently, protein expression data for the whole yeast genome was published. We investigate which properties of coding DNA sequences can be used to predict expression levels. The new algorithm by Carbone et al. for computing dominating codon bias in a genome is evaluated. It is concluded that it works at least as well as existing methods, and eliminates the need to arbitrarily choose a set of highly expressed genes. Also, the hypothesis that information on codon pair frequencies can be used to predict expression is investigated. Our conclusion is that codon pairs do not contribute more information than do single codon frequencies. Overall correlation between predicted and actual expression data using properties of coding DNA sequences is around 0.65. Hence, while being a useful source of information, the expression levels predicted by these methods should only be used as a rule of thumb.

Algorithms↗

Absence of immunoglobulin E synthesis and airway eosinophilia by vaccination with plasmid DNA encoding ProDer p 1.

BACKGROUND: Various studies have shown that immunization with naked DNA encoding allergens induces T helper 1(Th1)-biased non-allergic responses. OBJECTIVE: To evaluate the polarization of the immune responses induced by vaccinations with plasmid DNA encoding the major mite allergen precursor ProDer p 1. METHODS: A DNA vaccine was constructed on the basis of a synthetic cDNA encoding ProDer p 1 with optimized codon usage. The immunogenicity of ProDer p 1 DNA in CBA/J mice was compared with that of purified natural Der p 1 or recombinant ProDer p 1 adjuvanted with alum. Vaccinated mice were subsequently exposed to aerosolized house dust mite extracts to provoke airway inflammation. The presence of inflammatory cells was examined in bronchoalveolar lavage (BAL) fluids and allergen-specific T cell reactivity was measured. RESULTS: Naive mice immunized with ProDer p 1 DNA developed Th1 immune responses characterized by high titres of specific IgG2a antibodies, low titres of specific IgG1 and, remarkably, the absence of anti-ProDer p 1 IgE. No specific responses were observed in animals vaccinated with the blank DNA vector. By contrast, natural Der p 1 or recombinant ProDer p 1 adsorbed to alum induced pronounced Th2 allergic responses with strong specific IgG1 and IgE titres. Spleen cells from DNA ProDer p 1-vaccinated mice secreted high levels of IFN-gamma and low production of IL-5. Conversely, both adjuvanted allergens stimulated typical Th2-type cytokine profile characterized by high and low levels of IL-5 and IFN-gamma, respectively. Whereas BAL eosinophilia was clearly observed in Der p 1-immunized animals, ProDer p 1 DNA as well as ProDer p 1 vaccinations prevented airway eosinophil infiltrations. CONCLUSIONS: These results suggest that vaccination with DNA encoding ProDer p 1 effectively fails to induce the allergen-induced IgE synthesis and airway cell infiltration. Plasmid DNA encoding ProDer p 1 may provide a novel approach for the treatment of house dust mite allergy.

Animals↗

Codon usage patterns among genes for lepidopteran hemolymph proteins.

Patterns in codon usage were examined for the coding regions of the 23 known lepidopteran hemolymph proteins. Coding triplets are GC rich at the third position and a significant linear relationship between GC content of silent and nonsilent (replacement) sites was demonstrated. Intron GC content was significantly lower than in coding regions and no relationship between intron GC content and the same at silent and nonsilent sites was found. Though hemolymph proteins are all produced by the same tissue--fat body--significantly less bias was observed when all moth sequences were pooled than when sequences of the two major species were analyzed separately, as predicted by the genome hypothesis. In cases where no statistically significant bias was observed, polar or acidic/basic amino acids were almost exclusively involved. Calculation of codon adaptation indices (CAI) was of limited value in quantifying the degree of codon bias and probably reflects the complexity of multicellular-organism life cycles and the changing patterns of gene expression over different developmental stages.

Animals↗