Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Protein coding sequence identification by simultaneously characterizing the periodic and random features of DNA sequences.

Most codon indices used today are based on highly biased nonrandom usage of codons in coding regions. The background of a coding or noncoding DNA sequence, however, is fairly random, and can be characterized as a random fractal. When a gene-finding algorithm incorporates multiple sources of information about coding regions, it becomes more successful. It is thus highly desirable to develop new and efficient codon indices by simultaneously characterizing the fractal and periodic features of a DNA sequence. In this paper, we describe a novel way of achieving this goal. The efficiency of the new codon index is evaluated by studying all of the 16 yeast chromosomes. In particular, we show that the method automatically and correctly identifies which of the three reading frames is the one that contains a gene.

Journal Article↗

DNA sequences of yeast H3 and H4 histone genes from two non-allelic gene sets encode identical H3 and H4 proteins.

The complete DNA sequences of two loci encoding H3 and H4 histones in Saccharomyces cerevisiae have been determined. Each locus contains one H3 and one H4 gene. The genes at each locus are divergently transcribed and the coding sequences are separated by 646 base-pairs at one locus and 676 base-pairs at the other. The H3 genes code for identical histone H3 proteins and the H4 genes code for identical histone H4 proteins. The yeast proteins differ from histones H3 and H4 of calf by 15 and 8 amino acid substitutions, respectively, and these differences are largely confined to the carboxy-terminal halves of the proteins. The genes demonstrate a bias in synonymous codon usage similar to that noted for other yeast genes. This bias is confined to the coding sequences of the genes and is specific for the reading frame encoding the proteins. The coding sequence of each gene is flanked on both sides by DNA with an A + T content of 70 to 80%. Possible regulatory sequences are located relative to the 5' and 3'-termini of the histone H3 and H4 RNA transcripts.

Base Sequence↗

A problem in multivariate analysis of codon usage data and a possible solution.

Multivariate analyses are often used to identify major trends of variation in synonymous codon usage among genes. These analyses need to be performed on properly normalized codon usage data to avoid biases masking this synonymous variation, i.e., gene length, amino acid usage, and codon degeneracy; however, previous studies have failed to do so. In this paper, we demonstrate that the use of alternative normalized data (called 'relative adaptiveness' in the literature) can avoid all these biases and furthermore, can identify more trends of variation among genes, including GC-ending codon usage, GT-ending codon usage, and gene expression level.

Amino Acids↗

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids↗

Codon usage in Entamoeba histolytica.

The codon usage of 10 E. histolytica genes comprising 4455 codons was analysed. The codon usage revealed an extremely biased use of synonymous codons with a preference for NNU (44%) and NNA (41.4%) codons. Codons CGG (arg), AGG (arg) and CCG (pro) were absent in the E. histolytica genes examined. The codon usage of E. histolytica resembled that of Plasmodium falciparum.

Animals↗

Novel third-letter bias in Escherichia coli codons revealed by rigorous treatment of coding constraints.

A novel bias in codon third-letter usage was found in Escherichia coli genes with low fractions of "optimal codons", by comparing intact sequences with control random sequences. Third-letter usage has been found to be biased according to preference in codon usage and to doublet preference from the following first letter. The present study examines third-letter usage in the context of the nucleotide sequence when these preferences are considered. In order to exclude any influence by these factors, the random sequences were generated such that the amino acid sequence, codon usage, and the doublet frequency in each gene were all preserved. Comparison of intact sequences with these randomly generated sequences reveals that third letters of codons show a strong preference for the purine/pyrimidine pattern of the next codons: purine (R) is preferred to pyrimidine (Y) at the third site when followed by an R-Y-R codon, and pyrimidine is preferred when followed by an R-R-Y, an R-Y-Y or a Y-R-Y codon. This bias is probably related to interactions of tRNA molecules in the ribosome.

Amino Acids↗

Evolutionary forces in shaping the codon and amino acid usages in Blochmannia floridanus.

Endosymbiotic relationship has great effect on ecological system. Codon and amino acid usages bias of endosymbiotic bacteria Blochmannia floridanus (whose host is an ant Camponotus floridanus) was investigated using experimentally known genes of this organism. Correspondence Analysis on RSCU values show that there exists only one single explanatory major axis that is linked to the strand specific mutational biases. Majority of the genes have a tendency to concentrate on the leading strand, which may be related to the adaptive property related to the replication mechanisms. Amino acid usages were markedly different between the highly and lowly expressed genes in this organism and in particular, GC rich amino acids were found to occur significantly higher in highly expressed genes than the lowly expressed genes. Comparative analyses of the orthologous genes of Escherichia coli and Blochmannia floridanus show that highly expressed genes are significantly more conserved than lowly expressed genes. Based on our results we concluded that strand specific mutational bias is strongly operational in selecting the codon usage in this organism. Replicational-transcriptional selection can be invoked from the presence of majority of highly expressed genes in the leading strand. Conservation of GC rich amino acids in the highly expressed genes to its ancestor is the major source of variation in amino acid usages in the organism. Hydrophobicity of the genes is the second major source in differentiating the genes according to their amino acid usages in this organism.

Amino Acid Sequence↗

Codon usage in the A/T-rich bacterium Campylobacter jejuni.

Campylobacter jejuni is a Gram negative, microaerophilic pathogen that causes gastroenteritis in humans. The genome of C. jejuni is AT-rich, with a mol% G + C of 30.4. This high AT content was hypothesized to result in unique codon usage. In the present study, we analyzed the codon usage of sixty-seven C. jejuni genes and generated a codon frequency table. As predicted, the codon usage of C. jejuni revealed a strong bias towards codons ending in A or U. In addition to determining codon usage frequencies, the relative synonymous codon usage values were calculated to identify rare and optimal codons. Seventeen codons were identified as optimal and twelve codons as rare. Thirty-two codons exhibited little or no bias. A plot of the effective number of codons versus the third position %G + C values for the sixty-seven genes revealed that C. jejuni uses an average of 39 of the 61 codons to encode proteins. These data will be useful for various molecular analyses including selection of degenerate primers to screen C. jejuni-genomic DNA libraries.

Adenine↗

The 'effective number of codons' revisited.

Frank Wright [Gene 87 (1990) 23] derived a formula for calculation of a quantity termed the 'effective number of codons' (Nc) based on codon homozygosities. This quantity is a number between 20 and 61 and tells to what degree the codon usage in a gene is biased, i.e., it approaches 20 codons for the extremely biased genes, and approaches 61 for the genes where all possible codons are used with no preference. Among the different measures of codon bias Nc is considered the most useful and has found widespread use in papers dealing with codon usage phenomena. In this paper, the mathematical behaviours of codon homozygosities and Nc are evaluated, using Escherichia coli as the model organism. The results indicate that the classical formula for calculation of Nc could appropriately be substituted under circumstances, where there is bias discrepancy, i.e., when one amino acid (or more) within a degeneracy group is associated with strong codon bias while at the same time others in the same degeneracy group have little bias. An alternative estimator, termed Nc, is proposed and tested against Nc, and performs better when there is such bias discrepancy.

Codon↗

Cloning of two glutamate dehydrogenase cDNAs from Asparagus officinalis: sequence analysis and evolutionary implications.

Two different amplification products, termed c1 and c2, showing a high similarity to glutamate dehydrogenase sequences from plants, were obtained from Asparagus officinalis using two degenerated primers and RT-PCR (reverse transcriptase polymerase chain reaction). The genes corresponding to these cDNA clones were designated aspGDHA and aspGDHB. Screening of a cDNA library resulted in the isolation of cDNA clones for aspGDHB only. Analysis of the deduced amino acid (aa) sequence from the full-length cDNA suggests that the gene product contains all regions associated with metabolic function of NAD glutamate dehydrogenase (NAD-GDH). A first phylogenetic analysis including only GDHs from plants suggested that the two GDH genes of A. officinalis arose by an ancient duplication event, pre-dating the divergence of monocots and dicots. Codon usage analysis showed a bias towards A/T ending codons. This tendency is likely due to the biased nucleotide composition of the asparagus genome, rather than to the translational selection for specific codons. Using principal coordinate analysis, the evolutionary relatedness of plant GDHs with homologous sequences from a large spectrum of organisms was investigated. The results showed a closer affinity of plant GDHs to GDHs of thermophilic archaebacterial and eubacterial species, when compared to those of unicellular eukaryotic fungi. Sequence analysis at specific amino acid signatures, known to affect the thermal stability of GDH, and assays of enzyme activity at non-physiological temperatures, showed a greater adaptation to heat-stress conditions for the asparagus and tobacco enzymes compared with the Saccharomyces cerevisiae enzyme.

Amino Acid Sequence↗

Codon usage patterns in chromosomal and retrotransposon genes of the mosquito Anopheles gambiae.

Codon usage was compiled for fourteen chromosomal genes and four retrotransposons from the mosquito Anopheles gambiae. Variation exists among chromosomal genes in the degree of bias. The genes showing the highest bias are probably most highly expressed. In these genes, the base composition at the third codon position is much richer in G + C than is the overall coding sequence. Thus, codon usage is biased toward G- or C-ending codons. Codon usage in each retrotransposon is quite different, not only from chromosomal genes but also from the other retrotransposons. Codon usage comparisons among homologous genes from An. gambiae and two other Dipterans, the yellow fever mosquito Aedes aegypti and the fruitfly Drosophila melanogaster, show that while there are similarities, particularly between An. gambiae and D. melanogaster in the preference for G- and C-ending codons, each species has evolved a distinct pattern of codon usage.

Animals↗

De Novo Assembly and Comparative Analysis of the Complete Mitochondrial Genome of Mesenchytraeus (Annelida, Enchytraeidae).

The Changbai Mountain range is one of the key glacial refugia in Northeast Asia. Mesenchytraeus exhibits high species diversity, strong endemism, and widespread cryptic species in this region, for which mitogenomes provide useful molecular markers for exploring cryptic species complexes. This makes Mesenchytraeus an ideal model for studying mitogenome evolution among closely related lineages; however, no mitogenome data have been reported for this genus to date. In this study, we performed de novo assembly, annotation, and comparative analysis of the mitogenomes of 13 Mesenchytraeus species (14 individuals) from Changbai Mountain. All mitogenomes are typical circular molecules containing 37 genes, but putative control regions are rearranged and consistently located between ATP6 and trnR. All species exhibit annelid-specific strand nucleotide biases, characterized by negative GC skew and near-zero AT skew. Codon usage analysis reveals that codon families with wobble U are significantly biased toward mtDNA codons, whereas those with wobble C or G are biased toward non-mtDNA codons, suggesting a conserved mitochondrial codon usage pattern in annelids. All tRNAs form typical cloverleaf secondary structures except trnS2, which lacks the D-stem and the dihydrouridine (DHU) arm in some species. The putative control regions commonly contain complex palindromic repeats, hairpins, and repetitive elements, and may harbor dual replication origins. Phylogenetic analyses support the monophyly of Mesenchytraeus and reveal significant molecular divergence among morphologically cryptic species. This study provides the first mitogenome dataset for Mesenchytraeus and offers new insights into the evolution and replication mechanisms of mitogenomes in Clitellata and broader Annelida.

Mesenchytraeus↗

Evolution of codon usage patterns: the extent and nature of divergence between Candida albicans and Saccharomyces cerevisiae.

Codon usage in a sample of 28 genes from the pathogenic yeast Candida albicans has been analysed using multivariate statistical analysis. A major trend among genes, correlated with gene expression level, was identified. We have focussed on the extent and nature of divergence between C.albicans and the closely related yeast Saccharomyces cerevisiae. It was recently suggested that significant differences exist between the subsets of preferred codons in these two species [Brown et al. (1991) Nucleic Acids Res. 19, 4293]. Overall, the genes of C.albicans are more A + T-rich, reflecting the lower genomic G + C content of that species, and presumably resulting from a different pattern of mutational bias. However, in both species highly expressed genes preferentially use the same subset of 'optimal' codons. A suggestion that the low frequency of NCG codons in both yeast species results from selection against the presence of codons that are potentially highly mutable is discounted. Codon usage in C.albicans, as in other unicellular species, can be interpreted as the result of a balance between the processes of mutational bias and translational selection. Codon usage in two related Candida species, C.maltosa and C.tropicalis, is briefly discussed.

Biological Evolution↗

Mitochondrial genomes of Galathealinum, Helobdella, and Platynereis: sequence and gene arrangement comparisons indicate that Pogonophora is not a phylum and Annelida and Arthropoda are not sister taxa.

We report a contiguous region of more than half (> 7,500 nt) of the mitochondrial genomes for Platynereis dumerii (Annelida: Polychaeta), Helobdella robusta (Annelida: Hirudinida), and Galathealinum brachiosum (Pogonophora: Perviata). The relative arrangements of all 22 genes identified for Helobdella and Galathealinum are identical to one another and to their arrangements in the mtDNA of the previously studied oligochaete annelid Lumbricus. In contrast, Platynereis differs from these taxa in the positions of several tRNA genes and in having two additional tRNA genes (trnC and trnM) and a large noncoding sequence in this region. Comparisons of relative gene arrangements and of the nucleotide and inferred amino acid sequences among these and other published taxa provide strong support for an annelid-mollusk clade that excludes arthropods, and for the inclusion of pogonophorans within Annelida, rather than giving them separate phylum status. Gene arrangement comparisons include the first use of a recently described method on previously unpublished data. Although a variety of alternative initiation codons are typically used by mitochondrial protein-encoding genes, ATG appears to be the initiator for all but one reported here. The large noncoding region (1,091 nt) identified in Platynereis has no significant sequence similarity to the noncoding region of Lumbricus, although each contains runs of TA dinucleotides and of homopolymers, which could potentially serve as signaling elements. There is strong bias for synonymous codon usage in Helobdella and especially in Galathealinum. In this latter taxon, 5 codons are completely unused, 13 are used three or fewer times, and G appears at third codon positions in only 26 of the 2,236 codons. Nucleotide composition bias appears to influence amino acid composition of the proteins.

Amino Acid Sequence↗

Reduced synonymous substitution rate at the start of enterobacterial genes.

Synonymous codon usage is less biased at the start of Escherichia coli genes than elsewhere. The rate of synonymous substitution between E.coli and Salmonella typhimurium is substantially reduced near the start of the gene, which suggests the presence of an additional selection pressure which competes with the selection for codons which are most rapidly translated. Possible competing sources of selection are the presence of secondary ribosome binding sites downstream from the start codon, the avoidance of mRNA secondary structure near the start of the gene and the use of sub-optimal codons to regulate gene expression. We provide evidence against the last of these possibilities. We also show that there is a decrease in the frequency of A, and an increase in the frequency of G along the E.coli genes at all three codon positions. We argue that these results are most consistent with selection to avoid mRNA secondary structure.

Base Composition↗

The complete mitochondrial genome of the stomatopod crustacean Squilla mantis.

BACKGROUND: Animal mitochondrial genomes are physically separate from the much larger nuclear genomes and have proven useful both for phylogenetic studies and for understanding genome evolution. Within the phylum Arthropoda the subphylum Crustacea includes over 50,000 named species with immense variation in body plans and habitats, yet only 23 complete mitochondrial genomes are available from this subphylum. RESULTS: I describe here the complete mitochondrial genome of the crustacean Squilla mantis (Crustacea: Malacostraca: Stomatopoda). This 15994-nucleotide genome, the first described from a hoplocarid, contains the standard complement of 13 protein-coding genes, 22 transfer RNA genes, two ribosomal RNA genes, and a non-coding AT-rich region that is found in most other metazoans. The gene order is identical to that considered ancestral for hexapods and crustaceans. The 70% AT base composition is within the range described for other arthropods. A single unusual feature of the genome is a 230 nucleotide non-coding region between a serine transfer RNA and the nad1 gene, which has no apparent function. I also compare gene order, nucleotide composition, and codon usage of the S. mantis genome and eight other malacostracan crustaceans. A translocation of the histidine transfer RNA gene is shared by three taxa in the order Decapoda, infraorder Brachyura; Callinectes sapidus, Portunus trituberculatus and Pseudocarcinus gigas. This translocation may be diagnostic for the Brachyura. For all nine taxa nucleotide composition is biased towards AT-richness, as expected for arthropods, and is within the range reported for other arthropods. Codon usage is biased, and much of this bias is probably due to the skew in nucleotide composition towards AT-richness. CONCLUSION: The mitochondrial genome of Squilla mantis contains one unusual feature, a 230 base pair non-coding region has so far not been described in any other malacostracan. Comparisons with other Malacostraca show that all nine genomes, like most other mitochondrial genomes, share a bias toward AT-richness and a related bias in codon usage. The nine malacostracans included in this analysis are not representative of the diversity of the class Malacostraca, and additional malacostracan sequences would surely reveal other unusual genomic features that could be useful in understanding mitochondrial evolution in this taxon.

Animals↗

An experimental validation of orphan genes of Buchnera, a symbiont of aphids.

Although Buchnera sp. APS, an intracellular symbiont of pea aphids, is a close relative of Escherichia coli, its genome has been extensively modified because of its prolonged intracellular life. In our previous studies on the Buchnera genome, computer analysis predicted three "orphan" genes, yba2, yba3, and yba4, which are open reading frames (ORFs) with no homologs in the database. In this paper, we successfully validated all these orphan genes by RT-PCR and Northern hybridization. The present study also revealed that yba3 and yba4 formed an operon, suggesting that they function in concert. Sequences around transcriptional start sites suggests that these genes are under the control of sigma 70. In view of codon usage and AT bias observed in these genes, it is likely that Buchnera have maintained them for an evolutionarily long time.

Amino Acid Sequence↗

Selection conflicts, gene expression, and codon usage trends in yeast.

Synonymous codon usage in yeast appears to be influenced by natural selection on gene expression, as well as regional variation in compositional bias. Because of the large number of potential targets of selection (i.e., most of the codons in the genome) and presumed small selection coefficients, codon usage is an excellent model for studying factors that limit the effectiveness of selection. We use factor analysis to identify major trends in codon usage for 5836 genes in Saccharomyces cerevisiae. The primary factor is strongly correlated with gene expression, consistent with the model that a subset of codons allows for more efficient translation. The secondary factor is very strongly correlated with third codon position GC content and probably reflects regional variation in compositional bias. We find that preferred codon usage decreases in the face of three potential limitations on the effectiveness of selection: reduced recombination rate, increased gene length, and reduced intergenic spacing. All three patterns are consistent with the Hill-Robertson effect (reduced effectiveness of selection among linked targets). A reduction in gene expression in closely spaced genes may also reflect selection conflicts due to antagonistic pleiotropy.

Codon↗