Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Analysis of the codon usage pattern in the Vibrio cholerae genome.

The codon usage in the Vibrio cholerae genome is analyzed in this paper. Although there are much more genes on the chromosome 1 than on chromosome 2, the codon usage patterns of genes on the two chromosomes are quite similar, indicating that the two chromosomes may have coexisted in the same cell for a very long history. Unlike the base frequency pattern observed in other genomes, the G+C content at the third codon position of the V. cholerae genome varies in a rather small interval. The most notable feature of codon usage of V. cholerae genome is that there is a fraction of genes show significant bias in base choice at the second codon position. The 2,006 known genes can be classified into two clusters according to the base frequencies at this position. The smaller cluster contains 227 genes, most of which code for proteins involved in transport and binding functions. The encoding products of these genes have significant bias in amino acids composition as compared with other genes. The codon usage patterns for the 1,836 function unknown ORFs are also analyzed, which is useful to study their functions.

Amino Acids↗

Heterogeneity in codon usage in the flatworm Schistosoma mansoni.

Synonymous codon choices vary considerably among Schistosoma mansoni genes. Principal components analysis detects a single major trend among genes, which highly correlates with GC content in third codon positions and exons, but does not discriminate among putatively highly and lowly expressed genes. The effective number of codons used in each gene, and its distribution when plotted against GC3, suggests that codon usage is shaped mainly by mutational biases. The GC content of exons, GC3, 5', 3', and flanking (5' + 3' + introns) regions are all correlated among them, suggesting that variations in GC content may exist among different regions of the S. mansoni genome. We propose that this genome structure might be among the most important factors shaping codon usage in this species, although the action of selection on certain sequences cannot be excluded.

Animals↗

Understanding the adaptation of Halobacterium species NRC-1 to its extreme environment through computational analysis of its genome sequence.

The genome of the halophilic archaeon Halobacterium sp. NRC-1 and predicted proteome have been analyzed by computational methods and reveal characteristics relevant to life in an extreme environment distinguished by hypersalinity and high solar radiation: (1) The proteome is highly acidic, with a median pI of 4.9 and mostly lacking basic proteins. This characteristic correlates with high surface negative charge, determined through homology modeling, as the major adaptive mechanism of halophilic proteins to function in nearly saturating salinity. (2) Codon usage displays the expected GC bias in the wobble position and is consistent with a highly acidic proteome. (3) Distinct genomic domains of NRC-1 with bacterial character are apparent by whole proteome BLAST analysis, including two gene clusters coding for a bacterial-type aerobic respiratory chain. This result indicates that the capacity of halophiles for aerobic respiration may have been acquired through lateral gene transfer. (4) Two regions of the large chromosome were found with relatively lower GC composition and overrepresentation of IS elements, similar to the minichromosomes. These IS-element-rich regions of the genome may serve to exchange DNA between the three replicons and promote genome evolution. (5) GC-skew analysis showed evidence for the existence of two replication origins in the large chromosome. This finding and the occurrence of multiple chromosomes indicate a dynamic genome organization with eukaryotic character.

Adaptation, Biological↗

Synonymous codon usage in Pseudomonas aeruginosa PA01.

Pseudomonas aeruginosa PA01 has a large (6.7 Mbp) genome with a high (67%) G+C content. Codon usage in this species is dominated by this compositional bias, with the average G+C content at synonymously variable third positions of codons being 83%. Nevertheless, there is some variation of synonymous codon usage among genes. The nature and causes of this variation were investigated using multivariate statistical analyses. Three trends were identified. The major source of variation was attributable to genes with unusually low G+C content that are probably due to horizontal transfer. A lesser trend among genes was associated with the preferential use of putatively translationally optimal codons in genes expressed at high levels. In addition, genes on the leading strand of replication were on average more G+T-rich. Our findings contradict the results of two previous analyses, and the reasons for the discrepancies are discussed.

Amino Acids↗

Evolution in bacteria: evidence for a universal substitution rate in cellular genomes.

This paper constructs a temporal scale for bacterial evolution by tying ecological events that took place at known times in the geological past to specific branch points in the genealogical tree relating the 16S ribosomal RNAs of eubacteria, mitochondria, and chloroplasts. One thus obtains a relationship between time and bacterial RNA divergence which can be used to estimate times of divergence between other branches in the bacterial tree. According to this approach, Salmonella typhimurium and Escherichia coli diverged between 120 and 160 million years (Myr) ago, a date which fits with evidence that the chief habitats occupied now by these two enteric species became available that long ago. The median extent of divergence between S. typhimurium and E. coli at synonymous sites for 21 kilobases of protein-coding DNA is 100%. This implies a silent substitution rate of 0.7-0.8%/Myr--a rate remarkably similar to that observed in the nuclear genes of mammals, invertebrates, and flowering plants. Similarities in the substitution rates of eucaryotes and procaryotes are not limited to silent substitutions in protein-coding regions. The average substitution rate for 16S rRNA in eubacteria is about 1%/50 Myr, similar to the average rate for 18S rRNA in vertebrates and flowering plants. Likewise, we estimate a mean rate of roughly 1%/25 Myr for 5S rRNA in both eubacteria and eucaryotes. For a few protein-coding genes of these enteric bacteria, the extent of silent substitution since the divergence of S. typhimurium and E. coli is much lower than 100%, owing to extreme bias in the usage of synonymous codons. Furthermore, in these bacteria, rates of amino acid replacement were about 20 times lower, on average, than the silent rate. By contrast, for the mammalian genes studied to date, the average replacement rate is only four to five times lower than the rate of silent substitution.

Bacteria↗

Evolution of the syntrophic interaction between Desulfovibrio vulgaris and Methanosarcina barkeri: Involvement of an ancient horizontal gene transfer.

The sulfate reducing bacteria Desulfovibrio vulgaris and the methanogenic archaea Methanosarcina barkeri can grow syntrophically on lactate. In this study, a set of three closely located genes, DVU2103, DVU2104, and DVU2108 of D. vulgaris, was found to be up-regulated 2- to 4-fold following the lifestyle shift from syntroph to sulfate reducer; moreover, none of the genes in this gene set were differentially regulated when comparing gene expression from various D. vulgaris pure culture experiments. Although exact function of this gene set is unknown, the results suggest that it may play roles related to the lifestyle change of D. vulgaris from syntroph to sulfate reducer. This hypothesis is further supported by phylogenomic analyses showing that homologies of this gene set were only narrowly present in several groups of bacteria, most of which are restricted to a syntrophic lifestyle, such as Pelobacter carbinolicus, Syntrophobacter fumaroxidans, Syntrophomonas wolfei, and Syntrophus aciditrophicus. Phylogenetic analysis showed that all three individual genes in the gene set tended to be clustered with their homologies from archaeal genera, and they were rooted on archaeal species in the phylogenetic trees, suggesting that they were horizontally transferred from archaeal methanogens. In addition, no significant bias in codon and amino acid usages was detected between these genes and the rest of the D. vulgaris genome, suggesting the gene transfer may have occurred early in the evolutionary history so that sufficient time has elapsed to allow an adaptation to the codon and amino acid usages of D. vulgaris. This report provides novel insights into the origin and evolution of bacterial genes linked to the lifestyle change of D. vulgaris from a syntrophic to a sulfate-reducing lifestyle.

Amino Acid Sequence↗

Molecular evolution of dinoflagellate luciferases, enzymes with three catalytic domains in a single polypeptide.

Enzymes with multiple catalytic sites are rare, and their evolutionary significance remains to be established. This study of luciferases from seven dinoflagellate species examines the previously undescribed evolution of such proteins. All these enzymes have the same unique structure: three homologous domains, each with catalytic activity, preceded by an N-terminal region of unknown function. Both pairwise comparison and phylogenetic inference indicate that the similarity of the corresponding individual domains between species is greater than that between the three different domains of each polypeptide. Trees constructed from each of the three individual domains are congruent with the tree of the full-length coding sequence. Luciferase and ribosomal DNA trees both indicate that the Lingulodinium polyedrum luciferase diverged early from the other six. In all species, the amino acid sequence in the central regions of the three domains is strongly conserved, suggesting it as the catalytic site. Synonymous substitution rates also are greatly reduced in the central regions of two species but not in the other five. This lineage-specific difference in synonymous substitution rates in the central region of the domains correlates inversely with the content of GC3, which can be accounted for by the biased usage toward C-ending codons at the degenerate sites. RNA modeling of the central region of the L. polyedrum luciferase domain suggests a function of the constrained synonymous substitutions in the circadian-controlled protein synthesis.

Amino Acid Sequence↗

The complete sequence of the zebrafish (Danio rerio) mitochondrial genome and evolutionary patterns in vertebrate mitochondrial DNA.

We describe the complete sequence of the 16,596-nucleotide mitochondrial genome of the zebrafish (Danio rerio); contained are 13 protein genes, 22 tRNAs, 2 rRNAs, and a noncoding control region. Codon usage in protein genes is generally biased toward the available tRNA species but also reflects strand-specific nucleotide frequencies. For 19 of the 20 amino acids, the most frequently used codon ends in either A or C, with A preferred over C for fourfold degenerate codons (the lone exception was AUG: methionine). We show that rates of sequence evolution vary nearly as much within vertebrate classes as between them, yet nucleotide and amino acid composition show directional evolutionary trends, including marked differences between mammals and all other taxa. Birds showed similar compositional characteristics to the other nonmammalian taxa, indicating that the evolutionary trend in mammals is not solely due to metabolic rate and thermoregulatory factors. Complete mitochondrial genomes provide a large character base for phylogenetic analysis and may provide for robust estimates of phylogeny. Phylogenetic analysis of zebrafish and 35 other taxa based on all protein-coding genes produced trees largely, but not completely, consistent with conventional views of vertebrate evolution. It appears that even with such a large number of nucleotide characters (11,592), limited taxon sampling can lead to problems associated with extensive evolution on long phyletic branches.

Animals↗

Comparison of the amino acid sequences of the transacylase components of branched chain oxoacid dehydrogenase of Pseudomonas putida, and the pyruvate and 2-oxoglutarate dehydrogenases of Escherichia coli.

The nucleotide sequence of bkdB, the structural gene for E2b, the transacylase component of branched-chain-oxoacid dehydrogenase of Pseudomonas putida has been determined and translated into its amino acid sequence. The start of bkdB was identified from the N-terminal sequence of E2b isolated from branched-chain-oxoacid dehydrogenase of the closely related species, P. aeruginosa. The reading frame was composed of 65.5% G + C with 82.3% of the codons ending in G or C. There was no intergenic space between bkdA2 and bkdB. No codons requiring minor tRNAs were utilized and the codon bias index indicated a preferential codon usage. The bkdB gene encoded 423 amino acids although the N-terminal methionine was absent from E2b prepared from P. aeruginosa. The relative molecular mass of the encoded protein was 45,134 (45,003 minus methionine) vs 47,000 obtained by SDS/polyacrylamide gel electrophoresis. There was a single lipoyl domain in E2b compared to three lipoyl domains in E2p, and one domain in E2o, the transacylases of pyruvate and 2-oxoglutarate dehydrogenases of Escherichia coli respectively. There was significant similarity between the lipoyl domain of E2b and of E2p and E2o as well as between the E1-E2 binding domains of E2b, E2p and E2o. There was no similarity between the E3 binding domain of E2b to E2p and E2o which may reflect the uniqueness of the E3 component of branched-chain-oxoacid dehydrogenase of P. putida. The conclusions drawn from these comparisons are that the transacylases of prokaryotic pyruvate, 2-oxoglutarate and branched-chain-oxoacid dehydrogenases descended from a common ancestral protein probably at about the same time.

3-Methyl-2-Oxobutanoate Dehydrogenase (Lipoamide)↗

RPA190, the gene coding for the largest subunit of yeast RNA polymerase A.

Yeast RNA polymerases are being extensively studied at the gene level. The entire gene encoding the largest subunit of RNA polymerase A, A190, was isolated and characterized in detail. Southern hybridization and gene disruption experiments showed that the RPA190 gene is unique in the haploid yeast genome and essential for cell viability. Nuclease S1 mapping was used to identify mRNA 5' and 3' termini. RPA190 encodes a polypeptide chain of 186,270 daltons in a large uninterrupted reading frame. A dot matrix comparison of the deduced amino acid sequence of subunit A190 with Escherichia coli beta' and cognate subunits B220 and C160 from yeast RNA polymerases B and C showed a conserved pattern of homology regions (I-VI). A potential DNA-binding site (zinc-binding motif) is conserved in the N-terminal region I. Remarkably, the A190 subunit does not harbor the heptapeptide repeated sequence present in the B220 subunit. The sequence of the A190 subunit diverges from B220 and C160 by the presence of two hydrophilic domains inserted between homology regions I and II, and V and VI. From their codon usage and third base pyrimidine bias, RNA polymerase genes RPA190, RPB220, RPC160, and RPC40 fall among yeast genes expressed at an average level. The RPA190 5'-flanking region contains features present in other polymerase genes that might function in regulation.

Amino Acid Sequence↗

Amplification and characterization of cysteine proteinase genes from nematodes.

In order to isolate proteinase genes from parasitic nematodes by polymerase chain reaction (PCR) techniques, we employed a pair of consensus oligonucleotide primers designed to anneal to the active site cysteine (primer ncpC) and asparagine (primer ncpN) coding regions of cysteine proteinases. The primers were biased toward the nucleotide and codon usages of cysteine proteinase genes of nematodes and were based on the consensus nucleotide sequences flanking the active site residues of genes from Haemonchus contortus, Caenorhabditis elegans, and Ostertagia ostertagi. We employed 'touchdown' PCR conditions and were able to amplify novel cysteine proteinase gene fragments from the rodent parasite Strongyloides ratti, the human pathogen S. stercoralis, the canine hookworm Ancylostoma caninum, and from C. elegans. These clones are gene homologs of cathepsin B-like (lysosomal associated) proteases and will facilitate screening of both cDNA and genomic DNA libraries.

Amino Acid Sequence↗

Cloning and analysis of the DNA polymerase-encoding gene from Thermus filiformis.

The gene encoding Thermus filiformis (Tfi) DNA polymerase was cloned and its nucleotide sequence was determined. The primary structure of Tfi DNA polymerase was deduced from its nucleotide sequence. Tfi DNA polymerase is comprised of 833 amino acid residues and its molecular mass was determined to be 93,890 Da. The deduced amino acid sequence of Tfi DNA polymerase showed a high sequence homology to E. coli DNA polymerase I-like DNA polymerases: 78.5% homology to Taq DNA polymerase, 78.4% to Tca DNA polymerase, and 41.8% to E. coli DNA polymerase I. An extremely high sequence identity was observed in the region containing polymerase activity. The G + C content of the coding region for the Tfi DNA polymerase gene was 68.5%, which was higher than that of the chromosomal DNA (65%). The G + C contents in the first, second, and third positions of the codons used were 71.8%, 40.9%, and 92.7% respectively. Codon usage in Tfi DNA polymerase was heavily biased towards the use of G + C in the third position. Rare codons with U or A as the third base were sometimes used to avoid using GA(A/T) TC and TCGA sequences, as they are recognition sites for the restriction endonucleases TfiI and TaqI.

Amino Acids↗

Hill-Robertson interference in Drosophila melanogaster: reply to Marais, Mouchiroud and Duret.

The usage of preferred codons in Drosophila melanogaster is reduced in regions of lower recombination. This is consistent with population genetics theory, whereby the effectiveness of selection on multiple targets is limited by stochastic effects caused by linkage. However, because the selectively preferred codons in D. melanogaster end in C or G, it has been argued that base-composition-biasing effects of recombination can account for the observed relationship between preferred codon usage and recombination rate (Marais et al., 2003). Here, we show that the correlation between base composition (of protein-coding and intron regions) and recombination rate holds only for lower values of the latter. This is consistent with a Hill-Robertson interference model and does not support a model whereby the entire effect of recombination on codon usage can be attributed to its potential role in generating compositional bias.

Animals↗

Nucleotide sequence of gene pfkB encoding the minor phosphofructokinase of Escherichia coli K-12.

The nucleotide sequence of a 1.3-kb DNA fragment containing the entire pfkB gene which codes for Pfk-2 of Escherichia coli, a minor phosphofructokinase (Pfk) enzyme, is reported. The Pfk-2 protein subunit is encoded by 924 bp, has 308 amino acids and an Mr of 33 000. Like other weakly expressed E. coli genes the codon usage in the pfkB gene is random; there is no strong bias for the usage of major tRNA isoaccepting species, and the codon preference rules of Grosjean and Fiers [Gene, 18 (1982) 199-209] are followed. This is the first report of the complete gene sequence of a phosphofructokinase.

Amino Acid Sequence↗

Nucleotide sequence of an actin-encoding gene from Hydra attenuata: structural characteristics and evolutionary implications.

We have determined the complete nucleotide sequence of an actin-encoding gene from Hydra attenuata as well as partial sequences of cDNA clones from two additional actin-encoding genes. The gene from the genomic clone contains a single intron, and has promoter and polyadenylation signals similar to those found in other species. The hydra genome has a very A + T-rich base composition (71%). This is reflected in the codon usage of the actin-encoding genes, which is strongly biased towards codons having A or T in the third position. The hydra actin-encoding gene family consists of three or more transcribed genes, two of which are very closely related to each other and probably arose by a recent gene duplication. Hydra actin, like other invertebrate actins, is more similar to the non-muscle isotypes of vertebrates than to the vertebrate muscle actins. Hydra actin is more similar to animal actins than to those of plants or fungi, which is consistent with the view that all metazoans arose from a single protist ancestor.

Actins↗

Unusual usage of AGG and TTG codons in humans and their viruses.

Prior analysis on human protein-coding DNA sequences has identified local base composition as the primary predictor of synonymous codon usage. However, in many organisms, codon usage is influenced by natural selection, particularly for efficient expression of functional gene products. Because viruses are expected to evolve codon usage in the context of their host's molecular machinery, their genomes provide another window into the forces that guide their host's molecular evolution. Factor analysis was performed on codon usage of 16,654 genes annotated in Build 34 of the human genome, and the primary factor was correlated strongly with local base composition. However, two codons, AGG and TTG, rose in frequency as all other C- and G-ending codons decreased in frequency. These two codons were the only C- or G-ending codons with usages that negatively correlated with gene expression. Variation among viruses in codon usage also strongly reflects variation in base composition and, again, AGG and TTG decrease in frequency as all other C- and G-ending codons increase in frequency. It appears that usages of these two codons can not be explained by local compositional biases, implying a more direct role of natural selection on codon usage in humans.

Amino Acids↗

Influence of parasitic life style on the patterns of codon usage and base frequencies of Ancylostoma and Necator species.

Parametric analyses were used to investigate the nucleotide, codon, and amino acid composition of coding sequences corresponding to hook-worms. Ancylostoma caninum and Necator americanus. Although genomic research has become prevalent within the scientific community, few studies have dealt directly with parasitic species. Parasites have existed throughout the history of mankind due to their wide range of distribution in nature and their ability to evade immune detection. An AT nucleotide bias was identified in both A. caninum and N. americanus sequences. A similar AT bias was also identified in both datasets when considering relative synonymous codon usage. However, the codon bias was much more pronounced in N. americanus as compared to A. caninum. Bias was also present at the amino acid level, and appeared to be partially independent of the nucleotide-based biases. Analysis of parasite genomes will facilitate the development of vaccines against larval forms of parasites. Moreover, the examination of the parasite genes in general, will allow for a more in-depth understanding of the evolution of the parasites and parasitism.

Ancylostoma↗

Analyses of the gene and amino acid sequence of the Prevotella (Bacteroides) ruminicola 23 xylanase reveals unexpected homology with endoglucanases from other genera of bacteria.

The DNA sequence for the xylanase gene from Prevotella (Bacteroides) ruminicola 23 was determined. The xylanase gene encoded for a protein with a molecular weight of 65,740. An apparent leader sequence of 22 amino acids was observed. The promoter region for expression of the xylanase gene in Bacteroides species was identified with a promoterless chloramphenicol acetyltransferase gene. A region of high amino acid homology was found with the proposed catalytic domain of endoglucanases from several organisms, including Butyrivibrio fibrisolvens, Ruminococcus flavefaciens, and Clostridium thermocellum. The cloned xylanase was found to exhibit endoglucanase activity against carboxymethyl cellulose. Analysis of the codon usage for the xylanase gene found a bias towards G and C in the third position in 16 of 18 amino acids with degenerate codons.

Amino Acid Sequence↗