Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Codon usage in plant genes.

We have examined codon bias in 207 plant gene sequences collected from Genbank and the literature. When this sample was further divided into 53 monocot and 154 dicot genes, the pattern of relative use of synonymous codons was shown to differ between these taxonomic groups, primarily in the use of G + C in the degenerate third base. Maize and soybean codon bias were examined separately and followed the monocot and dicot codon usage patterns respectively. Codon preference in ribulose 1,5 bisphosphate and chlorophyll a/b binding protein, two of the most abundant proteins in leaves was investigated. These highly expressed are more restricted in their codon usage than plant genes in general.

Amino Acid Sequence↗

Cloning of the Zymomonas mobilis structural gene encoding alcohol dehydrogenase I (adhA): sequence comparison and expression in Escherichia coli.

Zymomonas mobilis ferments sugars to produce ethanol with two biochemically distinct isoenzymes of alcohol dehydrogenase. The adhA gene encoding alcohol dehydrogenase I has now been sequenced and compared with the adhB gene, which encodes the second isoenzyme. The deduced amino acid sequences for these gene products exhibited no apparent homology. Alcohol dehydrogenase I contained 337 amino acids, with a subunit molecular weight of 36,096. Based on comparisons of primary amino acid sequences, this enzyme belongs to the family of zinc alcohol dehydrogenases which have been described primarily in eucaryotes. Nearly all of the 22 strictly conserved amino acids in this group were also conserved in Z. mobilis alcohol dehydrogenase I. Alcohol dehydrogenase I is an abundant protein, although adhA lacked many of the features previously reported in four other highly expressed genes from Z. mobilis. Codon usage in adhA is not highly biased and includes many codons which were unused by pdc, adhB, gap, and pgk. The ribosomal binding region of adhA lacked the canonical Shine-Dalgarno sequence found in the other highly expressed genes from Z. mobilis. Although these features may facilitate the expression of high enzyme levels, they do not appear to be essential for the expression of Z. mobilis adhA.

Alcohol Dehydrogenase↗

Extraordinarily high evolutionary rate of pseudogenes: evidence for the presence of selective pressure against changes between synonymous codons.

Comparisons of nucleotide sequences of several pseudogenes described to date, including alpha- and beta-globin and immunoglobulin kappa-type variable domain pseudogenes, with those of functional counterparts revealed that pseudogenes accumulate mutations at an extremely high rate uniformly over their entirety. It is remarkable that the evolutionary rate exceeds the rate of changes between synonymous codons, the highest known rate, in functional genes. Because no pseudogenes appear to function, this result strongly supports the neutral theory. In addition this result apparently indicates the presence of selective pressure against changes between synonymous codons in functional genes. Close examinations of codon utilization patterns in pseudogenes and functional genes revealed a significant correlation between the rate of changes at synonymous codon sites and the strength of bias in code word usage. This implies that even synonymous codon changes are not completely free from selective pressure but are constrained in part, although presumably weakly, depending on the degree of bias in code word usage. We also reexamined alignment between mouse beta h3 (pseudogene) and beta maj sequences and found a unique structure of the beta h3 that is homologous in sequence to the beta maj gene overall but contains a long deletion (about 150 base pairs) in the middle of the gene.

Animals↗

Random sequence analysis of genomic DNA of a hyperthermophile: Aquifex pyrophilus.

Aquifex pyrophilus is one of the hyperthermophilic bacteria that can grow at temperatures up to 95 degrees C. To obtain information about its genomic structure, random sequencing was performed on plasmid libraries containing 0.5-2 kb genomic DNA fragments of A. pyrophilus. Comparison of the obtained sequence tags with known proteins revealed that 123 tags showed strong similarity to previously identified proteins in the PIR or Genebank databases. These included three proteases, two amino acid racemases, and three enzymes utilizing oxygen as substrate. Although the GC ratio of the genome is about 40%, the codon usage of A. pyrophilus showed biased occurrence of G and C at the third position of codons, especially those for amino acids such as asparagine, aspartic acid, cysteine, glutamine, glutamic acid, histidine, lysine, and tyrosine. A higher ratio of positively charged amino acids in A. pyrophilus proteins as compared with proteins from mesophiles suggested that Aquifex proteins might contain increased ion-pair interaction that could help to maintain heat stability.

Amino Acid Sequence↗

Evolutionary constraints on codon and amino acid usage in two strains of human pathogenic actinobacteria Tropheryma whipplei.

The factors governing codon and amino acid usages in the predicted protein-coding sequences of Tropheryma whipplei TW08/27 and Twist genomes have been analyzed. Multivariate analysis identifies the replicational-transcriptional selection coupled with DNA strand-specific asymmetric mutational bias as a major driving force behind the significant interstrand variations in synonymous codon usage patterns in T. whipplei genes, while a residual intrastrand synonymous codon bias is imparted by a selection force operating at the level of translation. The strand-specific mutational pressure has little influence on the amino acid usage, for which the mean hydropathy level and aromaticity are the major sources of variation, both having nearly equal impact. In spite of the intracellular lifestyle, the amino acid usage in highly expressed gene products of T. whipplei follows the cost-minimization hypothesis. The products of the highly expressed genes of these relatively A + T-rich actinobacteria prefer to use the residues encoded by GC-rich codons, probably due to greater conservation of a GC-rich ancestral state in the highly expressed genes, as suggested by the lower values of the rate of nonsynonymous divergences between orthologous sequences of highly expressed genes from the two strains of T. whipplei. Both the genomes under study are characterized by the presence of two distinct groups of membrane-associated genes, products of which exhibit significant differences in primary and potential secondary structures as well as in the propensity of protein disorder.

Actinobacteria↗

Identification of a novel operon in Lactococcus lactis encoding three enzymes for lactic acid synthesis: phosphofructokinase, pyruvate kinase, and lactate dehydrogenase.

The discovery of a novel multicistronic operon that encodes phosphofructokinase, pyruvate kinase, and lactate dehydrogenase in the lactic acid bacterium Lactococcus lactis is reported. The three genes in the operon, designated pfk, pyk, and ldh, contain 340, 502, and 325 codons, respectively. The intergenic distances are 87 bp between pfk and pyk and 117 bp between pyk and ldh. Plasmids containing pfk and pyk conferred phosphofructokinase and pyruvate kinase activity, respectively, on their host. The identity of ldh was established previously by the same approach (R. M. Llanos, A. J. Hillier, and B. E. Davidson, J. Bacteriol. 174:6956-6964, 1992). Each of the genes is preceded by a potential ribosome binding site. The operon is expressed in a 4.1-kb transcript. The 5' end of the transcript was determined to be a G nucleotide positioned 81 bp upstream from the pfk start codon. The pattern of codon usage within the operon is highly biased, with 11 unused amino acid codons. This degree of bias suggests that the operon is highly expressed. The three proteins encoded on the operon are key enzymes in the Embden-Meyerhoff pathway, the central pathway of energy production and lactic acid synthesis in L. lactis. For this reason, we have called the operon the las (lactic acid synthesis) operon.

Amino Acid Sequence↗

Codon usage in Pseudomonas aeruginosa.

We have generated a codon usage table for Pseudomonas aeruginosa. Codon usage in P. aeruginosa is extremely biased. In contrast to E. coli and yeast, P. aeruginosa preferentially uses those codons within a synonymous codon group with the strongest predicted codon-anticodon interaction. We were unable to correlate a particular codon usage pattern with predicted levels of mRNA expressivity. The choice of a third base reflects the high guanine plus cytosine content of the P. aeruginosa genome (67.2%) and cytosine is the preferred nucleotide for the third codon position.

Bacteriophages↗

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae↗

Quantifying the species-specificity in genomic signatures, synonymous codon choice, amino acid usage and G+C content.

Each prokaryote has a unique genomic signature as evidenced by a set of species-specific frequencies of short oligonucleotides. With respect to genomic signatures a bacterial genome is homogenous and the variation within a genome is smaller than the variations between genomes of different species. This study quantifies the species-specificity of genomic signatures in the complete genomes of 57 prokaryotes. The species-specificity in the genomic signature was related to the quantification of other sequence biases, such as G+C content, synonymous codon choice and amino acid usage. The results confirm that the genomic signature is genome-wide with high species-specificity in both coding and non-coding regions. In coding regions the species-specific bias in synonymous codon choice was comparable to the genomic signature, while the bias in amino acid usage only captured about 50% of the species-specific bias in the genomic signature. A correlation between the species-specificity in synonymous codon choice and amino acid usage was identified, in which proteins with species-specific amino acid usage were also coded with species-specific synonymous codon choice. However, we demonstrated that the G+C content captures only approximately 40% of the species-specificity in the genomic signature, and is insufficient to explain the species specificity in the non-coding regions. Thus, the species-specific bias in non-coding regions remains largely unknown. Further, we compared the genomic signature in relation to phylogenetic distance. This was performed in order to illustrate the feasibility of a hierarchical classification scheme in future applications of the described classification methodology in screening for horizontal gene transfer and biodiversity studies.

Amino Acids↗

Sequence analysis of the Alcaligenes eutrophus chromosomally encoded ribulose bisphosphate carboxylase large and small subunit genes and their gene products.

The nucleotide sequence of the chromosomally encoded ribulose bisphosphate carboxylase/oxygenase (RuBPCase) large (rbcL) and small (rbcS) subunit genes of the hydrogen bacterium Alcaligenes eutrophus ATCC 17707 was determined. We found that the two coding regions are separated by a 47-base-pair intergenic region, and both genes are preceded by plausible ribosome-binding sites. Cotranscription of the rbcL and rbcS genes has been demonstrated previously. The rbcL and rbcS genes encode polypeptides of 487 and 135 amino acids, respectively. Both genes exhibited similar codon usage which was highly biased and different from that of other organisms. The N-terminal amino acid sequence of both subunit proteins was determined by Edman degradation. No processing of the rbcS protein was detected, while the rbcL protein underwent a posttranslational loss of formylmethionyl. The A. eutrophus rbcL and rbcS proteins exhibited 56.8 to 58.3% and 35.6 to 38.5% amino acid sequence homology, respectively, with the corresponding proteins from cyanobacteria, eucaryotic algae, and plants. The A. eutrophus and Rhodospirillum rubrum rbcL proteins were only about 32% homologous. The N- and C-terminal sequences of both the rbcL and the rbcS proteins were among the most divergent regions. Known or proposed active site residues in other rbcL proteins, including Lys, His, Arg, and Asp residues, were conserved in the A. eutrophus enzyme. The A. eutrophus rbcS protein, like those of cyanobacteria, lacks a 12-residue internal sequence that is found in plant RuBPCase. Comparison of hydropathy profiles and secondary structure predictions by the method described by Chou and Fasman (P. Y. Chou and G. D. Fasman, Adv. Enzymol. 47:45-148, 1978) revealed striking similarities between A. eutrophus RuBPCase and other hexadecameric enzymes. This suggests that folding of the polypeptide chains is similar. The observed sequence homologies were consistent with the notion that both the rbcL and rbcS genes of the chemoautotroph A. eutrophus and the thus far characterized rbc genes of photosynthetic organisms have a common origin. This suggests that both subunit genes have a very ancient origin. The role of quaternary structure as a determinant of the rate of accepted amino acid substitution was examined. It is proposed that the sequence of the dimeric R. rubrum RuBPCase may be less conserved because there are fewer structural constraints for this RuBPCase than there are for hexadecameric enzymes.

Alcaligenes↗

Thermophilic prokaryotes have characteristic patterns of codon usage, amino acid composition and nucleotide content.

A number of recent studies have shown that thermophilic prokaryotes have distinguishable patterns of both synonymous codon usage and amino acid composition, indicating the action of natural selection related to thermophily. On the other hand, several other studies of whole genomes have illustrated that nucleotide bias can have dramatic effects on synonymous codon usage and also on the amino acid composition of the encoded proteins. This raises the possibility that the thermophile-specific patterns observed at both the codon and protein levels are merely reflections of a single underlying effect at the level of nucleotide composition. Moreover, such an effect at the nucleotide level might be due entirely to mutational bias. In this study, we have compared the genomes of thermophiles and mesophiles at three levels: nucleotide content, codon usage and amino acid composition. Our results indicate that the genomes of thermophiles are distinguishable from mesophiles at all three levels and that the codon and amino acid frequency differences cannot be explained simply by the patterns of nucleotide composition. At the nucleotide level, we see a consistent tendency for the frequency of adenine to increase at all codon positions within the thermophiles. Thermophiles are also distinguished by their pattern of synonymous codon usage for several amino acids, particularly arginine and isoleucine. At the protein level, the most dramatic effect is a two-fold decrease in the frequency of glutamine residues among thermophiles. These results indicate that adaptation to growth at high temperature requires a coordinated set of evolutionary changes affecting (i) mRNA thermostability, (ii) stability of codon-anticodon interactions and (iii) increased thermostability of the protein products. We conclude that elevated growth temperature imposes selective constraints at all three molecular levels: nucleotide content, codon usage and amino acid composition. In addition to these multiple selective effects, however, the genomes of both thermophiles and mesophiles are often subject to superimposed large changes in composition due to mutational bias.

Amino Acids↗

The nucleotide sequence of cDNA coding for the structural proteins of foot-and-mouth disease virus.

The complete nucleotide sequence of cDNA coding for the structural capsid polypeptides of foot-and-mouth disease virus (FMDV) (strain A(10)61) has been determined. Portions of the flanking sequence coding for the nonstructural proteins p20a and p52 are also provided. The three larger structural polypeptides VP1, VP2 and VP3 have unmodified Mrs of 23248, 24649 and 24213, respectively. The size of the smaller polypeptide, VP4, can only be estimated at 7360 because the 5'-limit of its coding region is not yet known with certainty. The sequence data for VP1 (the major immunising antigen) and the amino-terminal quarter of p52 are compared with the data of Kurz et al. (Nucl. Acids Res. 9 (1981) 1919-1931) for a different serotype (O1K). This shows that variation is much greater in the region coding for VP1 than in that coding for p52. This is reflected in the level of amino acid sequence variation predicted for the two proteins. Analysis of relative codon usage reveals a strong bias in favour of C and G over U and A in the third base position. The dinucleotide frequencies show a bias against A-U and U-A, and for A-C and C-A.

Aphthovirus↗

The complete genome of Bacillus subtilis: from sequence annotation to data management and analysis.

The completion of the entire 4.2-Mb genome sequence of the gram-positive bacterium Bacillus subtilis has been a milestone for biological studies on this model organism. This paper describes bioinformatics work related to this joint European and Japanese project: methods and strategies for gene annotation and detection of sequencing errors, using an integrated cooperative computer environment (Imagene); construction of a specialized database for data management and a WWW server for data retrieval (SubtiList); DNA sequence analysis, yielding striking results on oligonucleotide bias, repeated sequences, and codon usage, all landmarks of evolutionary events shaping the B. subtilis genome.

Amino Acid Sequence↗

An Acanthamoeba polyubiquitin gene and application of its promoter to the establishment of a transient transfection system.

We have isolated and sequenced a 2388 bp polyubiquitin encoding genomic DNA from Acanthamoeba encompassing two complete and one incomplete ubiquitin units. Codon usage frequency shows extreme bias. The deduced amino acid sequences of each unit are identical to each other and the same as that deduced from a previously sequenced Acanthamoeba castellanii cDNA. The upstream region of this gene, which contained some putative regulatory modules, was recovered by PCR (polymerase chain reaction) amplification and subcloning. This upstream fragment was ligated to the CAT (chloramphenicol acetyltransferase) gene in a eukaryotic expression plasmid and successfully applied to the establishment of an Acanthamoeba transient transfection system. Transfection was performed by electroporation and the optimal voltage was 4500 volts/cm at capacitance 25 microF. DEAE-dextran (25 microg/ml) added into the electroporation buffer increased the transfection efficiency by about 45%. The CAT activity was proportional to the amount of DNA transfected and reached the peak level 48 h after transfection. CAT assays showed that the polyubiquitin gene upstream fragment contains a functional promoter which is about 2.5 times as strong as a viral RSV-LTR promoter when driving CAT expression in Acanthamoeba.

Acanthamoeba↗

Ataxia-telangiectasia locus: sequence analysis of 184 kb of human genomic DNA containing the entire ATM gene.

Ataxia-telangiectasia (A-T) is an autosomal recessive disorder involving cerebellar degeneration, immunodeficiency, chromosomal instability, radiosensitivity, and cancer predisposition. The genomic organization of the A-T gene, designated ATM, was established recently. To date, more than 100 A-T-associated mutations have been reported in the ATM gene that do not support the existence of one or several mutational hotspots. To allow genotype/phenotype correlations it will be important to find additional ATM mutations. The nature and location of the mutations will also provide insights into the molecular processes that underly the disease. To facilitate the search for ATM mutations and to establish the basis for the identification of transcriptional regulatory elements, we have sequenced and report here 184,490 bp of genomic sequence from the human 11q22-23 chromosomal region containing the entire ATM gene, spanning 146 kb, and 10 kb of the 5'-region of an adjacent gene named E14/NPAT. The latter shares a bidirectional promoter with ATM and is transcribed in the opposite direction. The entire region is transcribed to approximately 85% and translated to 5%. Genome-wide repeats were found to constitute 37.2%, with LINE (17.1%) and Alu (14.6%) being the main repetitive elements. The high representation of LINE repeats is attributable to the presence of three full-length LINE-1s, inserted in the same orientation in introns 18 and 63 as well as downstream of the ATM gene. Homology searches suggest that ATM exon 2 could have derived from a mammalian interspersed repeat (MIR). Promoter recognition algorithms identified divergent promoter elements within the CpG island, which lies between the ATM and E14/NPAT genes, and provide evidence for a putative second ATM promoter located within intron 3, immediately upstream of the first coding exon. The low G+C level (38.1%) of the ATM locus is reflected in a strongly biased codon and amino acid usage of the gene.

Ataxia Telangiectasia↗

Cloning and nucleotide sequence of a leaf ferredoxin-nitrite reductase cDNA of rice.

A ferredoxin-nitrite reductase (EC 1.7.7.1) cDNA was isolated and sequenced from a lambda gt 11 cDNA library constructed from nitrate-induced greening shoots of rice (Oryza sativa L.) seedlings. The nucleotide sequence of the cDNA clone contains an open reading frame of 1788 nucleotides. There exists a strong bias for the third codon usage of G/C (95.5%) as in the case of the maize enzyme. The deduced amino acid sequence shows an overall homology to the maize (81%) and the dicot enzymes (70-74%), suggesting that the primary structure of ferredoxin-nitrite reductase is highly conserved in higher plants.

Amino Acid Sequence↗

Regulation of protein synthesis in Tetrahymena. RNA sequence sets of growing and starved cells.

The complexity of messenger RNA in growing or starved Tetrahymena thermophila is similar and unusually high (approximately 4.5 X 10(7) nucleotides). The complexity of nuclear RNA in growing cells (approximately 7.8 X 10(7) nucleotides) is only about 1.7 times that of mRNA. The concentration of complex class (rare) messages (approximately 53 copies/growing cell and approximately 11 copies/starved cell) is low in comparison to the size of the cell. The concentration of complex nuclear transcripts is also very low (approximately 0.7 copies/growing cell nucleus and approximately 2.6 copies/starved cell nucleus) considering that the macronucleus contains 45 to 90 copies of each single copy sequence. The complex sequence sets found on polysomes of growing and starved cells overlap about 80% and about 60% of the complex nuclear transcripts appear to be held in common. About 60% of macronuclear single copy DNA is transcribed in one or both physiological states. Although growing and starved cells have extremely different fractions of their messages loaded onto polysomes, within each cell type the complex messages in polysomal and nonpolysomal cytoplasmic fractions are indistinguishable, suggesting that exchange may occur between loaded and unloaded messages. Although T. thermophila DNA has an unusually low G + C content (23%), sequences coding for complex RNAs have base ratios similar to those of total DNA. Therefore, codon usage in Tetrahymena must be extremely biased towards adenine- and uridine-rich codons.

Animals↗

Hill-Robertson interference is a minor determinant of variations in codon bias across Drosophila melanogaster and Caenorhabditis elegans genomes.

According to population genetics models, genomic regions with lower crossing-over rates are expected to experience less effective selection because of Hill-Robertson interference (HRi). The effect of genetic linkage is thought to be particularly important for a selection of weak intensity such as selection affecting codon usage. Consistent with this model, codon bias correlates positively with recombination rate in Drosophila melanogaster and Caenorhabditis elegans. However, in these species, the G+C content of both noncoding DNA and synonymous sites correlates positively with recombination, which suggests that mutation patterns and recombination are associated. To remove this effect of mutation patterns on codon bias, we used the synonymous sites of lowly expressed genes that are expected to be effectively neutral sites. We measured the differences between codon biases of highly expressed genes and their lowly expressed neighbors. In D. melanogaster we find that HRi weakly reduces selection on codon usage of genes located in regions of very low recombination; but these genes only comprise 4% of the total. In C. elegans we do not find any evidence for the effect of recombination on selection for codon bias. Computer simulations indicate that HRi poorly enhances codon bias if the local recombination rate is greater than the mutation rate. This prediction of the model is consistent with our data and with the current estimate of the mutation rate in D. melanogaster. The case of C. elegans, which is highly self-fertilizing, is discussed. Our results suggest that HRi is a minor determinant of variations in codon bias across the genome.

Animals↗