Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Comparison of synonymous codon distribution patterns of bacteriophage and host genomes.

Synonymous codon usage patterns of bacteriophage and host genomes were compared. Two indexes, G + C base composition of a gene (fgc) and fraction of translationally optimal codons of the gene (fop), were used in the comparison. Synonymous codon usage data of all the coding sequences on a genome are represented as a cloud of points in the plane of fop vs. fgc. The Escherichia coli coding sequences appear to exhibit two phases, "rising" and "flat" phases. Genes that are essential for survival and are thought to be native are located in the flat phase, while foreign-type genes from prophages and transposons are found in the rising phase with a slope of nearly unity in the fgc vs. fop plot. Synonymous codon distribution patterns of genes from temperate phages P4, P2, N15 and lambda are similar to the pattern of E. coli rising phase genes. In contrast, genes from the virulent phage T7 or T4, for which a phage-encoded DNA polymerase is identified, fall in a linear curve with a slope of nearly zero in the fop vs. fgc plane. These results may suggest that the G + C contents for T7, T4 and E. coli flat phase genes are subject to the directional mutation pressure and are determined by the DNA polymerase used in the replication. There is significant variation in the fop values of the phage genes, suggesting an adjustment to gene expression level. Similar analyses of codon distribution patterns were carried out for Haemophilus influenzae, Bacillus subtilis, Mycobacterium tuberculosis and their phages with complete genomic sequences available.

Bacillus subtilis↗

Highly expressed and alien genes of the Synechocystis genome.

Comparisons of codon frequencies of genes to several gene classes are used to characterize highly expressed and alien genes on the SYNECHOCYSTIS: PCC6803 genome. The primary gene classes include the ensemble of all genes (average gene), ribosomal protein (RP) genes, translation processing factors (TF) and genes encoding chaperone/degradation proteins (CH). A gene is predicted highly expressed (PHX) if its codon usage is close to that of the RP/TF/CH standards but strongly deviant from the average gene. Putative alien (PA) genes are those for which codon usage is significantly different from all four classes of gene standards. In SYNECHOCYSTIS:, 380 genes were identified as PHX. The genes with the highest predicted expression levels include many that encode proteins vital for photosynthesis. Nearly all of the genes of the RP/TF/CH gene classes are PHX. The principal glycolysis enzymes, which may also function in CO(2) fixation, are PHX, while none of the genes encoding TCA cycle enzymes are PHX. The PA genes are mostly of unknown function or encode transposases. Several PA genes encode polypeptides that function in lipopolysaccharide biosynthesis. Both PHX and PA genes often form significant clusters (operons). The proteins encoded by PHX and PA genes are described with respect to functional classifications, their organization in the genome and their stoichiometry in multi-subunit complexes.

Codon↗

Analysis of messages expressed by Echinostoma paraensei miracidia and sporocysts, obtained by random EST sequencing.

A lambdaZAP Express cDNA library was constructed with mRNA obtained from immature miracidia within eggs, hatched miracidia, and sporocysts of Echinostoma paraensei. This cDNA library was amplified and 213 expressed sequence tag (EST) sequences (averaging 466 nucleotides in length) were obtained. The mean percentage of unresolved bases within the EST sequences was 0.4%, ranging from 0 to 4.6%. The 213 ESTs represent 151 unique messages. BLAST (version 2.0.8) analysis disclosed that 64 unique E. paraensei messages (42.4%) had significant similarities (BLAST score < or =e-5), at deduced amino acid or nucleotide levels, with known sequences in the nonredundant GenBank databases or the dbEST database (NCBI). The remainder, 57.6% of the unique EST-encoded messages, scored nonsignificant hits. Most of the E. paraensei messages that could be assigned a cellular role based on sequence similarities were involved in gene/protein expression. Several ESTs scored highest similarities with sequences obtained from trematode species. A total of 22,560 nucleotides present in open reading frames from ESTs that aligned with known sequences was used to determine codon usage for E. paraensei. Analysis of a subset of eight ESTs that contained full-length open reading frames did not reveal a bias in codon usage. Also, EST sequences were found to contain 3' untranslated regions with an average length of 69.9 +/- 88.4 nucleotides (n = 46). The EST sequences were submitted to GenBank/dbEST, adding to the 51 available Echinostoma-derived sequences, to provide reference information for both phylogenetic analysis and study of general trematode biology.

Animals↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

Fine structural features of the chloroplast genome: comparison of the sequenced chloroplast genomes.

The entire nucleotide sequences of the rice, tobacco and liverwort chloroplast genomes have been determined. We compared all the chloroplast genes, open reading frames and spacer regions in the plastid genomes of these three species in order to elucidate general structural features of the chloroplast genome. Analyses of homology, GC content and codon usage of the genes enabled us to classify them into two groups: photosynthesis genes and genetic system genes. Based on comparisons of homology, GC content and codon usage, unidentified ORFs can also be assigned to each of these groups such that it is possible to speculate about the functions of products which may be produced by these ORFs. The spacer regions and intron sequences were compared and found to have no obvious homology between rice and liverwort or between tobacco and liverwort.

Base Composition↗

Variation in G + C-content and codon choice: differences among synonymous codon groups in vertebrate genes.

The relationship between G + C-content and codon usage in genes of human, mus, rat, bovine and chicken nuclear genomes was investigated. Correlation and lineal regression analyses were carried out on plots that related the frequency of each codon within each synonymous codon group to the G + C-content of the coding sequence as a whole. Under GC pressure, in most of the quartet codon groups there is a preferential choice of the C-ending codon, except in leucine and valine codon groups where the choice of the G-ending codon is preferred. Among ducts, the choice of codons specifying phenylalanine and glutamate shows the strongest dependence on G + C-content. The relationship found between G + C-content and codon usage in these genomes correlate with taxonomic distance.

Animals↗

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids↗

Characterization of a highly expressed lignin peroxidase-encoding gene from the basidiomycete Phanerochaete chrysosporium.

The genomic clone, LG2, encoding LiP2, the major lignin peroxidase (LiP) isozyme from Phanerochaete chrysosporium strain OGC101, was isolated and characterized. The 5'-untranslated region of LG2 contains sequences similar to CRE and XRE promoter elements. Comparison with its transcript indicates that eight introns, each less than 59 bp, interrupt the coding sequence. Comparison with genes encoding other LiP isozymes shows five related patterns of intron location, whose incidence coincides with described LiP structural subfamilies. Codon bias indices calculated for all known P. chrysosporium genes, including trpC and genes encoding LiP, MnP, and exo-cellobiohydrolase I, demonstrate that LG2 has the most biased codon usage. We conclude that subdivisions of the LiP family may be based on intron location in the encoding genes, and that ranking of isozyme production levels can be estimated by the extent of bias in codon usage in the cognate gene.

Amino Acid Sequence↗

Increasing expression of P450 and P450-reductase proteins from monocots in heterologous systems.

Monocotyledonous crop plants are usually more resistant to herbicides than grass weeds and most dicots. Their resistance to herbicides is mediated in many cases by P450 oxygenases. Monocots thus constitute an appealing source of P450 enzymes for manipulating herbicide resistance and recombinant forms of the major xenobiotic metabolizing mooxygenases are potential tools for the optimization of new active molecules. We report here the isolation and functional characterization of the first P450 and P450 reductase coding sequences from wheat. The first attempts at expressing these cDNAs in yeast and tobacco led to levels of protein, which were extremely low, often not even detectable. The wheat P450 cDNAs were efficiently transcribed, but no protein or activity was found. Wheat coding sequences, like those of other monocots, are characterized by a high GC content and by a related strong bias of codon usage, different from that observed in yeast or dicots. Complete recoding of genes being costly, the reengineering their 5'-end using a single PCR megaprimer designed to comply with codon usage of the host was attempted. It was sufficient to relieve translation inhibition and to obtain good levels of protein expression. The same strategy also resulted in a dramatic increase in protein expression in tobacco. A basis for the success of such a partial recoding strategy, much easier and cheaper than complete recoding of the cDNA, is proposed.

Amino Acid Sequence↗

The nucleotide sequence of an Escherichia coli operon containing genes for the tRNA(m1G)methyltransferase, the ribosomal proteins S16 and L19 and a 21-K polypeptide.

The nucleotide sequence of a 4.6-kb SalI-EcoRI DNA fragment including the trmD operon, located at min 56 on the Escherichia coli K-12 chromosome, has been determined. The trmD operon encodes four polypeptides: ribosomal protein S16 (rpsP), 21-K polypeptide (unknown function), tRNA-(m1G)methyltransferase (trmD) and ribosomal protein L19 (rplS), in that order. In addition, the 4.6-kb DNA fragment encodes a 48-K and a 16-K polypeptide of unknown functions which are not part of the trmD operon. The mol. wt. of tRNA(m1G)methyltransferase determined from the DNA sequence is 28 424. The probable locations of promoter and terminator of the trmD operon are suggested. The translational start of the trmD gene was deduced from the known NH2-terminal amino acid sequence of the purified enzyme. The intercistronic regions in the operon vary from 9 to 40 nucleotides, supporting the earlier conclusion that the four genes are co-transcribed, starting at the major promoter in front of the rpsP gene. Since it is known that ribosomal proteins are present at 8000 molecules/genome and the tRNA-(m1G)methyltransferase at only approximately 80 molecules/genome in a glucose minimal culture, some powerful regulatory device must exist in this operon to maintain this non-coordinate expression. The codon usage of the two ribosomal protein genes is similar to that of other ribosomal protein genes, i.e., high preference for the most abundant tRNA isoaccepting species. The trmD gene has a codon usage typical for a protein made in low amount in accordance with the low number of tRNA-(m1G)methyltransferase molecules found in the cell.

Bacterial Proteins↗

The cytochrome b region in the mitochondrial DNA of the ant Tetraponera rufoniger: sequence divergence in Hymenoptera may be associated with nucleotide content.

Polymerase chain reaction (PCR) followed by sequencing of single-stranded DNA yielded sequence information from the cytochrome b (cyt b) region in mitochondrial DNA from the ant Tetraponera rufoniger. Compared with the cyt b genes from Apis mellifera, Drosophila melanogaster, and D. yakuba, the overall A+T content (A+T%) of that of T. rufoniger is lower (69.9% vs 80.7%, 74.2%, and 73.9%, respectively) than those of the other three. The codon usage in the cyt b gene of T. rufoniger is biased although not as much as in A. mellifera, D. melanogaster, and D. yakuba; T. rufoniger has eight unused codons whereas D. melanogaster, D. yakuba, and A. mellifera have 21, 20, and 23, respectively. The inferred cyt b polypeptide chain (PPC) of T. rufoniger has diverged at least as much from a common ancestor with D. yakuba as has that of A. mellifera (approximately 3.5 vs approximately 2.9). Despite the lower A+T%, the relative frequencies of amino acids in the cyt b PPC of T. rufoniger are significantly (P < 0.05) associated with the content of adenine and thymine (A+T%) and size of codon families. The mitochondrially located cytochrome oxidase subunit II genes (CO-II) of endopterygote insects have significantly higher average A+T% (approximately 75%) than those of exopterygous (approximately 69%) and paleopterous (approximately 69%) insects. The increase in A+T% of endopterygote insects occurred in Upper Carboniferous and coincided with a significant acceleration of PPC divergence. However, acceleration of PPC divergence is not significantly correlated with the increase of the A+T% (P > 0.1). The high A+T%, the biased codon usage, and the increased PPC divergence of Hymenoptera can in that respect most easily be explained by directional mutation pressure which began in the Upper Carboniferous and still occurs in most members of the order. Given the roughly identical A+T% of the cyt b and CO-II genes from the other insects whose DNA sequences are known (A. mellifera, D. melanogaster, and D. yakuba), it seems most likely that the A+T% of T. rufoniger declined secondarily within the last 100 Myr as a result of a reduced directional mutation pressure.

Amino Acid Sequence↗

Evolutionary lability of context-dependent codon bias in bacteria.

In bacteria, synonymous codon usage can be considerably affected by base composition at neighboring sites. Such context-dependent biases may be caused by either selection against specific nucleotide motifs or context-dependent mutation biases. Here we consider the evolutionary conservation of context-dependent codon bias across 11 completely sequenced bacterial genomes. In particular, we focus on two contextual biases previously identified in Escherichia coli; the avoidance of out-of-frame stop codons and AGG motifs. By identifying homologues of E. coli genes, we also investigate the effect of gene expression level in Haemophilus influenzae and Mycoplasma genitalium. We find that while context-dependent codon biases are widespread in bacteria, few are conserved across all species considered. Avoidance of out-of-frame stop codons does not apply to all stop codons or amino acids in E. coli, does not hold for different species, does not increase with gene expression level, and is not relaxed in Mycoplasma spp., in which the canonical stop codon, TGA, is recognized as tryptophan. Avoidance of AGG motifs shows some evolutionary conservation and increases with gene expression level in E. coli, suggestive of the action of selection, but the cause of the bias differs between species. These results demonstrate that strong context-dependent forces, both selective and mutational, operate on synonymous codon usage but that these differ considerably between genomes.

Codon↗

Effect of codon optimization on expression levels of a functionally folded malaria vaccine candidate in prokaryotic and eukaryotic expression systems.

We have produced two synthetic genes that code for the F2 domain located within region II of the 175-kDa Plasmodium falciparum erythrocyte binding antigen (EBA-175) to determine the effects of codon alteration on protein expression in homologous and heterologous host systems. EBA-175 plays a key role in the process of merozoite invasion into erythrocytes through a specific receptor-ligand interaction. The F2 domain of EBA-175 is the ligand that binds to the glycophorin A receptor on human erythrocytes and is therefore a target of vaccine development efforts. We designed synthetic genes based on P. falciparum, Escherichia coli, and Pichia codon usage and expressed recombinant F2 in E. coli and Pichia pastoris. Compared to the expression of the native F2 sequence, conversion to prokaryote (E. coli)- or eukaryote (Pichia)-based codon usage dramatically improved the levels of recombinant protein expression in both E. coli and P. pastoris. The majority of the protein expressed in E. coli, however, was produced as inclusion bodies. The protein expressed in P. pastoris, on the other hand, was expressed as a secreted, soluble protein. The P. pastoris-produced protein was superior to that produced in E. coli based on its ability to bind to red blood cells. Consistent with these observations, the antibodies generated against the Pichia-produced protein prevented the binding of recombinant EBA to red blood cells. These antibodies recognize EBA-175 present on merozoites as well as in sporozoites by immunofluorescence. Our results suggest that the Pichia-based EBA-F2 vaccine construct has further potential to be developed for clinical use.

Animals↗

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family↗

Chloramphenicol resistance in Campylobacter coli: nucleotide sequence, expression, and cloning vector construction.

A chloramphenicol-resistance determinant (CmR), originally cloned from Campylobacter coli plasmid pNR9589 in Japan, was isolated and the nucleotide sequence determined, which contained an open reading frame of 621 bp. The gene product was identified as Cm acetyltransferase (CAT), which had a putative amino acid sequence that showed 43% to 57% identity with other CAT proteins of both Gram+ and Gram- origin. Although expression of the cat gene was constitutive in both C. coli and Escherichia coli, results of primer extension experiments indicated that transcription was initiated at different sites in these two species. A kanamycin-resistance determinant, identified as the aphA-3 gene, was located downstream from the cat gene. The codon usage of the cat gene is very different from that used in E. coli, however, the CAT polypeptide was synthesized in large amounts in E. coli maxicells. Therefore, the codon usage bias is not one of the obstacles which affects Campylobacter spp. gene expression in E. coli. New Campylobacter cloning vectors were constructed in this study.

Amino Acid Sequence↗

Theory of degenerate coding and informational parameters of protein coding genes.

The theory of degenerate coding is presented in a way enabling further application to molecular biology. There are two kinds of redundancy of a degenerate code. The first is due to the excess in codon length and the second to the code degeneracy. If the code is asymmetrically degenerate, the second kind of redundancy can be profitable for control of error rate. This control can be performed just by selective synonymous codon usage. Utilisation of the genetic code is partially influenced by this theoretical possibility. In particular the degree of error protectivity is well correlated with deviation from equiprobability in synonymous codon usage. The biological significance of this fact is discussed.

Animals↗

A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes.

BACKGROUND: Correlations between genome composition (in terms of GC content) and usage of particular codons and amino acids have been widely reported, but poorly explained. We show here that a simple model of processes acting at the nucleotide level explains codon usage across a large sample of species (311 bacteria, 28 archaea and 257 eukaryotes). The model quantitatively predicts responses (slope and intercept of the regression line on genome GC content) of individual codons and amino acids to genome composition. RESULTS: Codons respond to genome composition on the basis of their GC content relative to their synonyms (explaining 71-87% of the variance in response among the different codons, depending on measure). Amino-acid responses are determined by the mean GC content of their codons (explaining 71-79% of the variance). Similar trends hold for genes within a genome. Position-dependent selection for error minimization explains why individual bases respond differently to directional mutation pressure. CONCLUSIONS: Our model suggests that GC content drives codon usage (rather than the converse). It unifies a large body of empirical evidence concerning relationships between GC content and amino-acid or codon usage in disparate systems. The relationship between GC content and codon and amino-acid usage is ahistorical; it is replicated independently in the three domains of living organisms, reinforcing the idea that genes and genomes at mutation/selection equilibrium reproduce a unique relationship between nucleic acid and protein composition. Thus, the model may be useful in predicting amino-acid or nucleotide sequences in poorly characterized taxa.

Amino Acids↗

Mutation and selection at silent and replacement sites in the evolution of animal mitochondrial DNA.

Two patterns are presented that illustrate the interaction of mutation and selection in the evolution of animal mtDNA: 1) variation among taxa in the ratio of polymorphism to divergence (rpd) at silent and replacement sites in protein-coding genes, and 2) strand-differences in polymorphism and divergence at 'silent' sites that suggest a mutation-selection balance in the evolution of codon usage. Cytochrome b data from GenBank show that about half of the species pairs tested have a significant excess of amino acid polymorphism, relative to divergence. The remaining half of species pairs do not depart from neutrality, but generally do show an excess of amino acid polymorphism. Sequences from Drosophila pseudoobscura displaying a signature of an expanding population show a slight, but non-significant, deficiency of amino acid polymorphism suggestive of recently intensified selection on mildly deleterious mutations. Genes whose reading frames lie on the major coding strand of Drosophila mtDNA show a preponderance of T- > C substitutions, while genes encoded on the minor strand experience more A- > G than T- > C substitutions between species at both silent and replacement sites. However, silent mutations at third codon positions are introduced into the population in proportions opposite to those observed as fixed differences between species (e.g., an excess of T- > C polymorphisms are found at the ND5 gene on the minor coding strand). The high A + T content of insect mtDNAs imposes strong codon usage bias favoring A-ending and T-ending codons resulting in a distinct mutation-selection balance for genes encoded on opposites strands. Thus, at both replacement and silent sites, mutations that appear to be constrained in terms of divergence between species are in excess within species. The data suggest that mildly deleterious mutations are common in mitochondrial genes. A test of this, and a competing, hypothesis is proposed that requires additional sequence surveys of polymorphism and divergence. An important challenge is to tease apart the impact of mutation and selection on levels of polymorphism versus divergence in a genome that does not generally recombine.

Animals↗