Search PubMedSearch

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Comparison of synonymous codon distribution patterns of bacteriophage and host genomes.

Synonymous codon usage patterns of bacteriophage and host genomes were compared. Two indexes, G + C base composition of a gene (fgc) and fraction of translationally optimal codons of the gene (fop), were used in the comparison. Synonymous codon usage data of all the coding sequences on a genome are represented as a cloud of points in the plane of fop vs. fgc. The Escherichia coli coding sequences appear to exhibit two phases, "rising" and "flat" phases. Genes that are essential for survival and are thought to be native are located in the flat phase, while foreign-type genes from prophages and transposons are found in the rising phase with a slope of nearly unity in the fgc vs. fop plot. Synonymous codon distribution patterns of genes from temperate phages P4, P2, N15 and lambda are similar to the pattern of E. coli rising phase genes. In contrast, genes from the virulent phage T7 or T4, for which a phage-encoded DNA polymerase is identified, fall in a linear curve with a slope of nearly zero in the fop vs. fgc plane. These results may suggest that the G + C contents for T7, T4 and E. coli flat phase genes are subject to the directional mutation pressure and are determined by the DNA polymerase used in the replication. There is significant variation in the fop values of the phage genes, suggesting an adjustment to gene expression level. Similar analyses of codon distribution patterns were carried out for Haemophilus influenzae, Bacillus subtilis, Mycobacterium tuberculosis and their phages with complete genomic sequences available.

Bacillus subtilis

Fine structural features of the chloroplast genome: comparison of the sequenced chloroplast genomes.

The entire nucleotide sequences of the rice, tobacco and liverwort chloroplast genomes have been determined. We compared all the chloroplast genes, open reading frames and spacer regions in the plastid genomes of these three species in order to elucidate general structural features of the chloroplast genome. Analyses of homology, GC content and codon usage of the genes enabled us to classify them into two groups: photosynthesis genes and genetic system genes. Based on comparisons of homology, GC content and codon usage, unidentified ORFs can also be assigned to each of these groups such that it is possible to speculate about the functions of products which may be produced by these ORFs. The spacer regions and intron sequences were compared and found to have no obvious homology between rice and liverwort or between tobacco and liverwort.

Base Composition

Variation in G + C-content and codon choice: differences among synonymous codon groups in vertebrate genes.

The relationship between G + C-content and codon usage in genes of human, mus, rat, bovine and chicken nuclear genomes was investigated. Correlation and lineal regression analyses were carried out on plots that related the frequency of each codon within each synonymous codon group to the G + C-content of the coding sequence as a whole. Under GC pressure, in most of the quartet codon groups there is a preferential choice of the C-ending codon, except in leucine and valine codon groups where the choice of the G-ending codon is preferred. Among ducts, the choice of codons specifying phenylalanine and glutamate shows the strongest dependence on G + C-content. The relationship found between G + C-content and codon usage in these genomes correlate with taxonomic distance.

Animals

Switches in species-specific codon preferences: the influence of mutation biases.

A model of synonymous codon usage is developed in which the most frequent codons are selectively advantageous because of their coadaptation with tRNA abundances. Random drift opposes the progress of this coevolution by pushing codon frequencies in the direction of the frequency that would result from mutation in the absence of selection. It is predicted that, within a certain range, an increased mutation bias away from an advantageous codon has little influence on its usage in highly expressed genes. However, a subsequent small increase in mutation bias over a critical range leads to a large reduction in the frequency of the codon. The switch in preference from one synonym to another is a sharp transition, with no stable intermediate state in which neither codon is advantageous. Codon usage patterns were compared among three related bacterial species of differing genomic G & C contents, Escherichia coli, Serratia marcescens, and Proteus vulgaris. It was found that although changes in mutation biases do not always result in switches in codon preferences, some switches have occurred in the direction of species-specific mutation biases. Fluctuating mutation biases may therefore be the main cause of differences between species in their codon preferences.

Amino Acids

Characterization of a highly expressed lignin peroxidase-encoding gene from the basidiomycete Phanerochaete chrysosporium.

The genomic clone, LG2, encoding LiP2, the major lignin peroxidase (LiP) isozyme from Phanerochaete chrysosporium strain OGC101, was isolated and characterized. The 5'-untranslated region of LG2 contains sequences similar to CRE and XRE promoter elements. Comparison with its transcript indicates that eight introns, each less than 59 bp, interrupt the coding sequence. Comparison with genes encoding other LiP isozymes shows five related patterns of intron location, whose incidence coincides with described LiP structural subfamilies. Codon bias indices calculated for all known P. chrysosporium genes, including trpC and genes encoding LiP, MnP, and exo-cellobiohydrolase I, demonstrate that LG2 has the most biased codon usage. We conclude that subdivisions of the LiP family may be based on intron location in the encoding genes, and that ranking of isozyme production levels can be estimated by the extent of bias in codon usage in the cognate gene.

Amino Acid Sequence

The nucleotide sequence of an Escherichia coli operon containing genes for the tRNA(m1G)methyltransferase, the ribosomal proteins S16 and L19 and a 21-K polypeptide.

The nucleotide sequence of a 4.6-kb SalI-EcoRI DNA fragment including the trmD operon, located at min 56 on the Escherichia coli K-12 chromosome, has been determined. The trmD operon encodes four polypeptides: ribosomal protein S16 (rpsP), 21-K polypeptide (unknown function), tRNA-(m1G)methyltransferase (trmD) and ribosomal protein L19 (rplS), in that order. In addition, the 4.6-kb DNA fragment encodes a 48-K and a 16-K polypeptide of unknown functions which are not part of the trmD operon. The mol. wt. of tRNA(m1G)methyltransferase determined from the DNA sequence is 28 424. The probable locations of promoter and terminator of the trmD operon are suggested. The translational start of the trmD gene was deduced from the known NH2-terminal amino acid sequence of the purified enzyme. The intercistronic regions in the operon vary from 9 to 40 nucleotides, supporting the earlier conclusion that the four genes are co-transcribed, starting at the major promoter in front of the rpsP gene. Since it is known that ribosomal proteins are present at 8000 molecules/genome and the tRNA-(m1G)methyltransferase at only approximately 80 molecules/genome in a glucose minimal culture, some powerful regulatory device must exist in this operon to maintain this non-coordinate expression. The codon usage of the two ribosomal protein genes is similar to that of other ribosomal protein genes, i.e., high preference for the most abundant tRNA isoaccepting species. The trmD gene has a codon usage typical for a protein made in low amount in accordance with the low number of tRNA-(m1G)methyltransferase molecules found in the cell.

Bacterial Proteins

The cytochrome b region in the mitochondrial DNA of the ant Tetraponera rufoniger: sequence divergence in Hymenoptera may be associated with nucleotide content.

Polymerase chain reaction (PCR) followed by sequencing of single-stranded DNA yielded sequence information from the cytochrome b (cyt b) region in mitochondrial DNA from the ant Tetraponera rufoniger. Compared with the cyt b genes from Apis mellifera, Drosophila melanogaster, and D. yakuba, the overall A+T content (A+T%) of that of T. rufoniger is lower (69.9% vs 80.7%, 74.2%, and 73.9%, respectively) than those of the other three. The codon usage in the cyt b gene of T. rufoniger is biased although not as much as in A. mellifera, D. melanogaster, and D. yakuba; T. rufoniger has eight unused codons whereas D. melanogaster, D. yakuba, and A. mellifera have 21, 20, and 23, respectively. The inferred cyt b polypeptide chain (PPC) of T. rufoniger has diverged at least as much from a common ancestor with D. yakuba as has that of A. mellifera (approximately 3.5 vs approximately 2.9). Despite the lower A+T%, the relative frequencies of amino acids in the cyt b PPC of T. rufoniger are significantly (P < 0.05) associated with the content of adenine and thymine (A+T%) and size of codon families. The mitochondrially located cytochrome oxidase subunit II genes (CO-II) of endopterygote insects have significantly higher average A+T% (approximately 75%) than those of exopterygous (approximately 69%) and paleopterous (approximately 69%) insects. The increase in A+T% of endopterygote insects occurred in Upper Carboniferous and coincided with a significant acceleration of PPC divergence. However, acceleration of PPC divergence is not significantly correlated with the increase of the A+T% (P > 0.1). The high A+T%, the biased codon usage, and the increased PPC divergence of Hymenoptera can in that respect most easily be explained by directional mutation pressure which began in the Upper Carboniferous and still occurs in most members of the order. Given the roughly identical A+T% of the cyt b and CO-II genes from the other insects whose DNA sequences are known (A. mellifera, D. melanogaster, and D. yakuba), it seems most likely that the A+T% of T. rufoniger declined secondarily within the last 100 Myr as a result of a reduced directional mutation pressure.

Amino Acid Sequence

Molecular Evolution and Expression Analysis of the ADH Gene Family in Apple Bud Mutants.

Alcohol dehydrogenase (ADH) catalyzes the reduction of aldehydes to alcohols, key precursor substrates for volatile ester biosynthesis, which determines the characteristic aroma of apple fruit. However, a comprehensive genome-wide investigation of the ADH gene family in apple has been lacking. In this study, we systematically identified ADH genes in the apple genome using integrated bioinformatics approaches, including phylogenetic analysis, synteny evaluation, promoter cis-element prediction, codon usage bias assessment, and protein interaction network modeling. Expression patterns were examined through transcriptomic data and validated by RT-qPCR analysis across different organs and among 'Red Delicious' and its four bud mutant lines. We identified 44 ADH genes, with 12 forming a prominent cluster on chromosome 1. RT-qPCR analysis revealed that MdADH20 was dramatically upregulated in the 'Red Chief' mutant (relative expression of 59.38), suggesting its pivotal role. Phylogenetic analysis revealed a close evolutionary relationship with wild strawberry. The encoded proteins were generally stable and predominantly localized to the cytoplasm. Promoter analysis showed enrichment of growth/development-related and ARE elements, while codon usage analysis identified AGA, GCU, GUU, and CUU as preferred codons. Protein interaction prediction suggested MdADH19 and MdADH20 as hub proteins. Expression profiling and RT-qPCR further identified MdADH20 as a core candidate gene, characterized by its stable and high expression, particularly in the 'Red Delicious' mutant. Its central position in the predicted protein-protein interaction network suggests a potential regulatory role in the aroma biosynthesis pathway of apple fruit. This study provides the first systematic genome-wide characterization of the apple ADH gene family, establishing a theoretical groundwork for deciphering aroma biosynthesis mechanisms and offering potential target genes for flavor improvement through bud mutation breeding strategies.

ADH gene family

Chloramphenicol resistance in Campylobacter coli: nucleotide sequence, expression, and cloning vector construction.

A chloramphenicol-resistance determinant (CmR), originally cloned from Campylobacter coli plasmid pNR9589 in Japan, was isolated and the nucleotide sequence determined, which contained an open reading frame of 621 bp. The gene product was identified as Cm acetyltransferase (CAT), which had a putative amino acid sequence that showed 43% to 57% identity with other CAT proteins of both Gram+ and Gram- origin. Although expression of the cat gene was constitutive in both C. coli and Escherichia coli, results of primer extension experiments indicated that transcription was initiated at different sites in these two species. A kanamycin-resistance determinant, identified as the aphA-3 gene, was located downstream from the cat gene. The codon usage of the cat gene is very different from that used in E. coli, however, the CAT polypeptide was synthesized in large amounts in E. coli maxicells. Therefore, the codon usage bias is not one of the obstacles which affects Campylobacter spp. gene expression in E. coli. New Campylobacter cloning vectors were constructed in this study.

Amino Acid Sequence

Theory of degenerate coding and informational parameters of protein coding genes.

The theory of degenerate coding is presented in a way enabling further application to molecular biology. There are two kinds of redundancy of a degenerate code. The first is due to the excess in codon length and the second to the code degeneracy. If the code is asymmetrically degenerate, the second kind of redundancy can be profitable for control of error rate. This control can be performed just by selective synonymous codon usage. Utilisation of the genetic code is partially influenced by this theoretical possibility. In particular the degree of error protectivity is well correlated with deviation from equiprobability in synonymous codon usage. The biological significance of this fact is discussed.

Animals

Mutation and selection at silent and replacement sites in the evolution of animal mitochondrial DNA.

Two patterns are presented that illustrate the interaction of mutation and selection in the evolution of animal mtDNA: 1) variation among taxa in the ratio of polymorphism to divergence (rpd) at silent and replacement sites in protein-coding genes, and 2) strand-differences in polymorphism and divergence at 'silent' sites that suggest a mutation-selection balance in the evolution of codon usage. Cytochrome b data from GenBank show that about half of the species pairs tested have a significant excess of amino acid polymorphism, relative to divergence. The remaining half of species pairs do not depart from neutrality, but generally do show an excess of amino acid polymorphism. Sequences from Drosophila pseudoobscura displaying a signature of an expanding population show a slight, but non-significant, deficiency of amino acid polymorphism suggestive of recently intensified selection on mildly deleterious mutations. Genes whose reading frames lie on the major coding strand of Drosophila mtDNA show a preponderance of T- > C substitutions, while genes encoded on the minor strand experience more A- > G than T- > C substitutions between species at both silent and replacement sites. However, silent mutations at third codon positions are introduced into the population in proportions opposite to those observed as fixed differences between species (e.g., an excess of T- > C polymorphisms are found at the ND5 gene on the minor coding strand). The high A + T content of insect mtDNAs imposes strong codon usage bias favoring A-ending and T-ending codons resulting in a distinct mutation-selection balance for genes encoded on opposites strands. Thus, at both replacement and silent sites, mutations that appear to be constrained in terms of divergence between species are in excess within species. The data suggest that mildly deleterious mutations are common in mitochondrial genes. A test of this, and a competing, hypothesis is proposed that requires additional sequence surveys of polymorphism and divergence. An important challenge is to tease apart the impact of mutation and selection on levels of polymorphism versus divergence in a genome that does not generally recombine.

Animals

Codon catalog usage is a genome strategy modulated for gene expressivity.

The nucleic acid sequence bank now contains 161 mRNAs, 43 new genes are added. One sequence, that of B. mori fibroin, is dropped due to uncertainty on the starting point for translation. Frequencies of all codons are given for each gene added and for each genome type in the total bank. A new series of correspondence analyses on codon use is presented, substantiating the genome hypothesis. Internal regulation of mRNA expression by different third base choices between quartet and duet codons is proposed for bacterial genes.

Amino Acid Sequence

The targeting of somatic hypermutation.

Somatic hypermutation does not occur randomly within immunoglobulin V genes but, rather, is preferentially targeted to certain nucleotide positions (hot spots) and away from others (cold spots). Cold spots often coincide with residues essential for V gene folding. Hotspots, which appear to be strategically located to favour affinity maturation, are most frequently located in the CDRs (particularly CDR1) though conserved hotspots are also found at the base of FR3. Hotspots are in part created by local DNA sequence and the strong biases of codon usage in V genes indicate that the genes have evolved such that somatic hypermutation is targeted to those parts of the V where it is likely to prove most useful. These features of mutational hotspots and biased codon usage are also evident in V genes of lower animals suggesting that diversification by strategic targeting of non-templated mutation may have evolved early in antigen receptor evolution.

Animals

Contrasting patterns of evolutionary divergence within the Acinetobacter calcoaceticus pca operon.

The six enzymes required for catabolism of protocatechuate to succinate and acetylCoA are encoded by the pca genes in the Gram-bacterium, Acinetobacter calcoaceticus. The clustered A. calcoaceticus cat genes encode an analogous set of enzymes associated with the metabolic dissimilation of catechol. The nucleotide (nt) sequences of pcaIJFB and pcaK, reported here, complete evidence showing that all of the pca structural genes are tightly grouped in the order pcaIJFBDKCHG within a single operon. The pcaIJF region is nearly identical in nt sequence to the A. calcoaceticus catIDJF region which exhibits a G+C content and a codon usage pattern exceptional for A. calcoaceticus. In contrast, pcaD, pcaC, pcaH and pcaG have diverged substantially from their evolutionary counterparts in the cat region; all of these divergent genes exhibit G+C contents and codon usage patterns that are typical for A. calcoaceticus. The pcaIJF and catIJF regions are known to exchange DNA sequence information, and this property may have contributed to their nt sequence conservation. The pcaK gene has no counterpart among known cat genes. The deduced amino-acid sequence of PcaK indicates that it may be a transmembrane protein associated with transport.

Acetyl Coenzyme A

On the informational content of overlapping genes in prokaryotic and eukaryotic viruses.

In genetic language a peculiar arrangement of biological information is provided by overlapping genes in which the same region of DNA can code for functionally unrelated messages. In this work, the informational content of overlapping genes belonging to prokaryotic and eukaryotic viruses was analyzed. Using information theory indices, we identified in the regions of overlap a first pattern, exhibiting a more uniform base composition and more severe constraints in base ordering with respect to the nonoverlapping regions. This pattern was found to be peculiar to coliphage, avian hepatitis B virus, human lentivirus, and plant luteovirus families. A second pattern, characterized by the occurrence of similar compositional constraints in both types of coding regions, was found to be limited to plant tymoviruses. At the level of codon usage, a low degree of correlation between overlapping and nonoverlapping coding regions characterized the first pattern, whereas a close link was found in tymoviruses, indicating a fine adaptation of the overlapping frame to the original codon choice of the virus. As a result of codon usage correlation analysis, deductions concerning the origin and evolution of several overlapping frames were also proposed. Comparison of amino acid composition revealed an increased frequency of amino acid residues with a high level of degeneracy (arginine, leucine, and serine) in the proteins encoded by overlapping genes; this peculiar feature of overlapping genes can be viewed as a way with which they may expand their coding ability and gain new, specialized functions.

Amino Acid Sequence

The nuclear genomes of African and American trypanosomes are strikingly different.

We have investigated the compositional distributions of exons and their different codon positions, as well as the codon usage and amino-acid (aa) composition of the nuclear genomes of the African and American trypanosomes Trypanosoma brucei and T. cruzi. Very large differences between the two species were found in all the properties investigated. The most striking differences concern the compositional distributions of third codon positions and the extremely large nucleotide divergence of third codon position for homologous genes encoding proteins that are highly conserved in their aa sequences. Moreover, if coding sequences from each species are divided into two groups according to the GC levels in third codon positions, very different codon usages and aa compositions are found. This indicates a compositional compartmentalization in both genomes which had previously been detected in T. brucei (and T. equiperdum) by compositional fractionation.

Animals

Evolution of chromosome bands: molecular ecology of noncoding DNA.

Giemsa dark bands, G-bands, are a derived chromatin character that evolved along the chromosomes of early chordates. They are facultative heterochromatin reflecting acquisition of a late replication mechanism to repress tissue-specific genes. Subsequently, R-bands, the primitive chromatin state, became directionally GC rich as evidenced by Q-banding of mammalian and avian chromosomes. Contrary to predictions from the neutral mutation theory, noncoding DNA is positionally constrained along the banding pattern with short interspersed repeats in R-bands and long interspersed repeats in G-bands. Chromosomes seem dynamically stable: the banding pattern and gene arrangement along several human and murine autosomes has remained constant for 100 million years, whereas much of the noncoding DNA, especially retroposons, has changed. Several coding sequence attributes and probably mutation rates are determined more by where a gene lives than by what it does. R-band exons in homeotherms but not G-band exons have directionally acquired GC-rich wobble bases and the corresponding codon usage: CpG islands in mammals are specific to R-band exons, exons not facultatively heterochromatinized, and are independent of the tissue expression pattern of the gene. The dynamic organization of noncoding DNA suggests a feedback loop that could influence codon usage and stabilize the chromosome's chromatin pattern: DNA sequences determine affinities of----proteins that together form----a chromatin that modulates----rate constants for DNA modification that determine----DNA sequences. Theories of hierarchical selection and molecular ecology show how selection can act on Darwinian units of noncoding DNA at the genome level thus creating positionally constrained DNA and contributing minimal genetic load at the individual level.

Base Sequence

Genetics of lactobacilli: plasmids and gene expression.

This paper reviews the present knowledge of the structure and properties of small (< 5 kb) plasmids present in Lactobacillus spp. The data show that plasmids from Lactobacillus spp., like many plasmids from other Gram-positive bacteria, display a modular organization and replicate by a mechanism of rolling circle replication. Structurally, plasmids from lactobacilli are closely related to plasmids from other Gram-positive bacteria. They contain elements (plus- and minus origin of replication, element(s) for control of plasmid replication, mobilization function) showing extensive similarity to analogous elements in plasmids from these other organisms. It is believed that lactobacilli have acquired such elements by intra- and/or intergenic transfer mechanisms. The first part of the review is concluded with a description of plasmid vectors with a Lactobacillus replicon and integrative vectors, including data concerning their structural and segregational stability. In the second part of this review we describe the progress that has been made during the last few years in identifying and characterizing elements that control expression of genetic information in lactobacilli. Based on the sequence of eleven identified and twenty presumed promoters, some preliminary conclusions can be drawn regarding the structure of Lactobacillus promoters. A typical Lactobacillus promoter shows significant similarity to promoters from E. coli and B. subtilis. An analysis of published sequences of seventy genes indicates that the region encompassing the translation start codon AUG also shows extensive similarity to that of E. coli and B. subtilis. Codon usage of Lactobacillus genes is not random and shows interspecies as well as intraspecies heterogeneity. Interspecies differences may, in part, be explained by differences in G+C content of different lactobacilli. Differences in gene expression levels can, to a large extent, account for intraspecies differences of codon usage bias. Finally, we review the knowledge that has become available concerning protein secretion and heterologous gene expression in lactobacilli. This part is concluded with a compilation of data on the expression in Lactobacillus of heterologous genes under the control of their own promoter or under control of a Lactobacillus promoter.

Bacillus subtilis