Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Codon usage in the vertebrate hemoglobins and its implications.

A study of codon usage in vertebrate hemoglobins revealed an evolutionary trend toward elevated numbers of CpG codon boundary pairs in mammalian hemoglobin alpha genes. Selection for CpG codon boundaries countering the generally observed CpG suppression is strongly suggested by these data. These observations parallel recently published experimental results that indicate that constitutive expression of the human alpha-globin gene appears to be determined by regulatory information encoded within the structural gene. The possibility is raised that, in the absence of selection, CpG decay can be used to date the evolutionary origin of a mammalian alpha pseudogene from its active alpha gene.

Animals

Nucleotide sequence of the Caulobacter crescentus flaF and flbT genes and an analysis of codon usage in organisms with G + C-rich genomes.

The Caulobacter crescentus flaFG region encodes trans-acting, regulatory factors that modulate flagellin synthesis during flagellum biogenesis. In this study, sequence analysis and experiments utilizing a promoterless cat gene demonstrated that the flaF and flbT genes have overlapping transcripts with the same orientation. In addition, the 5' ends of the flgL and flbA genes were located. A sequence resembling an Rho-factor-independent terminator was found in the 3' region of the flaF gene. This region was uniquely A + T-rich and the encoded mRNA contained an inverted repeat sequence which could form a stable stem-loop structure followed by nine U-residues. The codon usage of C. crescentus genes was examined and indicated a preference for specific codons from each of the synonymous codon groups. Furthermore, comparison to the codon usage of other organisms with G + C-rich genomes indicated a strong preference for the same codons preferred by C. crescentus.

Amino Acid Sequence

Essential factors determining codon usage in ubiquitin genes.

Ubiquitin is ubiquitous in all eukaryotes and its amino acid sequence shows extreme conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences coding for 13 ubiquitin genes from 11 species reported so far have been compiled and analyzed. The G + C content of codon third base reveals a positive linear correlation with the genome G + C content of the corresponding species. The slope strongly suggests that the overall G + C content of codons of polyubiquitin genes clearly reflects the genome G + C content by AT/GC substitutions at the codon third position. The G + C content of ubiquitin codon third base also shows a positive linear correlation with the overall G + C content of coding regions of compiled genes, indicating the codon choices among synonymous codons reflect the average codon usage pattern of corresponding species. On the other hand, the monoubiquitin gene, which is different from the polyubiquitin gene in gene organization, gene expression, and function of the encoding protein, shows a different codon usage pattern compared with that of the polyubiquitin gene. From comparisons of the levels of synonymous substitutions among ubiquitin repeats and the homology of the amino acid sequence of the tail of monomeric ubiquitin genes, we propose that the molecular evolution of ubiquitin genes occurred as follows: Plural primitive ubiquitin sequences were dispersed on genome in ancestral eukaryotes. Some of them situated in a particular environment fused with the tail sequence to produce monomeric ubiquitin genes that were maintained across species. After divergence of species, polyubiquitin genes were formed by duplication of the other primitive ubiquitin sequences on different chromosomes. Differences in the environments in which ubiquitin genes are embedded reflect the differences in codon choice and in gene expression pattern between poly- and monomeric ubiquitin genes.

Amino Acid Sequence

Codon usage in histone gene families of higher eukaryotes reflects functional rather than phylogenetic relationships.

The nucleic acid sequences coding for 23 H3 histone genes from a variety of species have been analyzed using a computer assisted alignment and analysis program. Although these histones are highly conserved within and between highly divergent species, they represent various classes of histones whose patterns of expression are distinctively regulated. Surprisingly, in dendrograms derived from these comparisons, H3 sequences cluster according to their modes of regulation rather than phylogenetically. These clusters are generated from highly distinctive patterns of codon usage within the functional gene classes. We suggest that one factor involved in specifying the differing codon usage patterns between functional classes is a difference in requirements for rapid translation of mRNA. In addition, the data presented here, together with structural and sequence information, suggest a heterodox evolutionary model in which genes related to the intron-bearing, basally expressed H3.3 vertebrate genes are the ancestors of the intronless H3.1 class of genes of higher eukaryotes. The H3.1 class must have arisen, therefore, following duplication of a primitive H3.3 gene, but prior to the plant-animal divergence. Implications of the data presented are discussed with regard to functional and evolutionary relationships.

Amino Acid Sequence

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence

Organization and codon usage of the streptomycin operon in Micrococcus luteus, a bacterium with a high genomic G + C content.

The DNA sequence of the Micrococcus luteus str operon, which includes genes for ribosomal proteins S12 (str or rpsL) and S7 (rpsG) and elongation factors (EF) G (fus) and Tu (tuf), has been determined and compared with the corresponding sequence of Escherichia coli to estimate the effect of high genomic G + C content (74%) of M. luteus on the codon usage pattern. The gene organization in this operon and the deduced amino acid sequence of each corresponding protein are well conserved between the two species. The mean G + C content of the M. luteus str operon is 67%, which is much higher than that of E. coli (51%). The codon usage pattern of M. luteus is very different from that of E. coli and extremely biased to the use of G and C in silent positions. About 95% (1,309 of 1,382) of codons have G or C at the third position. Codon GUG is used for initiation of S12, EF-G, and EF-Tu, and AUG is used only in S7, whereas GUG initiates only one of the EF-Tu's in E. coli. UGA is the predominant termination codon in M. luteus, in contrast to UAA in E. coli.

Amino Acid Sequence

Sequence analysis of the cDNA encoding human liver glycogen phosphorylase reveals tissue-specific codon usage.

We have cloned the cDNA encoding glycogen phosphorylase (1,4-alpha-D-glucan:orthophosphate alpha-D-glucosyl-transferase, EC 2.4.1.1) from human liver. Blot-hybridization analysis using a large fragment of the cDNA to probe mRNA from rabbit brain, muscle, and liver tissues shows preferential hybridization to liver RNA. Determination of the entire nucleotide sequence of the liver message has allowed a comparison with the previously determined rabbit muscle phosphorylase sequence. Despite an amino acid identity of 80%, the two cDNAs exhibit a remarkable divergence in G+C content. In the muscle phosphorylase sequence, 86% of the nucleotides at the third codon position are either deoxyguanosine or deoxycytidine residues, while in the liver homolog the figure is only 60%, resulting in a strikingly different pattern of codon usage throughout most of the sequence. The liver phosphorylase cDNA appears to represent an evolutionary mosaic; the segment encoding the N-terminal 80 amino acids contains greater than 90% G+C at the third codon position. A survey of other published mammalian cDNA sequences reveals that the data for liver and muscle phosphorylases reflects a bias in codon usage patterns in liver and muscle coding sequences in general.

Animals

Codon usage pattern in alpha 2(I) chain domain of chicken type I collagen and its implications for the secondary structure of the mRNA and the synthesis pauses of the collagen.

A stability map of local secondary structure of the mRNA of the triple-helical alpha 2(I) chain domain of chicken type I collagen was obtained by plotting the free energy of the optimal secondary structure of a local segment in mRNA against the segment position along a base sequence of the mRNA. It was found that the positions of the minima of free energy in the plot coincide with the positions where synthesis pauses of the alpha-chain polypeptides of the corresponding sizes translated from the mRNA have been reported to occur (1). The codon usage pattern of each of the three major amino acids of the alpha-chain domain of the collagen, Gly, Pro and Ala, fluctuates considerably along the base sequence segments of the mRNA and a deviation of the pattern from that of the average of the whole alpha 2(I) chain domain mRNA, particularly for Gly codons, leads to a loss of the stability of the local secondary structure of the mRNA. The results suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA of the chicken collagen alpha 2(I) chain domain which leads to a nonuniform polypeptide elongation pattern.

Animals

Evidence for selective evolution in codon usage in conserved amino acid segments of human alphaherpesvirus proteins.

The genomes of human viruses herpes simplex 1 (HSV1) and varicella zoster (VZV), although similar in biology, largely concordant in gene order, and identical in many amino acid segments, differ widely in their genomic G + C (abbreviated S) content, which is high in HSV1 (68%) and low in VZV (46%). This paper analyzes several striking codon usage contrasts. The S difference in coding regions is dramatically large in codon site 3, S3, about 42%. The large difference in S3 is maintained at the same level in a subset of closely similar genes and even in corresponding identical amino acid blocks. A similar difference in S levels in silent site 1 (S1) is found in leucine and arginine. The difference in S3 levels occurs in every gene and in every multicodon amino acid form. The S difference also exists in amino acid usage, with HSV1 using significantly more codon types SSN, while VZV uses more codon types WWN (where W stands for A or T). The nonoverlapping and narrow histograms of S3 gene frequencies in both viruses suggest that the difference has arisen and been maintained by a process of selective rather than nonselective effects. This is in sharp contrast to the relatively large variance seen for highly similar genes in the human versus yeast analysis. Interpretations and hypotheses to explain the HSV1 vs VZV codon usage disparity relate to virus-host interactions, to the role of viral genes in DNA metabolism, to availability of molecular resources (molecular Gause exclusion principle), and to differences in genomic structure.

Amino Acids

Contrasts in codon usage of latent versus productive genes of Epstein-Barr virus: data and hypotheses.

Epstein-Barr virus (EBV) has two different modes of existence: latent and productive. There are eight known genes expressed during latency (and hardly at all during the productive phase) and about 70 other ("productive") genes. It is shown that the EBV genes known to be expressed during latency display codon usage strikingly different from that of genes that are expressed during lytic growth. In particular, the percentage of S3 (G or C in codon site 3) is persistently lower (about 20%) in all latent genes than in nonlatent genes. Moreover, S3 is lower in each multicodon amino acid form. Also, the percentage of S in silent codon sites 1 of leucine and arginine is lower in latent than in nonlatent genes. The largest absolute differences in amino acid usage between latent and nonlatent genes emphasize codon types SSN and WWN (W means nucleotide A or T and N is any nucleotide). Two principal explanations to account for the EBV latent versus productive gene codon disparity are proposed. Latent genes have codon usage substantially different from that of host cell genes to minimize the deleterious consequences to the host of viral gene expression during latency. (Productive genes are not so constrained.) It is also proposed that the latency genes of EBV were acquired recently by the viral genome. Evidence and arguments for these proposals are presented.

Amino Acid Sequence

Translation of the downstream ORF from bicistronic mRNAs by human cells: Impact of codon usage and splicing in the upstream ORF.

Biochemistry textbooks describe eukaryotic mRNAs as monocistronic. However, increasing evidence reveals the widespread presence and translation of upstream open reading frames preceding the "main" ORF. DNA and RNA viruses infecting eukaryotes often produce polycistronic mRNAs and viruses have evolved multiple ways of manipulating the host's translation machinery. Here, we introduce an experimental model to study gene expression regulation from virus-like bicistronic mRNAs in human cells. The model consists of a short upstream ORF and a reporter downstream ORF encoding a fluorescent protein. We have engineered synonymous variants of the upstream ORF to explore large parameter space, including codon usage preferences, mRNA folding features, and splicing propensity. We show that human translation machinery can translate the downstream ORF from bicistronic mRNAs, albeit reporter protein levels are thousand times lower than those from the upstream ORF. Furthermore, synonymous recoding of the upstream ORF exclusively during elongation significantly influences its own translation efficiency, reveals cryptic splice signals, and modulates the probability of downstream ORF translation. Our results are consistent with a leaky scanning mechanism facilitating downstream ORF translation from bicistronic mRNAs in human cells, offering new insights into the role of upstream ORFs in translation regulation.

Humans

Nucleotide sequence of a macronuclear DNA molecule coding for alpha-tubulin from the ciliate Stylonychia lemnae. Special codon usage: TAA is not a translation termination codon.

The gene-sized macronuclear DNA of the hypotrichous ciliate Stylonychia lemnae contains two size classes of DNA molecules (1.85 and 1.73 kbp) coding for alpha-tubulin. Each macronucleus contains about 55000 copies of the 1.85 kbp molecules and about 17000 copies of the 1.73 kbp DNA molecules. Five macronuclear molecules of these sequences were cloned and sequenced, one, from the 1.85 kbp size class in its entirety. The 5 sequences fell into two classes suggesting that Stylonychia lemnae contains at least two different alpha-tubulin genes. All 5 clones show the codon TAA in the same nucleotide positions of the coding region. In this position the TAA codon cannot function as a translational stop codon and we suggest that this codon codes for the amino acid glutamine. The nucleotide sequence of the coding region as well as the encoded amino acid sequence is highly conserved compared to alpha-tubulin genes from vertebrates. The noncoding regions show several putative transcription-regulatory sequences as well as sequences presumably functioning as replication origins.

Base Sequence

Unconventional codon usage bias mediates mRNA translational dynamics in macrophages.

Macrophages require rapid and tightly controlled regulatory mechanisms to respond to environmental disruptions. While transcriptional regulation has been well characterized, the mechanisms underlying translational control in macrophages remain poorly understood. Here, we investigated the dynamics of mRNA translation in mouse macrophages during acute, intermediate, and prolonged LPS exposure. Our results reveal clear phase-specific translational regulation during macrophage polarization, which initially increases the synthesis of inflammatory mediators and cytokines, while simultaneously suppressing the expression of cell cycle-related genes. Mechanistically, we observed pervasive upstream translation in the 5' UTRs of cell cycle-related mRNAs, which contributes to cell cycle arrest during the early phase of inflammatory response. Notably, we identified a unique codon preference toward A/U in the third position of codons in macrophages, which contrasts with the G/C preference commonly observed in other tissues. AU codon preference increases the stability and translation efficiency of cell cycle-related mRNAs, promoting cell cycle restoration after extended LPS exposure. These findings reveal that uORF translation and codon usage bias are critical components of translational regulation during macrophage polarization, highlighting a potential therapeutic intervention for modulating immune activation via macrophage-specific codon optimization.

Animals

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Design, synthesis and expression of a human interleukin-2 gene incorporating the codon usage bias found in highly expressed Escherichia coli genes.

A synthetic gene encoding human interleukin-2 (IL-2) was designed such that the codon usage bias resembled that found in highly expressed Escherichia coli genes. The percentage of preferred codons was increased from 43% in the native cDNA sequence to 85% in the synthetic sequence. The cDNA and synthetic IL-2 genes were placed under the control of the trc promoter and expressed in E. coli JM101. While Northern blot analysis of IL-2 mRNA from each genetic construct demonstrated equivalent message half-lives, immunoblot and bioactivity analyses showed the synthetic gene to direct the synthesis of up to 16 times more IL-2 than the native cDNA sequence.

Amino Acid Sequence

Codon usage determines translation rate in Escherichia coli.

We wish to determine whether differences in translation rate are correlated with differences in codon usage or with differences in mRNA secondary structure. We therefore inserted a small DNA fragment in the lacZ gene either directly or flanked by a few frame-shifting bases, leaving the reading frame of the lacZ gene unchanged. The fragment was chosen to have "infrequent" codons in one reading frame and "common" codons in the other. The insert in these constructs does not seem to give mRNAs that are able to form extensive secondary structures. The translation time for these modified lacZ mRNAs was measured with a reproducibility better than plus or minus one second. We found that the mRNA with infrequent codons inserted has an approximately three-seconds longer translation time than the one with common codons. In another set of experiments we constructed two almost identical lacZ genes in which the lacZ mRNAs have the potential to generate stem structures with stabilities of about -75 kcal/mol. In this way we could investigate the influence of mRNA structure on translation rate. This type of modified gene was generated in two reading frames with either common or infrequent codons similar to our first experiments. We find that the yield of protein from these mRNAs is reduced, probably due to the action in vivo of an RNase. Nevertheless, the data do not indicate that there is any effect of mRNA secondary structure on translation rate. In contrast, our data persuade us that there is a difference in translation rate between infrequent codons and common codons that is of the order of sixfold.

Bacterial Proteins

Nucleotide sequence and codon usage of the elongation factor Tu(EF-Tu) gene from Mycoplasma pneumoniae.

The Mycoplasma pneumoniae tuf gene, encoding the elongation factor protein Tu, was cloned and sequenced. The nucleotide sequence of the mycoplasmal gene showed about 60% homology to the sequences of tuf genes of other prokaryotes, yeast mitochondria and Euglena gracilis chloroplasts, and about 75% similarity was found when comparing the deduced amino acid sequences of the various Tu proteins. The relatively low G + C content (40%) of the M. pneumoniae DNA was reflected in a low G + C content (44.6%) of the tuf gene, and in a preferential use of adenine and uracil at the third position of codons, yet codon usage analysis revealed the presence of almost all of the codons of the genetic code in the mycoplasmal gene. Southern blot hybridization of digested DNAs of 11 Mollicutes species with the entire M. pneumoniae tuf gene and with its 5' part suggested the presence of one copy only of this gene in the representative species of the Mollicutes. In this respect, the Mollicutes resemble Gram-positive bacteria and differ from the Gram-negative bacteria, which carry two copies of the tuf gene.

Amino Acid Sequence