Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Codon usage and gene expression level in Dictyostelium discoideum: highly expressed genes do 'prefer' optimal codons.

Codon usage patterns in the slime mould Dictyostelium discoideum have been re-examined (a total of 58 genes have been analysed). Considering the extreme A + T-richness of this genome (G + C = 22%), there is a surprising degree of codon usage variation among genes. For example, G + C content at silent sites varies from less than 10% to greater than 30%. It was previously suggested [Warrick, H.M. and Spudich, J.A. (1988) Nucleic Acids Res. 16: 6617-6635] that highly expressed genes contain fewer 'optimal' codons than genes expressed at lower levels. However, it appears that the optimal codons were misidentified. Multivariate statistical analysis shows that the greatest variation among genes is in relative usage of a particular subset of codons (about one per amino acid), many of which are C-ending. We have identified these as optimal codons, since (i) their frequency is positively correlated with gene expression level, and (ii) there is a strong mutation bias in this genome towards A and T nucleotides. Thus, codon usage in D. discoideum can be explained by a balance between the forces of mutational bias and translational selection.

Codon

Evaluation of foreign gene codon optimization in yeast: expression of a mouse IG kappa chain.

We have optimized the codons in an immunoglobulin kappa chain gene to those preferred in the yeast Saccharomyces cerevisiae. The mutant and wild type kappa chain genes were each fused with a synthetic invertase signal peptide that also contained only yeast-preferred codons, and expressed in the F762 yeast strain. The use of yeast-preferred codons resulted in a more than 5-fold increase in the rate of synthesis and at least a 50-fold increase in the steady state level of protein.

Animals

Natural Selection Drives Codon Usage Bias in the Mitochondrial Genome of Ligula intestinalis (Linnaeus, 1758) Gmelin, 1790 (Cestoda: Diphyllobothriidea): Insights from Comparative Genomics and Optimal Codon Identification.

Codon usage bias (CUB) is a useful indicator of evolutionary forces shaping mitochondrial genomes. Codon usage bias in mitochondrial genomes of Diphyllobothriidae and especially in Ligula intestinalis was characterized. The roles of natural selection and mutation pressure in framing this bias were evaluated on the basis of 12 protein-coding genes in Diphyllobothriidae. The complete mitogenome (13,725 bp) of L. intestinalis comprises 12 protein-coding genes (PCGs), 22 tRNAs, and two rRNAs, all positioned on the heavy strand, and contains an overall AT content of 66.15%. The mean CAI (0.176), CBI (-0.105), and ENC (45.33) and an evident preference for U-ending codons observed in all examined genes indicate weak CUB. Neutrality, ENC, and PR2 plots consistently demonstrate that natural selection is the predominant force driving CUB and contributes approximately 56% in L. intestinalis and 83% in other Diphyllobothriidea species, with mutation pressure playing a secondary role. Phylogenetic reconstruction supported the monophyly of Diphyllobothriidea, confirmed the paraphyly of Diphyllobothrium as traditionally defined, and placed Ligula and Digramma as sister taxa. These findings clarify the evolutionary constraints governing codon usage in cestode mitogenomes and provide practical resources for codon optimization in heterologous gene expression and genetic studies of this economically important parasite.

Diphyllobothriidea

Human hemoglobin expression in Escherichia coli: importance of optimal codon usage.

The overexpression of a nonfusion product of human beta-globin in Escherichia coli from its cDNA sequence has been accomplished for the first time. Expression of beta-globin from its native cDNA required the use of the strong bacteriophage T7 promoter. In this system, beta-globin accumulated to approximately 10% of total E. coli proteins. alpha-Globin was not expressed in the T7 system using the native cDNA. For the expression of alpha-globin, synthetic genes containing optimal E. coli codons were constructed. Neither synthetic alpha- nor beta-globin gene alone was expressed from the lac or tac promoter. Globin expression was achieved when the two synthetic alpha- and beta-globin genes were combined as an operon downstream of the lac promoter. The two proteins combined intracellularly with endogenous heme, which was concomitantly overproduced to yield tetrameric hemoglobin as roughly 5-10% of total E. coli protein. Cloning the alpha- and beta-globin cDNAs in a construct identical with the lac promoter did not yield globin production, establishing the requirement for optimal codon usage. The recombinant beta-globin from the T7 expression system was purified and reconstituted in vitro with heme and native alpha chains. N-terminal analyses showed that the beta-globin produced in the T7 system and the tetrameric hemoglobin produced from the synthetic genes contained an additional beta 1 methionine residue. Two additional mutants, beta 1 Val----Met and beta 1 Val----Ala were produced using the T7 system. Functional and structural properties of the purified hemoglobins will be discussed in the following papers.

Amino Acid Sequence

Transgene sequence codon optimization and composition determines replication competence of self-amplifying RNA.

Self-amplifying RNA (saRNA) is an emerging RNA therapeutic modality that can facilitate higher magnitude and more durable protein expression at substantially lower doses than nonreplicating mRNA. Unlike conventional messenger RNA (mRNA), alphavirus-derived saRNA must support a replicase-driven RNA amplification step in addition to translation, raising the possibility that transgene coding sequences impose sequence-level constraints on replication. Here, saRNA replication was found to be dependent on the codon composition of the transgene; multiple therapeutic transgenes were replication defective despite an intact Venezuelan Equine Encephalitis Virus (VEEV)-derived saRNA backbone. Replication defects were rescued by synonymous codon re-optimization of the same transgenes, indicating that nucleotide-level features of the coding sequence, rather than the encoded protein, govern replication competence. Comparative compositional analyses identified a distinct signature associated with productive replication, characterized by elevated GC (>53%) and GC3 (>63%) content, higher codon adaptation to human (>0.75), and reduced UpA (<43/kb) and UpU (<41/kb) dinucleotide density. Moreover, deliberate compositional perturbation of an otherwise replication-competent transgene shifted these features and abolished replication, supporting a causal and combinatorial role for sequence composition in defining saRNA replication outcome. These findings define an underappreciated constraint in saRNA therapeutics and motivate saRNA-specific payload design frameworks that incorporate alphavirus-associated compositional biases during transgene sequence optimization.

Codon

Codon usage in Aspergillus nidulans.

Synonymous codon usage in genes from the ascomycete (filamentous) fungus Aspergillus nidulans has been investigated. A total of 45 gene sequences has been analysed. Multivariate statistical analysis has been used to identify a single major trend among genes. At one end of this trend are lowly expressed genes, whereas at the other extreme lie genes known or expected to be highly expressed. The major trend is from nearly random codon usage (in the lowly expressed genes) to codon usage that is highly biased towards a set of 19-20 "optimal" codons. The G + C content of the A. nidulans genome is close to 50%, indicating little overall mutational bias, and so the codon usage of lowly expressed genes is as expected in the absence of selection pressure at silent sites. Most of the optimal codons are C- or G- ending, making highly expressed genes more G + C-rich at silent sites.

Aspergillus nidulans

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO

Codon usage of human DNA viruses and its similarity to certain host genes.

Codon usages of DNA viruses had previously been shown to associate with their genome size. Codon usage of various human DNA viruses was compared to those of human genes to further understand viral codon usage and its roles in viral-host interaction. Codon usage bias in both large and small genome human DNA viruses was dominantly driven by translation selection. Non-optimal codon usage in small DNA viruses showed similarity to cell cycle-related genes, whereas codon usage of large DNA viruses was more diverse, herpesviruses showed more heterogeneity than human adenoviruses, while poxviruses showed a clear bimodal pattern. Some of the large DNA viruses such as herpes simplex and molluscum contagiosum viruses showed more optimal codon usage. Enrichment analysis identified some groups of human genes with similar codon usage to each group of these viruses. These host genes with similarity in codon usages to those of viruses may be efficiently expressed in infected cells and involved in their life cycle, pathogenesis and/or immune evasion.

Humans

Synonymous codon preferences in bacteriophage T4: a distinctive use of transfer RNAs from T4 and from its host Escherichia coli.

Codon usage data of bacteriophage T4 genes were compiled and synonymous codon preferences were investigated in comparison with tRNA availabilities in an infected cell. Since the genome of T4 is highly AT rich and its codon usage pattern is significantly different from that of its host Escherichia coli, certain codons of T4 genes need to be translated by appropriate host transfer RNAs present in minor amounts. To avoid this predicament, T4 phage seems to direct the synthesis of its own tRNA molecules and these phage tRNAs are suggested to supplement the host tRNA population with isoacceptors that are normally present in minor amounts. A positive correlation was found in that the frequency of E. coli optimal codons in T4 genes increases as the number of protein monomers per phage particle increases. A negative correlation was also found between the number of protein monomers per phage and the frequency of "T4 optimal codons", which are defined as those codons that are efficiently recognized by T4 tRNAs. From these observations it was proposed that tRNAs from the host are predominantly used for translation of highly expressed T4 genes while tRNAs from T4 tend to be used for translation of weakly expressed T4 genes. This distinctive tRNA-usage in T4 may be an optimization of translational efficiency, and an adjustment of T4-encoded tRNAs to the synonymous codon preferences, which are largely influenced by the high genomic AT-content, would have occurred during evolution.

Bacteriophage T4

Efficient synthesis of secreted murine interleukin-2 by Saccharomyces cerevisiae: influence of 3'-untranslated regions and codon usage.

Several expression vectors were compared which directed the synthesis of secreted murine interleukin-2 (mIL2) in the culture medium of Saccharomyces cerevisiae. We used the prepro-sequence of the alpha 1 mating-factor precursor as a secretion signal in S. cerevisiae in combination with different promoters. The yield of mature mIL2 was significantly improved by deleting the major part of the 3'-untranslated region (UTR). In Northern-blotting experiments we showed that a destabilizing sequence present in the 3' UTR might be responsible for rapid degradation of the mIL2 mRNA. The highest expression (about 10 micrograms/ml) was obtained under control of the GAL1 promoter in an S. cerevisiae strain where the regulatory GAL4 gene was overexpressed. No difference in expression level was observed in a construct wherein twelve consecutive codons were replaced by optimal codons for S. cerevisiae.

Animals

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569&#x2009;bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A&#x2009;+&#x2009;T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU >&#x2009;1) with a pronounced preference for A/C-ending codons. The &#x394;RSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean&#x2009;=&#x2009;1.209&#x2009;&#xb1;&#x2009;0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Comprehensive analysis of synonymous codon usage bias and evolutionary dynamics in the chloroplast genomes of eight Coptis species.

Coptis is a medically important genus renowned for producing valuable isoquinoline alkaloids. Although its chloroplast genomes encode key components for photosynthesis and plastid gene expression, the evolutionary constraints acting on their coding sequences and synonymous codon usage remain poorly resolved. Here, we combined a transparent taxon-level sampling strategy with comparative analyses of chloroplast CDSs from eight Coptis taxa. We quantified nucleotide composition, relative synonymous codon usage, effective number of codons, neutrality and PR2 patterns, and correspondence analysis, and then integrated these results with a core-CDS distance analysis and gene-wise pairwise dN/dS estimates. The chloroplast genomes showed a conserved AT-rich composition, especially at the third codon position (GC3 approximately 30.3-30.8%), with a consistent GC1&#x2009;>&#x2009;GC2&#x2009;>&#x2009;GC3 trend. Thirty preferred codons were detected, 28 ending in A/T, and eleven optimal codons were shared across the genus. The core-CDS distance analysis recovered a close relationship between C. chinensis and C. chinensis var. brevisepala, whereas most coding genes showed dN/dS values below one, consistent with pervasive purifying constraint. Across 48 consistently filtered CDSs, GC3s was negatively associated with mean dN (Spearman rho = -0.404, P&#x2009;=&#x2009;0.00439) and CAI was positively associated with mean dN (rho&#x2009;=&#x2009;0.303, P&#x2009;=&#x2009;0.0361), whereas the remaining associations were not significant (all P&#x2009;>&#x2009;=&#x2009;0.0972). These results extend codon-usage analysis by linking synonymous-site composition to coding-sequence evolution within Coptis, while providing a hypothesis-generating resource for future plastid engineering studies.

Genome, Chloroplast

High-level expression of staphylococcal nuclease R gene in Escherichia coli.

Staphylococcal nuclease R, an analogue of nuclease A, was overproduced under the transcriptional control of the bacteriophage lambda PRPL promoters regulated by temperature sensitive repressors. The expression level reached 200-300 mg l-1 and showed little host dependence in different strains. The investigations of the recombinant nuclease R have revealed that the amino terminal formyl methionine residue of the nuclease is precisely processed, the protein consists of 155 amino acid residues. The experiment shows that the pBV221-DH5 alpha is a quite suitable vector-host system for high-level expression and precise processing of heterologous genes in Escherichia coli. The comparative studies between the codons used in the staphylococcal nuclease R gene and the optimal codon usage in E. coli indicate that high level expression of heterologous genes in E. coli may not always require a high degree of codon usage bias.

Amino Acid Sequence

Novel third-letter bias in Escherichia coli codons revealed by rigorous treatment of coding constraints.

A novel bias in codon third-letter usage was found in Escherichia coli genes with low fractions of "optimal codons", by comparing intact sequences with control random sequences. Third-letter usage has been found to be biased according to preference in codon usage and to doublet preference from the following first letter. The present study examines third-letter usage in the context of the nucleotide sequence when these preferences are considered. In order to exclude any influence by these factors, the random sequences were generated such that the amino acid sequence, codon usage, and the doublet frequency in each gene were all preserved. Comparison of intact sequences with these randomly generated sequences reveals that third letters of codons show a strong preference for the purine/pyrimidine pattern of the next codons: purine (R) is preferred to pyrimidine (Y) at the third site when followed by an R-Y-R codon, and pyrimidine is preferred when followed by an R-R-Y, an R-Y-Y or a Y-R-Y codon. This bias is probably related to interactions of tRNA molecules in the ribosome.

Amino Acids

Stable structure of thermophilic proton ATPase beta subunit.

F1-ATPase is the major enzyme for ATP synthesis in mitochondria, chloroplasts, and bacterial plasma membranes. F1-ATPase obtained from thermophilic bacterium PS3 (TF1) is the only ATPase which can be reconstituted from its primary structure. Its beta subunit constitutes the catalytic site, and is capable of forming hybrid F1's with E. coli alpha and gamma subunits. Since the stability of TF1 resides in its primary structure, we cloned a gene coding for TF1, and the primary structure of the beta subunit was deduced from the nucleotide sequence of the gene to compare the sequence with those of beta's of three major categories of F1's; prokaryotic membranes, chloroplasts, and mitochondria. The following results were obtained. Homology: The primary structure of the TF1 beta subunit (473 residues, Mr = 51,995.6) showed 89.3% homology with 270 residues which are identical in the beta subunits from human mitochondria, spinach chloroplasts, and E. coli. It contained regions homologous to several nucleotide-binding proteins. Secondary structure: The deduced alpha-helical (30.1%) and beta-sheet (22.3%) contents were consistent with those determined from the circular dichroism spectra. Residues forming reverse turns (Gly and Pro) were highly conserved among the F1 beta subunits. Substituted residues and stability of TF1: We compared the amino acid sequence of the TF1 beta subunit with those of the other F1 beta subunits mentioned above. The observed substitutions in the thermophilic subunit increased its propensities to form secondary structures, and its external polarity to form tertiary structure. Codon usage: The codon usage of the TF1 beta gene was found to be unique. The changes in codons that achieved these amino acid substitutions were much larger than those caused by minimal mutations, and the third letters of the optimal codons were either guanine or cytosine, except in codons for Gln, Lys, and Glu.

Amino Acid Sequence

Unconventional codon usage bias mediates mRNA translational dynamics in macrophages.

Macrophages require rapid and tightly controlled regulatory mechanisms to respond to environmental disruptions. While transcriptional regulation has been well characterized, the mechanisms underlying translational control in macrophages remain poorly understood. Here, we investigated the dynamics of mRNA translation in mouse macrophages during acute, intermediate, and prolonged LPS exposure. Our results reveal clear phase-specific translational regulation during macrophage polarization, which initially increases the synthesis of inflammatory mediators and cytokines, while simultaneously suppressing the expression of cell cycle-related genes. Mechanistically, we observed pervasive upstream translation in the 5' UTRs of cell cycle-related mRNAs, which contributes to cell cycle arrest during the early phase of inflammatory response. Notably, we identified a unique codon preference toward A/U in the third position of codons in macrophages, which contrasts with the G/C preference commonly observed in other tissues. AU codon preference increases the stability and translation efficiency of cell cycle-related mRNAs, promoting cell cycle restoration after extended LPS exposure. These findings reveal that uORF translation and codon usage bias are critical components of translational regulation during macrophage polarization, highlighting a potential therapeutic intervention for modulating immune activation via macrophage-specific codon optimization.

Animals

CasY7: An optimized Cas12i system for enhanced genome editing in monocot crops.

The CRISPR-Cas12 family nucleases, particularly the Cas12i subtypes, are considered promising alternatives to Cas9 for genome editing in plants. We previously developed a new Cas12i variant, CasY7, which has been successfully applied in clinical trials; its performance in plants remains to be investigated. Initial testing in stable transgenic maize and rice showed that the codon-optimized CasY7 (pCasY7e1) achieved average editing efficiencies of 58.7% and 62.3% across five target sites, respectively, outperforming the typical Cpf1 (pCpf1) control that targets the same sites. To further enhance activity, we fused T5 exonuclease to CasY7 (pCasY7e2), which shifted mutation profiles toward larger deletions, and subsequently integrated an MS2 aptamer into the crRNA scaffold (pCasY7e3). The optimized pCasY7e3 system increased editing efficiencies to 87.7% in maize and 82.9% in rice-approximately 2.7-fold higher than pCpf1. We further demonstrated multiplexed editing in maize, generating biallelic dwarf mutants, and validated functionality in hexaploid wheat with editing efficiencies up to 58.8%. Overall, our comprehensive validation across 942 transgenic plants confirmed robust editing in maize, rice, and wheat, establishing CasY7 as a high-efficiency addition to the CRISPR toolkit.

Zea mays

Evolution of codon usage patterns: the extent and nature of divergence between Candida albicans and Saccharomyces cerevisiae.

Codon usage in a sample of 28 genes from the pathogenic yeast Candida albicans has been analysed using multivariate statistical analysis. A major trend among genes, correlated with gene expression level, was identified. We have focussed on the extent and nature of divergence between C.albicans and the closely related yeast Saccharomyces cerevisiae. It was recently suggested that significant differences exist between the subsets of preferred codons in these two species [Brown et al. (1991) Nucleic Acids Res. 19, 4293]. Overall, the genes of C.albicans are more A + T-rich, reflecting the lower genomic G + C content of that species, and presumably resulting from a different pattern of mutational bias. However, in both species highly expressed genes preferentially use the same subset of 'optimal' codons. A suggestion that the low frequency of NCG codons in both yeast species results from selection against the presence of codons that are potentially highly mutable is discounted. Codon usage in C.albicans, as in other unicellular species, can be interpreted as the result of a balance between the processes of mutational bias and translational selection. Codon usage in two related Candida species, C.maltosa and C.tropicalis, is briefly discussed.

Biological Evolution