Search PubMedSearch

SEARCH · Search PubMed

Results for “codon optimization”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Binding of the bacteriophage T4 regA protein to mRNA targets: an initiator AUG is required.

Bacteriophage T4 regA protein translationally represses the synthesis of a subset of early phage-induced proteins. The protein binds to the translation initiation site of at least two mRNAs and prevents formation of the initiation complex. We show here that the protein binds to the translation initiation sites of other regA-sensitive mRNAs. Analysis of mRNA binding by filtration and nuclease protection assays shows that AUG is necessary but not sufficient for specific binding of regA protein to its mRNA targets. Anticipating the need for large quantities of regA protein for structural studies to further define the regA protein-RNA ligand interaction, we also report cloning the regA gene into a T4 overexpression system. The expression of regA protein in uninfected E. coli is lethal, so in our system regA driven by a strong T7 promoter is sequestered in a T4 phage until 'induction' by phage infection is desired. We have replaced the regA sensitive wild-type ribosome binding site with a strong insensitive ribosome binding site at an optimal distance from the regA initiation codon for maximizing expression. We have obtained large amounts of regA protein.

Base Sequence

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

CUG as a mutant start codon for cat-86 and xylE in Bacillus subtilis.

The cat-86 gene specifies chloramphenicol acetyltransferase (CAT). The cat-86 start codon is UUG, although related genes have AUG as the start codon. Changing the start codon to AUG increased expression of cat-86 by 36% in Bacillus subtilis. Changing the start codon to GUG and CUG decreased expression to 65% and 30%, respectively, of the level obtained when AUG was the start codon. CUG has not been previously shown to function as a start codon in B. subtilis. N-terminal sequencing of purified CAT protein specified by the CUG mutant, revealed that CUG was indeed the start codon and specified methionine. The gene xylE, which specifies catechol 2,3-dioxygenase, has AUG as its start codon. Changing the start codon for xylE to CUG decreased expression by 98%. However, when the ribosome-binding site sequence for xylE was optimized and the spacing between it and the start codon was increased to 8 nucleotides, xylE activity increased to 13% of the activity observed for AUG. CUG did not function efficiently as a start codon for cat-86 in Escherichia coli. These data suggest conditions under which CUG can function, with modest efficiency, as a start codon in B. subtilis.

Bacillus subtilis

Analysis of leaky viral translation termination codons in vivo by transient expression of improved beta-glucuronidase vectors.

Plant RNA viruses commonly exploit leaky translation termination signals in order to express internal protein coding regions. As a first step to elucidate the mechanism(s) by which ribosomes bypass leaky stop codons in vivo, we have devised a system in which readthrough is coupled to the transient expression of beta-glucuronidase (GUS) in tobacco protoplasts. GUS vectors that contain the stop codons and surrounding nucleotides from the readthrough regions of several different RNA viruses were constructed and the plasmids were tested for the ability to direct transient GUS expression. These studies indicated that ribosomes bypass the leaky termination sites at efficiencies ranging from essentially 0 to ca. 5% depending upon the viral sequence. The results suggest that the efficiency of readthrough is determined by the sequence surrounding the stop codon. We describe improved GUS expression vectors and optimized transfection conditions which made it possible to assay low-level translational events.

Base Sequence

Modified bacteriophage lambda promoter vectors for overproduction of proteins in Escherichia coli.

A new series of expression vectors that direct high-level overproduction of gene products in Escherichia coli is described. All contain strong bacteriophage lambda promoters, PR and PL, arranged in tandem so that both promote transcription into genes inserted into or between unique restriction sites. The vectors also direct expression of the lambda cI857 gene (from its natural promoter, PM), which enables their use in any E. coli host strain to effect controlled expression by shifting the temperature of cultures from 30 to 42 degrees C. The vectors pCE30, pND201, pPT150 and pMA200U are derivatives of the high-copy-number plasmid pUC9. Vector pCE33 is an analogous derivative of the heat-inducible runaway-replication plasmid, pMOB45, and directs overproduction of proteins by virtue of increase in both gene dosage and transcription following treatment at 42 degrees C. The vectors pND201 and pPT150 bear a ribosome-binding site (RBS) perfectly complementary to the 3' end of E. coli 16-S rRNA a few bp upstream from a unique HpaI site. Ways in which they may be used to improve the efficiency of translation of mRNA by substitution of a natural RBS with selection for optimal spacing from an ATG (or GTG) start codon are described. The phagemid vector pMA200U is a direct analog of pCE30 designed to facilitate preparation of single-stranded DNA templates for use in oligodeoxyribonucleotide-directed mutagenesis of overexpressed genes.

Bacteriophage lambda

Streptomycin causes misreading of natural messenger by interacting with ribosomes after initiation.

The induction of misreading by streptomycin in vitro, previously observed with synthetic messengers, is now demonstrated with natural (endogenous or viral) messenger by the use of extracts of temperature sensitive mutants lacking Glu--tRNA or Val--tRNA synthetase. With chain-elongating but noninitiating ribosomes (i.e., purified polysomes) deprived of an aminoacyl--tRNA, streptomycin and other aminoglycosides, over a wide range of concentrations, stimulate incorporation. With ribosomes initiating in the presence of streptomycin stimulation is also observed but it is restricted, just like phenotypic suppression in cells, to very low streptomycin concentrattions which evidently allow some ribosomes to initiate and later encounter them in the course of chain elongation. The stimulation is accompanied by an increase in the size of the products; hence, it is evidently due to substitution of an incorrect aminoacyl--tRNA for a missing one. The test introduced here also has revealed a misreading effect of streptomycin on resistant ribosomes. In addition, significant intrinsic misreading was observed without streptomycin, indicating that under optimal conditions for in vitro protein synthesis an empty codon is frequently read by an incorrect aminoacyl--tRNA.

Anti-Bacterial Agents

Co-expression of a precursor and the mature protein of wheat ribulose-1,5-bisphosphate carboxylase small subunit from a single gene in Escherichia coli.

The cDNA encoding a precursor of wheat ribulose-1,5-bisphosphate carboxylase/oxygenase was inserted in-phase with prokaryotic expression elements in four different vectors. Five expression vectors encoding the small subunit precursors were cloned in Escherichia coli. None of these constructs expressed detectable amounts of the precursor protein, but all directed synthesis of the mature small subunit. The expression of the small subunit was a consequence of an independent, intragenic Shine-Dalgarno sequence optimally located upstream from an ATG specifying the first codon of the mature small subunit portion in the precursor transcript. Similar internal translation signals have been identified in the nuclear-encoded cDNAs of the small-subunit precursors of numerous higher plant genes. The 5' end of the wheat small-subunit precursor was linked with a consensus E. coli DNA sequence such that the modified gene encoded a partial hybrid precursor carrying four additional residues at its amino terminus. The resultant construct, pEI-W3, directed abundant synthesis of both the partially hybrid small-subunit precursor and the mature small subunit, constituting as much as 10% of the total bacterial protein. The bacterially synthesized small subunit precursor was purified to homogeneity. The authenticity of the recombinant protein was verified by its size, immunological properties, amino-terminal sequence, and amino acid composition.

Amino Acid Sequence

Codon bias and gene expression.

The frequencies with which individual synonymous codons are used to code their cognate amino acids is quite variable from genome to genome and within genomes, from gene to gene. One particularly well documented codon bias is that associated with highly expressed genes in bacteria as well as in yeast; this is the so-called major codon bias. Here, it is suggested that the major codon bias is not an arrangement for regulating individual gene expression. Instead, the data suggest that this codon bias, which is correlated with a corresponding bias of tRNA abundance, is a global arrangement for optimizing the growth efficiency of cells. On the practical side, it is suggested that heterologous gene expression is not as sensitive to codon bias as previously thought, but that it is quite sensitive to other characteristics of the heterologous gene.

Codon

Detecting evolutionary trends from molecular data. 1. Some measures of compositional nonrandomness.

The measures of compositional nonrandomness to be discussed as to their physical significance and to their power of detecting evolutionary significant variations are (see article)(pi a priori probability for amino acid i, ni its number of occurrences in a protein of length L). As a concrete example, the pi are here supposed to represent equal frequencies of all non-stop codons. For each quantity, four levels are defined: The base level, with optimal (i.e. minimal nonrandomness) composition, admitting non-integer values of ni; the integer level with optimal integer composition; the noise level, represented by a typical random cain; and the real protein level. On all these levels, S, which is the measure with the most direct physical sense, shows the smoothest behavior with the smallest relative fluctuations and thus the highest resolution.

Albumins

The influence of ribosome-binding-site elements on translational efficiency in Bacillus subtilis and Escherichia coli in vivo.

A method is described to determine simultaneously the effect of any changes in the ribosome-binding site (RBS) of mRNA on translational efficiency in Bacillus subtilis and Escherichia coli in vivo. The approach was used to analyse systematically the influence of spacing between the Shine-Dalgarno sequence and the initiation codon, the three different initiation codons, and RBS secondary structure on translational yields in the two organisms. Both B. subtilis and E. coli exhibited similar spacing optima of 7-9 nucleotides. However, B. subtilis translated messages with spacings shorter than optimal much less efficiently than E. coli. In both organisms, AUG was the preferred initiation codon by two- to threefold. In E. coli GUG was slightly better than UUG while in B. subtilis UUG was better than GUG. The degree of emphasis placed on initiation codon type, as measured by translational yield, was dependent on the strength of the Shine-Dalgarno interaction in both organisms. B. subtilis was also much less able to tolerate secondary structure in the RBS than E. coli. While significant differences were found between the two organisms in the effect of specific RBS elements on translation, other mRNA components in addition to those elements tested appear to be responsible, in part, for translational species specificity. The approach described provides a rapid and systematic means of elucidating such additional determinants.

Bacillus subtilis

The Use of Deep Learning in RNA Therapeutic Development.

Ribonucleic acid (RNA)-based therapeutics have emerged as promising methods of disease treatment due to their ability to target the human genome and influence protein production, their versatility, and their relative lack of toxicity compared to other gene therapies. However, the RNA therapeutic design space is extremely large, encompassing multiple variables, including codon identities, secondary structure, and design of specific regions. RNA therapeutic optimization is difficult due to the impracticality of exploring such a vast design space experimentally. To address this limitation, deep learning methods have been employed to optimize RNA therapeutic development. In this review, we examine the application of deep learning models across three key aspects of RNA therapeutic development (RNA structure prediction, CRISPR activity, and RNA delivery), highlighting major contributions in these fields and analyzing how deep learning model architectures could affect model performance. We then discuss challenges associated with using deep learning for RNA therapeutics, such as computational and data limitations. Finally, we offer perspectives on areas for future exploration, such as emerging model architectures and methods of integration with more advanced high-throughput screening techniques. Ultimately, this review provides an overview of how deep learning is used in RNA therapeutic development and how it can evolve in the future.

Deep Learning

Computer-aided gene design.

A computer program, which runs on MS-DOS personal computers, is described that assists in the design of synthetic genes coding for proteins. The goal of the program is the design of a gene which (i) contains as many unique restriction sites as possible and (ii) uses a specific codon usage. The gene designed according to the criteria above is (i) suitable for 'modular mutagenesis' experiments and (ii) optimized for expression. The program 'reverse-translates' protein sequences into degenerated DNA sequences, generates a map of potential restriction sites and locates sequence positions where unique restriction sites can be accommodated. The nucleic acid sequence is then 'refined' according to a specific codon usage to remove any degeneration. Unique restriction sites, if potentially present, can be 'forced' into the degenerated nucleic acid sequence by using 'priority codes' assigned to different restriction sequences.

Amino Acid Sequence

Construction and expression of nonsense suppressor tRNAs which function in plant cells.

An Arabidopsis thaliana L. DNA containing the tRNA(TrpUGG) gene was isolated and altered to encode the amber suppressor tRNA(TrpUAG) or the ochre suppressor tRNA(TrpUAA). These DNAs were electroporated into carrot protoplasts and tRNA expression was demonstrated by the translational suppression of amber and ochre nonsense mutations in the chloramphenicol acetyltransferase (CAT) reporter gene. DNAs encoding tRNA(TrpUAG) and tRNA(TrpUAA) nonsense suppressor tRNAs caused suppression of their cognate nonsense codons in CAT mRNAs, with the tRNA(TrpUAG) gene exhibiting the greater suppression under optimal conditions for expression of CAT. The development of these translational suppressors which function in plant cells facilitates the study of plant tRNA gene expression and will make possible the manipulation of plant protein structure and function.

Anticodon

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing

Mutations affecting translational coupling between the rep genes of an IncB miniplasmid.

The nature of translational coupling between repB and repA, the overlapping rep genes of the IncB plasmid pMU720, was examined. Mutations in the start codon of the promoter proximal gene, repB, reduced the efficiency of translation of both rep genes. Moreover, there was no independent initiation of repA translation in the absence of repB translation. The position of the repB stop codon was crucial for the efficient expression of repA, with the wild-type positioning being optimal. Translational coupling was found to be totally dependent on the formation of a pseudoknot structure. A model which invokes formation of a pseudoknot to facilitate initiation of repA is proposed.

Bacterial Proteins

Characterization of an alpha 1----3-galactosyltransferase homologue on human chromosome 12 that is organized as a processed pseudogene.

UDP-Gal:Gal beta 1----4GlcNAc alpha 1----3-galactosyltransferase is a terminal glycosyltransferase that is widely expressed in a variety of mammalian species, with the notable exception of man, apes, and Old World monkeys. We recently reported the isolation of a bovine cDNA clone that contains the complete coding sequence for this enzyme (Joziasse, D. H., Shaper, J. H., Van den Eijnden, D. H., Van Tunen, A. J., and Shaper, N. L. (1989) J. Biol. Chem. 264, 14290-14297). Using this cDNA as a probe, we have demonstrated that, although transcripts cannot be detected in a variety of established human cell lines by Northern blot analysis, homologous sequences are present in human genomic DNA. To establish that these sequences represent a human homologue of alpha 1----3-galactosyltransferase, we have used the bovine cDNA as a probe to isolate two nonoverlapping clones (HGT-2 and HGT-10) from a human genomic DNA library. Clone HGT-2 contains a 1.5-kilobase uninterrupted linear sequence similar to bovine alpha 1----3-galactosyltransferase that is organized as a processed pseudogene. This sequence, flanked by Alu type repeats, contains a short 5'- and 3'-untranslated region and a complete recognizable coding region that is 81% similar at the nucleotide level to bovine alpha 1----3-galactosyltransferase. This putative coding region contains multiple frameshift mutations and nonsense codons in all three reading frames which precludes the synthesis of a functional enzyme. Nevertheless, after optimal alignment, translation predicts a polypeptide that is 68% similar at the amino acid level to the bovine enzyme. Based on Southern analysis and limited sequence analysis, clone HGT-10 contains coding sequences similar to the NH2-terminal region of bovine alpha 1----3-galactosyltransferase. By analysis of panels of human-rodent somatic cell hybrids we have established that the nonfunctional, processed pseudogene and the human homologue represented by HGT-10 are located on human chromosomes 12 and 9, respectively. Interestingly, a comparison of the predicted amino acid sequence of the carboxyl-terminal two-thirds of human alpha 1----3-galactosyltransferase, with the corresponding region of the human blood group A, UDP-GalNAc:[Fuc alpha 1----2]Gal beta 1----4GlcNAc alpha 1----3-GalNAc-transferase (Yamamoto, F., Marken, J., Tsuji, T., White, T., Clausen, H., and Hakomori, S. (1990a) J. Biol. Chem. 265, 1146-1151), reveals a significant similarity (39%) suggesting that these two enzymes may have arisen from the same ancestral gene as a result of gene duplication and subsequent divergence.

Amino Acid Sequence

Expression of human asparagine synthetase in Saccharomyces cerevisiae.

Human asparagine synthetase was expressed in the yeast Saccharomyces cerevisiae. The identity of the expressed protein was confirmed by immunoblotting and in vitro enzymatic activity. The recombinant enzyme was shown to have both the ammonia- and glutamine-dependent asparagine synthetase activity in vitro. In contrast to overproduction in Escherichia coli, the expressed protein was found to be soluble in the yeast cell. Furthermore, expression in yeast made it possible to isolate non-degraded human asparagine synthetase which had also the N-terminal methionine correctly processed. The yeast expression plasmid was constructed for optimal production of the recombinant enzyme. In addition, unique restriction enzyme sites that bracket the first five codons of the human asparagine synthetase gene were introduced. This will allow the use of oligonucleotide cassette mutagenesis to investigate the role of the N-terminal amino acids in asparagine synthetase enzymatic activity.

Amino Acid Sequence

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment