Search PubMedSearch

SEARCH · Search PubMed

Results for “codon optimization”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

The Use of Deep Learning in RNA Therapeutic Development.

Ribonucleic acid (RNA)-based therapeutics have emerged as promising methods of disease treatment due to their ability to target the human genome and influence protein production, their versatility, and their relative lack of toxicity compared to other gene therapies. However, the RNA therapeutic design space is extremely large, encompassing multiple variables, including codon identities, secondary structure, and design of specific regions. RNA therapeutic optimization is difficult due to the impracticality of exploring such a vast design space experimentally. To address this limitation, deep learning methods have been employed to optimize RNA therapeutic development. In this review, we examine the application of deep learning models across three key aspects of RNA therapeutic development (RNA structure prediction, CRISPR activity, and RNA delivery), highlighting major contributions in these fields and analyzing how deep learning model architectures could affect model performance. We then discuss challenges associated with using deep learning for RNA therapeutics, such as computational and data limitations. Finally, we offer perspectives on areas for future exploration, such as emerging model architectures and methods of integration with more advanced high-throughput screening techniques. Ultimately, this review provides an overview of how deep learning is used in RNA therapeutic development and how it can evolve in the future.

Deep Learning

Amino acid composition is correlated with protein abundance in Escherichia coli: can this be due to optimization of translational efficiency?

Amino acid occurrence frequencies were found for four groups of Escherichia coli proteins with different abundance levels in the cell. These frequencies decrease with increasing protein abundance for amino acids whose codons are translated by tRNAs present at low concentrations (e.g., Cys, Trp, Ser, etc.); the opposite tendency was observed for amino acids translated by abundant tRNAs (Lys, Val, etc.). The efficiency (rate and accuracy) of codon translation is expected to be proportional to the concentration of the cognate tRNA. Therefore, the observed constraints on amino acid composition may be explained as resulting from evolutionary pressure optimizing the translational efficiency of a gene (the same pressure is responsible for the nonrandom choice of synonymous codons).

Amino Acids

[Effectiveness of translation coupling in hybrid operons].

The possibility of creating artificial overlappons was studied on the model of two genes, that coding for the N-terminal part of lambda cro protein and the cat of E. coli. To test the dependence of translational coupling efficiency on the intercistronic region a series of recombinant DNA molecules carrying different hybrid operons with partially overlapping genes was constructed. The translational efficiency of the distal to the promoter gene was shown to depend on the intercistronic region structure: overlapping of the AUG codon with the terminating one of the proximal gene in the UGAUG manner is optimal for the translational coupling, and the displacement of AUG at several nucleotides in both directions decreases the translational reinitiation efficiency for the ribosomes, that have synthesised the first gene product.

Acetyltransferases

Truncation of the carboxy-terminal domain of yeast beta-tubulin causes temperature-sensitive growth and hypersensitivity to antimitotic drugs.

beta-tubulin of budding yeast Saccharomyces cerevisiae is a polypeptide of 457 amino acids encoded by the unique gene TUB2. We investigated the function of the carboxy-terminal part of yeast beta-tubulin corresponding to the carboxy-terminal variable domain of mammalian and avian beta-tubulins. The GAA codon for Glu-431 of TUB2 was altered to TAA termination codon by using in vitro site-directed mutagenesis so that the 27-amino acid residues of the carboxyl terminus was truncated when expressed. The mutagenized TUB2 gene (tub2(T430)) was introduced into a haploid strain in which the original TUB2 gene had been disrupted. The tub2(T430) haploid strain grows normally less than 30 but not at 37 degrees C. The truncation of the carboxyl terminus caused hypersensitivity to antimitotic drugs and low spore viability at the permissive temperature for vegetative growth. Immunofluorescence labeling with antitubulin antibody and DNA staining with 4',6'-diamidino-2-phenylindole showed that in these cells at 37 degrees C, formation of spindle microtubules and nuclear division was inhibited and cytoplasmic microtubule distribution was aberrant. These results suggest that functions of the carboxy-terminal domain of yeast beta-tubulin are necessary for cells growing under suboptimal growth conditions although it is not essential for growth under the optimal growth conditions. Cells bearing tub2(411), a tub2 gene in which the GAA codon for Glu-412 was altered to TAA were no more viable at any temperature. In addition, a haploid strain carrying two functional beta-tubulin genes is not viable.

Antineoplastic Agents

Computer-aided gene design.

A computer program, which runs on MS-DOS personal computers, is described that assists in the design of synthetic genes coding for proteins. The goal of the program is the design of a gene which (i) contains as many unique restriction sites as possible and (ii) uses a specific codon usage. The gene designed according to the criteria above is (i) suitable for 'modular mutagenesis' experiments and (ii) optimized for expression. The program 'reverse-translates' protein sequences into degenerated DNA sequences, generates a map of potential restriction sites and locates sequence positions where unique restriction sites can be accommodated. The nucleic acid sequence is then 'refined' according to a specific codon usage to remove any degeneration. Unique restriction sites, if potentially present, can be 'forced' into the degenerated nucleic acid sequence by using 'priority codes' assigned to different restriction sequences.

Amino Acid Sequence

Construction and expression of nonsense suppressor tRNAs which function in plant cells.

An Arabidopsis thaliana L. DNA containing the tRNA(TrpUGG) gene was isolated and altered to encode the amber suppressor tRNA(TrpUAG) or the ochre suppressor tRNA(TrpUAA). These DNAs were electroporated into carrot protoplasts and tRNA expression was demonstrated by the translational suppression of amber and ochre nonsense mutations in the chloramphenicol acetyltransferase (CAT) reporter gene. DNAs encoding tRNA(TrpUAG) and tRNA(TrpUAA) nonsense suppressor tRNAs caused suppression of their cognate nonsense codons in CAT mRNAs, with the tRNA(TrpUAG) gene exhibiting the greater suppression under optimal conditions for expression of CAT. The development of these translational suppressors which function in plant cells facilitates the study of plant tRNA gene expression and will make possible the manipulation of plant protein structure and function.

Anticodon

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing

cDNA cloning and functional expression in yeast Saccharomyces cerevisiae of beta-naphthoflavone-induced rabbit liver P-450 LM4 and LM6.

A cDNA library was constructed from liver mRNA of a beta-naphthoflavone-induced rabbit. Two clones pLM4-1 and pLM6-1 containing 2.2-kbp inserts that hybridized at low stringincy with a mouse P1 P-450 probe were selected. The clone pLM4-1 was fully sequenced and found to contain a full-length cDNA coding for cytochrome P-450 LM4. Partial sequence and restriction mapping made it possible to identify pLM6-1 as coding for the major part of cytochrome P-450 LM6. Cloned LM4-1 cDNA was reformed by deletion of the 5' and 3' non-coding regions before insertion into yeast expression vectors PYe DP1/10. A similar operation was performed on pLM6-1 cDNA after replacement of the missing N-terminus-coding sequences by homologous sequences form the pLM4-1 clone resulting in a chimeric cytochrome P-450 coding sequence. Expression of cloned rabbit cytochrome P-450 into transformed yeast was optimized by studying the effect of the nature of the DNA sequence just preceding the initiation codon on the level of cytochrome P-450 production. Yeast synthesized cytochromes P-450 were characterized by immunoblotting, spectra and catalytic activity determinations. Cloned cytochrome P-450 LM4 was found by all criteria to be identical to the authentic rabbit one. The chimeric cytochrome P-450 that contains the 143 N-terminal amino acids of cytochrome P-450 LM4 and the remaining 375 amino acids of cytochrome P-450 LM6 was found to exhibit most of the authentic cytochrome P-450 LM6 catalytic properties. Enzymatic and evolutionary implications of these results are discussed.

Amino Acid Sequence

Mutations affecting translational coupling between the rep genes of an IncB miniplasmid.

The nature of translational coupling between repB and repA, the overlapping rep genes of the IncB plasmid pMU720, was examined. Mutations in the start codon of the promoter proximal gene, repB, reduced the efficiency of translation of both rep genes. Moreover, there was no independent initiation of repA translation in the absence of repB translation. The position of the repB stop codon was crucial for the efficient expression of repA, with the wild-type positioning being optimal. Translational coupling was found to be totally dependent on the formation of a pseudoknot structure. A model which invokes formation of a pseudoknot to facilitate initiation of repA is proposed.

Bacterial Proteins

Characterization of an alpha 1----3-galactosyltransferase homologue on human chromosome 12 that is organized as a processed pseudogene.

UDP-Gal:Gal beta 1----4GlcNAc alpha 1----3-galactosyltransferase is a terminal glycosyltransferase that is widely expressed in a variety of mammalian species, with the notable exception of man, apes, and Old World monkeys. We recently reported the isolation of a bovine cDNA clone that contains the complete coding sequence for this enzyme (Joziasse, D. H., Shaper, J. H., Van den Eijnden, D. H., Van Tunen, A. J., and Shaper, N. L. (1989) J. Biol. Chem. 264, 14290-14297). Using this cDNA as a probe, we have demonstrated that, although transcripts cannot be detected in a variety of established human cell lines by Northern blot analysis, homologous sequences are present in human genomic DNA. To establish that these sequences represent a human homologue of alpha 1----3-galactosyltransferase, we have used the bovine cDNA as a probe to isolate two nonoverlapping clones (HGT-2 and HGT-10) from a human genomic DNA library. Clone HGT-2 contains a 1.5-kilobase uninterrupted linear sequence similar to bovine alpha 1----3-galactosyltransferase that is organized as a processed pseudogene. This sequence, flanked by Alu type repeats, contains a short 5'- and 3'-untranslated region and a complete recognizable coding region that is 81% similar at the nucleotide level to bovine alpha 1----3-galactosyltransferase. This putative coding region contains multiple frameshift mutations and nonsense codons in all three reading frames which precludes the synthesis of a functional enzyme. Nevertheless, after optimal alignment, translation predicts a polypeptide that is 68% similar at the amino acid level to the bovine enzyme. Based on Southern analysis and limited sequence analysis, clone HGT-10 contains coding sequences similar to the NH2-terminal region of bovine alpha 1----3-galactosyltransferase. By analysis of panels of human-rodent somatic cell hybrids we have established that the nonfunctional, processed pseudogene and the human homologue represented by HGT-10 are located on human chromosomes 12 and 9, respectively. Interestingly, a comparison of the predicted amino acid sequence of the carboxyl-terminal two-thirds of human alpha 1----3-galactosyltransferase, with the corresponding region of the human blood group A, UDP-GalNAc:[Fuc alpha 1----2]Gal beta 1----4GlcNAc alpha 1----3-GalNAc-transferase (Yamamoto, F., Marken, J., Tsuji, T., White, T., Clausen, H., and Hakomori, S. (1990a) J. Biol. Chem. 265, 1146-1151), reveals a significant similarity (39%) suggesting that these two enzymes may have arisen from the same ancestral gene as a result of gene duplication and subsequent divergence.

Amino Acid Sequence

Expression of human asparagine synthetase in Saccharomyces cerevisiae.

Human asparagine synthetase was expressed in the yeast Saccharomyces cerevisiae. The identity of the expressed protein was confirmed by immunoblotting and in vitro enzymatic activity. The recombinant enzyme was shown to have both the ammonia- and glutamine-dependent asparagine synthetase activity in vitro. In contrast to overproduction in Escherichia coli, the expressed protein was found to be soluble in the yeast cell. Furthermore, expression in yeast made it possible to isolate non-degraded human asparagine synthetase which had also the N-terminal methionine correctly processed. The yeast expression plasmid was constructed for optimal production of the recombinant enzyme. In addition, unique restriction enzyme sites that bracket the first five codons of the human asparagine synthetase gene were introduced. This will allow the use of oligonucleotide cassette mutagenesis to investigate the role of the N-terminal amino acids in asparagine synthetase enzymatic activity.

Amino Acid Sequence

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment

Restructuring the translation initiation region of the human parathyroid hormone gene for improved expression in Escherichia coli.

Overexpression of native human parathyroid hormone in Escherichia coli was achieved by a modification of the 5' end of the genomic gene sequence, thereby adapting this part of the translation initiation region to the bacterial host. Some simple rules abstracted from optimization studies of translation initiation of a beta-interferon gene were applied. These included (a) extending complementarity of the mRNA to the anticodon loop of tRNAfMet by use of a codon with a purine nucleotide directly following the ATG, (b) avoidance of stable secondary structure in the mRNA by use of synonymous A/U-rich codons, (c) elimination of a potential second Shine-Dalgarno sequence. The appropriate silent changes led to a 20-fold increase in parathyroid hormone production resulting in 4.3% of total soluble protein. This result proves the validity of our simple approach for optimization of foreign gene expression in E. coli.

Base Sequence

Inducible expression vectors incorporating the Escherichia coli atpE translational initiation region.

New expression vectors were constructed for use in strains of Escherichia coli. Their most important feature is a polylinker system that facilitates the insertion of a gene in an optimal relationship to the highly efficient E. coli atpE translational initiation region (from nucleotide -50 to the start codon). Three ATG-containing restriction endonuclease sites can be used for the insertion of the 5' end of a gene at, or near to, its translational initiation codon. These sites may alternatively be used for the creation of a suitable translational start codon. Transcription is started by the bacteriophage lambda major promoters pR and pL in tandem and terminated by the bacteriophage fd terminator. Transcriptional initiation is very effectively repressed at 28-30 degrees C by the product of the bacteriophage lambda cIts857 gene, which is also present on the vectors. Full induction is achieved by shifting the incubation temperature to 42 degrees C. The combination of highly efficient transcriptional and translational signals on these vectors allowed high-level expression of sequences encoding human interferon beta and interleukin 2 and of the E. coli atpA, sucC and sucD genes.

DNA Restriction Enzymes

On the information content of the genetic code.

In living organisms 20 amino acids along with the terminator value(s) are encoded by 64 codons giving a degeneracy of the codons as described by the genetic code. A basic theoretical problem of genetic codes is to explain the particular distribution of degeneracies of partitions involved in the codes. In this work the degeneracy problem is considered in the framework of information theory. It is shown by direct numerical evaluation of a certain degeneracy information function associated with the genetic code that the degeneracy of the codes is observed to be related to the optimization of this function.

Amino Acids

Hyperexpression of a Bacillus thuringiensis delta-endotoxin-encoding gene in Escherichia coli: properties of the product.

Conditions for hyperexpression, in Escherichia coli, of the Bacillus thuringiensis var, kurstaki gene, cryIA9(c)73, encoding an insecticidal crystal protein, CryIA(c)73, were investigated by varying the promoter type, host cell, plasmid copy number, the second codon and number of terminators. The cryIA(c)73 gene was cloned into three E. coli expression vectors, pKK223-3 (Ptac promoter), pET-3a (P phi 10 promoter), and pUC19 (Ptac promoter). The level of cryIA(c)73 expression was measured by ELISA and compared to total cellular protein over growth periods of 24 and 48 h. Maximum expression levels of 284 microgram CryIA(C)73/ml (48% of cellular protein) were obtained in shake flasks with the Ptac promoter in E. coli JM103. Optimal conditions were found to be low-copy-number plasmid (pBR322 ori), 48 h of growth, in lon+ cells. A change of the gene's second codon to AAA can improve expression by two to three fold but is undetectable in the presence of a strong E. coli promoter. The cryIA(c)73 gene product, in E. coli, formed crystals with the same lattice structure as the native crystals formed in B. thuringiensis (as visualized by electron microscopy). Bioassay results (insect toxicity and specificity) of the crystal produced in E. coli were similar to that produced in B. thuringiensis.

Bacillus thuringiensis

Optimization of the hygromycin B resistance-conferring gene as a dominant selectable marker in mammalian cells.

The HyR gene, conferring resistance to hygromycin B (Hy), has been modified for optimal expression in mammalian cells. Modifications to the HyR gene and its expression cassette include: (1) removal of all upstream start codons, (2) conversion of the region around the start codon to the consensus sequence associated with efficient translation initiation, and (3) removal of downstream splice donor and acceptor sequences. The resulting HyR gene is an efficient dominant selectable marker that is useful for studies requiring resistance from a low-copy-number gene driven by a promoter of moderate strength. The HyR gene was also tested for its compatibility with BPV vectors. Mouse C127 cells harboring pHyR-BPV plasmids exhibited properties of BPV-transformed cells and were resistant to toxic levels of Hy. The vectors were stable as episomes and present in high copy. The HyR gene thus joins the NmR (neo) gene as the only dominant selectable markers that are known to be compatible with BPV replication.

Animals