Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Enzymatic production of trans-4-hydroxy-L-proline by regio- and stereospecific hydroxylation of L-proline.

A proline 4-hydroxylase gene, which was cloned from Dactylosporangium sp. RH1, was overexpressed in Escherichia coli W1485 on a plasmid under a tryptophan tandem promoter after the codon usage of the 5' end of the gene was optimized. The proline 4-hydroxylase activity was l600-fold higher than that in Dactylosporangium sp. RH1. trans-4-Hydroxy-L-proline(Hyp) was produced and accumulated to 41 g/L (87% yield from L-proline) in 100 h when the recombinant E. coli was cultivated in a medium containing L-proline and glucose. 2-Oxoglutarate, which is necessary for the hydroxylation of L-proline by proline 4-hydroxylase, was apparently supplied from glucose through the cellular metabolic pathway. The putA mutant of W1485, which is not able to degrade L-proline, has allowed the quantitative conversion of L-proline to Hyp. The formation of other isomers of hydroxyproline was not observed. Productivity of Hyp was almost the same in a larger-scale culture. The method of manufacturing Hyp from L-proline was established.

Amino Acid Sequence↗

[Cloning of human beta-microglobulin gene and its high expression in Escherichia coli].

Human beta2-microglobulin (beta2m) is the light chain of major histocompatibility complex (MHC) class I molecule. High-yield production of this protein is a prerequisite to the preparation of MHC I tetramer. The present study aims to obtain recombinant human beta2m expressed in Escherichia coli (E. coli), for the purpose of preparing MHC class I tetramers. For cloning of human beta2m gene, a pair of specific primers was designed based on the published sequence of this gene and the cDNA of full coding region for beta2m precursor was obtained by RT-PCR from the total RNA of human leukocytes. The amplified cDNA was subsequently cloned and its sequence was confirmed by DNA sequencing analysis (the sequence has been deposited in GenBank with accession number of AY187687). The prokaryotic expression vector containing a gene encoding mature beta2m was constructed by inserting the DNA fragment, which was generated by PCR reaction with the cloned beta2m gene as template, into an IPTG-inducible expression vector pET-3c plasmid. The first eight codons for N terminal amino acid residues of beta2m were optimized for its expression in E. coli. The complete sequence of beta2 m gene in the expression vector was verified by DNA sequencing analysis. High-yield expression of beta2m was achieved in E. coli transformed with the expression vector, and most of the recombinant beta2m existed in the inclusion body after IPTG induction. The inclusion body was washed extensively and beta2m in the inclusion body was solublized with 8 mol/L urea. The beta2m was refolded by dialysis and purified by ion-exchange chromatography (Q-Sepharose). Western blotting assay indicated that the polyclonal antibody against human native beta2m could react specifically with the recombinant protein. The purified protein appeared as a single band on both SDS-PAGE and Western blotting, indicating that it was chemical and antigenic pure. This work establishes a convenient approach for renaturation and purification of large quantity of recombinant beta2m which is identical to the native protein without any tags fused except for a methionine residue at the amino terminus. This provides the basis for the preparation of MHC tetramers.

Base Sequence↗

Angiosperm divergence times: the effect of genes, codon positions, and time constraints.

An understanding of the evolution of modern terrestrial ecosystems requires an understanding of the dynamics associated with angiosperm evolution, including the timing of their origin and diversification into their extraordinary present-day diversity. Molecular estimates of angiosperm age have varied widely, and many substantially predate the Early Cretaceous fossil appearance of the group. In this study, the effect of different genes, codon positions, and chronological constraints on node ages are examined on divergence time estimates across seed plants, with a special focus on angiosperms. Penalized likelihood was used to estimate divergence times on a phylogenetic hypothesis for seed plants derived from Bayesian analysis, with branch lengths estimated with maximum likelihood. The plastid genes atpB, psaA, psbB, and rbcL were used individually and in combination, using first and second, third, and the three codon positions, including and excluding age constraints on 20 nodes derived from a critical examination of the land-plant fossil record. The optimal level of rate smoothing according to each unconstrained and constrained dataset was obtained with penalized likelihood. Tests for a molecular clock revealed significantly unclocklike rates in all datasets. Addition of fossil constraints resulted in even greater departures from constancy. Consistently with significant deviations from a clock, estimated optimal smoothing values were low, but a strict correlation between rate heterogeneity and optimal smoothing value was not found. Age estimates for nodes across the phylogeny varied, sometimes substantially, with gene and codon position. Nevertheless, estimates based on the four concatenated genes are very similar to the mean of the four individual gene estimates. For any given node, unconstrained age estimates are more variable than constrained estimates and are frequently younger than well-substantiated fossil members of the clade. Constrained estimates of ages of clades are older than unconstrained estimates and oldest fossil representatives, sometimes substantially so. Angiosperm age estimates decreased as rate smoothing increased. Whereas the range of unconstrained angiosperm age estimates spans the fossil age of the clade, the range of constrained estimates is narrower (and older) than the earliest angiosperm fossils. Results unambiguously indicate the relevance of constraints in reducing the variability of ages derived from different partitions of the data and diminishing the effect of the smoothing parameter. Constrained optimizations of divergence times and substitution rates across the phylogeny suggest appreciably different evolutionary dynamics for angiosperms and for gymnosperms. Whereas the gymnosperm crown group originated shortly after the origin of seed plants, a long time elapsed before the origin of crown group angiosperms. Although absolute age estimates of angiosperms and angiosperm clades are older than their earliest fossils, the estimated pace of phylogenetic diversification largely agrees with the rapid appearance of angiosperm lineages in stratigraphic sequences.

Base Sequence↗

Expression of a synthetic gene encoding a Tribolium castaneum carboxylesterase in Pichia pastoris.

This is the first report of an insect esterase efficiently expressed in the methylotrophic yeast Pichia pastoris (so far insect esterases have been produced only in the baculovirus system). Having isolated a Tribolium castaneum carboxylesterase cDNA (TCE), we were initially unable to express it in Escherichia coli or P. pastoris despite significant transcription levels. As codon usage bias is different in T. castaneum and P. pastoris, we assumed this was a possible explanation for the translational barrier observed in yeast. Accordingly, we designed and constructed by recursive PCR a synthetic TCE gene (synTCE) optimized for heterologous expression in P. pastoris, i.e., a gene in which certain TCE codons are replaced with synonymous codons 'preferred' in P. pastoris. When the altered gene was placed under the control of either the P. pastoris glyceraldehyde-3-phosphate dehydrogenase (GAP) promoter or the inducible alcohol oxidase (AOX1) promoter and introduced on an expression vector into P. pastoris, its product was produced intracellularly. We also successfully explored the possibility of obtaining a secreted product: P. pastoris cells expressing an in-frame fusion of synTCE with the alpha-factor secretion signal under the control of the GAP promoter were found to secrete the recombinant esterase into the external medium (to a concentration of 7 mg/L). In addition to this demonstration of TCE production in yeast, our results suggest that the GAP promoter could advantageously replace the AOX1 promoter as a driver of synTCE expression. TCE specific activity was approximately 5 U/mg when p-nitrophenyl acetate was used as substrate.

Animals↗

Detection of rare K-ras codon 12 mutations using allele-specific competitive blocker PCR.

Allele-specific competitive blocker PCR (ACB-PCR) is a sensitive allele-specific amplification method in which preferential amplification of the mutant allele occurs by using a primer that has more mismatches to the wild-type allele than to the mutant allele (mutant-specific primer, MSP). Additionally, a non-extendable primer with more mismatches to the mutant allele than to the wild-type allele (blocker primer, BP) competes with the MSP for binding to the wild-type allele, thereby reducing background amplification from the wild-type allele. ACB-PCR primer design is largely dependent upon the basepair substitution being measured, making it unclear if this method is broadly applicable. In an earlier study, an H-ras codon 61 CAA-->AAA mutation had been detected by ACB-PCR at a sensitivity of 10(-5). In this study, ACB-PCR was applied to two human K-ras codon 12 mutations: GGT-->GTT and GGT-->GAT. The method was optimized by systematically altering the concentrations of Perfect Match PCR Enhancer, MSP, BP, and dNTPs. For each mutation, mutant fractions as low as 10(-5) were detected, indicating that this assay can be used on a variety of base substitution mutations. In addition, the results suggest that the 3'-terminal mismatches between the MSP and wild-type allele may be used to predict the ACB-PCR conditions that will be appropriate for the detection of other base substitution mutations. The range of concentrations for each of these components is narrow, making this method relatively easy to apply to additional mutational targets.

Alleles↗

Translation of the F protein of hepatitis C virus is initiated at a non-AUG codon in a +1 reading frame relative to the polyprotein.

The hepatitis C virus (HCV) genome contains an internal ribosome entry site (IRES) followed by a large open reading frame coding for a polyprotein that is cleaved into 10 proteins. An additional HCV protein, the F protein, was recently suggested to result from a +1 frameshift by a minority of ribosomes that initiated translation at the HCV AUG initiator codon of the polyprotein. In the present study, we reassessed the mechanism accounting for the synthesis of the F protein by measuring the expression in cultured cells of a luciferase reporter gene with an insertion encompassing the IRES plus the beginning of the HCV-coding region preceding the luciferase-coding sequence. The insertion was such that luciferase expression was either in the +1 reading frame relative to the HCV AUG initiator codon, mimicking the expression of the F protein, or in-frame with this AUG, mimicking the expression of the polyprotein. Introduction of a stop codon at various positions in-frame with the AUG initiator codon and substitution of this AUG with UAC inhibited luciferase expression in the 0 reading frame but not in the +1 reading frame, ruling out that the synthesis of the F protein results from a +1 frameshift. Introduction of a stop codon at various positions in the +1 reading frame identified the codon overlapping codon 26 of the polyprotein in the +1 reading frame as the translation start site for the F protein. This codon 26(+1) is either GUG or GCG in the viral variants. Expression of the F protein strongly increased when codon 26(+1) was replaced with AUG, or when its context was mutated into an optimal Kozak context, but was severely decreased in the presence of low concentrations of edeine. These observations are consistent with a Met-tRNA(i)-dependent initiation of translation at a non-AUG codon for the synthesis of the F protein.

Base Sequence↗

Detecting evolutionary trends from molecular data. 1. Some measures of compositional nonrandomness.

The measures of compositional nonrandomness to be discussed as to their physical significance and to their power of detecting evolutionary significant variations are (see article)(pi a priori probability for amino acid i, ni its number of occurrences in a protein of length L). As a concrete example, the pi are here supposed to represent equal frequencies of all non-stop codons. For each quantity, four levels are defined: The base level, with optimal (i.e. minimal nonrandomness) composition, admitting non-integer values of ni; the integer level with optimal integer composition; the noise level, represented by a typical random cain; and the real protein level. On all these levels, S, which is the measure with the most direct physical sense, shows the smoothest behavior with the smallest relative fluctuations and thus the highest resolution.

Albumins↗

The influence of ribosome-binding-site elements on translational efficiency in Bacillus subtilis and Escherichia coli in vivo.

A method is described to determine simultaneously the effect of any changes in the ribosome-binding site (RBS) of mRNA on translational efficiency in Bacillus subtilis and Escherichia coli in vivo. The approach was used to analyse systematically the influence of spacing between the Shine-Dalgarno sequence and the initiation codon, the three different initiation codons, and RBS secondary structure on translational yields in the two organisms. Both B. subtilis and E. coli exhibited similar spacing optima of 7-9 nucleotides. However, B. subtilis translated messages with spacings shorter than optimal much less efficiently than E. coli. In both organisms, AUG was the preferred initiation codon by two- to threefold. In E. coli GUG was slightly better than UUG while in B. subtilis UUG was better than GUG. The degree of emphasis placed on initiation codon type, as measured by translational yield, was dependent on the strength of the Shine-Dalgarno interaction in both organisms. B. subtilis was also much less able to tolerate secondary structure in the RBS than E. coli. While significant differences were found between the two organisms in the effect of specific RBS elements on translation, other mRNA components in addition to those elements tested appear to be responsible, in part, for translational species specificity. The approach described provides a rapid and systematic means of elucidating such additional determinants.

Bacillus subtilis↗

P0 of beet Western yellows virus is a suppressor of posttranscriptional gene silencing.

Higher plants employ a homology-dependent RNA-degradation system known as posttranscriptional gene silencing (PTGS) as a defense against virus infection. Several plant viruses are known to encode proteins that can suppress PTGS. Here we show that P0 of beet western yellows virus (BWYV) displays strong silencing suppressor activity in a transient expression assay based upon its ability to inhibit PTGS of green fluorescent protein (GFP) when expressed in agro-infiltrated leaves of Nicotiana benthamiana containing a GFP transgene. PTGS suppressor activity was also observed for the P0s of two other poleroviruses, cucurbit aphid-borne yellows virus and potato leafroll virus. P0 is encoded by the 5'-proximal gene in BWYV RNA but does not accumulate to detectable levels when expressed from the genome-length RNA during infection. The low accumulation of P0 and the resulting low PTGS suppressor activity are in part a consequence of the suboptimal translation initiation context of the P0 start codon in viral RNA, although other factors, probably related to the viral replication process, also play a role. A mutation to optimize the P0 translation initiation efficiency in BWYV RNA was not stable during virus multiplication in planta. Instead, the P0 initiation codon in the progeny was frequently replaced by a less efficient initiation codon such as ACG, GTG, or ATA, indicating that there is selection against overexpression of P0 from the viral genome.

Amino Acid Sequence↗

Translation of the flagellar gene fliO of Salmonella typhimurium from putative tandem starts.

The flagellar gene fliO of Salmonella typhimurium can be translated from an AUG codon that overlaps the termination codon of fliN (K. Ohnishi et al., J. Bacteriol. 179:6092-6099, 1997). However, it had been concluded on the basis of complementation analysis that in Escherichia coli a second start codon 60 bp downstream was the authentic one (J. Malakooti et al., J. Bacteriol. 176:189-197, 1994). This raised the possibility of tandem translational starts, such as occur for the chemotaxis gene cheA; this possibility was increased by the existence of a stem-loop sequence covering the second start, a feature also found with cheA. Protein translated from the first start codon was detected regardless of whether the second start codon was present; it was also detected when the stem-loop structure was disrupted or deleted. Translation from the second start codon, either as the natural one (GUG) or as AUG, was not detected when the first start and intervening sequence were intact. Nor was it detected when the first codon was attenuated (by conversion of AUGAUG to AUAAUA; in S. typhimurium there is a second, adjacent, AUG) or eliminated (by conversion to CGCCGC); disruption of the stem-loop structure still did not yield detectable translation from the second start. When the entire sequence up to the second start was deleted, translation from the second start was detected provided the natural codon GUG had been converted to AUG. A fliO null mutant could be fully complemented in swarm assays whenever the first start and intervening sequence were present, regardless of the state of the second start. Reasonably good complementation occurred when the first start and intervening sequence were absent provided the second start was intact, either as AUG or as GUG; thus translation from the GUG codon must have been occurring even though protein levels were too low to be detected. The translated intervening sequence is rather divergent between S. typhimurium and E. coli and corresponds to a substantial cytoplasmic domain prior to the sole transmembrane segment, which is highly conserved; the sequence following the second start begins immediately prior to that transmembrane segment. The significance of the data for FliO is discussed and compared to the equivalent data for CheA. Attention is also drawn to the fact that given an optimal ribosome binding site, AUA can serve as a fairly efficient start codon even though it seldom if ever appears to be used in nature.

Amino Acid Sequence↗

Designing a neural network for the constraint optimization of the fitness functions devised based on the load minimization of the genetic code.

Nonrandom patterns in codon assignments are supported by many statistical and biochemical studies in the last two decades. The canonical genetic code is known to be highly efficient in minimizing the effects of mistranslational errors and point mutations, an ability, which in term is designated "load minimization". Prior studies have included many attempts at quantitative estimation of the fraction of randomly generated codes, which in terms of load minimization, score higher than the canonical genetic code. In this study, a neural network, which estimates a highly optimized genetic code in a relatively short period of time has been devised. Several fitness functions were used throughout this text. Meanwhile, we have made use of two cost measure matrices, PAM74-100 and mutation matrix.

Algorithms↗

Translation of chloroplast-encoded mRNA: potential initiation and termination signals.

A survey of 196 protein-coding chloroplast DNA sequences demonstrated the preference for AUG and UAA codons for initiation and termination of translation, respectively. As in prokaryotes at every nucleotide position from -25 to +25 (AUG is +1 to +3) and for 25 nucleotides 5' and 3' to the termination codon an A or U is predominant, except for C at +5 and G at +22. A Shine-Dalgarno (SD) sequence (GGAGG or tri- or tetranucleotide variant) was found within 100 bp 5' to the AUG codon in 92% of the genes. In 40% of these cases, the location of the SD sequence was similar to that of the consensus for prokaryotes (-12 to -7 5' to AUG), presumed to be optimal for translation initiation. A SD sequence could not be located in 6% of the chloroplast sequences. We propose that mRNA secondary structures may be required for the relocation of a distal SD sequences to within the optimal region (-12 to -7) for initiation of translation. We further suggest that termination at UGA codons in chloroplast genes may occur by a mechanism, involving 16S rRNA secondary structure, which has been proposed for UGA termination in E. coli.

Base Composition↗

How initiation factors maximize the accuracy of tRNA selection in initiation of bacterial protein synthesis.

During initiation of bacterial protein synthesis, messenger RNA and fMet-tRNAfMet bind to the 30S ribosomal subunit together with initiation factors IF1, IF2, and IF3. Docking of the 30S preinitiation complex to the 50S ribosomal subunit results in a peptidyl-transfer competent 70S ribosome. Initiation with an elongator tRNA may lead to frameshift and an aberrant N-terminal sequence in the nascent protein. We show how the occurrence of initiation errors is minimized by (1) recognition of the formyl group by the synergistic action of IF2 and IF1, (2) uniform destabilization of the binding of all tRNAs to the 30S subunit by IF3, and (3) an optimal distance between the Shine-Dalgarno sequence and the initiator codon. We suggest why IF1 is essential for E. coli, discuss the role of the G-C base pairs in the anticodon stem of some tRNAs, and clarify gene expression changes with varying IF3 concentration in the living cell.

Bacterial Proteins↗

The optimal use of IRES (internal ribosome entry site) in expression vectors.

In higher eucaryotes, natural bicistronic mRNA have been rarely found so far. The second cistron of constructed bicistronic mRNAs is generally considered as not translated unless special sequences named internal ribosome entry site (IRES) are added between the two cistrons. These sequences are believed to recruit ribosomes independently of a cap structure. In the present report, a new IRES found in the HTLV-1 genome is described. A systematic study revealed that this IRES, but also the poliovirus (polio) and the encephalomyocarditis virus (EMCV) IRES work optimally when they are added about 100 nucleotides after the termination codon of the first cistron. Unexpectedly, these IRES became totally inefficient when added after 300-500 nucleotide spacers. This result and others are not compatible with the admitted mechanism of IRES action. The IRES appear to be rather potent translation stimulators. Their effects are particularly emphasized in cells in which the normal mechanism of translation initiation is inhibited. For these reasons, we suggest to call IRES rescue translation stimulators (RTS).

Animals↗

The Use of Deep Learning in RNA Therapeutic Development.

Ribonucleic acid (RNA)-based therapeutics have emerged as promising methods of disease treatment due to their ability to target the human genome and influence protein production, their versatility, and their relative lack of toxicity compared to other gene therapies. However, the RNA therapeutic design space is extremely large, encompassing multiple variables, including codon identities, secondary structure, and design of specific regions. RNA therapeutic optimization is difficult due to the impracticality of exploring such a vast design space experimentally. To address this limitation, deep learning methods have been employed to optimize RNA therapeutic development. In this review, we examine the application of deep learning models across three key aspects of RNA therapeutic development (RNA structure prediction, CRISPR activity, and RNA delivery), highlighting major contributions in these fields and analyzing how deep learning model architectures could affect model performance. We then discuss challenges associated with using deep learning for RNA therapeutics, such as computational and data limitations. Finally, we offer perspectives on areas for future exploration, such as emerging model architectures and methods of integration with more advanced high-throughput screening techniques. Ultimately, this review provides an overview of how deep learning is used in RNA therapeutic development and how it can evolve in the future.

Deep Learning↗

Translational selection and yeast proteome evolution.

The primary structures of peptides may be adapted for efficient synthesis as well as proper function. Here, the Saccharomyces cerevisiae genome sequence, DNA microarray expression data, tRNA gene numbers, and functional categorizations of proteins are employed to determine whether the amino acid composition of peptides reflects natural selection to optimize the speed and accuracy of translation. Strong relationships between synonymous codon usage bias and estimates of transcript abundance suggest that DNA array data serve as adequate predictors of translation rates. Amino acid usage also shows striking relationships with expression levels. Stronger correlations between tRNA concentrations and amino acid abundances among highly expressed proteins than among less abundant proteins support adaptation of both tRNA abundances and amino acid usage to enhance the speed and accuracy of protein synthesis. Natural selection for efficient synthesis appears to also favor shorter proteins as a function of their expression levels. Comparisons restricted to proteins within functional classes are employed to control for differences in amino acid composition and protein size that reflect differences in the functional requirements of proteins expressed at different levels.

Adaptation, Physiological↗

[Overexpression of a sweet protein monellin in Escherichia coli].

According to the amino acid sequence of monellin, a single chain 294bp monellin gene was synthesized and inserted into vector pET-22b to yield the recombinant secretion plasmid pETMO. The single-chain monellin gene was designed based on the biased codons of E. coli so that its expression would be then optimized. Under the expressing conditions, monellin was produced accounting for 44.8% of total soluble proteins. The E. coli-expressed single-chain monellin is 3000 times sweeter than sucrose. The thermal-stability and acid-resistance of the protein are higher than the natural monellin.

Escherichia coli↗

Contribution of trans-splicing, 5' -leader length, cap-poly(A) synergism, and initiation factors to nematode translation in an Ascaris suum embryo cell-free system.

Trans-splicing introduces a common 5' 22-nucleotide sequence with an N-2,2,7-trimethylguanosine cap (m (2,2,7)(3)GpppG or TMG-cap) to more than 70% of transcripts in the nematodes Caenorhabditis elegans and Ascaris suum. Using an Ascaris embryo cell-free translation system, we found that the TMG-cap and spliced leader sequence synergistically collaborate to promote efficient translation, whereas addition of either a TMG-cap or spliced leader sequence alone decreased reporter activity. We cloned an A. suum embryo eIF4E homolog and demonstrate that this recombinant protein can bind m(7)G- and TMG-capped mRNAs in cross-linking assays and that binding is enhanced by eIF4G. Both the cap structure and the spliced leader (SL) sequence affect levels of A. suum eIF4E cross-linking to mRNA. Furthermore, the differential binding of eIF4E to a TMG-cap and to trans-spliced and non-trans-spliced RNAs is commensurate with the translational activity of reporter RNAs observed in the cell-free extract. Together, these binding data and translation assays with competitor cap analogs suggest that A. suum eIF4E-3 activity may be sufficient to mediate translation of both trans-spliced and non-trans-spliced mRNAs. Bioinformatic analyses demonstrate the SL sequence tends to trans-splice close to the start codon in a diversity of nematodes. This evolutionary conservation is functionally reflected in the optimal SL to AUG distance for reporter mRNA translation in the cell-free system. Therefore, trans-splicing of the SL1 leader sequence may serve at least two functions in nematodes, generation of an optimal 5'-untranslated region length and a specific sequence context (SL1) for optimal translation of trimethylguanosine capped transcripts.

5' Untranslated Regions↗