Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Codon usage pattern in alpha 2(I) chain domain of chicken type I collagen and its implications for the secondary structure of the mRNA and the synthesis pauses of the collagen.

A stability map of local secondary structure of the mRNA of the triple-helical alpha 2(I) chain domain of chicken type I collagen was obtained by plotting the free energy of the optimal secondary structure of a local segment in mRNA against the segment position along a base sequence of the mRNA. It was found that the positions of the minima of free energy in the plot coincide with the positions where synthesis pauses of the alpha-chain polypeptides of the corresponding sizes translated from the mRNA have been reported to occur (1). The codon usage pattern of each of the three major amino acids of the alpha-chain domain of the collagen, Gly, Pro and Ala, fluctuates considerably along the base sequence segments of the mRNA and a deviation of the pattern from that of the average of the whole alpha 2(I) chain domain mRNA, particularly for Gly codons, leads to a loss of the stability of the local secondary structure of the mRNA. The results suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA of the chicken collagen alpha 2(I) chain domain which leads to a nonuniform polypeptide elongation pattern.

Animals

[Physico-chemical basis of the genetic code origin: stereochemical analysis of interactions of amino acids and nucleotides based on the progene hypothesis].

A progene hypothesis has been proposed earlier to explain the mechanism of origin of the self-reproducing genetic system. Progenes (precursors of the genetic system) are mixed anhydrides of an amino acid and deoxyribotrinucleotide at the 3'-gamma-terminal phosphate (NpNpNppp-AA); they are produced from dinucleotides (NpNp) and 3'-gamma-aminoacylnucleotidylates (Nppp-AA) as a result of specific interaction between amino acid and dinucleotide. The postulated mechanism of progene formation accounts for the selection of substances, including chirality, the origin of the genetic code as well as for the mechanisms of formation, self-reproduction and evolution of the simpliest genetic system ("gene--polypeptide"). A stereochemical analysis of the progene formation mechanism has allowed us to support the main statements of the hypothesis that relate to the origin of the genetic code and to selection of substances. Atomic groups that could be responsible for the specificity of interaction between dinucleotides and amino acids in progene formation have been revealed. Stereochemical evidence for the physicochemical basis of the origin of the existing genetic code have been produced: 1) a special role of the second nucleotide in the codon is demonstrated in amino acid coding by the progene hypothesis principle; 2) an advantage of T against U in such coding is demonstrated; 3) for 16 amino acids out of 20 an agreement has been obtained between the optimal dinucleotide as revealed by the stereochemical analysis and the codon dinucleotides; 4) an explanation for the third nucleotide selection mechanism is offered. A restoration of the prebiotic code, based on these results, has indicated that the code contains 32 codons, is statistical and group-wise. It encodes 7 groups of isofunctional amino acids: 3 overlapping groups of non-polar amino acids 1) medium-size hydrophobic amino acids (chiefly Val, n-Val and a-But), 2) small and medium-size non-polar amino acids (chiefly Ala Val, n-Val a-But and Gly), 3) small non-polar amino acids (Gly, Ala, a-But) and 4 groups of polar amino acids--1) hydroxy--+dicarbonic (Asp, Glu, Ser and Thr), 2) dicarbonic (Asp and Glu), 3) hydroxy (Ser and Thr) and 4) basic (Arg and Lys). The code includes about 20 amino acids among which are 15-17 canonical and a few common non-canonical. The prebiotic code explains many properties of the existing genetic code and is capable of evolving into the latter by way of a gradual replacement of the physicochemical coding mechanism by the enzymatic coding mechanism.

Amino Acids

[tRNA adaptation and the optimization of translation].

The intracellular level of each tRNA species is adjusted to the codon frequency of the mRNA being decoded. This was first observed in such highly differentiated cells as the silk gland of Bombyx mori, which produces fibroin and sericin, and the rabbit reticulocyte. tRNA adaptation also occurs in other cell types from E. coli to mammalian cells. Regardless of the mechanism regulating tRNA biosynthesis, we believe that tRNA adaptation is the basic step optimizing chain elongation at the ribosomal level. We propose the system of trial and error as a working model for the ribosome. This model clarifies the correlations between iso-accepting tRNA levels and codon frequencies, as well as the effect of tRNA pool balance on mean elongation rate and non-uniform individual elongation rate (depending on whether codons are rare or abundant) for fibroin mRNA translated in a reticulocyte cell-free system.

Adaptation, Physiological

Modified bacteriophage lambda promoter vectors for overproduction of proteins in Escherichia coli.

A new series of expression vectors that direct high-level overproduction of gene products in Escherichia coli is described. All contain strong bacteriophage lambda promoters, PR and PL, arranged in tandem so that both promote transcription into genes inserted into or between unique restriction sites. The vectors also direct expression of the lambda cI857 gene (from its natural promoter, PM), which enables their use in any E. coli host strain to effect controlled expression by shifting the temperature of cultures from 30 to 42 degrees C. The vectors pCE30, pND201, pPT150 and pMA200U are derivatives of the high-copy-number plasmid pUC9. Vector pCE33 is an analogous derivative of the heat-inducible runaway-replication plasmid, pMOB45, and directs overproduction of proteins by virtue of increase in both gene dosage and transcription following treatment at 42 degrees C. The vectors pND201 and pPT150 bear a ribosome-binding site (RBS) perfectly complementary to the 3' end of E. coli 16-S rRNA a few bp upstream from a unique HpaI site. Ways in which they may be used to improve the efficiency of translation of mRNA by substitution of a natural RBS with selection for optimal spacing from an ATG (or GTG) start codon are described. The phagemid vector pMA200U is a direct analog of pCE30 designed to facilitate preparation of single-stranded DNA templates for use in oligodeoxyribonucleotide-directed mutagenesis of overexpressed genes.

Bacteriophage lambda

A hardware interpretation of the evolution of the genetic code.

A quantitative rationale for the evolution of the genetic code is developed considering the principle of minimal hardware. This principle defines an optimal code as one that minimizes for a given amount of information encoded, the product of the number of physical devices used by the average complexity of each device. By identifying the number of different amino acids, number of nucleotide positions per codon and number of base types that can occupy each such position with, respectively, the amount of information, number of devices and the complexity, we show that optimal codes occur for 3, 7 and 20 amino acids with codons having a single, two and three base positions per codon, respectively. The advantage of a code of exactly 4 symbols is deduced, as well as a plausible evolutionary pathway from a code of doublets to triplets. The present day code of 20 amino acids encoded by 64 codons is shown to be the most optimal in an absolute sense. Using a tetraplet code further evolution to a code in which there would be 55 amino acids is in principle possible, but such a code would deviate slightly more than the present day code from the minimal hardware configuration. The change from a triplet code to a tetraplet code would occur at about 32 amino acids. Our conclusions are independent of, but consistent with, the observed physico-chemical properties of the amino acids and codon structures. These correlations could have evolved within the constrains imposed by the minimal hardware principle.

Amino Acids

Streptomycin causes misreading of natural messenger by interacting with ribosomes after initiation.

The induction of misreading by streptomycin in vitro, previously observed with synthetic messengers, is now demonstrated with natural (endogenous or viral) messenger by the use of extracts of temperature sensitive mutants lacking Glu--tRNA or Val--tRNA synthetase. With chain-elongating but noninitiating ribosomes (i.e., purified polysomes) deprived of an aminoacyl--tRNA, streptomycin and other aminoglycosides, over a wide range of concentrations, stimulate incorporation. With ribosomes initiating in the presence of streptomycin stimulation is also observed but it is restricted, just like phenotypic suppression in cells, to very low streptomycin concentrattions which evidently allow some ribosomes to initiate and later encounter them in the course of chain elongation. The stimulation is accompanied by an increase in the size of the products; hence, it is evidently due to substitution of an incorrect aminoacyl--tRNA for a missing one. The test introduced here also has revealed a misreading effect of streptomycin on resistant ribosomes. In addition, significant intrinsic misreading was observed without streptomycin, indicating that under optimal conditions for in vitro protein synthesis an empty codon is frequently read by an incorrect aminoacyl--tRNA.

Anti-Bacterial Agents

Co-expression of a precursor and the mature protein of wheat ribulose-1,5-bisphosphate carboxylase small subunit from a single gene in Escherichia coli.

The cDNA encoding a precursor of wheat ribulose-1,5-bisphosphate carboxylase/oxygenase was inserted in-phase with prokaryotic expression elements in four different vectors. Five expression vectors encoding the small subunit precursors were cloned in Escherichia coli. None of these constructs expressed detectable amounts of the precursor protein, but all directed synthesis of the mature small subunit. The expression of the small subunit was a consequence of an independent, intragenic Shine-Dalgarno sequence optimally located upstream from an ATG specifying the first codon of the mature small subunit portion in the precursor transcript. Similar internal translation signals have been identified in the nuclear-encoded cDNAs of the small-subunit precursors of numerous higher plant genes. The 5' end of the wheat small-subunit precursor was linked with a consensus E. coli DNA sequence such that the modified gene encoded a partial hybrid precursor carrying four additional residues at its amino terminus. The resultant construct, pEI-W3, directed abundant synthesis of both the partially hybrid small-subunit precursor and the mature small subunit, constituting as much as 10% of the total bacterial protein. The bacterially synthesized small subunit precursor was purified to homogeneity. The authenticity of the recombinant protein was verified by its size, immunological properties, amino-terminal sequence, and amino acid composition.

Amino Acid Sequence

Expression of recombinant growth hormone in Escherichia coli: effect of the region between the Shine-Dalgarno sequence and the ATG initiation codon.

We constructed a synthetic Escherichia coli expression system in which various promoter elements can be changed easily. In this study we investigated the effect of a number of portable Shine-Dalgarno regions (SD regions) on the synthesis of two modified recombinant human growth hormones (hGH). The production of these modified hGH was measured during exponential growth and after the bacteria had reached stationary phase. The results show that the optimal distance between the SD region (AGGAGG) and the ATG start codon is approximately 11 nucleotides. However, the nucleotide sequence in this region also influences expression: 6-10 adenines result in comparable expression levels despite the varying lengths. Two overlapping SD regions reduce expression of the growth hormones considerably, whereas two potential ATG start codons do not affect expression. Having a SD-ATG region partly or totally complementary to the 5' end of the 16S ribosomal RNA does not alter translation efficiency. Estimation of the delta G values for the association between the 16S rRNA and the ribosome-binding region suggests that these are not indicators of expression efficiency.

Base Sequence

Codon bias and gene expression.

The frequencies with which individual synonymous codons are used to code their cognate amino acids is quite variable from genome to genome and within genomes, from gene to gene. One particularly well documented codon bias is that associated with highly expressed genes in bacteria as well as in yeast; this is the so-called major codon bias. Here, it is suggested that the major codon bias is not an arrangement for regulating individual gene expression. Instead, the data suggest that this codon bias, which is correlated with a corresponding bias of tRNA abundance, is a global arrangement for optimizing the growth efficiency of cells. On the practical side, it is suggested that heterologous gene expression is not as sensitive to codon bias as previously thought, but that it is quite sensitive to other characteristics of the heterologous gene.

Codon

Stoichiometry of GTP hydrolysis in a poly(U)-dependent cell-free translation system. Determination of GTP/peptide bond ratios during codon-specific elongation and misreading.

The stoichiometry of GTP hydrolysis during peptide elongation in the processes of codon-specific translation and misreading of polyuridylic acid was determined in a cell-free system in which all ribosomes were active in peptide synthesis. Ribosomes carrying oligophenylalanine presynthesized on poly(U) covalently bound to Sepharose were used. In the codon-specific translation of poly(Phe) on poly(U)-Sepharose at optimal Mg2+ concentration (6 mM MgCl2), the ratio of GTP cleaved to Phe polymerized was found to be about 2 (+/- 0.1). Under the same conditions but during misreading (elongation of polyleucine on poly(U)-Sepharose) the GTP/Leu ratio increased 10 times (from 16 to 25 in different experiments).

Codon

Presence of the hypermodified nucleotide N6-(delta 2-isopentenyl)-2-methylthioadenosine prevents codon misreading by Escherichia coli phenylalanyl-transfer RNA.

The overall structure of transfer RNA is optimized for its various functions by a series of unique post-transcriptional nucleotide modifications. Since many of these modifications are conserved from prokaryotes through higher eukaryotes, it has been proposed that most modified nucleotides serve to optimize the ability of the tRNA to accurately interact with other components of the protein synthesizing machinery. When a cloned synthetic Escherichia coli tRNAPhe gene was transfected into a bacterial host that carried a defective phenylalanine tRNA-synthetase gene, tRNAPhe was overexpressed by 11-fold. As a result of this overexpression, an undermodified tRNAPhe species was produced that lacked only N6-(delta 2-isopentenyl)-2-methylthioadenosine (ms2i6A), a hypermodified nucleotide found immediately 3' to the anticodon of all major E. coli tRNAs that read UNN codons. To investigate the role of ms2i6A in E. coli tRNA, we compared the aminoacylation kinetics and in vitro codon-reading properties of the ms2i6A-lacking and normal fully modified tRNAPhe species. The results of these experiments indicate that while ms2i6A is not required for normal aminoacylation of tRNAPhe, its presence stabilizes codon-anticodon interaction and thereby prevents misreading of the genetic code.

Adenosine

Detecting evolutionary trends from molecular data. 1. Some measures of compositional nonrandomness.

The measures of compositional nonrandomness to be discussed as to their physical significance and to their power of detecting evolutionary significant variations are (see article)(pi a priori probability for amino acid i, ni its number of occurrences in a protein of length L). As a concrete example, the pi are here supposed to represent equal frequencies of all non-stop codons. For each quantity, four levels are defined: The base level, with optimal (i.e. minimal nonrandomness) composition, admitting non-integer values of ni; the integer level with optimal integer composition; the noise level, represented by a typical random cain; and the real protein level. On all these levels, S, which is the measure with the most direct physical sense, shows the smoothest behavior with the smallest relative fluctuations and thus the highest resolution.

Albumins

The influence of ribosome-binding-site elements on translational efficiency in Bacillus subtilis and Escherichia coli in vivo.

A method is described to determine simultaneously the effect of any changes in the ribosome-binding site (RBS) of mRNA on translational efficiency in Bacillus subtilis and Escherichia coli in vivo. The approach was used to analyse systematically the influence of spacing between the Shine-Dalgarno sequence and the initiation codon, the three different initiation codons, and RBS secondary structure on translational yields in the two organisms. Both B. subtilis and E. coli exhibited similar spacing optima of 7-9 nucleotides. However, B. subtilis translated messages with spacings shorter than optimal much less efficiently than E. coli. In both organisms, AUG was the preferred initiation codon by two- to threefold. In E. coli GUG was slightly better than UUG while in B. subtilis UUG was better than GUG. The degree of emphasis placed on initiation codon type, as measured by translational yield, was dependent on the strength of the Shine-Dalgarno interaction in both organisms. B. subtilis was also much less able to tolerate secondary structure in the RBS than E. coli. While significant differences were found between the two organisms in the effect of specific RBS elements on translation, other mRNA components in addition to those elements tested appear to be responsible, in part, for translational species specificity. The approach described provides a rapid and systematic means of elucidating such additional determinants.

Bacillus subtilis

Translation of chloroplast-encoded mRNA: potential initiation and termination signals.

A survey of 196 protein-coding chloroplast DNA sequences demonstrated the preference for AUG and UAA codons for initiation and termination of translation, respectively. As in prokaryotes at every nucleotide position from -25 to +25 (AUG is +1 to +3) and for 25 nucleotides 5' and 3' to the termination codon an A or U is predominant, except for C at +5 and G at +22. A Shine-Dalgarno (SD) sequence (GGAGG or tri- or tetranucleotide variant) was found within 100 bp 5' to the AUG codon in 92% of the genes. In 40% of these cases, the location of the SD sequence was similar to that of the consensus for prokaryotes (-12 to -7 5' to AUG), presumed to be optimal for translation initiation. A SD sequence could not be located in 6% of the chloroplast sequences. We propose that mRNA secondary structures may be required for the relocation of a distal SD sequences to within the optimal region (-12 to -7) for initiation of translation. We further suggest that termination at UGA codons in chloroplast genes may occur by a mechanism, involving 16S rRNA secondary structure, which has been proposed for UGA termination in E. coli.

Base Composition

The Use of Deep Learning in RNA Therapeutic Development.

Ribonucleic acid (RNA)-based therapeutics have emerged as promising methods of disease treatment due to their ability to target the human genome and influence protein production, their versatility, and their relative lack of toxicity compared to other gene therapies. However, the RNA therapeutic design space is extremely large, encompassing multiple variables, including codon identities, secondary structure, and design of specific regions. RNA therapeutic optimization is difficult due to the impracticality of exploring such a vast design space experimentally. To address this limitation, deep learning methods have been employed to optimize RNA therapeutic development. In this review, we examine the application of deep learning models across three key aspects of RNA therapeutic development (RNA structure prediction, CRISPR activity, and RNA delivery), highlighting major contributions in these fields and analyzing how deep learning model architectures could affect model performance. We then discuss challenges associated with using deep learning for RNA therapeutics, such as computational and data limitations. Finally, we offer perspectives on areas for future exploration, such as emerging model architectures and methods of integration with more advanced high-throughput screening techniques. Ultimately, this review provides an overview of how deep learning is used in RNA therapeutic development and how it can evolve in the future.

Deep Learning

Amino acid composition is correlated with protein abundance in Escherichia coli: can this be due to optimization of translational efficiency?

Amino acid occurrence frequencies were found for four groups of Escherichia coli proteins with different abundance levels in the cell. These frequencies decrease with increasing protein abundance for amino acids whose codons are translated by tRNAs present at low concentrations (e.g., Cys, Trp, Ser, etc.); the opposite tendency was observed for amino acids translated by abundant tRNAs (Lys, Val, etc.). The efficiency (rate and accuracy) of codon translation is expected to be proportional to the concentration of the cognate tRNA. Therefore, the observed constraints on amino acid composition may be explained as resulting from evolutionary pressure optimizing the translational efficiency of a gene (the same pressure is responsible for the nonrandom choice of synonymous codons).

Amino Acids

[Effectiveness of translation coupling in hybrid operons].

The possibility of creating artificial overlappons was studied on the model of two genes, that coding for the N-terminal part of lambda cro protein and the cat of E. coli. To test the dependence of translational coupling efficiency on the intercistronic region a series of recombinant DNA molecules carrying different hybrid operons with partially overlapping genes was constructed. The translational efficiency of the distal to the promoter gene was shown to depend on the intercistronic region structure: overlapping of the AUG codon with the terminating one of the proximal gene in the UGAUG manner is optimal for the translational coupling, and the displacement of AUG at several nucleotides in both directions decreases the translational reinitiation efficiency for the ribosomes, that have synthesised the first gene product.

Acetyltransferases