Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Construction and expression of nonsense suppressor tRNAs which function in plant cells.

An Arabidopsis thaliana L. DNA containing the tRNA(TrpUGG) gene was isolated and altered to encode the amber suppressor tRNA(TrpUAG) or the ochre suppressor tRNA(TrpUAA). These DNAs were electroporated into carrot protoplasts and tRNA expression was demonstrated by the translational suppression of amber and ochre nonsense mutations in the chloramphenicol acetyltransferase (CAT) reporter gene. DNAs encoding tRNA(TrpUAG) and tRNA(TrpUAA) nonsense suppressor tRNAs caused suppression of their cognate nonsense codons in CAT mRNAs, with the tRNA(TrpUAG) gene exhibiting the greater suppression under optimal conditions for expression of CAT. The development of these translational suppressors which function in plant cells facilitates the study of plant tRNA gene expression and will make possible the manipulation of plant protein structure and function.

Anticodon

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins

Possibility of genetic coding of amino acid sequences by coherent electronic states in nucleotide chains.

The concept of coherent electronic states and coherent interactions in supramolecular structures is applied to the process of genetic information coding and its transcription from DNA to mRNA. A new genetic code is proposed based on the assumption of coherent electron states in linear chains of nucleotide bases. A new interpretation of codon equivalency (redundancy) is given. The number of existing amino acids is derived from the optimalization principle applied to the physical system storing the genetic information in the new code. The proposed code uses a variable number of positions or nucleotide bases along the DNA-mRNA structure to code a single amino acid in a protein. The average of this variable number must be equal to the base of natural logarithms (e = 2.7 . . .) in order to minimize the number of nucleotides required to code a sequence of amino acids.

Amino Acid Sequence

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing

cDNA cloning and functional expression in yeast Saccharomyces cerevisiae of beta-naphthoflavone-induced rabbit liver P-450 LM4 and LM6.

A cDNA library was constructed from liver mRNA of a beta-naphthoflavone-induced rabbit. Two clones pLM4-1 and pLM6-1 containing 2.2-kbp inserts that hybridized at low stringincy with a mouse P1 P-450 probe were selected. The clone pLM4-1 was fully sequenced and found to contain a full-length cDNA coding for cytochrome P-450 LM4. Partial sequence and restriction mapping made it possible to identify pLM6-1 as coding for the major part of cytochrome P-450 LM6. Cloned LM4-1 cDNA was reformed by deletion of the 5' and 3' non-coding regions before insertion into yeast expression vectors PYe DP1/10. A similar operation was performed on pLM6-1 cDNA after replacement of the missing N-terminus-coding sequences by homologous sequences form the pLM4-1 clone resulting in a chimeric cytochrome P-450 coding sequence. Expression of cloned rabbit cytochrome P-450 into transformed yeast was optimized by studying the effect of the nature of the DNA sequence just preceding the initiation codon on the level of cytochrome P-450 production. Yeast synthesized cytochromes P-450 were characterized by immunoblotting, spectra and catalytic activity determinations. Cloned cytochrome P-450 LM4 was found by all criteria to be identical to the authentic rabbit one. The chimeric cytochrome P-450 that contains the 143 N-terminal amino acids of cytochrome P-450 LM4 and the remaining 375 amino acids of cytochrome P-450 LM6 was found to exhibit most of the authentic cytochrome P-450 LM6 catalytic properties. Enzymatic and evolutionary implications of these results are discussed.

Amino Acid Sequence

Mutations affecting translational coupling between the rep genes of an IncB miniplasmid.

The nature of translational coupling between repB and repA, the overlapping rep genes of the IncB plasmid pMU720, was examined. Mutations in the start codon of the promoter proximal gene, repB, reduced the efficiency of translation of both rep genes. Moreover, there was no independent initiation of repA translation in the absence of repB translation. The position of the repB stop codon was crucial for the efficient expression of repA, with the wild-type positioning being optimal. Translational coupling was found to be totally dependent on the formation of a pseudoknot structure. A model which invokes formation of a pseudoknot to facilitate initiation of repA is proposed.

Bacterial Proteins

Characterization of an alpha 1----3-galactosyltransferase homologue on human chromosome 12 that is organized as a processed pseudogene.

UDP-Gal:Gal beta 1----4GlcNAc alpha 1----3-galactosyltransferase is a terminal glycosyltransferase that is widely expressed in a variety of mammalian species, with the notable exception of man, apes, and Old World monkeys. We recently reported the isolation of a bovine cDNA clone that contains the complete coding sequence for this enzyme (Joziasse, D. H., Shaper, J. H., Van den Eijnden, D. H., Van Tunen, A. J., and Shaper, N. L. (1989) J. Biol. Chem. 264, 14290-14297). Using this cDNA as a probe, we have demonstrated that, although transcripts cannot be detected in a variety of established human cell lines by Northern blot analysis, homologous sequences are present in human genomic DNA. To establish that these sequences represent a human homologue of alpha 1----3-galactosyltransferase, we have used the bovine cDNA as a probe to isolate two nonoverlapping clones (HGT-2 and HGT-10) from a human genomic DNA library. Clone HGT-2 contains a 1.5-kilobase uninterrupted linear sequence similar to bovine alpha 1----3-galactosyltransferase that is organized as a processed pseudogene. This sequence, flanked by Alu type repeats, contains a short 5'- and 3'-untranslated region and a complete recognizable coding region that is 81% similar at the nucleotide level to bovine alpha 1----3-galactosyltransferase. This putative coding region contains multiple frameshift mutations and nonsense codons in all three reading frames which precludes the synthesis of a functional enzyme. Nevertheless, after optimal alignment, translation predicts a polypeptide that is 68% similar at the amino acid level to the bovine enzyme. Based on Southern analysis and limited sequence analysis, clone HGT-10 contains coding sequences similar to the NH2-terminal region of bovine alpha 1----3-galactosyltransferase. By analysis of panels of human-rodent somatic cell hybrids we have established that the nonfunctional, processed pseudogene and the human homologue represented by HGT-10 are located on human chromosomes 12 and 9, respectively. Interestingly, a comparison of the predicted amino acid sequence of the carboxyl-terminal two-thirds of human alpha 1----3-galactosyltransferase, with the corresponding region of the human blood group A, UDP-GalNAc:[Fuc alpha 1----2]Gal beta 1----4GlcNAc alpha 1----3-GalNAc-transferase (Yamamoto, F., Marken, J., Tsuji, T., White, T., Clausen, H., and Hakomori, S. (1990a) J. Biol. Chem. 265, 1146-1151), reveals a significant similarity (39%) suggesting that these two enzymes may have arisen from the same ancestral gene as a result of gene duplication and subsequent divergence.

Amino Acid Sequence

DNA sequence of the Escherichia coli gene, gnd, for 6-phosphogluconate dehydrogenase.

Expression of gnd of Escherichia coli, which encodes 6-phosphogluconate dehydrogenase, an enzyme of the hexose monophosphate shunt, is subject to growth rate-dependent regulation and is gene dosage-dependent: the level of the enzyme increases in direct proportion to the cellular growth rate at both low and high gene copy numbers. We have determined the nucleotide sequence of gnd and flanking control regions, the 5'-end of in vivo gnd mRNA, and the start codon of the structural gene. Analysis of the sequence indicated that: (i) the gnd promoter is typical of other E. coli promoters and the structural gene is followed by a rho-independent transcription termination signal; (ii) the 56-nucleotide leader of gnd mRNA does not contain a rho-independent transcription termination signal, so growth rate-dependent regulation of 6-phosphogluconate dehydrogenase level is not carried out by an attenuation mechanism analogous to the one that controls expression of the E. coli ampC gene; (iii) the codon composition of the structural gene resembles that of other highly expressed E. coli genes and thus is not responsible for the regulation either; (iv) the structural gene is preceded at an optimal distance by a strong Shine-Dalgarno (SD) sequence, AGGAG ; (v) the leader region of the mRNA contains regions of dyad symmetry that have the potential to sequester the SD sequence and the start codon. This latter feature of the gene suggests that growth rate-dependent regulation may involve regulation of translation initiation frequency.

Amino Acid Sequence

Expression of human asparagine synthetase in Saccharomyces cerevisiae.

Human asparagine synthetase was expressed in the yeast Saccharomyces cerevisiae. The identity of the expressed protein was confirmed by immunoblotting and in vitro enzymatic activity. The recombinant enzyme was shown to have both the ammonia- and glutamine-dependent asparagine synthetase activity in vitro. In contrast to overproduction in Escherichia coli, the expressed protein was found to be soluble in the yeast cell. Furthermore, expression in yeast made it possible to isolate non-degraded human asparagine synthetase which had also the N-terminal methionine correctly processed. The yeast expression plasmid was constructed for optimal production of the recombinant enzyme. In addition, unique restriction enzyme sites that bracket the first five codons of the human asparagine synthetase gene were introduced. This will allow the use of oligonucleotide cassette mutagenesis to investigate the role of the N-terminal amino acids in asparagine synthetase enzymatic activity.

Amino Acid Sequence

Nucleotide sequence and structural analysis of the rat RT1.Eu and RT1.Aw3l genes, and of genes related to RT1.O and RT1.C.

A cDNA library was constructed using mRNA isolated from the R21 strain of rats which have the major histocompatibility complex (MHC) haplotype RT1.AlBlDlEu and the growth and reproduction complex (grc) genotype grc+. The cDNA clones that hybridized with the class I probes pAG64c and pARI.5 and were 1.3-1.7 kilobases were selected. Full-length clones were identified by sequencing partially the 5' and 3' ends of each clone, by the presence of a start codon at the 5' end, and by a polyadenylation sequence at the 3' end. The full-length cDNA clones were examined for in vitro transcription by transfection into human CIR cells using electroporation, and expression was detected by flow cytometry using monoclonal antibodies specific to the heavy chains and polyclonal antibody to beta 2-microglobulin. The RT1.Eu gene was transcribed and expressed optimally, and its nucleotide and deduced amino acid sequences differed significantly from the RT1.Aa, RT1.A(l), RT.Au, LW2, and 11/3R genes but only slightly from the RT1.K gene. The high level of sequence similarity between RT1.Eu and RT1.K suggests that the two genes may have originated from a common ancestral gene. In addition, three new genes (RT1.Aw3l, RT1.C-type, and RT1.O-type) were identified. The RT1.Aw3l gene is almost identical to RT1.A(l) with the exception of an in frame deletion of 21 nucleotides in exon 2 leading to a 7 amino acid deletion in the alpha 1 domain of the deduced amino acid sequence and 11 nucleotide substitutions and insertions in the rest of the sequence. It transcribed optimally, but no significant expression was detected. The RT1.C-type gene 119 is very similar (97%) to the LW2 gene in the 3' untranslated region, which suggests that it is in the RT1.C region. It transcribed optimally, but no significant expression was detected. The RT1.O-type gene 149 has all the features of a class Ib gene, but a premature stop codon in the alpha 1 domain causes incomplete translation. Its in vitro transcription was very low, and no expression was detected. These studies, combined with previous work, indicate that in the MHC of the R21 strain three class Ia genes (Eu, A(l), Aw3l) and three class Ib genes (C-type, O-type, N) are transcribed but only two class Ia genes (Eu, A(l)) are expressed.

Amino Acid Sequence

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment

Restructuring the translation initiation region of the human parathyroid hormone gene for improved expression in Escherichia coli.

Overexpression of native human parathyroid hormone in Escherichia coli was achieved by a modification of the 5' end of the genomic gene sequence, thereby adapting this part of the translation initiation region to the bacterial host. Some simple rules abstracted from optimization studies of translation initiation of a beta-interferon gene were applied. These included (a) extending complementarity of the mRNA to the anticodon loop of tRNAfMet by use of a codon with a purine nucleotide directly following the ATG, (b) avoidance of stable secondary structure in the mRNA by use of synonymous A/U-rich codons, (c) elimination of a potential second Shine-Dalgarno sequence. The appropriate silent changes led to a 20-fold increase in parathyroid hormone production resulting in 4.3% of total soluble protein. This result proves the validity of our simple approach for optimization of foreign gene expression in E. coli.

Base Sequence

Inducible expression vectors incorporating the Escherichia coli atpE translational initiation region.

New expression vectors were constructed for use in strains of Escherichia coli. Their most important feature is a polylinker system that facilitates the insertion of a gene in an optimal relationship to the highly efficient E. coli atpE translational initiation region (from nucleotide -50 to the start codon). Three ATG-containing restriction endonuclease sites can be used for the insertion of the 5' end of a gene at, or near to, its translational initiation codon. These sites may alternatively be used for the creation of a suitable translational start codon. Transcription is started by the bacteriophage lambda major promoters pR and pL in tandem and terminated by the bacteriophage fd terminator. Transcriptional initiation is very effectively repressed at 28-30 degrees C by the product of the bacteriophage lambda cIts857 gene, which is also present on the vectors. Full induction is achieved by shifting the incubation temperature to 42 degrees C. The combination of highly efficient transcriptional and translational signals on these vectors allowed high-level expression of sequences encoding human interferon beta and interleukin 2 and of the E. coli atpA, sucC and sucD genes.

DNA Restriction Enzymes

On the information content of the genetic code.

In living organisms 20 amino acids along with the terminator value(s) are encoded by 64 codons giving a degeneracy of the codons as described by the genetic code. A basic theoretical problem of genetic codes is to explain the particular distribution of degeneracies of partitions involved in the codes. In this work the degeneracy problem is considered in the framework of information theory. It is shown by direct numerical evaluation of a certain degeneracy information function associated with the genetic code that the degeneracy of the codes is observed to be related to the optimization of this function.

Amino Acids

Regulation of T-cell antigen receptor (TCR) alpha-chain expression by TCR beta-chain transcripts.

The TCR is an alpha beta heterodimer, a part of the multimeric structure through which physiological T-cell activation occurs. The expression of TCR alpha chain is greatly diminished in a beta-chain-deficient mutant Jurkat cell line (J.RT3-T3.5). The relationship between the expression of the TCR alpha and beta chains has been examined by stable transfection of a series of TCR beta-chain mutant constructs into this mutant cell line. The level of alpha-chain transcript was dramatically upregulated by the expression of the beta chain and specifically by a transcript of the beta-chain variable region alone, including a transcript in which the ATG start codon was mutated. The downregulation of the endogenous alpha-chain transcripts in mutants cells lacking complete beta-chain transcripts occurred primarily at the posttranscriptional level. This evidence for a regulatory function of the TCR beta-chain gene represents an unusual regulatory pathway in which the transcript of one gene is required for the optimal expression of another gene.

Amino Acid Sequence

[Origin and evolution of the genetic code].

We propose a quantitative model which suggests that the present genetic code appeared under the influence of mutations, while optimizing its own resistance against their effects. Its evolution was realized by successive steps in which the number of translated codons grew, whereas the number of terminators decreased. The main constraint of this model is selection against nonsense mutations: the competition among many primitive codes gives advantage to those which resist best to the occurrence of nonsense mutations. The structures of the selected systems converge towards that of the present genetic code. This one appears then as built so as to resist to errors, information noise, mutations.

Biological Evolution

Optimization of the synthesis of porcine somatotropin in Escherichia coli.

We report on the influence of choice of promoter and RNA polymerase, 5'-untranslated regions and ribosome binding sites, codon usage, leader peptide coding sequences and poly A tail in the 3'-untranslated region on the synthesis of porcine somatotropin (PST) in Escherichia coli. A total of 12 different constructs were tested in this study for the production of porcine somatotropin (PST) in E. coli. Several factors have significant effects on PST synthesis. In the presence of a strong promoter and a strong ribosome binding site, the next most important factor seems to be the combination of sequences at the 5'-end of the mRNA including both the 5'-untranslated region and the start of the coding sequence. Codon usage in the 5'-coding sequence per se is not important in determining the level of PST synthesis where high level expression is achieved from a strong ribosome binding site. However, where low level synthesis of recombinant PST (rPST) is achieved, codon usage in the 5'-coding sequence is important in determining the level of PST synthesis. Leader sequences dramatically reduce the level of PST synthesis. The presence of a poly A tail in the 3'-untranslated region has no significant effect on PST synthesis.

Animals

Codon equilibrium I: Testing for homogeneous equilibrium.

We present theoretical considerations that suggest that synonymous-codon usage might be expected to be close to an equilibrium distribution given a very homogeneous process of silent substitution. By homogeneous we mean that substitution depends only on the two bases involved, so that 12 base-substitution rates completely describe the silent substitution process. We have developed a method of statistically testing for such homogeneous equilibrium and applied it to reported data on the codon usages of different classes of organisms. Weakly expressed bacterial sequences and both mammalian and nonmammalian eukaryotic sequences deviate significantly from a random pattern of codon usage, in the direction of homogeneous equilibrium. On the other hand, highly expressed bacterial sequences do not exhibit homogeneous equilibrium, which may be correlated with recent experimental results showing that they are optimized to accept the most abundant tRNAs. To examine the effect of amino acid replacements on the homogeneous model of silent substitution, we divided the amino acids with degenerate codes into two classes, those with high mutabilities and those with low, and performed the same analysis on bacterial and eukaryotic data sets. The codon sets of the highly mutable class of amino acids are not further from homogeneous equilibrium than are the codon sets of the class with low mutabilities. We also found for the eukaryotic data that these independent classes of codon sets show very similar equilibrium patterns. The various results suggest a high level of uniformity in the process of silent fixation in the different synonymous-codon sets, especially in eukaryotes.

Amino Acid Sequence