Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Rapid construction of large synthetic genes: total chemical synthesis of two different versions of the bovine prochymosin gene.

We have tested several different synthesis designs and assembly methodologies to develop an improved gene synthesis strategy which enables significantly longer nucleotide sequences to be easily constructed. This strategy, based in part upon our ability to synthesize high-quality extended-length oligodeoxynucleotides (over 100-mer in length), together with the use of chemical 5'-phosphorylation, and simplified low-melting-temperature agarose gel purification methods, combines ease, speed and high overall efficiency. We show that it is now feasible to synthesize routinely even long genes (at least 1-2 kb). To demonstrate this capability we have chemically synthesized and assembled two different versions of the gene encoding the bovine enzyme prochymosin (prorennin). One gene is essentially the natural bovine prochymosin gene sequence. In the second gene the codons have been optimized with regard to the codon bias of highly expressed yeast genes. Each synthetic gene was in excess of 1100 bp, yet they were assembled from only 13 or 14 pairs of complementary oligodeoxynucleotides (oligos), the average lengths of which were 87 and 82 bp, respectively. The 'mutation' rate was low enough to assess that more than 75% of all such oligo pairs (160-170 total nt) were error-free.

Animals

Expression of heterologous genes in Saccharomyces cerevisiae from vectors utilizing the glyceraldehyde-3-phosphate dehydrogenase gene promoter.

The promoter region from the cloned glyceraldehyde-3-phosphate dehydrogenase (GPD) gene of Saccharomyces cerevisiae (Musti et al., 1983) has been characterized. A 653-bp TaqI restriction fragment with a 3' border 24 bp upstream from the ATG initiation codon was isolated and demonstrated to contain all sequences necessary for promoter function in vivo. This DNA segment was converted to a portable promoter by cloning it into M13mp9, and the entire nucleotide sequence of the portable promoter was determined. Two generalized yeast expression vectors have been constructed utilizing the GPD portable promoter. The expression vectors include the yeast 2 mu origin of replication and amplification functions, such that the plasmids are maintained at high copy number in ciro yeast hosts. These vectors direct synthesis of a consensus alpha-interferon (IFN-alpha Con1) as 1% of total cell protein. Hepatitis B surface antigen (HBsAg) was also expressed from these vectors. The 5' end of the HBsAg gene was replaced with a synthetic DNA segment which restored the deleted GPD untranslated leader and utilized optimal yeast codons for the first 30 amino acids. The partially synthetic gene resulted in a 10- to 15-fold increased expression level from GPD vectors yielding HBsAg polypeptide as 2-4% of total cell protein.

Base Sequence

Assessment of mRNA Decay and Calculation of Codon Occurrence to mRNA Stability Correlation Coefficients after 5-EU Metabolic Labeling.

mRNA translation and decay are tightly connected. This chapter describes a method to assess the influence of each codon identity on mRNA stability in cultured cells. The technique involves metabolic labeling of the nascent mRNAs by addition of the nucleoside analog 5-ethynyluridine (5-EU), purification of the RNA at different time-points after chase of the 5-EU, then biotinylation with Click chemistry, pull-down, and sequencing. The transcripts' half-lives are calculated from the expression level of each mRNA at the different time-points. Finally, the method describes the calculation of the Codon occurrence to mRNA Stability correlation Coefficient, or CSC, as a correlation between the codon occurrence in a transcript and the transcript half-life, for each codon.

RNA Stability

Sequence and translation of the murine coronavirus 5'-end genomic RNA reveals the N-terminal structure of the putative RNA polymerase.

A 28-kilodalton protein has been suggested to be the amino-terminal protein cleavage product of the putative coronavirus RNA polymerase (gene A) (M.R. Denison and S. Perlman, Virology 157:565-568, 1987). To elucidate the structure and mechanism of synthesis of this protein, the nucleotide sequence of the 5' 2.0 kilobases of the coronavirus mouse hepatitis virus strain JHM genome was determined. This sequence contains a single, long open reading frame and predicts a highly basic amino-terminal region. Cell-free translation of RNAs transcribed in vitro from DNAs containing gene A sequences in pT7 vectors yielded proteins initiated from the 5'-most optimal initiation codon at position 215 from the 5' end of the genome. The sequence preceding this initiation codon predicts the presence of a stable hairpin loop structure. The presence of an RNA secondary structure at the 5' end of the RNA genome is supported by the observation that gene A sequences were more efficiently translated in vitro when upstream noncoding sequences were removed. By comparing the translation products of virion genomic RNA and in vitro transcribed RNAs, we established that our clones encompassing the 5'-end mouse hepatitis virus genomic RNA encode the 28-kilodalton N-terminal cleavage product of the gene A protein. Possible cleavage sites for this protein are proposed.

Amino Acid Sequence

High-level production of fully active human alpha 1-antitrypsin in Escherichia coli.

The human alpha-1-antitrypsin (A1AT) gene expressed in Escherichia coli as a full-length, non-fusion gene product accumulates to a relatively low level approaching less than or equal to 0.1% of total cellular protein. In contrast, deletion of the first 5, 10 or 15 codons leads to production of truncated A1AT derivatives at levels between 10 and 30% of total cellular protein. The protein with the largest truncation was insoluble and inactive following solubilization by chaotropic agents. In contrast, the two derivatives with the smaller truncations were found to be soluble, and exhibit identical specific activities in both trypsin and elastase inhibition assays to authentic human A1AT. The expression of the full-length A1AT was also optimized by making silent third position mutations within its first 15 codons. These mutations were chosen to optimize codon usage and minimize the possibility of RNA secondary structure formation in this region. Via this approach, expression of full-length, authentic, fully active A1AT was increased at least 20-fold to 2% of total cellular protein. Optimal expression was obtained using as few as three silent mutations in the first five codons, confirming the importance of this 5'-terminal region as had been defined by our deletion mutants. Both the full-length derivatives as well as the small N-terminal deletion derivative can be readily purified from bacterial extracts in fully active form suitable for the examination of their potential therapeutic application.

Base Sequence

Codon usage and secondary structure of mRNA.

The specific codon usage pattern of the repetitive unit nucleotide sequence of silk fibroin mRNA suggests that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA. The correlation between the stability map of local secondary structure of type I collagen mRNA and the codon usage pattern and the translation rate of the collagen is also implied.

Animals

Specific codon usage pattern and its implications on the secondary structure of silk fibroin mRNA.

We have identified two distinctive regions of the repetitive unit nucleotide sequence of fibroin mRNA of Bombyx mori. The codon usage for the major amino acids, glycine, alanine and serine is distinctly different in these two regions, indicating that it is determined by the fibroin mRNA or gene structure but not by the tRNA population. Comparative computer analyses of nucleotide substitutions in the unit sequence suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the fibroin mRNA.

Amino Acid Sequence

Growth rate dependence of transfer RNA abundance in Escherichia coli.

We have tested the predictions of a model that accounts for the codon preferences of bacteria in terms of a growth maximization strategy. According to this model the tRNA species cognate to minor and major codons should be regulated differently under different growth conditions: the isoacceptors cognate to major codons should increase at fast growth rates while those cognate to minor codons should decrease at fast growth rates. We have used a quantitative Northern blotting technique to measure the abundance of the methionine and the leucine isoacceptor families over growth rates ranging from 0.5 to 2.1 doublings per hour. Five tRNA species that are cognate to major codons (tRNA(eMet), tRNA(1fMet), tRNA(2fMet), tRNA(1Leu) and tRNA(3Leu) increase both as a relative fraction of total tRNA and in absolute concentration with increasing growth rates. Three tRNA species that are cognate to minor codons (tRNA(2Leu), tRNA(4Leu) and tRNA(5Leu) decrease as a relative fraction of total RNA and in absolute concentration with increasing growth rates. These data suggest that the abundances of groups of tRNA species are regulated in different ways, and that they are not regulated simply according to isoacceptor specificity. In particular, the data support the growth optimization model for codon bias.

Base Sequence

Effect of ribosome binding site on gene expression in Escherichia coli.

Using the expression of human renin gene in Escherichia coli as a model, we have observed that a distance of 7-10 nucleotides between the Shine-Dalgarno site to the initiation codon ATG is optimal for translation initiation. The sequence ATA as the triplet preceding the ATG gives the best expression among those tested. These observations agree with the statistical bias observed for the genes of E. coli and its phages. We have also compared gene expression with three different Shine-Dalgarno sites. Expression of a given gene can be increased by as much as 1000-fold with slight modifications in the Shine-Dalgarno sequence.

Base Sequence

Secretion and export of IGF-1 in Escherichia coli strain JM101.

The processing of LamB-IGF-1 fusion protein and the export of processed IGF-1 (insulin-like growth-factor-1) into the growth medium was examined in the Escherichia coli host strain, JM101. Several strain or plasmid modifications were tried to increase export of periplasmic (processed) IGF-1 into the growth medium of JM101. These included: (1) use of a lon null mutant strain to increase accumulation levels of unprocessed LamB-IGF-1 fusion protein; (2) use of an alternative drug resistance marker on the expression plasmid rather than beta-lactamase, thereby reducing any competition for processing of LamB-IGF-1 by signal peptidase; (3) examination of whether phage M13 gene III protein expression caused more periplasmic IGF-1 to be exported into the growth medium due to increased outer membrane permeability; and (4) examination of the effect of E. coli or yeast optimized IGF-1 codons. None of these strain or plasmid modifications caused any significant increase in export of IGF-1 into the growth medium of JM101. Solubility studies of LamB-IGF-1 and processed IGF-1 showed that virtually all of the LamB-IGF-1 and IGF-1 remaining within the cell after a 2 h induction period was insoluble. This implied that only soluble LamB-IGF-1 was processed to IGF-1 and that only soluble IGF-1 was exported into the growth medium. Taken together, the results indicated that LamB-IGF-1 and IGF-1 solubility were the limiting factors in secretion of IGF-1 into the periplasm and export of IGF-1 into the growth medium.

Amino Acid Sequence

Nucleotide sequences of two serine tRNAs with a GGA anticodon: the structure-function relationships in the serine family of E. coli tRNAs.

We have determined the nucleotide sequence of the major species of E. coli tRNASer and of a minor species having the same GGA anticodon. These two tRNAs should recognize the UCC and UCU codons, the most widely used codons for serine in the highly expressed genes of E. coli. The two sequences differ in only one position of the D-loop. Neither tRNA has a modified adenosine in the position 3'-adjacent to the anticodon. This can be rationalized on the basis of a structural constraint in the anticodon stem and may be related to optimization of the codon-anticodon interaction. Comparison of all E.coli serine tRNAs (and that encoded by bacteriophage T4) reveals characteristic (possibly functional) features. Evolutionary analysis suggests an eubacterial origin of the T4 tRNASer gene and the existence of a recent common ancestor for the tRNASerGGA and tRNASerGUC genes.

Anticodon

Optimizing nucleotide mixtures to encode specific subsets of amino acids for semi-random mutagenesis.

In random mutagenesis, synthesis of an NNN triplet (i.e. equiprobable A, C, G, and T at each of the three positions in the codon) could be considered an optimal nucleotide mixture because all 20 amino acids are encoded. NN(G,C) might be considered a slightly more intelligent "dope" because the entire set of amino acids is still encoded using only half as many codons. Using a general algorithm described herein, it is possible to formulate more complex doping schemes which encode specific subsets of the twenty amino acids, excluding others from the mix. Maximizing the equiprobability of amino acid residues contributing to such a subset is suggested as an optimal basis for performing semi-random mutagenesis. This is important for reducing the nucleotide complexity of combinatorial cassettes so that "sequence space" can be searched more efficiently. Computer programs have been developed to provide tables of optimized dopes compatible with automated DNA synthesizers.

Amino Acid Sequence

Influence of the codon following the initiation codon on the expression of the lacZ gene in Saccharomyces cerevisiae.

A set of 32 different codons were introduced in a lacZ expression vector (pPTK400) immediately 3' from the AUG initiation codon. Expression of the lacZ gene was determined in Saccharomyces cerevisiae by measuring the amount of beta-galactosidase fusion protein using immuno-gel electrophoresis. A 5.3-fold difference in expression was found among the various constructs. It was found that there was no preference for a certain nucleotide in any position of the second codon and there was no distinct correlation between the level of tRNA corresponding to any particular second codon and expression. No correlation could be found between the local secondary structure and expression. When the overall codon usage in yeast and the codon usage in the second position of the mRNA is compared, there is no obvious significant difference in preference. This indicates that in yeast, in contrast to Escherichia coli, the codon choice at the beginning of the mRNA does not deviate from the one further downstream and is determined by the requirements for optimal translation elongation. Important determinants of the optimal context for an initiation codon in yeast therefore must be located mainly 5' from this codon.

Amino Acid Sequence

Codon usage in yeast: cluster analysis clearly differentiates highly and lowly expressed genes.

Codon usage data has been compiled for 110 yeast genes. Cluster analysis on relative synonymous codon usage revealed two distinct groups of genes. One group corresponds to highly expressed genes, and has much more extreme synonymous codon preference. The pattern of codon usage observed is consistent with that expected if a need to match abundant tRNAs, and intermediacy of tRNA-mRNA interaction energies are important selective constraints. Thus codon usage in the highly expressed group shows a higher correlation with tRNA abundance, a greater degree of third base pyrimidine bias, and a lesser tendency to the A+T richness which is characteristic of the yeast genome. The cluster analysis can be used to predict the likely level of gene expression of any gene, and identifies the pattern of codon usage likely to yield optimal gene expression in yeast.

Base Composition

Context effects and inefficient initiation at non-AUG codons in eucaryotic cell-free translation systems.

The context requirements for recognition of an initiator codon were evaluated in vitro by monitoring the relative use of two AUG codons that were strategically positioned to produce long (pre-chloramphenicol acetyl transferase [CAT]) and short versions of CAT protein. The yield of pre-CAT initiated from the 5'-proximal AUG codon increased, and synthesis of CAT from the second AUG codon decreased, as sequences flanking the first AUG codon increasingly resembled the eucaryotic consensus sequence. Thus, under prescribed conditions, the fidelity of initiation in extracts from animal as well as plant cells closely mimics what has been observed in vivo. Unexpectedly, recognition of an AUG codon in a suboptimal context was higher when the adjacent downstream sequence was capable of assuming a hairpin structure than when the downstream region was unstructured. This finding adds a new, positive dimension to regulation by mRNA secondary structure, which has been recognized previously as a negative regulator of initiation. Translation of pre-CAT from an AUG codon in a weak context was not preferentially inhibited under conditions of mRNA competition. That result is consistent with the scanning model, which predicts that recognition of the AUG codon is a late event that occurs after the competition-sensitive binding of a 40S ribosome-factor complex to the 5' end of mRNA. Initiation at non-AUG codons was evaluated in vitro and in vivo by introducing appropriate mutations in the CAT and preproinsulin genes. GUG was the most efficient of the six alternative initiator codons tested, but GUG in the optimal context for initiation functioned only 3 to 5% as efficiently as AUG. Initiation at non-AUG codons was artifactually enhanced in vitro at supraoptimal concentrations of magnesium.

Animals

Synthetic oligonucleotide probes deduced from amino acid sequence data. Theoretical and practical considerations.

Synthetic probes deduced from amino acid sequence data are widely used to detect cognate coding sequences in libraries of cloned DNA segments. The redundancy of the genetic code dictates that a choice must be made between (1) a mixture of probes reflecting all codon combinations, and (2) a single longer "optimal" probe. The second strategy is examined in detail. The frequency of sequences matching a given probe by chance alone can be determined and also the frequency of sequences closely resembling the probe and contributing to the hybridization background. Gene banks cannot be treated as random associations of the four nucleotides, and probe sequences deduced from amino acid sequence data occur more often than predicted by chance alone. Probe lengths must be increased to confer the necessary specificity. Examination of hybrids formed between unique homologous probes and their cognate targets reveals that short stretches of perfect homology occurring by chance make a significant contribution to the hybridization background. Statistical methods for improving homology are examined, taking human coding sequences as an example, and considerations of codon utilization and dinucleotide frequencies yield an overall homology of greater than 82%. Recommendations for probe design and hybridization are presented, and the choice between using multiple probes reflecting all codon possibilities and a unique optimal probe is discussed.

Amino Acid Sequence

Total DNA synthesis and cloning in Escherichia coli of a gene coding for the human growth hormone releasing factor.

A DNA containing a sequence coding for the human growth hormone releasing factor (hGRF) has been obtained by enzymatic assembly of chemically synthesized DNA fragments. The synthetic gene consists of a 140 base-pair fragment containing initiation and termination signals for translation and appropriate protruding ends for cloning into a newly constructed plasmid vector (pULB1219). Eleven oligodeoxyribonucleotides, from 14 to 31 bases in length, sharing pairwise stretches of complementary regions of at least 13 bases were prepared by phosphotriester solid-phase synthesis. The DNA sequence was designed to take into account the optimal use of E. coli codons. Oligomers were annealed in one step and assembled by ligation. The DNA fragment of the expected size (140 bp) was recovered and cloned into the pULB1219 vector. The expected sequence was confirmed by DNA sequencing.

Amino Acid Sequence

Role of an upstream open reading frame in the translation of polycistronic mRNAs in plant cells.

The influence of an upstream small open reading frame (URF) on the translation of two consecutive coding regions on an eukaryotic mRNA was studied. The cis effects of leader length, URF length, the sequences of the URF and neighboring regions, and the trans effects of the Cauliflower mosaic virus transactivator (TAV) were analyzed. Translation efficiency of the immediate downstream open reading frame (ORF) decreased with increasing URF length. Short URFs did not drastically inhibit translation of immediate downstream ORFs but supported far downstream translation in the presence of TAV. In the latter case, the optimal URF length was 30 codons.

Base Sequence