Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Total DNA synthesis and cloning in Escherichia coli of a gene coding for the human growth hormone releasing factor.

A DNA containing a sequence coding for the human growth hormone releasing factor (hGRF) has been obtained by enzymatic assembly of chemically synthesized DNA fragments. The synthetic gene consists of a 140 base-pair fragment containing initiation and termination signals for translation and appropriate protruding ends for cloning into a newly constructed plasmid vector (pULB1219). Eleven oligodeoxyribonucleotides, from 14 to 31 bases in length, sharing pairwise stretches of complementary regions of at least 13 bases were prepared by phosphotriester solid-phase synthesis. The DNA sequence was designed to take into account the optimal use of E. coli codons. Oligomers were annealed in one step and assembled by ligation. The DNA fragment of the expected size (140 bp) was recovered and cloned into the pULB1219 vector. The expected sequence was confirmed by DNA sequencing.

Amino Acid Sequence

Role of an upstream open reading frame in the translation of polycistronic mRNAs in plant cells.

The influence of an upstream small open reading frame (URF) on the translation of two consecutive coding regions on an eukaryotic mRNA was studied. The cis effects of leader length, URF length, the sequences of the URF and neighboring regions, and the trans effects of the Cauliflower mosaic virus transactivator (TAV) were analyzed. Translation efficiency of the immediate downstream open reading frame (ORF) decreased with increasing URF length. Short URFs did not drastically inhibit translation of immediate downstream ORFs but supported far downstream translation in the presence of TAV. In the latter case, the optimal URF length was 30 codons.

Base Sequence

[Rare initiation codons are regulators of expression of the rpoC gene].

Translation of the rpoC genes in Escherichia coli and Salmonella typhimurium is known to start from the GUG codon. Now, using toeprint analysis we have shown UUG to be the initiation codon of the Pseudomonas putida rpoC gene. IF3 does not seem to proofread initiation at the UUG codon. The rpoC genes of P. putida, E. coli, and S. typhimurium, which use rare start codons, have strong SD-domains AGGAGG (P. p.) and GGGAG (E. c., S. t.), optimal seven-nucleotide spacing between SD and start codons, and good second codon AAA. We suggest that rpoC presents an infrequent case of the regulation of translation initiation by selecting the start codon.

Base Sequence

Codon evolution and conservation of the reading phase in genetic code translation.

The description of the optimized evolution of a code based on 4 nucleotides involves a sequential transition of codons, formed firstly by monomers evolving to dimers and then to triplets, in accordance with the progressive increase of the number of amino acids to be coded. The successive increase in the size of these codons during evolution implies changes in the phase reading of the genetic message, which could become chaotic. In order to overcome this constraint, this paper proposes a codon evolution where two things occur simultaneously: codons change in size and there is an alternation of the molecule which holds the information. For example, the nucleotides of the original oligonucleotide are read as monomers when they are translated to an oligopeptide, but further on, this oligopeptide which is read as amino acid dimers, is translated to a nucleotide form (oligonucleotide). Finally, amino acids conforming a peptide are translated from this oligonucleotide, through a reading of triplets. Although plausible, this evolution is a low-probability process due to the fact that it requires a singular sequence of the oligonucleotide and oligopeptide involved. An alternative hypothesis of evolution is also discussed. It proposes that with the exclusion of the establishment of monomer and dimer codons, there is a direct generation of a code of trinucleotides which arises only when a certain number of amino acids has already been generated. Both hypotheses are discussed in terms of the development of a code in which an optimized hardware is maintained through out its evolution.

Amino Acids

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals

Efficient gene expression in mammalian cells from a dicistronic transcriptional unit in an improved retroviral vector.

We have studied the properties of dicistronic transcriptional units in retroviral vectors. In these vectors, the promoter in the 5' retroviral long terminal repeat (LTR) controls expression of both an upstream cistron (luc) encoding firefly luciferase and a downstream cistron (neo), a selectable marker encoding neomycin phosphotransferase (NPTII). By assaying for simultaneous expression of luc and neo after transfection or infection of hamster BHK, rat 208F, and mouse retroviral packaging cell lines, we have identified important factors that affect expression from the downstream cistron, including the presence of intercistronic ATG sequences, the length of the intercistronic sequence and conformity of the sequence surrounding the downstream start codon to the eukaryotic consensus sequence. Optimized dicistronic vectors produced amounts of NPTII comparable to a vector in which neo was driven by a strong internal promoter consisting of a modified Rous sarcoma virus LTR. Additionally, they produced higher virus titers and demonstrated improved stability of gene expression in the absence of selection. By virtue of their physical compactness and elimination of the need for a separate promoter for every gene, dicistronic transcriptional units allow the introduction of larger genes into retroviral vectors and may allow for more than two genes to be placed in a single vector.

Avian Sarcoma Viruses

Melibiose permease and alpha-galactosidase of Escherichia coli: identification by selective labeling using a T7 RNA polymerase/promoter expression system.

Identification and selective labeling of the melibiose permease and alpha-galactosidase in Escherichia coli, which are encoded by the melB and melA genes, respectively, have been accomplished by selectively labeling the two gene products with a T7 RNA polymerase expression system [Tabor, S., & Richardson, C. C. (1985) Proc. Natl. Acad. Sci. U.S.A. 82, 1074]. Following generation of a novel EcoRI restriction site in the intergenic sequence between the two genes of the mel operon by oligonucleotide-directed, site-specific mutagenesis, melA and melB were separately inserted into plasmid pT7-6 of the T7 expression system. Expression of melB was markedly enhanced by placing a strong, synthetic ribosome binding site at an optimal distance upstream from the initiation codon of melB. Expression of cloned gene products was characterized functionally and by performing autoradiographic analysis on total cell, inner membrane, and cytoplasmic proteins from cells pulse labeled with (35S)methionine in the presence of rifampicin and resolved by sodium dodecyl sulfate/polyacrylamide gel electrophoresis. The results first confirm that alpha-galactosidase is a cytoplasmic protein with an Mr of 50K; in contrast, the membrane-bound melibiose permease is identified as a protein with an apparent Mr of 39K, a value significantly higher than that of 30K previously suggested [Hanatani et al. (1984) J. Biol. Chem. 259, 1807].

Animals

Binding of the bacteriophage T4 regA protein to mRNA targets: an initiator AUG is required.

Bacteriophage T4 regA protein translationally represses the synthesis of a subset of early phage-induced proteins. The protein binds to the translation initiation site of at least two mRNAs and prevents formation of the initiation complex. We show here that the protein binds to the translation initiation sites of other regA-sensitive mRNAs. Analysis of mRNA binding by filtration and nuclease protection assays shows that AUG is necessary but not sufficient for specific binding of regA protein to its mRNA targets. Anticipating the need for large quantities of regA protein for structural studies to further define the regA protein-RNA ligand interaction, we also report cloning the regA gene into a T4 overexpression system. The expression of regA protein in uninfected E. coli is lethal, so in our system regA driven by a strong T7 promoter is sequestered in a T4 phage until 'induction' by phage infection is desired. We have replaced the regA sensitive wild-type ribosome binding site with a strong insensitive ribosome binding site at an optimal distance from the regA initiation codon for maximizing expression. We have obtained large amounts of regA protein.

Base Sequence

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

CUG as a mutant start codon for cat-86 and xylE in Bacillus subtilis.

The cat-86 gene specifies chloramphenicol acetyltransferase (CAT). The cat-86 start codon is UUG, although related genes have AUG as the start codon. Changing the start codon to AUG increased expression of cat-86 by 36% in Bacillus subtilis. Changing the start codon to GUG and CUG decreased expression to 65% and 30%, respectively, of the level obtained when AUG was the start codon. CUG has not been previously shown to function as a start codon in B. subtilis. N-terminal sequencing of purified CAT protein specified by the CUG mutant, revealed that CUG was indeed the start codon and specified methionine. The gene xylE, which specifies catechol 2,3-dioxygenase, has AUG as its start codon. Changing the start codon for xylE to CUG decreased expression by 98%. However, when the ribosome-binding site sequence for xylE was optimized and the spacing between it and the start codon was increased to 8 nucleotides, xylE activity increased to 13% of the activity observed for AUG. CUG did not function efficiently as a start codon for cat-86 in Escherichia coli. These data suggest conditions under which CUG can function, with modest efficiency, as a start codon in B. subtilis.

Bacillus subtilis

Analysis of leaky viral translation termination codons in vivo by transient expression of improved beta-glucuronidase vectors.

Plant RNA viruses commonly exploit leaky translation termination signals in order to express internal protein coding regions. As a first step to elucidate the mechanism(s) by which ribosomes bypass leaky stop codons in vivo, we have devised a system in which readthrough is coupled to the transient expression of beta-glucuronidase (GUS) in tobacco protoplasts. GUS vectors that contain the stop codons and surrounding nucleotides from the readthrough regions of several different RNA viruses were constructed and the plasmids were tested for the ability to direct transient GUS expression. These studies indicated that ribosomes bypass the leaky termination sites at efficiencies ranging from essentially 0 to ca. 5% depending upon the viral sequence. The results suggest that the efficiency of readthrough is determined by the sequence surrounding the stop codon. We describe improved GUS expression vectors and optimized transfection conditions which made it possible to assay low-level translational events.

Base Sequence

The extracellular portion of HLA-DR alpha chain is composed of two compactly folded domains.

A truncated form of the class II antigen DR alpha chain of the human major histocompatibility complex was produced in bacteria. A cDNA clone encoding the intact chain was modified so that the segment encoding the signal sequence was replaced by an ATG codon and the 3' region downstream to the part corresponding to the third exon was replaced by a stop codon. The new construct was put under the control of the Tac promoter in a bacterial expression vector. The distance between the Shine-Delgarno sequence and the initiation codon was randomized so that clones with optimal expression of the truncated DR alpha chain could be obtained after induced expression and immunoscreening. The truncated DR alpha chain was subjected to limited proteolysis with chymotrypsin, and the resulting cleavage products were analysed by sodium dodecyl sulphate-polyacrylamide gel electrophoresis. Two fragments were visualized by western blotting. Electrophoresis in the absence and presence of reducing agents suggested that one of the proteolytic fragments contained a disulphide bridge. It is concluded that the extracellular portion of the DR alpha chain is composed of two compactly folded domains connected by an extended stretch of the polypeptide chain.

Base Sequence

The bases of the tRNA anticodon loop are independent by genetic criteria.

We employed two methods to study the translational role of interactions between anticodon loop nucleotides. Starting with a set of previously constructed weakly-suppressing anticodon loop mutants of Su7, we searched for second-site revertants that increase amber suppressor efficiency. Though hundreds of revertants were characterized, no second-site revertants were found in the anticodon loop. Second site reversion was detected in the D-stem, thereby demonstrating the efficacy of the search method. As a second method for detecting interactions, we used site-directed mutagenesis to construct multiple mutations in the anticodon loop. These multiple mutants are very weak suppressors and have translational activities that are equal to or lower than that predicted for the independent action of single mutations. We conclude that although the anticodon loop sequence of Su7 has an optimal structure for the translation of amber codons, we find no evidence that interactions between loop bases can enhance translational efficiency.

Anticodon

The mouse int-2 gene exhibits basic fibroblast growth factor activity in a basic fibroblast growth factor-responsive cell line.

The int-2 protein is related to basic fibroblast growth factor (bFGF) by amino acid sequence homology. To assess its biological activity, we constructed retroviral vectors containing four variants of mouse int-2 complementary DNA under the transcriptional control of the beta-actin promoter and tested their effects on human SW13 adrenal cortical tumor cells. This cell line specifically requires bFGF, interleukin 1, or transforming growth factor e for anchorage-independent growth in soft agar. Despite encoding a signal sequence that should direct the protein to the secretory pathway, vectors containing unmodified int-2 complementary DNA, or a form optimized for translation initiation at the AUG codon, were incapable of inducing SW13 growth in soft agar. However, SW13 transfectants expressing a construct (pSP1), in which a mouse immunoglobulin signal peptide sequence is linked to the int-2 coding sequences, grew well in soft agar. The concentrated conditioned medium from these pSP1-transfected cells supported anchorage-independent growth of SW13 indicator cells and competed with bFGF for binding to receptors. Western blot analysis with an int-2-specific antiserum detected Mr 30,000-32,000 int-2 products in cell extracts and conditioned medium from pSP1-transfected clones, whereas the conditioned medium from these and other SW13 clones contained only low levels of bFGF as measured in a specific radioimmunoassay. These data suggest that the product of the int-2 gene can functionally replace bFGF in modulating the anchorage-independent growth of SW13 cells.

Amino Acid Sequence

Point mutations define a sequence flanking the AUG initiator codon that modulates translation by eukaryotic ribosomes.

By analyzing the effects of single base substitutions around the ATG initiator codon in a cloned preproinsulin gene, I have identified ACCATGG as the optimal sequence for initiation by eukaryotic ribosomes. Mutations within that sequence modulate the yield of proinsulin over a 20-fold range. A purine in position -3 (i.e., 3 nucleotides upstream from the ATG codon) has a dominant effect; when a pyrimidine replaces the purine in position -3, translation becomes more sensitive to changes in positions -1, -2, and +4. Single base substitutions around an upstream, out-of-frame ATG codon affect the efficiency with which it acts as a barrier to initiating at the downstream start site for preproinsulin. The optimal sequence for initiation defined by mutagenesis is identical to the consensus sequence that emerged previously from surveys of translational start sites in eukaryotic mRNAs. The mechanism by which nucleotides flanking the ATG codon might exert their effect is discussed.

Animals

Codon usage pattern in alpha 2(I) chain domain of chicken type I collagen and its implications for the secondary structure of the mRNA and the synthesis pauses of the collagen.

A stability map of local secondary structure of the mRNA of the triple-helical alpha 2(I) chain domain of chicken type I collagen was obtained by plotting the free energy of the optimal secondary structure of a local segment in mRNA against the segment position along a base sequence of the mRNA. It was found that the positions of the minima of free energy in the plot coincide with the positions where synthesis pauses of the alpha-chain polypeptides of the corresponding sizes translated from the mRNA have been reported to occur (1). The codon usage pattern of each of the three major amino acids of the alpha-chain domain of the collagen, Gly, Pro and Ala, fluctuates considerably along the base sequence segments of the mRNA and a deviation of the pattern from that of the average of the whole alpha 2(I) chain domain mRNA, particularly for Gly codons, leads to a loss of the stability of the local secondary structure of the mRNA. The results suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the mRNA of the chicken collagen alpha 2(I) chain domain which leads to a nonuniform polypeptide elongation pattern.

Animals