Search PubMedSearch

SEARCH · Search PubMed

Results for “codon optimization”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Cloning and sequence determination of a cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase.

A partial cDNA encoding Aspergillus nidulans calmodulin-dependent multifunctional protein kinase (ACMPK) was isolated from a lambda ZAP expression library by immunoselection using monospecific polyclonal antibodies to the enzyme. The sequence of both strands of the cDNA (CMKa) was determined. The deduced amino acid (aa) sequence contained all eleven consensus domains found in serine/threonine protein kinases [Hanks et al., Science 241 (1988) 42-52], as well as a putative calmodulin-binding domain. The cDNA contained an intron, lacked an in-frame start codon, and was not polyadenylated. A full-length copy of CMKa was subsequently isolated from a lambda gt10 library of A. nidulans cDNA using a restriction fragment of the first clone as a probe. It contained an in-frame start codon, an open reading frame (ORF) of 1242 bp and was polyadenylated. The ORF encoded a protein of 414 aa residues with an M(r) of 46,895 and an isoelectric point pI = 6.4. These values are in good agreement with that observed for the native enzyme [Bartelt et al., Proc. Natl. Acad. Sci. USA 85 (1988) 3279-3283]. When aligned to optimize homology, 29% of the predicted aa sequence of ACMPK is identical to that of the alpha-subunit of rat brain calmodulin-dependent protein kinase II. ACMPK shares 40 and 44% identity in aa sequence with YCMK1 and YCMK2, respectively, two Ca2+/calmodulin-dependent protein kinases recently cloned from Saccharomyces cerevisiae [Pausch et al., EMBO J. 10 (1991) 1511-1522]. Results of Southern analysis of restriction digests of genomic DNA indicate that ACMPK is encoded by a single-copy gene.

Amino Acid Sequence

Triphasic concentration effects of gentamicin on activity and misreading in protein synthesis.

Gentamicin is shown to exert a triphasic concentration effect on peptide synthesis in vitro with natural messengers. Low concentrations (up to 2 micron) caused slowing and a decrease in total synthesis, but little misreading (assayed with extracts lacking Glu-tRNA); the inhibition was greater with an initiating system (with phage RNA as messenger) than with pure chain elongation on purified endogenous polysomes of Escherichia coli. Moderate concentrations (up to 100 micron) slowed synthesis less, markedly increased its duration in the noninitiating system, and strongly stimulated misreading; at optimal concentrations total synthesis was even greater than normal. Moreover, with phage RNA these concentrations increased the synthesis of large polypeptides. We conclude that binding of gentamicin to its first site causes inhibition but little misreading; binding to additional site(s) partly reverses the inhibition by first-site binding and markedly stimulates misreading, and the misreading appears to favor "readthrough" of termination codons. In the third phase (greater than 100 micron) synthesis is slowed again but the pattern of misreading does not appear to be altered; this effect need not involve a specific further action on the ribosome.

Bacterial Proteins

The androgen receptor in LNCaP cells contains a mutation in the ligand binding domain which affects steroid binding characteristics and response to antiandrogens.

The human prostate tumor cell line LNCaP contains an abnormal androgen receptor system with broad steroid binding specificity. Progestagens, estradiol and several antiandrogens compete with androgens for binding to the androgen receptor in the cells to a higher extent than in other androgen sensitive systems. Optimal growth of LNCaP cells is observed after addition of the synthetic androgen R1881 (0.1 nM). In addition, estrogens, progestagens and several antiandrogens do not inhibit androgen responsive growth, but have striking growth stimulatory effects and increase EGF receptor level and acid phosphatase secretion. We have found that the androgen receptor in the LNCaP cells contains a single point mutation changing the sense of codon 868 (Thr to Ala) in the ligand binding domain. Expression vectors containing the normal or mutated androgen receptor sequence were transfected into COS or HeLa cells. Androgens, progestagens, estrogens and several antiandrogens bind the mutated androgen receptor protein and activate the expression of an androgen-regulated reporter gene (GRE-tk-CAT), indicating that the mutation directly affects both binding specificity and the induction of gene expression. Interestingly, the antiandrogen casodex showed antiandrogenic properties in growth studies of LNCaP cells and did not induce reporter gene activity in Hela cells transfected with the mutant receptor. The mutated androgen receptor of LNCaP cells is therefore a useful tool in the elucidation of different levels of action of steroids and antisteroids.

Binding Sites

Cloning and expression of the branching enzyme gene (glgB) from the cyanobacterium Synechococcus sp. PCC7942 in Escherichia coli.

Using the glgB gene from Escherichia coli as a hybridization probe, the gene encoding the branching enzyme of the cyanobacterium Synechococcus sp. PCC7942 has been identified on a 3.9-kb PstI fragment which was cloned into plasmid pUC9. Two types of plasmids have been isolated. Plasmid pKVN1 was expressing the Synechococcus sp. gene as was shown by complementation of the glgB mutation of E. coli KV832. Plasmid pKVN2, which carried the same insert in the opposite orientation was unable to complement E. coli KV832, indicating that the promoter of the cloned gene was either absent or was not recognized in E. coli. Determination of branching activity in extracts of Synechococcus sp. and E. coli KV832[pKVN1] showed that the enzyme was optimally active at approximately 35 degrees C. No significant activity was present at temperatures higher than 55 degrees C, reflecting the mesophilic nature of the cloned enzyme. In a cell-free coupled transcription-translation system the cloned gene specified two proteins of 84 kDa and 72 kDa, respectively, which are probably translated independently from the same gene by initiation at two different start codons.

1,4-alpha-Glucan Branching Enzyme

A standardized vector system for manipulation and enhanced expression of genes in Escherichia coli.

Different families of cloning and expression vectors were engineered on a standard plasmid. They contain several regulatory signals for transcription and/or translation initiation and termination. The plasmids in each series differ only in the number, type, and order of unique restriction cleavage sites clustered in front of a transcription terminator. The pLK30 plasmids are general cloning vectors and the corresponding pLK50 plasmids carry the lambda pL promoter. The pLK60 vectors carry the lambda pR promoter and translation initiation signals of the cro gene containing the Shine-Dalgarno sequence and initiation codon. The pLK70 series is similar to pLK60 except that additional 5'-translated cro sequences are included. The pLK80 plasmids have a lacZ gene fragment suitable for the construction of hybrid genes. The presence of translational stop signals in the pLK90 series facilitates the manipulation of genes truncated at the 3' end. This standardized pLK vector system offers great versatility in gene manipulation and in optimization of gene expression under the control of strong regulatable promoters. Measurement of expression levels under repressed conditions permits the identification of optimal promoter-gene configurations in constructions directing high-level expression.

Base Sequence

Location and sequence of the promoter of the gene for the NADH-dependent nitrite reductase of Escherichia coli and its regulation by oxygen, the Fnr protein and nitrite.

The DNA sequence containing the start of the Escherichia coli nirB gene is reported. The N-terminal amino acid sequence of purified NADH-dependent nitrite reductase coincided with that predicted from the DNA sequence, confirming that nirB is the structural gene for nitrite reductase apoprotein and identifying the translation start point. Using nuclease S1 mapping, the sole transcription startpoint for the nirB gene was found 23 or 24 base-pairs upstream from the ATG initiation codon. By subcloning successively smaller DNA fragments into a beta-galactosidase expression vector plasmid, we located the promoter within a sequence bounded by a TaqI site at +14 with respect to the transcription startpoint and a HpaII site at -208. Measurements in vivo of beta-galactosidase expression and RNA levels due to nirB promoter activity showed that this promoter was activated during anaerobic growth. Optimal activity was found only after anaerobic growth in the presence of nitrite. The sequence of the nirB promoter is compared with sequences found at other anaerobically activated promoters.

Bacterial Proteins

The positive regulatory function of the 5'-proximal open reading frames in GCN4 mRNA can be mimicked by heterologous, short coding sequences.

Translational control of GCN4 expression in the yeast Saccharomyces cerevisiae is mediated by multiple AUG codons present in the leader of GCN4 mRNA, each of which initiates a short open reading frame of only two or three codons. Upstream AUG codons 3 and 4 are required to repress GCN4 expression in normal growth conditions; AUG codons 1 and 2 are needed to overcome this repression in amino acid starvation conditions. We show that the regulatory function of AUG codons 1 and 2 can be qualitatively mimicked by the AUG codons of two heterologous upstream open reading frames (URFs) containing the initiation regions of the yeast genes PGK and TRP1. These AUG codons inhibit GCN4 expression when present singly in the mRNA leader; however, they stimulate GCN4 expression in derepressing conditions when inserted upstream from AUG codons 3 and 4. This finding supports the idea that AUG codons 1 and 2 function in the control mechanism as translation initiation sites and further suggests that suppression of the inhibitory effects of AUG codons 3 and 4 is a general consequence of the translation of URF 1 and 2 sequences upstream. Several observations suggest that AUG codons 3 and 4 are efficient initiation sites; however, these sequences do not act as positive regulatory elements when placed upstream from URF 1. This result suggests that efficient translation is only one of the important properties of the 5' proximal URFs in GCN4 mRNA. We propose that a second property is the ability to permit reinitiation following termination of translation and that URF 1 is optimized for this regulatory function.

Base Sequence

The proteomic origin of the genetic code.

INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.

Genetic Code

Effect of deletions in the 5'-noncoding region on the translational efficiency of phosphoglycerate kinase mRNA in yeast.

Deletions of various sizes were introduced into the region of the yeast PGK gene encoding the 5'-nontranslated portion of the phosphoglycerate kinase (PGK) mRNA. The effect of these deletions on the translational efficiency of the mutant transcripts was analysed by assaying the levels of mutant PGK mRNA and PGK protein in cells transformed with the mutant genes. Quantification of transcript levels by either Northern analysis or a reverse transcription assay demonstrated that there were no significant differences in the levels of mutant PGK mRNA between the various mutants. Thus, the leader sequence does not appear to play a role in determining the relatively long half-life of yeast PGK mRNA. Analysis of PGK protein levels in the various mutants revealed no effect when the length of the leader was reduced from 45 to 27 nucleotides (nt). Protein levels dropped by about a factor 2, however, upon a further decrease to 21 nt. Additional shortening did not cause a further dramatic reduction in translational yield. Even an mRNA containing a leader of only 7 nt was still translated at about 50% of the optimal rate. Therefore, while optimal translation of a yeast mRNA requires a leader length of at least some 30 nt, shorter leaders still allow considerable translation to take place.

Base Sequence

Cloning and expression in Escherichia coli of two additional amylase genes of a strictly anaerobic thermophile, Dictyoglomus thermophilum, and their nucleotide sequences with extremely low guanine-plus-cytosine contents.

An obligately anaerobic and extremely thermophilic bacterium, Dictyoglomus thermophilum, produces multiple extracellular amylases. In addition to one of the amylase genes, amyA, which we previously cloned and characterized, we have cloned two additional genes, amyB and amyC, coding for amylases of this thermophile, into Escherichia coli and determined their nucleotide sequences. The two amylase genes were expressed under the control of E. coli promoters. Almost all activity was detected in the intracellular fraction in the E. coli cells. The molecular mass and NH2-terminal amino acid sequence of the AmyB enzyme, which was purified from an E. coli transformant containing the amyB gene, confirmed that the reading frame of amyB consisted of 562 amino acids (Mr 67,000). The molecular mass of the AmyC enzyme, estimated by activity staining of a crude extract of E. coli containing amyC, confirmed that AmyC consisted of 498 amino acids (Mr 59,000). The optimal temperatures for AmyB and AmyC activities on soluble starch were 80 degrees C and 70 degrees C, respectively. Both AmyB and AmyC showed a pH optimum of 5.5. AmyB and AmyC showed a different pattern of starch hydrolysis when examined by thin-layer chromatography. Some homology in the amino acid sequences with the functional regions of Taka-amylase A was found in both AmyB and AmyC. The codon usage in the amyA, amyB and amyC genes was highly biased, which reflects the fact that the guanine-plus-cytosine (G + C) content of DNA of D. thermophilum is 29 mol%. The distribution of G and C at each position of the codons was non-random; the G + C content of the first position of codons is significantly high, whereas that of the third position is somewhat low. In addition, codons consisting only of A and T were preferentially used in this thermophile.

Amino Acid Sequence

Lipofectin enhances cellular uptake of antisense DNA while inhibiting tumor cell growth.

A natural DNA oligomer (15-mer) was synthesized with a sequence complementary to the translation initiation codon region of the human TGF-alpha mRNA and mixed with Lipofectin to form unilamellar complexes. It was found that tumor cell growth was inhibited when HCT116 cells were treated with Lipofectin-DNA oligomer complexes or with Lipofectin alone. Uptake of 32P-labeled 15-mers into colon tumor cells was compared in the presence and absence of Lipofectin. The amount of labeled oligomer found in cells that received optimal ratios of Lipofectin to DNA was 4- to 10-fold higher than the amount found in cells that received 32P-labeled DNA alone. Although Lipofectin-antisense DNA oligomer treatment of HCT116 cells caused a dose-dependent inhibition of cell growth, there was a subsequent rise in target mRNA product. Because the mechanism of growth inhibition could not involve an inhibition of TGF-alpha expression, it was concluded that Lipofectin probably exerts a nonspecific, detergent-like effect upon the cell membrane, producing an enhancement of TGF-alpha processing and release.

Base Sequence

Compositional patterns in vertebrate genomes: conservation and change in evolution.

The evolution of vertebrate genomes can be investigated by analyzing their regional compositional patterns, namely the compositional distributions of large DNA fragments (in the 30-100-kb size range), of coding sequences, and of their different codon positions. This approach has shown the existence of two evolutionary modes. In the conservative mode, compositional patterns are maintained over long times (many million years), in spite of the accumulation of enormous numbers of base substitutions. In the transitional, or shifting, mode, compositional patterns change into new ones over much shorter times. The conservation of compositional patterns, which has been investigated in mammalian genomes, appears to be due in part to some measure of compositional conservation in the base substitution process, and in part to negative selection acting at regional (isochore) levels in the genome and eliminating deviations from a narrow range of values, presumably corresponding to optimal functional properties. On the other hand, shifts of compositional patterns, such as those that occurred between cold-blooded and warm-blooded vertebrates, appear to be due essentially to both negative and positive selection again operating at the isochore level, largely under the influence of changes in environmental conditions, and possibly taking advantage of mutational biases in the replication/repair enzymes and/or in the enzyme make-up of nucleotide precursor pools. Other events (like translocations and changes in chromosomal structure) also play a role in the transitional mode of genome evolution. The present findings (1) indicate that isochores, which correspond to the DNA segments of individual or contiguous chromatin domains, represent selection units in the vertebrate genome; and (2) shed new light on the selectionist-neutralist controversy.

Animals

The promoter of the late p10 gene in the insect nuclear polyhedrosis virus Autographa californica: activation by viral gene products and sensitivity to DNA methylation.

In lepidopteran insect cells infected with the baculovirus Autographa californica nuclear polyhedrosis virus (AcNPV), two major late viral gene products are expressed: the polyhedrin, a 28 000 mol. wt. protein which makes up the mass of the nuclear inclusion bodies, and a 10 000 mol. wt. protein (p10) whose function is unknown. The nucleotide sequences of these strong promoters conform to those of other eukaryotic promoters and are rich in AT base pairs. We used the pSVO-CAT construct containing the prokaryotic gene chloramphenicol acetyl transferase (CAT) to study the function of the p10 gene promoter in insect and mammalian cells. Upon transfection of the pAcp10-CAT construct, which contained 402 bp of the p10 gene of AcNPV DNA in the HindIII site of pSVO-CAT, CAT activity was determined. The p10 gene promoter was inactive in human HeLa cells and in uninfected Spodoptera frugiperda insect cells. The same promoter was active, however, in AcNPV-infected S. frugiperda cells and exhibited optimal activity when cells were transfected 18 h after infection with the insect virus. This finding demonstrated directly that the p10 gene promoter required other viral gene products for its activity in insect cells. The nature of these products was unknown. The p10 gene promoter sequence contained one 5'-CCGG-3' site 40 bp upstream from the cap site of the gene and two such sites 178 and 192 bp downstream from the ATG initiation codon of the gene. Since Drosophila DNA or S. frugiperda DNA contained no 5-methylcytosine or extremely small amounts of it, we were interested in determining the effect of site-specific methylations on the p10 gene insect virus promoter. Methylation at the 5'-CCGG-3' sites led to a block of this promoter.(ABSTRACT TRUNCATED AT 250 WORDS)

Acetyltransferases

A Proteogenomic Pipeline for the Analysis of Protein Biosynthesis Errors in the Human Pathogen Candida albicans.

Candida albicans is a diploid pathogen known for its ability to live as a commensal fungus in healthy individuals but causing both superficial infections and disseminated candidiasis in immunocompromised patients where it is associated with high morbidity and mortality. Its success in colonizing the human host is attributed to a wide range of virulence traits that modulate interactions between the host and the pathogen, such as optimal growth rate at 37 °C, the ability to switch between yeast and hyphal forms, and a remarkable genomic and phenotypic plasticity. A fascinating aspect of its biology is a prominent heterogeneous proteome that arises from frequent genomic rearrangements, high allelic variation, and high levels of amino acid misincorporations in proteins. This leads to increased morphological and physiological phenotypic diversity of high adaptive potential, but the scope of such protein mistranslation is poorly understood due to technical difficulties in detecting and quantifying amino acid misincorporation events in complex protein samples. We have developed and optimized mass spectrometry and bioinformatics pipelines capable of identifying rare amino acid misincorporation events at the proteome level. We have also analyzed the proteomic profile of an engineered C. albicans strain that exhibits high level of leucine misincorporation at protein CUG sites and employed an in vivo quantitative gain-of-function fluorescence reporter system to validate our LC-MS/MS data. C. albicans misincorporates amino acids above the background level at protein sites of diverse codons, particularly at CUG, confirming our previous data on the quantification of leucine incorporation at single CUG sites of recombinant reporter proteins, but increasing misincorporation of Leucine at these sites does not alter the translational fidelity of the other codons. These findings indicate that the C. albicans statistical proteome exceeds prior estimates, suggesting that its highly plastic phenome may also be modulated by environmental factors due to translational ambiguity.

Candida albicans

Purification and characterization of an endoglucanase from Streptomyces lividans 66 and DNA sequence of the gene.

The endoglucanase isolated from culture filtrates of Streptomyces lividans IAF74 was shown to have an Mr of 46,000 and a pI of 3.3. The specific enzyme activity of 539 IU/mg, determined by the reducing assay method on carboxymethyl cellulose, is among the highest reported in the literature. The cellulase showed typical endo-type activity when reacting on oligocellodextrins. Optimal enzyme activity was obtained at 50 degrees C and pH 5.5. The kinetic constants for this endoglucanase, determined with carboxymethyl cellulose as the substrate, were a Vmax of 24.9 IU/mg of enzyme and a Km of 4.2 mg/ml. Activity was found against neither methylumbelliferyl- nor p-nitrophenyl-cellobiopyranoside nor with xylan. The DNA sequence contains one possible reading frame validated by the N terminus of the mature purified protein. However, neither ATG nor GTG starting codons were identified near the ribosome-binding site. A putative TTG codon was found as a good candidate for the start codon. Comparison of the primary amino acid sequence of the endoglucanase of S. lividans revealed that the N terminus contains a bacterial cellulose-binding domain. The catalytic domain at the C terminus showed similarity to endoglucanases from a Bacillus sp. Thus, the endoglucanase CelA belongs to family A of cellulases as described before (N. R. Gilkes, B. Henrissat, D. G. Kilburn, R. C. Miller, Jr., and R. A. J. Warren, Microbiol. Rev. 55:303-315, 1991.

Amino Acid Sequence

Isolation of a cDNA coding for human galactosyltransferase.

Human milk galactosyltransferase (EC 2.4.1.22) was purified to homogeneity using affinity chromatography. Edman degradation was used to determine the amino acid sequences of eight peptide fragments isolated from the purified enzyme. A 60-mer "optimal" oligonucleotide probe that corresponded to the amino acid sequence of one of the galactosyltransferase peptide fragments was constructed and used to screen a lambda gt10 cDNA library. Two hybridization-positive recombinant phages, each with a 1.7 Kbp insert, were detected among 3 X 10(6) recombinant lambda gt10 phages. Sequencing of one of the cDNA inserts revealed a 783 bp galactosyltransferase coding sequence. The remainder of the sequence corresponded to the 3'-region of the mRNA downstream from the termination codon.

Amino Acid Sequence

Selecting high-affinity binding proteins by monovalent phage display.

Variants of human growth hormone (hGH) with increased affinity and specificity for the hGH receptor were isolated using an improved phage display system. Nearly one million random mutants of hGH were generated at 12 sites previously shown to modulate binding to the hGH receptor or human prolactin (hPRL) receptor. The mutant hormones were displayed in a monovalent fashion from filamentous phage particles as fusions to the gene III product of M13 packaged within each particle. After three to six cycles of enrichment for hGH-phage particles that bound to hGH receptor beads, we isolated hGH mutants that exhibited consensus binding sequences for the hGH receptor. Residues previously identified as important for hGH receptor binding by alanine-scanning mutagenesis were more highly conserved by this selection method. However, other residues nearby were not optimal, and by mutating them, hormone variants having greater affinity and selectivity for the hGH receptor were isolated. This approach should be useful for those who wish to modify and understand the energetics of protein-ligand interfaces.

Amino Acid Sequence

Site dependent time optimization of protein synthesis with special regard to accuracy.

The efficiency of protein synthesis is determined by its rate, accuracy, and energy consumption. With the energy consumption fixed, we optimize the system with respect to time and accuracy. Using an analytic model for a simple system and computer simulations for more complex systems, where also the possibility of errors is included, we demonstrate how different parts of the messenger RNA influence the protein production rate differently. The first part of the coding sequence is of major importance, since the availability of empty initiation sites is crucial, and queuing back to that region may interfere with initiation. The elongation rate at different positions depends on codon usage, on the concentrations of substrate and co-factors, and on the kinetic rate constants, including those of the proofreading branch(es). Ribosomal proofreading is a time consuming process and by allowing for more errors in the beginning of a protein, it is possible to increase the production rate of that protein. We calculate the mean translation time per functioning protein for various translation accuracies, and discuss the different strategies open to living cells.

Algorithms