Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

High-level expression of human superoxide dismutase in the cyanobacterium Anacystis nidulans 6301.

A chemically synthesized gene encoding human CuZn superoxide dismutase (hSOD) was cloned into the shuttle vector pBAX18R and expressed in Anacystis nidulans 6301 (Synechococcus sp. strain PCC 6301) under the control of a ribulose-1,5-bisphosphate carboxylase/oxygenase gene (rbc) promoter derived from A. nidulans 6301. The sequences immediately upstream from the hSOD coding region and the distances between the ribosomal binding site and ATG initiation codon strongly affected the expression of the hSOD gene in A. nidulans cells. Optimal expression of hSOD was obtained with the expression vector pBAXSOD8-I, which contained a GGAGAG sequence. In defined conditions, irradiation with light increased hSOD enzyme activity in the transformants > 18-fold and the level of the hSOD protein reached a value of about 3% of the total soluble protein. The transformants that expressed hSOD acquired the ability to extenuate photooxidative damage induced by methyl viologen.

Amino Acid Sequence↗

Transient transfection and expression of firefly luciferase in Giardia lamblia.

We have developed a gene transfer system for the protozoan parasite Giardia lamblia. This organism is responsible for many cases of diarrhea worldwide and is considered to be one of the most primitive eukaryotes. Expression of a heterologous gene was detected in this parasite after electroporation with appropriate DNA constructs. We constructed a series of transfection plasmids using flanking sequences of the Giardia glutamate dehydrogenase (GDH) gene to drive expression of the firefly luciferase reporter gene. The optimal construct consisted of a GDH/luciferase fusion gene in which the first 18 codons of the GDH gene immediately preceded the luciferase gene; this fusion gene was flanked by the upstream and downstream sequences of the GDH gene. Electroporation of this construct into Giardia yielded luciferase activity that was 3000- to 50,000-fold above background. Removal of either the 5' or 3' GDH flanking sequences from this construct resulted in significantly reduced luciferase activity, and removal of both flanking sequences reduced luciferase activity to background levels. Luciferase activity was proportional to the amount of DNA electroporated and was maximal at 6 hr after electroporation.

Animals↗

Improved fluorescence and dual color detection with enhanced blue and green variants of the green fluorescent protein.

The green fluorescent protein (GFP) from the jellyfish Aequorea victoria is a versatile reporter protein for monitoring gene expression and protein localization in a variety of systems. Applications using GFP reporters have expanded greatly due to the availability of mutants with altered spectral properties, including several blue emission variants, all of which contain the single point mutation Tyr-66 to His in the chromophore region of the protein. However, previously described "BFP" reporters have limited utility, primarily due to relatively dim fluorescence and low expression levels attained in higher eukaryotes with such variants. To improve upon these qualities, we have combined a blue emission mutant of GFP containing four point mutations (Phe-64 to Leu, Ser-65 to Thr, Tyr-66 to His, and Tyr-145 to Phe) with a synthetic gene sequence containing codons preferentially found in highly expressed human proteins. These mutations were chosen to optimize expression of properly folded fluorescent protein in mammalian cells cultured at 37 degreesC and to maximize signal intensity. The combination of improved fluorescence and higher expression levels yield an enhanced blue fluorescent protein that provides greater sensitivity and is suitable for dual color detection with green-emitting fluorophores.

Fluorescence↗

Evolutionary dynamics of insertion sequences in Helicobacter pylori.

Prokaryotic insertion sequence (IS) elements behave like parasites in terms of their ability to invade and proliferate in microbial gene pools and like symbionts when they coevolve with their bacterial hosts. Here we investigated the evolutionary history of IS605 and IS607 of Helicobacter pylori, a genetically diverse gastric pathogen. These elements contain unrelated transposase genes (orfA) and also a homolog of the Salmonella virulence gene gipA (orfB). A total of 488 East Asian, Indian, Peruvian, and Spanish isolates were screened, and 18 and 14% of them harbored IS605 and IS607, respectively. IS605 nucleotide sequence analysis (n = 42) revealed geographic subdivisions similar to those of H. pylori; the geographic subdivision was blurred, however, due in part to homologous recombination, as indicated by split decomposition and homoplasy tests (homoplasy ratio, 0.56). In contrast, the IS607 populations (n = 44) showed strong geographic subdivisions with less homologous recombination (homoplasy ratio, 0.2). Diversifying selection (ratio of nonsynonymous change to synonymous change, >>1) was evident in approximately 15% of the IS605 orfA codons analyzed but not in the IS607 orfA codons. Diversifying selection was also evident in approximately 2% of the IS605 orfB and approximately 10% of the IS607 orfB codons analyzed. We suggest that the evolution of these elements reflects selection for optimal transposition activity in the case of IS605 orfA and for interactions between the OrfB proteins and other cellular constituents that potentially contribute to bacterial fitness. Taken together, similarities in IS elements and H. pylori population genetic structures and evidence of adaptive evolution in IS elements suggest that there is coevolution between these elements and their bacterial hosts.

DNA Transposable Elements↗

[Studies on fermentation conditions and purification of mutant human interleukin-2 expressed in Pichia pastoris].

Interleukin-2 (IL-2) was initially isolated as a T cell growth factor and had been shown to direct the expansion and differentiation of several hematopoietic cell types. Clinical studies using IL-2 in the treatment of AIDS have been encouraging, due to its critical role as a proliferative signal for activated T-lymphocytes. IL-2 has also undergone trials in the treatment of several types of cancer, based on its stimulation of cytotoxic, antitumor cells. Today, human IL-2 is produced completely by genetically engineered method, and it has been proved that genetically engineered recombinant human IL-2 has almost the same function and clinical effect as wild IL-2. In the former study, recombinant human IL-2 usually comes from E. coli, in this paper the mutant IL-2 was successfully expressed and purified in Pichia pastoris for the first time. As a eukaryote, Pichia pastoris has many of the advantages of higher eukaryotic expression systems such as protein processing, protein folding, and posttranslational modification, while being as easy to manipulate as E. coli or Saccharomyces cerevisiae. It is faster, easier, and less expensive to use than other eukaryotic expression systems such as baculovirus or mammalian tissue culture, and generally gives higher expression level. Expression conditions of human mutant interleukin-2(the codon for cysteine-125 of human IL-2 with alanine; the codon for leucine-18 with methionine; the codon for leucine-19 with serine) in the recombinant Pichia pastoris strain were optimized via test of some factors such as the rate of aeration, the inductive duration, the initial pH and the concentration of methanol. The results from tests showed that the most important parameter for efficient expression of interleukin-2 in recombinant Pichia pastoris strain is adequate aeration during methanol induction, and the optimum inductive condition for interleukin-2 expression was: more than 80% aeration, 2 days for induction, the initial pH of 6.0, the final methanol concentration of 1.0%. With this condition, the expressed IL-2 was secreted into fermentation broth and reached a yield of 30%, approximately 200 mg/L. Expressed interleutin-2 (MvIL-2) was isolated and purified by centrifugation, millipore filtration to concentration, Econo-PacS strongly acidic cation exchanger cartridge and molecular sieve chromatography and the yield of MvIL-2 was 27%. MvIL-2 was purified to electrophoretic purity by SDS-PAGE and only one peak being loaded on HPLC. Purified MvIL-2 protein had stimulating activity similar to the wild type of IL-2 as assayed by IL-2-dependent CTLL-2 cells. However, the stability of MvIL-2 was superior than that of IL-2 at different temperatures. The activity of obtained MvIL-2 was 4 - 5 times of the wild type of IL-2, So MvIL-2 had an advantage over wild type of rhIL-2 in storage stability and activity.

Fermentation↗

Computer-aided gene design.

A computer program, which runs on MS-DOS personal computers, is described that assists in the design of synthetic genes coding for proteins. The goal of the program is the design of a gene which (i) contains as many unique restriction sites as possible and (ii) uses a specific codon usage. The gene designed according to the criteria above is (i) suitable for 'modular mutagenesis' experiments and (ii) optimized for expression. The program 'reverse-translates' protein sequences into degenerated DNA sequences, generates a map of potential restriction sites and locates sequence positions where unique restriction sites can be accommodated. The nucleic acid sequence is then 'refined' according to a specific codon usage to remove any degeneration. Unique restriction sites, if potentially present, can be 'forced' into the degenerated nucleic acid sequence by using 'priority codes' assigned to different restriction sequences.

Amino Acid Sequence↗

Sense codons are found in specific contexts.

The sequence environment of codons in structural genes has been investigated statistically, using computer methods. A set of Escherichia coli genes with abundant products was compared with a set having low gene product levels, in order to detect potential differences associated with expression. The results show striking non-randomness in the nucleotides occurring near codons. These effects are, unexpectedly, very much larger and more homogeneous among the genes with rare products. The intensity of effects in weakly expressed genes suggests that such non-random sequence environments decrease expression. In the weakly expressed set of genes, the 5' neighbor of a codon, and all positions of the 3' neighbor codon are biased. In the highly expressed genes, the first nucleotide of the next codon is a uniquely affected site. The distribution of non-randomness in weakly expressed genes suggests that sequence bias is primarily due to a constraint acting directly on the secondary or tertiary structure of the codon/anticodon. In highly expressed genes, the observed bias suggests an interaction between the codon/anticodon and a site outside the codon/anticodon. Much of the tendency to non-random near-neighbor sequences in weakly expressed genes can be ascribed to a correlation between nearby nucleotides and the wobble nucleotide of the codon, despite the fact that selection of such correlations will alter the amino acid sequence. The favored pattern, in genes expressed at low level, is R YYR or Y RRY. R indicates purine, Y indicates pyrimidine; the space is the boundary between codons. It seems likely that this preference for nearby sequences is the physical basis of the genetic context effect. Under this assumption such sequence biases will affect expression. On this basis, we predict new sites for contextual mutations which decrease expression, and suggest strategy for the design of messages having optimal translational activity.

Amino Acids↗

Translational pauses during the synthesis of proteins and mRNA structure.

Translational pauses are observed during a spider fibroin synthesis (1,2). The spider major ampullate (dragline) silk of the spider Nephila clavipes is composed of multiple proteins. The amino acid sequences of the partial cDNA clones for the two major dragline silk fibroin components (Spidroin 1 and 2) exhibit repetitive motifs (3,4). Our detailed inspection of the nucleotide sequences of the repetitive motifs revealed highly selective site-specific codon usage patterns within a motif, suggesting that the secondary structure of the spider fibroin mRNA is optimized by the nucleotide sequence of the fibroin gene. The results, combined with our preceding results on silk fibroin from Bombyx mori (5) suggest that translational pauses of spider silk are interpreted in terms of the mRNA secondary structure.

Amino Acid Sequence↗

Evolutionary divergence and salinity-mediated selection in halophilic archaea.

Halophilic (literally salt-loving) archaea are a highly evolved group of organisms that are uniquely able to survive in and exploit hypersaline environments. In this review, we examine the potential interplay between fluctuations in environmental salinity and the primary sequence and tertiary structure of halophilic proteins. The proteins of halophilic archaea are highly adapted and magnificently engineered to function in an intracellular milieu that is in ionic balance with an external environment containing between 2 and 5 M inorganic salt. To understand the nature of halophilic adaptation and to visualize this interplay, the sequences of genes encoding the L11, L1, L10, and L12 proteins of the large ribosome subunit and Mn/Fe superoxide dismutase proteins from three genera of halophilic archaea have been aligned and analyzed for the presence of synonymous and nonsynonymous nucleotide substitutions. Compared to homologous eubacterial genes, these halophilic genes exhibit an inordinately high proportion of nonsynonymous nucleotide substitutions that result in amino acid replacement in the encoded proteins. More than one-third of the replacements involve acidic amino acid residues. We suggest that fluctuations in environmental salinity provide the driving force for fixation of the excessive number of nonsynonymous substitutions. Tinkering with the number, location, and arrangement of acidic and other amino acid residues influences the fitness (i.e., hydrophobicity, surface hydration, and structural stability) of the halophilic protein. Tinkering is also evident at halophilic protein positions monomorphic or polymorphic for serine; more than one-third of these positions use both the TCN and the AGY serine codons, indicating that there have been multiple nonsynonymous substitutions at these positions. Our model suggests that fluctuating environmental salinity prevents optimization of fitness for many halophilic proteins and helps to explain the unusual evolutionary divergence of their encoding genes.

Archaea↗

Construction and expression of nonsense suppressor tRNAs which function in plant cells.

An Arabidopsis thaliana L. DNA containing the tRNA(TrpUGG) gene was isolated and altered to encode the amber suppressor tRNA(TrpUAG) or the ochre suppressor tRNA(TrpUAA). These DNAs were electroporated into carrot protoplasts and tRNA expression was demonstrated by the translational suppression of amber and ochre nonsense mutations in the chloramphenicol acetyltransferase (CAT) reporter gene. DNAs encoding tRNA(TrpUAG) and tRNA(TrpUAA) nonsense suppressor tRNAs caused suppression of their cognate nonsense codons in CAT mRNAs, with the tRNA(TrpUAG) gene exhibiting the greater suppression under optimal conditions for expression of CAT. The development of these translational suppressors which function in plant cells facilitates the study of plant tRNA gene expression and will make possible the manipulation of plant protein structure and function.

Anticodon↗

Inorganic polyphosphate interacts with ribosomes and promotes translation fidelity in vitro and in vivo.

Inorganic polyphosphate is a biological macromolecule consisting of multiple phosphates linked by high-energy bonds. Polyphosphate occurs in cells from all domains of life, and is known to play roles in a diverse collection of cellular functions. Here we examine the relationship between polyphosphate and protein synthesis in Escherichia coli. We report that polyphosphate associates with E. coli ribosomes in vitro. Characterization of this interaction reveals that both long-chain and short-chain polyphosphates interact with the ribosome. Intact 70S ribosomes, as well as 50S and 30S subunits, display a specific interaction with polyphosphate that is mediated primarily by contacts with ribosomal proteins. Additionally, we examined functional consequences of a ppk mutation, which severely reduces levels of intracellular polyphosphate. Extracts from ppk mutants contain lower levels of polysomes than wild-type cells, suggesting a defect in mRNA utilization or the mRNA-ribosome interaction. Ribosomes from wild-type and ppk mutant cells were isolated, and their activities were compared using a polyU RNA in vitro translation assay. While rates of polyphenylalanine synthesis are similar, use of ribosomes from ppk cells results in a misincorporation rate about five times higher compared with the rate observed when ribosomes from wild-type cells are used. Mistranslation rates in vivo were measured directly, and ppk mutants displayed higher readthrough frequencies for two different stop codons. Taken together, these results indicate that polyphosphate plays an important role in maintaining optimal translation efficiency in vivo and in vitro.

Escherichia coli↗

Sequence analysis, expression, and binding activity of recombinant major outer sheath protein (Msp) of Treponema denticola.

The gene encoding the major outer sheath protein (Msp) of the oral spirochete Treponema denticola ATCC 35405 was cloned, sequenced, and expressed in Escherichia coli. Preliminary sequence analysis showed that the 5' end of the msp gene was not present on the 5.5-kb cloned fragment described in a recent study (M. Haapasalo, K. H. Müller, V. J. Uitto, W. K. Leung, and B. C. McBride, Infect. Immun. 60:2058-2065,1992). The 5' end of msp was obtained by PCR amplification from a T. denticola genomic library, and an open reading frame of 1,629 bp was identified as the coding region for Msp by combining overlapping sequences. The deduced peptide consisted of 543 amino acids and had a molecular mass of 58,233 Da. The peptide had a typical prokaryotic signal sequence with a potential cleavage site for signal peptidase 1. Northern (RNA) blot analysis showing the msp transcript to be approximately 1.7 kb was consistent with the identification of a promoter consensus sequence located optimally upstream of msp and a transcription termination signal found downstream of the stop codon. The entire msp sequence was amplified from T. denticola genomic DNA and cloned in E. coli by using a tightly regulated T7 RNA polymerase vector system. Expression of Msp was toxic to E. coli when the entire msp gene was present. High levels of Msp were produced as inclusion bodies when the putative signal peptide sequence was deleted and replaced by a vector-encoded T7 peptide sequence. Recombinant Msp purified to homogeneity from a clone containing the full-length msp gene adhered to immobilized laminin and fibronectin but not to bovine serum albumin. Attachment of recombinant Msp was decreased in the presence of soluble substrate. Attachment of T. denticola to immobilized laminin and fibronectin was increased by pretreatment of the substrate with recombinant Msp. These studies lend further support to the hypothesis that Msp mediates the extracellular matrix binding activity of T. denticola.

Amino Acid Sequence↗

Monitoring of residual disease in non-Hodgkin's lymphomas by quantitative PCR (preliminary report).

Highly sensitive PCR techniques are often used in molecular monitoring of hematological malignancies, and a quantification of residual disease is important for further prognosis. Here, the limiting dilution methodology and the multiplex IgH/ras PCR are proposed as approaches to molecular monitoring of NHLs. Applying the limiting dilution methodology as a simple dose-response assay for the translocation t(14,18) and CDR3 clonal rearrangement of IgH, critical amounts of total cells determined with stored consecutive diagnostic samples in the same PCR run are compared. Assuming that specific targets are diluted proportionally in dilution of total genomic DNA, the samples showing lower critical concentrations of total DNA are considered as containing higher portion of cells possessing the specific disease marker and vice versa. So far, the correlation of results with the disease outcome confirmed that this simple semi-quantitative approach may in some cases substitute laborious precisely quantifying techniques in the monitoring of the disease. In optimized multiplex IgH/ras PCR co-amplifying clonal CDR3 rearrangement of IgH and the codon 61 of Hras 1 gene, the amount of CDR3 product as the disease marker is related to the ras product as a standard marker of all cells, and quantitative results are obtained by software analyses of detecting gels. Presumably, both approaches may provide clinically useful information on the disease activity and treatment outcome.

Blood Cells↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Bacterial start site prediction.

With the growing number of completely sequenced bacterial genes, accurate gene prediction in bacterial genomes remains an important problem. Although the existing tools predict genes in bacterial genomes with high overall accuracy, their ability to pinpoint the translation start site remains unsatisfactory. In this paper, we present a novel approach to bacterial start site prediction that takes into account multiple features of a potential start site, viz., ribosome binding site (RBS) binding energy, distance of the RBS from the start codon, distance from the beginning of the maximal ORF to the start codon, the start codon itself and the coding/non-coding potential around the start site. Mixed integer programing was used to optimize the discriminatory system. The accuracy of this approach is up to 90%, compared to 70%, using the most common tools in fully automated mode (that is, without expert human post-processing of results). The approach is evaluated using Bacillus subtilis, Escherichia coli and Pyrococcus furiosus. These three genomes cover a broad spectrum of bacterial genomes, since B.subtilis is a Gram-positive bacterium, E.coli is a Gram-negative bacterium and P. furiosus is an archaebacterium. A significant problem is generating a set of 'true' start sites for algorithm training, in the absence of experimental work. We found that sequence conservation between P. furiosus and the related Pyrococcus horikoshii clearly delimited the gene start in many cases, providing a sufficient training set.

Algorithms↗

Maximum likelihood estimation on large phylogenies and analysis of adaptive evolution in human influenza virus A.

Algorithmic details to obtain maximum likelihood estimates of parameters on a large phylogeny are discussed. On a large tree, an efficient approach is to optimize branch lengths one at a time while updating parameters in the substitution model simultaneously. Codon substitution models that allow for variable nonsynonymous/synonymous rate ratios (omega = d(N)/d(S)) among sites are used to analyze a data set of human influenza virus type A hemagglutinin (HA) genes. The data set has 349 sequences. Methods for obtaining approximate estimates of branch lengths for codon models are explored, and the estimates are used to test for positive selection and to identify sites under selection. Compared with results obtained from the exact method estimating all parameters by maximum likelihood, the approximate methods produced reliable results. The analysis identified a number of sites in the viral gene under diversifying Darwinian selection and demonstrated the importance of including many sequences in the data in detecting positive selection at individual sites.

Algorithms↗

Evolutionary protein stabilization in comparison with computational design.

Two major strategies are currently used for stabilizing proteins: in vitro evolution and computational design. Here, we used gene libraries of the beta1 domain of the streptococcal protein G (Gbeta1) and Proside, an in vitro selection method, to identify stabilized variants of this protein. In the Gbeta1 libraries, the codons for the four boundary positions 16, 18, 25, and 29 were randomized. Many Gbeta1 variants with strongly increased thermal stabilities were found in 11 selections performed with five independent libraries. Previously, Mayo and co-workers used computational design to stabilize Gbeta1 by sequence optimization at the same positions. Their best variant ranked third within the panel of the selected variants. None of the ten computed sequences was found in the Proside selections, because several computed residues for positions 18 and 29 were not optimal for stability.

Bacterial Proteins↗