Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Sense codons are found in specific contexts.

The sequence environment of codons in structural genes has been investigated statistically, using computer methods. A set of Escherichia coli genes with abundant products was compared with a set having low gene product levels, in order to detect potential differences associated with expression. The results show striking non-randomness in the nucleotides occurring near codons. These effects are, unexpectedly, very much larger and more homogeneous among the genes with rare products. The intensity of effects in weakly expressed genes suggests that such non-random sequence environments decrease expression. In the weakly expressed set of genes, the 5' neighbor of a codon, and all positions of the 3' neighbor codon are biased. In the highly expressed genes, the first nucleotide of the next codon is a uniquely affected site. The distribution of non-randomness in weakly expressed genes suggests that sequence bias is primarily due to a constraint acting directly on the secondary or tertiary structure of the codon/anticodon. In highly expressed genes, the observed bias suggests an interaction between the codon/anticodon and a site outside the codon/anticodon. Much of the tendency to non-random near-neighbor sequences in weakly expressed genes can be ascribed to a correlation between nearby nucleotides and the wobble nucleotide of the codon, despite the fact that selection of such correlations will alter the amino acid sequence. The favored pattern, in genes expressed at low level, is R YYR or Y RRY. R indicates purine, Y indicates pyrimidine; the space is the boundary between codons. It seems likely that this preference for nearby sequences is the physical basis of the genetic context effect. Under this assumption such sequence biases will affect expression. On this basis, we predict new sites for contextual mutations which decrease expression, and suggest strategy for the design of messages having optimal translational activity.

Amino Acids↗

Translational pauses during the synthesis of proteins and mRNA structure.

Translational pauses are observed during a spider fibroin synthesis (1,2). The spider major ampullate (dragline) silk of the spider Nephila clavipes is composed of multiple proteins. The amino acid sequences of the partial cDNA clones for the two major dragline silk fibroin components (Spidroin 1 and 2) exhibit repetitive motifs (3,4). Our detailed inspection of the nucleotide sequences of the repetitive motifs revealed highly selective site-specific codon usage patterns within a motif, suggesting that the secondary structure of the spider fibroin mRNA is optimized by the nucleotide sequence of the fibroin gene. The results, combined with our preceding results on silk fibroin from Bombyx mori (5) suggest that translational pauses of spider silk are interpreted in terms of the mRNA secondary structure.

Amino Acid Sequence↗

Evolutionary divergence and salinity-mediated selection in halophilic archaea.

Halophilic (literally salt-loving) archaea are a highly evolved group of organisms that are uniquely able to survive in and exploit hypersaline environments. In this review, we examine the potential interplay between fluctuations in environmental salinity and the primary sequence and tertiary structure of halophilic proteins. The proteins of halophilic archaea are highly adapted and magnificently engineered to function in an intracellular milieu that is in ionic balance with an external environment containing between 2 and 5 M inorganic salt. To understand the nature of halophilic adaptation and to visualize this interplay, the sequences of genes encoding the L11, L1, L10, and L12 proteins of the large ribosome subunit and Mn/Fe superoxide dismutase proteins from three genera of halophilic archaea have been aligned and analyzed for the presence of synonymous and nonsynonymous nucleotide substitutions. Compared to homologous eubacterial genes, these halophilic genes exhibit an inordinately high proportion of nonsynonymous nucleotide substitutions that result in amino acid replacement in the encoded proteins. More than one-third of the replacements involve acidic amino acid residues. We suggest that fluctuations in environmental salinity provide the driving force for fixation of the excessive number of nonsynonymous substitutions. Tinkering with the number, location, and arrangement of acidic and other amino acid residues influences the fitness (i.e., hydrophobicity, surface hydration, and structural stability) of the halophilic protein. Tinkering is also evident at halophilic protein positions monomorphic or polymorphic for serine; more than one-third of these positions use both the TCN and the AGY serine codons, indicating that there have been multiple nonsynonymous substitutions at these positions. Our model suggests that fluctuating environmental salinity prevents optimization of fitness for many halophilic proteins and helps to explain the unusual evolutionary divergence of their encoding genes.

Archaea↗

Construction and expression of nonsense suppressor tRNAs which function in plant cells.

An Arabidopsis thaliana L. DNA containing the tRNA(TrpUGG) gene was isolated and altered to encode the amber suppressor tRNA(TrpUAG) or the ochre suppressor tRNA(TrpUAA). These DNAs were electroporated into carrot protoplasts and tRNA expression was demonstrated by the translational suppression of amber and ochre nonsense mutations in the chloramphenicol acetyltransferase (CAT) reporter gene. DNAs encoding tRNA(TrpUAG) and tRNA(TrpUAA) nonsense suppressor tRNAs caused suppression of their cognate nonsense codons in CAT mRNAs, with the tRNA(TrpUAG) gene exhibiting the greater suppression under optimal conditions for expression of CAT. The development of these translational suppressors which function in plant cells facilitates the study of plant tRNA gene expression and will make possible the manipulation of plant protein structure and function.

Anticodon↗

Sequence analysis, expression, and binding activity of recombinant major outer sheath protein (Msp) of Treponema denticola.

The gene encoding the major outer sheath protein (Msp) of the oral spirochete Treponema denticola ATCC 35405 was cloned, sequenced, and expressed in Escherichia coli. Preliminary sequence analysis showed that the 5' end of the msp gene was not present on the 5.5-kb cloned fragment described in a recent study (M. Haapasalo, K. H. Müller, V. J. Uitto, W. K. Leung, and B. C. McBride, Infect. Immun. 60:2058-2065,1992). The 5' end of msp was obtained by PCR amplification from a T. denticola genomic library, and an open reading frame of 1,629 bp was identified as the coding region for Msp by combining overlapping sequences. The deduced peptide consisted of 543 amino acids and had a molecular mass of 58,233 Da. The peptide had a typical prokaryotic signal sequence with a potential cleavage site for signal peptidase 1. Northern (RNA) blot analysis showing the msp transcript to be approximately 1.7 kb was consistent with the identification of a promoter consensus sequence located optimally upstream of msp and a transcription termination signal found downstream of the stop codon. The entire msp sequence was amplified from T. denticola genomic DNA and cloned in E. coli by using a tightly regulated T7 RNA polymerase vector system. Expression of Msp was toxic to E. coli when the entire msp gene was present. High levels of Msp were produced as inclusion bodies when the putative signal peptide sequence was deleted and replaced by a vector-encoded T7 peptide sequence. Recombinant Msp purified to homogeneity from a clone containing the full-length msp gene adhered to immobilized laminin and fibronectin but not to bovine serum albumin. Attachment of recombinant Msp was decreased in the presence of soluble substrate. Attachment of T. denticola to immobilized laminin and fibronectin was increased by pretreatment of the substrate with recombinant Msp. These studies lend further support to the hypothesis that Msp mediates the extracellular matrix binding activity of T. denticola.

Amino Acid Sequence↗

Monitoring of residual disease in non-Hodgkin's lymphomas by quantitative PCR (preliminary report).

Highly sensitive PCR techniques are often used in molecular monitoring of hematological malignancies, and a quantification of residual disease is important for further prognosis. Here, the limiting dilution methodology and the multiplex IgH/ras PCR are proposed as approaches to molecular monitoring of NHLs. Applying the limiting dilution methodology as a simple dose-response assay for the translocation t(14,18) and CDR3 clonal rearrangement of IgH, critical amounts of total cells determined with stored consecutive diagnostic samples in the same PCR run are compared. Assuming that specific targets are diluted proportionally in dilution of total genomic DNA, the samples showing lower critical concentrations of total DNA are considered as containing higher portion of cells possessing the specific disease marker and vice versa. So far, the correlation of results with the disease outcome confirmed that this simple semi-quantitative approach may in some cases substitute laborious precisely quantifying techniques in the monitoring of the disease. In optimized multiplex IgH/ras PCR co-amplifying clonal CDR3 rearrangement of IgH and the codon 61 of Hras 1 gene, the amount of CDR3 product as the disease marker is related to the ras product as a standard marker of all cells, and quantitative results are obtained by software analyses of detecting gels. Presumably, both approaches may provide clinically useful information on the disease activity and treatment outcome.

Blood Cells↗

Codon distribution in vertebrate genes may be used to predict gene length.

I have analysed the coding regions of 96 eukaryotic genes for their use of iso-coding codons. Specific codons occur more frequently in specific positions in all members of some gene families than would be expected if codon choice was determined solely by the frequency of codon usage. In the absence of evidence a priori for selection for particular codons at particular positions, I term such co-occurring codons "coincident codons". Coincident codons are not confined to particular regions of genes, and their occurrence is not detectably linked with the location of introns in the genomic sequence. Their presence is partly but not completely explained by the exchange of sequence between similar functional genes within a species: homologous genes from different organisms also possess the same codons at some sites with greater than expected frequencies. The relative excess of coincident codons correlates well with the overall length of the genes analysed, but not with the length of mRNA or coding regions, or with qualitative features of gene structure or expression. This, and the unusual sequence environment of coincident codons, suggests that they are a feature of the overall secondary structure of the heterogeneous nuclear RNA. Such considerations suggest approaches for optimizing the expression of exogenous genes in eukaryotic systems, and for predicting the structure of genes for which only partial sequence data is available.

Actins↗

Bacterial start site prediction.

With the growing number of completely sequenced bacterial genes, accurate gene prediction in bacterial genomes remains an important problem. Although the existing tools predict genes in bacterial genomes with high overall accuracy, their ability to pinpoint the translation start site remains unsatisfactory. In this paper, we present a novel approach to bacterial start site prediction that takes into account multiple features of a potential start site, viz., ribosome binding site (RBS) binding energy, distance of the RBS from the start codon, distance from the beginning of the maximal ORF to the start codon, the start codon itself and the coding/non-coding potential around the start site. Mixed integer programing was used to optimize the discriminatory system. The accuracy of this approach is up to 90%, compared to 70%, using the most common tools in fully automated mode (that is, without expert human post-processing of results). The approach is evaluated using Bacillus subtilis, Escherichia coli and Pyrococcus furiosus. These three genomes cover a broad spectrum of bacterial genomes, since B.subtilis is a Gram-positive bacterium, E.coli is a Gram-negative bacterium and P. furiosus is an archaebacterium. A significant problem is generating a set of 'true' start sites for algorithm training, in the absence of experimental work. We found that sequence conservation between P. furiosus and the related Pyrococcus horikoshii clearly delimited the gene start in many cases, providing a sufficient training set.

Algorithms↗

Maximum likelihood estimation on large phylogenies and analysis of adaptive evolution in human influenza virus A.

Algorithmic details to obtain maximum likelihood estimates of parameters on a large phylogeny are discussed. On a large tree, an efficient approach is to optimize branch lengths one at a time while updating parameters in the substitution model simultaneously. Codon substitution models that allow for variable nonsynonymous/synonymous rate ratios (omega = d(N)/d(S)) among sites are used to analyze a data set of human influenza virus type A hemagglutinin (HA) genes. The data set has 349 sequences. Methods for obtaining approximate estimates of branch lengths for codon models are explored, and the estimates are used to test for positive selection and to identify sites under selection. Compared with results obtained from the exact method estimating all parameters by maximum likelihood, the approximate methods produced reliable results. The analysis identified a number of sites in the viral gene under diversifying Darwinian selection and demonstrated the importance of including many sequences in the data in detecting positive selection at individual sites.

Algorithms↗

Evolutionary protein stabilization in comparison with computational design.

Two major strategies are currently used for stabilizing proteins: in vitro evolution and computational design. Here, we used gene libraries of the beta1 domain of the streptococcal protein G (Gbeta1) and Proside, an in vitro selection method, to identify stabilized variants of this protein. In the Gbeta1 libraries, the codons for the four boundary positions 16, 18, 25, and 29 were randomized. Many Gbeta1 variants with strongly increased thermal stabilities were found in 11 selections performed with five independent libraries. Previously, Mayo and co-workers used computational design to stabilize Gbeta1 by sequence optimization at the same positions. Their best variant ranked third within the panel of the selected variants. None of the ten computed sequences was found in the Proside selections, because several computed residues for positions 18 and 29 were not optimal for stability.

Bacterial Proteins↗

A stable disulfide-free gene-3-protein of phage fd generated by in vitro evolution.

Disulfide bonds provide major contributions to the conformational stability of proteins, and their cleavage often leads to unfolding. The gene-3-protein of the filamentous phage fd contains two disulfides in its N1 domain and one in its N2 domain, and these three disulfide bonds are essential for the stability of this protein. Here, we employed in vitro evolution to generate a disulfide-free variant of the N1-N2 protein with a high conformational stability. The gene-3-protein is essential for the phage infectivity, and we exploited this requirement for a proteolytic selection of stabilized protein variants from phage libraries. First, optimal replacements for individual disulfide bonds were identified in libraries, in which the corresponding cysteine codons were randomized. Then stabilizing amino acid replacements at non-cysteine positions were selected from libraries that were created by error-prone PCR. This stepwise procedure led to variants of N1-N2 that are devoid of all three disulfide bonds but stable and functional. The best variant without disulfide bonds showed a much higher conformational stability than the disulfide-containing wild-type form of the gene-3-protein. Despite the loss of all three disulfide bonds, the midpoints of the thermal transitions were increased from 48.5 degrees C to 67.0 degrees C for the N2 domain and from 60.0 degrees C to 78.7 degrees C for the N1 domain. The major loss in conformational stability caused by the removal of the disulfides was thus over-compensated by strongly improved non-covalent interactions. The stabilized variants were less infectious than the wild-type protein, probably because the domain mobility was reduced. Only a small fraction of the sequence space could be accessed by using libraries created by error-prone PCR, but still many strongly stabilized variants could be identified. This is encouraging and indicates that proteins can be stabilized by mutations in many different ways.

Bacteriophage M13↗

Approaches to enhance the efficacy of DNA vaccines.

DNA vaccines consist of antigen-encoding bacterial plasmids that are capable of inducing antigen-specific immune responses upon inoculation into a host. This method of immunization is advantageous in terms of simplicity, adaptability, and cost of vaccine production. However, the entry of DNA vaccines and expression of antigen are subjected to physical and biochemical barriers imposed by the host. In small animals such as mice, the host-imposed impediments have not prevented DNA vaccines from inducing long-lasting, protective humoral, and cellular immune responses. In contrast, these barriers appear to be more difficult to overcome in large animals and humans. The focus of this article is to summarize the limitations of DNA vaccines and to provide a comprehensive review on the different strategies developed to enhance the efficacy of DNA vaccines. Several of these strategies, such as altering codon bias of the encoded gene, changing the cellular localization of the expressed antigen, and optimizing delivery and formulation of the plasmid, have led to improvements in DNA vaccine efficacy in large animals. However, solutions for increasing the amount of plasmid that eventually enters the nucleus and is available for transcription of the transgene still need to be found. The overall conclusions from these studies suggest that, provided these critical improvements are made, DNA vaccines may find important clinical and practical applications in the field of vaccination.

Animals↗

PCR-based gene synthesis as an efficient approach for expression of the A+T-rich malaria genome.

The A+T-rich genome of the human malaria parasite Plasmodium falciparum encodes genes of biological importance that cannot be expressed efficiently in heterologous eukaryotic systems, owing to an extremely biased codon usage and the presence of numerous cryptic polyadenylation sites. In this work we have optimized an assembly polymerase chain reaction (PCR) method for the fast and extremely accurate synthesis of a 2.1 kb Plasmodium falciparum gene (pfsub-1) encoding a subtilisin-like protease. A total of 104 oligonucleotides, designed with the aid of dedicated computer software, were assembled in a single-step PCR. The assembly was then further amplified by PCR to produce a synthetic gene which has been cloned and successfully expressed in both Pichia pastoris and recombinant baculovirus-infected High Five(TM) cells. We believe this strategy to be of special interest as it is simple, accessible and has no limitation with respect to the size of the gene to be synthesized. Used as a systematic approach for the malarial genome or any other A + T-rich organism, the method allows the rapid synthesis of a nucleotide sequence optimized for expression in the system of choice and production of sufficiently large amounts of biological material for complete molecular and structural characterization.

Amino Acid Sequence↗

Enhanced heterologous expression of two Streptomyces griseolus cytochrome P450s and Streptomyces coelicolor ferredoxin reductase as potentially efficient hydroxylation catalysts.

The herbicide-inducible, soluble cytochrome P450s CYP105A1 and CYP105B1 and their adjacent ferredoxins, Fd1 and Fd2, of Streptomyces griseolus were expressed in Escherichia coli to high levels. Conditions for high-level expression of active enzyme able to catalyze hydroxylation have been developed. Analysis of the expression levels of the P450 proteins in several different E. coli expression hosts identified E. coli BL21 Star(DE3)pLysS as the optimal host cell to express CYP105B1 as judged by CO difference spectra. Examination of the codons used in the CYP1051A1 sequence indicated that it contains a number of codons corresponding to rare E. coli tRNA species. The level of its expression was improved in the modified forms of E. coli BL21(DE3), which contain extra copies of rare codon E. coli tRNA genes. The activity of correctly folded cytochrome P450s was further enhanced by cloning a ferredoxin reductase from Streptomyces coelicolor downstream of CYP105A1 and CYP105B1 and their adjacent ferredoxins. Expression of CYP105A1 and CYP105B1 was also achieved in Streptomyces lividans 1326 by cloning the P450 genes and their ferredoxins into the expression vector pBW160. S. lividans 1326 cells containing CYP105A1 or CYP105B1 were able efficiently to dealkylate 7-ethoxycoumarin.

Bacterial Proteins↗

Antiretroviral resistance mutations in human immunodeficiency virus type 1 infected patients enrolled in genotype testing at the Central Public Health Laboratory, São Paulo, Brazil: preliminary results.

Antiretroviral resistance mutations (ARM) are one of the major obstacles for pharmacological human immunodeficiency virus (HIV) suppression. Plasma HIV-1 RNA from 306 patients on antiretroviral therapy with virological failure was analyzed, most of them (60%) exposed to three or more regimens, and 28% of them have started therapy before 1997. The most common regimens in use at the time of genotype testing were AZT/3TC/nelfinavir, 3TC/D4T/nelfinavir and AZT/3TC/efavirenz. The majority of ARM occurred at protease (PR) gene at residue L90 (41%) and V82 (25%); at reverse transcriptase (RT) gene, mutations at residue M184 (V/I) were observed in 64%. One or more thymidine analogue mutations were detected in 73%. The number of ARM at PR gene increased from a mean of four mutations per patient who showed virological failure at the first ARV regimens to six mutations per patient exposed to six or more regimens; similar trend in RT was also observed. No differences in ARM at principal codon to the three drug classes for HIV-1 clades B or F were observed, but some polymorphisms in secondary codons showed significant differences. Strategies to improve the cost effectiveness of drug therapy and to optimize the sequencing and the rescue therapy are the major health priorities.

Adolescent↗

Cardiac troponin I sense-antisense RNA duplexes in the myocardium.

Natural antisense RNA is now thought to regulate, at least in part, a growing number of eukaryotic genes. It is becoming increasingly apparent that such endogenous antisense RNA molecules may modulate gene expression in a manner analogous to synthetic oligomers. Here, we report the detection of antisense-orientated RNA transcripts of cardiac specific troponin I in rat and human myocardium. Interestingly, the different sizes of the rat and human antisense cTNI transcripts suggest species-specific reverse transcription initiation sites. Moreover, for the first time in cardiomyocytes, we could demonstrate in vivo duplex formation between sense and antisense transcripts. The existence of antisense-sense duplexes represents compelling evidence and a potential mechanism for endogenous antisense transcript-mediated modulation of mRNA translation. The potential effect of attenuating translation was illustrated by in vitro and in vivo model systems. Testing several oligonucleotides based on the natural antisense sequences, the optimal region for inhibition of translation was identified as being close to the translational start codon.

Adult↗

Possibility of genetic coding of amino acid sequences by coherent electronic states in nucleotide chains.

The concept of coherent electronic states and coherent interactions in supramolecular structures is applied to the process of genetic information coding and its transcription from DNA to mRNA. A new genetic code is proposed based on the assumption of coherent electron states in linear chains of nucleotide bases. A new interpretation of codon equivalency (redundancy) is given. The number of existing amino acids is derived from the optimalization principle applied to the physical system storing the genetic information in the new code. The proposed code uses a variable number of positions or nucleotide bases along the DNA-mRNA structure to code a single amino acid in a protein. The average of this variable number must be equal to the base of natural logarithms (e = 2.7 . . .) in order to minimize the number of nucleotides required to code a sequence of amino acids.

Amino Acid Sequence↗