Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Maximum likelihood estimation on large phylogenies and analysis of adaptive evolution in human influenza virus A.

Algorithmic details to obtain maximum likelihood estimates of parameters on a large phylogeny are discussed. On a large tree, an efficient approach is to optimize branch lengths one at a time while updating parameters in the substitution model simultaneously. Codon substitution models that allow for variable nonsynonymous/synonymous rate ratios (omega = d(N)/d(S)) among sites are used to analyze a data set of human influenza virus type A hemagglutinin (HA) genes. The data set has 349 sequences. Methods for obtaining approximate estimates of branch lengths for codon models are explored, and the estimates are used to test for positive selection and to identify sites under selection. Compared with results obtained from the exact method estimating all parameters by maximum likelihood, the approximate methods produced reliable results. The analysis identified a number of sites in the viral gene under diversifying Darwinian selection and demonstrated the importance of including many sequences in the data in detecting positive selection at individual sites.

Algorithms↗

Evolutionary protein stabilization in comparison with computational design.

Two major strategies are currently used for stabilizing proteins: in vitro evolution and computational design. Here, we used gene libraries of the beta1 domain of the streptococcal protein G (Gbeta1) and Proside, an in vitro selection method, to identify stabilized variants of this protein. In the Gbeta1 libraries, the codons for the four boundary positions 16, 18, 25, and 29 were randomized. Many Gbeta1 variants with strongly increased thermal stabilities were found in 11 selections performed with five independent libraries. Previously, Mayo and co-workers used computational design to stabilize Gbeta1 by sequence optimization at the same positions. Their best variant ranked third within the panel of the selected variants. None of the ten computed sequences was found in the Proside selections, because several computed residues for positions 18 and 29 were not optimal for stability.

Bacterial Proteins↗

A stable disulfide-free gene-3-protein of phage fd generated by in vitro evolution.

Disulfide bonds provide major contributions to the conformational stability of proteins, and their cleavage often leads to unfolding. The gene-3-protein of the filamentous phage fd contains two disulfides in its N1 domain and one in its N2 domain, and these three disulfide bonds are essential for the stability of this protein. Here, we employed in vitro evolution to generate a disulfide-free variant of the N1-N2 protein with a high conformational stability. The gene-3-protein is essential for the phage infectivity, and we exploited this requirement for a proteolytic selection of stabilized protein variants from phage libraries. First, optimal replacements for individual disulfide bonds were identified in libraries, in which the corresponding cysteine codons were randomized. Then stabilizing amino acid replacements at non-cysteine positions were selected from libraries that were created by error-prone PCR. This stepwise procedure led to variants of N1-N2 that are devoid of all three disulfide bonds but stable and functional. The best variant without disulfide bonds showed a much higher conformational stability than the disulfide-containing wild-type form of the gene-3-protein. Despite the loss of all three disulfide bonds, the midpoints of the thermal transitions were increased from 48.5 degrees C to 67.0 degrees C for the N2 domain and from 60.0 degrees C to 78.7 degrees C for the N1 domain. The major loss in conformational stability caused by the removal of the disulfides was thus over-compensated by strongly improved non-covalent interactions. The stabilized variants were less infectious than the wild-type protein, probably because the domain mobility was reduced. Only a small fraction of the sequence space could be accessed by using libraries created by error-prone PCR, but still many strongly stabilized variants could be identified. This is encouraging and indicates that proteins can be stabilized by mutations in many different ways.

Bacteriophage M13↗

Approaches to enhance the efficacy of DNA vaccines.

DNA vaccines consist of antigen-encoding bacterial plasmids that are capable of inducing antigen-specific immune responses upon inoculation into a host. This method of immunization is advantageous in terms of simplicity, adaptability, and cost of vaccine production. However, the entry of DNA vaccines and expression of antigen are subjected to physical and biochemical barriers imposed by the host. In small animals such as mice, the host-imposed impediments have not prevented DNA vaccines from inducing long-lasting, protective humoral, and cellular immune responses. In contrast, these barriers appear to be more difficult to overcome in large animals and humans. The focus of this article is to summarize the limitations of DNA vaccines and to provide a comprehensive review on the different strategies developed to enhance the efficacy of DNA vaccines. Several of these strategies, such as altering codon bias of the encoded gene, changing the cellular localization of the expressed antigen, and optimizing delivery and formulation of the plasmid, have led to improvements in DNA vaccine efficacy in large animals. However, solutions for increasing the amount of plasmid that eventually enters the nucleus and is available for transcription of the transgene still need to be found. The overall conclusions from these studies suggest that, provided these critical improvements are made, DNA vaccines may find important clinical and practical applications in the field of vaccination.

Animals↗

PCR-based gene synthesis as an efficient approach for expression of the A+T-rich malaria genome.

The A+T-rich genome of the human malaria parasite Plasmodium falciparum encodes genes of biological importance that cannot be expressed efficiently in heterologous eukaryotic systems, owing to an extremely biased codon usage and the presence of numerous cryptic polyadenylation sites. In this work we have optimized an assembly polymerase chain reaction (PCR) method for the fast and extremely accurate synthesis of a 2.1 kb Plasmodium falciparum gene (pfsub-1) encoding a subtilisin-like protease. A total of 104 oligonucleotides, designed with the aid of dedicated computer software, were assembled in a single-step PCR. The assembly was then further amplified by PCR to produce a synthetic gene which has been cloned and successfully expressed in both Pichia pastoris and recombinant baculovirus-infected High Five(TM) cells. We believe this strategy to be of special interest as it is simple, accessible and has no limitation with respect to the size of the gene to be synthesized. Used as a systematic approach for the malarial genome or any other A + T-rich organism, the method allows the rapid synthesis of a nucleotide sequence optimized for expression in the system of choice and production of sufficiently large amounts of biological material for complete molecular and structural characterization.

Amino Acid Sequence↗

Enhanced heterologous expression of two Streptomyces griseolus cytochrome P450s and Streptomyces coelicolor ferredoxin reductase as potentially efficient hydroxylation catalysts.

The herbicide-inducible, soluble cytochrome P450s CYP105A1 and CYP105B1 and their adjacent ferredoxins, Fd1 and Fd2, of Streptomyces griseolus were expressed in Escherichia coli to high levels. Conditions for high-level expression of active enzyme able to catalyze hydroxylation have been developed. Analysis of the expression levels of the P450 proteins in several different E. coli expression hosts identified E. coli BL21 Star(DE3)pLysS as the optimal host cell to express CYP105B1 as judged by CO difference spectra. Examination of the codons used in the CYP1051A1 sequence indicated that it contains a number of codons corresponding to rare E. coli tRNA species. The level of its expression was improved in the modified forms of E. coli BL21(DE3), which contain extra copies of rare codon E. coli tRNA genes. The activity of correctly folded cytochrome P450s was further enhanced by cloning a ferredoxin reductase from Streptomyces coelicolor downstream of CYP105A1 and CYP105B1 and their adjacent ferredoxins. Expression of CYP105A1 and CYP105B1 was also achieved in Streptomyces lividans 1326 by cloning the P450 genes and their ferredoxins into the expression vector pBW160. S. lividans 1326 cells containing CYP105A1 or CYP105B1 were able efficiently to dealkylate 7-ethoxycoumarin.

Bacterial Proteins↗

Antiretroviral resistance mutations in human immunodeficiency virus type 1 infected patients enrolled in genotype testing at the Central Public Health Laboratory, São Paulo, Brazil: preliminary results.

Antiretroviral resistance mutations (ARM) are one of the major obstacles for pharmacological human immunodeficiency virus (HIV) suppression. Plasma HIV-1 RNA from 306 patients on antiretroviral therapy with virological failure was analyzed, most of them (60%) exposed to three or more regimens, and 28% of them have started therapy before 1997. The most common regimens in use at the time of genotype testing were AZT/3TC/nelfinavir, 3TC/D4T/nelfinavir and AZT/3TC/efavirenz. The majority of ARM occurred at protease (PR) gene at residue L90 (41%) and V82 (25%); at reverse transcriptase (RT) gene, mutations at residue M184 (V/I) were observed in 64%. One or more thymidine analogue mutations were detected in 73%. The number of ARM at PR gene increased from a mean of four mutations per patient who showed virological failure at the first ARV regimens to six mutations per patient exposed to six or more regimens; similar trend in RT was also observed. No differences in ARM at principal codon to the three drug classes for HIV-1 clades B or F were observed, but some polymorphisms in secondary codons showed significant differences. Strategies to improve the cost effectiveness of drug therapy and to optimize the sequencing and the rescue therapy are the major health priorities.

Adolescent↗

Cardiac troponin I sense-antisense RNA duplexes in the myocardium.

Natural antisense RNA is now thought to regulate, at least in part, a growing number of eukaryotic genes. It is becoming increasingly apparent that such endogenous antisense RNA molecules may modulate gene expression in a manner analogous to synthetic oligomers. Here, we report the detection of antisense-orientated RNA transcripts of cardiac specific troponin I in rat and human myocardium. Interestingly, the different sizes of the rat and human antisense cTNI transcripts suggest species-specific reverse transcription initiation sites. Moreover, for the first time in cardiomyocytes, we could demonstrate in vivo duplex formation between sense and antisense transcripts. The existence of antisense-sense duplexes represents compelling evidence and a potential mechanism for endogenous antisense transcript-mediated modulation of mRNA translation. The potential effect of attenuating translation was illustrated by in vitro and in vivo model systems. Testing several oligonucleotides based on the natural antisense sequences, the optimal region for inhibition of translation was identified as being close to the translational start codon.

Adult↗

Possibility of genetic coding of amino acid sequences by coherent electronic states in nucleotide chains.

The concept of coherent electronic states and coherent interactions in supramolecular structures is applied to the process of genetic information coding and its transcription from DNA to mRNA. A new genetic code is proposed based on the assumption of coherent electron states in linear chains of nucleotide bases. A new interpretation of codon equivalency (redundancy) is given. The number of existing amino acids is derived from the optimalization principle applied to the physical system storing the genetic information in the new code. The proposed code uses a variable number of positions or nucleotide bases along the DNA-mRNA structure to code a single amino acid in a protein. The average of this variable number must be equal to the base of natural logarithms (e = 2.7 . . .) in order to minimize the number of nucleotides required to code a sequence of amino acids.

Amino Acid Sequence↗

Inhibition of influenza virus replication in cultured cells by RNA-cleaving DNA enzyme.

Influenza virus replication has been effectively inhibited by antisense phosphothioate oligonucleotides targeting the AUG initiation codon of PB2 mRNA. We designed RNA-cleaving DNA enzymes from 10-23 catalytic motif to target PB2-AUG initiation codon and measured their RNA-cleaving activity in vitro. Although the RNA-cleaving activity was not optimal under physiological conditions, DNA enzymes inhibited viral replication in cultured cells more effectively than antisense phosphothioate oligonucleotides. Our data indicated that DNA enzymes could be useful for the control of viral infection.

Animals↗

Protein evolution drives the evolution of the genetic code and vice versa.

A model for the developmental pathway of the genetic code, grounded on group theory and the thermodynamics of codon-anticodon interaction is presented. At variance with previous models, it takes into account not only the optimization with respect to amino acid attributes but, also physicochemical constraints and initial conditions. A 'simple-first' rule is introduced after ranking the amino acids with respect to two current measures of chemical complexity. It is shown that a primeval code of only seven amino acids is enough to build functional proteins. It is assumed that these proteins drive the further expansion of the code. The proposed primeval code is compared with surrogate codes randomly generated and with another proposal for primeval code found in the literature. The departures from the 'universal' code, observed in many organisms and cellular compartments, fit naturally in the proposed evolutionary scheme. A strong correlation is found between, on one side, the two classes of aminoacyl-tRNA synthetases, and on the other, the amino acids grouped by end-atom-type and by codon type. An inverse of Davydov's rules, to associate the amino acid end atoms (O/N and non-O/non-N) of 18 amino acids with codons containing a weak base (A/U), extended to the 20 amino acids, is derived.

Amino Acid Sequence↗

The phylogenetic utility of the codon-degeneracy model.

The codon-degeneracy model (CDM) predicts relative frequencies of substitution for any set of homologous protein-coding DNA sequences based on patterns of nucleotide degeneracy, codon composition, and the assumption of selective neutrality. However, at present, the CDM is reliant on outside estimates of transition bias. A new method by which the power of the CDM can be used to find a synonymous transition bias that is optimal for any given phylogenetic tree topology is presented. An example is illustrated that utilizes optimized transition biases to generate CDM GF-scores for every possible phylogenetic tree for pocket gophers of the genus Orthogeomys. The resulting distribution of CDM GF-scores is compared and contrasted with the results of maximum parsimony and maximum likelihood methods. Although convergence on a single tree topology by the CDM and another method indicates greater support for that particular tree, the value of CDM GF-score as the sole optimality criterion for phylogeny reconstruction remains to be determined. It is clear, however, that the a priori estimation of an optimum transition bias from codon composition has a direct application to differentiating between alternative trees.

Animals↗

On the optimality of the genetic code, with the consideration of coevolution theory by comparison of prominent cost measure matrices.

Statistical and biochemical studies have revealed non-random patterns in codon assignments. The canonical genetic code is known to be highly efficient in minimizing the effects of mistranslation errors and point mutations, since it is known that when an amino acid is converted to another due to error, the biochemical properties of the resulted amino acid are usually very similar to those of the original one. In this study, using altered forms of the fitness functions used in the prior studies, we have optimized the parameters involved in the calculation of the error minimizing property of the genetic code so that the genetic code outscores the random codes as much as possible. This work also compares two prominent matrices, the Mutation Matrix and Point Accepted Mutations 74-100 (PAM(74-100)). It has been resulted that the hypothetical properties of the coevolution theory of the genetic code are already considered in PAM(74-100), giving more evidence on the existence of bias towards the genetic code in this matrix. Furthermore, our results indicate that PAM(74-100) is biased towards the single base mistranslation occurrences in second codon position as well as the frequency of amino acids. Thus PAM(74-100) is not a suitable substitution matrix for the studies conducted on the evolution of the genetic code.

Animals↗

A plasmid system for optimization of Fab' production in Escherichia coli: importance of balance of heavy chain and light chain synthesis.

We demonstrate the importance of optimizing the balance of light chain (LC) and heavy chain (HC) expression to achieve high level production of Fab' fragments in the Escherichia coli periplasm. The LC:HC balance has been controlled by varying the codon usage of the signal peptide (SP) and 5' mature domain coding regions. Different SP coding regions have been identified from a codon wobble-based library using alkaline phosphatase (AP) as a reporter gene. A plasmid system that enables random combination of these variant SP coding regions is used to construct optimized Fab' expression plasmids. These small plasmid libraries facilitated selection of optimal Fab' expression plasmids and resulted in increases of periplasmic yield, up to 580 mgL(-1) from E. coli fermentations and will enable rapid variable region subcloning and selection of future Fab(') expression plasmids.

Base Sequence↗

Effects of codon usage versus putative 5'-mRNA structure on the expression of Fusarium solani cutinase in the Escherichia coli cytoplasm.

Matching the codon usage of recombinant genes to that of the expression host is a common strategy for increasing the expression of heterologous proteins in bacteria. However, while developing a cytoplasmic expression system for Fusarium solani cutinase in Escherichia coli, we found that altering codons to those preferred by E. coli led to significantly lower expression compared to the wild-type fungal gene, despite the presence of several rare E. coli codons in the fungal sequence. On the other hand, expression in the E. coli periplasm using a bacterial PhoA leader sequence resulted in high levels of expression for both the E. coli optimized and wild-type constructs. Sequence swapping experiments as well as calculations of predicted mRNA secondary structure provided support for the hypothesis that differential cytoplasmic expression of the E. coli optimized versus wild-type cutinase genes is due to differences in 5(') mRNA secondary structures. In particular, our results indicate that increased stability of 5(') mRNA secondary structures in the E. coli optimized transcript prevents efficient translation initiation in the absence of the phoA leader sequence. These results underscore the idea that potential 5(') mRNA secondary structures should be considered along with codon usage when designing a synthetic gene for high level expression in E. coli.

Amino Acid Sequence↗

Investigation on the causes of codon and amino acid usages variation between thermophilic Aquifex aeolicus and mesophilic Bacillus subtilis.

Base composition, codon usages and amino acid usages have been analyzed by taking 529 orthologous sequences of Aquifex aeolicus and Bacillus subtilis, having different optimal growth temperatures. These two bacteria do not have significant difference in overall GC composition, but GC(1+2) and GC3 levels were found to vary significantly. Significant increments in purine content and GC3 composition have been observed in the coding sequences of Aquifex aeolicus than its Bacillus subtilis counterparts. Correspondence analyses on codon and amino acid usages reveal that variation in base composition actually influences their codon and amino acid usages. Two selection pressures acting on the nucleotide level (GC3 and purine enrichment), causes variation in the amino acid usage differently in different protein secondary structures. Our results suggest that adaptation of amino acid usages in coil structure of Aquifex aeolicus proteins is under the control of both purine increment and GC3 composition, whereas the adaptation of the amino acids in the helical region of thermophilic bacteria is strongly influenced by the purine content. Evolutionary perspectives concerning the temperature adaptation of DNA and protein molecules of these two bacteria have been discussed on the basis of these results.

Amino Acids↗

Engineered Lactiplantibacillus plantarum and Levilactobacillus brevis utilizing ribonucleoprotein-mediated editing for inactivation of hemolysin gene.

Lactiplantibacillus plantarum and Levilactobacillus brevis are widely used probiotics with significant potential as chassis organisms for probiotic engineering. However, their bioengineering remains underdeveloped compared to that of other probiotic bacteria due to the limited availability of genetic tools. Although CRISPR-Cas systems have shown promise for genome editing in Lactobacillus species, strain- or site-specific targeting challenges must be overcome to enhance their broader applicability. This study aimed to develop a novel editing system with reduced dependency on plasmids and antibiotics in L. plantarum WCFS1, L. plantarum SPC 72 - 1 and L. brevis SPC-SNU 70 - 2 using a Cas9-gRNA ribonucleoprotein (RNP) complex. Although the hlyIII gene has been annotated as a hemolysin-related gene in several Lactobacillus genomes, no functional hemolytic activity has been definitively demonstrated to date. In this study, hlyIII was selected as a target to evaluate genome editing efficiency and to assess its potential relevance to strain safety. To construct ΔhlyIII strains, the RNP complex targeting hlyIII was separately transformed with recombinase RecE/T and double-stranded donor DNA. As a result, ΔhlyIII mutants were obtained under optimized electroporation conditions. Sequencing analysis revealed a 50 bp deletion and the introduction of a stop codon in hlyIII across all mutant strains. The hemolytic activity test showed a reduction in free hemoglobin levels in the ΔhlyIII strains compared to the wild type: 27.0%, 74.3%, and 5.0% in L. plantarum WCFS1, L. plantarum SPC 72 - 1, and L. brevis SPC-SNU 70 - 2, respectively. These results suggest strain-dependent differences in hemolytic activity and indicate that inactivation of hlyIII may contribute to reduced hemolysis, although further validation is needed to clarify its functional role. In conclusion, the hlyIII gene was successfully edited in L. plantarum and L. brevis using Cas9-gRNA ribonucleoprotein-mediated editing, demonstrating the feasibility of this genome editing platform for application in probiotic strains.

Gene Editing↗

Conformational preferences of the base substituent in hypermodified nucleotide queuosine 5'-monophosphate 'pQ' and protonated variant 'pQH+'.

Conformational preferences of the base substituent in hypermodified nucleotide queuosine 5'-monophosphate 'pQ' and its protonated form 'pQH+' have been studied using quantum chemical Perturbative Configuration Interaction with Localized Orbitals PCILO method. The salient points have also been examined using molecular mechanics force field MMFF, parameterized modified neglect of differential overlap PM3 and Hartree Fock-Density Functional Theory HF DFT (pBP/DN*) approaches. Aqueous solvation of pQ and pQH+ has also been studied using molecular dynamics simulations. Consistent with the observed crystal structure, in isolated protonated form pQH+, the quaternary amine HN(13)(+)H, of the sidechain having 7-aminomethyl linkage, hydrogen bonds with the carbonyl oxygen O(10) of the base. However, N(13)H-O(10) hydrogen bonding is not preferred for unprotonated pQ, whether isolated or hydrated. Interaction between the 5'-phosphate and the 7-aminomethyl group is more likely for isolated pQ. The cyclopentenediol hydroxyl group O4"H may hydrogen bond with the O(10) in isolated pQ as well as in pQH+. The O4"H may hydrogen bond with the 5'-phosphate as well. The presence of -CH2-NH- and O"H groups in pQ and pQH+ allows interesting possibilities for intranucleotide hydrogen bonds and interactions across the anticodon loop. Simultaneous hydrogen bonds O2P-HN(13)+H-O(10) are indicated for hydrated pQH+. Unlike weak involvement of O4"H, these interactions also persist in hydrated pQH+ and may much reduce backbone flexibility. Resulting sub-optimal Q:C base pairing leads to unbiased reading of U or C as the third codon letter. Cyclopentenediol hydroxyl groups may interact with other biomolecules, allowing specific recognition. Prospective pQ(34) and pQ(34)H+ sites for codon-anticodon base pairing remain unhindered, but non canonical Q:G base pairing (amber-suppression) is ruled out.

Anticodon↗