Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

A quality control pathway that down-regulates aberrant T-cell receptor (TCR) transcripts by a mechanism requiring UPF2 and translation.

Nonsense-mediated decay (NMD) is an RNA surveillance pathway that degrades mRNAs containing premature termination codons (PTC). T-cell receptor (TCR) and immunoglobulin (Ig) transcripts, which are encoded by genes that very frequently acquire PTCs during lymphoid ontogeny, are down-regulated much more dramatically in response to PTCs than are other known transcripts. Another feature unique to TCR, Ig, and a subset of other mRNAs is that they are down-regulated in response to nonsense codons in the nuclear fraction of cells. This is paradoxical, as the only well recognized entity that recognizes nonsense codons is the cytoplasmic translation apparatus. Therefore, we investigated whether translation is responsible for this nuclear-associated mechanism. We found that the down-regulation of TCR-beta transcripts in response to nonsense codons requires several features of translation, including an initiator ATG and the ability to scan. We also found that optimal down-regulation depends on a Kozak consensus sequence surrounding the initiator ATG and that it can be initiated by an internal ribosome entry site, neither of which has been demonstrated before for any other PTC-bearing mRNA. At least a portion of this down-regulatory response is mediated by the NMD pathway as antisense hUPF2 transcripts increased the levels of PTC-bearing TCR-beta transcripts in the nuclear fraction of cells. We conclude that a hUPF2-dependent RNA surveillance pathway with translation-like features operating in the nuclear fraction of cells prevents the expression of potentially deleterious truncated proteins encoded by non-productively rearranged TCR genes.

Adaptor Proteins, Signal Transducing↗

Optimality of the genetic code with respect to protein stability and amino-acid frequencies.

BACKGROUND: The genetic code is known to be efficient in limiting the effect of mistranslation errors. A misread codon often codes for the same amino acid or one with similar biochemical properties, so the structure and function of the coded protein remain relatively unaltered. Previous studies have attempted to address this question quantitatively, by estimating the fraction of randomly generated codes that do better than the genetic code in respect of overall robustness. We extended these results by investigating the role of amino-acid frequencies in the optimality of the genetic code. RESULTS: We found that taking the amino-acid frequency into account decreases the fraction of random codes that beat the natural code. This effect is particularly pronounced when more refined measures of the amino-acid substitution cost are used than hydrophobicity. To show this, we devised a new cost function by evaluating in silico the change in folding free energy caused by all possible point mutations in a set of protein structures. With this function, which measures protein stability while being unrelated to the code's structure, we estimated that around two random codes in a billion (109) are fitter than the natural code. When alternative codes are restricted to those that interchange biosynthetically related amino acids, the genetic code appears even more optimal. CONCLUSIONS: These results lead us to discuss the role of amino-acid frequencies and other parameters in the genetic code's evolution, in an attempt to propose a tentative picture of primitive life.

Amino Acid Substitution↗

Nucleotide sequence and structural analysis of the rat RT1.Eu and RT1.Aw3l genes, and of genes related to RT1.O and RT1.C.

A cDNA library was constructed using mRNA isolated from the R21 strain of rats which have the major histocompatibility complex (MHC) haplotype RT1.AlBlDlEu and the growth and reproduction complex (grc) genotype grc+. The cDNA clones that hybridized with the class I probes pAG64c and pARI.5 and were 1.3-1.7 kilobases were selected. Full-length clones were identified by sequencing partially the 5' and 3' ends of each clone, by the presence of a start codon at the 5' end, and by a polyadenylation sequence at the 3' end. The full-length cDNA clones were examined for in vitro transcription by transfection into human CIR cells using electroporation, and expression was detected by flow cytometry using monoclonal antibodies specific to the heavy chains and polyclonal antibody to beta 2-microglobulin. The RT1.Eu gene was transcribed and expressed optimally, and its nucleotide and deduced amino acid sequences differed significantly from the RT1.Aa, RT1.A(l), RT.Au, LW2, and 11/3R genes but only slightly from the RT1.K gene. The high level of sequence similarity between RT1.Eu and RT1.K suggests that the two genes may have originated from a common ancestral gene. In addition, three new genes (RT1.Aw3l, RT1.C-type, and RT1.O-type) were identified. The RT1.Aw3l gene is almost identical to RT1.A(l) with the exception of an in frame deletion of 21 nucleotides in exon 2 leading to a 7 amino acid deletion in the alpha 1 domain of the deduced amino acid sequence and 11 nucleotide substitutions and insertions in the rest of the sequence. It transcribed optimally, but no significant expression was detected. The RT1.C-type gene 119 is very similar (97%) to the LW2 gene in the 3' untranslated region, which suggests that it is in the RT1.C region. It transcribed optimally, but no significant expression was detected. The RT1.O-type gene 149 has all the features of a class Ib gene, but a premature stop codon in the alpha 1 domain causes incomplete translation. Its in vitro transcription was very low, and no expression was detected. These studies, combined with previous work, indicate that in the MHC of the R21 strain three class Ia genes (Eu, A(l), Aw3l) and three class Ib genes (C-type, O-type, N) are transcribed but only two class Ia genes (Eu, A(l)) are expressed.

Amino Acid Sequence↗

Codon usage decreases the error minimization within the genetic code.

The genetic code is not random but instead is organized in such a way that single nucleotide substitutions are more likely to result in changes between similar amino acids. This fidelity, or error minimization, has been proposed to be an adaptation within the genetic code. Many models have been proposed to measure this adaptation within the genetic code. However, we find that none of these consider codon usage differences between species. Furthermore, use of different indices of amino acid physicochemical characteristics leads to different estimations of this adaptation within the code. In this study, we try to establish a more accurate model to address this problem. In our model, a weighting scheme is established for mistranslation biases of the three different codon positions, transition/transversion biases, and codon usage. Different indices of amino acids' physicochemical characteristics are also considered. In contrast to pervious work, our results show that the natural genetic code is not fully optimized for error minimization. The genetic code, therefore, is not the most optimized one for error minimization, but one that balances between flexibility and fidelity for different species.

Amino Acid Substitution↗

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment↗

Kinetic properties of mutant deoxyguanosine kinase in a case of reversible hepatic mtDNA depletion.

DGUOK [dG (deoxyguanosine) kinase] is one of the two mitochondrial deoxynucleoside salvage pathway enzymes involved in precursor synthesis for mtDNA (mitochondrial DNA) replication. DGUOK is responsible for the initial rate-limiting phosphorylation of the purine deoxynucleosides, using a nucleoside triphosphate as phosphate donor. Mutations in the DGUOK gene are associated with the hepato-specific and hepatocerebral forms of MDS (mtDNA depletion syndrome). We identified two missense mutations (N46S and L266R) in the DGUOK gene of a previously reported child, now 10 years old, who presented with an unusual revertant phenotype of liver MDS. The kinetic properties of normal and mutant DGUOK were studied in mitochondrial preparations from cultured skin fibroblasts, using an optimized methodology. The N46S/L266R DGUOK showed 14 and 10% residual activity as compared with controls with dG and deoxyadenosine as phosphate acceptors respectively. Similar apparent negative co-operativity in the binding of the phosphate acceptors to the wild-type enzyme was found for the mutant. In contrast, abnormal bimodal kinetics were shown with ATP as the phosphate donor, suggesting an impairment of the ATP binding mode at the phosphate donor site. No kinetic behaviours were found for two other patients with splicing defects or premature stop codon. The present study represents the first characterization of the enzymatic kinetic properties of normal and mutant DGUOK in organello and our optimized protocol allowed us to demonstrate a residual activity in skin fibroblast mitochondria from a patient with a revertant phenotype of MDS. The residual DGUOK activity may play a crucial role in the phenotype reversal.

Cells, Cultured↗

Analysis of a shift in codon usage in Drosophila.

In order to gain further insight into a shift in codon usage first observed in Drosophila willistoni we have analyzed seven genes in six species in the lineage leading to D. willistoni. This lineage contains the willistoni and saltans species groups. Sequences were obtained from GenBank or newly sequenced for this study. All species studied showed significant difference in codon usage compared to D. melanogaster for about one third of all amino acids. Within the willistoni/saltans lineage, codon usage is homogeneous, indicating that the shift in codon usage occurred prior to the diversification of extant species in this lineage which we estimate to date to about 20 million years ago. Thus the shift is old and has been stable. We also examined introns from these genes and the G/C composition at four-fold degenerate sites in an effort to detect a change in mutation bias. There is little or no evidence for a difference in mutation bias compared to D. melanogaster. We also considered whether relaxed selection (possibly due to reduced population sizes) or reduced recombination (due to numerous naturally occurring inversions) could account for the shift and concluded these factors alone are insufficient to explain the patterns observed. A change in the relative abundance of isoaccepting tRNAs is one of the few explanations that can account for the observations. Particularly intriguing is the fact that the greatest changes in codon usage have occurred for amino acids with two-fold C/T ending codons for which it is known that posttranscriptional modification occurs in tRNAs from a G in the wobble position to Queuosine that changes optimal binding from C to a slight preference for U. However, we do not argue that this shift was adaptive in nature, rather it may be an example of a "frozen accident."

Amino Acids↗

ANGLE: a sequencing errors resistant program for predicting protein coding regions in unfinished cDNA.

In the process of making full-length cDNA, predicting protein coding regions helps both in the preliminary analysis of genes and in any succeeding process. However, unfinished cDNA contains artifacts including many sequencing errors, which hinder the correct evaluation of coding sequences. Especially, predictions of short sequences are difficult because they provide little information for evaluating coding potential. In this paper, we describe ANGLE, a new program for predicting coding sequences in low quality cDNA. To achieve error-tolerant prediction, ANGLE uses a machine-learning approach, which makes better expression of coding sequence maximizing the use of limited information from input sequences. Our method utilizes not only codon usage, but also protein structure information which is difficult to be used for stochastic model-based algorithms, and optimizes limited information from a short segment when deciding coding potential, with the result that predictive accuracy does not depend on the length of an input sequence. The performance of ANGLE is compared with ESTSCAN on four dataset each of them having a different error rate (one frame-shift error or one substitution error per 200-500 nucleotides) and on one dataset which has no error. ANGLE outperforms ESTSCAN by 9.26% in average Matthews's correlation coefficient on short sequence dataset (< 1000 bases). On long sequence dataset, ANGLE achieves comparable performance.

Algorithms↗

Dual expression system suitable for high-throughput fluorescence-based screening and production of soluble proteins.

Many studies that aim to characterize the proteome structurally or functionally require the production of pure protein in a high-throughput format. We have developed a fast and flexible integrated system for cloning, protein expression in Escherichia coli, solubility screening and purification that can be completely automated in a 96-well microplate format. We used recombination cloning in custom-designed vectors including (i) a (His)(6) tag-encoding sequence, (ii) a variable solubilizing partner gene, (iii) the DNA sequence corresponding to the TEV protease cleavage site, (iv) the gene (or DNA fragment) of interest, (v) a suppressible amber stop codon, and (vi) an S.tag peptide-encoding sequence. First, conditions of bacterial culture in microplates (250 microL) were optimized to obtain expression and solubility patterns identical to those obtained in a 1-L flask (100-mL culture). Such conditions enabled the screening of various parameters in addition to the fusion partners (E. coli strains, temperature, inducer...). Second, expression of fusion proteins in amber suppressor strains allowed quantification of soluble and insoluble proteins by fluorescence through the detection of the S.tag. This technique is faster and more sensitive than other commonly used methods (dot blots, Western blots, SDS-PAGE). The presence of the amber suppressor tRNA was shown to affect neither the expression pattern nor the solubility of the target proteins. Third, production of the most interesting soluble fusion proteins, as detected by our screening method, could be performed in nonsuppressor strains. After cleavage with the TEV protease, the target proteins were obtained in a native form with a unique additional N-terminal glycine.

Blotting, Western↗

Restructuring the translation initiation region of the human parathyroid hormone gene for improved expression in Escherichia coli.

Overexpression of native human parathyroid hormone in Escherichia coli was achieved by a modification of the 5' end of the genomic gene sequence, thereby adapting this part of the translation initiation region to the bacterial host. Some simple rules abstracted from optimization studies of translation initiation of a beta-interferon gene were applied. These included (a) extending complementarity of the mRNA to the anticodon loop of tRNAfMet by use of a codon with a purine nucleotide directly following the ATG, (b) avoidance of stable secondary structure in the mRNA by use of synonymous A/U-rich codons, (c) elimination of a potential second Shine-Dalgarno sequence. The appropriate silent changes led to a 20-fold increase in parathyroid hormone production resulting in 4.3% of total soluble protein. This result proves the validity of our simple approach for optimization of foreign gene expression in E. coli.

Base Sequence↗

Inducible expression vectors incorporating the Escherichia coli atpE translational initiation region.

New expression vectors were constructed for use in strains of Escherichia coli. Their most important feature is a polylinker system that facilitates the insertion of a gene in an optimal relationship to the highly efficient E. coli atpE translational initiation region (from nucleotide -50 to the start codon). Three ATG-containing restriction endonuclease sites can be used for the insertion of the 5' end of a gene at, or near to, its translational initiation codon. These sites may alternatively be used for the creation of a suitable translational start codon. Transcription is started by the bacteriophage lambda major promoters pR and pL in tandem and terminated by the bacteriophage fd terminator. Transcriptional initiation is very effectively repressed at 28-30 degrees C by the product of the bacteriophage lambda cIts857 gene, which is also present on the vectors. Full induction is achieved by shifting the incubation temperature to 42 degrees C. The combination of highly efficient transcriptional and translational signals on these vectors allowed high-level expression of sequences encoding human interferon beta and interleukin 2 and of the E. coli atpA, sucC and sucD genes.

DNA Restriction Enzymes↗

On the information content of the genetic code.

In living organisms 20 amino acids along with the terminator value(s) are encoded by 64 codons giving a degeneracy of the codons as described by the genetic code. A basic theoretical problem of genetic codes is to explain the particular distribution of degeneracies of partitions involved in the codes. In this work the degeneracy problem is considered in the framework of information theory. It is shown by direct numerical evaluation of a certain degeneracy information function associated with the genetic code that the degeneracy of the codes is observed to be related to the optimization of this function.

Amino Acids↗

Clinical value of K-ras codon 12 analysis and endobiliary brush cytology for the diagnosis of malignant extrahepatic bile duct stenosis.

Extrahepatic biliary stenosis can be caused by benign and malignant disorders. In most cases, a tissue diagnosis is needed for optimal management of patients, but the sensitivity of biliary cytology for the diagnosis of a malignancy is relatively low. The additional diagnostic value of K-ras mutational analysis of endobiliary brush cytology was assessed. Endobiliary brush cytology specimens obtained during endoscopic retrograde cholangiopancreaticography were prospectively collected from 312 consecutive patients with extrahepatic biliary stenosis. The results of conventional light microscopic cytology and K-ras codon 12 mutational analysis were compared and evaluated in view of the final diagnosis made by histological examination of the stenotic lesion and/or patient follow-up. The sensitivities of cytology and mutational analysis to detect malignancy were 36 and 42%, respectively. When both tests were combined, the sensitivity increased to 62%. The specificity of cytology was 98%, and the specificity of the mutational analysis and of both tests combined was 89%. Positive predictive values for cytology, mutational analysis, and both tests combined were 98, 92, and 94%, whereas the corresponding negative predictive values were 34, 34, and 44%, respectively. The sensitivity of K-ras mutational analysis was 63% for pancreatic carcinomas compared to 27% for bile duct, gallbladder, and ampullary carcinomas. K-ras mutational analysis can be considered supplementary to conventional light microscopy of endobiliary brush cytology to diagnose patients with malignant extrahepatic biliary stenosis, particularly in the case of pancreatic cancer. The presence of a K-ras codon 12 mutation in endobiliary brush cytology per se supports a clinical suspicion of malignancy, even when the conventional cytology is negative or equivocal.

Bile Duct Neoplasms↗

Synthesis and sequence optimization of GFP mutants containing aromatic non-natural amino acids at the Tyr66 position.

In order to alter the fluorescence properties of green fluorescent protein (GFP), aromatic non-natural amino acids were introduced into the Tyr66 position of GFP in a cell-free translation system using a four-base codon method. Two non-natural mutants (O-methyltyrosine and p-aminophenylalanine mutants) out of 18 mutants showed blue-shifted but weak fluorescence compared with wild-type GFP. Then the aminophenylalanine mutant was sequence optimized by introducing random mutations around the Tyr66 site. For this purpose, a method for random mutation of non-natural proteins in a cell-free system was developed. Three aminophenylalanine mutants with Y145F, Y145L and Y145 M mutations were obtained, which exhibited increased fluorescence by 1.5-, 3- and 4-fold, respectively. These results indicate that random mutation around non-natural amino acids is useful strategy in order to improve protein functions that are reduced by non-natural amino acid incorporation. The method described here will be applicable to other non-natural mutant proteins in a high-throughput manner.

Cell-Free System↗

Cooperative effects by the initiation codon and its flanking regions on translation initiation.

The purine-rich Shine-Dalgarno (SD) sequence located a few bases upstream of the mRNA initiation codon supports translation initiation by complementary binding to the anti-SD in the 16S rRNA, close to its 3' end. AUG is the canonical initiation codon but the weaker UUG and GUG codons are also used for a minority of genes. The codon sequence of the downstream region (DR), including the +2 codon immediately following the initiation codon, is also important for initiation efficiency. We have studied the interplay between these three initiation determinants on gene expression in growing Escherichia coli. One optimal SD sequence (SD(+)) and one lacking any apparent complementarity to the anti-SD in 16S rRNA (SD(-)) were analyzed. The SD(+) and DR sequences affected initiation in a synergistic manner and large differences in the effects were found. The gene expression level associated with the most efficient of these DRs together with SD(-) was comparable to that of other DRs together with SD(+). The otherwise weak initiation codon UUG, but not GUG, was comparable with AUG in strength, if placed in the context of two of the DRs. The +2 codon was one, but not the only, determinant for this unexpectedly high efficiency of UUG.

Base Sequence↗

Folding of the MS2 coat protein in Escherichia coli is modulated by translational pauses resulting from mRNA secondary structure and codon usage: a hypothesis.

Possible translational pauses within the coat protein of the RNA bacteriophage MS2 were located on the basis of a distribution plot of rare codons and RNA secondary structure. It appeared that the position of certain codon pauses corresponds with the size of some nascent polypeptide intermediates, which have been isolated from MS2-infected cells. Other accumulated polypeptide intermediates seemed to be related to RNA regions, where double-stranded secondary structures occur, which probably impede the movement of ribosomes during chain elongation. We assume that a discontinuous translation rate is designed to allow optimal folding of this (and other) polypeptide(s).

Capsid↗

Rapid evolution of translational control mechanisms in RNA genomes.

We have introduced 13 base substitutions into the coat protein gene of RNA bacteriophage MS2. The mutations, which are clustered ahead of the overlapping lysis cistron, do not change the amino acid sequence of the coat protein, but they disrupt a local hairpin, which is needed to control translation of the lysis gene. The mutations decreased the phage titer by four orders of magnitude but, upon passaging, the virus accumulated suppressor mutations that raised the fitness to almost wild-type level. Analysis of the pseudorevertants showed that the disruption of the local hairpin, controlling expression of the lysis gene, had apparently been so complete that its restoration by chance mutations could not be achieved. Instead, alternative foldings initiated by the starting mutations were further stabilized and optimized. Strikingly, in the pseudorevertants analyzed, translational control of the lysis gene had been restored. This feat was accomplished by, on average, four suppressor mutations that generally occurred at codon wobble positions. We also introduced 11 mutations in a hairpin more upstream in the coat protein gene and not implicated in lysis control. Here the titer dropped by three logs, but pseudorevertants with a fitness close to wild-type were soon generated. These pseudorevertants again were the result of the optimization of alternative foldings induced by the mutations. The transition of the secondary structure from wild-type to pseudorevertant could be visualized by structure probing. Our study shows that the folding of the RNA is an important phenotypic property of RNA viruses. However, its distortion can easily be overcome by optimizing alternative base-pairings. These new structures are not qualitatively equivalent to the original one, since they do not successfully compete with the wild-type.

Base Sequence↗

Rapid detection of ethambutol-resistant Mycobacterium tuberculosis strains by PCR-RFLP targeting embB codons 306 and 497 and iniA codon 501 mutations.

Mutations at embB gene codons 306 and 497 and iniA gene codon 501 occur frequently in ethambutol (EMB)-resistant Mycobacterium tuberculosis strains worldwide. The identification of these mutations in resistant strains has been achieved by labor-intensive DNA sequencing or by tedious amplification protocols followed by restriction endonuclease digestion. In this report, we describe PCR-restriction fragment length polymorphism (RFLP)-based methods for determining substitutions at embB codons 306 and 497 and iniA codon 501 directly in BACTEC cultures of M. tuberculosis isolates. The wild-type and mutant alleles are revealed by easily interpretable and different RFLP patterns. The methods optimized initially on reference strains were tested directly on BACTEC cultures of 25 randomly selected clinical M. tuberculosis isolates, seven of which were determined to contain EMB-resistant strains by phenotypic drug susceptibility testing. The PCR-RFLP methods identified mutations in four of seven EMB-resistant strains with three isolates containing mutated embB codon 306 and one isolate containing mutated embB codon 497. The results of PCR-RFLP were confirmed by DNA sequencing. The worldwide prevalence figures for mutations at embB codons 306 and 497 and iniA codon 501 suggest that nearly half of EMB-resistant M. tuberculosis strains could be identified within one working day even in developing countries equipped with simple PCR technology instead of weeks required for phenotypic drug susceptibility testing. Further, since EMB resistance is also associated with multiple-drug resistance from some geographical locations, detection of EMB resistance may also lead to rapid identification of multidrug-resistant strains of M. tuberculosis.

Codon↗