Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

High expression of synthetic human interferon-gamma cDNA in E. coli.

Human interferon-gamma (IFN-gamma) cDNA was synthesized, and it makes the usage of favorable codons in E. coli. The authors got 9 different expression plasmids which contain the synthetic IFN-gamma-cDNA and have different spaces between SD sequence and ATG. The free energies G0f298 in the formation of stable secondary structure in the translation initiation region (TIR) are different in various expression plasmids. One of them, pLY4-gamma 5, can highly yield INF-gamma which will be about 60%-80% of the total bacterial proteins, such a high expression was hardly noted in literature. The reasons of high expression in this work are optimal spaces between SD and ATC, favorable delta G0f298, favourable condons usage for E. coli.

Amino Acid Sequence

Green fluorescent protein expressed by recombinant pseudorabies virus as an in vivo marker for viral replication.

We isolated and characterized a pseudorabies virus (PrV) mutant expressing an engineered green fluorescent protein (GFP) optimized for expression in human cells. The GFP DNA was inserted in the non-essential glycoprotein G (gG) gene of the attenuated PrV strain Bartha. The coding sequence was cloned in frame behind the first seven codons of the gG gene under control of the strong gG promotor. On excitation with blue light, live cells infected with the recombinant PrV B80eGFP exhibited bright fluorescence when examined microscopically using filters for FITC fluorescence. In fixed samples detection sensitivity was increased by immunofluorescence using an anti-GFP antibody. Specifically labelled PrV mutants have been used successfully as transsynaptic circuit tracers for definition of central command neurons in the brain (Jansen et al., 1995. Central command neurons of the sympathetic nervous system: basis of the fight-or-flight response. Science 270, 644-646). Availability of this recombinant allows the study of even more complex interactions using differentially labelled PrV mutants, and provides a means to monitor viral replication and spread without destruction of the cell.

Animals

Optimized genomic editing of a common Duchenne muscular dystrophy mutation in patient-derived muscle cells and a new humanized mouse model.

Duchenne muscular dystrophy (DMD) is a fatal X-linked, recessive disease caused by mutations in the DMD gene encoding dystrophin, a membrane-associated protein necessary for maintaining muscle structure and function. One of the common DMD mutations is the deletion of exon 52 (Δ52), which introduces a premature stop codon in exon 53, preventing the expression of functional dystrophin protein. Patients with this mutation could benefit from skipping or reframing exon 53 to restore the dystrophin open reading frame. In this study, we investigated the efficacy of single-cut CRISPR gene editing with Staphylococcus pyogenes Cas9 (SpCas9)-LRVQR to restore dystrophin expression in patient-derived induced pluripotent stem cells (iPSCs) and a newly generated humanized DMD mouse model. We compared two injection routes for adeno-associated virus (AAV) serotype 9 to deliver gene-editing components to neonatal mice: intraperitoneal (IP) and facial vein (FV) injection. We observed efficient restoration of dystrophin protein expression across multiple skeletal muscle groups and the heart. The AAV9-mediated CRISPR single-cut approach ameliorated key DMD hallmarks, including histopathological phenotypes, impaired grip strength, and elevated serum creatine kinase levels. Our optimized strategies for dystrophin restoration in humanized DMD mice with exon 52 deletion represent a promising treatment for DMD.

AAV

Optimizing the expression in E. coli of a synthetic gene encoding somatomedin-C (IGF-I).

Double-stranded DNA encoding the human hormone somatomedin-C (SMC) has been synthesized. This synthetic gene has been inserted into a plasmid bearing the strong leftward promoter (PL) of bacteriophage lambda and expressed in E. coli. Codons for the N-terminal region of SMC which maximized the hormone's synthesis were chosen in an SMC-lac z fusion assay. The amounts of SMC accumulated in E. coli were influenced by mutations at two chromosomal loci, lon and htpR.

ATP-Dependent Proteases

Structural requirements for initiation of translation by internal ribosome entry within genome-length hepatitis C virus RNA.

Cap-independent translation of hepatitis C virus (HCV) RNA is mediated by an internal ribosomal entry segment (IRES) located within the 5' nontranslated RNA (5'NTR), but previous studies provide conflicting views of the viral sequences which are required for translation initiation. These discrepancies could have resulted from the inclusion of less than full-length 5'NTR in constructs studied for translation or destabilization of RNA secondary structure due to fusion of the 5'NTR to heterologous reporter sequences. In an effort to resolve this confusion, we constructed a series of mutations within the 5'NTR of a nearly full-length 9.5-kb HCV cDNA clone and examined the impact of these mutations on HCV translation in vitro in rabbit reticulocyte lysates and in transfected Huh-T7 cells. The inclusion of the entire open reading frame in HCV transcripts did not lead to an increase in IRES-directed translation of the capsid and E1 proteins, suggesting that the nonstructural proteins of HCV do not include a translational transactivator. However, in reticulocyte lysates programmed with full-length transcripts, there were multiple aberrent translation initiation sites resembling those identified in some picornaviruses. The deletion of nucleotides (nt) 28-69 of the 5'NTR (stem-loop IIa) sharply reduced capsid translation both in vitro and in vivo. A small deletion mutation involving nt 328-334, immediately upstream of the initiator AUG at nt 342, also resulted in a nearly complete inhibition of translation, as did the deletion of multiple intervening structural elements. An in-frame 12-nt insertion placed within the capsid-coding region 9 nt downstream of the initiator AUG strongly inhibited translation both in vitro and in vivo, while multiple silent mutations within the first 42 nt of the open reading frame also reduced translation in reticulocyte lysates. Thus, domains II and III of the 5'NTR are both essential to activity of the IRES, while conservation of sequence downstream of the initiator AUG is required for optimal IRES-directed translation.

Amino Acid Sequence

A segment-based dynamic programming algorithm for predicting gene structure.

An algorithm called segment-based dynamic programming is described for predicting gene structure from a sequence of genomic DNA. The algorithm explores the space of gene structures that satisfy junctional and frame constraints and finds the gene structure that optimizes the sum of junctional and segmental scoring functions. Junctional constraints specify acceptable sites of initiation, termination, and splicing, whereas frame constraints ensure that the total exon length is a multiple of three and that no in-frame stop codons occur within exons or at exon-exon junctions. By computing over segments, segment-based dynamic programming maintains reading frame and phase information for each segment, it can assemble exons in-frame as well as score them in-frame. The algorithm is used to quantify the computational power of constraints. Experimental results show that frame constraints reduce the size of the search space by several orders of magnitude and that cardinality constraints place an asymptotic limit on the size of the search space. The algorithm is also used to compare the accuracy of different methods for assembly and scoring. A scoring scheme based on fifth-order Markov hexamer frequencies is presented and used in three objective functions, corresponding to in-frame, frame-independent, and frame-maximal scoring strategies. Experimental results show that in-frame assembly improves specificity only slightly over frame-independent assembly, whereas in-frame scoring improves specificity substantially over frame-independent and frame-maximal scoring.

Algorithms

Nitrile hydratase gene from Rhodococcus sp. N-774 requirement for its downstream region for efficient expression.

For improvement of the production of nitrile hydratase (NHase) from Rhodococcus sp. N-774 by recombinant DNA techniques, several plasmids, each of which had a deletion of the upstream or downstream region of the genes encoding the alpha and beta subunits of NHase, were constructed. Enzyme assays of recombinant R. rhodochrous and Escherichia coli cells showed that a downstream region of the NHase genes was indispensable for the production of active NHase in both cells, but for the production of the active amidase, no genes other than the amidase structural gene were required. The nucleotide sequence of the downstream region contained a single open reading frame (Orf1188) with 396 amino acids. Orf1188 showed similarity in amino acid sequence to P47K, an open reading frame found downstream of the NHase genes from Pseudomonas chlororaphis B23, and also to the cobW gene product, which may be involved in cobalamin biosynthesis in Pseudomonas denitrificans. Because the distance between the TGA stop codon for the NHase beta-subunit and the ATG codon for Orf1188 is only 98 bp, and because production of both Orf1188 and NHase is dependent on a promoter upstream of the amidase gene, these genes appear to be co-transcribed in a polycistronic manner, forming an operon. By optimization of the culture conditions of R. rhodochrous carrying pKRNH2, which contained the amidase, NHase, and Orf1188 genes, the transformant showed the NHase activity 6-fold higher than that of the original strain, Rhodococcus sp. N-774.

Amino Acid Sequence

[Computer programs for the analysis of nucleotide sequences (MALK)].

A system for the computer analysis of nucleic acid and protein sequences ("Helix") is described. Format of the DNA sequences is EMBL--compatible and may be easily commented with the help of convenient menus. "Helix" has also following possibilities: an effective alignment of gele reading data and formation of the final sequence; simple making of recombined molecules "in calcular"; calculations of nucleotide and dinucleotide distribution along the sequence; looking for coding frames; calculations percentage of codons and amino acids in coding frames; searching for direct and inverted repeats; sequences alignment; protein secondary structure prediction; restriction mapping; DNA--protein translation. "Helix" also contain programs for RNA-structure prediction, looking for homologies throughover the EMAL bank, choosing optimal sequence for probes and searching promoters. All the programs are written at FORTRAN-77 and automatically translated into FORTRAN-4. "Helix" require only 64 kbite.

Base Sequence

Recognition of genes in human DNA sequences.

A new approach to computer-assisted gene recognition in higher eukaryote DNA is suggested. It allows one to use not only linear functions for scoring structures, but all functions satisfying natural monotonicity conditions. The algorithm constructs the set of structures guaranteed to contain an optimal structure for every function. So, it uncouples the time-consuming step of generation of this set from the fast step of structure scoring, thus making it simple to experiment with different functions. One particular scoring function, taking into account only codon usage and positional nucleotide frequencies of the splicing sites, has been implemented in the Genome Recognition and Exon Assembly Tool program, and has been tested on an independent sample of human genes, yielding 88% sensitivity and 79% specificity.

Algorithms

Transcriptional regulation of the carbon monoxide dehydrogenase gene (cdhA) in Methanosarcina thermophila.

The mechanisms of gene regulation in the phylogenetic domain Archaea are not yet understood. To examine the expression of a gene encoding a highly regulated catabolic enzyme from the methanogenic archaea, a Methanosarcina thermophila lambda gt11 chromosomal library was probed with antiserum prepared against the 89-kDa subunit of carbon monoxide dehydrogenase, an enzyme which is required for growth and methanogenesis from acetate. A 2.3-kilobase DNA fragment was isolated that encoded 300 bases of the 5'-end of cdhA, the gene which encodes the 89-kDa subunit, and 2 kilobases upstream of cdhA that included an upstream open reading frame (ORF1). Primer extension analyses determined that cdhA and ORF1 each had a single transcriptional initiation site located 370 and 9 nucleotides, respectively, 5' of the putative translation initiation codons for cdhA and ORF1. Each promoter element had sequence similarity to a consensus archaeal promoter sequence. Three discrete mRNA cdhA transcripts of 9.5, 5.6, and 4.8 kilobases and one mRNA ORF1 transcript of < 2 kilobases were identified. All four transcripts were optimally expressed in cells grown with acetate, while growth with the more energetically favorable substrates methanol and trimethylamine caused a significant reduction in levels of the cdhA and ORF1 mRNA's. The half-lives of the 5' ends of the three cdhA transcripts and entire ORF1 mRNA transcript were approximately 2 min upon addition of methanol to cells growing exponentially in medium that contained acetate. Results of this study demonstrate that transcription of both cdhA and ORF1 is highly regulated in response to substrate by this methanogenic archaeon.

Aldehyde Oxidoreductases

Sequence and linkage analysis of the Coxiella burnetii citrate synthase-encoding gene.

The nucleotide (nt) sequence of the Coxiella burnetii citrate synthase-encoding gene (gltA), previously cloned in Escherichia coli, was determined. The nt sequence analysis revealed an open reading frame (ORF) of 1290 bp capable of coding for a protein of 430 amino acids (aa) with a deduced Mr of 48,633. Preceding an ATG start codon, a possible transcription start point (tsp) with homology to the E. coli promoter consensus was detected. A poly-purine-rich region occurred immediately upstream from the gltA reading frame and potentially serves as a ribosome-binding site. Additionally, a G + C-rich region of dyad symmetry 3' to the translational stop codon was found that could possibly function as a Rho-independent transcriptional termination signal. A large, nearly perfect, inverted repeat was identified upstream from the gltA tsp and was shown by Southern analysis to be present in multiple copies in the C. burnetii genome. The deduced aa sequence of C. burnetii GltA was optimally aligned with enzymes from various prokaryotic sources and one eukaryotic source (pig heart). Using perfect aa identity, the C. burnetii enzyme demonstrated the greatest homology with GltA from Acinetobacter anitratum (65%). Although only 26% aa identity was seen with the pig heart enzyme, many of the residues identified in ligand binding appear to be conserved. Sequencing studies of a region centered approx. 5.6 kb upstream from gltA revealed an ORF read with opposite polarity that encodes a peptide highly homologous to the C terminus of the flavoprotein subunit of E. coli succinate dehydrogenase. This report represents the first nt sequence analysis of a gene of known function from the obligate intracellular parasite, C. burnetii.

Amino Acid Sequence

Optimizing the promoter and ribosome binding sequence for expression of human single chain urokinase-like plasminogen activator in Escherichia coli and stabilization of the product by avoiding heat shock response.

The expression of recombinant single-chain urokinase-like plasminogen activator (rscuPA) in Escherichia coli was optimized by fusing the puk gene to different promoters and ribosome binding sequences. Comparison of the tac, trp and lambda PL promoters showed that expression was maximal under tac control. Variation in the ribosome binding sequence and its distance to the AUG start codon yielded a further slight improvement of expression. The largest increase in rscuPA expression was achieved by variations in the host strain and growth conditions. In E. coli DG75 grown at 37 degrees C maximal expression was achieved 30 min after induction and decreased gradually until 240 min after induction. Growth at 30 degrees C yielded maximal expression 60 min after induction and resulted in reduced activity at longer times. Western blot analysis of the products showed that degradation of rscuPA was much larger at 37 degrees C than at 30 degrees C. Using E. coli CAG630 carrying the htpR mutation, which avoids heat shock response, for expression of rscuPA eliminated the instability of the product at both temperatures. Expression in this strain was even more efficient than in E. coli JM101 carrying the lon mutation. It is concluded that induction of the general heat-shock response in E. coli must be avoided to obtain stabilization of rscuPA. This drastically improves the overall yield of rscuPA from recombinant E. coli strains.

Base Sequence

A comprehensive set of DnaA-box mutations in the replication origin, oriC, of Escherichia coli.

We probed the complex between the replication origin, oriC, and the initiator protein DnaA using different types of mutations in the five binding sites for DnaA, DnaA boxes R1-R4 and M: (i) point mutations in individual DnaA boxes and combinations of them; (ii) replacement of the DnaA boxes by a scrambled 9 bp non-box motif; (iii) positional exchange; and (iv) inversion of the DnaA boxes. For each of the five DnaA boxes we found at least one type of mutation that resulted in a phenotype. This demonstrates that all DnaA boxes in oriC have a function in the initiation process. Most mutants with point mutations retained some origin activity, and the in vitro DnaA-binding capacity of these origins correlated well with their replication proficiency. Inversion or scrambling of DnaA boxes R1 or M inactivated oriC-dependent replication of joint replicons or minichromosomes under all conditions, demonstrating the importance of these sites. In contrast, mutants with inverted or scrambled DnaA boxes R2 or R4 could not replicate in wild-type hosts but gave transformants in host strains with deleted or compromised chromosomal oriC at elevated DnaA concentrations. We conclude that these origins require more DnaA per origin for initiation than does wild-type oriC. Mutants in DnaA box R3 behaved essentially like wild-type oriC, except for those in which the low-affinity box R3 was replaced by the high-affinity box R1. Apparently, initiation is possible without DnaA binding to box R3, but high-affinity DnaA binding to DnaA box R3 upsets the regulation. Taken together, these results demonstrate that there are finely tuned DnaA binding requirements for each of the individual DnaA boxes for optimal build-up of the initiation complex and replication initiation in vivo.

Bacterial Proteins

Early assignments of the genetic code dependent upon protein structure.

Orgel (1972) has suggested that polynucleotides with sequences of alternating purine/pyrimidine are likely to have predominated in prebiotic conditions. Therefore, in any early template-directed protein synthesis, the number of available codons would have been limited. However, for any self-organizing system to survive and propagate, some feedback must occur from the products of the synthesis to the control of the synthetic procedure itself; i.e. the protein synthesized should have catalysed some step in the initation of template-directed synthesis. A given protein structure with a characteristic conformation and function would be optimal for the product of such a synthesis, and this in turn would limit the number of nucleotide sequences of those available able to give rise to a functioning synthetic assembly. A possible candidate for such an early polypeptide is the ferredoxin group of proteins and it is shown that with the present-day code the corresponding nucleotides do have a high percentage of alternating purine/pyrimidine sequences. Hence these combined restraints on the primitive synthetic machinery would direct the possible assignments of the genetic code helping to explain its regularity and universality.

Amino Acid Sequence

Structural and functional effects of mutations altering the subunit interface of mitochondrial malate dehydrogenase.

Among highly conserved residues in eucaryotic mitochondrial malate dehydrogenases are those with roles in maintaining the interactions between identical monomeric subunits that form the dimeric enzymes. The contributions of two of these residues, Asp-43 and His-46, to structural stability and catalytic function were investigated by construction of mutant enzymes containing Asn-43 and Leu-46 substitutions using in vitro mutagenesis of the Saccharomyces cerevisiae gene (MDH1) encoding mitochondrial malate dehydrogenase. The mutant enzymes were expressed in and purified from a yeast strain containing a disruption of the chromosomal MDH1 locus. The enzyme containing the H46L substitution, as compared to the wild type enzyme, exhibits a dramatic shift in the pH profile for catalysis toward an optimum at low pH values. This shift corresponds with an increased stability of the dimeric form of the mutant enzyme, suggesting that His-46 may be the residue responsible for the previously described pH-dependent dissociation of mitochondrial malate dehydrogenase. The D43N substitution results in a mutant enzyme that is essentially inactive in in vitro assays and that tends to aggregate at pH 7.5, the optimal pH for catalysis for the dimeric wild type enzyme.

Acetates

Detection of K-ras mutations in stools of patients with colorectal cancer by mutant-enriched PCR.

Mutant-enriched PCR was applied to the detection of mutations at codons 12 and 13 of K-ras genes in the stools of patients with colorectal cancer. Mutations were analyzed in stool samples obtained prior to surgery. Resected tumor specimens were screened for K-ras mutations by PCR-mediated RFLP analysis. Using normal stool samples, assay conditions were adjusted to optimal sensitivity and specificity. The following specimens were included in the study: 16 stool samples corresponding to carcinomas in which K-ras mutations had been identified; 7 randomly selected stool samples corresponding to carcinomas which were negative for K-ras mutations; 1 stool sample from a patient with non-Hodgkin's lymphoma. In 13 of the 16 stool samples (81%) corresponding to tumors in which K-ras mutations had been identified previously, K-ras mutations were detected. In 2 of the 7 stool samples corresponding to tumors in which K-ras mutations had not been detected by previous PCR-mediated RFLP analysis, K-ras mutations were also present. Reanalyses of the tumors corresponding to these 2 positive stool samples by mutant-enriched PCR revealed a K-ras mutation in one of the tumors. The stool and tumor of the patient with non-Hodgkins lymphoma were negative for K-ras mutations. DNA sequence analysis revealed that, for each of the K-ras mutations identified in stool samples, identical base substitutions were present in the corresponding tumor tissue. The results indicate that tumor cells harboring K-ras mutations can be detected in the stools of patients with colorectal cancer by mutant-enriched PCR with high sensitivity and specificity. Because of the simplicity of the technique, it may be suitable for screening of stool samples for mutations of the K-ras gene.

Adult

Progress in genetic screening of multiple endocrine neoplasia type 2A: is calcitonin testing obsolete?

BACKGROUND: Recent identification of RET mutations in multiple endocrine neoplasia type 2A (MEN 2A) allows a DNA-based approach to diagnosis in lieu of calcitonin sampling. To prospectively evaluate the efficacy of mutational analysis, genetic screening was performed in 124 patients (53 male, 71 female; age, 1 month to 80 years) at risk for MEN 2A referred over 3 months. METHODS: Analysis used genomic DNA and a polymerase chain reaction-based denaturing gradient gel electrophoresis strategy for mutation detection at RET codons 609, 611, 618, 620, and 634. Ninety-three of 124 patients were from established MEN 2A kindreds (group A), and screening replaced calcitonin testing. Twenty-one of 124 patients (group B) represented index cases of medullary thyroid carcinoma (MTC), and DNA analysis was performed to distinguish sporadic from hereditary disease. Ten patients (group C) had modest calcitonin elevations or had undergone thyroidectomy without confirming pathologic results, and testing was undertaken to clarify status. RESULTS: Group A: RET mutations occurred in 29 (median age, 10 years) of 93 patients, 14 of whom underwent thyroidectomy. No false-positive results were observed. Group B: five (24%) of 21 patients with seemingly sporadic MTC had RET mutations at codons 618 (one), 620 (one), or 634 (three). Group C: Nine of 10 patients with alleged MEN 2A had genetically negative results. CONCLUSIONS: Denaturing gradient gel electrophoresis reliably detects MEN 2A. Modest calcitonin elevations may lead to a false-positive diagnosis of MTC. DNA testing is the optimal approach to evaluating MEN 2A. Index cases of sporadic MTC should also undergo DNA analysis.

Adolescent

A convenient and adaptable package of DNA sequence analysis programs for microcomputers.

We describe a package of DNA data handling and analysis programs designed for microcomputers. The package is convenient for immediate use by persons with little or no computer experience, and has been optimized by trial in our group for a year. By typing a single command, the user enters a system which asks questions or gives instructions in English. The system will enter, alter, and manage sequence files or a restriction enzyme library. It generates the reverse complement, translates, calculates codon usage, finds restriction sites, finds homologies with various degrees of mismatch, and graphs amino acid composition or base frequencies. A number of options for data handling and printing can be used to produce figures for publication. The package will be available in ANSI Standard FORTRAN for use with virtually any FORTRAN compiler.

Amino Acid Sequence