Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Kozak sequence polymorphism of the glycoprotein (GP) Ibalpha gene is a major determinant of the plasma membrane levels of the platelet GP Ib-IX-V complex.

Despite the known importance of the sequences surrounding ATG start codons (Kozak sequences) for efficient translation of proteins, few reports have appeared that describe the natural variations in these sequences. Here, we report a human polymorphism in the Kozak sequence of the platelet adhesion receptor, glycoprotein (GP) Ibalpha, a component of the GP Ib-IX-V complex, which mediates the initial adhesion of platelets to the blood vessel wall following injury. The polymorphism is based on the presence of either thymine (T) or cytosine (C) at position -5 from the initiator ATG in the GP Ibalpha gene. The less common allele, -5C, represented 8% to 17% of the alleles in four ethnic populations surveyed. This allele more closely resembles the sequence considered optimal for efficient initiation of protein translation and is associated with increased expression of the receptor on the cell membrane, both in transfected cells and in the platelets of individuals carrying the allele. In vitro transcription/translation studies indicate that the increased expression results from more efficient translation of the -5C form of the GP Ibalpha mRNA. Other mutations made to approximate more closely the consensus sequence described by Kozak did not increase expression of the receptor. This is the first known description of Kozak sequence polymorphism as a determinant of the surface levels of a cell adhesion receptor. This polymorphism may influence an individual's susceptibility for the development of cardiovascular disease.

Alleles↗

Capped mRNA with a single nucleotide leader is optimally translated in a primitive eukaryote, Giardia lamblia.

The 5'-untranslated region (5'-UTR) of an mRNA plays an important role in translation initiation in eukaryotes. A minimal length of about 20 nucleotides is required to prevent leaky ribosome scanning. In one of the most primitive eukaryotes, Giardia lamblia, however, the mRNAs have 5'-UTRs mostly in the range of 0 to 14 nucleotides without a conserved sequence, which raises the question on how the ribosome could effectively scan such short 5'-UTRs for an accurate initiation of translation. In the present study, we expressed capped transcripts of luciferase gene in Giardia trophozoites via transfection and observed that when the 5'-UTR of the transcript was lengthened from 9 to 21 nucleotides, there was a corresponding decrease of translation efficiency. Conversely, shortening of the 5'-UTR from nine nucleotides down to a single nucleotide did not result in any reduced translation or leaky scanning. Translation appeared to initiate exclusively from the first initiation codon located downstream from the cap. Experimental evidence indicated also that a stem-loop structure immediately downstream from the initiation codon exerted significant inhibition on translation initiation when the 5'-UTR consisted of less than seven nucleotides. This inhibitory effect was abolished by increasing the distance between the stem-loop and the cap-G structure either upstream or downstream from the start codon, thus suggesting a spatial requirement for effective ribosome recruitment. Overall, our results suggest an absence of ribosome scanning for AUG in initiating translation in Giardia. A capped mRNA with a single nucleotide leader is apparently sufficient for recruiting ribosome and initiating translation.

5' Untranslated Regions↗

A program for selecting DNA fragments to detect mutations by denaturing gel electrophoresis methods.

A computer program was developed to automate the selection of DNA fragments for detecting mutations within a long DNA sequence by denaturing gel electrophoresis methods. The program, MELTSCAN, scans through a user specified DNA sequence calculating the melting behavior of overlapping DNA fragments covering the sequence. Melting characteristics of the fragments are analyzed to determine the best fragment for detecting mutations at each base pair position in the sequence. The calculation also determines the optimal fragment for detecting mutations within a user specified mutational hot spot region. The program is built around the statistical mechanical model of the DNA melting transition. The optimal fragment for a given position is selected using the criteria that its melting curve has at least two steps, the base pair position is in the fragment's lowest melting domain, and the melting domain has the smallest number of base pairs among fragments that meet the first two criteria. The program predicted fragments for detecting mutations in the cDNA and genomic DNA of the human p53 gene.

Algorithms↗

Use of a multi-way method to analyze the amino acid composition of a conserved group of orthologous proteins in prokaryotes.

BACKGROUND: Amino acids in proteins are not used equally. Some of the differences in the amino acid composition of proteins are between species (mainly due to nucleotide composition and lifestyle) and some are between proteins from the same species (related to protein function, expression or subcellular localization, for example). As several factors contribute to the different amino acid usage in proteins, it is difficult both to analyze these differences and to separate the contributions made by each factor. RESULTS: Using a multi-way method called Tucker3, we have analyzed the amino composition of a set of 64 orthologous groups of proteins present in 62 archaea and bacteria. This dataset corresponds to essential proteins such as ribosomal proteins, tRNA synthetases and translational initiation or elongation factors, which are common to all the species analyzed. The Tucker3 model can be used to study the amino acid variability within and between species by taking into consideration the tridimensionality of the data set. We found that the main factor behind the amino acid composition of proteins is independent of the organism or protein function analyzed. This factor must be related to the biochemical characteristics of each amino acid. The difference between the non-ribosomal proteins and the ribosomal proteins (which are rich in arginine and lysine) is the main factor behind the differences in amino acid composition within species, while G+C content and optimal growth temperature are the main factors behind the differences in amino acid usage between species. CONCLUSION: We show that a multi-way method is useful for comparing the amino acid composition of several groups of orthologous proteins from the same group of species. This kind of dataset is extremely useful for detecting differences between and within species.

Amino Acids↗

Molecular characterization of the alpha-glucosidase gene (malA) from the hyperthermophilic archaeon Sulfolobus solfataricus.

Acidic hot springs are colonized by a diversity of hyperthermophilic organisms requiring extremes of temperature and pH for growth. To clarify how carbohydrates are consumed in such locations, the structural gene (malA) encoding the major soluble alpha-glucosidase (maltase) and flanking sequences from Sulfolobus solfataricus were cloned and characterized. This is the first report of an alpha-glucosidase gene from the archaeal domain. malA is 2,083 bp and encodes a protein of 693 amino acids with a calculated mass of 80.5 kDa. It is flanked on the 5' side by an unusual 1-kb intergenic region. Northern blot analysis of the malA region identified transcripts for malA and an upstream open reading frame located 5' to the 1-kb intergenic region. The malA transcription start site was located by primer extension analysis to a guanine residue 8 bp 5' of the malA start codon. Gel mobility shift analysis of the malA promoter region suggests that sequences 3' to position -33, including a consensus archaeal TATA box, play an essential role in malA expression. malA homologs were detected by Southern blot analysis in other S. solfataricus strains and in Sulfolobus shibatae, while no homologs were evident in Sulfolobus acidocaldarius, lending further support to the proposed revision of the genus Sulfolobus. Phylogenetic analyses indicate that the closest S. solfataricus alpha-glucosidase homologs are of mammalian origin. Characterization of the recombinant enzyme purified from Escherichia coli revealed differences from the natural enzyme in thermostability and electrophoretic behavior. Glycogen is a substrate for the recombinant enzyme. Unlike maltose hydrolysis, glycogen hydrolysis is optimal at the intracellular pH of the organism. These results indicate a unique role for the S. solfataricus alpha-glucosidase in carbohydrate metabolism.

Amino Acid Sequence↗

Coincident indices of exons and introns.

In this paper, the coincident index, proposed by W. F. Friedman in cryptology, is made use of in DNA sequence analysis and exon prediction. The coincident index of exons exceeds that of introns by many times, and is mainly affected by window length, which is correlated negatively with the coincident index. An optimal exon prediction scheme was obtained by experimental analysis with an orthogonal table. Besides exons, many other special sites such as tandem repeats can be identified by using the coincident index approach. The application of this approach to the ARV-2 (AIDS associated retrovirus 2) genome found three new possible coding regions and some unusual base composition regions which are probably related to definite biological functions.

Base Composition↗

The influence of mRNA primary and secondary structure on human IFN-gamma gene expression in E. coli.

Parameters influencing the efficiency of expression of the human immune interferon (IFN-gamma) gene in E. coli were studied by comparing a series of eight in vitro-derived gene variants. These contained all possible combinations of silent mutations in the first three codons of the mature IFN-gamma polypeptide coding sequence. Expression levels varied up to 50-fold among the different constructions. Comparison of messenger RNA secondary structure models for each variant suggested that the presence of stem-loop structures blocking the translation initiation signals could drastically decrease the efficiency of IFN-gamma synthesis. With variants displaying no stable mRNA secondary structure in the region, a C----U transition at position +11 after the AUG resulted in a 5-fold increase in expression indicating that RNA primary structure also plays an important role in expression. In addition we demonstrate that, in this system, a spacing of 8 nucleotides between the Shine-Dalgarno region and AUG was optimal for gene expression and that the steady-state production level of IFN-gamma rose exponentially with increasing rate of synthesis.

Base Sequence↗

Rapid direct detection of multiple rifampin and isoniazid resistance mutations in Mycobacterium tuberculosis in respiratory samples by real-time PCR.

Rapid detection of resistance in Mycobacterium tuberculosis can optimize the efficacy of antituberculous therapy and control the transmission of resistant M. tuberculosis strains. Real-time PCR has minimized the time required to obtain the susceptibility pattern of M. tuberculosis strains, but little effort has been made to adapt this rapid technique to the direct detection of resistance from clinical samples. In this study, we adapted and evaluated a real-time PCR design for direct detection of resistance mutations in clinical respiratory samples. The real-time PCR was evaluated with (i) 11 clinical respiratory samples harboring bacilli resistant to isoniazid (INH) and/or rifampin (RIF), (ii) 10 culture-negative sputa spiked with a set of strains encoding 14 different resistance mutations in 10 independent codons, and (iii) 16 sputa harboring susceptible strains. The results obtained with this real-time PCR design completely agreed with DNA sequencing data. In all sputa harboring resistant M. tuberculosis strains, the mutation encoding resistance was successfully detected. No mutation was detected in any of the susceptible sputa. The test was applied only to smear-positive specimens and succeeded in detecting a bacterial load equivalent to 10(3) CFU/ml in sputum samples (10 acid-fast bacilli/line). The analytical specificity of this method was proved with a set of 14 different non-M. tuberculosis bacteria. This real-time PCR design is an adequate method for the specific and rapid detection of RIF and INH resistance in smear-positive clinical respiratory samples.

Antibiotics, Antitubercular↗

A complete sequence of the T. tengcongensis genome.

Thermoanaerobacter tengcongensis is a rod-shaped, gram-negative, anaerobic eubacterium that was isolated from a freshwater hot spring in Tengchong, China. Using a whole-genome-shotgun method, we sequenced its 2,689,445-bp genome from an isolate, MB4(T) (Genbank accession no. AE008691). The genome encodes 2588 predicted coding sequences (CDS). Among them, 1764 (68.2%) are classified according to homology to other documented proteins, and the rest, 824 CDS (31.8%), are functionally unknown. One of the interesting features of the T. tengcongensis genome is that 86.7% of its genes are encoded on the leading strand of DNA replication. Based on protein sequence similarity, the T. tengcongensis genome is most similar to that of Bacillus halodurans, a mesophilic eubacterium, among all fully sequenced prokaryotic genomes up to date. Computational analysis on genes involved in basic metabolic pathways supports the experimental discovery that T. tengcongensis metabolizes sugars as principal energy and carbon source and utilizes thiosulfate and element sulfur, but not sulfate, as electron acceptors. T. tengcongensis, as a gram-negative rod by empirical definitions (such as staining), shares many genes that are characteristics of gram-positive bacteria whereas it is missing molecular components unique to gram-negative bacteria. A strong correlation between the G + C content of tDNA and rDNA genes and the optimal growth temperature is found among the sequenced thermophiles. It is concluded that thermophiles are a biologically and phylogenetically divergent group of prokaryotes that have converged to sustain extreme environmental conditions over evolutionary timescale.

Bacillaceae↗

Integrating alternative splicing detection into gene prediction.

BACKGROUND: Alternative splicing (AS) is now considered as a major actor in transcriptome/proteome diversity and it cannot be neglected in the annotation process of a new genome. Despite considerable progresses in term of accuracy in computational gene prediction, the ability to reliably predict AS variants when there is local experimental evidence of it remains an open challenge for gene finders. RESULTS: We have used a new integrative approach that allows to incorporate AS detection into ab initio gene prediction. This method relies on the analysis of genomically aligned transcript sequences (ESTs and/or cDNAs), and has been implemented in the dynamic programming algorithm of the graph-based gene finder EuGENE. Given a genomic sequence and a set of aligned transcripts, this new version identifies the set of transcripts carrying evidence of alternative splicing events, and provides, in addition to the classical optimal gene prediction, alternative optimal predictions (among those which are consistent with the AS events detected). This allows for multiple annotations of a single gene in a way such that each predicted variant is supported by a transcript evidence (but not necessarily with a full-length coverage). CONCLUSIONS: This automatic combination of experimental data analysis and ab initio gene finding offers an ideal integration of alternatively spliced gene prediction inside a single annotation pipeline.

Algorithms↗

Multiple mutation analyses in single tumor cells with improved whole genome amplification.

Combining whole genome amplification (WGA) methods with novel laser-based microdissection techniques has made it possible to exploit recent progress in molecular knowledge of cancer development and progression. However, WGA of one or a few cells has not yet been optimized and systematically evaluated for samples routinely processed in tumor pathology. We therefore studied the value of established WGA protocols in comparison to an improved PEP (I-PEP) PCR method in defined numbers of flow-sorted and microdissected tumor cells obtained both from frozen as well as formalin-fixed and paraffin-embedded tissue sections. In addition, the feasibility of I-PEP-PCR for mutation analysis was tested using clusters of 50-100 unfixed tumor cells obtained by touch preparation of ten breast carcinomas by conventional sequencing of exon 7 and 8 of the p53 gene. Finally, immunocytochemically stained microdissected single disseminated tumor cells from bone marrow aspirates were investigated with respect to mutations in codon 12 of Ki-ras by restriction fragment length polymorphism (RFLP)-PCR after I-PEP-PCR. The modified I-PEP-PCR protocol was superior to the original PEP-PCR and DOP-PCR protocols concerning amplification of DNA from one cell (efficiency rate I-PEP-PCR 40% versus PEP-PCR 15% and DOP-PCR 30%) and five cells (efficiency rate I-PEP-PCR 100% versus PEP-PCR 33% and DOP-PCR 20%). Preamplification by I-PEP allowed 100% sequence accuracy in > 4000 sequenced base pairs and Ki-ras mutation detection in isolated single disseminated tumor cells. For reliable microsatellite analysis of I-PEP-preamplified DNA, at least 10 unfixed cells from fluorescence-activated cell sorting, 10 cells from frozen tissue, or at least 30 cells from formalin-fixed and paraffin-embedded tissue sections were required. Thus, I-PEP-PCR allowed multiple reliable microsatellite analyses suited for microsatellite instability and losses of heterozygosity and mutation analysis even at the single cell level, rendering this technique a powerful new tool for molecular analyses in diagnostic and experimental tumor pathology.

Bone Marrow Neoplasms↗

[Restriction enzyme mismatch polymerase chain reaction in the demonstration of Ki-ras-oncogene mutation in carcinoma of the pancreas].

A pilot study was undertaken to test whether combining the polymerase chain reaction with restriction fragment length polymorphism is suitable for the routine diagnosis of carcinoma of the pancreas. The method makes it possible to recognize point mutations in codon 12 of the Ki-ras oncogene. 60 cytological specimens from the pancreatobiliary tract, the bronchopulmonary system (by bronchoalveolar lavage or from pleural effusion) and ascites were tested. Results from nine pancreas carcinoma cell lines served as control. A PCR product was successfully amplified in all cell lines and 47 of the clinical specimens. In all eight samples in which a mutated Ki-ras oncogene was demonstrated, there was at least a suspicion of malignancy by cytological examination. A mutation was also found in four of five pancreas carcinomas. But no mutation was found in one, clinically certain, case of pancreas carcinoma. The described method is an elegant and, most of all, rapid means to complement and optimize any cytological diagnosis.

Ascitic Fluid↗

Determinants of hepatitis C translational initiation in vitro, in cultured cells and mice.

Hepatitis C virus (HCV) is an RNA virus infecting 1 in every 40 people worldwide. Development of new therapeutics for treating HCV has been hampered by the lack of small-animal models. We have adapted existing hydrodynamic transfection methods to optimize the delivery of RNAs to the cytoplasm of mouse liver cells in vivo. Transfected HCV genomic RNA failed to replicate in mouse liver, suggesting a post-entry block to viral replication. Real-time imaging of HCV internal ribosome entry site (IRES) firefly luciferase reporter mRNA translation in living mice demonstrated that the HCV IRES was functional in mouse liver. We then used this system as a model for studying HCV RNA translation in mice. We compared translation by several mutant HCV IRES variants in cell lysates, cultured cells, and mouse liver. We measured the contribution to translation of a cap, HCV 3'-untranslated region (UTR), poly(A) tail, domains II, IIIb, IIIabc, IIIabcd, IIId, and the initiator codon. Efficient translation required a 3'-UTR in mice and HeLa cells, but not in rabbit reticulocyte lysates. Translational regulation of transfected RNAs was stringent in mice. The method we describe could be useful for studies in mice of antisense or ribozyme inhibitors targeting the IRES as well as other RNA biochemical studies in vivo.

3' Untranslated Regions↗

Engineering hyperexpression of bacteriophage Mu C protein by removal of secondary structure at the translation initiation region.

The structure at the translation initiation region (TIR) of mRNA has pronounced regulatory effects on gene expression. Our attempts to overexpress the C gene of bacteriophage Mu in a variety of expression vectors resulted in low yields of protein. Analysis of Mu C mRNA shows the potential to form a secondary structure involving a ribosome binding site and AUG codon. We have engineered the overproduction of the protein using a PCR-aided cloning approach to remove the sequences involved in the formation of this secondary structure. The overexpressing clone, under the control of T7 gene 10 promoter in a T7 expression system yielded > 30% of total cell protein. The difference in mRNA structure between expressing and non-expressing clones was confirmed by electrophoretic analysis of run-off transcripts. The overexpressed protein was purified in a single step by site-specific DNA affinity chromatography. The purified recombinant protein was active in band shift assays. DNA binding activity required Mg2+ and was weak in the presence of Mn2+. Cd2+ or Zn2+ could not support DNA binding. Under optimal conditions, the equilibrium binding constant (Kapp) was determined to be 2 x 10(12) M-1.

Bacteriophage mu↗

[High expression of LIP1 in Pichia pastoris].

The gene of LIP1, the most important isoenzyme of Candida rugosa lipase (CRL), was artificially synthesized according to its mature peptide sequence. It consisted of 20 codons of preference in Pichia pastoris. The artificial gene was cloned into methanol-inducible expression vector pPICZalphaA, and constitutive expression vector pGAPZalphaA, respectively. The linearized recombinant plasmids were transformed into chromosome of Pichia pastoris SMD1168H strain by electroporation. The abilities of expressing LIP1 in both transformed yeasts had been compared, and the yeast transformed with pGAPZalphaA was more efficient than pPICZalphaA. A recombinant yeast strain named CHT-II expressed LIP1 constitutively, and was the most efficient one. Some enzymatic properties of the recombinant LIP1 were also determined. CHT-II secreted LIP1 into supernate at a level of 2.00x10(5) u/L after 72 h (the cells had been transferred to fresh culture medium at the 24th h). After optimizing the conditions for high cell-density fermentation, the selected yeast strain could secrete LIP1 into supernate at a level of 1.395x10(6) u/L after 72 h. The results indicated that this modification of lip1 gene was successful.

Amino Acid Sequence↗

Accuracy of DNA amplification from archival hematological slides for use in genetic biomarker studies.

Archival slides are a potentially useful source of DNA for mutation analyses in large population-based studies. However, it is unknown whether specimen age or histological stains alter the accuracy of Taq polymerase or induce secondary mutations in sample DNA. To address this question, we evaluated five methods for extraction of genomic DNA from archival bone marrow slides of 17 leukemia patients and analyzed exons 1 and 2 of the N- and K-ras genes for the presence of mutations. Of the five methods, optimal DNA purification was achieved by boiling and phenol:chloroform extraction. N-and K-ras exons 1 and 2 were independently amplified using 35 cycles of PCR, and 6-12 clones for each exon were isolated and individually sequenced for each patient. Mutations were confirmed by repeat extraction, cloning, and sequencing. Sixteen of 17 patient samples were successfully amplified (94%), including slides up to 29 years old. Twelve slides had been stained with Wright-Giemsa, I stained with toluidine blue, and 4 were unstained. A total of 16 single-base mutations were identified of 33,840 nucleotides sequenced. No insertions or deletions were identified. Six of 16 single-base mutations were previously described activating mutations in codon 13 of N-ras exon 1. The 10 other mutations were in other regions of the N- and K-ras genes and were not reproduced after repeat extraction, cloning, and sequencing. The frequency of these other alterations was I of 3384 bp. This value is comparable with the inherent error frequency for Taq polymerase. Our findings suggest that high fidelity DNA amplification can be achieved using archival hematological slides as old as 29 years and can be reliably used in genetic analyses.

Archives↗

Murray valley encephalitis virus envelope protein antigenic variants with altered hemagglutination properties and reduced neuroinvasiveness in mice.

Neutralization escape variants of Murray Valley encephalitis virus were selected using a type-specific, neutralizing, and passively protective anti-envelope protein (E) monoclonal antibody (4B6C-2) which defines epitope E-1c. Nucleotide sequence analysis revealed single nucleotide changes in the E genes of 15 variants resulting in nonconservative amino acid substitutions in all cases. One variant had a three-nucleotide deletion in the E gene which resulted in loss of serine at residue 277. Changes were clustered into two separate regions of the E polypeptide (residues 126-128 and 274-277), indicating that E-1c is a discontinuous epitope. One variant (BHv1), altered at residue 277 (Ser-->Ile), failed to hemagglutinate across the pH range 5.5-7.5, in contrast to parental virus and the other escape variants which hemagglutinated at an optimal pH of 6.6. BHv1 was also of reduced neuroinvasiveness in 21-day-old mice following intraperitoneal inoculation compared to the other viruses. Parental virus and the neutralization escape variants grew equally well in both vertebrate and invertebrate cell cultures, indicating that the reduced neuroinvasiveness of BHv1 was not due to a major abnormality of replication.

Amino Acid Sequence↗

Purification, properties, and genetic location of Escherichia coli cytidine 5'-monophosphate N-acetylneuraminic acid synthetase.

N-Acetylneuraminic acid cytidylyltransferase (EC 2.7.7.43) (CMP-NeuAc synthetase) catalyzes the formation of cytidine monophosphate N-acetylneuraminic acid. We have purified CMP-NeuAc synthetase from an Escherichia coli O18:K1 cytoplasmic fraction to apparent homogeneity by ion exchange chromatography and affinity chromatography on CDP-ethanolamine linked to agarose. The enzyme has a specific activity of 2.1 mumol/mg/min and migrates as a single protein and activity band on nondenaturing polyacrylamide gel electrophoresis. The enzyme has a requirement for Mg2+ or Mn2+ and exhibits optimal activity between pH 9.0 and 10. The apparent Michaelis constants for the CTP and NeuAc are 0.31 and 4 mM, respectively. The CTP analogues 5-mercuri-CTP and CTP-2',3'-dialdehyde are inhibitors. The purified CMP-N-acetylneuraminic acid synthetase has a molecular weight of approximately 50,000 on sodium dodecyl sulfate-polyacrylamide gel electrophoresis. The gene encoding CMP-N-acetylneuraminic acid synthetase is located on a 3.3-kilobase HindIII fragment. The purified enzyme appears to be identical to the 50,000 Mr polypeptide encoded by this gene based on insertion mutations that result in the loss of detectable enzymatic activity. The amino-terminal sequence of the purified protein was used to locate the start codon for the CMP-NeuAc synthetase gene. Both the enzyme and the 50,000 Mr polypeptide have the same NH2-terminal amino acid sequence. Antibodies prepared to a peptide derived from the NH2-terminal amino acid sequence bind to purified CMP-NeuAc synthetase.

Amino Acid Sequence↗