Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

A small mitochondrial double-stranded (ds) RNA element associated with a hypovirulent strain of the chestnut blight fungus and ancestrally related to yeast cytoplasmic T and W dsRNAs.

A small double-stranded (ds) RNA element was isolated from a moderately hypovirulent strain of the chestnut blight fungus Cryphonectria parasitica (Murr.) Barr. from eastern New Jersey. Virulence was somewhat lower in the dsRNA-containing strain than in a virulent dsRNA-free control strain, but colony morphology and sporulation levels were comparable. A library of cDNA clones was constructed, and overlapping clones representing the entire genome were sequenced. The 2728-bp dsRNA was considerably smaller than previously characterized C. parasitica dsRNAs, which are 12-13 kb and ancestrally related to the Potyviridae family of plant viruses. Sequence analysis revealed one large open reading frame, but only if mitochondrial codon usage (UGA = Trp) was invoked. Nuclease assays of purified mitochondria confirmed that the dsRNA was localized within mitochondria. Assuming mitochondrial translation, the deduced amino acid sequence had landmarks typical of RNA-dependent RNA polymerases. Alignments of the conserved regions indicate that this dsRNA is more closely related to yeast T and W dsRNAs and single-stranded RNA bacteriophages such as Q beta than to other hypovirulence-associated dsRNAs.

Amino Acid Sequence↗

Gene recognition via spliced sequence alignment.

Gene recognition is one of the most important problems in computational molecular biology. Previous attempts to solve this problem were based on statistics, and applications of combinatorial methods for gene recognition were almost unexplored. Recent advances in large-scale cDNA sequencing open a way toward a new approach to gene recognition that uses previously sequenced genes as a clue for recognition of newly sequenced genes. This paper describes a spliced alignment algorithm and software tool that explores all possible exon assemblies in polynomial time and finds the multiexon structure with the best fit to a related protein. Unlike other existing methods, the algorithm successfully recognizes genes even in the case of short exons or exons with unusual codon usage; we also report correct assemblies for genes with more than 10 exons. On a test sample of human genes with known mammalian relatives, the average correlation between the predicted and actual proteins was 99%. The algorithm correctly reconstructed 87% of genes and the rare discrepancies between the predicted and real exon-intron structures were caused either by short (less than 5 amino acids) initial/terminal exons or by alternative splicing. Moreover, the algorithm predicts human genes reasonably well when the homologous protein is nonvertebrate or even prokaryotic. The surprisingly good performance of the method was confirmed by extensive simulations: in particular, with target proteins at 160 accepted point mutations (PAM) (25% similarity), the correlation between the predicted and actual genes was still as high as 95%.

Algorithms↗

Mercuric ion reduction and resistance in transgenic Arabidopsis thaliana plants expressing a modified bacterial merA gene.

With global heavy metal contamination increasing, plants that can process heavy metals might provide efficient and ecologically sound approaches to sequestration and removal. Mercuric ion reductase, MerA, converts toxic Hg2+ to the less toxic, relatively inert metallic mercury (Hg0) The bacterial merA sequence is rich in CpG dinucleotides and has a highly skewed codon usage, both of which are particularly unfavorable to efficient expression in plants. We constructed a mutagenized merA sequence, merApe9, modifying the flanking region and 9% of the coding region and placing this sequence under control of plant regulatory elements. Transgenic Arabidopsis thaliana seeds expressing merApe9 germinated, and these seedlings grew, flowered, and set seed on medium containing HgCl2 concentrations of 25-100 microM (5-20 ppm), levels toxic to several controls. Transgenic merApe9 seedlings evolved considerable amounts of Hg0 relative to control plants. The rate of mercury evolution and the level of resistance were proportional to the steady-state mRNA level, confirming that resistance was due to expression of the MerApe9 enzyme. Plants and bacteria expressing merApe9 were also resistant to toxic levels of Au3+. These and other data suggest that there are potentially viable molecular genetic approaches to the phytoremediation of metal ion pollution.

Amino Acid Sequence↗

RNA editing in Arabidopsis mitochondria effects 441 C to U changes in ORFs.

On the basis of the sequence of the mitochondrial genome in the flowering plant Arabidopsis thaliana, RNA editing events were systematically investigated in the respective RNA population. A total of 456 C to U, but no U to C, conversions were identified exclusively in mRNAs, 441 in ORFs, 8 in introns, and 7 in leader and trailer sequences. No RNA editing was seen in any of the rRNAs or in several tRNAs investigated for potential mismatch corrections. RNA editing affects individual coding regions with frequencies varying between 0 and 18.9% of the codons. The predominance of RNA editing events in the first two codon positions is not related to translational decoding, because it is not correlated with codon usage. As a general effect, RNA editing increases the hydrophobicity of the coded mitochondrial proteins. Concerning the selection of RNA editing sites, little significant nucleotide preference is observed in their vicinity in comparison to unedited C residues. This sequence bias is, per se, not sufficient to specify individual C nucleotides in the total RNA population in Arabidopsis mitochondria.

Arabidopsis↗

Molecular cloning and characterization of AqpZ, a water channel from Escherichia coli.

The aquaporin family of molecular water channels is widely expressed throughout the plant and animal kingdoms. No bacterial aquaporins are known; however, sequence-related bacterial genes have been identified that encode glycerol facilitators (glpF). By homology cloning, a novel aquaporin-related DNA (aqpZ) was identified that contained no surface N-glycosylation consensus. The aqpZ RNA was not identified in mammalian mRNA by Northern analysis and exhibited bacterial codon usage preferences. Southern analysis failed to demonstrate aqpZ in mammalian genomic DNA, whereas a strongly reactive DNA was present in chromosomal DNA from Escherichia coli and other bacterial species and did not correspond to glpF. The aqpZ DNA isolated from E. coli contained a 693-base pair open reading frame encoding a polypeptide 28-38% identical to known aquaporins. When compared with other aquaporins, aqpZ encodes a 10-residue insert preceding exofacial loop C, truncated NH2 and COOH termini, and no cysteines at known mercury-sensitive sites. Expression of aqpZ cRNA conferred Xenopus oocytes with a 15-fold increase in osmotic water permeability, which was maximal after 5 days of expression, was not inhibited with HgCl2, exhibited a low activation energy (Ea = 3.8 kcal/mol), and failed to transport nonionic solutes such as urea and glycerol. In contrast, oocytes expressing glpF transported glycerol but exhibited limited osmotic water permeability. Phylogenetic comparison of aquaporins and homologs revealed a large separation between aqpZ and glpF, consistent with an ancient gene divergence.

Amino Acid Sequence↗

Rapid identification of yeast proteins on two-dimensional gels.

This work describes a rapid and sensitive technique for the identification of Saccharomyces cerevisiae proteins on two-dimensional gels based on the determination of their amino acid ratios. Specific double labeling with 3H and 14C or 35S-labeled amino acids, chosen among those that are specifically incorporated into proteins without interconversion, allowed an accurate measurement of different amino acid ratios for 200 proteins. A computer program was developed to screen a yeast data base containing 1700 protein sequences and to identify proteins matching the measured Mr, pI, and amino acid ratios. The method, tested with 45 reference proteins, allowed 79 new identifications corresponding to abundant proteins belonging to a few functional families. Some protein spots correspond to homologs of mammalian proteins or to uncharacterized open reading frames. Remarkably, among identified proteins of similar abundance, the organellar proteins have a markedly lower codon usage bias than the cytosolic ones. The double labeling technique is particularly suited to the analysis, on a single two-dimensional gel, of the influence of physiological or genetic changes on yeast protein content.

Amino Acids↗

Cloning, expression, and characterization of polyphosphate glucokinase from Mycobacterium tuberculosis.

Polyphosphate glucokinase from Mycobacterium tuberculosis catalyzes the phosphorylation of glucose using polyphosphate or ATP as the phosphoryl donor. The M. tuberculosis H37Rv gene encoding this enzyme has been cloned, sequenced, and expressed in Escherichia coli. The gene contains an open reading frame for 265 amino acids with a calculated mass of 27,400 daltons. The recombinant polyphosphate glucokinase was purified 189-fold to homogeneity and shown to contain dual enzymatic activities, similar to the native enzyme from H37Ra strain. The high G+C content in the codon usage (64.5%) of the gene and the absence of an E. coli-like promoter consensus sequence are consistent with other mycobacterial genes. Two phosphate binding domains conserved in the eukaryotic hexokinase family were identified in the polyphosphate glucokinase sequence, however, "adenosine" and "glucose" binding motifs were not apparent. In addition, a putative polyphosphate binding region is also proposed for the polyphosphate glucokinase enzyme.

Adenosine↗

Disulfide bond assignment in human interleukin-7 by matrix-assisted laser desorption/ionization mass spectroscopy and site-directed cysteine to serine mutational analysis.

Interleukin-7 (IL-7) is a proteinaceous biological response modifier that has a bioactive tertiary structure dependent on disulfide bond formation. Disulfide bond assignments in human (h)IL-7 are based upon the results of matrix-assisted laser desorption/ionization (MALDI) mass spectroscopy and Cys to Ser mutational analyses. A gene encoding the hIL-7 was synthesized incorporating Escherichia coli codon usage bias and was used to express biologically active protein as determined by stimulation of precursor B-cell proliferation. MALDI mass spectroscopic analysis of trypsin-digested hIL-7 was performed and compared with the anticipated results of a simulated tryptic digestion. Many of the anticipated hIL-7 tryptic fragments were detected including one with a molecular mass equivalent to the sum of two polypeptides linked through a disulfide bond formed from Cys residues (Cys3 and Cys142). Subsequently, Cys to Ser substitution mutational analyses were performed. A hIL-7 variant with all six Cys substituted with Ser was found to be biologically inactive (EC50 > 1 x 10(-7) M). In contrast, a family of single disulfide bond-forming variants of hIL-7 were constructed by reintroducing Cys pairs (Cys3-Cys142, Cys35-Cys130, and Cys48-Cys93), and each could stimulate cell proliferation with an EC50 of 4 x 10(-9), 2 x 10(-8), and 2 x 10(-9) M, respectively. In single disulfide bond-forming mutants of hIL-7, the ability to stimulate cell proliferation was abolished in the presence of 2 mM dithiothreitol. The results presented strongly suggest that only a single disulfide bond is required for hIL-7 to form a tertiary structure capable of stimulating precursor B-cell proliferation.

Amino Acid Sequence↗

Expression of the mitochondrial ADP/ATP carrier in Escherichia coli. Renaturation, reconstitution, and the effect of mutations on 10 positive residues.

Previously, the role of residues in the ADP/ATP carrier (AAC) from Saccharomyces cerevisiae has been studied by mutagenesis, but the dependence of mitochondrial biogenesis on functional AAC impedes segregation of the mutational effects on transport and biogenesis. Unlike other mitochondrial carriers, expression of the AAC from yeast or mammalians in Escherichia coli encountered difficulties because of disparate codon usage. Here we introduce the AAC from Neurospora crassa in E. coli, where it is accumulated in inclusion bodies and establish the reconstitution conditions. AAC expressed with heat shock vector gave higher activity than with pET-3a. Transport activity was absolutely dependent on cardiolipin. The 10 single mutations of intrahelical positive residues and of the matrix repeat (+X+) motif resulted in lower activity, except of R245A. R143A had decreased sensitivity toward carboxyatractylate. The ATP-linked exchange is generally more affected than ADP exchange. This reflects a charge network that propagates positive charge defects to ATP(4-) more strongly than to ADP(3-) transport. Comparison to the homologous mutants of yeast AAC2 permits attribution of the roles of these residues more to ADP/ATP transport or to AAC import into mitochondria.

Amino Acid Sequence↗

The typically mitochondrial DNA-encoded ATP6 subunit of the F1F0-ATPase is encoded by a nuclear gene in Chlamydomonas reinhardtii.

The atp6 gene, encoding the ATP6 subunit of F(1)F(0)-ATP synthase, has thus far been found only as an mtDNA-encoded gene. However, atp6 is absent from mtDNAs of some species, including that of Chlamydomonas reinhardtii. Analysis of C. reinhardtii expressed sequence tags revealed three overlapping sequences that encoded a protein with similarity to ATP6 proteins. PCR and 5'- and 3'-RACE were used to obtain the complete cDNA and genomic sequences of C. reinhardtii atp6. The atp6 gene exhibited characteristics of a nucleus-encoded gene: Southern hybridization signals consistent with nuclear localization, the presence of introns, and a codon usage and a polyadenylation signal typical of nuclear genes. The corresponding ATP6 protein was confirmed as a subunit of the mitochondrial F(1)F(0)-ATP synthase from C. reinhardtii by N-terminal sequencing. The predicted ATP6 polypeptide has a 107-amino acid cleavable mitochondrial targeting sequence. The mean hydrophobicity of the protein is decreased in those transmembrane regions that are predicted not to participate directly in proton translocation or in intersubunit contacts with the multimeric ring of c subunits. This is the first example of a mitochondrial protein with more than two transmembrane stretches, directly involved in proton translocation, that is nucleus-encoded.

Adenosine Triphosphatases↗

Type I shorthorn sculpin antifreeze protein: recombinant synthesis, solution conformation, and ice growth inhibition studies.

A number of structurally diverse classes of "antifreeze" proteins that allow fish to survive in sub-zero ice-laden waters have been isolated from the blood plasma of cold water teleosts. However, despite receiving a great deal of attention, the one or more mechanisms through which these proteins act are not fully understood. In this report we have synthesized a type I antifreeze polypeptide (AFP) from the shorthorn sculpin Myoxocephalus scorpius using recombinant methods. Construction of a synthetic gene with optimized codon usage and expression as a glutathione S-transferase fusion protein followed by purification yielded milligram amounts of polypeptide with two extra residues appended to the N terminus. Circular dichroism and NMR experiments, including residual dipolar coupling measurements on a 15N-labeled recombinant polypeptide, show that the polypeptides are alpha-helical with the first four residues being more flexible than the remainder of the sequence. Both the recombinant and synthetic polypeptides modify ice growth, forming facetted crystals just below the freezing point, but display negligible thermal hysteresis. Acetylation of Lys-10, Lys-20, and Lys-21 as well as the N terminus of the recombinant polypeptide gave a derivative that displays both thermal hysteresis (0.4 degrees C at 15 mg/ml) and ice crystal faceting. These results confirm that the N terminus of wild-type polypeptide is functionally important and support our previously proposed mechanism for all type I proteins, in which the hydrophobic face is oriented toward the ice at the ice/water interface.

Amino Acid Sequence↗

Determination of the functionality of common APOA5 polymorphisms.

Common variants of APOA5 have consistently shown association with differences in plasma triglyceride (TG) levels. These single nucleotide polymorphisms (SNPs) fall into three common haplotypes: APOA5*1, with common alleles at all sites; APOA5*2, with rare alleles of -1131T--> C, -3A--> G, 751G--> T, and 1891T--> C; and APOA5*3, distinguished by the c56C--> G (S19W). Molecular modeling of the apoAV signal peptide (SP) showed an increased angle of insertion (65 degrees ) at the lipid/water interface of Trp-19 SP compared with Ser-19 SP (40 degrees ), predicting reduced translocation. This was confirmed by 50% reduction of Trp-19-encoded SP.secretory alkaline phosphatase (SEAP) fusion protein secreted into the medium from HepG2 cells compared with the Ser-19.SEAP fusion protein (p < 0.002). Considering APOA5*2 SNPs, there was no significant difference in the relative luciferase expression in Huh7 cells transiently transfected with a -1131T construct compared with the -1131C (fragments -1177 to -516 or -1177 to -3). Similarly, for the -3A--> G in the Kozak sequence, in vitro transcription/translation assays and primer extension inhibition assays showed no alternate AUG initiation codon usage, demonstrating that -3A--> G did not influence translation efficiency. Although 1891T--> C in the 3'-untranslated region disrupts a putative Oct-1 transcription factor binding site, when inserted 3' of the luciferase gene the T--> C change demonstrated no significant difference in luciferase expression. Thus, association of APOA5*2 SNPs with TG levels is not due to the individual effects of any of these SNPs, although cooperativity between the SNPs cannot be excluded. Alternatively, the effect on TG levels may reflect the strong linkage disequilibrium with the functional APOC3 SNPs.

3' Untranslated Regions↗

A maize homologue of the bacterial CMP-3-deoxy-D-manno-2-octulosonate (KDO) synthetases. Similar pathways operate in plants and bacteria for the activation of KDO prior to its incorporation into outer cellular envelopes.

The eight-carbon acid sugar 3-deoxy-d-manno-2-octulosonate (KDO) is an essential component of Gram-negative bacterial cell walls and capsular polysaccharides. KDO is incorporated into these polymers as CMP-KDO, which is produced in an unusual activation step catalyzed by the enzyme CMP-KDO synthetase. CMP-KDO synthetase activity has traditionally been considered exclusive to Gram-negative bacteria. CMP-KDO synthetase inhibitors attract great interest owing to their potential as selective bactericides. The sugar KDO is also a component of the rhamnogalacturonan II pectin fraction of the primary cell walls of most higher plants and of the cell wall polysaccharides of some green algae. However, the metabolic pathway leading to its incorporation into the plant cell wall is unknown. This paper describes the isolation and characterization of a maize gene, which codes for a protein very similar in sequence and activity to prokaryotic CMP-KDO synthetases. Remarkably, the maize gene can complement a CMP-KDO synthetase (kdsB) Salmonella typhimurium mutant defective in cell wall synthesis. ZmCKS activity is novel in eukaryotes. The evolutionary origin of ZmCKS is discussed in relation to the high degree of conservation between the plant and bacterial genes and its atypical codon usage in maize.

Amino Acid Sequence↗

Phylogeny and evolution of SARS-CoV-2 during Delta and Omicron variant waves in India.

SARS-CoV-2 evolution has continued to generate variants, responsible for new pandemic waves locally and globally. Varying disease presentation and severity has been ascribed to inherent variant characteristics and vaccine immunity. This study analyzed genomic data from 305 whole genome sequences from SARS-CoV-2 patients before and through the third wave in India. Delta variant was reported in patients without comorbidity (97%), while Omicron BA.2 was reported in patients with comorbidity (77%). Tissue adaptation studies brought forth higher propensity of Omicron variants to bronchial tissue than lung, contrary to observation in Delta variants from Delhi. Study of codon usage pattern distinguished the prevalent variants, clustering them separately, Omicron BA.2 isolated in February grouped away from December strains, and all BA.2 after December acquired a new mutation S959P in ORF1b (44.3% of BA.2 in the study) indicating ongoing evolution. Loss of critical spike mutations in Omicron BA.2 and gain of immune evasion mutations including G142D, reported in Delta but absent in BA.1, and S371F instead of S371L in BA.1 could explain very brief period of BA.1 in December 2021, followed by complete replacement by BA.2. Higher propensity of Omicron variants to bronchial tissue, probably ensured increased transmission while Omicron BA.2 became the prevalent variant possibly due to evolutionary trade-off. Virus evolution continues to shape the epidemic and its culmination.Communicated by Ramaswamy H. Sarma.

SARS-CoV-2↗

Complete mitochondrial DNA sequence and amino acid analysis of the cytochrome C oxidase subunit I (COI) from Aedes aegypti.

The complete sequence of the yellow fever mosquito, Aedes aegypti, mitochondrial cytochrome c oxidase subunit 1 gene has been identified. The nucleotide sequence codes for a 512 amino acid peptide. The AeCOI sequence is A + T rich (68.6%) and the codon usage is highly biased toward a preference for A- or T-ending triplets. The A. aegypti COI peptide shows high homology, up to 93% identity, with several other insect sequences and a phylogenetic analysis indicates that the A. aegypti sequence is closely related to two other mosquito species, Anopheles gambiae and A. quadrimaculatus. Comparisons of the nucleotide sequence for four A. aegypti laboratory strains revealed single nucleotide polymorphisms, with 25 nucleotide sites showing SNPs between strains. All SNPs occurred as synonymous transitions such that the peptide sequence is conserved among A. aegypti strains. RT-PCR analysis showed that COI is expressed at similar levels in all developmental stages and tissues.

AT Rich Sequence↗

Cloning and sequence analysis of putative glyceraldehyde-3-phosphate dehydrogenase gene from Monascus purpureus KCCM11832.

Using a synthetic oligonucleotide probe, glyceraldehyde-3-phosphate dehydrogenase gene (gpd1) was cloned from Monascus purpureus KCCM11832. The 2834 bp EcoRV-HindIII region harbored 1183 bp 5'-UTR containing such regulatory elements as CT box, common in fungal gpd's, and gpd box previously found exclusively in Aspergillus gpd's. Full-length cDNA was cloned by PCR, and its sequence was determined. Transcription starting point was located 88 bp upstream from start codon. Polyadenylation signal sequence occurred 201 bp downstream from stop codon. Region from start codon ATG to stop codon TAA including introns showed 62 approximately 69% nucleotide sequence identity to those of Aspergillus gpd's. Significant bias in third position, with pyrimidines favored over purines, was observed in codon usage. The deduced amino acid sequence had 81 approximately 85% identity to Aspergillus gpd's. Monascus purpureus GPD was located at the same clade with Aspergillus GPD's.

Amino Acid Sequence↗

Expression of human metallothionein III and its metalloabsorption capability in Escherichia coli.

Human metallothionein III (MT III) gene was synthesized with Escherichia coli preference codon usage and expressed in E. coli in glutathione-S-transferase (GST) fusion form. The recombinant MT III was released by proteinase Factor Xa digestion and purified with the yield of 2 mg/L culture, and its specific Cd2+ binding capability was confirmed. E. coli strain BL21(DE3), expressing MT III, showed metal tolerance between 0.1 and 0.5 mM Cd2+ and bacterial growth was inhibited at 1 mM Cd2+. MT III expressing E. coli strain showed binding discrimination between different metal ions in combination use, with the preference order of Cd2+ > Cu2+ > Zn2+. It absorbed different metal ions with relatively constant ratio and showed a cumulative absorption capability for mixed heavy metals.

Cloning, Molecular↗

Three differentially expressed Na,K-ATPase alpha subunit isoforms: structural and functional implications.

We have characterized cDNAs coding for three Na,K-ATPase alpha subunit isoforms from the rat, a species resistant to ouabain. Northern blot and S1-nuclease mapping analyses revealed that these alpha subunit mRNAs are expressed in a tissue-specific and developmentally regulated fashion. The mRNA for the alpha 1 isoform, approximately equal to 4.5 kb long, is expressed in all fetal and adult rat tissues examined. The alpha 2 mRNA, also approximately equal to 4.5 kb long, is expressed predominantly in brain and fetal heart. The alpha 3 cDNA detected two mRNA species: a approximately equal to 4.5 kb mRNA present in most tissues and a approximately equal to 6 kb mRNA, found only in fetal brain, adult brain, heart, and skeletal muscle. The deduced amino acid sequences of these isoforms are highly conserved. However, significant differences in codon usage and patterns of genomic DNA hybridization indicate that the alpha subunits are encoded by a multigene family. Structural analysis of the alpha subunits from rat and other species predicts a polytopic protein with seven membrane-spanning regions. Isoform diversity of the alpha subunit may provide a biochemical basis for Na,K-ATPase functional diversity.

Amino Acid Sequence↗