Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Comparative study of translation termination sites and release factors (RF1 and RF2) in procaryotes.

Translation termination is catalyzed by release factors that recognize stop codons. However, previous works have shown that in some bacteria, the termination process also involves bases around stop codons. Recently, Ito et al. analyzed release factors and identified the amino acids therein that recognize stop codons. However, the amino acids that recognize bases around stop codons remain unclear. To identify the candidate amino acids that recognize the bases around stop codons, we aligned the protein sequences of the release factors of various bacteria and searched for amino acids that were conserved specifically in the sequence of bacteria that seemed to regulate translation termination by bases around stop codons. As a result, species having several highly conserved residues in RF1 and RF2 showed positive correlations between their codon usage bias and conservation of the bases around the stop codons. In addition, some of the residues were located very close to the SPF motif, which deciphers stop codons. These results suggest that these conserved amino acids enable the release factors to recognize the bases around the stop codons.

Amino Acid Sequence↗

Gene expression, amino acid conservation, and hydrophobicity are the main factors shaping codon preferences in Mycobacterium tuberculosis and Mycobacterium leprae.

Mycobacterium tuberculosis and Mycobacterium leprae are the ethiological agents of tuberculosis and leprosy, respectively. After performing extensive comparisons between genes from these two GC-rich bacterial species, we were able to construct a set of 275 homologous genes. Since these two bacterial species also have a very low growth rate, translational selection could not be so determinant in their codon preferences as it is in other fast-growing bacteria. Indeed, principal-components analysis of codon usage from this set of homologous genes revealed that the codon choices in M. tuberculosis and M. leprae are correlated not only with compositional constraints and translational selection, but also with the degree of amino acid conservation and the hydrophobicity of the encoded proteins. Finally, significant correlations were found between GC3 and synonymous distances as well as between synonymous and nonsynonymous distances.

Amino Acid Sequence↗

The plastid genome of the critically endangered Valeriana trinervis (= Centranthus trinervis) and insights from comparison with other Valeriana plastomes (Caprifoliaceae).

The first complete plastid genome of the critically endangered species Valeriana trinervis was sequenced, assembled and compared with other published Valeriana plastomes. In this study, we assembled the plastid genome of the critically endangered, endemic species Valeriana trinervis (= Centranthus trinervis) and compare it with all published plastomes of Valeriana. We found not only differences in the inverted repeats boundaries, in the type and abundance of repeats, but also similarities in codon usage and microsatellite numbers. We detected non-canonical start codons in several genes and identified variation in several regions that could be useful for phylogenetic and phylogeographic studies. The phylogenetic tree inference based on both full plastomes and coding sequence data indicated that V. trinervis is sister to all Eurasian Valeriana accessions confirming the phylogenetic position recently investigated. This is the first plastome available for a species of the Mediterranean clade of Valeriana previously known as Centranthus, and it adds further data to understand the evolution and diversification of this systematically debated genus.

Genome, Plastid↗

Nucleotide sequence of the mitochondrial structural gene for subunit 9 of yeast ATPase complex.

We have determined the nucleotide sequence of a segment of Saccharomyces mtDNA that contains the structural gene for one of the subunits (the dicyclohexylcarbodiimide-binding protein) of the mitochondrial ATPase complex. The sequence fits the known amino acid sequence of this protein with the exception of one amino acid. Codon usage is biased in favor of A + T-rich codons. On both sides of the gene, the nucleotide sequence contains less than 4% (mol/mol) G + C for at least 180 nucleotides; these A + T sequences show no evidence of internal repetition. The gene and all the A + T-rich sequence preceding the gene are present in a 12S RNA that is the major transcript of this segment of mtDNA. The nature of the sequences responsible for binding ribosomes to mitochondrial mRNA and for termination of RNA synthesis is considered.

Adenosine Triphosphatases↗

Structure and regulation of the anthranilate synthase genes in Pseudomonas aeruginosa: I. Sequence of trpG encoding the glutamine amidotransferase subunit.

We have determined the DNA sequence of the distal 148 codons of trpE and all of trpG in Pseudomonas aeruginosa. These genes encode, respectively, the large and small (glutamine amidotransferase) subunits of anthranilate synthase, the first enzyme in the tryptophan synthetic pathway. The sequenced region of trpE is homologous with the distal portion of E. coli and Bacillus subtilis trpE, whereas the trpG sequence is homologous to the glutamine amidotransferase subunit genes of a number of bacterial and fungal anthranilate synthases. The two coding sequences overlap by 23 bp. Codon usage in these Pseudomonas genes shows a marked preference for codons ending in G or C, thereby resembling that of trpB, trpA, and several other chromosomal loci from this species and others with a high G + C content in their DNA. The deduced amino acid sequence for the P. aeruginosa trpG gene product differs to a surprising extent from the directly determined amino acid sequence of the glutamine amidotransferase subunit of P. putida anthranilate synthase (Kawamura et al. 1978). This suggests that these two proteins are encoded by loci that duplicated much earlier in the phylogeny of these organisms but have recently assumed the same function. We have also determined 490 bp of DNA sequence distal to trpG but have not ascertained the function of this segment, though it is rich in dyad symmetries.

Amino Acid Sequence↗

Cloning of the Zymomonas mobilis structural gene encoding alcohol dehydrogenase I (adhA): sequence comparison and expression in Escherichia coli.

Zymomonas mobilis ferments sugars to produce ethanol with two biochemically distinct isoenzymes of alcohol dehydrogenase. The adhA gene encoding alcohol dehydrogenase I has now been sequenced and compared with the adhB gene, which encodes the second isoenzyme. The deduced amino acid sequences for these gene products exhibited no apparent homology. Alcohol dehydrogenase I contained 337 amino acids, with a subunit molecular weight of 36,096. Based on comparisons of primary amino acid sequences, this enzyme belongs to the family of zinc alcohol dehydrogenases which have been described primarily in eucaryotes. Nearly all of the 22 strictly conserved amino acids in this group were also conserved in Z. mobilis alcohol dehydrogenase I. Alcohol dehydrogenase I is an abundant protein, although adhA lacked many of the features previously reported in four other highly expressed genes from Z. mobilis. Codon usage in adhA is not highly biased and includes many codons which were unused by pdc, adhB, gap, and pgk. The ribosomal binding region of adhA lacked the canonical Shine-Dalgarno sequence found in the other highly expressed genes from Z. mobilis. Although these features may facilitate the expression of high enzyme levels, they do not appear to be essential for the expression of Z. mobilis adhA.

Alcohol Dehydrogenase↗

Archaeal grpE: transcription in two different morphologic stages of Methanosarcina mazei and comparison with dnaK and dnaJ.

Transcription of the heat shock gene grpE was studied in two different morphologic stages of the archaeon Methanosarcina mazei S-6 that differ in resistance to physical and chemical traumas: single cells and packets. While single cells are directly exposed to environmental changes, such as temperature elevations, cells in packets are surrounded by intercellular and peripheral material that keeps them together in a globular structure which can reach several millimeters in diameter. grpE transcript levels determined by Northern (RNA) blotting peaked after a 15-min heat shock in single cells. In contrast, the highest transcript levels in packets were observed after the longest heat shock tested, 60 min. The same response profiles were demonstrated by primer extension experiments and S1 nuclease analysis. A comparison of the grpE response to heat shock with those of dnaK and dnaJ showed that the grpE transcript level was the most increased, closely followed by that of the dnaK transcript, with that of the dnaJ gene being the least augmented. Transcription of grpE started at the same site under normal and heat shock temperatures, and the transcript was consistently approximately 700 bases long. Codon usage patterns revealed that the three archaeal genes use most codons and have the same codon preference for 61% of the amino acids.

Bacterial Proteins↗

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae↗

Heuristic approach to deriving models for gene finding.

Computer methods of accurate gene finding in DNA sequences require models of protein coding and non-coding regions derived either from experimentally validated training sets or from large amounts of anonymous DNA sequence. Here we propose a new, heuristic method producing fairly accurate inhomogeneous Markov models of protein coding regions. The new method needs such a small amount of DNA sequence data that the model can be built 'on the fly' by a web server for any DNA sequence >400 nt. Tests on 10 complete bacterial genomes performed with the GeneMark.hmm program demonstrated the ability of the new models to detect 93.1% of annotated genes on average, while models built by traditional training predict an average of 93.9% of genes. Models built by the heuristic approach could be used to find genes in small fragments of anonymous prokaryotic genomes and in genomes of organelles, viruses, phages and plasmids, as well as in highly inhomogeneous genomes where adjustment of models to local DNA composition is needed. The heuristic method also gives an insight into the mechanism of codon usage pattern evolution.

Codon↗

Increased levels of glycine tRNA associated with collagen synthesis.

Analysis of codon usage for chick Type I collagen indicates that 89% of glycine codons are GGU/C. Since collagens are one-third glycine, chick Type I collagen synthesis should require large amounts of tRNAGly with the anticodon GCC. Earlier chromatographic studies of chick tRNA had indicated that connective tissues showed altered tRNAGly isoacceptor profiles [P. J. Christner and J. Rosenbloom (1976) Arch. Biochem. Biophys. 172, 399-409; H. J. Drabkin and L. N. Lukens (1978) J. Biol. Chem. 253, 6233-6241]. We have therefore used both two-dimensional gel electrophoresis and hybridization analysis to investigate whether collagen synthesis in chick connective tissues is associated with expression of a novel tRNAGly. Liver and calvaria tRNAs produced qualitatively similar patterns when separated on 2-D gels. Northern blots of 2-D-separated tRNAs from liver and calvaria, when hybridized to genes for vertebrate tRNAGly isoacceptors with GCC or UCC anticodons, showed hybridization to the same tRNAs in both tissues. Quantitation of tRNA species by dot blot hybridization indicated an increase in levels of the tRNAGly isoacceptor with anticodon GCC. Tissues synthesizing Type I collagen had a two- to threefold increase in this tRNA while tissues synthesizing Type II collagen showed a more modest increase. We conclude that elevated tRNAGly levels associated with collagen synthesis are due to increased amounts of the same isoacceptor which is the major tRNAGly in other tissues.

Animals↗

Characterization of two divergent beta-tubulin genes from Colletotrichum graminicola.

We have cloned and sequenced two beta-tubulin genes, TUB1 and TUB2, from the phytopathogenic fungus, Colletotrichum graminicola. The nucleotide sequences of the coding regions of the two genes are only 72.8% homologous. This divergence is reflected in the deduced amino acid (aa) sequences which differ at 94 aa residues. Comparison with the aa sequences of other fungal beta-tubulins indicates that the C. graminicola TUB2 gene encodes a conserved isotype, whereas the C. graminicola TUB1 product is highly divergent. Both genes contain six identically placed introns and the position of each intron is conserved in other fungal beta-tubulin genes. Also typical of other fungal beta-tubulin genes, there is a pronounced bias in codon usage in the C. graminicola TUB2 gene; there is a lesser codon bias in TUB1 from C. graminicola. Both C. graminicola beta-tubulin genes are transcribed and yield similar sized messages.

Amino Acid Sequence↗

Cloning and sequence of several alpha 2u-globulin cDNAs.

We describe a simple cloning procedure for alpha 2u-globulin that requires neither enrichment of mRNA for cloning nor purification of a specific probe for screening recombinant colonies. Total adult male liver poly(A)+RNA was used as template for cloning, and the subsequent recombinant colonies were screened by comparing hybridization to radioactive cDNA probes prepared from hepatic male and female mRNA, respectively. Almost all of the selected "male-specific" clones were later shown to contain alpha 2u-globulin sequences. This cloned alpha 2u-globulin cDNA has been shown to specifically hybridize to male rat liver RNA, which, when isolated and translated in vitro, codes for a 21,000-dalton protein (pro-alpha 2u-globulin) immunologically identical to alpha 2u-globulin. When translation occurs in the presence of pancreatic microsomes this in vitro synthesized pro-alpha 2u-globulin is processed to the 19,000-dalton mature form of alpha 2u-globulin. The nucleotide sequence of the alpha 2u-globulin cDNA has been determined, thus elucidating the complete amino acid sequence of alpha 2u-globulin and most of the hydrophobic "leader" sequence of pro-alpha 2u-globulin. The amino acid sequence deduced from the cDNA is in agreement with the partial sequence that we previously determined by sequential Edman degradation of the purified protein. alpha 2u-Globulin cDNA clones contain within the 3'-untranslated region one or both of the two putative polyadenylylation/transcription termination sites (A-A-T-A-A-A and A-A-T-T-A-A-A). Either of these can be used, generating alpha 2u-globulin mRNA species of two lengths. A codon usage analysis of the cDNA showed that, although all six leucine codons are used for the 14 leucine residues in mature alpha 2u-globulin, the seven leucines in the partial leader sequence reported are all encoded by the same codon, CTG. The primary amino acid sequence contains a unique Asn-Gly-Ser sequence, likely to be in beta-turn conformation, as the probable site of glycosylation for this glycoprotein.

Alpha-Globulins↗

Multiple-alphabet amino acid sequence comparisons of the immunoglobulin kappa-chain constant domain.

We compare the amino acid sequences of the constant domains of the immunoglobulin kappa chain of human, mouse, and rabbit by using four classification schemes ("alphabets") of the 20 amino acids based on their chemical, functional, charge, and structural properties. The comparison reveals three regions of pronounced similarity across the three species, independent of allotype. Two of these regions (residues 65-73 and 99-103) entail a high degree of identity at the DNA level and are distinguished from the rest of the constant domain in codon usage and in the dinucleotide sequence at abutting sites of adjacent codons. Residues 22-29 are highly conserved among the three species in the chemical and functional alphabets but do not show any three-sequence significant amino acid block identities. These results are discussed in terms of transcript processing, effector functions, and structural interactions within the constant domain and with the heavy chain.

Amino Acid Sequence↗

Molecular evolution between Drosophila melanogaster and D. simulans: reduced codon bias, faster rates of amino acid substitution, and larger proteins in D. melanogaster.

Both natural selection and mutational biases contribute to variation in codon usage bias within Drosophila species. This study addresses the cause of codon bias differences between the sibling species, Drosophila melanogaster and D. simulans. Under a model of mutation-selection-drift, variation in mutational processes between species predicts greater base composition differences in neutrally evolving regions than in highly biased genes. Variation in selection intensity, however, predicts larger base composition differences in highly biased loci. Greater differences in the G+C content of 34 coding regions than 46 intron sequences between D. melanogaster and D. simulans suggest that D. melanogaster has undergone a reduction in selection intensity for codon bias. Computer simulations suggest at least a fivefold reduction in Nes at silent sites in this lineage. Other classes of molecular change show lineage effects between these species. Rates of amino acid substitution are higher in the D. melanogaster lineage than in D. simulans in 14 genes for which outgroup sequences are available. Surprisingly, protein sizes are larger in D. melanogaster than in D. simulans in the 34 genes compared between the two species. A substantial fraction of silent, replacement, and insertion/deletion mutations in coding regions may be weakly selected in Drosophila.

Amino Acids↗

Nucleotide sequence of Candida pelliculosa beta-glucosidase gene.

The nucleotide sequence of the DNA fragment containing the beta-glucosidase gene of Candida pelliculosa was determined. Analysis of the sequence revealed three open reading frames which could encode 65,825, and 412 amino acid residues. The presence of the second frame was found to be sufficient for the expression of the beta-glucosidase gene in a heterologous host Saccharomyces cerevisiae. Putative protein encoded by this gene had hydrophobic amino acids, resembling a signal peptide, at its N-terminal region and 19 potential glycosylation sites. Codon usage of Candida genes had the similar pattern shown in S.cerevisiae. Codon bias of the beta-glucosidase gene of Candida was relatively low, compared with that of the highly expressed genes of S. cerevisiae.

Amino Acid Sequence↗

Purification, cloning, and sequence of outer membrane protein P1 of Haemophilus influenzae type b.

Outer membrane protein P1 from Haemophilus influenzae type b MinnA was purified and partially characterized. Antiserum was generated against the purified protein and was used to immunologically screen a lamba EMBL3 genomic library prepared from strain MinnA DNA. A 4.2-kilobase-pair EcoRI-BamHI fragment containing the P1 gene was subcloned into pBR322. The recombinant protein was synthesized by Escherichia coli K-12, in which it localized to the outer membrane. The N-terminal sequence of the purified protein was determined and found to correspond to residues 23 through 36. The 22-amino-acid leader peptide had a typical structure, with two lysine residues near the amino terminus, a stretch of hydrophobic residues, and alanine residues at positions 20 and 22. The Mr of the processed protein was 47,752, which is in good agreement with the estimate of 50,000 from sodium dodecyl sulfate-polyacrylamide gel electrophoresis. Putative -35 and -10 promoter sequences were identified upstream from the translational start site. Codon usage was examined and determined to be substantially different than the codon preference in E. coli.

Amino Acid Sequence↗

Insertion of N-linked glycosylation sites in the variable regions of the human immunodeficiency virus type 1 surface glycoprotein through AAT triplet reiteration.

Variable regions with sequence length variation in the human immunodeficiency virus type 1 envelope exhibit an unusual pattern of codon usage with AAT, ACT, and AGT together composing > 70% of all codons used. We postulate that this distribution is caused by insertion of AAT triplets followed by point mutations and selection. Accumulation of the encoded amino acids (asparagine, serine, and threonine) leads to the creation of new N-linked glycosylation sites, which helps the virus to escape from the immune pressure exerted by virus-neutralizing antibodies.

Base Sequence↗

Amplification and molecular cloning of the IMP dehydrogenase gene of Leishmania donovani.

A mutant (MPA100) strain of Leishmania donovania was generated from a wild type (D1700) population by virtue of its ability to survive the selective pressure of gradually increasing concentrations of mycophenolic acid (MPA), an inhibitor of IMP dehydrogenase (IMPDH) activity. Comparative growth experiments revealed that the MPA100 strain was 100-fold more resistant to MPA toxicity and cross-resistant to ribavarin, another inhibitor of IMPDH. A direct comparison of IMPDH levels in D1700 and MPA100 cells showed that the latter expressed at least 20-fold higher enzyme activity. In order to evaluate the mechanism by which MPA100 cells overexpressed IMPDH, the leishmanial gene encoding IMPDH was isolated from a genomic library in EMBL3 by cross-hybridization to a mouse IMPDH cDNA, and a 2.3-kilobase EcoRV-PstI fragment was subcloned into a Bluescript vector and sequenced. The EcoRV-PstI fragment contained an open reading frame of 514 amino acids that encompassed the entire leishmanial IMPDH coding sequence. The predicted amino acid sequence showed a 52.5% identity with that of the corresponding human IMPDH. The codon usage of the leishmanial IMPDH gene reflected a strong bias toward codons containing either G or C in the wobble position. The EcoRV-PstI fragment hybridized to a 3.0-kilobase mRNA that was expressed at 10-20-fold greater levels in the MPA100 cells. Using the EcoRV-PstI fragment as a probe, the increased amount of IMPDH activity and IMPDH mRNA in the MPA100 cells could be attributed to an approximately 10-20-fold amplification of the leishmanial IMPDH gene.

Amino Acid Sequence↗