Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Sequences of the E. coli uvrC gene and protein.

We have determined the sequence of a 2400 bp region of E. coli chromosomal DNA containing the uvrC gene. The coding region of uvrc is 2267 bp in length, encodes a polypeptide with a calculated molecular weight of 66,038 daltons, and is preceded by a typical E. coli ribosome binding site. By constructing deletion derivatives we have established that a uvrC promoter lies within the 113 bp region preceding the translational start of uvrC. The codon usage in uvrC is strongly biased in favor of codons used infrequently in E. coli, which may contribute to the relatively low intracellular concentration of uvrC protein.

Amino Acid Sequence↗

Nucleotide sequence of the LuxC gene and the upstream DNA from the bioluminescent system of Vibrio harveyi.

The nucleotide sequence of the luxC gene (1431 bp) and the upstream DNA (1049 bp) of the luminescent bacterium Vibrio harveyi has been determined. The luxC gene can be translated into a polypeptide of 55 kDa in excellent agreement with the molecular mass of the reductase polypeptide required for synthesis of the aldehyde substrate for the bioluminescent reaction. Analyses of codon usage showed a high frequency (1.9%) of the isoleucine codon, AUA, in the luxC gene compared to that found in Escherichia coli genes (0.2%) and its absence in the luxA, B and D genes. The low G/C content of the luxC gene and upstream DNA (38-39%) compared to that found in the other lux genes of V. harveyi (45%) was primarily due to a stretch of 500 nucleotides with only a 24% G/C content, extending from 200 bp inside lux C to 300 bp upstream. Moreover, an open reading frame did not extend for more than 48 codons between the luxC gene and 600 bp upstream at which point a gene transcribed in the opposite direction started. As the lux system in the luminescent bacterium, V. fischeri, contains a regulatory gene immediately upstream of luxC transcribed in the same direction, these results show that the organization and regulation of the lux genes have diverged in different luminescent bacteria.

Amino Acid Sequence↗

TransTerm, the translational signal database, extended to include full coding sequences and untranslated regions.

TransTerm is a database of mRNA sequences and parameters useful for detecting translational control signals in general. TransTerm-98 has been expanded beyond previous years to include full coding sequences and UTRs, while retaining the original small contexts about the coding sequence start- and stop-codons. The database contains more than 130 000 non-redundant coding sequences with associated untranslated regions (UTRs) from over 450 species. This includes the complete genomes of 12 prokaryotic and one eukaryotic organism. Several coding sequence parameters are available: coding sequence length, Nc, GC3 and, when it is computable, Codon Adaptation Index (CAI). Codon usage tables and summaries of start- and stop-codon contexts are also included. TransTerm-98 has both a relational database form with a WWW interface and a flatfile format, also available by Internet browser. TransTerm is available at: http://biochem.otago.ac.nz:800/Transterm/homepage.h tml

Codon↗

The complete mitochondrial sequence of Tarsius bancanus: evidence for an extensive nucleotide compositional plasticity of primate mitochondrial DNA.

Inconsistencies between phylogenetic interpretations obtained from independent sources of molecular data occasionally hamper the recovery of the true evolutionary history of certain taxa. One prominent example concerns the primate infraordinal relationships. Phylogenetic analyses based on nuclear DNA sequences traditionally represent Tarsius as a sister group to anthropoids. In contrast, mitochondrial DNA (mtDNA) data only marginally support this affiliation or even exclude Tarsius from primates. Two possible scenarios might cause this conflict: a period of adaptive molecular evolution or a shift in the nucleotide composition of higher primate mtDNAs through directional mutation pressure. To test these options, the entire mt genome of Tarsius bancanus was sequenced and compared with mtDNA of representatives of all major primate groups and mammals. Phylogenetic reconstructions at both the amino acid (AA) and DNA level of the protein-coding genes led to faulty tree topologies depending on the algorithms used for reconstruction. We propose that these artifactual affiliations rather reflect the nucleotide compositional similarity than phylogenetic relatedness and favor the directional mutation pressure hypothesis because: (1) the overall nucleotide composition changes dramatically on the lineage leading to higher primates at both silent and nonsilent sites, and (2) a highly significant correlation exists between codon usage and the nucleotide composition at the third, silent codon position. Comparisons of mt genes with mt pseudogenes that presumably transferred to the nucleus before the directional mutation pressure took place indicate that the ancestral DNA composition is retained in the relatively fossilized mtDNA-like sequences, and that the directed acceleration of the substitution rate in higher primates is restricted to mtDNA.

Animals↗

Nucleotide sequence and characteristics of the gene for L-lactate dehydrogenase of Thermus caldophilus GK24 and the deduced amino-acid sequence of the enzyme.

The gene for L-lactate dehydrogenase (LDH) (EC 1.1.1.27) of Thermus caldophilus GK24 was cloned in Escherichia coli using synthetic oligonucleotides as hybridization probes. The nucleotide sequence of the cloned DNA was determined. The primary structure of the LDH was deduced from the nucleotide sequence. The deduced amino acid sequence agreed with the NH2-terminal and COOH-terminal sequences previously reported and the determined amino acid sequences of the peptides obtained from trypsin-digested T. caldophilus LDH. The LDH comprised 310 amino acid residues and its molecular mass was determined to be 32,808. On alignment of the whole amino acid sequences, the T. caldophilus LDH showed about 40% identity with the Bacillus stearothermophilus, Lactobacillus casei and dogfish muscle LDHs. The T. caldophilus LDH gene was expressed with the E. coli lac promoter in E. coli, which resulted in the production of the thermophilic LDH. The gene for the T. caldophilus LDH showed more than 40% identity with those for the human and mouse muscle LDHs on alignment of the whole nucleotide sequences. The G + C content of the coding region for the T. caldophilus LDH was 74.1%, which was higher than that of the chromosomal DNA (67.2%). The G + C contents in the first, second and third positions of the codons used were 77.7%, 48.1% and 95.5% respectively. The high G + C content in the third base caused extremely non-random codon usage in the LDH gene. About half (48.7%) the codons in the LDH gene started with G, and hence there were relatively high contents of Val, Ala, Glu and Gly in the LDH. The contents of Pro, Arg, Ala and Gly, which have high G + C contents in their codons, were also high. Rare codons with U or A as the third base were sometimes used to avoid the TCGA sequence, the recognition site for the restriction endonuclease, TaqI. Two TCGA sequences were found only in the sequence of CTCGAG (XhoI site) in the sequenced region of the T. caldophilus DNA. There were three segments with similar sequences in the two 5' non-coding regions, probably the promoter and ribosome-binding regions, of the genes for the T. caldophilus LDH and the Thermus thermophilus 3-isopropylmalate dehydrogenase.

Amino Acid Sequence↗

Expression of the strA-strB streptomycin resistance genes in Pseudomonas syringae and Xanthomonas campestris and characterization of IS6100 in X. campestris.

Expression of the strA-strB streptomycin resistance (SMr) genes was examined in Pseudomonas syringae pv. syringae and Xanthomonas campestris pv. vesicatoria. The strA-strB genes in P. syringae and X. campestris were encoded on elements closely related to Tn5393 from Erwinia amylovora and designated Tn5393a and Tn5393b, respectively. The putative recombination site (res) and resolvase-repressor (tnpR) genes of Tn5393 from E. amylovora, P syringae, and X. campestris were identical; however, IS6100 mapped within tnpR in X. campestris, and IS1133 was previously located downstream of tnpR in E. amylovora (C.-S Chiou and A. L. Jones, J. Bacteriol. 175:732-740, 1993). Transcriptional fusions (strA-strB::uidA) indicated that a strong promoter sequence was located within res in Tn5393a. Expression from this promoter sequence was reduced when the tnpR gene was present in cis position relative to the promoter. In X. campestris pv. vesicatoria, analysis of promoter activity with transcriptional fusions indicated that IS6100 increased the expression of strA-strB. Analysis of codon usage patterns and percent G+C in the third codon position indicated that IS6100 could have originated in a gram-negative bacterium. The data obtained in the present study help explain differences observed in the levels of SMr expressed by three genera which share common genes for resistance. Furthermore, the widespread dissemination of Tn5393 and derivatives in phytopathogenic prokaryotes confirms the importance of these bacteria as reservoirs of antibiotic resistance in the environment.

Base Sequence↗

Characterization of the rfc region of Shigella flexneri.

The O antigen of the Shigella flexneri lipopolysaccharide (LPS) is an important virulence determinant and immunogen. We have isolated S. flexneri mutants which produce a semi-rough LPS by using an O-antigen-specific phage, Sf6c. Western immunoblotting was used to show that the LPS produced by the semi-rough mutants contained only one O-antigen repeat unit. Thus, the mutants are deficient in production of the O-antigen polymerase and were termed rfc mutants. Complementation experiments were used to locate the rfc adjacent to the rfb genes on plasmid clones previously isolated and containing this region (D. F. Macpherson, R. Morona, D. W. Beger, K.-C. Cheah, and P. A. Manning, Mol. Microbiol 5:1491-1499, 1991). A combination of deletions and subcloning analysis located the rfc gene as spanning a 2-kb region. Insertion of a kanamycin resistance cartridge into a SalI site in this region inactivated the rfc gene. The DNA sequence of the rfc region was determined. An open reading frame spanning the SalI site was identified and encodes a protein with a predicted molecular mass of 43.7 kDa. The predicted protein is highly hydrophobic and showed little sequence homology with any other protein. Comparison of its hydropathy plot with that of other Rfc proteins from Salmonella enterica (typhimurium) and Salmonella enterica (muenchen) revealed that the profiles were similar and that the proteins have 12 or more potential membrane-spanning segments. A comparison of the S. flexneri rfc gene and protein product with other rfc and rfc-like proteins revealed that they have a similarly low percentage of G + C content and have similar codon usage, and all have a high percentage of rare codons. An attempt to identify the S. flexneri Rfc protein was unsuccessful, although proteins encoded upstream and downstream of the rfc gene could be identified. Examination of the distribution of rare or minor codons in the rfc gene revealed that it has several minor codons within the first 25 amino acids. This is in contrast to the upstream gene rfbG, which also has a high percentage of rare codons but whose gene product could be detected. The positioning of the rare codons in the rfc gene may restrict translation and suggests that minor isoaccepting tRNA species may be involved in translational regulation of rfc expression. The low percentage of G + C content of rfc genes may be a consequence of the selection pressure to maintain this form of control.

Amino Acid Sequence↗

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗

The compositional distribution of coding sequences and DNA molecules in humans and murids.

The compositional distributions of coding sequences and DNA molecules (in the 50-100-kb range) are remarkably narrower in murids (rat and mouse) compared to humans (as well as to all other mammals explored so far). In murids, both distributions begin at higher and end at lower GC values. A comparison of homologous coding sequences from murids and humans revealed that their different compositional distributions are due to differences in GC levels in all three codon positions, particularly of genes located at both ends of the distribution. In turn, these differences are responsible for differences in both codon usage and amino acids. When GC levels at first + second codon positions and third codon positions, respectively, of murid genes are plotted against corresponding GC levels of homologous human genes, linear relationships (with very high correlation coefficients and slopes of about 0.78 and 0.60, respectively) are found. This indicates a conservation of the order of GC levels in homologous genes from humans and murids. (The same comparison for mouse and rat genes indicates a conservation of GC levels of homologous genes.) A similar linear relationship was observed when plotting GC levels of corresponding DNA fractions (as obtained by density gradient centrifugation in the presence of a sequence-specific ligand) from mouse and human. These findings indicate that orderly compositional changes affecting not only coding sequences but also noncoding sequences took place since the divergence of murids. Such directional fixations of mutations point to the existence of selective pressures affecting the genome as a whole.

Amino Acid Sequence↗

Sequence diversity and molecular evolution of the merozoite surface antigen 2 of Plasmodium falciparum.

Eleven new alleles of the Plasmodium falciparum merozoite surface antigen 2 (MSA2) from Papua New Guinea were analyzed by direct sequencing of polymerase chain reaction (PCR) products. We have used the sequence information to trace the molecular evolution of MSA2. The repeats of ten alleles belonging to the 3D7 allelic family differed considerably in size, nucleotide sequence, and repeat copy number. In the repeat region of these new alleles, codon usage was extremely biased with an exclusive use of NNT codons. Another new allele sequenced belonged to the FC27 family and confirmed the family-specific conserved structure of 96 and 36 bp repeats. In order to assess sequence microheterogeneity within samples defined as the same genotype by restriction fragment length polymorphism (RFLP), we have analyzed single-strand conformation polymorphism (SSCP) of different samples of the most frequent allele (D10 of the FC27 family) in the study population. No sequence heterogeneity could be detected within the repeat region. Based on analysis of the repeat regions in both allelic families, we discuss the hypothesis of a different evolutionary strategy being represented by each of the allelic families. Kew words: Merozoite surface antigen 2 - Nucleotide sequence comparisons - Molecular evolution

Amino Acid Sequence↗

Gene expression, amino acid conservation, and hydrophobicity are the main factors shaping codon preferences in Mycobacterium tuberculosis and Mycobacterium leprae.

Mycobacterium tuberculosis and Mycobacterium leprae are the ethiological agents of tuberculosis and leprosy, respectively. After performing extensive comparisons between genes from these two GC-rich bacterial species, we were able to construct a set of 275 homologous genes. Since these two bacterial species also have a very low growth rate, translational selection could not be so determinant in their codon preferences as it is in other fast-growing bacteria. Indeed, principal-components analysis of codon usage from this set of homologous genes revealed that the codon choices in M. tuberculosis and M. leprae are correlated not only with compositional constraints and translational selection, but also with the degree of amino acid conservation and the hydrophobicity of the encoded proteins. Finally, significant correlations were found between GC3 and synonymous distances as well as between synonymous and nonsynonymous distances.

Amino Acid Sequence↗

The plastid genome of the critically endangered Valeriana trinervis (= Centranthus trinervis) and insights from comparison with other Valeriana plastomes (Caprifoliaceae).

The first complete plastid genome of the critically endangered species Valeriana trinervis was sequenced, assembled and compared with other published Valeriana plastomes. In this study, we assembled the plastid genome of the critically endangered, endemic species Valeriana trinervis (= Centranthus trinervis) and compare it with all published plastomes of Valeriana. We found not only differences in the inverted repeats boundaries, in the type and abundance of repeats, but also similarities in codon usage and microsatellite numbers. We detected non-canonical start codons in several genes and identified variation in several regions that could be useful for phylogenetic and phylogeographic studies. The phylogenetic tree inference based on both full plastomes and coding sequence data indicated that V. trinervis is sister to all Eurasian Valeriana accessions confirming the phylogenetic position recently investigated. This is the first plastome available for a species of the Mediterranean clade of Valeriana previously known as Centranthus, and it adds further data to understand the evolution and diversification of this systematically debated genus.

Genome, Plastid↗

Nucleotide sequence of the mitochondrial structural gene for subunit 9 of yeast ATPase complex.

We have determined the nucleotide sequence of a segment of Saccharomyces mtDNA that contains the structural gene for one of the subunits (the dicyclohexylcarbodiimide-binding protein) of the mitochondrial ATPase complex. The sequence fits the known amino acid sequence of this protein with the exception of one amino acid. Codon usage is biased in favor of A + T-rich codons. On both sides of the gene, the nucleotide sequence contains less than 4% (mol/mol) G + C for at least 180 nucleotides; these A + T sequences show no evidence of internal repetition. The gene and all the A + T-rich sequence preceding the gene are present in a 12S RNA that is the major transcript of this segment of mtDNA. The nature of the sequences responsible for binding ribosomes to mitochondrial mRNA and for termination of RNA synthesis is considered.

Adenosine Triphosphatases↗

Structure and regulation of the anthranilate synthase genes in Pseudomonas aeruginosa: I. Sequence of trpG encoding the glutamine amidotransferase subunit.

We have determined the DNA sequence of the distal 148 codons of trpE and all of trpG in Pseudomonas aeruginosa. These genes encode, respectively, the large and small (glutamine amidotransferase) subunits of anthranilate synthase, the first enzyme in the tryptophan synthetic pathway. The sequenced region of trpE is homologous with the distal portion of E. coli and Bacillus subtilis trpE, whereas the trpG sequence is homologous to the glutamine amidotransferase subunit genes of a number of bacterial and fungal anthranilate synthases. The two coding sequences overlap by 23 bp. Codon usage in these Pseudomonas genes shows a marked preference for codons ending in G or C, thereby resembling that of trpB, trpA, and several other chromosomal loci from this species and others with a high G + C content in their DNA. The deduced amino acid sequence for the P. aeruginosa trpG gene product differs to a surprising extent from the directly determined amino acid sequence of the glutamine amidotransferase subunit of P. putida anthranilate synthase (Kawamura et al. 1978). This suggests that these two proteins are encoded by loci that duplicated much earlier in the phylogeny of these organisms but have recently assumed the same function. We have also determined 490 bp of DNA sequence distal to trpG but have not ascertained the function of this segment, though it is rich in dyad symmetries.

Amino Acid Sequence↗

Cloning of the Zymomonas mobilis structural gene encoding alcohol dehydrogenase I (adhA): sequence comparison and expression in Escherichia coli.

Zymomonas mobilis ferments sugars to produce ethanol with two biochemically distinct isoenzymes of alcohol dehydrogenase. The adhA gene encoding alcohol dehydrogenase I has now been sequenced and compared with the adhB gene, which encodes the second isoenzyme. The deduced amino acid sequences for these gene products exhibited no apparent homology. Alcohol dehydrogenase I contained 337 amino acids, with a subunit molecular weight of 36,096. Based on comparisons of primary amino acid sequences, this enzyme belongs to the family of zinc alcohol dehydrogenases which have been described primarily in eucaryotes. Nearly all of the 22 strictly conserved amino acids in this group were also conserved in Z. mobilis alcohol dehydrogenase I. Alcohol dehydrogenase I is an abundant protein, although adhA lacked many of the features previously reported in four other highly expressed genes from Z. mobilis. Codon usage in adhA is not highly biased and includes many codons which were unused by pdc, adhB, gap, and pgk. The ribosomal binding region of adhA lacked the canonical Shine-Dalgarno sequence found in the other highly expressed genes from Z. mobilis. Although these features may facilitate the expression of high enzyme levels, they do not appear to be essential for the expression of Z. mobilis adhA.

Alcohol Dehydrogenase↗

Archaeal grpE: transcription in two different morphologic stages of Methanosarcina mazei and comparison with dnaK and dnaJ.

Transcription of the heat shock gene grpE was studied in two different morphologic stages of the archaeon Methanosarcina mazei S-6 that differ in resistance to physical and chemical traumas: single cells and packets. While single cells are directly exposed to environmental changes, such as temperature elevations, cells in packets are surrounded by intercellular and peripheral material that keeps them together in a globular structure which can reach several millimeters in diameter. grpE transcript levels determined by Northern (RNA) blotting peaked after a 15-min heat shock in single cells. In contrast, the highest transcript levels in packets were observed after the longest heat shock tested, 60 min. The same response profiles were demonstrated by primer extension experiments and S1 nuclease analysis. A comparison of the grpE response to heat shock with those of dnaK and dnaJ showed that the grpE transcript level was the most increased, closely followed by that of the dnaK transcript, with that of the dnaJ gene being the least augmented. Transcription of grpE started at the same site under normal and heat shock temperatures, and the transcript was consistently approximately 700 bases long. Codon usage patterns revealed that the three archaeal genes use most codons and have the same codon preference for 61% of the amino acids.

Bacterial Proteins↗

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae↗

Heuristic approach to deriving models for gene finding.

Computer methods of accurate gene finding in DNA sequences require models of protein coding and non-coding regions derived either from experimentally validated training sets or from large amounts of anonymous DNA sequence. Here we propose a new, heuristic method producing fairly accurate inhomogeneous Markov models of protein coding regions. The new method needs such a small amount of DNA sequence data that the model can be built 'on the fly' by a web server for any DNA sequence >400 nt. Tests on 10 complete bacterial genomes performed with the GeneMark.hmm program demonstrated the ability of the new models to detect 93.1% of annotated genes on average, while models built by traditional training predict an average of 93.9% of genes. Models built by the heuristic approach could be used to find genes in small fragments of anonymous prokaryotic genomes and in genomes of organelles, viruses, phages and plasmids, as well as in highly inhomogeneous genomes where adjustment of models to local DNA composition is needed. The heuristic method also gives an insight into the mechanism of codon usage pattern evolution.

Codon↗