Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Accessing and exploring the unusual chemistry by radical SAM-RiPP enzymes.

Radical SAM enzymes involved in the biosynthesis of ribosomally synthesized and post-translationally modified peptides catalyze unusual transformations that lead to unique peptide scaffolds and building blocks. Several natural products from these pathways show encouraging antimicrobial activities and represent next-generation therapeutics for infectious diseases. These systems are uniquely configured to benefit from genome-mining approaches because minimal substrate and cognate modifying enzyme expression can reveal unique, chemically complex transformations that outperform late-stage chemical reactions. This report highlights the main strategies used to reveal these enzymatic transformations, which have relied mainly on genome mining using enzyme-first approaches. We describe the general biosynthetic components for rSAM enzymes and highlight emerging approaches that may broaden the discovery and study of rSAM-RiPP enzymes. The large number of uncharacterized rSAM proteins, coupled with their unpredictable transformations, will continue to be an essential and exciting resource for enzyme discovery.

S-Adenosylmethionine↗

Logical Exploration of Cinnamoyl-Containing Nonribosomal Peptides via Metabologenomic Targeting and Regulator Overexpression.

A targeted method for discovering cinnamoyl-containing nonribosomal peptides (CCNPs), a unique class of bioactive compounds, was devised by using cinnamoyl isomerase, a key enzyme in the biosynthesis of the cinnamoyl moiety, as a genome mining probe. A total of 39 hit strains were obtained, including 35 from polymerase chain reaction-based screening of the in-house bacterial library (2.5% of 1400 strains) targeting the cinnamoyl isomerase-encoding gene and 4 from the genome mining of online databases. Sequence similarity networking and phylogenetic analyses of the isomerase amplicons (∼530 bp) classified the CCNPs into three major substructure-based groups (Z-, E-, and M-type CCNPs) and revealed distinct clade-structure relationships (13 clades). To overcome the challenge of silent biosynthetic gene clusters, we activated these clusters by overexpressing conserved cluster-situated LuxR regulators combined with extensive culture optimization. CCNP production was metabolomically detected in the bacterial extracts by using the characteristic UV absorption and MS/MS fragments of cinnamoyl moieties. CCNP production was observed in 20 of the 39 hit strains, resulting in the isolation of 6 new CCNPs, including oxy-skyllamycin B (2), gwanacinnamycin (3), and luxocinnamycins A-D (4-7), with high structural novelty. Their structures were elucidated using comprehensive spectroscopic analyses and multiple-step chemical derivatizations, and the putative biosynthetic pathways were bioinformatically proposed. Gwanacinnamycin (3) exhibited significant antimycobacterial activity, whereas luxocinnamycin A (4) displayed moderate antiproliferative activity against stomach cancer cells. Our findings highlight a targeted metabologenomic approach combined with transcriptional regulator overexpression as a logical and efficient platform for the discovery of bioactive compounds from nature.

Peptides↗

Repeats in genomic DNA: mining and meaning.

For hundreds of millions of years, perhaps from the very beginning of their evolutionary history, eukaryotic cells have been habitats and junkyards for countless generations of transposable elements, preserved in repetitive DNA sequences. Analysis of these sequences, combined with experimental research, reveals a history of complex 'intracellular ecosystems' of transposable elements that are inseparably associated with genomic evolution.

Animals↗

The University of Minnesota Biocatalysis/Biodegradation Database: post-genomic data mining.

The University of Minnesota Biocatalysis/Biodegradation Database (UM-BBD, http://umbbd.ahc.umn.edu/) provides curated information on microbial catabolism and related biotransformations, primarily for environmental pollutants. Currently, it contains information on over 130 metabolic pathways, 800 reactions, 750 compounds and 500 enzymes. In the past two years, it has increased its breath to include more examples of microbial metabolism of metals and metalloids; and expanded the types of information it includes to contain microbial biotransformations of, and binding interactions with many chemical elements. It has also increased the ways in which this data can be accessed (mined). Structure-based searching was added, for exact matches, similarity, or substructures. Analysis of UM-BBD reactions has lead to a prototype, guided, pathway prediction system. Guided prediction means that the user is shown all possible biotransformations at each step and guides the process to its conclusion. Mining the UM-BBD's data provides a unique view into how the microbial world recycles organic functional groups. UM-BBD users are encouraged to comment on all aspects of the database, including the information it contains and the tools by which it can be mined. The database and prediction system develop under the direction of the scientific community.

Biodegradation, Environmental↗

Utility of the Trypanosoma cruzi sequence database for identification of potential vaccine candidates by in silico and in vitro screening.

Glycosylphosphatidylinositol (GPI)-anchored proteins are abundantly expressed in the infective and intracellular stages of Trypanosoma cruzi and are recognized as antigenic targets by both the humoral and cellular arms of the immune system. Previously, we demonstrated the efficacy of genes encoding GPI-anchored proteins in eliciting partially protective immunity to T. cruzi infection and disease, suggesting their utility as vaccine candidates. For the identification of additional vaccine targets, in this study we screened the T. cruzi expressed sequence tag (EST) and genomic sequence survey (GSS) databases. By applying a variety of web-based genome-mining tools to the analysis of approximately 2,500 sequences, we identified 348 (37.6%) EST and 260 (17.4%) GSS sequences encoding novel parasite-specific proteins. Of these, 19 sequences exhibited the characteristics of secreted and/or membrane-associated GPI proteins. Eight of the selected sequences were amplified to obtain genes TcG1, TcG2, TcG3, TcG4, TcG5, TcG6, TcG7, and TcG8 (TcG1-TcG8) which are expressed in different developmental stages of the parasite and conserved in the genome of a variety of T. cruzi strains. Flow cytometry confirmed the expression of the antigens encoded by the cloned genes as surface proteins in trypomastigote and/or amastigote stages of T. cruzi. When delivered as a DNA vaccine, genes TcG1-TcG6 elicited a parasite-specific antibody response in mice. Except for TcG5, antisera to genes TcG1-TcG6 exhibited trypanolytic activity against the trypomastigote forms of T. cruzi, a property known to correlate with the immune control of T. cruzi. Taken together, our results validate the applicability of bioinformatics in genome mining, resulting in the identification of T. cruzi membrane-associated proteins that are potential vaccine candidates.

Animals↗

Accessing Underexplored Biosynthetic Potential by Initiation Unit Engineering of Nonribosomal Peptide Synthetases in Proteobacteria.

Nonribosomal peptide synthetases (NRPSs) represent a valuable yet underexplored resource for producing bioactive natural products. However, most NRPSs remain silenced potentially due to factors such as dysfunction of the initiation unit. The starter condensation (Cs) domain of the initiation unit catalyzes the lipoinitiation of nonribosomal peptides via the incorporation of an N-terminal fatty acyl chain. The concept of initiation unit engineering introduced herein encompasses the replacement of the native initiation unit of NRPSs with a foreign and well-characterized Cs domain-containing initiation unit to activate the NRPS and optimize its expression. This strategy was employed herein to successfully access three of the six previously silent NRPS pathways in Mycetohabitans rhizoxinica HKI 454, a bacterium of the class β-proteobacteria, resulting in the identification of three classes of lipopeptides. This strategy was then extended to access two NRPS pathways in Pseudomonas syringae (γ-proteobacteria) and obtain novel lipopeptides, thereby establishing a feasible complement to existing genome mining strategies for natural product discovery. Furthermore, change of the initiation regions of biosynthetic pathways of nonlipidated chitinimide (β-proteobacteria) and pseudotetraivprolide (γ-proteobacteria) with heterologous Cs-containing initiation units enabled the successful incorporation of fatty acyl chains into the N-terminus of both peptide backbones, launching a workable approach to create artificial lipopeptides. Overall, this study provides a practical strategy for the rational recovery of silent BGCs and introduction of fatty acyl chains into nonribosomal peptides, at least in Proteobacteria, thereby enriching genome mining and combinatorial biosynthesis approaches for accessing the underexplored biosynthetic potential of NRPSs from various bacteria.

Proteobacteria↗

Nerpa 2: probabilistic linking of biosynthetic gene clusters to nonribosomal peptides.

MOTIVATION: Nonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways. RESULTS: We present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC-product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry. AVAILABILITY AND IMPLEMENTATION: Nerpa 2 is freely available at https://github.com/gurevichlab/nerpa.

Multigene Family↗

Mining bacterial genomes for antimicrobial targets.

The elucidation of whole-genome sequences is expected to have a revolutionary impact on the discovery of novel medicines. With the availability of complete genome sequences of more than 30 different species, the field of antimicrobial drug discovery has the opportunity to access a remarkable diversity of genomic information. In this review, I summarize how microbial genomics has changed strategies of drug discovery by applying bioinformatics, novel genetic approaches and genomics-based technologies, including analysis of gene expression using DNA microarrays.

Anti-Bacterial Agents↗

Global protein function annotation through mining genome-scale data in yeast Saccharomyces cerevisiae.

As we are moving into the post genome-sequencing era, various high-throughput experimental techniques have been developed to characterize biological systems on the genomic scale. Discovering new biological knowledge from the high-throughput biological data is a major challenge to bioinformatics today. To address this challenge, we developed a Bayesian statistical method together with Boltzmann machine and simulated annealing for protein functional annotation in the yeast Saccharomyces cerevisiae through integrating various high-throughput biological data, including yeast two-hybrid data, protein complexes and microarray gene expression profiles. In our approach, we quantified the relationship between functional similarity and high-throughput data, and coded the relationship into 'functional linkage graph', where each node represents one protein and the weight of each edge is characterized by the Bayesian probability of function similarity between two proteins. We also integrated the evolution information and protein subcellular localization information into the prediction. Based on our method, 1802 out of 2280 unannotated proteins in yeast were assigned functions systematically.

Bayes Theorem↗

Genome-Wide Mining of lncRNAs Reveals Their Potential Regulatory Role in the Evolution of Viviparity.

Reproduction in vertebrates usually involves egg-laying (oviparity) or live-bearing (viviparity). Oviparity is the ancestral trait from which viviparity has independently evolved more than 100 times in squamate reptiles. This transition involves a series of physiological and structural changes, including the degeneration of eggshell and the evolution of a placenta and differences in the temporal and spatial expression patterns of some functional genes that drive the structural transformation. Long non-coding RNAs (lncRNAs) play important roles in the regulation of gene expression, yet it remains unclear whether they participate in gene expression shifts during the transition from oviparity to viviparity, and if so how. Therefore, we employ deep mining to identify novel lncRNAs of a closely related oviparous-viviparous pair of lizards (Phrynocephalus przewalskii and P. vlangalii). We construct cis- and trans-regulatory networks between lncRNAs and target genes using the transcriptomic data of oviduct or uteri tissues across reproductive periods. Results show that lncRNAs that regulate eggshell gland developmental genes in the oviparous lizard are lost or less expressed in the viviparous lizard. A number of lncRNAs involved in the regulation of placental development and embryo attachment in viviparous species have no orthologs in oviparous species, and others show little or no expression. Accordingly, lncRNAs may play important regulatory roles in the physiological and structural changes in the transition from oviparity to viviparity. These results open doors to the further elucidation of genetic regulatory networks.

Animals↗

Serpins in the Caenorhabditis elegans genome.

Data mining in genome sequences can identify distant homologues of known protein families, and is most powerful if solved structures are available to reveal the three-dimensional implications of very dissimilar sequences. Here we describe putative serpin sequences identified with very high statistical significance in the Caenorhabditis elegans genome. When mapped onto vertebrate serpins such as alpha1-antitrypsin, they suggest novel structural features. Some appear complete, some show extensive deletions, and others appear to contain only the C-terminal part of the known serpin fold, probably in partnership with N-terminal regions that have conformations unlike those of known serpins. The observation of such striking sequence similarity, in proteins that must have significantly different overall structures, substantially extends the structural characteristics of the serpin family of proteins.

Amino Acid Sequence↗