Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Mining genomes and mapping proteomes: identification and characterization of protein subunit vaccines.

Currently, there is an extensive and unprecedented effort to obtain the complete nucleotide sequence of the complex genomes of many micro-organisms. In this post-genomic era, based on the availability of the entire genome sequence of an organism, three new disciplines of molecular biology have emerged: genomics, transcriptional profiling and proteomics. All these technologies have the potential to accelerate the process of identifying protective protein antigens as subunit vaccine targets as well as validating and extending the range of available candidate antigens. The progress of these technologies has led to the origination of the science of bioinformatics for management and critical evaluation of the large amount of information generated. Although genomics, transcriptional profiling and proteomics are each based on different principles, there is considerable synergy between them. Appropriate application of any one, or a combination of two or more of these approaches, coupled with bioinformatics, would allow identification of a short-list of vaccine candidates from the entire list of several hundreds to thousands of proteins encoded by the genome. These candidates would then require usual channelling through the subsequent process involving recombinant expression, purification and testing for immunogenicity and protective efficacy.

Animals↗

Introduction to the special issue on advances in clinical and health-care knowledge management.

Clinical and health-care knowledge management (KM) as a discipline has attracted increasing worldwide attention in recent years. The approach encompasses a plethora of interrelated themes including aspects of clinical informatics, clinical governance, artificial intelligence, privacy and security, data mining, genomic mining, information management, and organizational behavior. This paper introduces key manuscripts which detail health-care and clinical KM cases and applications.

Confidentiality↗

DAGchainer: a tool for mining segmental genome duplications and synteny.

SUMMARY: Given the positions of protein-coding genes along genomic sequence and probability values for protein alignments between genes, DAGchainer identifies chains of gene pairs sharing conserved order between genomic regions, by identifying paths through a directed acyclic graph (DAG). These chains of collinear gene pairs can represent segmentally duplicated regions and genes within a single genome or syntenic regions between related genomes. Automated mining of the Arabidopsis genome for segmental duplications illustrates the use of DAGchainer.

Algorithms↗

Molecular complexity of sexual development and gene regulation in Plasmodium falciparum.

The malaria parasite, Plasmodium falciparum, has a complex life cycle which alternates between the vertebrate host and the invertebrate vector. Various morphological changes as well as stage-specific transcripts and gene expression profiles that accompany parasite's asexual and sexual life cycle suggest that gene regulation is crucial for the parasite's continual adaptations to survive the changing environments as well as for pathogenesis. Development of sexual stages is crucial for malaria transmission and relatively little is known about the role of specific gene products during asexual to sexual differentiation and further development. Therefore, in order to have a full understanding of the biology of the malaria parasite, gene regulation on a genome-wide global level must be understood, an area remaining to be elucidated in P. falciparum. Parasite features, such as A-T bias, difficulties in cloning, labor-intensive culture and purification of specific stages of the parasite, all contribute to the difficulties to investigate many aspects of parasite biology. However, despite these challenges, limited studies have revealed a number of parallelisms with eukaryotic transcription. For example, the parasite's genes are organised in a similar fashion, contain promoter elements and upstream activation sequences, as shown by structural searches and functional assays, and some of the basal machinery and general transcription factors have been found in Plasmodium. The completion of the full genome sequence of P. falciparum and other species of Plasmodium has resulted in the search for specific transcription factors through genome mining. Although genome mining may identify some of the factors, search for these factors solely by primary sequence homology would result in a non-comprehensive list for transcription factors present in the genome. Here, we present further discussion on putative transcription factors like activities detected in the asexual and sexual stages of P. falciparum.

Animals↗

Data mining parasite genomes.

The term 'data mining' can be used to describe any process where useful information is extracted from data with a large background of 'noise'. In the context of a genome project, several stages involve data mining. Amongst the sequence data, 'signals' need to be detected that indicate the presence of interesting features. Often this involves differentiating between transcribed and non-transcribed bases to predict coding regions. After detection, defining the roles of these sequences involves sifting through multiple lines of evidence. If these roles are accurately reflected in genome annotation, they can be used by researchers to frame queries and interrogate the data further.

Animals↗

Identification of a novel human fibroblast growth factor and characterization of its role in oncogenesis.

The fibroblast growth factor (FGF) family of signaling molecules has been implicated in normal developmental and physiological processes, as well as in human malignancy. Using a homology-based genomic DNA mining process, we identified a human gene encoding a novel member of the FGF family, that we designate FGF-20. The FGF-20 cDNA was isolated, and its sequence confirmed the gene prediction. FGF-20 is expressed in normal brain, particularly the cerebellum, and in some cancer cell lines. Recombinant FGF-20 protein induces DNA synthesis in a variety of cell types and is recognized by multiple FGF receptors. Ectopic expression of FGF-20 in NIH 3T3 cells renders the cells transformed in vitro and tumorigenic in nude mice. These results underscore the utility of mining genomic DNA databases and reveal FGF-20 to be a novel oncogene that may play a role in human cancer.

3T3 Cells↗

Mass spectrometric genomic data mining: Novel insights into bioenergetic pathways in Chlamydomonas reinhardtii.

A new high-throughput computational strategy was established that improves genomic data mining from MS experiments. The MS/MS data were analyzed by the SEQUEST search algorithm and a combination of de novo amino acid sequencing in conjunction with an error-tolerant database search tool, operating on a 256 processor computer cluster. The error-tolerant search tool, previously established as GenomicPeptideFinder (GPF), enables detection of intron-split and/or alternatively spliced peptides from MS/MS data when deduced from genomic DNA. Isolated thylakoid membranes from the eukaryotic green alga Chlamydomonas reinhardtii were separated by 1-D SDS gel electrophoresis, protein bands were excised from the gel, digested in-gel with trypsin and analyzed by coupling nano-flow LC with MS/MS. The concerted action of SEQUEST and GPF allowed identification of 2622 distinct peptides. In total 448 peptides were identified by GPF analysis alone, including 98 intron-split peptides, resulting in the identification of novel proteins, improved annotation of gene models, and evidence of alternative splicing.

Algorithms↗

Resistance Gene-Guided Discovery of a Fungal Spirotetramate as an Acetolactate Synthase Inhibitor.

Biosynthetic gene clusters (BGCs) of bioactive natural products occasionally encode resistant versions of the proteins they inhibit, offering opportunities for resistance gene-guided genome mining to uncover natural products with predictable modes of action. In this study, we developed a genome mining tool designed to identify fungal BGCs harboring putative resistance genes. Applying this tool to approximately 2500 fungal genomes, we identified a BGC designated as the pts cluster, which encodes an acetolactate synthase (ALS) homologue. Functional characterization of the pts cluster resulted in the identification of pterrespiramide A (1), featuring unique spirotetramate and cis-decalin moieties. Consistent with the predicted activity, 1 was confirmed as an ALS inhibitor and exhibited both antifungal and herbicidal activities. This study illuminates the potential of resistance gene-guided genome mining as a powerful strategy for accelerating the discovery of previously undescribed bioactive natural products.

Acetolactate Synthase↗

Genomic exploration and in silico prioritization of putative COX-2-targeting metabolites from Streptomyces sp. VITGV156 (MCC 4965).

INTRODUCTION: Streptomyces species represent an important source of bioactive natural products, yet systematic genome-guided prioritization of metabolites targeting cyclooxygenase-2 (COX-2/PTGS2) remains limited. This study aimed to investigate the biosynthetic potential of Streptomyces sp. VITGV156 (MCC 4965) using an integrated genome mining and computational drug discovery pipeline. METHODS: Whole-genome sequencing, functional annotation, antiSMASH v7.0.1-based biosynthetic gene cluster (BGC) prediction, LC-MS/MS metabolomic profiling, SwissADME analysis, target prediction, disease association mapping, molecular docking against PTGS2 (PDB: 5IKR), and PASS bioactivity prediction were performed to prioritize putative bioactive metabolites. RESULTS: Genome analysis identified 29 predicted biosynthetic gene clusters, including clusters associated with geosmin, ectoine, albaflavenone, hopene, coelichelin, and SapB, together with several cryptic clusters exhibiting low similarity to known pathways. LC-MS/MS metabolomic profiling provided experimental support for active secondary metabolite production under the cultivation conditions employed. Computational prioritization identified PTGS2 (COX-2) as a biologically relevant target. Molecular docking demonstrated favorable binding affinities and interaction profiles for several predicted metabolites within the PTGS2 catalytic pocket. PASS analysis further suggested potential anticancer-related biological activities that require experimental validation. DISCUSSION: These findings demonstrate the utility of integrating genome mining, metabolomic profiling, and computational drug discovery for prioritizing natural-product candidates. Streptomyces sp. VITGV156 (MCC 4965) represents a promising source of biosynthetic diversity and provides a genome-guided framework for identifying putative COX-2-targeting natural products for future experimental validation rather than confirming metabolite production or biological activity.

COX-2 (PTGS2)↗

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics↗

Mining nematode genome data for novel drug targets.

Expressed sequence tag projects have currently produced over 400 000 partial gene sequences from more than 30 nematode species and the full genomic sequences of selected nematodes are being determined. In addition, functional analyses in the model nematode Caenorhabditis elegans have addressed the role of almost all genes predicted by the genome sequence. This recent explosion in the amount of available nematode DNA sequences, coupled with new gene function data, provides an unprecedented opportunity to identify pre-validated drug targets through efficient mining of nematode genomic databases. This article describes the various information sources available and strategies that can expedite this process.

Animals↗

Conserved features of type III secretion.

Type III secretion systems (TTSSs) are essential mediators of the interaction of many Gram-negative bacteria with human, animal or plant hosts. Extensive sequence and functional similarities exist between components of TTSS from bacteria as diverse as animal and plant pathogens. Recent crystal structure determinations of TTSS proteins reveal extensive structural homologies and novel structural motifs and provide a basis on which protein interaction networks start to be drawn within the TTSSs, that are consistent with and help rationalize genetic and biochemical data. Such studies, along with electron microscopy, also established common architectural design and function among the TTSSs of plant and mammalian pathogens, as well as between the TTSS injectisome and the flagellum. Recent comparative genomic analysis, bioinformatic genome mining and genome-wide functional screening have revealed an unsuspected number of newly discovered effectors, especially in plant pathogens and uncovered a wider distribution of TTSS in pathogenic, symbiotic and commensal bacteria. Functional proteomics and analysis further reveals common themes in TTSS effector functions across phylogenetic host and pathogen boundaries. Based on advances in TTSS biology, new diagnostics, crop protection and drug development applications, as well as new cell biology research tools are beginning to emerge.

Amino Acid Sequence↗

Two Saccharopolyspora isolates from archaeological excavation sites: polyphasic taxonomy, biosynthetic potential, bioactivity profiling and description of Saccharopolyspora antiqui sp. nov.

Archaeological excavation sites represent underexplored microbial habitats with the potential to recover taxonomically and biotechnologically valuable actinomycetes. In this study, two Saccharopolyspora strains, 5N708T and 5N102, were isolated from soil samples collected from the Gaziantep-Doliche-Dülük and Bitlis-Ahlat-Selçuklu Cemetery archaeological excavation sites in Türkiye. A polyphasic taxonomic approach, including 16S rRNA gene sequencing, phylogenetic and phylogenomic analyses, average nucleotide identity, digital DNA-DNA hybridization, phenotypic characterization, and chemotaxonomic analyses, showed that strain 5N708T represents a novel species of the genus Saccharopolyspora, for which the name Saccharopolyspora antiqui sp. nov. is proposed, whereas strain 5N102 was assigned to Saccharopolyspora elongata. Both isolates were further evaluated for their antimicrobial, antioxidant, and cytotoxic activities, and their biosynthetic potential was investigated by genome mining. Both strains showed activity against Staphylococcus aureus, with strain 5N708T producing the larger inhibition zone. Strain 5N102 exhibited markedly stronger antioxidant activity than strain 5N708T in radical scavenging, ferric reducing antioxidant power, and reducing power assays. In contrast, strain 5N708T showed more promising cytotoxic activity, with relative selectivity toward MIA PaCa-2 pancreatic cancer cells compared with HEK293 cells after prolonged incubation. Genome mining revealed multiple biosynthetic gene clusters in both isolates, supporting their capacity to produce secondary metabolites. These findings indicate that archaeological soils are promising reservoirs of taxonomically novel and biologically active Saccharopolyspora strains.

Saccharopolyspora↗

Automating genomic data mining via a sequence-based matrix format and associative rule set.

There is an enormous amount of information encoded in each genome--enough to create living, responsive and adaptive organisms. Raw sequence data alone is not enough to understand function, mechanisms or interactions. Changes in a single base pair can lead to disease, such as sickle-cell anemia, while some large megabase deletions have no apparent phenotypic effect. Genomic features are varied in their data types and annotation of these features is spread across multiple databases. Herein, we develop a method to automate exploration of genomes by iteratively exploring sequence data for correlations and building upon them. First, to integrate and compare different annotation sources, a sequence matrix (SM) is developed to contain position-dependant information. Second, a classification tree is developed for matrix row types, specifying how each data type is to be treated with respect to other data types for analysis purposes. Third, correlative analyses are developed to analyze features of each matrix row in terms of the other rows, guided by the classification tree as to which analyses are appropriate. A prototype was developed and successful in detecting coinciding genomic features among genes, exons, repetitive elements and CpG islands.

Base Sequence↗

Genome-derived vaccines.

Vaccine research entered a new era when the complete genome of a pathogenic bacterium was published in 1995. Since then, more than 97 bacterial pathogens have been sequenced and at least 110 additional projects are now in progress. Genome sequencing has also dramatically accelerated: high-throughput facilities can draft the sequence of an entire microbe (two to four megabases) in 1 to 2 days. Vaccine developers are using microarrays, immunoinformatics, proteomics and high-throughput immunology assays to reduce the truly unmanageable volume of information available in genome databases to a manageable size. Vaccines composed by novel antigens discovered from genome mining are already in clinical trials. Within 5 years we can expect to see a novel class of vaccines composed by genome-predicted, assembled and engineered T- and Bcell epitopes. This article addresses the convergence of three forces--microbial genome sequencing, computational immunology and new vaccine technologies--that are shifting genome mining for vaccines onto the forefront of immunology research.

Animals↗

Genome data mining of lactic acid bacteria: the impact of bioinformatics.

Lactic acid bacteria (LAB) have been widely used in food fermentations and, more recently, as probiotics in health-promoting food products. Genome sequencing and functional genomics studies of a variety of LAB are now rapidly providing insights into their diversity and evolution and revealing the molecular basis for important traits such as flavor formation, sugar metabolism, stress response, adaptation and interactions. Bioinformatics plays a key role in handling, integrating and analyzing the flood of 'omics' data being generated. Reconstruction of metabolic potential using bioinformatics tools and databases, followed by targeted experimental verification and exploration of the metabolic and regulatory network properties, are the present challenges that should lead to improved exploitation of these versatile food bacteria.

Adaptation, Biological↗

High-level terpene production via a novel Actinomycetota-derived MVA pathway in E. coli.

The heterologous production of terpene in microbial hosts is often limited by inefficient and unstable pathway expression, creating a major bottleneck for industrial-scale synthesis. While E. coli as a chassis offers significant advantages, such as rapid growth, ease of cultivation, and genetic tractability. Its endogenous supply of terpenoid precursors remains a critical constraint, fundamentally restricting high-yield production. To address this challenge, we developed a genomically integrated Mevalonate (MVA) pathway from Actinomycetota in E. coli BL21(DE3) to enhance terpene precursor supply. Our approach began with an in silico multi-layer global genome mining analysis of 25,261 Actinomycetota genomes to identify a series of MVA pathway enzymes with potentially high catalytic efficiency, created a high-efficiency chassis E. coli MVA platform (ecMVA-1 and ecMVA-2) for terpene precursor synthesis. Its functionality was validated by testing eight distinct TSs. Among them, the fermentation of artemisinin precursor amorphadiene using a 5-liter bioreactor yielded 947.80 mg/L. These results indicated that E. coli (MVA) is well-suited for TS studies in the laboratory as well as holding significant promise for industrial applications. In addition, this in silico approach offers a new perspective for metabolic engineering and provides potential reservoir of diverse chassis for the industrial production of terpenoid-derived compounds.

Actinomycetota↗

Data mining parasite genomes: haystack searching with a computer.

A number of genomes of parasitic organisms are presently being sequenced in the public domain, including Plasmodium falciparum, Leishmania major and Trypanosoma brucei with the likelihood of at least expressed sequence tag (EST) projects for several filarial and apicomplexan species. The early and timely release of sequence data to the community via the World Wide Web (www), and the public databases, (EMBL and GENBANK), forms an invaluable resource. Data mining, or 'haystack searching' this resource is becoming more fruitful to all members of the scientific community as the volume of data, diversity of genomes sampled, and accessibility increase.

Animals↗