Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genome mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Accurate prediction of protein functional class from sequence in the Mycobacterium tuberculosis and Escherichia coli genomes using data mining.

The analysis of genomics data needs to become as automated as its generation. Here we present a novel data-mining approach to predicting protein functional class from sequence. This method is based on a combination of inductive logic programming clustering and rule learning. We demonstrate the effectiveness of this approach on the M. tuberculosis and E. coli genomes, and identify biologically interpretable rules which predict protein functional class from information only available from the sequence. These rules predict 65% of the ORFs with no assigned function in M. tuberculosis and 24% of those in E. coli, with an estimated accuracy of 60-80% (depending on the level of functional assignment). The rules are founded on a combination of detection of remote homology, convergent evolution and horizontal gene transfer. We identify rules that predict protein functional class even in the absence of detectable sequence or structural homology. These rules give insight into the evolutionary history of M. tuberculosis and E. coli.

Amino Acid Sequence↗

The Computational Revolution in Natural Product Research: A Data-Driven Roadmap for Next-Generation Drug Development.

Natural products (NPs) have historically provided the foundational scaffolds for drug development, yet traditional bioprospecting faces critical limitations: high rediscovery rates, laborious isolation workflows, and substantial attrition during clinical translation. The emergence of big data technologies is fundamentally transforming this landscape, enabling a shift from serendipity-based discovery toward systematic, data-driven approaches. This review examines how the integration of artificial intelligence (AI), machine learning (ML), and multi-omics datasets is accelerating natural product research across three key domains: (1) genome mining for biosynthetic gene cluster identification using platforms such as antiSMASH, (2) cheminformatics-driven prediction of structure-activity relationships and ADMET properties, and (3) metabolomics-guided dereplication to prioritize novel bioactive scaffolds. We evaluate the convergence of genomics, metabolomics, and computational chemistry in enabling in silico lead optimization and the discovery of cryptic metabolites from previously inaccessible microbial taxa. While challenges in data standardization and scalability persist, the synergy between big data and NP research is accelerating clinical translation. Despite persistent challenges in data standardization, scalability, and equitable benefit-sharing, the convergence of big data and NP research is poised to redefine drug development. These advances position computational NP research as a cornerstone of next-generation drug development.

big data analytics↗

South African Myxococcota: an untapped resource for microbial ecolo gy and biotechnology.

An extraordinary multicellular life cycle, ecological versatility, and prolific production of bioactive secondary metabolites characterise the phylum Myxococcota. While research has predominantly focused on Myxococcota in Asia, Europe, and North America, their potential occurrence in Sub-Saharan Africa remains largely unexplored. To date, only one study has isolated Myxococcota in South Africa, with additional findings limited to incidental detection through metagenomic studies. Considering South Africa's ecological diversity, its biomes may represent promising but under-examined environments for systematic bioprospecting aimed at discovering novel Myxococcota with ecological or biotechnological potential. The recent reclassification of Myxococcota from the former Deltaproteobacteria has provided a more coherent taxonomic framework to guide future ecological and systematic studies. This review presents an overview of the taxonomic revision and explores the potential occurrence of Myxococcota in South African biomes. It covers the challenges associated with conventional culture-based isolation methods and highlights potential genome- and metagenome-based approaches, including the use of metagenome-assembled genomes (MAGs) to identify cryptic biosynthetic gene clusters (BGCs), while acknowledging current limitations. Considering the increasing resistance to chemical fungicides in South African agriculture, this review further explores the potential of Myxococcota-derived secondary metabolites as candidate bioprotective alternatives. By identifying current research gaps, it aims to support future efforts towards systematic bioprospecting to investigate the ecological and biotechnological potential of Myxococcota in South Africa. KEY POINTS: • South African biomes may harbour novel Myxococcota with biosynthetic potential. • Genome mining could reveal cryptic biosynthetic gene clusters (BGCs). • Myxococcota metabolites may help control resistant fungal phytopathogens.

South Africa↗

Systematic association of genes to phenotypes by genome and literature mining.

One of the major challenges of functional genomics is to unravel the connection between genotype and phenotype. So far no global analysis has attempted to explore those connections in the light of the large phenotypic variability seen in nature. Here, we use an unsupervised, systematic approach for associating genes and phenotypic characteristics that combines literature mining with comparative genome analysis. We first mine the MEDLINE literature database for terms that reflect phenotypic similarities of species. Subsequently we predict the likely genomic determinants: genes specifically present in the respective genomes. In a global analysis involving 92 prokaryotic genomes we retrieve 323 clusters containing a total of 2,700 significant gene-phenotype associations. Some clusters contain mostly known relationships, such as genes involved in motility or plant degradation, often with additional hypothetical proteins associated with those phenotypes. Other clusters comprise unexpected associations; for example, a group of terms related to food and spoilage is linked to genes predicted to be involved in bacterial food poisoning. Among the clusters, we observe an enrichment of pathogenicity-related associations, suggesting that the approach reveals many novel genes likely to play a role in infectious diseases.

Bacteria↗

Molecular biology and integrated strategies for activating cryptic biosynthetic gene clusters toward next-generation antibiotic discovery.

Antimicrobial resistance (AMR) has been identified as one of the 21st century's severest global public health crises. AMR led to an estimated 4.95 million deaths in 2019 and will claim 10 million lives a year by 2050 in the absence of targeted interventions. During the same period, the number of novel antibiotics discovered has decreased drastically as many researchers are rediscovering known antibiotics, non-model microorganisms are poorly understood or difficult to culture and antibiotic research and development investment has declined drastically. However, high-throughput whole genome sequencing and the subsequent application of bioinformatics in bacterial and fungal genomes have shown that a numerous of cryptic or silent biosynthetic gene clusters (BGCs) remain latent at ambient laboratory conditions since their genes are transcriptionally inactive. Cryptic BGCs represent a vast source of unique secondary metabolites, many of which may yield novel antibacterial, antifungal, anti-cancer and other potentially valuable natural products. This review discusses the biological relevance of cryptic BGCs, the major limiting factors that restricts their activation and novel strategies that have been employed to activate them and exploit their potential to produce novel natural products. The review focuses on biological approaches including CRISPR-Cas mediation for the activation of cryptic BGCs, promoter engineering, pathway refactoring, and heterologous expression; biochemical strategies such as Osman, OsMAC, Precursor Feeding, Chemical Elicitation, Epigenetic Regulation and Co-cultivation and technology-based strategies such as Genome mining, Microfluidic Cultivation systems, High-Throughput Screening, Metabolomics, Molecular Networking and Artificial Intelligence and Machine Learning based prediction of BGCs and their metabolites. The use of multi-omics technologies combined with synthetic biology to achieve better discovery, characterization and large-scale production of novel natural products is also discussed herein. Finally, we will talk about the ecological significance and evolutionary advantage of cryptic BGCs' role in interactions between microorganisms, such as competition, communication, symbiosis and environmental adaptability, so as to provide a useful background for accelerating next-generation antibiotics.

CRISPR-Cas activation↗

Unveiling Aziridine-Containing Natural Products by Genomic and Spectroscopic Approaches.

Aziridine-containing natural products are prized for their potent bioactivities, yet their scarcity and poorly understood biosynthesis have limited systematic exploration. Here, we address this by integrating genome mining with a 1H-13C coupled HSQC metabolomic approach that exploits the distinctive NMR signatures of aziridines, enabling their direct detection from complex extracts. This strategy unveiled the desertolides, the first macrolides incorporating a rare terminal 2-methyl-aziridine-2-carboxylate moiety. Genetic and isotopic studies identified a dedicated biosynthetic subcluster (desA-desN) that assembles and installs this unit from glutamate, and heterologous expression confirmed the self-sufficiency of this subcluster. Direct MS evidence reveals the aziridine moiety covalently bound to the active-site Cys113 of DesN, establishing this KAS III homolog as the first dedicated aziridine-transferase and a promising tool for polyketide engineering. Bioinformatic analysis uncovered over 50 biosynthetic gene clusters, suggesting that this aziridine-associated biosynthetic logic may be more widespread than currently appreciated. This work establishes a tractable platform for the targeted discovery and engineered biosynthesis of aziridine-containing natural products, opening this underexplored pharmacophore to systematic interrogation.

Aziridines↗

Application of genomics and proteomics for identification of bacterial gene products as potential vaccine candidates.

The ability of bioinformatics to characterize genomic sequences from pathogenic bacteria for prediction of genes that may encode vaccine candidates, e.g. surface localized proteins, has been evaluated. By applying appropriate tools for genomic mining to the published sequence of Haemophilus influenzae Rd genome, it was possible to identify a putative vaccine candidate, the outer membrane lipoprotein, P6. Proteomics complements genomics by offering abilities to rapidly identify the products of predicted genes, e.g. proteins in outer membrane preparations. The ability to identify the P6 protein uniquely from entries in a sequence database from the expected peptide-mass fingerprint of P6 demonstrates the power of proteomics. The application of proteomics for identification of vaccine candidates for another pathogenic bacterium, Helicobacter pylori using two different approaches is described. The first involves rapid identification of a series of monoclonal antibody reactive proteins from N-terminal sequence tags. The other approach involves identification of proteins in outer membrane preparations by 2-D electrophoresis followed by trypsin digestion and peptide mass map analysis. Our combined studies demonstrate that utilization of genome sequences by application of bioinformatics through genomics and proteomics can expedite the vaccine discovery process by rapidly providing a set of potential candidates for further testing.

Antibodies, Bacterial↗

Activity, structure, and diversity of Type II proline-rich antimicrobial peptides from insects.

Apidaecin 1b (Api), the first characterized Type II Proline-rich antimicrobial peptide (PrAMP), is encoded in the honey bee genome. It inhibits bacterial growth by binding in the nascent peptide exit tunnel of the ribosome after the release of the completed protein and trapping the release factors. By genome mining, we have identified 71 PrAMPs encoded in insect genomes as pre-pro-polyproteins. Having chemically synthesized and tested the activity of 26 peptides, we demonstrate that despite significant sequence variation in the N-terminal sequence, the majority of the PrAMPs that retain the conserved C-terminal sequence of Api are able to trap the ribosome at the stop codons and induce stop codon readthrough-all hallmarks of Type II PrAMP mode of action. Some of the characterized PrAMPs exhibit superior antibacterial activity in comparison with Api. The newly solved crystallographic structures of the ribosome complexed with Api and with the more active peptide Fva1 from the stingless bee demonstrate the universal placement of the PrAMPs' C-terminal pharmacophore in the post-release ribosome despite variations in their N-terminal sequence.

Animals↗

Comparative genomics using data mining tools.

We have analysed the genomes of representatives of three kingdoms of life, namely, archaea, eubacteria and eukaryota using data mining tools based on compositional analyses of the protein sequences. The representatives chosen in this analysis were Methanococcus jannaschii, Haemophilus influenzae and Saccharomyces cerevisiae. We have identified the common and different features between the three genomes in the protein evolution patterns. M. jannaschii has been seen to have a greater number of proteins with more charged amino acids whereas S. cerevisiae has been observed to have a greater number of hydrophilic proteins. Despite the differences in intrinsic compositional characteristics between the proteins from the different genomes we have also identified certain common characteristics. We have carried out exploratory Principal Component Analysis of the multivariate data on the proteins of each organism in an effort to classify the proteins into clusters. Interestingly, we found that most of the proteins in each organism cluster closely together, but there are a few 'outliers'. We focus on the outliers for the functional investigations, which may aid in revealing any unique features of the biology of the respective organisms

Archaeal Proteins↗

Genomes of 211 Actinomycete Strains from Diverse Environments.

Actinomycetes are a highly diverse group of microorganisms that have long been recognized as a valuable source of antibiotics and other bioactive metabolites. Recent advances in genome mining have revealed a wealth of previously unexplored silent secondary metabolite biosynthetic gene clusters (smBGCs) in actinomycete genomes, underscoring their untapped bioactivity potential. Here, we present the genome sequences of 211 actinomycete strains isolated from various environmental sources, generated through high-throughput sequencing. The resulting genome assemblies exhibit high completeness and accuracy, offering high-quality data for downstream analyses and biological resource exploration.

Actinobacteria↗

Deciphering gene expression regulatory networks.

In the past year, great strides have been made in our understanding of the regulatory networks that control gene expression in the model eukaryote Saccharomyces cerevisiae. The development and use of a number of genomic tools, including genome-wide location and expression analysis, has fueled this progress. In addition, a variety of computational algorithms have been devised to mine genomic sequence for conserved regulatory motifs in co-regulated genes. The recent description of the genetic network controlling the cell cycle illustrates the tremendous potential of these approaches for deciphering gene expression regulatory networks in eukaryotic cells.

Algorithms↗

Neobacillus driksii sp. nov. isolated from a Mars 2020 spacecraft assembly facility and genomic potential for lasso peptide production in Neobacillus.

UNLABELLED: During microbial surveillance of the Mars 2020 spacecraft assembly facility, two novel bacterial strains, potentially capable of producing lasso peptides, were identified. Characterization using a polyphasic taxonomic approach, whole-genome sequencing and phylogenomic analyses revealed a close genetic relationship among two strains from Mars 2020 cleanroom floors (179-C4-2-HS, 179-J1A1-HS), one strain from the Agave plant (AT2.8), and another strain from wheat-associated soil (V4I25). All four strains exhibited high 16S rRNA gene sequence similarity (>99.2%) and low average nucleotide identity (ANI) with Neobacillus niacini NBRC 15566T, delineating new phylogenetic branches within the genus. Detailed molecular analyses, including gyrB (90.2%), ANI (86.4%), average amino acid identity (87.8%) phylogenies, digital DNA-DNA hybridization (32.6%), and percentage of conserved proteins (77.7%) indicated significant divergence from N. niacini NBRC 15566T. Consequently, these strains have been designated Neobacillus driksii sp. nov., with the type strain 179-C4-2-HST (DSM 115941T = NRRL B-65665T). N. driksii grew at 4°C to 45°C, pH range of 6.0 to 9.5, and 0.5% to 5% NaCl. The major cellular fatty acids are iso-C15:0 and anteiso-C15:0. The dominant polar lipids include diphosphatidylglycerol, phosphatidylglycerol, phosphatidylethanolamine, and an unidentified aminolipid. Metagenomic analysis within NASA cleanrooms revealed that N. driksii is scarce (17 out of 236 samples). Genes encoding the biosynthesis pathway for lasso peptides were identified in all N. driksii strains and are not commonly found in other Neobacillus species, except in 7 out of 26 recognized species. This study highlights the unique metabolic capabilities of N. driksii, underscoring their potential in antimicrobial research and biotechnology. IMPORTANCE: The microbial surveillance of the Mars 2020 assembly cleanroom led to the isolation of novel N. driksii with potential applications in cleanroom environments, such as hospitals, pharmaceuticals, semiconductors, and aeronautical industries. N. driksii genomes were found to possess genes responsible for producing lasso peptides, which are crucial for antimicrobial defense, communication, and enzyme inhibition. Isolation of N. driksii from cleanrooms, Agave plants, and dryland wheat soils, suggested niche-specific ecology and resilience under various environmentally challenging conditions. The discovery of potent antimicrobial agents from novel N. driksii underscores the importance of genome mining and the isolation of rare microorganisms. Bioactive gene clusters potentially producing nicotianamine-like siderophores were found in N. driksii genomes. These siderophores can be used for bioremediation to remove heavy metals from contaminated environments, promote plant growth by aiding iron uptake in agriculture, and treat iron overload conditions in medical applications.

Phylogeny↗

Secondary metabolite profiling of rare Micromonospora spp. from cold desert of NW Himalayas via multi-omics analysis.

INTRODUCTION: The genus Micromonospora is a prolific producer of specialized metabolites with pharmacological and agronomic relevance. Natural products derived from the genus Micromonospora have a distinctive chemical diversity and enormous therapeutic potential, thus represent a potential source for drugs and drug leads. OBJECTIVE: To explore the biosynthetic potential of four Micromonospora strains isolated from cold desert of NW Himalayas through genome mining and to correlate predicted biosynthetic gene clusters with chemical features detected by untargeted LC-HRMS metabolomics. METHOD: High-quality genomes were annotated for BGCs and matched against untargeted LC-HRMS features (peak picking, alignment, and annotation to chemical classes). Each isolate was grown in triplicate, and fermented broth was pooled for further metabolomic studies. RESULTS: By integrating genomic and metabolomic approaches, specialized biosynthetic gene clusters and strain-based putative metabolite classes were identified. LRS1 showed elevated xanthines (RiPP/siderophore), LRS3 had phenolic glycosides (hybrid PKS/NRPS), LRS4 showed 70-fold hydroxycinnamate enrichment (Type II PKS), and LRS5 displayed p-benzoquinone enrichment (Type III PKS). The metabolite profile of each strain aligned with its predicted biosynthetic gene cluster composition. CONCLUSION: Under a single growth regime, each Micromonospora strain exhibits a distinct metabolomic profile. This metabologenomics workflow can be further explored to isolate specialized metabolites with potential therapeutic and agricultural value.

Micromonospora↗

The hidden code in genomics: a tool for gene discovery.

Among new insights coming from the completion of sequencing of the human genome, reported in Nature and Science, are clues of how evolution has increased the complexity of species, and in particular how the genetic code has enabled this process. It is clear that life has not only evolved by increasing the number of genes, but also by ingeniously evolving an efficient code for expressing diversity in the building blocks (i.e. the amino acids). The rules of nucleic acid base pairing and the classification of amino acids according to hydrophobicity/hydrophilicity relationships define a binary DNA code, which determines the general biophysical characteristics of proteins. Sense and antisense strands can encode protein segments having inverted and complementary hydropathy. The underlying binary code controls association and dissociation of proteins and presumably represents a primordial code that might have emerged in the early stages of self-organizing biochemical cycles. It is the purpose of this communication to provide a perspective of the code in the context of a binary language from its primordial origin to its present day format and to propose to use this code as a genomic mining tool.

Expressed Sequence Tags↗

Genome sequence data of the chitinase-producing bacterium Paenibacillus mucilaginosus YWY-5.1.

Paenibacillus mucilaginosus is a beneficial bacterium widely applied as a biofertilizer in agriculture. To date, genomic information on this species remains limited; however, no genome assemblies from Vietnam have been reported. This work presented the draft genome of P. mucilaginosus YWY-5.1, a promising strain with strong chitin-degrading capability and agricultural potential, isolated from Yok Don National Park, Vietnam, using Illumina technology. Results showed that the assembled genome comprised 48 contigs with 4,076,146 bp and 73.8% GC-content. Genome annotation identified 3,611 protein-coding genes, 2 rRNA genes, and 53 tRNA genes. A total of 150 carbohydrate-active enzyme-related genes were predicted from the genome; among them, seven putative chitinolytic genes were identified, including 4 genes related to family 18 chitinase, 2 genes to family 20 β-N-acetylglucosaminidase, and one gene to auxiliary activity family 10. In addition, at least 32 genes related to plant growth-promoting functions were identified, including those associated with indole-3-acetic acid production, phosphate and potassium solubilization, siderophore biosynthesis, iron uptake, ACC metabolism, and nitrate transport and reduction. Furthermore, genome mining identified 4 biosynthetic gene clusters probably involved in secondary metabolite production, of which 3 displayed no similarity to previously reported clusters, indicating potential for novel bioactive compounds. These genomic data improved our understanding of the biodegradation capacity and agricultural potential of P. mucilaginosus YWY-5.1 isolated from Vietnam, and provided a valuable genomic resource for future functional and biotechnological investigations toward crop production and related fields.

Chitinases↗

Exploration of antimicrobial potential in LAB by genomics.

A tremendous flow of information has been created through various genome sequencing projects worldwide. So far, 128 bacterial genome sequences have been completed and 391 are under way. Many of these bacteria, including several lactic acid bacteria (LAB), are used in the production and preservation of food and feed. The major antimicrobial and biopreservative substance produced by LAB is organic acid; however, some LAB produce additional antimicrobial compounds. Among these, the bacteriocins have demonstrated great potential as food preservatives. Additionally, antimicrobial compounds different from the bacteriocins have recently been identified, of which several display strong antifungal activity. The information obtained from genomics and related technologies will have great impact on the future identification and development of new antimicrobial agents. Developments will include the identification of pathways for the production of antimicrobials and genome mining for new antimicrobial peptides.

Amino Acid Sequence↗

Structure of trichamide, a cyclic peptide from the bloom-forming cyanobacterium Trichodesmium erythraeum, predicted from the genome sequence.

A gene cluster for the biosynthesis of a new small cyclic peptide, dubbed trichamide, was discovered in the genome of the global, bloom-forming marine cyanobacterium Trichodesmium erythraeum ISM101 because of striking similarities to the previously characterized patellamide biosynthesis cluster. The tri cluster consists of a precursor peptide gene containing the amino acid sequence for mature trichamide, a putative heterocyclization gene, an oxidase, two proteases, and hypothetical genes. Based upon detailed sequence analysis, a structure was predicted for trichamide and confirmed by Fourier transform mass spectrometry. Trichamide consists of 11 amino acids, including two cysteine-derived thiazole groups, and is cyclized by an N C terminal amide bond. As the first natural product reported from T. erythraeum, trichamide shows the power of genome mining in the prediction and discovery of new natural products.

Amino Acid Sequence↗

Programmable enzymes for targeted gene insertion.

Genome editing technologies have advanced from nuclease-based reagents that generate programmed DNA double-strand breaks, which can cause deleterious effects, to next-generation reagents that perform controlled DNA modification through double-strand break-independent mechanisms, such as base editing and prime editing. Although these approaches enable precise small-scale sequence changes, methods for programmable insertion of large DNA cargos have been limited. The ability to write entire genes or large regions into the genome could transform the treatment of genetically heterogeneous disorders, for which numerous pathogenic variants underlie a common disease and mutation-specific editing strategies are impractical. Recent advances in computational genome mining have accelerated the discovery of naturally occurring enzymes with novel biochemical and functional properties, including recombinases and transposases capable of large-scale modifications. Moreover, directed evolution, rational engineering and expanded homologue discovery are enabling the repurposing and optimization of these systems for genome engineering. Here we review recent technology development efforts that harness diverse enzymes for kilobase-scale genome engineering, with a particular focus on CRISPR-associated transposase systems.

Journal Article↗