Search PubMedSearch

SEARCH · Search PubMed

Results for “Biosynthetic gene clusters”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multichassis Expression of Cyanobacterial and Other Bacterial Biosynthetic Gene Clusters.

Heterologous expression of biosynthetic gene clusters (BGCs) is a powerful strategy for natural product (NP) discovery, yet achieving consistent expression across microbial hosts remains challenging. Here, we developed cross-phyla vector systems enabling the expression of BGCs from cyanobacteria and other bacterial origins in Gram-negative Escherichia coli, Gram-positive Bacillus subtilis, and two model cyanobacterial strains including unicellular Synechocystis PCC 6803 and filamentous Anabaena sp. PCC 7120. Following validation using constitutive and inducible expression of the enhanced yellow fluorescent protein (eYFP), we applied these vectors to express the shinorine and violacein BGCs in all four hosts. Promoter tuning, substrate feeding, BGC refactoring, and inducible control enhanced NP production and mitigated host toxicity. Notably, we demonstrated that B. subtilis can serve as a chassis for cyanobacterial NP BGC expression. Our results provide versatile expression platforms for probing BGC function and accelerating natural product discovery from diverse cyanobacterial and other bacterial lineages.

Multigene Family

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family

Comparative Genomics of Paenibacillus Secondary Metabolism: Unveiling the Putative Biosynthetic Gene Cluster for Paenialvins in Paenibacillus Alvei Strain 32.

In this study, we used comparative genomics and culture-based methods to investigate Biosynthetic Gene Clusters (BGCs) responsible for the production of antimicrobial peptides. Paenibacillus alvei strain 32 was isolated from a cystic fibrosis sputum. Its genome was sequenced using Illumina, showing a size of 6,584,590 bp with 239 contigs assembled in 26 scaffolds, an average coverage of 243X, and 6,832 coding sequences. ANI analysis and in silico DNA-DNA hybridization showed its affiliation inside Paenibacillus alvei, with a clear separation from other related strains, leading us to propose a distinct species-level genomic clade (genomospecies) within this group. AntiSMASH analysis predicted 22 putative BGCs in the genome of strain 32. Its culture supernatant exhibited inhibitory activity against Gram-positive pathogens, including methicillin-resistant Staphylococcus aureus (MRSA), Bacillus cereus, and Enterococcus faecalis. By comparing in silico BGC predictions with activities described in the literature, we propose that strain 32 harbours a specific 110-kb cluster (cluster 6.2) with five non-ribosomal peptide synthetase (NRPS) genes. These synthetases are predicted to direct the assembly of a 16-amino acid backbone that correlates with the structure of paenialvins, which are known anti-MRSA molecules. This study describes the putative biosynthetic pathway of the paenialvins and explains structural variations, bringing useful data on Paenibacillus secondary metabolism for future antibiotic development.

Paenibacillus alvei

Nerpa 2: probabilistic linking of biosynthetic gene clusters to nonribosomal peptides.

MOTIVATION: Nonribosomal peptides (NRPs) are bioactive microbial metabolites with high pharmaceutical potential. Although genome mining enables large-scale detection of biosynthetic gene clusters (BGCs) predicted to encode NRPs, reliably linking these clusters to their chemical products remains challenging due to the flexible and heterogeneous organization of NRP assembly pathways. RESULTS: We present Nerpa 2, a probabilistic framework for accurate and scalable linking of NRP BGCs to candidate chemical structures. The method represents assembly lines as hidden Markov models (HMMs) that capture uncertainty and alternative biosynthetic routes. On curated datasets of experimentally validated BGC-product pairs, our tool outperforms existing methods in linking accuracy and pathway reconstruction. When applied to large genome mining datasets, Nerpa 2 efficiently identifies BGCs likely associated with known compounds and highlights potential producers of novel chemistry. AVAILABILITY AND IMPLEMENTATION: Nerpa 2 is freely available at https://github.com/gurevichlab/nerpa.

Multigene Family

Discovery of antimicrobial peptides from incomplete biosynthetic gene clusters to combat multidrug-resistant bacteria.

The escalating crisis of multidrug-resistant bacteria necessitates innovative antibiotic discovery platforms. Conventional antimicrobial peptide (AMP) mining often relies on complete biosynthetic gene clusters (BGCs), leaving fragmented genomic resources underexplored. Here, we present an evolution-inspired approach to reconstruct and predict AMPs from partial BGCs. Applying this strategy to 954 Paenibacillus genomes identifies five polymyxin-like peptides, NP001-NP005, with broad in vitro activity. Crucially, in murine models of polymyxin-resistant infection, NP001 reduced bacterial burdens by up to 1,000-fold in a thigh infection model and improved survival (50% vs. 0%) in a lethal peritonitis model. Structural simulations and biophysical assays revealed that NP001 maintains high affinity for bacterial membranes and effectively binds to MCR-1-modified lipid A, a key colistin-resistance mechanism. Moreover, Leu at position 10 of NP001 plays a key role in antibacterial activity against MCR-1-resistant bacteria. Our work establishes a generalizable framework for AMP discovery and introduces a promising therapeutic candidate, NP001, which effectively counteracts polymyxin-resistant pathogens.

Multigene Family

Genome mining of alkaliphilic cyanobacterial consortia: identification of biosynthetic gene clusters in Sodalinema and associated heterotrophs.

Alkaline soda lakes are high-pH environments that host specialized microbial communities with potential for biotechnology and natural product discovery. We characterized three Sodalinema-dominated cyanobacterial consortia enriched from Canadian soda lakes over 510 days. Using hybrid metagenomic sequencing and metatranscriptomics across pH, alkalinity, and temperature gradients, we reconstructed high-quality metagenome-assembled genomes and assessed functional activity. All consortia converged toward cyanobacteria dominance and exhibited temperature optima between 21°C and 30°C. Phylogenetic analysis placed Sodalinema genomes within a distinct clade affiliated with Candidatus Sodalinema alkaliphilum. Genomic analysis indicated complete biosynthetic pathways for vitamin B5, vitamin B7, and the molybdenum cofactor, but incomplete pathways for vitamins B1, B9, and B12, consistent with patterns observed in Sodalinema yuhuli. Metatranscriptomic profiles showed increased expression of genes involved in phycocyanin and carotenoid biosynthesis at pH 10.2 relative to pH 8.5. Biosynthetic gene cluster analysis revealed that most secondary metabolic potential resided in heterotrophic community members. Roseinatronobacter encoded pathways for N-acyl homoserine lactones, osmoprotectants, betalactones, and prodigiosin, while Alkalimonas, Wenzhouxiangella, and members of the Kiloniellales encoded clusters for lanthipeptides, cyclodipeptides, hydrogen cyanide, and pyrroloquinoline quinone. These findings indicate functional partitioning within the consortia and highlight the contribution of heterotrophs to secondary metabolism.IMPORTANCEAlkaline soda lakes contain microbial communities adapted to high pH that remain underexplored for biotechnology. This study focuses on Sodalinema, a filamentous cyanobacterium that dominates enriched consortia from Canadian soda lakes, and its associated heterotrophic partners. We show that while Sodalinema drives primary productivity, heterotrophic bacteria encode most of the pathways for antimicrobial and signaling compounds. These interactions may support community stability and defense against competing microorganisms. By linking genomic potential with gene expression, this work identifies alkaline cyanobacterial consortia as a source of bioactive compounds and provides a framework for exploring extremophilic microbial communities for natural product discovery.

Sodalinema

Molecular biology and integrated strategies for activating cryptic biosynthetic gene clusters toward next-generation antibiotic discovery.

Antimicrobial resistance (AMR) has been identified as one of the 21st century's severest global public health crises. AMR led to an estimated 4.95 million deaths in 2019 and will claim 10 million lives a year by 2050 in the absence of targeted interventions. During the same period, the number of novel antibiotics discovered has decreased drastically as many researchers are rediscovering known antibiotics, non-model microorganisms are poorly understood or difficult to culture and antibiotic research and development investment has declined drastically. However, high-throughput whole genome sequencing and the subsequent application of bioinformatics in bacterial and fungal genomes have shown that a numerous of cryptic or silent biosynthetic gene clusters (BGCs) remain latent at ambient laboratory conditions since their genes are transcriptionally inactive. Cryptic BGCs represent a vast source of unique secondary metabolites, many of which may yield novel antibacterial, antifungal, anti-cancer and other potentially valuable natural products. This review discusses the biological relevance of cryptic BGCs, the major limiting factors that restricts their activation and novel strategies that have been employed to activate them and exploit their potential to produce novel natural products. The review focuses on biological approaches including CRISPR-Cas mediation for the activation of cryptic BGCs, promoter engineering, pathway refactoring, and heterologous expression; biochemical strategies such as Osman, OsMAC, Precursor Feeding, Chemical Elicitation, Epigenetic Regulation and Co-cultivation and technology-based strategies such as Genome mining, Microfluidic Cultivation systems, High-Throughput Screening, Metabolomics, Molecular Networking and Artificial Intelligence and Machine Learning based prediction of BGCs and their metabolites. The use of multi-omics technologies combined with synthetic biology to achieve better discovery, characterization and large-scale production of novel natural products is also discussed herein. Finally, we will talk about the ecological significance and evolutionary advantage of cryptic BGCs' role in interactions between microorganisms, such as competition, communication, symbiosis and environmental adaptability, so as to provide a useful background for accelerating next-generation antibiotics.

CRISPR-Cas activation

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500 m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed > 99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071ᵀ (= ATCC 10145ᵀ), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33 Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8 kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~ 22 kb, ~ 17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family

Genomic Insights Into Multidrug-Resistant Foodborne Serratia liquefaciens Strains Carrying mcr-9 and Comparative Genomic Analysis of Novel Biosynthetic Gene Clusters.

Serratia liquefaciens is an opportunistic nosocomial pathogen with a wide range of antibiotic resistance patterns. This study reports the characterization of the first mcr-9-positive S. liquefaciens strains, 35E-19E1 and CST-066, isolated from meat products in Japan. The strains were screened for the presence of β-lactamases, plasmid-mediated mobile colistin resistance (mcr) genes, and carbapenemase-encoding genes using PCR. Antimicrobial susceptibility was tested using the broth microdilution method. The strains exhibited multidrug resistance (MDR) phenotypes to third-generation cephalosporins, cephamycin, fosfomycin, and other clinically important antimicrobials. Genomic DNA sequencing showed that the genome sizes of CST-066 and 35E-19E1 are 5,529,704 and 5,261,506 bps, respectively. mcr-9 was identified on a chromosome within a genetic environment that included the two-component system qseBC, which plays a key role in the signaling network that triggers colistin resistance in Enterobacterales. Downstream genome analysis revealed a 1695-bp eptB-like kdo2-lipid phosphoethanolamine transferase, which is involved in intrinsic polymyxin resistance mechanisms in Serratia spp. The strain 35E-19E1 carries five CRISPR-Cas enzymes that are essential for adaptive immunity in bacteria, allowing defense against invading elements. Functional analysis using subsystem technology revealed that both strains possess subsystem features responsible for invasion and adhesion within the host biomes. Genome mining using antiSMASH and BAGL4 revealed various biosynthetic gene clusters, responsible for secondary metabolite synthesis. Notably, we identified novel gene clusters, mainly nonribosomal peptide synthetases, in both the strains, indicating their potential to produce bioactive compounds. Although the presence of mcr-9 in Serratia may not be of clinical significance because of natural resistance of the strain to polymyxins, we shed light on the genomic characteristics of this MDR pathogen and the potential spread of mcr-9 among other bacterial species. The emergence of mcr-9 in drug-resistant S. liquefaciens provides significant insights, underscoring the need for increased surveillance of this pathogen.

biosynthetic gene cluster

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family

Comparative metagenomic assessment of Illumina-compatible library preparation methods, short-read lengths, and PacBio HiFi sequencing reveals differences in microbial and functional diversity recovery from a complex environmental sample.

UNLABELLED: Metagenomics enables comprehensive exploration of microbial communities but is influenced by library preparation and sequencing technologies, affecting recovery of microbial genomes and proteins. Here, we benchmarked six Illumina-compatible short-read library preparation conditions in triplicate at 2 × 150 bp and 2 × 250 bp read lengths alongside PacBio HiFi long-read sequencing using a composite environmental sample of marine mangrove sediment and terrestrial palm tree soil. Longer short reads (2 × 250 bp) combined with optimal library preparation approaches improved assembly quality, protein detection, and metagenome-assembled genome (MAG) recovery, achieving results approaching those of long-read sequencing. TruSeq libraries at 2 × 250 bp recovered more than sevenfold more unique proteins than the same kit at 2 × 150 bp (811,701 vs 110,108) using the same number of sequencing reads, while recovering a comparable number of high-quality MAGs to PacBio HiFi long-read sequencing (11 vs 18) and surpassing it in protein discovery by almost 10-fold (811,701 vs 87,745) at less than half of the sequencing cost. Furthermore, biosynthetic gene cluster analysis identified 46 biosynthetic gene clusters in TruSeq-250PE assemblies compared to 38 in PacBio HiFi, with several showing no close match in the MIBiG database. Although long reads yield more contiguity and complete genomes, longer short reads offer a cost-effective, scalable alternative for uncovering microbial and functional diversity. These findings provide critical guidance for metagenomic experimental design, demonstrating that strategic selection of library preparation chemistry and sequencing parameters can reveal more unknown microbial information in complex biomes without requiring additional sequencing depth. IMPORTANCE: Metagenomic outcomes are strongly influenced by library preparation and sequencing strategies, yet their combined effects in complex environmental samples remain poorly defined. Here, we provide the first direct comparison of Illumina NovaSeq short-read metagenomic sequencing at 2 × 150 bp and 2 × 250 bp across multiple library preparation kits, alongside PacBio HiFi long-read sequencing. We show that sequencing read length and library preparation critically shape assembly quality, protein recovery, and metagenome-assembled genome (MAG) reconstruction. These findings demonstrate that short-read sequencing at 2 × 250 bp, with appropriate library preparation, can match long-read technologies in MAG recovery while substantially surpassing them in protein discovery. With less than half of the sequencing price and a 3.5-fold reduction in cost per gigabase of usable data, this method facilitates more accessible large-scale metagenomic analysis within complex environmental systems.

Metagenomics

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software

Biosynthetic potential of the culturable foliar fungi associated with field-grown lettuce.

Fungal endophytes and epiphytes associated with plant leaves can play important ecological roles through the production of specialized metabolites encoded by biosynthetic gene clusters (BGCs). However, their functional capacity, especially in crops like lettuce (Lactuca sativa L.), remains poorly understood. We sequenced the genomes of nine fungal isolates, representing Fusarium sp., Fulvia sp., Alternaria alternata, and Alternaria postmessia, from leaves of lettuce grown under field conditions in Arizona, USA. We used antibiotics and secondary metabolite analysis shell (antiSMASH) and the database for automated carbohydrate-active enzyme annotation (dbCAN3), to predict BGCs and carbohydrate-active enzymes (CAZymes) for each strain, and then compared them to conspecific strains from other environments and substrates. Foliar lettuce-associated fungi featured 39-95 BGCs per genome, with substantial overlap between isolates occurring in association with lettuce leaves vs. from other substrates. Species identity was a significant determinant of BGC count, while host type, isolation source, and lifestyle were not. Several BGCs, including those for alternariol and 1,3,6,8-Tetrahydroxynaphthalene (T4HN), showed 100% similarity to characterized minimum information about a biosynthetic gene cluster (MIBiG) clusters based on antiSMASH predictions. Although analysis by biosynthetic gene similarity clustering and prospecting engine (BiG-SCAPE) identified gene cluster families (GCFs) across the dataset, these reference-matching clusters were not always grouped, reflecting methodological differences in how the tools assess similarity. Comparative CAZyme analysis in a focal species (Fulvia sp.) revealed higher gene counts in a foliar lettuce-derived isolate than in tomato (Solanum lycopersicum)-associated strains, challenging assumptions about host chemical complexity. These results highlight the importance of phylogenetic context in shaping fungal functional potential and suggest that selection on microbial traits in edible leafy crops may be more subtle and species-specific than previously assumed. KEY POINTS: • Lettuce-associated fungi feature diverse biosynthetic potential • Phylogeny predicts fungal BGC content more strongly than ecological lifestyle • Findings support genome-informed microbiome strategies for leafy crops.

Lactuca

Secondary metabolite profiling of rare Micromonospora spp. from cold desert of NW Himalayas via multi-omics analysis.

INTRODUCTION: The genus Micromonospora is a prolific producer of specialized metabolites with pharmacological and agronomic relevance. Natural products derived from the genus Micromonospora have a distinctive chemical diversity and enormous therapeutic potential, thus represent a potential source for drugs and drug leads. OBJECTIVE: To explore the biosynthetic potential of four Micromonospora strains isolated from cold desert of NW Himalayas through genome mining and to correlate predicted biosynthetic gene clusters with chemical features detected by untargeted LC-HRMS metabolomics. METHOD: High-quality genomes were annotated for BGCs and matched against untargeted LC-HRMS features (peak picking, alignment, and annotation to chemical classes). Each isolate was grown in triplicate, and fermented broth was pooled for further metabolomic studies. RESULTS: By integrating genomic and metabolomic approaches, specialized biosynthetic gene clusters and strain-based putative metabolite classes were identified. LRS1 showed elevated xanthines (RiPP/siderophore), LRS3 had phenolic glycosides (hybrid PKS/NRPS), LRS4 showed 70-fold hydroxycinnamate enrichment (Type II PKS), and LRS5 displayed p-benzoquinone enrichment (Type III PKS). The metabolite profile of each strain aligned with its predicted biosynthetic gene cluster composition. CONCLUSION: Under a single growth regime, each Micromonospora strain exhibits a distinct metabolomic profile. This metabologenomics workflow can be further explored to isolate specialized metabolites with potential therapeutic and agricultural value.

Micromonospora

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics

Microbe Profile: Streptomyces formicae KY5: an ANT-ibiotic factory.

Streptomyces formicae KY5 was isolated from a Tetraponera penzigi plant-ant nest. It is primarily known for its production of the formicamycins, antibiotics with potent activity against Gram-positive pathogens including methicillin-resistant Staphylococcus aureus, and additionally produces an antifungal compound that inhibits multi-drug-resistant fungal pathogens including Lomentospora prolificans. S. formicae is genetically tractable using CRISPR-Cas9 gene editing, allowing for detailed analysis of the formicamycin biosynthetic gene cluster. AntiSMASH analysis predicts the genome to encode at least 45 secondary metabolite biosynthetic gene clusters, many of which appear to encode novel compounds. Current research efforts are focussing on characterising the regulation of secondary metabolism at a global level in order to switch on pathways that are not typically expressed under standard laboratory conditions with the aim of identifying novel antimicrobials.

Streptomyces

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products