Search PubMedSearch

SEARCH · Search PubMed

Results for “antiSMASH”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Biosynthetic potential of the culturable foliar fungi associated with field-grown lettuce.

Fungal endophytes and epiphytes associated with plant leaves can play important ecological roles through the production of specialized metabolites encoded by biosynthetic gene clusters (BGCs). However, their functional capacity, especially in crops like lettuce (Lactuca sativa L.), remains poorly understood. We sequenced the genomes of nine fungal isolates, representing Fusarium sp., Fulvia sp., Alternaria alternata, and Alternaria postmessia, from leaves of lettuce grown under field conditions in Arizona, USA. We used antibiotics and secondary metabolite analysis shell (antiSMASH) and the database for automated carbohydrate-active enzyme annotation (dbCAN3), to predict BGCs and carbohydrate-active enzymes (CAZymes) for each strain, and then compared them to conspecific strains from other environments and substrates. Foliar lettuce-associated fungi featured 39-95 BGCs per genome, with substantial overlap between isolates occurring in association with lettuce leaves vs. from other substrates. Species identity was a significant determinant of BGC count, while host type, isolation source, and lifestyle were not. Several BGCs, including those for alternariol and 1,3,6,8-Tetrahydroxynaphthalene (T4HN), showed 100% similarity to characterized minimum information about a biosynthetic gene cluster (MIBiG) clusters based on antiSMASH predictions. Although analysis by biosynthetic gene similarity clustering and prospecting engine (BiG-SCAPE) identified gene cluster families (GCFs) across the dataset, these reference-matching clusters were not always grouped, reflecting methodological differences in how the tools assess similarity. Comparative CAZyme analysis in a focal species (Fulvia sp.) revealed higher gene counts in a foliar lettuce-derived isolate than in tomato (Solanum lycopersicum)-associated strains, challenging assumptions about host chemical complexity. These results highlight the importance of phylogenetic context in shaping fungal functional potential and suggest that selection on microbial traits in edible leafy crops may be more subtle and species-specific than previously assumed. KEY POINTS: • Lettuce-associated fungi feature diverse biosynthetic potential • Phylogeny predicts fungal BGC content more strongly than ecological lifestyle • Findings support genome-informed microbiome strategies for leafy crops.

Lactuca

Exploring biosynthetic potential of the endophytic Penicillium turbatum BLH34 using whole-genome sequence analysis and molecular networking.

An in-depth genomic and metabolomic investigation was conducted on the endophytic fungus Penicillium turbatum BLH34, isolated from Macleaya cordata. Hybrid sequencing (Illumina-Nanopore) generated a high-quality 27.9 Mb genome (GC 48.6%) encoding 9798 proteins, with functional annotation linking 5350 genes to the NCBI non-redundant database and 3404 to KEGG pathways. AntiSMASH analysis uncovered 35 biosynthetic gene clusters (BGCs), 23 of which lacked homology to known pathways, highlighting BLH34's potential for novel metabolite discovery. Molecular networking (GNPS) and LC-MS/MS identified 19 specialised metabolites, including antimicrobial polyketides. Bioassays demonstrated potent inhibition against Staphylococcus aureus (36 mm), Bacillus subtilis (28 mm) and Escherichia coli (24 mm), underscoring its pharmaceutical relevance.

Penicillium

Whole-Genome Analysis of Bacillus Licheniformis Ali5 and Synthesis of Lichenysin via Genome Shuffling.

Whole-genome sequencing of Bacillus licheniformis Ali5 was performed via MGI-seq PE150 and Nanopore single-molecule real-time sequencing. The strain has a 4,114,664 bp circular genome encoding 4030 protein-coding genes. Functional annotation across NR, COG, GO, KEGG, CARD, BacMet, and CAZy databases identified 4025, 2812, 988, 1242, 72, 69, and 94 corresponding genes, respectively, and antiSMASH 6.0 revealed multiple antimicrobial biosynthetic gene clusters, including intact lichenysin and lichenicidin VK21 A1/A2 gene clusters. Three rounds of recursive protoplast fusion-based genome shuffling, paired with a dual-index screening system, significantly improved strain growth and lichenysin biosynthesis. Recombinants exhibited shortened lag phase, enhanced proliferation, improved stationary-phase stability, and higher diauxic peak biomass. PP3-176 and PP3-186 showed 4.6%-8.1% higher 12-h shake-flask titer and 3.1%-4.0% higher maximum titer than the parental average, with excellent fermentation stability. 1-L bioreactor validation confirmed strong scale-up potential. PP3-186 achieved 27.2% and 31.6% titer increases at 12 h and 20 h, while PP3-176 yielded 20.4% and 14.6% improvements with robust metabolic performance. This study validates genome shuffling as an effective strategy for enhancing lichenysin production, providing candidate strains and technical support for industrial application.

Bacillus licheniformis

Characterization and genomic analysis of Bacillus halotolerans G3-2: a potential biocontrol agent against apple Alternaria leaf blotch disease.

BACKGROUND: Apple Alternaria leaf blotch (ALB) is a devastating disease threatening the apple industry worldwide. Biocontrol offers an effective and environmentally friendly alternative for disease management. RESULTS: Bacillus strain G3-2 exhibits strong antagonistic activity against Alternaria alternata (a major causal pathogen of ALB). In dual-culture assays, G3-2 inhibited A. alternata by 88.39%; in detached-leaf inoculation assays, it reduced the lesion area by >88%. 16S rRNA sequencing and phylogenetic analysis identified this strain as Bacillus halotolerans. Oxford Nanopore Technology (ONT) sequencing generated a 4.18-Mb complete genome (43.8% G + C) containing 4149 protein-coding genes, 30 rRNAs and 86 tRNAs. CAZy annotation identified 182 genes encoding carbohydrate-active enzymes (CAZymes), including glycoside hydrolases, glycosyltransferase, and carbohydrate esterases, suggesting potential for glycosylated secondary metabolite production. AntiSMASH analysis detected nine biosynthetic gene clusters, including those for surfactin, fengycin, bacillaene and laterocidine. Plate assays confirmed that G3-2 has the ability to produce protease, cellulase and siderophore. Moreover, it exhibits ~70% inhibition against several other phytopathogenic fungi. CONCLUSIONS: These findings demonstrate that G3-2 suppresses A. alternata through antibiosis (lipopeptides and polyketides), nutrient competition (siderophores) and cell-wall degradation (proteases and cellulases). Moreover, our study revealed that it has great potential to be used as a broad-spectrum, environmentally friendly biocontrol agent. © 2026 Society of Chemical Industry.

Alternaria

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500 m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed > 99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071ᵀ (= ATCC 10145ᵀ), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33 Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8 kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~ 22 kb, ~ 17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family

Comparative Genomics of Paenibacillus Secondary Metabolism: Unveiling the Putative Biosynthetic Gene Cluster for Paenialvins in Paenibacillus Alvei Strain 32.

In this study, we used comparative genomics and culture-based methods to investigate Biosynthetic Gene Clusters (BGCs) responsible for the production of antimicrobial peptides. Paenibacillus alvei strain 32 was isolated from a cystic fibrosis sputum. Its genome was sequenced using Illumina, showing a size of 6,584,590 bp with 239 contigs assembled in 26 scaffolds, an average coverage of 243X, and 6,832 coding sequences. ANI analysis and in silico DNA-DNA hybridization showed its affiliation inside Paenibacillus alvei, with a clear separation from other related strains, leading us to propose a distinct species-level genomic clade (genomospecies) within this group. AntiSMASH analysis predicted 22 putative BGCs in the genome of strain 32. Its culture supernatant exhibited inhibitory activity against Gram-positive pathogens, including methicillin-resistant Staphylococcus aureus (MRSA), Bacillus cereus, and Enterococcus faecalis. By comparing in silico BGC predictions with activities described in the literature, we propose that strain 32 harbours a specific 110-kb cluster (cluster 6.2) with five non-ribosomal peptide synthetase (NRPS) genes. These synthetases are predicted to direct the assembly of a 16-amino acid backbone that correlates with the structure of paenialvins, which are known anti-MRSA molecules. This study describes the putative biosynthetic pathway of the paenialvins and explains structural variations, bringing useful data on Paenibacillus secondary metabolism for future antibiotic development.

Paenibacillus alvei

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9 Mb chromosome with a G + C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4

Genomic analysis of regulatory mechanisms governing EPS66A biosynthesis in Streptomyces changanensis HL-66.

Streptomyces changanensis HL-66 produces the α-(1,4)/(1,6)-glucan exopolysaccharide EPS66A, a potent plant immune elicitor with promising applications in plant protection. However, its low native fermentation yield limits large-scale application. To investigate the biosynthetic potential and regulatory mechanisms underlying EPS66A production, the whole genome of HL-66 was sequenced and analyzed. The HL-66 genome is 6.82 Mb in size, with a GC content of 74%, and encodes 6081 predicted functional genes. Among these, 1390 genes were annotated to Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, 4187 were assigned to Gene Ontology (GO) terms, and 143 were classified into Clusters of Orthologous Groups (COG) categories. antiSMASH analysis identified 22 secondary metabolite biosynthetic gene clusters, including multiple polyketide synthase (PKS) and nonribosomal peptide synthetase (NRPS) clusters. Functional analyses revealed that the glycosyltransferase gene (GTy) and the global regulatory gene (bldD) are involved in EPS66A biosynthesis. bldD is involved in morphological development and EPS66A production, whereas GTy specifically regulates EPS66A production without affecting growth or development. In both in vivo and potted-plant experiments, EPS66A (200 μg/mL) significantly reduced the severity of tobacco mosaic virus, apple anthracnose leaf spot, walnut bacterial leaf spot, and jujube anthracnose, achieving control efficacies of 90.21%, 87.95%, 77.41%, and 68.55%, respectively, and outperforming a commercial chitosan oligosaccharide control. These findings provide new insights into the genetic architecture and regulatory mechanisms of EPS66A biosynthesis and support its development as a polysaccharide-based green pesticide.

Streptomyces

Discovery of Glycosylated β-Amino Acid-Containing Macrolactams from Nonomuraea sp. 0L2P via Genome Mining.

β-Amino acid-containing macrolactams (β-AACMs) are a class of bioactive natural products characterized by nitrogen-containing starter units within polyketide-derived macrocycles. Here, we report four previously undescribed macrolactams, gruelactams A-D (1-4), from Nonomuraea sp. 0L2P, discovered through an integrated approach combining genome mining, 15N-labeling, and antibacterial screening. Their planar structures were elucidated by comprehensive spectroscopic analyses, including 1D and 2D NMR and HRESI-MS, and their configurations were partially assigned based on ROESY data and bioinformatic analysis. Genome sequencing and antiSMASH analysis identified a putative type I polyketide synthase (PKS) biosynthetic gene cluster, enabling the proposal of a biosynthetic pathway. Bioactivity assays showed that gruelactam D (4) exhibits antibacterial activity against Bacillus cereus and Staphylococcus aureus, with MIC values of 8 and 16 μg/mL, respectively. These findings expand the chemical diversity of β-AACMs and demonstrate the utility of genome-guided approaches for discovering bioactive natural products from rare actinomycetes.

Anti-Bacterial Agents

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus

Microbe Profile: Streptomyces formicae KY5: an ANT-ibiotic factory.

Streptomyces formicae KY5 was isolated from a Tetraponera penzigi plant-ant nest. It is primarily known for its production of the formicamycins, antibiotics with potent activity against Gram-positive pathogens including methicillin-resistant Staphylococcus aureus, and additionally produces an antifungal compound that inhibits multi-drug-resistant fungal pathogens including Lomentospora prolificans. S. formicae is genetically tractable using CRISPR-Cas9 gene editing, allowing for detailed analysis of the formicamycin biosynthetic gene cluster. AntiSMASH analysis predicts the genome to encode at least 45 secondary metabolite biosynthetic gene clusters, many of which appear to encode novel compounds. Current research efforts are focussing on characterising the regulation of secondary metabolism at a global level in order to switch on pathways that are not typically expressed under standard laboratory conditions with the aim of identifying novel antimicrobials.

Streptomyces

Genomic Insights Into Multidrug-Resistant Foodborne Serratia liquefaciens Strains Carrying mcr-9 and Comparative Genomic Analysis of Novel Biosynthetic Gene Clusters.

Serratia liquefaciens is an opportunistic nosocomial pathogen with a wide range of antibiotic resistance patterns. This study reports the characterization of the first mcr-9-positive S. liquefaciens strains, 35E-19E1 and CST-066, isolated from meat products in Japan. The strains were screened for the presence of β-lactamases, plasmid-mediated mobile colistin resistance (mcr) genes, and carbapenemase-encoding genes using PCR. Antimicrobial susceptibility was tested using the broth microdilution method. The strains exhibited multidrug resistance (MDR) phenotypes to third-generation cephalosporins, cephamycin, fosfomycin, and other clinically important antimicrobials. Genomic DNA sequencing showed that the genome sizes of CST-066 and 35E-19E1 are 5,529,704 and 5,261,506 bps, respectively. mcr-9 was identified on a chromosome within a genetic environment that included the two-component system qseBC, which plays a key role in the signaling network that triggers colistin resistance in Enterobacterales. Downstream genome analysis revealed a 1695-bp eptB-like kdo2-lipid phosphoethanolamine transferase, which is involved in intrinsic polymyxin resistance mechanisms in Serratia spp. The strain 35E-19E1 carries five CRISPR-Cas enzymes that are essential for adaptive immunity in bacteria, allowing defense against invading elements. Functional analysis using subsystem technology revealed that both strains possess subsystem features responsible for invasion and adhesion within the host biomes. Genome mining using antiSMASH and BAGL4 revealed various biosynthetic gene clusters, responsible for secondary metabolite synthesis. Notably, we identified novel gene clusters, mainly nonribosomal peptide synthetases, in both the strains, indicating their potential to produce bioactive compounds. Although the presence of mcr-9 in Serratia may not be of clinical significance because of natural resistance of the strain to polymyxins, we shed light on the genomic characteristics of this MDR pathogen and the potential spread of mcr-9 among other bacterial species. The emergence of mcr-9 in drug-resistant S. liquefaciens provides significant insights, underscoring the need for increased surveillance of this pathogen.

biosynthetic gene cluster

Complete genome sequence of Streptomyces californicus ADR1, an anti-infective, anti-biofilm and anti-oxidant producing endophyte isolated from the medicinal plant Datura metel.

OBJECTIVE: Streptomyces californicus strain ADR1 is an endophytic actinobacterium isolated from Datura metel that produces secondary metabolites with potent antibacterial and anti-biofilm activities against WHO-listed high-priority Gram-positive pathogens. While anti-bacterial and antioxidant potential of the strain ADR1 has been extensively characterized, its complete genome sequence remains to be investigated for further insights into its biosynthetic potential. This study presents the complete genome sequence analysis of the strain ADR1 to provide a robust genomic foundation for understanding its metabolic versatility and biosynthesis of compounds with therapeutic significance. DATA DESCRIPTION: The ADR1 genome was sequenced using Illumina HiSeq. The assembly comprised 262 scaffolds with a total genome size of 8.4 Mb and G + C content of 72.5%, containing 7427 protein-coding genes. AntiSMASH and IIT-Hyderabad novelBGC analysis revealed 39 biosynthetic gene clusters, including non-ribosomal peptide synthetases, type I polyketide synthases, terpene and melanin clusters, correlating with the diverse therapeutic compounds previously identified through GC-MS analysis. This high-quality genome provides crucial insights into the biosynthetic potential underlying potent antimicrobial and antioxidant activities of the strain ADR1.

Streptomyces

Integrated functional genomics and safety assessment of plant-growth-promoting Caryophanales from post-maize-cultivation soils.

This study aimed to evaluate six environmental bacterial strains isolated from post-maize cultivation soils as candidates for agricultural biopreparation development, using an integrated functional genomic and safety assessment framework. Building on experimental validation of plant-growth-promoting activities, the analysis included: plant-growth-promoting traits (PGPT-Pred) using PLABase; carbohydrate-active enzymes (CAZymes) relevant for lignocellulosic crop residue degradation (dbCAN3); secondary metabolite profiles (antiSMASH); and screening for virulence factors and antibiotic resistance genes (ABRicate, BTyper3).All analyzed strains possess 1,449-1,617 predicted PGPT-encoding genes (24.1-35.9% of total genes), which are strongly shaped by taxonomic relatedness, as confirmed by congruence testing against ANI-based genomic divergence. Paenibacillus amylolyticus 5mez and Priestia megaterium 7psych showed distinct functional profiles compared to Bacillus spp., while Bacillus subtilis sensu lato strains were most similar to each other. Genomic predictions suggest involvement in nutrient acquisition (N, P, K, Fe) and stress mitigation. Secondary metabolite analysis revealed high biosynthetic potential, with non-Bacillus species harbouring a large proportion of unknown gene clusters, indicating underexplored metabolite diversity. CAZyme profiling identified P. amylolyticus 5mez as the most enzyme-rich strain, while B. cereus s.s. zielonkawy showed ligninolytic potential despite low overall CAZyme abundance. The safety assessment identified B. cereus s.s. zielonkawy as toxigenic and unsuitable for use. Of the remaining strains, P. amylolyticus 5mez and Pr. megaterium 7psych demonstrated the most favourable safety profiles, exhibiting no detectable virulence factors or antibiotic resistance genes, justifying their priority use in agricultural biopreparations, pending phenotypic validation. Given the high-dimensional, low-sample-size nature of multi-trait datasets in applied microbial genomics, tailored statistical approaches, including noise-reduction-validated PCA and distance-based congruence testing, were applied; their rationale and limitations are discussed.

Soil Microbiology

Metagenomic analyses reveal E. coli-derived siderophores as potential signatures for breast cancer.

BACKGROUND: Breast cancer remains a leading cause of cancer-related mortality in women. Recent evidence implicates the gut microbiome and metabolites in breast cancer pathogenesis. This study explores associations between gut microbial species, their predicted metabolites, and breast cancer to uncover potential mechanistic insights. METHODS: Comprehensive metagenomic analyses were conducted on the gut microbiome of pre- and postmenopausal breast cancer patients, where microbial species were profiled through AMPHORA2 and metabolites were predicted through antiSMASH. Multivariate association analysis was used to identify significant associations between specific microbial species, predicted metabolites, and breast cancer status. A custom ensemble machine learning classifier was developed to classify pre- and postmenopausal breast cancer cases and controls based on microbial and predicted metabolite features. Additionally, a synthetic microbiome dataset was generated through MIDASim to validate the reproducibility of the ML results. Using our results, we explored the underlying dynamics of identified taxa and metabolite in breast cancer through literature and statistical support. RESULTS: Our analysis identified 471 microbial species and predicted 40 key metabolites in the metagenomic data. Multivariate analysis identified significant positive associations (p-value&#x2009;<&#x2009;0.05) of E. coli, siderophore, and thiopeptide with breast cancer. The custom ensemble model achieved accuracy and AUC as high as 78% and 90%, respectively, in classifying pre- and postmenopausal cases and controls. The high-ranking features i.e., E. coli, siderophore, and thiopeptide were consistent with the results of the multivariate association analysis, thereby substantiating their biological significance. Using these findings, we propose a mechanistic model in which E. coli secretes siderophores under iron-limited conditions in breast cancer patients, for iron sequestration from the host, which can potentially promote angiogenesis and tumor progression. CONCLUSION: Our findings suggest that microbial iron acquisition mechanisms may play a critical role in breast cancer pathophysiology. Functional validation of these mechanisms is needed to assess therapeutic potential. This study highlights gut microbiota and their metabolites as promising targets for breast cancer research and intervention.

Breast Neoplasms

Genomic exploration and in silico prioritization of putative COX-2-targeting metabolites from Streptomyces sp. VITGV156 (MCC 4965).

INTRODUCTION: Streptomyces species represent an important source of bioactive natural products, yet systematic genome-guided prioritization of metabolites targeting cyclooxygenase-2 (COX-2/PTGS2) remains limited. This study aimed to investigate the biosynthetic potential of Streptomyces sp. VITGV156 (MCC 4965) using an integrated genome mining and computational drug discovery pipeline. METHODS: Whole-genome sequencing, functional annotation, antiSMASH v7.0.1-based biosynthetic gene cluster (BGC) prediction, LC-MS/MS metabolomic profiling, SwissADME analysis, target prediction, disease association mapping, molecular docking against PTGS2 (PDB: 5IKR), and PASS bioactivity prediction were performed to prioritize putative bioactive metabolites. RESULTS: Genome analysis identified 29 predicted biosynthetic gene clusters, including clusters associated with geosmin, ectoine, albaflavenone, hopene, coelichelin, and SapB, together with several cryptic clusters exhibiting low similarity to known pathways. LC-MS/MS metabolomic profiling provided experimental support for active secondary metabolite production under the cultivation conditions employed. Computational prioritization identified PTGS2 (COX-2) as a biologically relevant target. Molecular docking demonstrated favorable binding affinities and interaction profiles for several predicted metabolites within the PTGS2 catalytic pocket. PASS analysis further suggested potential anticancer-related biological activities that require experimental validation. DISCUSSION: These findings demonstrate the utility of integrating genome mining, metabolomic profiling, and computational drug discovery for prioritizing natural-product candidates. Streptomyces sp. VITGV156 (MCC 4965) represents a promising source of biosynthetic diversity and provides a genome-guided framework for identifying putative COX-2-targeting natural products for future experimental validation rather than confirming metabolite production or biological activity.

COX-2 (PTGS2)

Comparative genomics reveals hidden biosynthetic diversity in Streptomyces spp. and metal-dependent regulatory features associated with untapped specialized metabolites.

The genus Streptomyces is one of the richest sources of bioactive natural products; however, a substantial proportion of its biosynthetic gene clusters (BGCs) remain cryptic and their metabolic products are unresolved. Advances in genome mining and computational prediction now enable comprehensive exploration of this hidden biosynthetic repertoire. In this study, whole-genome sequencing and comparative genomic analyses were performed on three three newly isolated Streptomyces strains to evaluate their specialized metabolic potential. Genome assemblies were annotated and systematically analyzed using antiSMASH, DeepBGC, GECCO, and PRISM to identify, cross-validate, and functionally characterize BGCs while predicting their associated secondary metabolite scaffolds. Taxonomic analyses based on Average Nucleotide Identity (ANI), phylogenomics, and BLAST identified the isolates as Streptomyces thinghirensis, Streptomyces novocaesareae, and Streptomyces griseorubens. Applying the consensus framework across the three Streptomyces genomes yielded 43 cryptic BGCs, lacking close similarity to reference BGCs in the MIBiG database, of which 26 were classified as HIGH, 10 as MEDIUM, and 7 as LOW confidence. Notably, numerous BGCs exhibited low abundance to characterized reference clusters, indicating a high potential for previously undescribed biosynthetic pathways and novel metabolite scaffolds. Comparative analyses further revealed strain-specific biosynthetic architectures together with putative metal-responsive regulatory systems; Fur, Zur, and Nur, which were frequently associated with specialized metabolite biosynthetic loci. Collectively, these findings demonstrate the effectiveness of integrated genome-mining strategies for prioritizing cryptic biosynthetic gene clusters and highlight the remarkable biosynthetic potential of newly identified Streptomyces isolates as a source of novel natural products.

comparative genomics