Search PubMedSearch

PubMed · 41863303

Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2.

Abstract

MOTIVATION: Infectious diseases continue to be a leading cause of mortality and pose a significant global health threat. Thus, the development of tools for surveillance and early detection of emerging pathogens is needed. RESULTS: We introduce PathogenFinder2, a novel, alignment-free, taxonomy-agnostic model for predicting bacterial pathogenic capacity in humans using protein language models. It outperforms previous methods, particularly for novel taxa, and provides interpretable outputs by highlighting proteins most relevant to pathogenic potential. These insights aid the identification of virulence factors, vaccine targets, and infection-related metabolic pathways. Furthermore, we introduce the Bacterial Pathogenic Capacity Landscape, which reveals patterns linked to host condition, infection site, microbial antagonism, and environmental origin. AVAILABILITY: The model is freely available online at https://genepi.dk/pathogenfinder2, or as a standalone program (https://github.com/genomicepidemiology/PathogenFinder2).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alfred Ferrer Florensa, Jose Juan Almagro Armenteros, Rolf Sommer Kaas, Philip Thomas Lanken Conradsen Clausen, Henrik Nielsen, Burkhard Rost, Frank Møller Aarestrup. 2026-05-03. Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2.. https://doi.org/10.1093/bioinformatics%2Fbtag129

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

SURE-Pipe: a pipeline to compare genomes and extract shared and unique regions.

Identification of unique and shared genomic regions between organisms has substantial translational potential for the development of marker-based diagnostic assays and sequence homology-driven taxonomic classification. An automated pipeline capable of performing genome comparisons at both the intra- and inter-species levels with minimal computational requirements can significantly advance genome-driven translational research. Species-specific genomic regions are particularly valuable for sequence-based species identification and for developing DNA amplification- or hybridization-based diagnostic assays. Here, we present SURE-Pipe, an automated and flexible pipeline for genome comparison and extraction of unique and shared genomic regions (https://github.com/BPaul-bioinfoLAB/SURE-Pipe). Benchmarking of this pipeline using simulated datasets demonstrated high accuracy for shared and unique region identification. Using the pairwise genome comparison module, six genome pairs from diverse microorganisms were analysed, and identified the unique and shared regions. In addition, the multigenome comparison module was applied to 96 genomes representing 24 Bacillus species and identified species-specific genomic regions. These regions were highly conserved among four strains of a species (>98% sequence identity) and exhibit little to no similarity with other species. Species-specific primers designed for all 24 Bacillus species showed no off-target amplification in in-silico polymerase chain reaction analysis, indicating their specificity. Overall, SURE-Pipe provides a robust and multipurpose framework for comparative genomics, and the outcomes can be used for species identification and the development of genome-based diagnostic approaches.

Genome, Bacterial

Complete genomes from a xenic Dolichospermum flosaquae FBCC-A233 culture reveal genome-inferred metabolic asymmetry with associated bacteria.

Cyanobacteria form phycosphere communities with associated bacteria, but genome-resolved resources are needed to formulate testable hypotheses about their metabolic interactions. Here, we reconstructed three complete circular genomes from a unialgal xenic culture, including Dolichospermum flosaquae FBCC-A233 and two associated alphaproteobacterial genomes assigned to Sphingorhabdus sp. and Brevundimonas sp. Genome-wide read mapping and genome-quality assessment supported the three recovered genomes as high-quality circular reconstructions. Comparative genome analysis placed the cyanobacterial genome within the Dolichospermum flosaquae species cluster under the GTDB framework, while the associated bacterial genomes represented Sphingorhabdus sp. and a putative undescribed Brevundimonas species-level lineage. Genome architecture analysis indicated reduced genome size and gene content in Brevundimonas relative to genus-level references although additional metrics did not support a strong conclusion of classical genome streamlining. Selected KEGG module and KO-level reconstructions indicated genome-inferred metabolic asymmetries across the consortium. FBCC-A233 encoded photosynthesis- and nitrogen-related modules and a BioU-mediated de novo biotin biosynthesis route, whereas the associated bacteria lacked complete de novo biotin biosynthesis but retained biotin-dependent carboxylase genes. FBCC-A233 also encoded extensive anaerobic corrinoid biosynthesis potential; however, canonical DMB-containing cobalamin completion, cobamide identity, and complete transporter systems were not resolved. Together, these complete genomes provide a genome-resolved resource for investigating genome-inferred metabolic differentiation and ecological interactions in cyanobacteria-associated bacterial consortia.IMPORTANCEPhycosphere interactions between cyanobacteria and associated bacteria can shape aquatic microbial communities, but many proposed interactions remain difficult to evaluate without genome-resolved resources. This study provides three complete circular genomes from a unialgal xenic Dolichospermum flosaquae culture, capturing the cyanobacterium and two co-maintained bacterial associates. Our analysis identifies genome-inferred metabolic asymmetries, particularly in biotin- and cobamide-related pathways. D. flosaquae FBCC-A233 encoded candidate de novo biotin and corrinoid biosynthesis capacity, whereas the associated bacteria lacked complete de novo pathways but retained cofactor-dependent enzymes. These findings nominate cofactor-related dependencies as experimentally testable hypotheses while emphasizing unresolved uptake, export, cobamide identity, and growth-dependence mechanisms. The complete genomes and KO-level reconstructions generated here provide a resource for future studies of cyanobacteria-associated consortia.

Genome, Bacterial

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial