Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.

Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

Gene Regulatory Networks↗

Molecular biology and integrated strategies for activating cryptic biosynthetic gene clusters toward next-generation antibiotic discovery.

Antimicrobial resistance (AMR) has been identified as one of the 21st century's severest global public health crises. AMR led to an estimated 4.95 million deaths in 2019 and will claim 10 million lives a year by 2050 in the absence of targeted interventions. During the same period, the number of novel antibiotics discovered has decreased drastically as many researchers are rediscovering known antibiotics, non-model microorganisms are poorly understood or difficult to culture and antibiotic research and development investment has declined drastically. However, high-throughput whole genome sequencing and the subsequent application of bioinformatics in bacterial and fungal genomes have shown that a numerous of cryptic or silent biosynthetic gene clusters (BGCs) remain latent at ambient laboratory conditions since their genes are transcriptionally inactive. Cryptic BGCs represent a vast source of unique secondary metabolites, many of which may yield novel antibacterial, antifungal, anti-cancer and other potentially valuable natural products. This review discusses the biological relevance of cryptic BGCs, the major limiting factors that restricts their activation and novel strategies that have been employed to activate them and exploit their potential to produce novel natural products. The review focuses on biological approaches including CRISPR-Cas mediation for the activation of cryptic BGCs, promoter engineering, pathway refactoring, and heterologous expression; biochemical strategies such as Osman, OsMAC, Precursor Feeding, Chemical Elicitation, Epigenetic Regulation and Co-cultivation and technology-based strategies such as Genome mining, Microfluidic Cultivation systems, High-Throughput Screening, Metabolomics, Molecular Networking and Artificial Intelligence and Machine Learning based prediction of BGCs and their metabolites. The use of multi-omics technologies combined with synthetic biology to achieve better discovery, characterization and large-scale production of novel natural products is also discussed herein. Finally, we will talk about the ecological significance and evolutionary advantage of cryptic BGCs' role in interactions between microorganisms, such as competition, communication, symbiosis and environmental adaptability, so as to provide a useful background for accelerating next-generation antibiotics.

CRISPR-Cas activation↗

Effectidor II: a pan-genomic AI-based algorithm for the prediction of type III secretion system effectors.

MOTIVATION: Type III secretion systems are used by many Gram-negative bacteria to inject type 3 effectors (T3Es) directly into eukaryotic cells, promoting disease or provoking immune response. Because of these opposing evolutionary forces, T3E repertoires often vary within taxonomic groups. Identifying the full effector gene repertoire in genomes of related individuals is crucial for determining core and specialized effectors, understanding the disease dynamics, and developing appropriate management strategies against pathogens. It can also help uncover novel T3Es that have recently emerged in a population. Our previously published Effectidor web server successfully addressed the challenge of identifying T3Es in a single bacterial genome. Here, we enriched the web server with various novel capabilities, including the identification of T3Es from multiple genome sequences simultaneously. RESULTS: We present Effectidor II, a web server that relies on machine learning to predict T3E-encoding genes within bacterial pan-genomes. We demonstrate the benefit of learning based on features extracted from the entire sequences comprising the pan-genome and report a novel T3E discovered by it in Xanthomonas euroxanthea. AVAILABILITY AND IMPLEMENTATION: Effectidor II is available at: https://effectidor.tau.ac.il and the source code is available at: https://github.com/naamawagner/Effectidor. A stand-alone version of Effectidor II is available at: https://github.com/naamawagner/Effectidor/tree/StandAlone. The source code for the standalone version and the data used in this work are also provided in https://doi.org/10.5281/zenodo.15081636.

Type III Secretion Systems↗

Investigating genetic, antigenic, and structural diversity in the Neisseria gonorrhoeae outer membrane protein, PorB: implications for vaccine design.

UNLABELLED: Vaccines targeting Neisseria gonorrhoeae are needed to reduce disease burden and help address the problem of antimicrobial resistance, with an understanding of relationships between gonococcal genetics and molecules influencing diversity, infection, and the immune response essential for developing effective vaccine formulations. Whole-genome sequence data can be used to investigate these relationships among thousands of gonococcal isolates, allowing the study of antigenic diversity on a population scale. Such analyses typically examine antigenic diversity occurring in complete protein sequences, generating mean diversity indices and phylogenetic analyses that can inform on vaccine potential; however, to detect and measure the immune responses elicited, epitope characterization within an antigen helps guide vaccine formulations, with epitopes commonly located in surface-exposed regions of a protein. Here, we analyzed the genetic diversity of the major gonococcal antigen, PorB, in WGS from 22,227 N. gonorrhoeae isolates. We characterized the diversity of all eight surface-exposed outer membrane loops, or variable regions (VRs), and generated a PorB VR subtyping scheme to facilitate the global and temporal detection of circulating PorB subtypes. These analyses identified the presence of dominant VR combinations that persisted over time, indicative of (i) epistatic interactions between VRs and (ii) positive selection. Strain-specific, anti-PorB IgG responses directed toward distinct VR subtypes were detected in sera obtained from participants vaccinated with 4CMenB. The deconstruction of PorB into each surface-exposed loop provides a powerful approach for evaluating vaccine candidates: the methods used here allow immunodominant regions to be detected, which is invaluable for further vaccine investigations. IMPORTANCE: In the context of rising global gonorrhea cases, the development of vaccines becomes a priority; however, N. gonorrhoeae antigenic diversity and its ability to evade the immune system complicate vaccine development. This study characterizes the genetic diversity of the outer membrane protein, PorB, a key component of the outer membrane and a major gonococcal antigen. Using genomics and machine-learning techniques, this research identified dominant PorB variants that drive the immune response, proposing potential vaccine candidates and improving our understanding of the evolutionary forces maintaining genome structure and biological fitness. Understanding these processes is crucial for designing vaccines that effectively target N. gonorrhoeae and combat the spread of multidrug-resistant gonococci.

Neisseria gonorrhoeae↗

Identification and analysis of metabolic reprogramming-related genes in triple-negative breast cancer.

Triple-negative breast cancer (TNBC) is notorious for its rapid progression, tendency to metastasize, high recurrence rates, dismal outcomes, and limited treatment options, underscoring the urgent need to uncover new biomarkers and molecular pathways to enhance diagnosis, prognosis, and therapeutic strategies. Metabolic reprogramming continues to play a role throughout the life cycle of cancer, evolving and adapting. In this study, we aimed to identify specific genes associated with metabolic reprogramming in TNBC, which can potentially become unique biomarkers of this cancer. TNBC datasets retrieved from the Gene Expression Omnibus were employed to pinpoint genes exhibiting altered expression linked to tumor metabolic reprogramming. Key genes were accurately screened through machine learning algorithms, and then externally verified using the TBNC dataset based on the Cancer Genome Atlas database. Finally, immunohistochemical methods were used to clinically confirm the differential expression and trends of these key genes. Our analysis accurately identified four genes-CLEC7A, IRS1, RSPO3, and ALB-that are closely correlated with the metabolic reprogramming characteristics of cancer, and could be regarded as innovative biomarkers for TNBC. This opens a new avenue for further investigation into the mechanisms of metabolic reprogramming in TNBC and new treatment strategies.

Humans↗

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9±10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics↗

Sparse deconvolution of cell type medleys in spatial transcriptomics.

Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling However, current deconvolution models aim to estimate the proportions of distinct cell types in each spatial transcriptomics spot by integrating reference single-cell data. These models often assume strong overlap between the reference and spatial datasets, neglecting biology-grounded constraints such as sparsity and cell-type variations, as well as technical sparsity. As a result, these methods rely on over-permissive algorithms that ignore given constraints leading to inaccurate predictions, particularly in heterogeneous or unmatched datasets. We introduce Weight-Induced Sparse Regression (WISpR), a machine learning algorithm that integrates spot-specific hyperparameters and sparsity-driven modeling. Unlike conventional approaches that neglect biology-grounded constraints, WISpR accurately predicts cell-type distributions while preserving biological coherence, i.e., spatially and functionally consistent cell-type localization, even in unmatched datasets. Benchmarking against five alternative methods across ten datasets, WISpR consistently outperformed competitors and predicted cellular landscapes in both normal and cancerous tissues. By leveraging sparse cell-type arrangements, WISpR provides biologically informed, high-resolution cellular maps. Its ability to decode tissue organization in both healthy and diseased states highlights WISpR's practical utility for spatial transcriptomics, particularly in challenging settings involving noise, sparsity, or reference mismatches.

Humans↗

Adapting systems biology to address the complexity of human disease in the single-cell era.

Systems biology aims to achieve holistic insights into the molecular workings of cellular systems through iterative loops of measurement, analysis and perturbation. This framework has had remarkable success in unicellular model organisms, and recent experimental and computational advances - from single-cell and spatial profiling to CRISPR genome editing and machine learning - have raised the exciting possibility of leveraging such strategies to prevent, diagnose and treat human diseases. However, adapting systems-inspired approaches to dissect human disease complexity is challenging, given that discrepancies between the biological features of human tissues and the experimental models typically used to probe function (which we term 'translational distance') can confound insight. Here we review how samples, measurements and analyses can be contextualized within overall multiscale human disease processes to mitigate data and representation gaps. We then examine ways to bridge the translational distance between systems-inspired human discovery loops and model system validation loops to empower precision interventions in the era of single-cell genomics.

Humans↗

Cooccurrence of Homologous Recombination Deficiency and Mismatch Repair Deficiency in Colorectal Cancer.

Homologous recombination deficiency (HRD) in colorectal cancer (CRC) remains largely unexplored. In contrast, mismatch repair deficiency (dMMR) occurs in ∼15% of patients with CRC. Although HRD and dMMR have historically been regarded as mutually exclusive, emerging evidence suggests that this mutual exclusivity may not be absolute. Here, we conducted a retrospective cohort study utilizing genomic and transcriptomic data to define HRD status in a Chinese dMMR CRC cohort (n = 99). Multiple machine learning approaches were employed to analyze the expression profiles of these tumors and to develop a classifier distinguishing HRD from homologous recombination proficiency (HRP) in dMMR CRCs. In the Chinese dMMR CRC cohort, 66% of tumors were classified as HRD. Compared with the HRP group, the HRD group had a significantly higher tumor mutational burden and better outcomes. The derived expression signature, comprising eight genes, successfully predicted HRD status in dMMR tumors with high accuracy in the training set (AUC = 0.88, Naïve Bayes) and the test set (AUC = 0.87). In this study, a subset of dMMR CRC tumors with co-occurring HRD was identified, which may have potential implications for patient stratification and the application of targeted therapies, such as PARP inhibitors, in this molecular subgroup.

colorectal cancer↗

Cis-regulatory control of transcriptional timing and noise in response to estrogen.

Cis-regulatory elements control transcription levels, temporal dynamics, and cell-cell variation or transcriptional noise. However, the combination of regulatory features that control these different attributes is not fully understood. Here, we used single-cell RNA-seq during an estrogen treatment time course and machine learning to identify predictors of expression timing and noise. We found that genes with multiple active enhancers exhibit faster temporal responses. We verified this finding by showing that manipulation of enhancer activity changes the temporal response of estrogen target genes. Analysis of transcriptional noise uncovered a relationship between promoter and enhancer activity, with active promoters associated with low noise and active enhancers linked to high noise. Finally, we observed that co-expression across single cells is an emergent property associated with chromatin looping, timing, and noise. Overall, our results indicate a fundamental tradeoff between a gene's ability to quickly respond to incoming signals and maintain low variation across cells.

Humans↗

Circulating microRNA panels for multi-cancer detection and gastric cancer screening: leveraging a network biology approach.

BACKGROUND: Screening tests, particularly liquid biopsy with circulating miRNAs, hold significant potential for non-invasive cancer detection before symptoms manifest. METHODS: This study aimed to identify biomarkers with high sensitivity and specificity for multiple and specific cancer screening. 972 Serum miRNA profiles were compared across thirteen cancer types and healthy individuals using weighted miRNA co-expression network analysis. To prioritize miRNAs, module membership measure and miRNA trait significance were employed. Subsequently, for specific cancer screening, gastric cancer was focused on, using a similar strategy and a further step of preservation analysis. Machine learning techniques were then applied to evaluate two distinct miRNA panels: one for multi-cancer screening and another for gastric cancer classification. RESULTS: The first panel (hsa-miR-8073, hsa-miR-614, hsa-miR-548ah-5p, hsa-miR-1258) achieved 96.1% accuracy, 96% specificity, and 98.6% sensitivity in multi-cancer screening. The second panel (hsa-miR-1228-5p, hsa-miR-1343-3p, hsa-miR-6765-5p, hsa-miR-6787-5p) showed promise in detecting gastric cancer with 87% accuracy, 90% specificity, and 89% sensitivity. CONCLUSIONS: Both panels exhibit potential for patient classification in diagnostic and prognostic applications, highlighting the significance of liquid biopsy in advancing cancer screening methodologies.

Neoplasms↗

Nature-based meaning-focused photography intervention enhances subjective well-being: A three-arm randomized controlled study.

Gaining meaning from nature contact can promote subjective well-being. However, few studies have validated the effectiveness of nature-based meaning interventions in enhancing subjective well-being. This study consisted of a 7-day online intervention to examine the effects of nature-based meaning-focused photography on well-being by comparing a photo-only group, a photo + writing group, and a waiting list control group and how meaning in life mediates the relationship between nature contact and well-being. A pre-registered three-arm randomized controlled trial (groups: photo + writing group vs. photo-only group vs. control group) * (time: pre-test vs. post-test vs. 1-month follow-up) was conducted with 219 college students. In the photo + writing group, participants captured nature scenes and wrote 100-word reflections. The photo-only group only took nature photos. The primary outcomes were meaning in life and well-being, and the secondary outcome was life satisfaction. A conservative Bayesian causal forest analysis based on machine learning was used to detect both treatment and heterogeneous intervention effects. Compared with the control group, the photo + writing group showed positive effects on meaning in life, subjective well-being, and life satisfaction, with average treatment effects of 0.36, 0.27, and 0.66 standard deviations (SD), respectively. The photo-only group also showed generally positive effects on these outcomes, with average treatment effects of 0.27, 0.24, and 0.54 SD, respectively. However, these effects were not sustained after 1 month. The intervention was especially beneficial for participants from lower subjective socioeconomic status, with limited prior nature exposure, or lower baseline psychological well-being. Importantly, enhanced meaning in life helped explain how the intervention improved well-being and life satisfaction. This study also demonstrated that combining nature-based photography and reflective writing can improve well-being.

Humans↗

Assessing data size requirements for training generalizable sequence-based TCR specificity models via pan-allelic MHC-I point-mutation ligandome evaluation.

Rapid identification of T cell receptors (TCRs) that specifically bind patient-unique neoepitopes is a critical challenge for personalized TCR-based therapies in oncology. Due to enormous diversity of both TCR and neoepitope repertoires, a machine learning predictor of TCR-pMHC specificity for personalized therapy must generalize to TCRs and epitopes not seen in the training data. We estimate the necessary size of such training data. We first confirm that published models fail to generalize beyond a single-residue dissimilarity to the epitope training set distribution. We then impute the point-mutation ligandome across the 34 most prevalent human MHC alleles and represent it as a graph based on our established dissimilarity cutoff. By finding the dominating set of this graph, we estimate that between one and 100 million epitopes are required to train a generalizable sequence-based TCR specificity prediction model-1000 times the size of current public data.

Humans↗

Digital profiling of dysarthria in late- and early-onset Parkinson's disease.

BackgroundDigital speech analysis affords robust markers of Parkinson's disease (PD). However, most studies target late-onset PD (LOPD), neglecting early-onset PD (EOPD) -an increasingly prevalent subtype. This proof-of-concept study tackles such gap.MethodsWe used machine learning to discriminate persons with EOPD (with symptom onset before age 50) and LOPD (with symptom onset after age 50) from healthy controls (HCs) through prosodic and articulatory features from natural speech.ResultsMaximal classification between patients and HCs was afforded by combined prosodic and articulation features in LOPD (AUC&#x2009;=&#x2009;0.90) and by articulation alone in EOPD (AUC&#x2009;=&#x2009;0.79), with chance-level discrimination between patient groups (AUC&#x2009;=&#x2009;0.55). Motor severity (MDS-UPDRS-III) scores predicted by these features correlated with actual motor severity scores in both LOPD (r&#x2009;=&#x2009;0.52, p&#x2009;<&#x2009;0.001) and EOPD (r&#x2009;=&#x2009;0.27, p&#x2009;<&#x2009;0.001).ConclusionsDigital speech markers offer markers of PD irrespective of age of onset.Plain language summary titleVoice recordings capture motor symptoms in Parkinson's disease irrespective of age of onset.

Humans↗

A Multi-omics Regulated Cell Death Framework Defines Immune Phenotypes and Guides Precision Therapy in Colorectal Cancer.

Colorectal cancer (CRC) is molecularly and immunologically heterogeneous, contributing to variable treatment response. Because regulated cell death (RCD) intersects with tumor metabolism, immune regulation, and therapeutic susceptibility, we built an RCD-centered framework for CRC stratification. Multi-cohort transcriptomic data were used to infer RCD subtypes with non-negative matrix factorization (NMF) and non-negative least squares (NNLS). Genomic, bulk RNA-seq, single-cell RNA-seq, and spatial transcriptomic datasets were integrated to characterize subtype-associated biology. Machine-learning models were developed for immunotherapy response and survival-risk estimation. Candidate compounds were screened by GDSC2-based drug-sensitivity modeling and molecular docking, and FSTL3 was functionally assessed in vitro. The framework separated CRC samples into two RCD-related phenotypes resembling immune-hot and immune-cold states. RCD1 showed immune activation and higher mutational burden, whereas RCD2 showed immune-suppressed features, intratumoral heterogeneity, and aggressive biology. RCD-associated signatures showed potential for predicting immunotherapy response and survival risk. Dasatinib was prioritized for immune-cold, high-risk tumors, with preliminary evidence supporting its activity in CRC cells, while functional assays suggested a role for FSTL3 in growth, invasion, epithelial-mesenchymal transition, and apoptosis regulation. These findings suggest that RCD-based multi-omics analysis may refine CRC stratification and help generate therapeutic hypotheses.

Colorectal cancer↗

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified &#x223c;218,000 OGs supported by expression evidence, while &#x223c;330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans↗

Spatial concordance metrics and related risk factors of brain-peripheral barrier axes: unveiling distinct concordance patterns for mental and neurological axes.

Numerous studies have documented bidirectional interactions between the central nervous system and barrier organs (skin, gut, and lung). While genome-wide association studies have revealed shared genetic factors across brain-peripheral barrier axes, investigating these connections from an environmental perspective in large populations remains difficult. Using data from the Global Burden of Disease (GBD) 2023, I extracted annual incidence rates for 56 diseases related to brain-peripheral barrier axes and exposure rates for the 70 most detailed risk factors across 204 countries and territories. By categorizing regional incidence rates into four quartiles for each disease, I pinpointed regions with concordance of these axes and constructed a spatial atlas of disease concordance within the brain-peripheral barrier axis from a macro-epidemiologic view. Subsequently, I calculated global spatial concordance percentages for each axis, which allowed the comparatively assessment of concordance patterns across different axes, specific diseases, and their variations over time, across the lifespan, and by gender. Finally, I applied machine learning models and Shapley additive explanations to identify risk factors related to the spatial concordance of each axis. From 1990 to 2023, the overall trend for most brain-peripheral barrier axis pairs remained stable. Spatial concordance patterns showed dynamic fluctuations across the lifespan, followed by a convergence toward stability in older age. Several risk factors are related to most brain-peripheral barrier axes. Notably, the mental and neurological axes exhibited distinct concordance patterns. Compared with neurological axes, concordance within mental axes showed a broader and more dispersed geographic distribution, with greater variation across sexes and over time. Furthermore, concordance percentages of mental and neurological axes exhibited opposing age-related trends, contrasting disease spectra for peripheral conditions, and inverse relationships with alcohol and sodium consumption. Those divergences suggest distinct mechanisms underlying the brain-peripheral barrier axes in mental and neurological diseases. Related risk factors offer population-based hypotheses for further investigation in individual-level studies.

Humans↗

Fine-grained structural classification of biosynthetic gene cluster-encoded products.

MOTIVATION: Biosynthetic gene clusters (BGCs) are responsible the biosynthesis of many natural products, including a multitude of effective therapeutics and their precursors. Advances in genomic data collection as well as computational techniques have made it possible to identify BGCs at scale. However, accurately determining the types of BGC-encoded products from genomic content remains elusive. RESULTS: Here, we introduce BGC annotation tool (BGCat), a machine learning method for fine-grained structural classification of BGC-encoded products, leveraging the NPClassifier natural product nomenclature. Our method leverages a pre-trained protein language model for creating meaningful gene representations and a deep neural network for class label prediction. We show the method outperforms state-of-the-art approaches in coarse-grained product classification and is effective for detailed classification. We implement a clustering-based augmentation strategy for BGC-product relationships, addressing a crucial gap in the available datasets. We then introduce the concept of product class profiles of gene cluster families (GCFs), associating each GCF with a probabilistic distribution of product types and offering a new perspective on GCF functions. Lastly, we use BGCat to provide new product class labels for over 100k BGCs in antiSMASH DB that presently have minimal information about their products. AVAILABILITY AND IMPLEMENTATION: The source code and trained model weights are freely available at https://github.com/HassounLab/BGCat.

Multigene Family↗