Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n = 38, 74%). Hierarchical clustering (n = 20) and K-means clustering (n = 14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans↗

Moving pieces in a taxonomic puzzle: venom 2D-LC/MS and data clustering analyses to infer phylogenetic relationships in some scorpions from the Buthidae family (Scorpiones).

The Buthidae is the most clinically important scorpion family, with over 500 species distributed worldwide. Taxonomical positions and phylogenetic relationships concerning the representative genera and species of this family have been mostly inferred based upon comparisons between morphological characters. Yet, some authors have performed such inferences by comparing some structural properties of a few selected molecules found in the venoms from these scorpions. Here, we propose a novel methodology pipeline designed to address these issues. We have analyzed the whole venoms from some species that exemplify peculiar cases in the Buthidae family (Tityus stigmurus, Tityus serrulatus, Tityus bahiensis, Leiurus quinquestriatus quinquestriatus and Leiurus quinquestriatus hebraeus), by means of a proteomic approach using a 2D-LC/MS technique. The molecules found in these venoms were clustered according to their physicochemical properties (molecular mass and hydrophobicity), by using the machine learning-based Weka software. The clusters assessment, along with the number of molecules found in a given cluster for each scorpion, which assigns for the venom and structural family complexities, respectively, was used to generate a phenetic correlation tree for positioning these species. Our results were in accordance with the classical taxonomy viewpoint, which places T. serrulatus and T. stigmurus as very close species, T. bahiensis as a less related species in the Tityus genus and L. q. quinquestriatus and L. q. hebraeus with small differences within the same species (L. quinquestriatus). Therefore, we believe that this is a well-suited method to determine venom complexities that reflect the scorpions' evolutionary history, which can be crucial to reconstruct their phylogeny through the molecular evolution of their venoms.

Animals↗

Saturating the eQTL map in Drosophila: Genome-wide patterns of cis and trans regulation of transcriptional variation in outbred populations.

Most genetic polymorphisms associated with complex traits are found in non-coding regions of the genome. Characterizing their effect presents a formidable challenge, and expression quantitative trait locus (eQTLs) mapping has been a key approach to do so. As comprehensive eQTL maps are available only for a few species, here we developed the Drosophila outbred synthetic population (Dros-OSP) and used it to characterize the landscape of transcriptional regulation in Drosophila melanogaster. We collected head and body transcriptomes and genomes from 1,286 outbred flies and mapped local and distant eQTLs for 98% of the genes. We characterized the network organization of the transcriptome across tissues and described the properties of local and distal eQTLs in terms of genetic diversity, heritability, connectivity, and pleiotropy. These results provide new insights into the genetic basis of transcriptional regulation in the fruit fly and offer a new mapping resource that will expand the possibilities currently available for the Drosophila community.

Animals↗

Attribution of PM2.5-Induced Transcriptomic Perturbation to Toxic Components.

Ambient fine particulate matter (PM2.5) is a chemically complex mixture whose health impacts are not fully captured by particle mass. Here, we developed an interpretable chemotranscriptomic framework to attribute PM2.5-induced molecular perturbations to toxicity-relevant components. PM2.5 collected from urban roadside and coastal environments was separated into whole, extractable, and unextractable fractions, characterized by LC/GC × GC-HRMS-based nontarget analysis and inductively coupled plasma mass spectrometry (ICP-MS), and evaluated using cytotoxicity testing and transcriptomic profiling in human bronchial epithelial cells. Urban PM2.5 exhibited greater cytotoxic potency per unit mass than coastal PM2.5, with extractable fractions accounting for most cytotoxic and pathway-level responses. Transcriptomics revealed distinct site-specific modes of action: urban PM2.5 preferentially induced oxidative stress, xenobiotic metabolism, and cell cycle suppression, consistent with acute, nonapoptotic injury, whereas coastal PM2.5 elicited weaker cytotoxicity but stronger interferon-mediated immune and apoptosis-related signaling. Integrating chemical abundance with pathway activity using random forest regression, SHAP interpretation, and mechanistic corroboration reduced 5,033 detected features to 444 pathway-linked candidate drivers. Fewer than 5% of features explained ∼95% of cumulative model contribution. Standard-confirmed contributors included plasticizer-related compounds, aromatic and heteroaromatic combustion products, and copper for urban PM2.5 and secondary/aged organics and nickel for coastal PM2.5. These findings support mechanism-informed prioritization of hazardous PM2.5 components beyond mass-based assessment.

Particulate Matter↗

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine↗

Clinical translation of senescence-related pan-cancer multi-omics: tools for assessment and immunotherapy prediction.

Cellular senescence (CS) exerts dual roles in tumorigenesis, yet its pan-cancer molecular characteristics and clinical value remain unclear, hindering its translation to oncology and personalized therapy. To address the lack of specific and universal tools for senescence assessment and immunotherapy response prediction, this study systematically analyzed 1259 CS-related genes from the CellAge database across 31 cancer types by integrating multi-omics data, including bulk RNA-seq, single-cell/spatial transcriptomics, and CRISPR screening. We developed a rank-based algorithm SenScoreR (publicly available at https://gxhub.shinyapps.io/SenScoreR/ ) for senescence quantification, validated with 10 independent datasets, and constructed a machine learning-based predictive model CS.Sig for immunotherapy response. Results showed that tumors had significantly lower Rank-based Senescence Score (RSS) than normal tissues across 31 cancers (average diagnostic AUC = 0.895), with low RSS linked to poor survival; high RSS correlated with reduced genomic instability, enriched CD8⁺ T/NK cell/macrophage infiltration, upregulated PD-L1 expression, and elevated immune cytolytic activity. CS.Sig demonstrated robust performance in predicting ICI response (AUC = 0.716 across 10 cohorts), outperforming 13 existing signatures, while CRISPR screening identified 17 senescence-related targets (e.g., CEP55, PPP1CC) whose knockout enhanced anti-tumor immunity. Our findings clarify CS's role in maintaining tumor genomic stability and shaping immune microenvironments, and the developed SenScoreR, CS.Sig, and identified targets bridge basic CS research with clinical oncology, providing a translational resource and hypothesis basis for future experimental and clinical validation.

Journal Article↗

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans↗

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

Organ-delimited gene regulatory networks provide high accuracy in candidate transcription factor selection across diverse processes.

Organ-specific gene expression datasets that include hundreds to thousands of experiments allow the reconstruction of organ-level gene regulatory networks (GRNs). However, creating such datasets is greatly hampered by the requirements of extensive and tedious manual curation. Here, we trained a supervised classification model that can accurately classify the organ-of-origin for a plant transcriptome. This K-Nearest Neighbor-based multiclass classifier was used to create organ-specific gene expression datasets for the leaf, root, shoot, flower, and seed in Arabidopsis thaliana. A GRN inference approach was used to determine the: i. influential transcription factors (TFs) in each organ and, ii. most influential TFs for specific biological processes in that organ. These genome-wide, organ-delimited GRNs (OD-GRNs), recalled many known regulators of organ development and processes operating in those organs. Importantly, many previously unknown TF regulators were uncovered as potential regulators of these processes. As a proof-of-concept, we focused on experimentally validating the predicted TF regulators of lipid biosynthesis in seeds, an important food and biofuel trait. Of the top 20 predicted TFs, eight are known regulators of seed oil content, e.g., WRI1, LEC1, FUS3. Importantly, we validated our prediction of MybS2, TGA4, SPL12, AGL18, and DiV2 as regulators of seed lipid biosynthesis. We elucidated the molecular mechanism of MybS2 and show that it induces purple acid phosphatase family genes and lipid synthesis genes to enhance seed lipid content. This general approach has the potential to be extended to any species with sufficiently large gene expression datasets to find unique regulators of any trait-of-interest.

Arabidopsis↗

Stage-specific ROMO1 in rheumatoid arthritis: predictive immune insights into the MIF pathway and HLA-DR/IL2RA axis via integrated GWAS, transcriptomic, single-cell, and spatial profiling.

Emerging evidence links reactive oxygen species modulator 1 (ROMO1), a key mitochondrial ROS regulator, to rheumatoid arthritis (RA) pathogenesis. However, its exact mechanism remains elusive given the conflicting evidence about its specific function. We used a four-level integrative framework combining multi-omics data and literature‑supported mechanistic inference. At the genetic level, Mendelian randomization (MR) was performed to explore potential causal relationships between ROMO1, IL2RA, HLA-DR, MIF, and RA risk, followed by differential expression analysis and machine learning-based feature selection to identify key mROS genes. The temporal expression dynamics of ROMO1 were assessed in RA progression. At the cellular and tissue levels, we integrated single-cell RNA sequencing and spatial transcriptomics to map cell-type-specific expression and synovial localization of ROMO1-related immune cells and pathways. Finally, our multi-omics findings were contextualized with literature-supported mechanistic inference. (1) MR results were consistent with a potential protective effect of ROMO1 on RA (OR = 0.52) and its potential regulation of risk factors IL2RA (OR = 0.46) and HLA-DR (OR = 0.40). Conversely, IL2RA (OR = 1.42), HLA-DR (OR = 1.88), and MIF (OR = 1.17) were positively associated with RA risk. Additionally, ROMO1 was identified as a top candidate diagnostic predictor with stage-specific dynamics: downregulated in the early but upregulated in the late/remission stages. (2) Single-cell RNA sequencing showed ROMO1's cell-specific expression in CD14+ HLA-DR+ CD74+ monocytes and CD4+ IL2RA+ T cells. Cell communication analysis further suggested that these cells may participate in MIF pathway regulation. Spatial transcriptomics subsequently identified that ROMO1-related cells localized to synovial pathological regions, with MIF pathway changes correlated with RA progression. (3) Finally, literature-supported mechanistic inference suggests that ROMO1 may modulate mROS levels to promote anti-inflammatory M2 macrophage polarization, which could theoretically contribute to reduced systemic inflammation and the alleviation of multi-organ decline in RA. This integrated multi-omics investigation, supported by literature-based mechanistic inference, suggests ROMO1 as a stage-dependent biomarker candidate and potential immune regulator in RA.

Humans↗

Clinical trial design for microarray predictive marker discovery and assessment.

Transcriptional profiling technologies that simultaneously measure the expression of thousands of mRNA species represent a powerful new clinical research tool. Similar to previous laboratory analytical methods including immunohistochemistry, PCR and in situ hybridization, this new technology may also find its niche in routine diagnostics. Outcome predictors discovered by these methods may be quite different from previous single-gene markers. These novel tests will probably combine the information embedded in the expression of multiple genes with mathematical prediction algorithms to formulate classification rules and predict outcome. The performance of machine learning-algorithm-based diagnostic tests may improve as they are trained on larger and larger sets of samples, and several generations of tests with improving accuracy may be introduced sequentially. Several gene-expression profiling-technology platforms are mature enough for clinical testing. The most important next step that is needed for further progress is the development and validation of multigene predictors in prospectively designed clinical trials to determine the true accuracy and clinical value of this new technology. This manuscript reviews methodological and statistical issues relevant to clinical trial design to discover and validate multigene predictors of response to therapy.

Clinical Trials as Topic↗

Predicting protein--protein interactions from primary structure.

MOTIVATION: An ambitious goal of proteomics is to elucidate the structure, interactions and functions of all proteins within cells and organisms. The expectation is that this will provide a fuller appreciation of cellular processes and networks at the protein level, ultimately leading to a better understanding of disease mechanisms and suggesting new means for intervention. This paper addresses the question: can protein-protein interactions be predicted directly from primary structure and associated data? Using a diverse database of known protein interactions, a Support Vector Machine (SVM) learning system was trained to recognize and predict interactions based solely on primary structure and associated physicochemical properties. RESULTS: Inductive accuracy of the trained system, defined here as the percentage of correct protein interaction predictions for previously unseen test sets, averaged 80% for the ensemble of statistical experiments. Future proteomics studies may benefit from this research by proceeding directly from the automated identification of a cell's gene products to prediction of protein interaction pairs.

Artificial Intelligence↗

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation↗

Hierarchical metabolic engineering for rewiring cellular metabolism.

Metabolic engineering is a key enabling technology for rewiring cellular metabolism to enhance production of chemicals, biofuels, and materials from renewable resources. However, how to make cells into efficient factories is still challenging due to its robust metabolic networks. To open this door, metabolic engineering has realized great breakthroughs through three waves of technological research and innovations, especially the third wave. To understand the third wave of metabolic engineering better, we discuss its mainstream strategies and examples of its application at five hierarchies, including part, pathway, network, genome, and cell level, and provide insights as to how to rewire cellular metabolism in the context of maximizing product titer, yield, and productivity. Finally, we highlight future perspectives on metabolic engineering for the successful development of cell factories.

Metabolic Engineering↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans↗

Transcriptomic analysis identifies novel ferroptosis-related biomarkers and therapeutic targets in pulmonary arterial hypertension.

BACKGROUND: Ferroptosis plays a significant role in pulmonary arterial hypertension (PAH), although its underlying mechanisms and key pathogenic genes remain unclear. METHODS: Transcriptomic data from human PAH and control lung tissue were obtained from the Gene Expression Omnibus (GEO) database, whereas ferroptosis-related genes (FRGs) were sourced from the MsigDb and FerrDb databases. Differentially expressed FRGs (DE-FRGs) were identified through the intersection of FRGs with differentially expressed genes (DEGs). Functional enrichment analysis was performed using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Key hub genes were identified through Least Absolute Shrinkage and Selection Operator (LASSO), support vector machine-recursive feature elimination (SVM-RFE), and weighted correlation network analysis (WGCNA). Gene set enrichment analysis (GSEA) was conducted to explore the functional roles and associated pathways of hub genes. The relationship between hub genes and immune infiltration was investigated. Expression levels of potential biomarkers were validated via Quantitative real-time polymerase chain reaction (qRT-PCR) and immunohistochemistry (IHC) in two PAH animal models (monocrotaline-induced and Sugen5416 plus hypoxia-induced PAH). Finally, molecular docking was employed to screen potential therapeutic compounds. RESULTS: A total of 133 DE-FRGs were identified, with KEGG and GO analyses highlighting their involvement in intracellular iron homeostasis and ferroptosis. Hub genes, notably FZD7 and NFE2, were identified using LASSO, SVM-RFE, and WGCNA. Immune infiltration analysis suggested that monocytes and neutrophils play key roles in PAH pathogenesis. Validation in PAH animal models showed significant upregulation of Fzd7 and downregulation of Nfe2 in lung tissues of both MCT- and SuHx-induced PAH models. Molecular docking identified tetrachlorodibenzodioxin (TCDD) has good binding affinity. CONCLUSION: In summary, we investigated two ferroptosis-related biomarkers, FZD7 and NFE2, in PAH using transcriptomics, offering new insights into molecular mechanisms and potential targeted therapies for the disease.

Ferroptosis↗

Synthetic community Hi-C benchmarking provides a baseline for virus-host inferences.

Microbiomes influence diverse ecosystems, and viruses increasingly appear to impose key constraints. While viromics has expanded genomic catalogs, host identification for these viruses remains challenging due to the limitations in scaling cultivation-based approaches and the uncertain reliability and relative low resolution of in silico predictions - particularly for understudied viral taxa. Towards this, Hi-C proximity ligation uses sequenced, cross-linked virus and host genomic fragments to infer virus-host linkages and has now been applied in at least ten studies. However, its accuracy remains unknown. Here we assess Hi-C performance in recovering virus-host interactions using synthetic communities (SynComs) composed of four marine bacterial strains and nine phages with known interactions and then apply optimized bioinformatic protocols to natural soil samples. In SynComs, standard Hi-C sample preparations and analyses showed poor normalized contact score performance (26% specificity, 100% sensitivity, incorrect matches up to class level) that could be dramatically improved by Z-score filtering (Z ≥ 0.5, 99% specificity), though at reduced sensitivity (62% down from 100%). Detection limits were established as reproducibility was poor below minimal phage abundances of 105 PFU/mL. Applying optimized bioinformatic protocols to natural soil samples, we compared virus-host linkages inferred from proximity-ligated Hi-C sequencing with predictions generated by in silico homology-based and machine learning-based bioinformatic approaches. Prior to Z-score thresholding, agreement was relatively high at the phylum to family levels (72%), but not at the genus (43%) or species (15%) levels. Z-score thresholding reduced sensitivity (only 34% of predictions were retained), with only modest improvements in congruence with bioinformatic methods (48% or 18% at genus or species levels, respectively). Regardless, this led to 79 genus-level-congruent virus-host linkages and 293 new ones revealed by Hi-C alone - i.e., providing many new virus-host interactions to explore in already well-studied climate-critical soils. Overall, these findings provide empirical benchmarks and methodological guidelines to improve the accuracy and reliability of Hi-C for virus-host linkage studies in complex microbial communities.

Genomics↗