Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

KAVAS-2: Knowledge Acquisition, Visualization and Assessment System.

The objective of KAVAS-2 is the development of a tool, named KAVIAR, with which domain experts can make their knowledge explicit. It contains components for (computer assisted) knowledge elicitation and for machine learning. A key issue in KAVAS is the assessment of the quality of the classification and domain models built. Various quality measures are available and implemented in KAVIAR to assess the quality of models, specifically those developed from data bases by machine learning techniques.

Computer Simulation

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

Representation for discovery of protein motifs.

There are several dimensions and levels of complexity in which information on protein motifs may be available. For example, one-dimensional sequence motifs may be associated with secondary structure identifiers. Alternatively, three-dimensional information on polypeptide segments may be used to induce prototypical three-dimensional structure templates. This paper surveys various representations encountered in the protein motif discovery literature. Many of the representations are based on incompatible semantics, making difficult the comparison and combination of previous results. To make better use of machine learning techniques and to provide for an integrated knowledge representation framework, a general representation language--in which all types of motifs can be encoded and given a uniform semantics--is required. In this paper we propose such a model, called a spatial description logic, and present a machine learning approach based on the model.

Amino Acids

Integrating single-cell transcriptomics to construct an oncogene-driven prognostic model and elucidate metabolic-immune crosstalk in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a leading cause of cancer-related deaths, its progression and treatment heterogeneity are mainly influenced by driver gene and tumor micro-environment (TME) interactions. Nevertheless, the mechanisms of this process at the single-cell level remain unclear. This study integrated TCGA and multi-center single-cell transcriptome data to identify a 575 genes HCC-specific core set, developing a single-cell "oncogene scoring" system to quantify individual carcinogenic activity. This score is significantly elevated in malignant and proliferative T cells and is closely associated with metabolic reprogramming, aberrant cell‒cell communication, and immunosuppressive phenotypes. Based on these characteristics, we constructed a machine learning-based Random Survival Forest (RSF) prognostic model validated in multiple independent cohorts, which classifies patients into distinct risk subtypes. The high-risk group exhibits genomic instability, increased tumor stemness, and immune evasion, while the low-risk group was more sensitive to drugs such as sorafenib. This study highlights the potential pathways by which high oncogenic activity is associated with HCC progression, suggesting a profound link with single-cell metabolic‒immune crosstalk. The constructed RSF model offers a promising computational framework for risk stratification and provides hypothesis-generating insights that may inform future personalized treatment strategies for HCC patients.

Hepatocellular carcinoma

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

Proteomic and machine learning analysis predicts treatment response signatures in Myasthenia Gravis.

BACKGROUND: Myasthenia gravis (MG) is a prototypical antibody-mediated autoimmune disease with variable treatment responses with a need for biomarkers to guide therapeutic decision making. Proteomic profiling, coupled with machine learning, offers a hypothesis-free approach to identify multi-protein signatures associated with treatment response. METHODS: We analyzed sera collected at entry (baseline) from participants in a phase 3 trial randomized trial comparing thymectomy plus prednisone versus prednisone alone, along with matched controls using liquid chromatography-mass spectrometry. We derived disease-specific proteomic signatures and evaluated associations between baseline proteins and 6-month clinical outcomes using multiple machine-learning approaches with internal validation. RESULTS: Baseline serum proteomes distinguished MG from controls, with pathway enrichment implicating complement activation, immunoglobulin production, and T-cell receptor signaling. Distinct protein panels predicted 6-month clinical improvement within each treatment arm. In the thymectomy-plus-prednisone group, models captured non-linear relationships of predictive proteins in contrast with the predominant additive patterns observed in the prednisone-alone group. Predictive proteins were enriched for T-cell signaling and leukocyte trafficking functions, providing insight into treatment-specific biology. CONCLUSIONS: Baseline serum proteomics captures core disease characteristics of MG and predicts short-term clinical response in a treatment-specific manner. While our results require validation in independent cohorts, these findings could enable biomarker-guided selection of thymectomy, refine risk stratification, and furnish mechanistic readouts for future MG trials and clinical care. We aim to conduct future studies using -omic approaches to validate these baseline predictive biomarkers and pathways of treatment response in patients with MG.

Adult

Developing a machine learning-based prognosis and immunotherapeutic response signature in colorectal cancer: insights from ferroptosis, fatty acid dynamics, and the tumor microenvironment.

INSTRUCTION: Colorectal cancer (CRC) poses a challenge to public health and is characterized by a high incidence rate. This study explored the relationship between ferroptosis and fatty acid metabolism in the tumor microenvironment (TME) of patients with CRC to identify how these interactions impact the prognosis and effectiveness of immunotherapy, focusing on patient outcomes and the potential for predicting treatment response. METHODS: Using datasets from multiple cohorts, including The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO), we conducted an in-depth multi-omics study to uncover the relationship between ferroptosis regulators and fatty acid metabolism in CRC. Through unsupervised clustering, we discovered unique patterns that link ferroptosis and fatty acid metabolism, and further investigated them in the context of immune cell infiltration and pathway analysis. We developed the FeFAMscore, a prognostic model created using a combination of machine learning algorithms, and assessed its predictive power for patient outcomes and responsiveness to treatment. The FeFAMscore signature expression level was confirmed using RT-PCR, and ACAA2 progression in cancer was further verified. RESULTS: This study revealed significant correlations between ferroptosis regulators and fatty acid metabolism-related genes with respect to tumor progression. Three distinct patient clusters with varied prognoses and immune cell infiltration were identified. The FeFAMscore demonstrated superior prognostic accuracy over existing models, with a C-index of 0.689 in the training cohort and values ranging from 0.648 to 0.720 in four independent validation cohorts. It also responses to immunotherapy and chemotherapy, indicating a sensitive response of special therapies (e.g., anti-PD-1, anti-CTLA4, osimertinib) in high FeFAMscore patients. CONCLUSION: Ferroptosis regulators and fatty acid metabolism-related genes not only enhance immune activation, but also contribute to immune escape. Thus, the FeFAMscore, a novel prognostic tool, is promising for predicting both the prognosis and efficacy of immunotherapeutic strategies in patients with CRC.

Ferroptosis

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-α/NF-κB, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Boolean matrix logic programming for active learning of gene functions in genome-scale metabolic network models.

Reasoning about hypotheses and updating knowledge through empirical observations are central to scientific discovery. In this work, we applied logic-based machine learning methods to drive biological discovery by guiding experimentation. Genome-scale metabolic network models (GEMs) - comprehensive representations of metabolic genes and reactions - are widely used to evaluate genetic engineering of biological systems. However, GEMs often fail to accurately predict the behaviour of genetically engineered cells, primarily due to incomplete annotations of gene interactions. The task of learning the intricate genetic interactions within GEMs presents computational and empirical challenges. To efficiently predict using GEM, we describe a novel approach called Boolean Matrix Logic Programming (BMLP) by leveraging Boolean matrices to evaluate large logic programs. We developed a new system, [Formula: see text], which guides cost-effective experimentation and uses interpretable logic programs to encode a state-of-the-art GEM of a model bacterial organism. Notably, [Formula: see text] successfully learned the interaction between a gene pair with fewer training examples than random experimentation, overcoming the increase in experimental design space. [Formula: see text] enables rapid optimisation of metabolic models to reliably engineer biological systems for producing useful compounds. It offers a realistic approach to creating a self-driving lab for biological discovery, which would then facilitate microbial engineering for practical applications.

Active learning

Essence: A benchmarking-validated transformer framework for early diagnosis of Parkinson's disease using cerebrospinal fluid protein biomarkers.

Parkinson's disease (PD) is a progressive neurodegenerative disorder characterized by motor and non-motor symptoms. The lack of objective molecular biomarkers limits early diagnosis and personalized treatment. Here, we propose Essence, a benchmarking-validated framework integrating cerebrospinal fluid (CSF) proteomics with traditional and deep learning models to identify robust protein signatures for PD. Using data from two independent cohorts, 1266 high-confidence proteins are quantified, among which 178 exhibit differential abundance between PD and healthy controls (HC). Through systematic benchmarking of ten machine learning algorithms and four neural architectures, the Transformer model consistently outperforms alternatives across multiple feature selection strategies, achieving an area under the receiver operating characteristic curve (AUC) of 1.0000 with only 35 features. Functional analyses of the top-ranked 35 proteins reveal enrichment in neuroinflammatory, synaptic, and oxidative stress-related pathways. Importantly, spatial transcriptomic profiling based on the Allen Brain Atlas shows region-specific expression of these biomarkers in PD-relevant brain structures, including the striatum, subthalamic nucleus, hippocampus, and white matter tracts. This anatomical alignment supports the functional relevance of the identified markers and highlights their potential utility in early-stage diagnosis and mechanistic understanding of PD.

Benchmarking

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans

The Continuity Trap in Data Science Health Research.

Secondary use is now the ordinary condition of data science health research rather than an exception to it. Electronic health records collected for clinical care become prediction tools and inputs for generative AI; imaging archives become foundation-model corpora; genomic datasets become resources for polygenic risk scores; and legacy biospecimens become renewable, indefinitely distributable cell lines. Governance has responded by emphasizing verifiable instruments such as provenance logs, repository approvals, broad-consent forms, data-use agreements, model cards, records of processing, and locality-preserving architectures. These instruments are necessary, and they answer real questions about lineage, privacy, institutional responsibility, and accountability, but they are not sufficient to establish that a present use remains ethically justified. We define ethical continuity as the persistence of normatively relevant relationships between the original conditions of data generation or material collection and subsequent downstream uses, such that current uses remain justifiable in light of the expectations, permissions, meanings, and relational obligations present at entrustment. We then define the Continuity Trap as a review-stage governance error in which a salient signal of continuity in one domain is treated as sufficient evidence of ethical continuity overall, causing inquiry into the remaining domains to close prematurely. The trap is not ordinary noncompliance, ethics creep, or a demand for universal rereview; it is a cross-domain inference error that can arise even in careful, good-faith review. We distinguish it from proxy closure, of which it is a continuity-specific subtype, and from Goodhart's and Campbell's laws, which describe how measures degrade once they become targets. We operationalize ethical continuity across 4 domains: provenance, semantics, authorization, and relational standing, developed in our Representational Veracity framework, and we show that these domains can diverge as data are linked, transformed, modeled, and redeployed. We identify the institutional mechanisms-provenance privilege, descriptor sedimentation, authorization fossilization, and community effacement-that cause auditable signals to be overread, and we examine how the US Health Insurance Portability and Accountability Act (HIPAA) of 1996, the General Data Protection Regulation, the European Health Data Space, US Food and Drug Administration guidance, the US National Institute of Standards and Technology (NIST) AI Risk Management Framework, and federated-learning governance can reduce risk while still inducing continuity traps. We apply the framework to consent and nonconsent settings, including public health, immunization, syndromic, and wastewater surveillance, polygenic risk scores, induced pluripotent stem cells, federated learning, and health-related large language models. The policy implication is trigger-based continuity review: rather than rereviewing every reuse, investigators and reviewers should identify the weakest continuity domain at the present data stage and impose a domain-matched safeguard, recorded in a short continuity statement. This reframing is intended for the committees, repositories, funders, and governance bodies that decide whether reuse may proceed, and it matters most in cross-border and low-resource settings. Provenance should begin ethical review; it should not end it.

Data Science

ProMeta: a meta-learning framework for robust disease diagnosis and prediction from plasma proteomics.

MOTIVATION: The plasma proteome offers a dynamic window of human health, capturing the real-time intersections between genetics and physiology. However, the application of deep learning to proteomics is currently hindered by a reliance on large-scale labeled datasets, rendering standard models ineffective for rare or novel diseases where patient samples are inherently scarce. RESULTS: Here, we present ProMeta, a meta-learning framework designed to enable robust disease modeling under extreme data restrictions. By integrating knowledge-guided pathway encoding with bi-level meta-optimization, ProMeta projects unstructured proteomic profiles into biologically interpretable functional tokens. This architecture allows the model to learn a global initialization containing transferable biological priors from biobank-scale data, facilitating rapid adaptation to novel tasks. Through comprehensive benchmark experiments, ProMeta consistently outperformed transfer learning and traditional machine learning baselines in both disease diagnosis and prediction tasks. In the most challenging 4-shot scenarios (utilizing only 2 cases and 2 controls), the model achieved robust generalization with an average AUROC of &#x223c;0.69, representing a 24.6% relative improvement over the best-performing baseline methods. Mechanistic investigation revealed that ProMeta disentangles cases from controls in the latent space prior to task-specific adaptation, confirming the acquisition of universal biological rules rather than rote memorization. Furthermore, gradient-based interpretation identified disease-specific protein biomarkers and functional pathways consistent with known pathophysiology. Collectively, ProMeta overcomes the data-scarcity bottleneck in precision medicine, providing a scalable, interpretable framework for characterizing the full spectrum of human diseases, particularly for rare conditions lacking extensive clinical cohorts. AVAILABILITY AND IMPLEMENTATION: The source code of ProMeta is available at GitHub (https://github.com/lihan97/ProMeta).

Proteomics

Water source, latrine type, and rainfall are associated with detection of non-optimal and enteric bacteria in the vaginal microbiome: a prospective observational cohort study nested within a cluster randomized controlled trial.

BACKGROUND: Less than one-third of sub-Saharan Africans have access to improved water sources. In US, Indian, and African studies, Bacterial vaginosis (BV) is increased among women with poor water, sanitation, and hygiene (WASH). We examined water source, sanitation (latrine type), and rainfall in relation to the vaginal microbiome (VMB). METHODS: In a cluster randomized controlled trial of menstrual cups and cash transfer, we measured the impact of cups on VMB via 16S rRNA gene amplicon sequencing in a subset of 436 adolescent girls. We analyzed how self-reported water source and latrine type at home related to VMB over 18-months, examining community state type I (CST-I, L. crispatus dominant) vs. other CST; alpha diversity; targeted taxa (coliform and other water-related pathogens); and non-targeted taxa via machine learning approaches. Mixed effects multivariable longitudinal models were adjusted for intervention arm, age, socioeconomic status, sexual activity, and cluster-level school WASH and rainfall (in millimeters). RESULTS: Adjusting for all covariates in all models: (1) the odds of CST-I were increased among participants with piped water (vs. pond), and decreased with traditional pit latrine vs. flush toilet. (2) Alpha diversity varied by water source and latrine type without consistent trends. (3) Coliform bacteria relative abundance (RA) was higher among participants with traditional pit or ventilated improved pit latrines vs. flush toilet, and higher among participants relying on stream vs. pond water. Streptococcus agalactiae RA was higher among participants with non-flush toilets, while Bacteroides fragilis RA was lower with non-flush toilets. (4) Key taxa from non-targeted analyses associated with water source and latrine type included typical vaginal bacteria, opportunistic pathogens, and urinary tract pathobionts. (6) Increased rainfall was associated with decreased odds of CST-I. TRIAL REGISTRATION: ClinicalTrials.gov NCT03051789, February 14, 2017.

Adolescent