Search PubMedSearch

SEARCH · Search PubMed

Results for “PROBAST”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5 recordsLinked to original sources

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95 % CI 0.85-0.94; 95 % prediction interval 0.62-0.98), with sensitivity of 0.80 (95 % CI 0.77-0.83) and specificity of 0.87 (95 % CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Risk prediction models for blood transfusion in patients undergoing total hip and knee arthroplasty: a systematic review and meta-analysis.

OBJECTIVE: To systematically review and evaluate published risk prediction models for perioperative blood transfusion in patients undergoing total hip or knee arthroplasty (THA/TKA). METHODS: We systematically searched PubMed, Web of Science, the Cochrane Library, and Embase from inception to May 31, 2025. Two researchers independently screened the literature, extracted data, and assessed the risk of bias and applicability using the Prediction model Risk Of Bias Assessment Tool (PROBAST). The area under the receiver operating characteristic curve (AUC) values were pooled via a meta-analysis using Stata 18.0. RESULTS: d Fourteen studies containing 36 prediction models were included. The incidence of blood transfusion among THA/TKA patients ranged from 3.2% to 30.8%. Preoperative hemoglobin (Hb) level, tranexamic acid (TXA) use, operative duration, intraoperative blood loss, and age were the most frequently incorporated predictors. Model sensitivity ranged from 58% to 94.5%, and specificity ranged from 71.3% to 94%. Meta-analysis showed that the pooled AUC value of the 13 validated models was 0.87 (95% CI: 0.85-0.90), suggesting good discriminatory performance. All models were rated as having a high risk of bias. The applicability of four studies was rated as unclear. CONCLUSION: Although the included studies demonstrated promising discriminative ability of prediction models for blood transfusion in THA/TKA, all were assessed as having a high risk of bias using the PROBAST tool. Therefore, future research should prioritize the development of models with larger sample sizes, rigorous study designs, and multicenter external validation.

Humans

Artificial intelligence-derived myocardial fibrosis on cardiac magnetic resonance for prognosis in cardiomyopathy: A systematic review of a sparse evidence base.

BACKGROUND: Myocardial fibrosis on cardiovascular magnetic resonance (CMR), assessed by late gadolinium enhancement (LGE) and parametric mapping, is an established predictor of adverse events in cardiomyopathy. We assessed whether artificial intelligence (AI) quantification of fibrosis adds independent prognostic value. METHODS: We searched six databases, a clinical-trials register, and a preprint server from inception to 13 June 2026. Eligible studies used AI to generate a fibrosis marker in adults with ischemic or nonischemic cardiomyopathy, with covariate-adjusted outcomes over ≥12 months. Risk of bias was assessed using PROBAST, PROBAST+AI, and QUIPS. Fewer than three comparable studies precluded meta-analysis; certainty was rated using GRADE. RESULTS: Of 448 records (381 after de-duplication), 18 full texts were reviewed and two included, one peer-reviewed and one preprint. In an ischemic-cardiomyopathy registry (Ghanbari et al.; n = 216 analytic, 26 events), AI-derived dense LGE scar predicted arrhythmic events (univariable hazard ratio [HR] 2.35, 95% CI 1.33-4.15), and AI-derived but not manual scar improved discrimination beyond guideline criteria (area under the curve 0.63 to 0.68; p = 0.02). In a nonischemic dilated-cardiomyopathy preprint (Kim et al.; n = 347, 119 events), automated extracellular volume ≥30% predicted cardiovascular death or heart-failure hospitalization (adjusted HR 2.00, 95% CI 1.32-3.03). Both were at high risk of bias, with data-derived thresholds and no external validation. CONCLUSIONS: Across only two studies, AI-derived fibrosis was independently associated with adverse cardiovascular events, but its added value over manual quantification remains unproven. Certainty was very low. The evidence base is sparse and not yet ready for clinical use.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning