Search PubMedSearch

SEARCH · Search PubMed

Results for “clinical decision support Intelligent medicine”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,936 records · Page 2Linked to original sources

The future of precision oncology and artificial intelligence in Belgium: scenarios and policy responses.

PURPOSE: Precision medicine, also known as personalized medicine, enables the provision of tailored health services to patients. In the prevention, early detection, and treatment of cancers, precision medicine is highly promising, given the increasing use of genomic profiling for diagnosis and adapting therapies in several tumor types. Artificial Intelligence (AI) can support this process by analyzing vast amounts of relevant data. However, high-quality data and financial investments in the health system are essential for the implementation of precision medicine and AI solutions in routine cancer care. DESIGN/METHODOLOGY/APPROACH: Building on the quantitative outcomes of a foresight exercise published in another study, this article collects qualitative data to gain more detailed insights into the future of precision oncology in Belgium and discusses the role of AI in this field. It reports the results of a series of expert workshops, focusing on four hypothetical future scenarios that are centered around technological and economic issues that must be overcome for the widespread use of precision oncology in Belgium. FINDINGS: The study concludes that all four scenarios discussed in the workshops would require supportive policy measures in Belgium, which should go beyond mere technological and economic considerations, such as involving patient associations and the public in policy design or creating multi-disciplinary expert groups for precision medicine. ORIGINALITY/VALUE: To the best of our knowledge, this is the first study to employ foresight methodology to illustrate possible future scenarios, scrutinize feasible approaches for implementing precision oncology in Belgium, and discuss the use of AI in this context.

Belgium

Applications of artificial intelligence in robot-assisted surgery: a systematic review.

To characterize applications of artificial intelligence (AI) in robot-assisted surgery, summarize technical and clinical performance, and assess the quality of the available evidence. PubMed, Web of Science Core Collection, and Scopus were searched for English-language journal articles published from 1 January 2020 through 31 October 2025. Randomized, observational, model-development, validation, and feasibility studies evaluating AI in robot-assisted surgery or closely related image-guided minimally invasive workflows were eligible. Two reviewers independently performed study selection, data extraction, and risk-of-bias assessment. Owing to heterogeneity in surgical procedures, AI tasks, analytical units, validation strategies, and outcomes, findings were synthesized descriptively without statistical pooling. The review was registered in the International Prospective Register of Systematic Reviews (CRD420251175699). Seventeen studies were included: seven clinical prediction or decision-support studies, eight intraoperative recognition, segmentation, or image-guided studies, and two training or workflow studies. Five prediction studies reported area-under-the-curve values of 0.74-0.95. Technical studies reported F1 or Dice scores of 0.525-0.995 and task-specific accuracies of 0.840-0.998. Two randomized studies suggested benefits for personalized suturing feedback and automated camera control, but neither established improved patient outcomes. Only one study had low overall risk of bias; the remaining studies were at high or unclear risk or raised some concerns. AI applications in robot-assisted surgery show promise for prediction, intraoperative perception, training, and workflow support. Evidence primarily demonstrates technical feasibility rather than established clinical effectiveness. Independent multicenter validation and prospective evaluation of patient, educational, and workflow outcomes are required before widespread implementation.

Robotic Surgical Procedures

Pharmacogenomic and drug interactions risk in cardio-oncology: A precision medicine perspective for India.

Cardio-oncology patients may face complex treatment regimens due to the concurrent existence of cancer and cardiovascular disease, leading to a considerable polypharmacy burden. This significantly increases the prospect of drug-drug interactions (DDIs) and gene-drug interactions. The majority of these interactions arise from comparable pharmacokinetic and pharmacological pathways associated with drug transporters and cytochrome P450 enzymes. The significance of pharmacogenomics in tailored treatment strategies are emphasised by the fact that genetic variability enhances individual differences in drug response, safety, and efficacy. This narrative review focus on the effects of key genetic polymorphisms (e.g., DPYD, CYP2C19, and CYP2C9) on the metabolism and efficacy of commonly prescribed anticancer and cardiovascular medications such as fluoropyrimidines, clopidogrel, and warfarin. In addition it explore the role of pharmacogenomic variants on drug-drug interactions within the field of cardio-oncology. The study ultimately emphasizes the necessity of precision medicine in India to address the genetic diversity and underrepresentation in global genomic databases. The absence of pharmacogenomic testing, infrastructural deficiencies, financial constraints, and insufficient clinical integration hinder the widespread use of this technology in India. The Genome India Project and other national initiatives establish the foundation for pharmacogenomic-guided therapy. Utilizing genetic data, together with artificial intelligence-based predictive tools, for clinical decision-making may enhance medication safety and yield optimal outcomes in Indian cardio-oncology patients.

Humans

Artificial intelligence-supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs.

BACKGROUND: Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers. Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways. PURPOSE: To synthesize prospective or program-embedded evaluations of AI conducted within European-style population screening programs and to estimate exploratory program-level absolute risk differences (RDs) per 1000 examinations for cancer detection rate (CDR) and recall. MATERIALS AND METHODS: We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation). Outcomes were harmonized as AI-control RDs per 1000 examinations. Random-effects pooling used Hartung-Knapp-Sidik-Jonkman models. For the paired-reader design, sensitivity analyses applied a Kish effective sample-size approach across plausible within-examination correlations (ρ = 0.3-0.8). Positive predictive value (PPV) and workflow/time outcomes were summarized descriptively. RESULTS: Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I2 ≈ 12%), consistent with a modest directional increase with borderline statistical uncertainty. The pooled recall RD was -0.6 per 1000 (95% CI -3.1 to +2.1; I2 ≈ 41-43%), indicating no consistent recall increase across screening programs. Where reported, PPV was higher with AI-supported screening. Efficiency signals included 44.3% fewer total readings in MASAI and shorter reading times for AI-normal examinations in PRAIM; in PRAIM, a program-level safety-net mechanism recovered 204 cancers that would otherwise have been missed. CONCLUSION: In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals. These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution.

Humans

Meniscal preservation in the age of biologics: toward a quantitative decision algorithm for personalized repair.

BACKGROUND: Despite advances in arthroscopic repair and biologic augmentation, surgical indication for meniscal tears remains heterogeneous. No standardized framework currently integrates biomechanical, clinical, and biological determinants to guide repair versus resection. PURPOSE: To develop a quantitative decision model-the Meniscal Preservation Score (MPS)-that unifies biomechanical and biological evidence to stratify reparability potential and standardize treatment selection in meniscal surgery. METHODS: A systematic evidence synthesis conducted in accordance with PRISMA 2020 reporting standards of studies published from 2000 to 2025 in PubMed, Embase, and Scopus identified key determinants of meniscal healing. Five consistent predictors-patient age, vascularity, tear morphology, associated pathology, and activity profile-were weighted through a two-round modified Delphi consensus among ten experienced knee surgeons. The resulting 0-9-point MPS was incorporated into a stepwise decision tree linking lesion morphology, biological context, and surgical strategy. Conceptual validation used 50 simulated cases and a retrospective cohort of 45 patients to test agreement between algorithm recommendations and expert surgical decisions. RESULTS: The MPS achieved 86% concordance with expert judgment in simulation and 84% agreement in clinical validation. In this retrospective exploratory cohort, cases in which surgical management was concordant with MPS recommendations demonstrated higher mean IKDC scores at 24 months and lower observed reoperation rates. These findings should be interpreted as associative rather than causal, as treatment allocation was not controlled and discordant cases may have represented inherently more complex pathology. CONCLUSION: The MPS represents an evidence-informed decision-support framework designed to systematize reparability assessment. While exploratory analyses suggest structural coherence with expert reasoning, prospective implementation and external validation are required before clinical adoption as a predictive tool. LEVEL OF EVIDENCE: conceptual model with exploratory validation.

Humans

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

ORBIT: Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space for cancer driver gene identification.

Accurate identification of cancer driver genes is crucial for precision oncology but remains challenging due to the complexity of integrating heterogeneous data and modeling dynamic biological systems. To address these limitations, we propose ORBIT (Oncogenic Representation Learning via Bi-Prototype Contrastive Learning in Hyperbolic Space). Our framework synergistically fuses multi-omics profiles with functional network data using a context-adaptive graph reweighting mechanism to capture cancer-specific dynamics. The model employs a bi-prototype contrastive learning strategy within hyperbolic space, which aligns gene representations around distinct driver and non-driver semantic anchors while preserving the intrinsic hierarchy of biological networks. Comprehensive evaluations demonstrate that ORBIT achieves highly competitive stability in pan-cancer analysis while consistently outperforming state-of-the-art methods in cancer-specific predictions. Furthermore, functional enrichment analysis confirms that the model effectively segregates core cancer pathways, and drug sensitivity profiling validates the clinical relevance of the identified drivers. By integrating hyperbolic geometry with context-adaptive learning, ORBIT offers a robust and interpretable paradigm for precision medicine. The source codes and datasets are publicly accessible at https://github.com/spcho-dev/ORBIT.

Humans

Automated CEAP Classification of Venous Duplex Reports Using Multimodal Artificial Intelligence.

OBJECTIVE: To develop and internally validate a prototype multimodal artificial intelligence system for automated CEAP (Clinical, Etiological, Anatomical and Pathophysiological) classification of venous duplex ultrasound (VDUS) reports, integrating natural language processing of free-text components with computer vision analysis of hand-drawn anatomical diagrams. METHODS: Single centre retrospective observational study using routinely collected clinical data. One thousand consecutive venous duplex ultrasound reports from Cambridge University Hospitals NHS Foundation Trust, UK (July 2024 - May 2025) were labelled according to the CEAP classification, excluding the Etiological component, which could not be reliably determined from duplex reports alone. Transfer learning was applied using ClinicalBERT for text and MobileNetV3 for diagrammatic data. Clinical classes were predicted from request line text. Text- and image-based pathophysiological models were developed for four anatomical territories (Great Saphenous Vein, Small Saphenous Vein, Deep system, Perforators), combined using late fusion with probability averaging. RESULTS: The clinical CEAP model achieved accuracy of 0.91, macro-F1 of 0.82, and macro-AUC of 0.98. Pathophysiological prediction varied, with text models broadly outperforming image models. Fusion yielded heterogeneous benefits, improving SSV performance but reducing Deep system accuracy. The performance of the final pathophysiological CEAP fusion models varied across anatomical territories: accuracy ranged from 0.70-0.92 and macro-AUC from 0.80-0.92. CONCLUSION: This study demonstrates the feasibility of automated CEAP classification from VDUS reports. Despite class imbalance affecting minority class predictions, the strong discriminatory performance validates this multimodal ML model for extracting clinically meaningful information from real-world data. This approach offers potential, pending external validation, to streamline vascular services through automated triage and guideline-compliant decision making.

Artificial intelligence

Clinical applications of digital twin technology in In Vitro Fertilisation.

BACKGROUND: Digital twin technology, originating from aerospace and manufacturing industries, has emerged as a transformative tool in healthcare. In vitro fertilisation (IVF) faces persistent challenges including suboptimal embryo selection, unpredictable treatment outcomes, and limited personalisation of protocols. Despite advances in assisted reproductive technology, existing literature exhibits fragmentation: artificial intelligence applications in embryo selection, ovarian stimulation, and endometrial assessment have been developed independently without systematic integration into comprehensive treatment frameworks. Digital twin technology offers unprecedented opportunities to create virtual replicas of biological systems, enabling real-time monitoring, predictive modelling, and personalised treatment strategies. AIM: This narrative review aims to critically examine the current applications of digital twin technology in IVF, evaluate its potential benefits and limitations, synthesize existing evidence into an integrative conceptual model, and identify future directions for implementation in reproductive medicine. METHOD: A comprehensive narrative review was conducted using PubMed, Scopus, Web of Science, and IEEE Xplore databases. A narrative review approach was selected over systematic review to accommodate the heterogeneity of evidence types in this emerging field, including theoretical frameworks, simulation studies, and proof-of-concept implementations that would be excluded from systematic reviews. Search terms included "digital twin," "IVF," "in vitro fertilisation," "assisted reproductive technology," "embryo selection," and "predictive modelling." Studies published between 2015 and 2025 were included, focusing on original research articles, systematic reviews, and proof-of-concept studies describing digital twin applications in reproductive medicine. RESULTS: Digital twin technology in IVF demonstrates significant potential across multiple domains including embryo development simulation, ovarian response prediction, endometrial receptivity modelling, and personalised stimulation protocols. Current applications integrate artificial intelligence, machine learning algorithms, time-lapse imaging, and omics data to create comprehensive virtual models. Early evidence suggests improvements in embryo selection accuracy, ovarian response prediction, and treatment protocol optimization, though large-scale randomized controlled trials remain limited. Implementation challenges include data integration complexity, computational requirements, regulatory considerations, and validation requirements. CONCLUSION: Digital twin technology represents a paradigm shift in IVF practice, offering personalised, predictive, and precision medicine approaches. This review synthesizes existing evidence to propose an integrative conceptual model for digital twin implementation across the IVF treatment spectrum, identifies critical knowledge gaps, and establishes research priorities to advance clinical translation. Despite current limitations, continued advancement promises improved success rates and patient outcomes.

Humans

Cost-Effectiveness and the Economics of Genomic Testing and Molecularly Matched Therapies.

Cost-effectiveness analysis of precision oncology can help guide value-driven care. Next-generation sequencing is increasingly cost-efficient over single gene testing because diagnostic algorithms require multiple individual gene tests to determine biomarker status. Matched targeted therapy is often not cost-effective due to the high cost associated with drug treatment. However, genomic profiling can promote cost-effective care by identifying patients who are unlikely to benefit from therapy. Additional applications of genomic profiling such as universal testing for hereditary cancer syndromes and germline testing in patients with cancer may represent cost-effective approaches compared with traditional history-based diagnostic methods.

Humans

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

Long-term microbiome and clinical effects of a microbiome-guided personalized diet versus low-FODMAP diet in irritable bowel syndrome: A 12-month follow-up randomized controlled trial.

Dietary therapy is central to irritable bowel syndrome (IBS) management, yet the long-term durability of the low-FODMAP diet (LFD), and of microbiome-guided personalization, remains unclear. We assessed the long-term clinical and gut-microbiome effects of a microbiome-guided personalized diet (PD) compared with a standard LFD in adults meeting Rome IV criteria for IBS. In this multicenter, open-label randomized controlled trial with blinded outcome assessment, participants who completed a 6-week dietary intervention (PD or LFD) were followed at 6 and 12 months without further dietary intervention. Outcomes included the IBS Severity Scoring System (IBS-SSS), IBS Quality of Life (IBS-QOL), and the Hospital Anxiety and Depression Scale (HADS); gut microbiota were profiled by 16S rRNA sequencing. Longitudinal changes were evaluated using linear mixed-effects models, responder analyses, PERMANOVA, and PERMDISP. Both diets reduced IBS-SSS at 6 weeks. PD maintained symptom improvement at 6 and 12 months (-82.0 and -78.3 points from baseline), whereas LFD benefits regressed by 12 months (+29.3 points; between-group p&#x2009;=&#x2009;0.001). At 12 months, IBS-SSS responder rates were higher with PD than LFD (62.5% vs 34.5%; absolute risk difference&#x2009;+28.0%, 95% CI 4.2-47.7; Fisher p&#x2009;=&#x2009;0.029), and IBS-QOL, HADS-anxiety, and HADS-depression showed more favourable trajectories with PD. PD was associated with sustained Shannon alpha-diversity gains (+0.488 at 6 weeks;&#x2009;+0.205 at 12 months; both p&#x2009;<&#x2009;0.01). A modest between-group beta-diversity difference at 6 months (R2&#x2009;=&#x2009;0.035; p&#x2009;=&#x2009;0.011) was not significant at 12 months. This hypothesis-generating follow-up suggests more durable benefit with PD; larger trials powered for long-term clinical and microbiome outcomes are warranted.

Humans

Criteria for Safe Hospital Discharge in Bronchiolitis: A Systematic Review.

Bronchiolitis is the leading cause of hospital presentation and admission for infants in Australasia. We aimed to synthesise current evidence on the effect of discharge criteria for infants (aged <&#x2009;12&#x2009;months) who are presenting to or are admitted to hospital with bronchiolitis, to inform a binational guideline recommendation update. Systematic searches were conducted on MEDLINE, EMBASE, PubMed, Cochrane Library and CINAHL (last search 19 February 2025) for non-randomised studies evaluating hospital discharge criteria in bronchiolitis. The primary outcomes were length of stay (LOS) and readmission rates. The risk of bias (ROBINS-I) and certainty of the evidence (GRADE) were appraised, and findings were narratively synthesised. GRADE evidence-to-decision methodology, expert consensus voting and interest-holder consultation were used to finalise the recommendation update. Two retrospective observational studies were included (N&#x2009;=&#x2009;2697) (low to very low quality), reporting on unique discharge criteria. In both studies, use of the discharge criteria was associated with a significant reduction in LOS relative to alternative protocols. There was no significant difference in readmission rates observed in either study. There was low to very low certainty evidence across outcomes due to risk of bias, indirectness and imprecision. The review findings informed a recommendation update for safe discharge criteria in the 2025 Australasian Bronchiolitis Guideline update. Updated, prescriptive discharge criteria and flow chart were developed, covering clinical stability, oxygen saturation/support, feeding difficulties, caregiver confidence and education on deterioration, social factors and follow-up. The revised criteria provide clinicians with increased certainty in decision-making in bronchiolitis, albeit with further research needed.

Humans

Toward personalized interventions for preventing depression in primary care: Qualitative and quantitative findings from the e-predictD pilot study.

BACKGROUND: The predictD intervention, delivered by family physicians (FPs), has demonstrated effectiveness and cost-efficiency in preventing depression and anxiety. The e-predictD study aims to design, develop, and evaluate a novel personalized intervention for depression prevention by integrating information and communication technologies (ICTs), risk prediction algorithms, and decision support systems (DSS) for both patients and FPs. OBJECTIVE: To evaluate the satisfaction, usability, and acceptability, of a beta version of the e-predictD intervention in primary care settings. METHODS: The e-predictD intervention follows a biopsychosocial approach, including an initial patient-FP interview, specific FP training, and an app. A &#x3b2;-version was tested in a pilot study without a control group over three months. The app integrates a validated depression risk prediction algorithm, decision algorithms, and a monitoring system supporting the DSS. The DSS generates a personalized prevention plan (PPP) from eight intervention modules: physical exercise, social relationships, problem-solving, communication skills, decision-making, assertiveness, sleep improvement, and cognitive restructuring. Patients and FPs discussed the PPP in a 15-minute baseline interview, selecting modules for implementation over three months. Semi-structured interviews gathered feedback. Assessments included depression (PHQ-9), anxiety (GAD-7), quality of life (SF-12), and major depression risk (predictD algorithm). RESULTS: Six FPs from six Spanish cities enrolled 56 non-depressed patients at moderate-to-high risk of depression; 47 (84%) completed follow-up. The app was used for a median of six days (interquartile range: 1-30). Both FPs and patients expressed satisfaction, leading to incorporated improvements. After three months, significant reductions in major depression risk and anxiety symptoms were observed, alongside improved mental quality of life. However, no significant changes were found in depressive symptoms or physical quality of life. CONCLUSION: This pilot study supports the feasibility and acceptability of the e-predictD &#x3b2;-version, despite lower-than-expected app usability. Health improvements were observed, warranting confirmation in a randomized controlled trial. TRIAL REGISTRATION: ClinicalTrials.gov NCT03990792.

Adult

Improving Community-Based Care for Adolescents with ADHD: a Randomized Controlled Trial of Artificial Intelligence-Assisted Fidelity Supports.

Cognitive-behavioral treatments (CBTs) for adolescents with ADHD demonstrate promise of long-term effects on outcome. However, their implementation in routine care community clinics faces barriers that impact quantity, efficiency, and quality of delivery, as well as client outcomes. This study is a randomized controlled trial designed to evaluate the impact of an AI-assisted service delivery model on therapist implementation of Supporting Teens' Autonomy Daily (STAND), a CBT blended with Motivational Interviewing (MI) for adolescents with ADHD. Adolescents with ADHD (N&#x2009;=&#x2009;51), who were clients at three community mental health agencies, received treatment from 23 therapists. There was randomization of adolescents and therapists to AI-assisted or standard implementation supports. In addition to standard supports (i.e., training, standard facilitation resources, technical assistance, case supervision), AI-assisted support package included digitized facilitation resources housed in a clinical dashboard (Care4), feedback on content fidelity, and AI-generated feedback on MI implementation quality. The AI-assisted group was associated with more efficient treatment delivery and lower number of appointments attended by the adolescent. There was also a significant decrement in MI quality over time in the AI-assisted group compared to the standard support group. Feedback in focus groups indicated that therapists perceived a task-oriented mindset to be associated with receipt of the AI-assisted support package, leading therapists to prioritize efficiency over relational aspects of therapy. Following the results of this trial, a future, larger RCT should examine the impact of the AI-assisted implementation model on mental health outcomes and cost savings to organizations, third party payers, and clients. Trial registration number: NCT05135065; https://www.clinicaltrials.gov ; Registered September 2021.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

The Role of Artificial Intelligence for Intimate Partner Violence Prevention: A Systematic Review.

INTRODUCTION: Intimate partner violence (IPV), encompassing physical, sexual, emotional and economic abuse, remains a pervasive global health concern. Traditional prevention efforts face obstacles such as underreporting, delayed detection and limited personalised support. Emerging artificial intelligence (AI) approaches offer new opportunities to enhance IPV prevention. AIM: This systematic review maps and synthesises evidence on AI-driven tools in IPV prevention based on studies published between 2004 and 2024. METHODS: Following PRISMA 2020 guidelines and PROSPERO registration, we searched PubMed, Embase, CINAHL, PsycINFO, IEEE Xplore and Web of Science. Eligible studies explicitly evaluated AI technologies targeting IPV prediction, screening, intervention or support delivery. Study quality was appraised using the Mixed Methods Appraisal Tool (MMAT). RESULTS: Of 1304 records initially identified, 41 studies met eligibility criteria. AI applications ranged from machine learning (ML) for risk prediction and natural language processing (NLP) for IPV detection in clinical and social media data, to image analysis for forensic evaluation and chatbot-based support. Predictive modelling demonstrated strong discriminative performance, while NLP-based screening detected IPV with notable sensitivity. Chatbots showed feasibility and user acceptability, but evidence of their direct impact on reducing IPV incidence was limited, with one randomised controlled trial showing a modest reduction. Key challenges identified included algorithmic bias, data privacy risks and barriers to integration across health and social care systems. DISCUSSION: AI-informed interventions show promise for improving IPV detection, risk assessment, and scalable support, but questions remain about long-term effectiveness, ethical fairness, transparency and equitable implementation. Future interdisciplinary research should address these concerns to responsibly deploy AI in IPV prevention. RELEVANCE TO CLINICAL PRACTICE: The findings highlight the importance of trauma-informed, culturally responsive care and provider training in AI applications. Nurse-led innovation and policy advocacy will be crucial for safe, equitable integration of AI in IPV prevention.

Artificial Intelligence

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans