Search PubMedSearch

SEARCH · Search PubMed

Results for “explainable AI in healthcare”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,957 records · Page 2Linked to original sources

Artificial intelligence in genitourinary oncology: publication trends and systematic review.

OBJECTIVE: To conduct an analysis of publication trends and a systematic review of randomized controlled trials (RCTs) to characterize the current state of artificial intelligence (AI) use in genitourinary (GU) oncology, as AI has emerged as a transformative tool in healthcare with potential applications in diagnostics, treatment planning, and prognostication. METHODS: We searched the Medical Literature Analysis and Retrieval System Online (MEDLINE), Excerpta Medica dataBASE (EMBASE; Ovid), and Cumulative Index to Nursing and Allied Health Literature (CINAHL) Ultimate for studies related to AI and GU oncology, excluding non-English papers, non-human studies, review articles, and articles using AI solely for manuscript writing. Publication trends were analysed from 2013 to 2023 and categorized by study design and cancer type. RCTs were evaluated through systematic review using Covidence (Veritas Health Innovation Ltd, Melbourne, Victoria, Australia) for screening and data extraction. Two reviewers independently assessed all studies, with risk of bias (RoB) evaluated using the Cochrane RoB 2.0 tool. RESULTS: Of 2409 articles identified, 1220 met inclusion criteria. These included 962 retrospective articles, 175 prospective studies, 79 studies with combined retrospective/prospective methods, and four RCTs. Studies most commonly addressed prostate (n = 923), renal (n = 274), and urothelial (n = 194) cancers. Publications grew from 14 in 2013 to 362 in 2023, with substantial acceleration in 2019. Four RCTs were identified - one in urothelial cancer and three in prostate cancer. Two RCTs evaluated AI-based diagnostics, demonstrating improved performance over conventional methods; the remaining two RCTs evaluated AI in prognostication and treatment planning, showing improved gains in imaging interpretation and operational efficiency. RoB varied across studies, primarily related to randomisation and deviations from intended interventions. CONCLUSIONS: Artificial intelligence research in GU oncology has grown, although high-level evidence from RCTs remains limited. Existing trials underscore AI's promise in diagnostics, prognostication, and treatment planning, and the rapidly evolving nature of this field warrants continued prospective investigation.

Humans

The impact of artificial intelligence on critical thinking and clinical reasoning in health professions education: A systematic review and meta-analysis.

BACKGROUND: Critical thinking and clinical reasoning underpin healthcare professionals' ability to navigate uncertainties and deliver safe and effective care. With artificial intelligence (AI) advancement and growing adoption, AI-based educational tools are increasingly used to support these cognitive competencies' development. OBJECTIVE: To synthesize randomised and controlled clinical trials on AI-based educational tools in health professions education and examine their effects on critical thinking and clinical reasoning among health professions students. METHODS: Six electronic databases were searched from January 1, 2014 to July 28, 2025 was reviewed: PubMed, Cochrane Central Register of Controlled Trials, CINAHL, Scopus, Embase and Web of Science. Two independent reviewers performed data extraction and quality assessment using standardized JBI checklists. The GRADE approach was used to assess the certainty of evidence. Studies were pooled via random-effects meta-analyses or narrative syntheses. RESULTS: Fourteen randomised controlled trials and seven controlled clinical trials were included (n = 21). Meta-analyses revealed small to medium effect sizes for the surrogate clinical reasoning outcomes of performance-based assessment scores (SMD 0.68; 95% CI [0.38, 0.98], p-value = 0.00; I2 = 38%) and knowledge test scores (SMD 0.39; 95% CI [0.09, 0.69], p-value = 0.01; I2 = 79%). Critical thinking and clinical reasoning skills and dispositions were narratively synthesized, with majority of included studies favouring AI-based interventions but the evidence had low to very low certainty. CONCLUSION: AI-based educational interventions may improve critical thinking and clinical reasoning among health profession students, but the evidence is very uncertain. This review offers preliminary insights but does not allow identification of optimal interventions or discipline-specific recommendations due to small sample sizes and substantial intervention heterogeneity. Further research is required to draw definitive conclusions. PROTOCOL REGISTRATION: CRD42025634074.

Humans

Artificial Intelligence Technologies in Nursing Clinical Decision-Making: An Umbrella Review.

AIM: To describe contemporary peer-reviewed literature on artificial intelligence in nurses' clinical decision-making. METHODS: An umbrella review of literature reviews. DATA SOURCES: Four major databases were searched for reviews published between 2019 and 2024. RESULTS: Sixteen literature reviews reported on 965 nursing artificial intelligence primary studies. The studies focused on technology development and emerging performance evaluations, whilst real-world testing or implementation in nursing clinical settings was rare. Rigorous comparative analyses were lacking. While artificial intelligence demonstrates promise in decision-making, challenges such as a lack of controlled studies, algorithmic bias, limited reproducibility and insufficient clinical trials hinder its practical impact. Ethical concerns, transparency and patient data privacy issues pose barriers to AI integration in nursing practice. Ethical and legal guidelines for patient privacy are needed and should be taught along with AI literacy training for nurses. CONCLUSIONS: Artificial intelligence has the potential to enhance clinical nursing decision-making, although evidence is limited by too few examples of nurse participation during development. Underutilisation in administrative nursing functions hinders implementation. Nurses should assume a central role in the design and development of AI applications to ensure that these technologies address the realities of nursing practice. With such improvements, artificial intelligence can transform nursing practice, improve nurses' clinical decision-making and ultimately enhance consumer healthcare outcomes. PATIENT OR PUBLIC INVOLVEMENT: No Patient or Public Involvement. REPORTING METHOD: While there is no reporting checklist for umbrella reviews, the PRISMA guide for systematic reviews was followed.

Artificial Intelligence

Artificial intelligence enabled social robotic interventions (PARO) in Australian dementia care: A systematic review and meta-analysis.

BACKGROUND: Although there is a growing body of research indicating that Personal Robot/Social Robot could be used in various aspects of care for individuals with dementia, little is known about how well these types of interventions work in an actual hospital setting in Australia. AIMS & OBJECTIVES: The objective of the present systematic review and meta-analysis is to assess the effectiveness of PARO-based socially assistive robotic intervention in terms of its effectiveness outcomes towards the reduction of dementia-related behavioural and psychological symptoms in Australian based healthcare settings. METHODS: A systematic search was conducted across five electronic databases, including MEDLINE (PubMed), EMBASE, CINAHL, PsycINFO, and the Cochrane Library, to identify randomised controlled trials (RCTs) investigating PARO-based socially assistive robotic interventions for dementia in Australian healthcare settings. This review was registered with PROSPERO (CRD420251251916) and followed the PRISMA 2020 guidelines. In addition, the Cochrane Risk of Bias tool (RoB 2) was used to evaluate the risk of bias across all studies. Pooled standardised mean differences (SMD) with 95 % confidence intervals (CI) were calculated for agitation, anxiety, and depression. Heterogeneity across studies was evaluated using the I2 statistic. RESULTS: Six RCTs involving 1444 participants were identified for inclusion in this review. AI-enabled socially assistive robotic interventions, specifically the PARO therapeutic robot, significantly reduced agitation and anxiety when compared to standard treatment or control conditions. The pooled analysis showed that agitation [SMD = -0.44 (95 % CI: -0.70, -0.18) p = 0.0008] and anxiety [SMD = -0.59 (95 % CI: -0.91, -0.27) p = 0.0003] were reduced significantly, while the decrease in depression [SMD = -0.44 (95 % CI: -0.95, -0.07) p = 0.09] scores was non-significant among dementia patients receiving PARO-based socially assistive robotic interventions as compared to the control. The overall risk of bias across all six studies was considered low to moderate. CONCLUSION: PARO-based socially assistive robotic interventions may provide preliminary evidence of effectiveness in reducing agitation and anxiety in individuals with dementia in Australian healthcare, but the evidence regarding the reduction of depression remains unclear. Therefore, additional high-quality trials with consistent methodology and extended follow-up will be necessary to determine both the short-term and long-term clinical efficacy and practicality of implementing these interventions into practice.

Humans

Factors associated with successful integration of pharmacists into residential aged care teams: A qualitative study.

BACKGROUND: Australia's Aged Care Onsite Pharmacist program aims to support quality use of medicines in residential aged care homes. This is a novel role introduced into existing teams in a complex environment. Factors associated with successful integration since implementation are currently unknown. AIM: This study aims to explore the perspectives of pharmacists and other stakeholders within aged care homes regarding successful integration of the novel aged care pharmacist service into healthcare teams. METHODS: A qualitative approach, using interpretive descriptive methodology, was used to explore perspectives. Semi-structured focus groups and interviews with pharmacists, nursing and care staff, allied health professionals, general practitioners, residents, and family members were undertaken. Data were collected via Zoom™, audio- and video-recorded, and transcribed verbatim. Two researchers undertook inductive thematic analysis to identify key themes. RESULTS: 30 participants across focus groups, focus-group interviews, interviews, and member-checking processes contributed. An overarching theme of proactivity and showing a genuine interest in others underpinned three key themes. Theme 1: Pharmacists needed to be seen, through physical presence and availability, as well as developing a distinct identity. Theme 2: Pharmacists needed to build trust, through collaboration in real time and demonstrating value. Theme 3: Pharmacists needed to develop an understanding of the aged care home environment, including social and contextual norms, as well as procedures, routines and roles. DISCUSSION: This research complements existing understandings of interprofessional collaboration and teamwork amongst healthcare professionals. Themes were interlinked; we used a sensitising framework, social cognitive theory, to present and explain the findings and interactions that can support pharmacist integration into existing teams. Aged care services should structure onboarding to prioritise early visibility, clarify roles and organisational needs, and foster in-person collaboration. Pharmacists should demonstrate proactivity and an authentic interest in all staff, residents and families.

Humans

Identifying and Prioritizing Core Components of Relationship Education Programs: a Case Study of an Artificial Intelligence (AI) Assisted Systematic Review.

The field of prevention science seeks to identify and implement effective strategies to address social, emotional, and health challenges. A critical aspect of this endeavor is determining the core components of prevention programs that drive positive outcomes. This article presents a case study utilizing artificial intelligence (AI)-assisted systematic review methods to identify key components of healthy marriage and relationship education programs. Given the growing body of research in this domain, AI tools offer a promising means to enhance the efficiency and accuracy of literature reviews. This study employed AI to screen, code, and validate research articles, demonstrating its effectiveness in expediting systematic reviews while maintaining high accuracy in inclusion screening. This case study involved a systematic review of 22,028 resources (identified from PsycINFO, Academic Search Ultimate, and Google) and a final data set of 268 relevant studies. AI screening was integral in effectively conducting multiple rounds of screening. However, findings also highlight challenges in AI-assisted qualitative data abstraction, underscoring the continued need for human expertise in complex coding tasks. The study contributes to the ongoing discourse on integrating AI into prevention science methodologies and offers insights for optimizing AI applications in systematic reviews.

Artificial Intelligence

Psychological consequences of AI-assisted training and the buffering role of mindfulness.

The integration of artificial intelligence (AI) into athletic training is accelerating, yet its psychological implications for athletes remain insufficiently understood. Drawing on the transactional model of stress and the stress-buffering framework of mindfulness, this study examined whether mindfulness training can mitigate adverse psychological responses associated with AI-assisted training. Using a randomized controlled factorial design, 160 collegiate athletes were assigned to AI-assisted training or standard training, with or without concurrent mindfulness intervention, and assessed at baseline, week 4, and week 8. Athletes exposed to AI-assisted training without psychological support exhibited increases in perceived stress and AI dependence over time. In contrast, these stress increases were substantially attenuated when mindfulness training was implemented alongside AI-assisted training. A significant AI × Mindfulness × Time interaction emerged for perceived stress at post-intervention, and difference-in-differences analyses corroborated a robust buffering effect. Mediation analyses further indicated that mindfulness training reduced stress partially through enhancing mindful awareness; a three-wave cross-lagged analysis showed that mindful awareness and stress were reciprocally related over time, with the hypothesized awareness-to-stress pathway remaining robust. Together, these findings suggest that AI-assisted training introduces a distinct form of evaluative pressure, and that mindfulness training may serve as an effective psychological buffer during the adoption of continuous algorithmic performance evaluation systems.

Humans

AI echo INSIGHT study: A prospective blinded randomized trial of artificial intelligence echocardiogram interpretation.

BACKGROUND: Transthoracic echocardiography (TTE) is the most commonly performed cardiac imaging modality with over 30 million studies annually. Demand for timely expert interpretation continues to outpace capacity, creating diagnostic delays and inter-observer variability that impact patient care. Recent research has suggested computer vision artificial intelligence (AI) models can generate accurate preliminary comprehensive TTE reports, however, prospective evaluation is needed to determine whether AI-assisted TTE interpretation can improve clinician efficiency while preserving diagnostic accuracy. METHODS: AI ECHO INSIGHT is a prospective randomized blinded clinical trial conducted at Kaiser Permanente Northern California that will evaluate 1200 historical TTE studies (1000 consecutive unselected studies plus 200 with moderate or greater valvular disease) interpreted using three workflows: (1) AI-generated preliminary report finalized by a blinded cardiologist (AI-assisted); (2) cardiologist-generated preliminary report finalized by a blinded cardiologist (cardiologist-assisted); and (3) sonographer-generated preliminary report finalized by a blinded cardiologist (sonographer-assisted). The primary outcome is the rate of substantial change between preliminary and final reports, comparing the AI-assisted workflow to the pooled cardiologist-assisted and sonographer-assisted workflows. Secondary outcomes include cardiologist interpretation time for report finalization, superiority testing for diagnostic accuracy, and reporting consistency. CONCLUSION: AI ECHO INSIGHT is a prospective randomized blinded clinical trial evaluating the clinical impact of AI-assisted TTE interpretation on diagnostic accuracy, cardiologist efficiency, and reporting consistency in real-world echocardiography workflows. TRIAL REGISTRATION: ClinicalTrials.gov registration number NCT07229300.

Humans

AI Health message intervention: The role of message customization and message source in breast cancer screening among women of color.

OBJECTIVES: To examine the effectiveness of breast cancer screening messages with varying levels of customization (generic, targeted, and tailored) and to compare AI-generated versus human-generated messages. METHODS: A between-subjects experimental design with a control condition was employed. Message content followed a standardized structure and varied by level of customization: generic, targeted (demographic-based), and tailored (perceived susceptibility- and barrier-based). Messages were developed by either the authors or GenAI (ChatGPT-4o). A total of 391 participants recruited via Prolific were randomly assigned to five groups (generic, targeted-human, targeted-AI, tailored-human, and tailored-AI). Self-efficacy, behavioral intentions, attitudes, and message believability were measured using different scales. RESULTS: Customized (tailoring and targeting) health messages performed comparably to generic messages in shaping positive health outcomes. GenAI-generated messages also produced outcomes comparable to those of human-generated messages under standardized conditions. Significant negative indirect effects through message believability for the human-tailored condition was found relative to the generic condition. CONCLUSIONS: GenAI may be a useful tool for developing and customizing scalable health messages. Its effectiveness depends not only on customization but also on maintaining message quality, including readability, clarity, coherence, naturalness, and credibility. PRACTICAL IMPLICATIONS: GenAI may support health practitioners in developing customized and scalable breast cancer messages. However, professional review remains necessary to ensure that the message is culturally appropriate, responsive to patient concerns, and suitable for use alongside patient-provider communication.

Humans

AI-enabled viral genomics: from virus discovery to host prediction and emerging variant forecasting.

The rapid expansion of metagenomic sequencing has generated vast repositories of viral sequence data that far outpace our capacity to interpret them using conventional approaches. Highly divergent sequences, sparse functional annotation, and taxonomically uneven sampling present fundamental challenges for reference-dependent methods, which lose sensitivity precisely for novel and understudied viruses with high public health relevance. Artificial intelligence (AI) provides a new avenue to address these challenges by enabling predictive inference from viral genomes and proteins while reducing dependence on sequence similarity. In this Review, we discuss representative advances in AI for virus discovery, taxonomic classification and functional annotation, prediction of host range and zoonotic potential, and efforts toward forecasting emerging variants. These advances are transforming viral genomics from a largely descriptive discipline into one with increasing predictive capability. We also critically assess the major challenges that constrain current approaches, including the availability of high-quality and representative datasets, rigorous model evaluation, biological interpretability and responsible governance for increasingly capable AI models.

Artificial Intelligence

Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis.

BACKGROUND: Emergency radiographic interpretation for fractures is prone to missed or misdiagnoses. Artificial intelligence (AI) is expected to become a powerful tool to assist clinicians in fracture detection. PURPOSE: A systematic review and meta-analysis was performed to assess whether AI improves clinicians' ability to detect fractures on radiographs. MATERIALS AND METHODS: A literature search was conducted in PubMed, Web of Science, and Cochrane Library for studies published between January 1, 2010, and October 10, 2025. A meta-analysis of diagnostic accuracy studies was performed using a Summary Receiver Operating Characteristic (SROC) curve. The quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Subgroup analysis and meta-regression were conducted to explore potential sources of heterogeneity. RESULTS: A total of 26 studies were included . The pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94). The corresponding AUC values were 0.90 (95% CI: 0.87-0.92) before and 0.95 (95% CI: 0.93-0.97) after AI assistance. Eight studies were rated as high risk of bias. Subgroup analysis and meta-regression identified potential sources of heterogeneity, including fracture location, AI model type, high risk of bias, and reference standards. CONCLUSION: AI assistance significantly improves clinicians' diagnostic performance in detecting fractures on radiographs for extremity and trunk fractures.

Humans

Challenges and future directions in AI-driven biomaterials for microbiome-associated oral infectious diseases: A systematic review.

Oral biofilm-induced antimicrobial resistance is the core pathogenic mechanism of microbiome-associated oral infectious diseases (dental caries, periodontitis, peri-implantitis, and endodontic infection). Traditional therapies and biomaterials are limited by poor biofilm penetration, drug resistance induction, single functionality, and inadequate adaptation to dynamic oral microenvironmental changes (e.g., pH fluctuations, salivary rinsing, masticatory stimulation). Artificial intelligence (AI) has transformed the field by integrating materials science, microbiology, and stomatology data. Via machine learning, deep learning, and multi-physics simulation, AI optimizes biomaterial physicochemical properties, decodes microenvironmental signals, constructs precise sensing-response loops, and supports the full chain of material design, performance prediction, and action simulation, advancing treatment from empirical intervention to precision regulation. This systematic review retrieved literature from PubMed, Embase, and Web of Science (January 2016-January 2026) using keywords across three dimensions: AI, biomaterials, and oral microbiome. Following inclusion/exclusion criteria, 99 articles were included. It elaborates on five core mechanisms of AI-driven oral biomaterials (precise oral microbiome analysis, targeted material design/optimization, performance prediction/simulation, targeted delivery/intervention, effect evaluation/dynamic regulation), analyzes their applications in microbiome-targeted biomaterial research and development (R&D) and clinical practice for the four major oral infectious diseases, addresses technical bottlenecks (insufficient targeting specificity and precision of biomaterials, poor stability and durability in complex oral microenvironments, inadequate biofilm disruption capacity, and clinical translation obstacles), and proposes future directions (multimodal design to enhance targeting specificity, structural and component optimization to improve stability/durability, development of multi-mechanism synergistic biofilm disruption strategies, strengthening translational research for clinical application, and deep integration of AI in the full chain of biomaterial R&D). This work provides comprehensive theoretical and practical support for the R&D, optimization, and clinical translation of AI-driven microbiome-targeted oral biomaterials.

Humans

AI-driven snapshot hyperspectral imaging for on-line sorting systems in food industry: From real-time sensing to intelligent decision-making.

High-throughput food sorting requires rapid, non-destructive detection of external defects, foreign materials, and internal quality attributes in heterogeneous food matrices. Conventional scanning hyperspectral imaging may suffer from motion-induced spatial-spectral mismatches, whereas snapshot hyperspectral imaging (S-HSI) captures spectral images within a single integration time. However, its advantage is limited by trade-offs in resolution, signal-to-noise ratio (SNR), reconstruction uncertainty, and calibration stability, which are further amplified by variable tissue structure, surface reflection, moisture, and fat distribution in foods. This review critically examines artificial intelligence (AI)-driven S-HSI for on-line food sorting within a sensing-representation-decision-execution framework. Compact architectures are compared according to their physical constraints, food-sorting suitability, and ability to support mapping between spectral responses and physicochemical quality attributes. AI strategies are reviewed for spectral reconstruction, image restoration, spatial-spectral representation, band selection, uncertainty-aware decision-making, and edge implementation. AI can partially compensate for snapshot-specific limitations, but current evidence remains largely limited to laboratory or prototype studies. Future work should link system performance to food safety and quality outcomes by reporting throughput, decision latency, calibration drift, missed-detection risk, false-rejection cost, and closed-loop sorting success.

Hyperspectral Imaging

An AI-assisted Clinical Decision Support System for Green Classification of Cystocele on Dynamic Transperineal Ultrasound.

Green classification of cystocele on dynamic transperineal ultrasound (TPUS) remains operator-dependent because it requires manual frame selection and landmark-based assessment of the Valsalva maneuver. We developed a workflow-oriented AI-assisted clinical decision support system for automated urethrovesical junction localization and dynamic Green classification and prospectively evaluated its standalone and reader-support performance. This diagnostic accuracy and reader study included 881 patients from a tertiary referral hospital, comprising a retrospective development cohort (n = 688) and an independent prospective test cohort (n = 193). A nested subset of 67 prospective patients was used for a reader study involving two junior and two intermediate radiologists under unaided and AI-assisted conditions. In the complete prospective test cohort, Green-AttGRU achieved a macro-averaged AUC of 0.939 (95% CI, 0.897-0.971) and an overall accuracy of 0.902 (95% CI, 0.860-0.943). In the reader study, overall accuracy increased from 0.761 to 0.821 without AI to 0.851-0.881 with AI, while macro-F1 increased from 0.660 to 0.777 to 0.820-0.860. Overall inter-reader agreement increased from a Fleiss' κ of 0.453 to 0.786, and pooled median interpretation time decreased from 26.7 s to 9.9 s. These findings support the preliminary feasibility of the system as a workflow-oriented decision-support tool for dynamic TPUS interpretation.

Humans

Manual, digital, and AI tumour-infiltrating lymphocyte scoring: a secondary analysis of the APHINITY randomised trial.

BACKGROUND: Stromal tumour-infiltrating lymphocytes (sTILs) are prognostic in early-stage HER2-positive breast cancer, but their role in the context of dual HER2 blockade remains undefined. We evaluated manual, digital, and artificial intelligence (AI)-based sTIL quantification, together with AI-derived spatial metrics, for prognostic and treatment-benefit stratification using tumour samples from the phase 3 APHINITY trial. METHODS: In the APHINITY trial, 4805 patients were randomly assigned to receive chemotherapy plus trastuzumab with pertuzumab or chemotherapy plus trastuzumab with placebo. Median follow-up was 74&#xb7;1 months (IQR 68&#xb7;3-75&#xb7;4). We analysed 4262 haematoxylin and eosin-stained images using manual assessment, an automated digital approach, AI-based lymphocyte quantification (AI percentage lymphocytes), and two AI-derived spatial features (AI-TIL and immune hotspot). Interobserver reproducibility was assessed in 262 randomly chosen tumour samples scored independently by five pathologists. Multivariable Cox models were used to assess associations between TIL levels and invasive disease-free survival (primary outcome in APHINITY), distant recurrence-free interval, and overall survival. The heterogeneity of pertuzumab benefit was evaluated using subgroup analyses, subpopulation treatment effect pattern plot analyses, and nested Cox models with treatment-by-biomarker interaction terms. FINDINGS: Manual scoring showed high interobserver reproducibility (intraclass correlation coefficient 0&#xb7;84 [95% CI 0&#xb7;79-0&#xb7;88]). Concordance between manual and automated methods was modest. AI-based scoring (AI percentage lymphocytes) reclassified 120 (11&#xb7;6%) of 1035 node-positive tumours from immune-low (by manual scoring) to immune-high; this subgroup of patients showed greater separation of 5-year invasive disease-free survival curves between pertuzumab and placebo groups compared with patients whose tumours were concordantly classified as immune-low by both manual and AI-based approaches. Higher levels of TILs were associated with improved invasive disease-free survival for all sTIL measurement approaches and spatial measurements (hazard ratios [HRs] 0&#xb7;41-0&#xb7;93). Pertuzumab was associated with improved invasive disease-free survival at higher sTIL levels across all measurement approaches (HRs 0&#xb7;36-0&#xb7;48), but was not associated with higher values of spatial measures. The largest 6-year absolute improvements with pertuzumab were observed in patients with node-positive disease whose tumours scored in the highest level of immune infiltration of manual sTIL scoring (&#x2265;70&#xb7;0%; mean absolute improvement 12&#xb7;1 percentage points [SD 2&#xb7;8]). In nested prognostic and predictive models, AI-based immune hotspot scores provided the most consistent additional information when combined with any sTIL measurement (all p<0&#xb7;010). INTERPRETATION: Standardised manual sTIL scoring was reproducible, and digital and AI-based methods showed consistent prognostic stratification and potential for treatment-benefit stratification despite only modest correlation between platforms. AI spatial metrics provided complementary information beyond sTIL density and could support more scalable immune assessment. Future studies are needed to validate these approaches in independent cohorts and to clarify their clinical utility for stratifying contemporary HER2-directed therapies. FUNDING: None.

Humans

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis.

BACKGROUND: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. OBJECTIVE: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. METHODS: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (&#x2265;18 years of age), reporting mean polyp detection counts stratified by size (&#x2264;5 mm, 6-9 mm, and &#x2265;10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and &#x3c4;2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. RESULTS: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (&#x2264;5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI -1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI -0.02 to 0.06, 95% PI -0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI -0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. CONCLUSIONS: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment-prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932.

Colonoscopy

User Engagement and Feature Preferences in an AI-Powered mHealth Intervention for Diabetes Prevention: Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Prediabetes is highly prevalent and increasing globally, yet lifestyle interventions remain underused. AI-driven mobile health (mHealth) tools can help scale diabetes prevention efforts, but the key factors driving their success are not well understood. OBJECTIVE: This post hoc secondary analysis of a randomized controlled trial (RCT) aimed to characterize the most valued features and the role of user engagement in outcomes of a fully automated mHealth intervention for diabetes prevention. METHODS: Data from 151 participants with prediabetes and overweight or obesity who were assigned to an AI-based diabetes prevention program (Sweetch) in a parent RCT (NCT05056376) were analyzed. Engagement (defined as the total number of days the app was used) was categorized into tertiles (low, medium, and high). Baseline characteristics were compared across engagement groups using ANOVA, Kruskal-Wallis, and chi-square tests, and regression models assessed the association between engagement and achievement of diabetes risk reduction outcomes (&#x2265;5% weight loss, &#x2265;4% weight loss with &#x2265;150 min/week of physical activity, or &#x2265;0.2 percentage point reduction in hemoglobin A1c [HbA1c] at 12 months). Perceived usefulness of intervention features was surveyed at 12 months. RESULTS: Median engagement was 98 (IQR 34-232) days. Older age (P<.001) and lower baseline BMI (P=.04) were significantly associated with higher engagement. Compared with low engagement, high engagement was associated with greater odds of achieving the composite diabetes risk reduction outcome (odds ratio [OR] 2.59, 95% CI 1.11-6.01; P=.03), &#x2265;5% weight loss (OR 3.31, 95% CI 1.16-9.42; P=.03), and &#x2265;0.2 percentage point reduction in HbA1c (OR 3.57, 95% CI 1.19-10.75; P=.02). Participants most frequently rated weight tracking, physical activity tracking, and the digital body weight scale as the features that were most helpful for achieving their health goals. CONCLUSIONS: Higher engagement with an AI-driven intervention requiring no human intervention was associated with improved diabetes risk reduction. Contrary to concerns about lower digital literacy, older adults engaged with the intervention more than younger adults. Features related to weight and physical activity tracking were most valued by patients in the program. TRIAL REGISTRATION: ClinicalTrials.gov NCT05056376; https://clinicaltrials.gov/study/NCT05056376.

Humans