Search PubMedSearch

SEARCH · Search PubMed

Results for “AI-assisted meta analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,623 recordsLinked to original sources

Identifying and Prioritizing Core Components of Relationship Education Programs: a Case Study of an Artificial Intelligence (AI) Assisted Systematic Review.

The field of prevention science seeks to identify and implement effective strategies to address social, emotional, and health challenges. A critical aspect of this endeavor is determining the core components of prevention programs that drive positive outcomes. This article presents a case study utilizing artificial intelligence (AI)-assisted systematic review methods to identify key components of healthy marriage and relationship education programs. Given the growing body of research in this domain, AI tools offer a promising means to enhance the efficiency and accuracy of literature reviews. This study employed AI to screen, code, and validate research articles, demonstrating its effectiveness in expediting systematic reviews while maintaining high accuracy in inclusion screening. This case study involved a systematic review of 22,028 resources (identified from PsycINFO, Academic Search Ultimate, and Google) and a final data set of 268 relevant studies. AI screening was integral in effectively conducting multiple rounds of screening. However, findings also highlight challenges in AI-assisted qualitative data abstraction, underscoring the continued need for human expertise in complex coding tasks. The study contributes to the ongoing discourse on integrating AI into prevention science methodologies and offers insights for optimizing AI applications in systematic reviews.

Artificial Intelligence

Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis.

BACKGROUND: Emergency radiographic interpretation for fractures is prone to missed or misdiagnoses. Artificial intelligence (AI) is expected to become a powerful tool to assist clinicians in fracture detection. PURPOSE: A systematic review and meta-analysis was performed to assess whether AI improves clinicians' ability to detect fractures on radiographs. MATERIALS AND METHODS: A literature search was conducted in PubMed, Web of Science, and Cochrane Library for studies published between January 1, 2010, and October 10, 2025. A meta-analysis of diagnostic accuracy studies was performed using a Summary Receiver Operating Characteristic (SROC) curve. The quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Subgroup analysis and meta-regression were conducted to explore potential sources of heterogeneity. RESULTS: A total of 26 studies were included . The pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94). The corresponding AUC values were 0.90 (95% CI: 0.87-0.92) before and 0.95 (95% CI: 0.93-0.97) after AI assistance. Eight studies were rated as high risk of bias. Subgroup analysis and meta-regression identified potential sources of heterogeneity, including fracture location, AI model type, high risk of bias, and reference standards. CONCLUSION: AI assistance significantly improves clinicians' diagnostic performance in detecting fractures on radiographs for extremity and trunk fractures.

Humans

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis.

BACKGROUND: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. OBJECTIVE: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. METHODS: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (≥18 years of age), reporting mean polyp detection counts stratified by size (≤5 mm, 6-9 mm, and ≥10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and τ2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. RESULTS: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (≤5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI -1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI -0.02 to 0.06, 95% PI -0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI -0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. CONCLUSIONS: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment-prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932.

Colonoscopy

Psychological consequences of AI-assisted training and the buffering role of mindfulness.

The integration of artificial intelligence (AI) into athletic training is accelerating, yet its psychological implications for athletes remain insufficiently understood. Drawing on the transactional model of stress and the stress-buffering framework of mindfulness, this study examined whether mindfulness training can mitigate adverse psychological responses associated with AI-assisted training. Using a randomized controlled factorial design, 160 collegiate athletes were assigned to AI-assisted training or standard training, with or without concurrent mindfulness intervention, and assessed at baseline, week 4, and week 8. Athletes exposed to AI-assisted training without psychological support exhibited increases in perceived stress and AI dependence over time. In contrast, these stress increases were substantially attenuated when mindfulness training was implemented alongside AI-assisted training. A significant AI × Mindfulness × Time interaction emerged for perceived stress at post-intervention, and difference-in-differences analyses corroborated a robust buffering effect. Mediation analyses further indicated that mindfulness training reduced stress partially through enhancing mindful awareness; a three-wave cross-lagged analysis showed that mindful awareness and stress were reciprocally related over time, with the hypothesized awareness-to-stress pathway remaining robust. Together, these findings suggest that AI-assisted training introduces a distinct form of evaluative pressure, and that mindfulness training may serve as an effective psychological buffer during the adoption of continuous algorithmic performance evaluation systems.

Humans

AI echo INSIGHT study: A prospective blinded randomized trial of artificial intelligence echocardiogram interpretation.

BACKGROUND: Transthoracic echocardiography (TTE) is the most commonly performed cardiac imaging modality with over 30 million studies annually. Demand for timely expert interpretation continues to outpace capacity, creating diagnostic delays and inter-observer variability that impact patient care. Recent research has suggested computer vision artificial intelligence (AI) models can generate accurate preliminary comprehensive TTE reports, however, prospective evaluation is needed to determine whether AI-assisted TTE interpretation can improve clinician efficiency while preserving diagnostic accuracy. METHODS: AI ECHO INSIGHT is a prospective randomized blinded clinical trial conducted at Kaiser Permanente Northern California that will evaluate 1200 historical TTE studies (1000 consecutive unselected studies plus 200 with moderate or greater valvular disease) interpreted using three workflows: (1) AI-generated preliminary report finalized by a blinded cardiologist (AI-assisted); (2) cardiologist-generated preliminary report finalized by a blinded cardiologist (cardiologist-assisted); and (3) sonographer-generated preliminary report finalized by a blinded cardiologist (sonographer-assisted). The primary outcome is the rate of substantial change between preliminary and final reports, comparing the AI-assisted workflow to the pooled cardiologist-assisted and sonographer-assisted workflows. Secondary outcomes include cardiologist interpretation time for report finalization, superiority testing for diagnostic accuracy, and reporting consistency. CONCLUSION: AI ECHO INSIGHT is a prospective randomized blinded clinical trial evaluating the clinical impact of AI-assisted TTE interpretation on diagnostic accuracy, cardiologist efficiency, and reporting consistency in real-world echocardiography workflows. TRIAL REGISTRATION: ClinicalTrials.gov registration number NCT07229300.

Humans

An AI-assisted Clinical Decision Support System for Green Classification of Cystocele on Dynamic Transperineal Ultrasound.

Green classification of cystocele on dynamic transperineal ultrasound (TPUS) remains operator-dependent because it requires manual frame selection and landmark-based assessment of the Valsalva maneuver. We developed a workflow-oriented AI-assisted clinical decision support system for automated urethrovesical junction localization and dynamic Green classification and prospectively evaluated its standalone and reader-support performance. This diagnostic accuracy and reader study included 881 patients from a tertiary referral hospital, comprising a retrospective development cohort (n = 688) and an independent prospective test cohort (n = 193). A nested subset of 67 prospective patients was used for a reader study involving two junior and two intermediate radiologists under unaided and AI-assisted conditions. In the complete prospective test cohort, Green-AttGRU achieved a macro-averaged AUC of 0.939 (95% CI, 0.897-0.971) and an overall accuracy of 0.902 (95% CI, 0.860-0.943). In the reader study, overall accuracy increased from 0.761 to 0.821 without AI to 0.851-0.881 with AI, while macro-F1 increased from 0.660 to 0.777 to 0.820-0.860. Overall inter-reader agreement increased from a Fleiss' κ of 0.453 to 0.786, and pooled median interpretation time decreased from 26.7 s to 9.9 s. These findings support the preliminary feasibility of the system as a workflow-oriented decision-support tool for dynamic TPUS interpretation.

Humans

Improving Community-Based Care for Adolescents with ADHD: a Randomized Controlled Trial of Artificial Intelligence-Assisted Fidelity Supports.

Cognitive-behavioral treatments (CBTs) for adolescents with ADHD demonstrate promise of long-term effects on outcome. However, their implementation in routine care community clinics faces barriers that impact quantity, efficiency, and quality of delivery, as well as client outcomes. This study is a randomized controlled trial designed to evaluate the impact of an AI-assisted service delivery model on therapist implementation of Supporting Teens' Autonomy Daily (STAND), a CBT blended with Motivational Interviewing (MI) for adolescents with ADHD. Adolescents with ADHD (N = 51), who were clients at three community mental health agencies, received treatment from 23 therapists. There was randomization of adolescents and therapists to AI-assisted or standard implementation supports. In addition to standard supports (i.e., training, standard facilitation resources, technical assistance, case supervision), AI-assisted support package included digitized facilitation resources housed in a clinical dashboard (Care4), feedback on content fidelity, and AI-generated feedback on MI implementation quality. The AI-assisted group was associated with more efficient treatment delivery and lower number of appointments attended by the adolescent. There was also a significant decrement in MI quality over time in the AI-assisted group compared to the standard support group. Feedback in focus groups indicated that therapists perceived a task-oriented mindset to be associated with receipt of the AI-assisted support package, leading therapists to prioritize efficiency over relational aspects of therapy. Following the results of this trial, a future, larger RCT should examine the impact of the AI-assisted implementation model on mental health outcomes and cost savings to organizations, third party payers, and clients. Trial registration number: NCT05135065; https://www.clinicaltrials.gov ; Registered September 2021.

Humans

Effectiveness and usability of artificial intelligence-powered assistive technologies in Supporting daily activities of children with cerebral palsy: a systematic review.

BACKGROUND: Cerebral Palsy (CP) is the main cause of motor disabilities in childhood, necessitating innovative approaches to rehabilitation and assistive technology (AT). Simultaneously, artificial intelligence (AI) is increasingly being integrated into devices to create more adaptive, personalized, and effective AT. This systematic review aimed to evaluate the effectiveness and usability of AI-powered assistive technologies designed to support daily activities and rehabilitation in children with CP. MATERIALS AND METHODS: Five databases, including Scopus, Web of Science, PubMed, Embase, and IEEE Xplore, were systematically searched, and 23 articles were included in the final analysis. Articles were identified, selected, and categorized into emerging thematic areas based on the primary function and application of the technology. RESULTS: Five key thematic topics were identified: 1) AI-driven motor rehabilitation and gait training for functional mobility; 2) intelligent assessment and monitoring systems for clinical decision support; 3) AI-supported communication, social interaction, and intention recognition tools; 4) gamified and virtual reality-based interventions to enhance engagement and usability; and 5) smart assistive systems supporting daily living and independent mobility. The findings demonstrate a strong trend toward the application of AI technologies in personalized, engaging, and data-driven interventions for children with CP. However, the field is predominantly in the proof-of-concept stage, with limitations including small sample sizes, lack of long-term clinical validation, challenges in user-centered design, and usability for children with CP. CONCLUSION: AI-powered assistive technologies hold significant potential for transforming the care of children with CP by enabling highly personalized and engaging interventions. To actualize this potential, future work must realize that practical application remains challenging owing to limited clinical validation, technological integration, and usability barriers for children with CP. Future research must prioritize user-centered design and multidisciplinary collaboration to ensure that AI and robotic advancements improve the usability and quality of life for children with CP.

Humans

Development and Crossover Evaluation of an Artificial Intelligence-Assisted System for Solid Pancreatic Lesion Detection and Pancreatic Parenchyma Recognition in Endoscopic Ultrasonography (With Video).

BACKGROUND AND STUDY AIMS: Pancreatobiliary endoscopic ultrasonography (EUS) is technically demanding, and supervised training opportunities are limited. We developed an artificial intelligence (AI) overlay system for detecting solid pancreatic lesions (SPL) and recognizing pancreatic parenchyma (PP) and evaluated its effect on reader performance. PATIENTS AND METHODS: Across six centers, two deep learning-based models were trained using expert-annotated EUS frames. We then conducted a randomized, two-sequence, two-period crossover reader study in which eight endosonographers (five novices and three experts) interpreted image sets with and without AI assistance. The primary endpoint was superiority of sensitivity for SPL detection among novices; key secondary endpoints included specificity and PP recognition. RESULTS: From 118 patients, 120 SPL-positive/negative image sets and 160 PP-positive/negative image sets were constructed. Among novices, AI assistance improved SPL detection sensitivity (88.7% vs. 76.8%, p&#x2009;<&#x2009;0.001) and accuracy (86.4% vs. 78.7%), while specificity met the predefined noninferiority criterion (84.2% vs. 80.5%, p&#x2009;<&#x2009;0.001). For PP recognition, sensitivity increased numerically (86.3% vs. 83.3%) but did not meet the predefined superiority criterion (p&#x2009;=&#x2009;0.095); specificity met the noninferiority criterion (87.8% vs. 81.0%), and accuracy increased from 82.1% to 87.0%. Among experts, sensitivity was maintained for both tasks, whereas specificity increased with AI assistance. CONCLUSIONS: AI assistance improved SPL detection among novice endosonographers. For PP recognition, sensitivity increased without reaching statistical superiority, whereas specificity met the predefined noninferiority criterion. These findings support a potential adjunctive role for AI in EUS interpretation.

Humans

Applications of artificial intelligence in robot-assisted surgery: a systematic review.

To characterize applications of artificial intelligence (AI) in robot-assisted surgery, summarize technical and clinical performance, and assess the quality of the available evidence. PubMed, Web of Science Core Collection, and Scopus were searched for English-language journal articles published from 1 January 2020 through 31 October 2025. Randomized, observational, model-development, validation, and feasibility studies evaluating AI in robot-assisted surgery or closely related image-guided minimally invasive workflows were eligible. Two reviewers independently performed study selection, data extraction, and risk-of-bias assessment. Owing to heterogeneity in surgical procedures, AI tasks, analytical units, validation strategies, and outcomes, findings were synthesized descriptively without statistical pooling. The review was registered in the International Prospective Register of Systematic Reviews (CRD420251175699). Seventeen studies were included: seven clinical prediction or decision-support studies, eight intraoperative recognition, segmentation, or image-guided studies, and two training or workflow studies. Five prediction studies reported area-under-the-curve values of 0.74-0.95. Technical studies reported F1 or Dice scores of 0.525-0.995 and task-specific accuracies of 0.840-0.998. Two randomized studies suggested benefits for personalized suturing feedback and automated camera control, but neither established improved patient outcomes. Only one study had low overall risk of bias; the remaining studies were at high or unclear risk or raised some concerns. AI applications in robot-assisted surgery show promise for prediction, intraoperative perception, training, and workflow support. Evidence primarily demonstrates technical feasibility rather than established clinical effectiveness. Independent multicenter validation and prospective evaluation of patient, educational, and workflow outcomes are required before widespread implementation.

Robotic Surgical Procedures

Artificial intelligence-assisted detection and optical differentiation of colorectal lesions in Lynch syndrome surveillance (CADLY2): a multicentre, open-label, randomised controlled superiority trial.

BACKGROUND: Artificial intelligence (AI)-based computer-aided detection (CADe) systems improve adenoma detection in average-risk colorectal cancer screening. Meanwhile, evidence in Lynch syndrome surveillance is sparse and inconsistent. We assessed the effect of CADe on adenoma detection during Lynch syndrome surveillance. Computer-aided optical diagnosis (CADx) performance for optical differentiation of colorectal lesions was evaluated as a secondary aim. METHODS: CADLY2 was an international, multicentre, open-label, randomised controlled superiority trial at nine specialised hereditary cancer surveillance centres in Belgium, Germany, the Netherlands, and Spain. Adults aged 18 years or older with genetically confirmed Lynch syndrome scheduled for surveillance colonoscopy were randomly assigned (1:1) to high-definition white-light (HD-WL) colonoscopy alone or to HD-WL colonoscopy with computer-aided assistance from CAD EYE (Fujifilm, Tokyo, Japan). CAD EYE was used for CADe during withdrawal and for CADx after lesion detection. Randomisation was done centrally through a secure web-based system using Pocock's minimisation algorithm with a stochastic component and was stratified by centre, sex, previous colorectal cancer, underlying pathogenic variant, and interval since previous colonoscopy. Allocation concealment was ensured through the centralised web-based system. Patients were masked to group allocation until the start of withdrawal in procedures with mild sedation, or until completion of the procedure in procedures with propofol-based sedation. Endoscopists were not masked. The primary outcome was adenoma detection rate, defined as the proportion of patients with at least one histopathologically confirmed adenoma, analysed in the full analysis set (defined as all randomly allocated patients with available data for the primary outcome). The diagnostic performance of the CADx system was evaluated as a secondary outcome. The safety analysis set comprised all randomly allocated patients who underwent a study colonoscopy. This study is registered with the German Clinical Trials Register, DRKS00030695, and is completed. FINDINGS: Between May 9, 2023, and Oct 30, 2025, 757 patients were randomly allocated to HD-WL colonoscopy (377 patients) or to AI-assisted colonoscopy (380 patients); 733 patients were included in the full analysis set (369 HD-WL and 364 AI-assisted). The median age was 49 years (IQR 38-59) in the HD-WL group and 50 years (38-59) in the AI-assisted group; 213 (58%) were female and 156 (42%) male in the HD-WL group, and 207 (57%) were female and 157 (43%) male in the AI-assisted group. The adenoma detection rate was 30&#xb7;9% (114 of 369 patients) with HD-WL versus 33&#xb7;8% (123 of 364 patients) with CADe assistance (odds ratio 1&#xb7;14 [95% CI 0&#xb7;83-1&#xb7;57], p=0&#xb7;41). For CADx differentiation of neoplastic versus non-neoplastic lesions in the paired lesion-level analysis, with histopathology as the reference standard and sessile serrated lesions and traditional serrated adenomas classified as non-neoplastic, CADx sensitivity was 85&#xb7;9% (95% CI 82&#xb7;0-89&#xb7;1) and specificity was 91&#xb7;4% (89&#xb7;4-93&#xb7;0). Three adverse events occurred in the AI-assisted group: two mild post-polypectomy bleedings and one serious pulmonary embolism or deep venous thrombosis unrelated to the procedure. No adverse events occurred in the HD-WL group. INTERPRETATION: CADe-assisted colonoscopy did not show the absolute improvement in adenoma detection rate that was assumed in the prespecified sample-size calculation. CADx did not clearly improve lesion differentiation beyond expert optical diagnosis in expert Lynch syndrome surveillance settings. FUNDING: Third-party research funding of the National Center for Hereditary Tumor Syndromes, University Hospital Bonn.

Humans

Artificial intelligence enabled social robotic interventions (PARO) in Australian dementia care: A systematic review and meta-analysis.

BACKGROUND: Although there is a growing body of research indicating that Personal Robot/Social Robot could be used in various aspects of care for individuals with dementia, little is known about how well these types of interventions work in an actual hospital setting in Australia. AIMS & OBJECTIVES: The objective of the present systematic review and meta-analysis is to assess the effectiveness of PARO-based socially assistive robotic intervention in terms of its effectiveness outcomes towards the reduction of dementia-related behavioural and psychological symptoms in Australian based healthcare settings. METHODS: A systematic search was conducted across five electronic databases, including MEDLINE (PubMed), EMBASE, CINAHL, PsycINFO, and the Cochrane Library, to identify randomised controlled trials (RCTs) investigating PARO-based socially assistive robotic interventions for dementia in Australian healthcare settings. This review was registered with PROSPERO (CRD420251251916) and followed the PRISMA 2020 guidelines. In addition, the Cochrane Risk of Bias tool (RoB 2) was used to evaluate the risk of bias across all studies. Pooled standardised mean differences (SMD) with 95&#xa0;% confidence intervals (CI) were calculated for agitation, anxiety, and depression. Heterogeneity across studies was evaluated using the I2 statistic. RESULTS: Six RCTs involving 1444 participants were identified for inclusion in this review. AI-enabled socially assistive robotic interventions, specifically the PARO therapeutic robot, significantly reduced agitation and anxiety when compared to standard treatment or control conditions. The pooled analysis showed that agitation [SMD&#xa0;=&#xa0;-0.44 (95&#xa0;% CI: -0.70, -0.18) p&#xa0;=&#xa0;0.0008] and anxiety [SMD&#xa0;=&#xa0;-0.59 (95&#xa0;% CI: -0.91, -0.27) p&#xa0;=&#xa0;0.0003] were reduced significantly, while the decrease in depression [SMD&#xa0;=&#xa0;-0.44 (95&#xa0;% CI: -0.95, -0.07) p&#xa0;=&#xa0;0.09] scores was non-significant among dementia patients receiving PARO-based socially assistive robotic interventions as compared to the control. The overall risk of bias across all six studies was considered low to moderate. CONCLUSION: PARO-based socially assistive robotic interventions may provide preliminary evidence of effectiveness in reducing agitation and anxiety in individuals with dementia in Australian healthcare, but the evidence regarding the reduction of depression remains unclear. Therefore, additional high-quality trials with consistent methodology and extended follow-up will be necessary to determine both the short-term and long-term clinical efficacy and practicality of implementing these interventions into practice.

Humans

The application of artificial intelligence in healthcare practice: A mapping review of systematic reviews.

Artificial intelligence (AI) is rapidly transforming healthcare practice, with growing evidence supporting its use in diagnosis, prognosis, treatment planning, and operational decision-making. The proliferation of systematic reviews in recent years underscores the need for an updated synthesis of the literature to inform research, policy, and practice. We searched PubMed, Web of Science, Scopus, IEEE Xplore, and CINAHL for systematic reviews and meta-analyses published between 2019 and February 2026. Eligible reviews focused on AI applications in healthcare practice, were peer-reviewed, and written in English. A total of 368 reviews met the inclusion criteria. Publication volume increased steadily, peaking in 2025. AI research was concentrated in high-density domains, such as radiology, oncology, and critical care. Across reviews, diagnostic imaging, electronic health record (EHR) data, and biomarkers/laboratory results accounted for 68% of training data sources, though newer data types, such as wearable device and sensor data, emerged from 2022 onward. Diagnosis, prognosis, and treatment comprised over 80% of AI applications, with novel uses emerging in recent years, such as AI-assisted clinical documentation (e.g., ambient documentation tools) and patient education. Ethical concerns were reported in 78.5% of reviews, with privacy, model accuracy, data and algorithmic bias, and explainability as recurrent themes. The proportion of reviews reporting ethical concerns increased from 2021 to 2025. AI applications in healthcare are expanding in scope, diversifying in data sources, and evolving toward novel clinical and operational uses. The human-centered AI or augmented intelligence paradigm, integrating computational precision with clinical expertise, holds significant promise but will require parallel advances in governance, regulatory frameworks, and ethical oversight to ensure safe adoption.

Artificial Intelligence

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis.

BACKGROUND: Current endoscopic retrograde cholangiopancreatography (ERCP) and cholangioscopic-based diagnostic sampling for indeterminant biliary strictures remain suboptimal. Artificial intelligence (AI)-based algorithms by means of computer vision in machine learning have been applied to cholangioscopy in an effort to improve diagnostic yield. The aim of this study was to perform a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based diagnostic performance of AI-associated cholangioscopic diagnosis of indeterminant or malignant biliary strictures. METHODS: Individualized searches were developed in accordance with PRISMA and MOOSE guidelines, and meta-analysis according to Cochrane Diagnostic Test Accuracy working group methodology. A bivariate model was used to compute pooled sensitivity and specificity, likelihood ratio, diagnostic odds ratio, and summary receiver operating characteristics curve (SROC). RESULTS: Five studies (n=675 lesions; a total of 2,685,674 cholangioscopic images) were included. All but one study analyzed a deep learning AI-based system using a convoluted neural network (CNN) with an average image processing speed of 30 to 60 frames per second. The pooled sensitivity and specificity were 95% (95% CI: 85-98) and 88% (95% CI: 76-94), with a diagnostic accuracy (SROC) of 97% (95% CI: 95-98). Sensitivity analysis of CNN studies (4 studies, 538 patients) demonstrated a pooled sensitivity, specificity, and accuracy (SROC) of 95% (95% CI: 82-99), 88% (95% CI: 72-95), and 97% (95% CI: 95-98), respectively. CONCLUSIONS: Artificial intelligence-based machine learning of cholangioscopy images appears to be a promising modality for the diagnosis of indeterminant and malignant biliary strictures.

Humans

Experimental validation of an AI-driven digital healthcare platform for oral health behavior and plaque assessment among vietnamese children.

BACKGROUND: Oral health among children in developing countries, including Vietnam, remains a significant public health concern. Innovative approaches leveraging artificial intelligence AI-based digital health platforms may offer effective strategies for managing dental plaque and promoting better oral hygiene behaviors among school-aged children. This study aimed to evaluate the effectiveness of an AI-driven oral healthcare platform (Denti-i Vietnam) in improving oral hygiene and behavioral outcomes among Vietnamese primary school students. METHODS: A total of 204 primary school students aged 8-10&#xa0;years in Hanoi, Vietnam, participated in this experimental study. Participants were randomly assigned to an intervention group (n&#xa0;=&#xa0;107), which used the AI-driven oral healthcare platform, and a comparison group (n&#xa0;=&#xa0;97), which received traditional oral health education via pamphlets. Oral health behaviors, dental plaque levels (Simplified Oral Hygiene Index; OHI-S), and caries indices (dft/DMFT) were assessed at baseline and after the intervention period. RESULTS: The intervention group demonstrated a significant reduction in the OHI-S score compared to baseline (2.49&#xa0;&#xb1;&#xa0;0.60 to 1.70&#xa0;&#xb1;&#xa0;0.76, p&#xa0;<&#xa0;0.001), particularly in the debris component, indicating enhanced plaque control. Notable improvements were also observed in oral hygiene behaviors, including increased frequency of toothbrushing before and after breakfast (p&#xa0;<&#xa0;0.01) and more frequent parental assistance during brushing (p&#xa0;=&#xa0;0.03). Furthermore, parental awareness of dental caries significantly increased in the intervention group (p&#xa0;=&#xa0;0.001). CONCLUSIONS: The AI-driven oral healthcare platform significantly improved both oral hygiene behaviors and plaque control among Vietnamese primary school children. These findings suggest that AI-driven digital health tools can serve as practical and scalable solutions for promoting oral health in developing countries.

Humans

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis.

BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of &#x2265;5, &#x2265;15, and &#x2265;30 events/hour, with emphasis on models using non-PSG-derived inputs. METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2&#xd7;2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of &#x2265;5, &#x2265;15, and &#x2265;30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Humans