Search PubMedSearch

SEARCH · Search PubMed

Results for “human-AI interaction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,422 records · Page 3Linked to original sources

The Role of Artificial Intelligence for Intimate Partner Violence Prevention: A Systematic Review.

INTRODUCTION: Intimate partner violence (IPV), encompassing physical, sexual, emotional and economic abuse, remains a pervasive global health concern. Traditional prevention efforts face obstacles such as underreporting, delayed detection and limited personalised support. Emerging artificial intelligence (AI) approaches offer new opportunities to enhance IPV prevention. AIM: This systematic review maps and synthesises evidence on AI-driven tools in IPV prevention based on studies published between 2004 and 2024. METHODS: Following PRISMA 2020 guidelines and PROSPERO registration, we searched PubMed, Embase, CINAHL, PsycINFO, IEEE Xplore and Web of Science. Eligible studies explicitly evaluated AI technologies targeting IPV prediction, screening, intervention or support delivery. Study quality was appraised using the Mixed Methods Appraisal Tool (MMAT). RESULTS: Of 1304 records initially identified, 41 studies met eligibility criteria. AI applications ranged from machine learning (ML) for risk prediction and natural language processing (NLP) for IPV detection in clinical and social media data, to image analysis for forensic evaluation and chatbot-based support. Predictive modelling demonstrated strong discriminative performance, while NLP-based screening detected IPV with notable sensitivity. Chatbots showed feasibility and user acceptability, but evidence of their direct impact on reducing IPV incidence was limited, with one randomised controlled trial showing a modest reduction. Key challenges identified included algorithmic bias, data privacy risks and barriers to integration across health and social care systems. DISCUSSION: AI-informed interventions show promise for improving IPV detection, risk assessment, and scalable support, but questions remain about long-term effectiveness, ethical fairness, transparency and equitable implementation. Future interdisciplinary research should address these concerns to responsibly deploy AI in IPV prevention. RELEVANCE TO CLINICAL PRACTICE: The findings highlight the importance of trauma-informed, culturally responsive care and provider training in AI applications. Nurse-led innovation and policy advocacy will be crucial for safe, equitable integration of AI in IPV prevention.

Artificial Intelligence

Privacy, security, and reliability risks of artificial intelligence in healthcare: a systematic review of empirical evidence.

BACKGROUND: Artificial intelligence (AI) is increasingly integrated into healthcare information systems, supporting clinical decision-making, imaging analysis, and predictive modeling. While these applications offer operational and clinical benefits, they also introduce emerging risks to patient privacy, data security, and system reliability. OBJECTIVE: To systematically review empirical evidence on privacy breaches, security vulnerabilities, and misuse associated with AI applications in healthcare settings. METHODS: PubMed, Embase, Web of Science, Scopus, IEEE Xplore, and ACM Digital Library were searched for empirical studies published between January 2015 and November 2025 that evaluated AI use or misuse in clinical diagnosis, treatment, or decision-making. Two reviewers independently screened studies and extracted data using a standardized form. Findings were synthesized narratively due to heterogeneity in study designs, AI methods, and reported outcomes. RESULTS: Of 7,285 records identified through database searches and 205 through citation screening, 22 empirical studies met the inclusion criteria, spanning multiple clinical domains and data modalities, predominantly medical imaging applications. Five recurring threat categories were identified: patient re-identification, membership inference, unauthorized access and adversarial exploitation, input manipulation, and misuse or overinterpretation of AI outputs. Across studies, AI models were shown to encode latent biometric signals across diverse data types, limiting the effectiveness of traditional anonymization and synthetic data approaches. Adversarial attacks and input manipulation were also shown to compromise diagnostic performance and system integrity. CONCLUSION: This systematic review provides empirical evidence suggesting that contemporary AI systems in healthcare introduce privacy and security risks that may challenge traditional assumptions about data protection. These findings underscore the need for privacy- and security-by-design approaches and governance frameworks that address risks across the AI lifecycle.

Humans

Improving Community-Based Care for Adolescents with ADHD: a Randomized Controlled Trial of Artificial Intelligence-Assisted Fidelity Supports.

Cognitive-behavioral treatments (CBTs) for adolescents with ADHD demonstrate promise of long-term effects on outcome. However, their implementation in routine care community clinics faces barriers that impact quantity, efficiency, and quality of delivery, as well as client outcomes. This study is a randomized controlled trial designed to evaluate the impact of an AI-assisted service delivery model on therapist implementation of Supporting Teens' Autonomy Daily (STAND), a CBT blended with Motivational Interviewing (MI) for adolescents with ADHD. Adolescents with ADHD (N = 51), who were clients at three community mental health agencies, received treatment from 23 therapists. There was randomization of adolescents and therapists to AI-assisted or standard implementation supports. In addition to standard supports (i.e., training, standard facilitation resources, technical assistance, case supervision), AI-assisted support package included digitized facilitation resources housed in a clinical dashboard (Care4), feedback on content fidelity, and AI-generated feedback on MI implementation quality. The AI-assisted group was associated with more efficient treatment delivery and lower number of appointments attended by the adolescent. There was also a significant decrement in MI quality over time in the AI-assisted group compared to the standard support group. Feedback in focus groups indicated that therapists perceived a task-oriented mindset to be associated with receipt of the AI-assisted support package, leading therapists to prioritize efficiency over relational aspects of therapy. Following the results of this trial, a future, larger RCT should examine the impact of the AI-assisted implementation model on mental health outcomes and cost savings to organizations, third party payers, and clients. Trial registration number: NCT05135065; https://www.clinicaltrials.gov ; Registered September 2021.

Humans

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis.

BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG-derived inputs. METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Humans

The application of artificial intelligence in healthcare practice: A mapping review of systematic reviews.

Artificial intelligence (AI) is rapidly transforming healthcare practice, with growing evidence supporting its use in diagnosis, prognosis, treatment planning, and operational decision-making. The proliferation of systematic reviews in recent years underscores the need for an updated synthesis of the literature to inform research, policy, and practice. We searched PubMed, Web of Science, Scopus, IEEE Xplore, and CINAHL for systematic reviews and meta-analyses published between 2019 and February 2026. Eligible reviews focused on AI applications in healthcare practice, were peer-reviewed, and written in English. A total of 368 reviews met the inclusion criteria. Publication volume increased steadily, peaking in 2025. AI research was concentrated in high-density domains, such as radiology, oncology, and critical care. Across reviews, diagnostic imaging, electronic health record (EHR) data, and biomarkers/laboratory results accounted for 68% of training data sources, though newer data types, such as wearable device and sensor data, emerged from 2022 onward. Diagnosis, prognosis, and treatment comprised over 80% of AI applications, with novel uses emerging in recent years, such as AI-assisted clinical documentation (e.g., ambient documentation tools) and patient education. Ethical concerns were reported in 78.5% of reviews, with privacy, model accuracy, data and algorithmic bias, and explainability as recurrent themes. The proportion of reviews reporting ethical concerns increased from 2021 to 2025. AI applications in healthcare are expanding in scope, diversifying in data sources, and evolving toward novel clinical and operational uses. The human-centered AI or augmented intelligence paradigm, integrating computational precision with clinical expertise, holds significant promise but will require parallel advances in governance, regulatory frameworks, and ethical oversight to ensure safe adoption.

Artificial Intelligence

Artificial intelligence in genitourinary oncology: publication trends and systematic review.

OBJECTIVE: To conduct an analysis of publication trends and a systematic review of randomized controlled trials (RCTs) to characterize the current state of artificial intelligence (AI) use in genitourinary (GU) oncology, as AI has emerged as a transformative tool in healthcare with potential applications in diagnostics, treatment planning, and prognostication. METHODS: We searched the Medical Literature Analysis and Retrieval System Online (MEDLINE), Excerpta Medica dataBASE (EMBASE; Ovid), and Cumulative Index to Nursing and Allied Health Literature (CINAHL) Ultimate for studies related to AI and GU oncology, excluding non-English papers, non-human studies, review articles, and articles using AI solely for manuscript writing. Publication trends were analysed from 2013 to 2023 and categorized by study design and cancer type. RCTs were evaluated through systematic review using Covidence (Veritas Health Innovation Ltd, Melbourne, Victoria, Australia) for screening and data extraction. Two reviewers independently assessed all studies, with risk of bias (RoB) evaluated using the Cochrane RoB 2.0 tool. RESULTS: Of 2409 articles identified, 1220 met inclusion criteria. These included 962 retrospective articles, 175 prospective studies, 79 studies with combined retrospective/prospective methods, and four RCTs. Studies most commonly addressed prostate (n = 923), renal (n = 274), and urothelial (n = 194) cancers. Publications grew from 14 in 2013 to 362 in 2023, with substantial acceleration in 2019. Four RCTs were identified - one in urothelial cancer and three in prostate cancer. Two RCTs evaluated AI-based diagnostics, demonstrating improved performance over conventional methods; the remaining two RCTs evaluated AI in prognostication and treatment planning, showing improved gains in imaging interpretation and operational efficiency. RoB varied across studies, primarily related to randomisation and deviations from intended interventions. CONCLUSIONS: Artificial intelligence research in GU oncology has grown, although high-level evidence from RCTs remains limited. Existing trials underscore AI's promise in diagnostics, prognostication, and treatment planning, and the rapidly evolving nature of this field warrants continued prospective investigation.

Humans

Artificial intelligence-supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programs.

BACKGROUND: Most European population mammography screening programs rely on double reading with arbitration, a model that delivers mortality benefit but is increasingly challenged by radiologist workload, variable specificity, and interval cancers. Artificial intelligence (AI) is being evaluated to support or optimize these established European screening pathways. PURPOSE: To synthesize prospective or program-embedded evaluations of AI conducted within European-style population screening programs and to estimate exploratory program-level absolute risk differences (RDs) per 1000 examinations for cancer detection rate (CDR) and recall. MATERIALS AND METHODS: We performed a prespecified, focused evidence synthesis of three large studies embedded within routine population screening programs operating under European-relevant workflows: MASAI (randomized AI-supported risk triage within a national program), ScreenTrustCAD (prospective paired-reader evaluation with AI as an independent reader in a double-reading framework), and PRAIM (nationwide decision-referral implementation). Outcomes were harmonized as AI-control RDs per 1000 examinations. Random-effects pooling used Hartung-Knapp-Sidik-Jonkman models. For the paired-reader design, sensitivity analyses applied a Kish effective sample-size approach across plausible within-examination correlations (ρ = 0.3-0.8). Positive predictive value (PPV) and workflow/time outcomes were summarized descriptively. RESULTS: Across 597,419 examinations, the pooled CDR RD was +0.9 per 1000 (95% CI -0.0 to +1.8; I2 ≈ 12%), consistent with a modest directional increase with borderline statistical uncertainty. The pooled recall RD was -0.6 per 1000 (95% CI -3.1 to +2.1; I2 ≈ 41-43%), indicating no consistent recall increase across screening programs. Where reported, PPV was higher with AI-supported screening. Efficiency signals included 44.3% fewer total readings in MASAI and shorter reading times for AI-normal examinations in PRAIM; in PRAIM, a program-level safety-net mechanism recovered 204 cancers that would otherwise have been missed. CONCLUSION: In European population screening programs characterized by double reading and arbitration, prospective program-embedded evidence suggests that AI integration may yield a small absolute increase in cancer detection (≈1/1000) without a consistent increase in recall, alongside improved PPV and efficiency signals. These findings suggestAI primarily as a complementary reader within European screening workflows, with implementation requiring explicit quality assurance and monitoring of interval cancers and stage distribution.

Humans

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans

Artificial intelligence-derived myocardial fibrosis on cardiac magnetic resonance for prognosis in cardiomyopathy: A systematic review of a sparse evidence base.

BACKGROUND: Myocardial fibrosis on cardiovascular magnetic resonance (CMR), assessed by late gadolinium enhancement (LGE) and parametric mapping, is an established predictor of adverse events in cardiomyopathy. We assessed whether artificial intelligence (AI) quantification of fibrosis adds independent prognostic value. METHODS: We searched six databases, a clinical-trials register, and a preprint server from inception to 13 June 2026. Eligible studies used AI to generate a fibrosis marker in adults with ischemic or nonischemic cardiomyopathy, with covariate-adjusted outcomes over ≥12 months. Risk of bias was assessed using PROBAST, PROBAST+AI, and QUIPS. Fewer than three comparable studies precluded meta-analysis; certainty was rated using GRADE. RESULTS: Of 448 records (381 after de-duplication), 18 full texts were reviewed and two included, one peer-reviewed and one preprint. In an ischemic-cardiomyopathy registry (Ghanbari et al.; n = 216 analytic, 26 events), AI-derived dense LGE scar predicted arrhythmic events (univariable hazard ratio [HR] 2.35, 95% CI 1.33-4.15), and AI-derived but not manual scar improved discrimination beyond guideline criteria (area under the curve 0.63 to 0.68; p = 0.02). In a nonischemic dilated-cardiomyopathy preprint (Kim et al.; n = 347, 119 events), automated extracellular volume ≥30% predicted cardiovascular death or heart-failure hospitalization (adjusted HR 2.00, 95% CI 1.32-3.03). Both were at high risk of bias, with data-derived thresholds and no external validation. CONCLUSIONS: Across only two studies, AI-derived fibrosis was independently associated with adverse cardiovascular events, but its added value over manual quantification remains unproven. Certainty was very low. The evidence base is sparse and not yet ready for clinical use.

Humans

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial.

BACKGROUND: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin. METHODS: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0-10 points), with a prespecified non-inferiority margin of -0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment. RESULTS: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was -0.335 points (95% CI, -1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of -0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group. CONCLUSIONS: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.

Humans

Development and Crossover Evaluation of an Artificial Intelligence-Assisted System for Solid Pancreatic Lesion Detection and Pancreatic Parenchyma Recognition in Endoscopic Ultrasonography (With Video).

BACKGROUND AND STUDY AIMS: Pancreatobiliary endoscopic ultrasonography (EUS) is technically demanding, and supervised training opportunities are limited. We developed an artificial intelligence (AI) overlay system for detecting solid pancreatic lesions (SPL) and recognizing pancreatic parenchyma (PP) and evaluated its effect on reader performance. PATIENTS AND METHODS: Across six centers, two deep learning-based models were trained using expert-annotated EUS frames. We then conducted a randomized, two-sequence, two-period crossover reader study in which eight endosonographers (five novices and three experts) interpreted image sets with and without AI assistance. The primary endpoint was superiority of sensitivity for SPL detection among novices; key secondary endpoints included specificity and PP recognition. RESULTS: From 118 patients, 120 SPL-positive/negative image sets and 160 PP-positive/negative image sets were constructed. Among novices, AI assistance improved SPL detection sensitivity (88.7% vs. 76.8%, p&#x2009;<&#x2009;0.001) and accuracy (86.4% vs. 78.7%), while specificity met the predefined noninferiority criterion (84.2% vs. 80.5%, p&#x2009;<&#x2009;0.001). For PP recognition, sensitivity increased numerically (86.3% vs. 83.3%) but did not meet the predefined superiority criterion (p&#x2009;=&#x2009;0.095); specificity met the noninferiority criterion (87.8% vs. 81.0%), and accuracy increased from 82.1% to 87.0%. Among experts, sensitivity was maintained for both tasks, whereas specificity increased with AI assistance. CONCLUSIONS: AI assistance improved SPL detection among novice endosonographers. For PP recognition, sensitivity increased without reaching statistical superiority, whereas specificity met the predefined noninferiority criterion. These findings support a potential adjunctive role for AI in EUS interpretation.

Humans

The Impact of Chatbot Type and Normative Messaging on Chatbot Usage Intention Based on the Health Technology Acceptance Model: Randomized Controlled Trial.

BACKGROUND: Digital health tools, such as health chatbots, may improve access to scalable health support, but adoption remains inconsistent. Existing models do not fully integrate technology acceptance factors with health motivation factors relevant to digital health use. OBJECTIVE: This study proposed and tested the health technology acceptance model and examined whether normative message framing and chatbot type were associated with health motivation, technology acceptance, and intention to use a health chatbot. METHODS: In October 2025, we conducted a 4 &#xd7; 2 between-participants online experiment with 1000 US adults recruited from a nationally representative YouGov panel. Participants were randomized to 1 of 8 conditions varying norm message type (self-oriented, peer-oriented, expert-oriented, or family-oriented) and chatbot type (AI-powered or rule-based) in a cancer prevention and genetic risk information scenario. Outcomes included descriptive norms, injunctive norms, perceived susceptibility, perceived severity, perceived benefits, self-efficacy, perceived ease of use, trust, privacy concerns, and usage intention. Data were analyzed using a multivariate ANOVA with Bonferroni-adjusted post hoc tests and multiple linear regression. RESULTS: Peer-oriented and family-oriented messages produced higher usage intention than expert-oriented messages, and peer-oriented messages also increased descriptive norms, injunctive norms, self-efficacy, and trust. AI-powered chatbots were associated with higher usage intention (P=.02) and greater trust (P=.008) than rule-based chatbots. In regression analyses, the model explained 50.8% of the variance in usage intention. Usage intention was positively associated with descriptive norms (&#x3b2;=0.087; P=.003), injunctive norms (&#x3b2;=0.078; P=.009), perceived susceptibility (&#x3b2;=0.051; P=.03), perceived benefits (&#x3b2;=0.253; P<.001), and trust (&#x3b2;=0.33; P<.001), and negatively associated with perceived severity (&#x3b2;=-0.047; P=.049) and privacy concerns (&#x3b2;=-0.11; P<.001). Perceived ease of use and self-efficacy were not significant predictors. CONCLUSIONS: The health technology acceptance model was a useful framework for explaining the intention to use a health chatbot by combining technology acceptance and health motivation constructs. Both social design features and chatbot design features shaped adoption-related beliefs, with peer-oriented and family-oriented framing and AI-powered chatbots showing particular promise. Trust and privacy concerns remained central determinants of intended use.

Humans

Effectiveness of artificial intelligence in nursing simulation education: A systematic review, meta-analysis and bibliometric visualization analysis.

OBJECTIVES: To synthesize the roles and core functions of AI in nursing simulation education for nursing students via systematic review, quantitatively evaluate its effects on students' knowledge and skill outcomes through meta-analysis, and map the research landscape and development trends of this field through bibliometric visualization analysis. DESIGN: Systematic review, meta-analysis and bibliometric visualization analysis. DATA SOURCES: Eight electronic databases: PubMed, Web of Science, MEDLINE, ERIC, Academic Search Complete, China National Knowledge Infrastructure (CNKI), Wanfang Database, VIP Chinese Science and Technology Journal Database (VIP) were employed to search studies from the time of construction to 16 December 2025. REVIEW METHODS: Studies meeting the inclusion criteria were screened. The revised Cochrane Risk of Bias tool (ROB 2) and Joanna Briggs Institute (JBI) critical appraisal checklists were used for quality assessment. Meta-analysis was performed with Review Manager 5.4, and bibliometric visualization analysis was conducted using VOSviewer 1.6.20 and Bibliometrix (based on R4.4.3). RESULTS: A total of 61 studies were included. AI primarily played two roles in nursing simulation education: peer-type new subject (n&#xa0;=&#xa0;24) and direct mediator (n&#xa0;=&#xa0;22). Meta-analysis showed that AI interventions significantly improved nursing students' knowledge (SMD&#xa0;=&#xa0;1.49, 95% CI [0.55,2.43], p&#xa0;=&#xa0;0.002) and skills (SMD&#xa0;=&#xa0;0.66, 95% CI [0.02,1.31], p&#xa0;=&#xa0;0.04). Bibliometric analysis identified that the United States of America and China were the two main contributing countries in this field, and the key motor themes included generative artificial intelligence, virtual patients, and geriatric care. CONCLUSIONS: AI exerts positive effects on nursing students' knowledge acquisition and skill enhancement in simulation education, with peer-type new subject and direct mediator as the dominant roles. Future research should focus on expanding AI applications in multi-specialty simulation scenarios, activating the data-driven value of machine learning, and strengthening international collaboration and standardization construction, so as to promote the sustainable development of AI-integrated nursing simulation education.

Humans

The Impact of Baseline Negative Emotions on Postoperative Quality of Life in Adolescent Idiopathic Scoliosis Patients: A 2-Year Follow-Up Study.

OBJECTIVE: Adolescent idiopathic scoliosis (AIS) is a three-dimensional spinal deformity that develops during puberty without a clear etiology. Beyond physical manifestations, AIS severely impacts adolescents' psychological and social well-being, leading to anxiety, depression, and low self-esteem. While advancements in surgical techniques have enhanced objective outcomes, existing studies on AIS have primarily focused on objective indices, with limited attention to the long-term impact of preoperative negative emotions on patient-reported subjective quality of life. METHODS: This was a retrospective cohort study. A total of 112 eligible AIS patients who underwent posterior spinal correction surgery between April and August 2023 were enrolled. Inclusion criteria included confirmed AIS, completion of 2-year follow-up, and informed consent; exclusion criteria included missing imaging/questionnaire data, comorbid psychiatric/neurological diseases, or prior spinal surgery. Patients were grouped using the Hospital Anxiety and Depression Scale (HADS) administered on admission. Quality of life was assessed preoperatively and 2&#x2009;years postoperatively using the Scoliosis Research Society-22 (SRS-22, evaluating self-image, mental health, pain, function, treatment satisfaction) and Short Form 36 Health Survey (SF-36, assessing 8 physical and mental health dimensions). Statistical analysis was performed via SPSS, using independent t-tests, paired t-tests, Mann-Whitney U test, and chi-square test. p&#x2009;<&#x2009;0.05 was considered significant. RESULTS: There were no significant differences in baseline characteristics (age, gender, BMI, surgical parameters, scoliosis type, preoperative/postoperative Cobb angles) between the two groups (all p&#x2009;>&#x2009;0.05). Preoperatively, SRS-22 and SF-36 scores showed no inter-group differences (all p&#x2009;>&#x2009;0.05). Postoperatively, the Negative Emotion Group had significantly lower scores in SRS-22 mental health (3.9&#x2009;&#xb1;&#x2009;0.3 vs. 4.5&#x2009;&#xb1;&#x2009;0.2) and treatment satisfaction (4.0&#x2009;&#xb1;&#x2009;0.3 vs. 4.6&#x2009;&#xb1;&#x2009;0.7), as well as SF-36 general health (68.6&#x2009;&#xb1;&#x2009;6.4 vs. 79.7&#x2009;&#xb1;&#x2009;13.3), role-emotional (61.3&#x2009;&#xb1;&#x2009;9.3 vs. 70.8&#x2009;&#xb1;&#x2009;9.7), and mental health (61.8&#x2009;&#xb1;&#x2009;14.3 vs. 68.9&#x2009;&#xb1;&#x2009;10.7) (all p&#x2009;<&#x2009;0.05); no inter-group differences were observed in physical function-related dimensions. Both groups showed significant improvements in physical function-related dimensions postoperatively. The Non-Negative Emotion Group also exhibited significant improvements in SRS-22 self-image/pain and SF-36 bodily pain (all p&#x2009;<&#x2009;0.05), while the Negative Emotion Group showed no significant improvements in these dimensions. CONCLUSIONS: Preoperative anxiety and depression do not affect the recovery of physical function in AIS patients after spinal correction surgery but significantly impede improvements in subjective quality of life dimensions, including mental health and treatment satisfaction. These findings highlight the need to integrate psychological assessment and targeted interventions into the perioperative management of AIS. Such a patient-centered approach will help optimize both physical and psychological outcomes, ultimately achieving comprehensive rehabilitation for AIS adolescents.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95&#xa0;% CI 0.85-0.94; 95&#xa0;% prediction interval 0.62-0.98), with sensitivity of 0.80 (95&#xa0;% CI 0.77-0.83) and specificity of 0.87 (95&#xa0;% CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

The impact of artificial intelligence on critical thinking and clinical reasoning in health professions education: A systematic review and meta-analysis.

BACKGROUND: Critical thinking and clinical reasoning underpin healthcare professionals' ability to navigate uncertainties and deliver safe and effective care. With artificial intelligence (AI) advancement and growing adoption, AI-based educational tools are increasingly used to support these cognitive competencies' development. OBJECTIVE: To synthesize randomised and controlled clinical trials on AI-based educational tools in health professions education and examine their effects on critical thinking and clinical reasoning among health professions students. METHODS: Six electronic databases were searched from January 1, 2014 to July 28, 2025 was reviewed: PubMed, Cochrane Central Register of Controlled Trials, CINAHL, Scopus, Embase and Web of Science. Two independent reviewers performed data extraction and quality assessment using standardized JBI checklists. The GRADE approach was used to assess the certainty of evidence. Studies were pooled via random-effects meta-analyses or narrative syntheses. RESULTS: Fourteen randomised controlled trials and seven controlled clinical trials were included (n&#xa0;=&#xa0;21). Meta-analyses revealed small to medium effect sizes for the surrogate clinical reasoning outcomes of performance-based assessment scores (SMD 0.68; 95% CI [0.38, 0.98], p-value&#xa0;=&#xa0;0.00; I2&#xa0;=&#xa0;38%) and knowledge test scores (SMD 0.39; 95% CI [0.09, 0.69], p-value&#xa0;=&#xa0;0.01; I2&#xa0;=&#xa0;79%). Critical thinking and clinical reasoning skills and dispositions were narratively synthesized, with majority of included studies favouring AI-based interventions but the evidence had low to very low certainty. CONCLUSION: AI-based educational interventions may improve critical thinking and clinical reasoning among health profession students, but the evidence is very uncertain. This review offers preliminary insights but does not allow identification of optimal interventions or discipline-specific recommendations due to small sample sizes and substantial intervention heterogeneity. Further research is required to draw definitive conclusions. PROTOCOL REGISTRATION: CRD42025634074.

Humans

Phase 1 Study Evaluating Gefurulimab Pharmacokinetics and Safety Following Delivery Via Autoinjector or Prefilled Syringe With Needle Safety Device in Healthy Adults.

PURPOSE: Gefurulimab, a novel dual-binding nanobody targeting complement component 5 (C5), is in clinical development for anti-acetylcholine receptor antibody-positive generalized myasthenia gravis. Gefurulimab has a low molecular weight, enabling subcutaneous (SC) self-administration by autoinjector (AI) or prefilled syringe with needle safety device (PFS-SD). We compared gefurulimab pharmacokinetic (PK) exposure and safety in healthy adults following a single SC dose administered by AI versus PFS-SD. METHODS: In this phase 1, open-label, randomized, parallel-group study (NCT06208488), healthy participants aged 18 to 65 years were stratified by weight and randomized equally to 1 of 6 combination groups of device and injection site (abdomen/thigh/upper arm). Participants received a single SC dose of gefurulimab on day 1 and were assessed throughout the 92-day evaluation period. Primary endpoints were PK parameters for each device: maximum observed concentration (Cmax) and area under the serum concentration-time curve (AUCinf, AUClast). PK across injection sites, pharmacodynamics, safety, immunogenicity, and device performance were also assessed. FINDINGS: Overall, 175 participants were randomized: AI (n = 87), PFS-SD (n = 88). Geometric least squares mean ratios (90% CI) comparing AI/PFS-SD for Cmax, AUCinf, and AUClast were 97.6% (94.5-100.8), 99.6% (96.1-103.3), and 98.8% (95.2&#x2012;102.6), respectively. Secondary analyses found no meaningful differences in PK parameters across injection sites. Serum-free C5 concentrations over time, treatment-emergent adverse event (TEAE) profiles, and antidrug antibody responses were similar between cohorts. Most TEAEs were mild; none led to study discontinuation. IMPLICATIONS: SC administration of gefurulimab by AI and PFS-SD was well tolerated with comparable exposure, meeting bioequivalence criteria.

Humans

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis.&#xa0;A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST&#x2009;+&#x2009;AI for prediction model studies.&#xa0;Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST&#x2009;+&#x2009;AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection.&#xa0;AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans