Search PubMedSearch

SEARCH · Search PubMed

Results for “AI-assisted literature review”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,207 records · Page 5Linked to original sources

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial.

BACKGROUND: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin. METHODS: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0-10 points), with a prespecified non-inferiority margin of -0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment. RESULTS: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was -0.335 points (95% CI, -1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of -0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group. CONCLUSIONS: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.

Humans

Pharmacogenomic and drug interactions risk in cardio-oncology: A precision medicine perspective for India.

Cardio-oncology patients may face complex treatment regimens due to the concurrent existence of cancer and cardiovascular disease, leading to a considerable polypharmacy burden. This significantly increases the prospect of drug-drug interactions (DDIs) and gene-drug interactions. The majority of these interactions arise from comparable pharmacokinetic and pharmacological pathways associated with drug transporters and cytochrome P450 enzymes. The significance of pharmacogenomics in tailored treatment strategies are emphasised by the fact that genetic variability enhances individual differences in drug response, safety, and efficacy. This narrative review focus on the effects of key genetic polymorphisms (e.g., DPYD, CYP2C19, and CYP2C9) on the metabolism and efficacy of commonly prescribed anticancer and cardiovascular medications such as fluoropyrimidines, clopidogrel, and warfarin. In addition it explore the role of pharmacogenomic variants on drug-drug interactions within the field of cardio-oncology. The study ultimately emphasizes the necessity of precision medicine in India to address the genetic diversity and underrepresentation in global genomic databases. The absence of pharmacogenomic testing, infrastructural deficiencies, financial constraints, and insufficient clinical integration hinder the widespread use of this technology in India. The Genome India Project and other national initiatives establish the foundation for pharmacogenomic-guided therapy. Utilizing genetic data, together with artificial intelligence-based predictive tools, for clinical decision-making may enhance medication safety and yield optimal outcomes in Indian cardio-oncology patients.

Humans

A Sentiment-Based Comparison of AI- and Physician-Generated Empathic Statements in Palliative Care.

CONTEXT: Empathic communication promotes trust in patient-provider relationships. As healthcare integrates artificial intelligence (AI) into patient communication, we have yet to understand how these models' communication compares to that of physicians. OBJECTIVES: Our primary objectives were to examine patient preferences for AI-generated vs. palliative care physician-generated empathic statements addressing fear and anxiety around cancer treatment, and to analyze associations between linguistic features and patient preferences. METHODS: We conducted a secondary analysis of the PALL-AI trial, a randomized controlled survey comparing cancer patients' preferences of AI- to physician-generated empathic statements. Physicians and AI were provided the same prompt with a maximum sentence length. Patient preferences for each statement were measured in blinded surveys. We analyzed sentiment of the statements using the Valence Aware Dictionary and Sentiment Reasoner (VADER) and the National Research Council Canada (NRC) Emotion Lexicon. We evaluated associations between sentiment scores and patient preferences using Spearman's correlation coefficients. RESULTS: A total of 105 patients completed blinded surveys, preferring the AI-generated statement 72.4% of the time. VADER sentiment analysis showed all three AI statements displayed positive sentiment, while all three physician statements displayed negative sentiment. Controlling for statement length, AI statements used twice as many positive words as human statements. However, they contained a similar number of negative words. Of the eight NRC emotions, "trust" and "joy" demonstrated the strongest correlations with patient preference. CONCLUSION: Patients preferred AI-generated statements around cancer care over those from palliative care physicians when standardized for prompt and statement length. Analysis shows AI-generated statements contain more positive language which may be the factor driving patient preference toward AI.

Humans

Orofacial Cleft Disparities in American Indian and Alaska Native Populations: A Systematic Review and Meta-Analysis.

ObjectiveTo evaluate the prevalence, access to care, and health outcomes of orofacial clefts (OFCs) among American Indian and Alaska Native (AI/AN) populations through a systematic review and meta-analysis.DesignSystematic review and meta-analysis performed in accordance with PRISMA 2020 guidelines and registered with PROSPERO (CRD420251035364).SettingUS-based population registries, hospital databases, and institutional or community-level retrospective studies involving AI/AN populations.Patients and ParticipantsAI/AN individuals with OFCs compared with non-Hispanic White patients.InterventionsPrimary cleft lip and palate repair, secondary cleft-related procedures, and multidisciplinary cleft care.Main Outcome Measure(s)Prevalence of OFCs, timing of cleft surgery, discharge disposition, access to specialists, and qualitative determinants of disparities.ResultsEighteen studies including more than 1985 AI/AN patients were identified. Meta-analysis of 5 studies estimated a pooled OFC prevalence of 15 per 10 000 live births (95% confidence interval: 5-49), with substantial heterogeneity (I2 = 99.8%). Individual studies reported significantly higher OFC prevalence in AI/AN populations compared to non-Hispanic Whites (odds ratio range: 1.44-2.68). Geographic maldistribution of craniofacial-trained surgeons, increased odds of nonhome discharge, and delayed cleft palate repair were consistently observed barriers. Qualitative analyses highlighted structural inequities, perceived racism, and lack of culturally responsive care as major contributors to disparities.ConclusionsAI/AN populations face a disproportionately high burden of OFCs alongside structural barriers to timely, culturally competent care. Addressing these disparities requires community-engaged, multidisciplinary interventions that improve geographic access and integrate culturally responsive approaches to care.

Humans

Food-derived extracellular vesicles as delivery platforms for medicine-food homology components in metabolic syndrome.

Diet-induced obesity and associated metabolic syndromes have become major global public health challenge, highlighting the urgent need for safe and effective strategies. Recently, food-derived extracellular vesicles (FDEVs) have garnered increasing attention as natural nanocarriers due to their excellent biocompatibility and specific targeted delivery capabilities. FDEVs can efficiently deliver medicine-food homology components (MFHCs) to precisely regulate lipid metabolism, inflammatory responses, and insulin sensitivity, thereby improving obesity and its metabolic abnormalities. This systematic review summarizes recent advances in the use of FDEVs as delivery vehicles for MFHCs to suppress diet-induced obesity and metabolic syndrome, with a particular focus on the underlying molecular mechanisms, including signaling pathway regulation and cellular metabolic remodeling. In addition, the clinical translational potential and industrial application prospects of FDEVs are evaluated, and key challenges related to preparation techniques, safety assessment, and large-scale production are discussed. By integrating current evidence, this review aims to provide theoretical framework and future perspectives for the development of FDEVs as a novel targeted delivery platform and treatment of metabolic diseases.

Extracellular Vesicles

Manual, digital, and AI tumour-infiltrating lymphocyte scoring: a secondary analysis of the APHINITY randomised trial.

BACKGROUND: Stromal tumour-infiltrating lymphocytes (sTILs) are prognostic in early-stage HER2-positive breast cancer, but their role in the context of dual HER2 blockade remains undefined. We evaluated manual, digital, and artificial intelligence (AI)-based sTIL quantification, together with AI-derived spatial metrics, for prognostic and treatment-benefit stratification using tumour samples from the phase 3 APHINITY trial. METHODS: In the APHINITY trial, 4805 patients were randomly assigned to receive chemotherapy plus trastuzumab with pertuzumab or chemotherapy plus trastuzumab with placebo. Median follow-up was 74&#xb7;1 months (IQR 68&#xb7;3-75&#xb7;4). We analysed 4262 haematoxylin and eosin-stained images using manual assessment, an automated digital approach, AI-based lymphocyte quantification (AI percentage lymphocytes), and two AI-derived spatial features (AI-TIL and immune hotspot). Interobserver reproducibility was assessed in 262 randomly chosen tumour samples scored independently by five pathologists. Multivariable Cox models were used to assess associations between TIL levels and invasive disease-free survival (primary outcome in APHINITY), distant recurrence-free interval, and overall survival. The heterogeneity of pertuzumab benefit was evaluated using subgroup analyses, subpopulation treatment effect pattern plot analyses, and nested Cox models with treatment-by-biomarker interaction terms. FINDINGS: Manual scoring showed high interobserver reproducibility (intraclass correlation coefficient 0&#xb7;84 [95% CI 0&#xb7;79-0&#xb7;88]). Concordance between manual and automated methods was modest. AI-based scoring (AI percentage lymphocytes) reclassified 120 (11&#xb7;6%) of 1035 node-positive tumours from immune-low (by manual scoring) to immune-high; this subgroup of patients showed greater separation of 5-year invasive disease-free survival curves between pertuzumab and placebo groups compared with patients whose tumours were concordantly classified as immune-low by both manual and AI-based approaches. Higher levels of TILs were associated with improved invasive disease-free survival for all sTIL measurement approaches and spatial measurements (hazard ratios [HRs] 0&#xb7;41-0&#xb7;93). Pertuzumab was associated with improved invasive disease-free survival at higher sTIL levels across all measurement approaches (HRs 0&#xb7;36-0&#xb7;48), but was not associated with higher values of spatial measures. The largest 6-year absolute improvements with pertuzumab were observed in patients with node-positive disease whose tumours scored in the highest level of immune infiltration of manual sTIL scoring (&#x2265;70&#xb7;0%; mean absolute improvement 12&#xb7;1 percentage points [SD 2&#xb7;8]). In nested prognostic and predictive models, AI-based immune hotspot scores provided the most consistent additional information when combined with any sTIL measurement (all p<0&#xb7;010). INTERPRETATION: Standardised manual sTIL scoring was reproducible, and digital and AI-based methods showed consistent prognostic stratification and potential for treatment-benefit stratification despite only modest correlation between platforms. AI spatial metrics provided complementary information beyond sTIL density and could support more scalable immune assessment. Future studies are needed to validate these approaches in independent cohorts and to clarify their clinical utility for stratifying contemporary HER2-directed therapies. FUNDING: None.

Humans

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

User Engagement and Feature Preferences in an AI-Powered mHealth Intervention for Diabetes Prevention: Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Prediabetes is highly prevalent and increasing globally, yet lifestyle interventions remain underused. AI-driven mobile health (mHealth) tools can help scale diabetes prevention efforts, but the key factors driving their success are not well understood. OBJECTIVE: This post hoc secondary analysis of a randomized controlled trial (RCT) aimed to characterize the most valued features and the role of user engagement in outcomes of a fully automated mHealth intervention for diabetes prevention. METHODS: Data from 151 participants with prediabetes and overweight or obesity who were assigned to an AI-based diabetes prevention program (Sweetch) in a parent RCT (NCT05056376) were analyzed. Engagement (defined as the total number of days the app was used) was categorized into tertiles (low, medium, and high). Baseline characteristics were compared across engagement groups using ANOVA, Kruskal-Wallis, and chi-square tests, and regression models assessed the association between engagement and achievement of diabetes risk reduction outcomes (&#x2265;5% weight loss, &#x2265;4% weight loss with &#x2265;150 min/week of physical activity, or &#x2265;0.2 percentage point reduction in hemoglobin A1c [HbA1c] at 12 months). Perceived usefulness of intervention features was surveyed at 12 months. RESULTS: Median engagement was 98 (IQR 34-232) days. Older age (P<.001) and lower baseline BMI (P=.04) were significantly associated with higher engagement. Compared with low engagement, high engagement was associated with greater odds of achieving the composite diabetes risk reduction outcome (odds ratio [OR] 2.59, 95% CI 1.11-6.01; P=.03), &#x2265;5% weight loss (OR 3.31, 95% CI 1.16-9.42; P=.03), and &#x2265;0.2 percentage point reduction in HbA1c (OR 3.57, 95% CI 1.19-10.75; P=.02). Participants most frequently rated weight tracking, physical activity tracking, and the digital body weight scale as the features that were most helpful for achieving their health goals. CONCLUSIONS: Higher engagement with an AI-driven intervention requiring no human intervention was associated with improved diabetes risk reduction. Contrary to concerns about lower digital literacy, older adults engaged with the intervention more than younger adults. Features related to weight and physical activity tracking were most valued by patients in the program. TRIAL REGISTRATION: ClinicalTrials.gov NCT05056376; https://clinicaltrials.gov/study/NCT05056376.

Humans

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype&#x2011;dependent opioid consumption over 72&#xa0;h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non&#x2011;carriers, despite reporting similar subjective pain scores. This consistent genotype&#x2011;dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

Effectiveness of an AI-based home exercise app for rehabilitation of rotator cuff-related shoulder pain: A randomized controlled trial.

BACKGROUND: Rotator cuff-related shoulder pain contributes to disability and healthcare use. Although therapeutic exercise is first-line treatment, limited supervision and adherence may reduce its effectiveness; digital rehabilitation with real-time feedback may address these limitations. OBJECTIVES: To evaluate the effectiveness of adding a digital rehabilitation program to standard physiotherapy on pain, function, fear-avoidance beliefs, and healthcare utilization. DESIGN: Single-center, assessor-blinded, randomized controlled trial with two parallel groups. METHOD: Forty-six adults (mean age 59 years) with rotator cuff-related shoulder pain were randomized to 12 weeks of conventional physiotherapy or physiotherapy plus an AI-based digital rehabilitation program using computer vision for real-time feedback and performance monitoring. Outcomes were assessed at baseline and at 2, 4, and 12 weeks. Pain intensity (NPRS) was primary outcome; secondary outcomes included upper limb function (QuickDASH), fear-avoidance beliefs (FABQ), and post-intervention healthcare utilization. Analyses followed an intention-to-treat approach. RESULTS: Pain reduction exceeded the MCID (1.3) at 4 and 12 weeks. Between-group differences favoured the intervention at Weeks 2 and 4 (MD -0.7; 95% CI -1.13 to -0.14 and MD -1.01; 95% CI -1.8 to -0.2, respectively). Upper limb function improved more at Week 4 (MD -7.3; 95% CI -12.3 to -2.2). FABQ scores decreased more at Week 12 (MD -7.6; 95% CI -14 to -0.5). Fewer participants in the experimental group required post-intervention healthcare (3 vs 10; p&#x202f;=&#x202f;0.02). CONCLUSION: Adding AI-based home exercise app to conventional treatment improve pain and may improve function and reduce healthcare utilization in rotator cuff-related shoulder pain.

Humans

Intervention Without Borders - an Automated Self-Guided AI-Enhanced Psychoeducation Intervention for Dementia Caregivers: Parallel-Group Randomized Waitlist-Controlled Trial.

OBJECTIVE: To examine whether a fully automated, self-guided intervention (PDC30) could improve caregiver well-being over a 1-month waitlist control in an international sample. DESIGN: Randomized waitlist-controlled trial. SETTING: Web-based platform accessible globally. PARTICIPANTS: 441 individuals responded to study promotion on the internet, of whom 274 from 43 countries met the study criteria and were randomized. Eligible participants were adults providing &#x2265;10 care hours weekly to community-dwelling relatives with dementia, scoring &#x2265;5 on Patient Health Questionnaire-9 (PHQ-9), and without recent caregiver intervention. INTERVENTION: Available 24/7, PDC30 is a self-guided, automated intervention consisting of a Guidebook, an AI-powered counseling chatbot, and interactive applications for cognitive-behavioral techniques, relaxation, and caregiver-recipient bonding. MEASUREMENTS: At baseline and follow-ups at 1, 2, and 3 months, depression was assessed by PHQ-9. Secondary outcomes were measured with validated brief versions of anxiety, burden, and positive gains. RESULTS: Intent-to-treat analysis using mixed-effects regression showed treatment x time2 effects on all outcomes except anxiety. At 1-month follow-up, coinciding with exclusive access to PDC30, intervention caregivers showed significant improvements in depression (d = -0.37), burden (d = -0.34), and positive gains (d = 0.42). The differences mostly disappeared after control participants received the intervention, while improvements in both groups were sustained thereafter. Participants reported using the website several times weekly, were generally satisfied with it, and found the chatbot most helpful. CONCLUSIONS: The effects on depression and other outcomes were consistent with those observed for in-person programs, suggesting the viability of well-designed automated intervention. The study demonstrates the feasibility, acceptability, and potential global health impact of PDC30.

Humans

How Following Medical Artificial Intelligence Advice Can Mitigate Malpractice Liability: Cross-National Insights from a Randomized Trial.

Artificial intelligence (AI) increasingly influences clinical decision-making, yet its recommendations may diverge from standard care. Although malpractice concerns are thought to discourage physicians from following AI advice, experimental evidence from the United States suggests the opposite: lay jurors are more likely to hold physicians liable when they reject AI recommendations. Whether this pattern extends to systems in which court-appointed experts, not lay jurors, determine liability remains unknown. Methods: To examine how physicians and laypeople in expert-based and lay-juror legal systems evaluate physicians' acceptance or rejection of AI recommendations, particularly when those recommendations deviate from standard care, we designed a randomized vignette study: a 2 &#xd7; 2 factorial design varying the AI recommendation (standard vs. nonstandard care) and a fictional physician's decision (accept vs. reject). The study was conducted online in 2023 among nationally representative samples of U.S. and German adults and from 2023 to 2024 among German physicians. In total, 387 German physicians, 2291 U.S. adults, and 2283 German adults participated; those not completing the survey or failing attention checks were excluded per preregistered criteria. Participants were randomly assigned to 1 of 4 vignettes, varying the AI recommendation (standard vs. nonstandard care) and physician's decision (accept vs. reject). The reasonableness of the fictional physician's decision was measured, rated by participants on a Likert scale. Results: Analysis, following preregistered exclusion criteria, included 248 German physicians, 1202 U.S. adults, and 1358 German adults. Physicians accepting standard-care AI recommendations were rated more reasonable than those rejecting them (U.S. laypeople: t = 5.36; 95% CI, 0.45-0.97; P < 0.001; German physicians: t = 2.47; 95% CI, 0.14-1.30; P = 0.02; German laypeople: t = 4.14; 95% CI, 0.27-0.76; P < 0.001). Ratings of physicians accepting versus rejecting AI nonstandard-care recommendations were statistically equivalent. Equivalence was tested at an &#x3b1;-value of 0.05 using a two 1-sided tests procedure, reported with 90% CIs per standard convention (U.S. laypeople: t = -4.90; 90% CI, -0.1 to 0.36; P < 0.001; German physicians: t = -1.76; 90% CI, -0.12 to 0.67; P = 0.04; German laypeople: t = 5.35; 90% CI, -0.35 to 0.06; P < 0.001). Conclusion: Across the United States and Germany, samples representative of lay jurors and court-appointed experts viewed accepting standard-care AI advice as more reasonable, whereas accepting or rejecting nonstandard-care AI advice was judged similarly. Contrary to predictions, malpractice liability regimes do not necessarily pose a barrier to AI use in precision medicine.

Artificial Intelligence

Intravenous thrombolysis for ischemic stroke in extended time window selected with CT perfusion: a systematic review and meta-analysis.

PURPOSE: Recent randomized controlled trials (RCTs) have provided new evidence regarding the efficacy and safety of intravenous thrombolysis (IVT) in patients with acute ischemic stroke (AIS) presenting within the extended time window (ETW). We performed a systematic review and meta-analysis to evaluate the efficacy and safety of IVT, in patients treated within the ETW and selected with perfusion imaging criteria, predominantly computed tomography perfusion (CTP). METHODS: A systematic review and meta-analysis, registered in PROSPERO, was conducted including all available RCTs comparing IVT with best medical treatment (BMT) in patients with AIS within the ETW, selected using advanced perfusion imaging criteria. The predefined efficacy outcomes were excellent functional outcome and good functional outcome at 3 months. The safety endpoints included symptomatic intracranial hemorrhage (sICH) and all-cause mortality at 90 days. RESULTS: Six RCTs, including 1182 patients treated with IVT and 1176 patients receiving BMT, were included. IVT was associated with a higher likelihood of achieving excellent and good functional outcomes at 3 months. Exploratory subgroup analyses by treatment timing suggested consistent findings up to 24 hours. No significant difference in 90-day mortality was observed between groups, whereas IVT was associated with an increased risk of sICH. CONCLUSION: Treatment with IVT in the ETW (4.5-24 h) in patients selected using advanced perfusion imaging, predominantly CTP, may be associated with improved functional outcomes in patients with AIS. Although IVT was associated with an increased risk of sICH, no significant increase in 90-day mortality was observed. PROSPERO REGISTRATION: CRD420261304314.

Aged

Beyond da Vinci: a systematic review of next-generation multiport robotic platforms in pediatric surgery.

As robotic surgery expands beyond the da Vinci platform, the relevance of new-generation systems to pediatric patients remains uncertain. This review examined the technical characteristics, applications, and perioperative outcomes of alternative multiport robotic platforms in children. PubMed/MEDLINE, Web of Science, Scopus, and the Cochrane Library were searched through 28 February 2026 in accordance with PRISMA 2020. Owing to clinical heterogeneity, findings were synthesized narratively, with Wilson 95% confidence intervals for key binary outcomes. Six studies reported 166 patients across 27 procedure types. Senhance accounted for 164 patients, while Hugo RAS and Hinotori were each represented by one patient. Ages ranged from 15&#xa0;days to 17&#xa0;years and weights from 3.8 to more than 100&#xa0;kg. Senhance was the only platform used with 3-mm robotic instruments. Conversion or planned escalation occurred in 19 patients (11.4%; 95% CI, 7.5-17.2%). Four intraoperative complications were reported (2.4%; 95% CI, 0.9-6.0%). Across all reports, 20 patients experienced postoperative complications (12.0%; 95% CI, 7.9-17.9%). In the largest cohort, seven patients required reintervention (4.6%; 95% CI, 2.2-9.2%) and seven were readmitted (4.6%; 95% CI, 2.2-9.2%). No deaths or comparative pediatric studies were reported. Published experience remains sparse and is almost entirely limited to Senhance. Current evidence describes early clinical use without establishing comparative safety, effectiveness, or platform superiority. Prospective multicenter studies with standardized reporting are needed.

Humans

A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.

BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.

Humans