Search PubMedSearch

PubMed · 42607666

A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.

Abstract

BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.

Explore related subjects

Keep this discovery

BibTeXRIS

Peixing Wan, Zigeng Huang, Haoquan Huang, Guihan Liang, Mingyan Guo, Sulin Xu, Xinti Sun, Wenjun Tang, Dajun Pei, Jing Chen, Yulan Nie, Shaofen Deng, Yizhi Zhou, Hongru Duan, Rong Zhang, Xiaoying Xu, Erfu Hai, Minghui Cao, Erping Long. 2026-08-17. A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.. https://doi.org/10.1016/j.medj.2026.101263

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Pedagogical Efficacy of LLM-Generated Synthetic Data Versus Real-World Clinical Records: A Randomized Controlled Non-Inferiority Trial.

BACKGROUND: Expert-reviewed clinical cases generated by large language models (LLMs) may supplement case resources in medical education, but their short-term educational performance relative to real-case-derived teaching materials remains uncertain. We compared immediate post-training test performance after teaching with the two types of case materials and assessed non-inferiority against a prespecified margin. METHODS: We conducted a prospective, parallel-group, randomized non-inferiority trial. Through the Wenjuanxing online platform, participants were randomized 1:1 to learn with either real-case-derived teaching cases compiled by clinicians and reviewed by experts or AI-generated clinical cases produced by Gemini 3.0 Pro from fully de-identified matched real cases and reviewed by three senior general surgery specialists with full-professor rank. The primary outcome was the total score on an independent 10-item immediate post-training test (0-10 points), with a prespecified non-inferiority margin of -0.5 points. Secondary outcomes included the training-phase performance score, learning efficiency index, single-item mental effort rating, case realism, and case-source judgment. RESULTS: A total of 403 participants were randomized, of whom 386 were included in the modified intention-to-treat analysis: 192 in the real-case group and 194 in the AI-generated case group. The mean post-training test score was 4.95 (SD, 3.35) in the real-case group and 4.61 (SD, 3.35) in the AI-generated case group. The mean difference (AI-generated minus real-case group) was -0.335 points (95% CI, -1.006 to 0.337). Because the lower bound of the confidence interval was below the prespecified non-inferiority margin of -0.5 points, non-inferiority was not demonstrated (one-sided P = 0.314). No significant between-group differences were observed in the training-phase performance score, learning efficiency index, or single-item mental effort rating. AI-generated cases received lower realism ratings for Level 3 cases. The proportion of participants with at least one high-confidence completely incorrect response was 1.6% in the real-case group and 2.1% in the AI-generated case group. CONCLUSIONS: In this short-term, text-based online case-learning setting, no statistically significant between-group difference was observed in immediate post-training test performance; however, non-inferiority of AI-generated clinical cases relative to real-case-derived teaching materials was not demonstrated.

Humans

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Evaluating a culturally adapted question prompt list to improve end-of-life communication among indonesian migrant caregivers: A randomized controlled trial with qualitative insights.

OBJECTIVE: Indonesian caregivers serve as essential providers of end-of-life (EOL) care in Taiwan. But often face communication challenges due to language, cultural, and hierarchical barriers. This study evaluated the effectiveness of a culturally adapted Question Prompt List (QPL). METHODS: This study employed a two-arm randomized controlled trial design supplemented with qualitative interviews. The study was conducted in a hospice ward and home care setting within a medical center in Taiwan. A total of sixty Indonesian caregivers were recruited and randomly assigned to either the intervention group (n&#x202f;=&#x202f;30) or the control group (n&#x202f;=&#x202f;30). The intervention group received routine end-of-life (EOL) education along with a culturally adapted Question Prompt List (QPL), which consisted of 37 items covering domains including the dying process, emotional support, communication, symptom management, and care decision-making. The control group received routine EOL education. Outcome measures included caregiving preparedness, communication self-efficacy, satisfaction, and question-asking behavior. In addition, semi-structured interviews were conducted with eight participants, and the data were analyzed using thematic content analysis. RESULTS: Analysis of covariance revealed no statistically significant between-group differences in caregiving preparedness (F = 1.58, p&#x202f;=&#x202f;.215 [-0.41, 0.44]) or communication selfefficacy (F = 0.83, p&#x202f;=&#x202f;.366 [-0.44, 0.79]). However, communication satisfaction was significantly higher in the intervention group (F = 4.19, p&#x202f;<&#x202f;.05 [0.04, 0.44]). The number of questions asked was also significantly higher in the intervention group (t&#x202f;=&#x202f;-4.35, p&#x202f;<&#x202f;.001 [-5.41, -1.98]). Thematic analysis of qualitative data identified 4 themes and 14 subthemes, illustrating how the QPL reduced anxiety, clarified care needs, and improved confidence. CONCLUSIONS: A culturally adapted QPL can enhance communication engagement and satisfaction among migrant caregivers. PRACTICE IMPLICATIONS: Integrating culturally tailored QPLs into caregiver education and palliative care practice may promote more inclusive and effective communication.

Humans