Search PubMedSearch

SEARCH · Search PubMed

Results for “Speech-language therapy”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

634 recordsLinked to original sources

Novel Proactive Speech-Language Intervention Is More Effective Than Usual Care: Randomized Controlled Trial of Babble Boot Camp for Infants With Classic Galactosemia.

PURPOSE: Speech and language disorders cannot be diagnosed and treated until children are approximately 2-4 years old. To investigate whether these disorders can be prevented, we developed and trialed Babble Boot Camp (BBC), the first proactive sustained intervention starting with precursor skills including cooing and babbling. METHOD: Participants were two randomly assigned groups of 22 infants with classic galactosemia, a metabolic disease with known risks for severe speech and language disorders. One group started BBC at under 6 months of age, and the other started at 15 months of age, both completing BBC at 24 months of age. Coached by a speech-language pathologist in weekly telehealth sessions, caregivers implemented BBC activities and routines daily at home. A typical control group and a group of children with classic galactosemia who received usual care participated as well. All children completed standardized assessments of speech and language at postintervention. RESULTS: Assessment scores showed that BBC was more effective than usual care for both intervention groups. Greatest benefits were seen in the group that started at or before 6 months of age, with a proportion of clinically concerning scores equal to that in the typically developing peers. No effects of sex, genotype, or milk consumption were evident in the outcomes. CONCLUSIONS: Findings motivate a paradigm shift from deficit-based to proactive approaches for infants with classic galactosemia. BBC is extensible to many other disorders, with trials currently underway for infants with Down syndrome and infants born preterm.

Humans

Comparing Traditional Motor Speech Practice to Contextualized Speech Practice in Preschoolers With Childhood Apraxia of Speech.

PURPOSE: The aim of this study was to compare retention of real-word targets across practice conditions (contextualized vs. motor-only) within a modified integral stimulation treatment for preschoolers with childhood apraxia of speech (CAS). METHOD: A single-subject experimental design with alternating treatments was used with matched target sets randomly assigned to contextualized practice, motor-only practice, or no treatment. Three preschoolers with CAS completed 18 therapy sessions, each consisting of two 25-min blocks: one contextualized practice and one motor-only practice. Order of practice was randomized each visit. Changes in percent phonemes correct (PPC) and lexical stress accuracy, derived from blinded transcription, were explored with visual analysis and effect sizes (standardized mean difference, d statistic). RESULTS: Meaningful improvements (d > 1) were observed in PPC across words treated in contextualized practice for all three children immediately posttreatment and for two of three children at the 1-month follow-up. Meaningful improvements in the motor-only condition were observed in two of three children immediately posttreatment and at follow-up. No meaningful changes were observed in lexical stress across any conditions in any participant. CONCLUSIONS: This study provides preliminary support for the feasibility of a modified integral stimulation therapy that incorporates elements of linguistically grounded therapies (linguistic retrieval, recasts, expansions) that may facilitate target retention in some preschoolers with CAS. However, other elements should be explored in conjunction with integral stimulation to maximize clinical outcomes. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.33228981.

Humans

Modulating sentence comprehension in people with aphasia through anodal tDCS: A double-blind randomized cross-over study.

This double-blind randomized cross-over study investigated the effects of perilesional anodal transcranial direct current stimulation (AtDCS) combined with speech-language therapy on sentence comprehension in eight individuals with chronic nonfluent agrammatic aphasia. The behavioral therapy consisted of an intensive comprehension treatment including drilling in sentence-to-picture matching and Mapping Therapy. Each participant underwent both the anodal tDCS and sham stimulation conditions (five received sham first followed by real stimulation, and the remaining three the reverse sequence), with each condition paired with the same behavioral treatment and separated by a four-month washout period. Stimulation was applied over the perilesional area (left BA6) for 20 min during daily 40-min therapy sessions over four consecutive weeks. Sentence comprehension was assessed with the RiComprendo battery and functional communication with the Communicative Effectiveness Index (CETI). Data were analyzed using paired t-tests, Bayesian analyses, and linear mixed-effects models to control for baseline performance and individual variability. Both stimulation conditions produced significant pre-to-post improvements in sentence comprehension, particularly for syntactically complex structures such as passives and center-embedded object relatives. However, gains were overall greater following AtDCS, as reflected in larger effect sizes, stronger Bayes factors, and a significant treatment effect in the mixed-effects models. Only the AtDCS condition yielded significant improvements in self-perceived comprehension abilities on the CETI. These findings suggest that AtDCS over perilesional cortical areas may boost the effects of traditional language therapy on sentence comprehension, supporting its feasibility and potential as an adjuvant intervention in post-stroke aphasia rehabilitation.

Humans

Effectiveness of Caregiver-Mediated Spoken Language Interventions for Children Under Five at Risk of Developmental Language Disorder: A Systematic Review and Meta-Analysis.

BACKGROUND AND AIMS: Caregiver-mediated interventions are commonly used by Speech and Language Therapists to support early language development. Developmental Language Disorder (DLD) is associated with reduced quality of life throughout the lifespan. Understanding factors that predict intervention success is essential for developing appropriate, cost-effective therapy provision for the approximately 12% of preschool children who present with early markers for Developmental Language Disorder (DLD). This systematic review and meta-analysis examined the effectiveness of caregiver-mediated spoken language interventions for under-fives at risk of DLD, and factors influencing intervention effectiveness. METHODS: A systematic review following PRISMA guidelines was conducted. Five electronic databases were searched to identify experimental studies comparing caregiver-mediated spoken language interventions to control conditions in under-fives presenting with risk factors for DLD. Risk factors included prematurity, socioeconomic factors, caregiver language development concerns, and formal or informal language screening or assessment scores. Twenty-six experimental studies with 1407 child participants were included in qualitative synthesis. Meta-analysis was performed on nine Randomised Controlled Trials involving 947 children. RESULTS: Effectiveness was examined for outcomes including child language gains, child wellbeing, inclusion and attainment. Meta-analysis indicated a significant effect of caregiver-mediated spoken language interventions on language outcomes compared to treatment-as-usual, non-language intervention or waitlist control conditions. Non-language outcomes were evaluated via qualitative synthesis. Interventions significantly improved language development trajectories for under-fives presenting with risk factors or early markers for DLD. CONCLUSION AND IMPLICATIONS: This review contributes to the growing evidence base demonstrating that caregiver-mediated interventions can positively impact language development and wellbeing outcomes for children under five at risk of DLD. These findings support the implementation of caregiver-mediated environmental language interventions in clinical practice to maximise accessibility and cost-effectiveness while delivering optimal outcomes for vulnerable populations. WHAT THIS PAPER ADDS: What is already known on this subject Previous research on caregiver-mediated spoken language interventions has highlighted gaps in the evidence regarding the impact of risk factors, demographic characteristics, dosage and intervention components on child language outcomes. Developmental Language Disorder has relatively high population prevalence, estimated at 7%. Prevalence is associated with risk factors including low household socioeconomic status (SES), prematurity and late language emergence. In contrast to its prevalence, there is low public and professional awareness of DLD and a low diagnostic rate. Therefore, a strengthened evidence base and additional insights into the factors affecting success of family-based interventions is important in order to increase the effectiveness of service provision and care planning for this underserved population. Timely and effective intervention with young children presenting with early markers for DLD has the potential to offer lifelong improvement to their wellbeing, inclusion and attainment outcomes. Recent systematic reviews of the effectiveness of caregiver-mediated language interventions had differences in population age range and diagnostic inclusion criteria. What this paper adds to existing knowledge Our review examines the effectiveness of caregiver-mediated early spoken language interventions on child language, attainment and wellbeing, and on caregiver self-efficacy and adherence to language support strategies. Our population was children under five presenting with risk factors for Developmental Language Disorder, in the absence of other neurodevelopmental or genetic conditions such as intellectual disability or autism. This review adds depth and detail to the evidence base supporting the effectiveness of caregiver-mediated spoken language interventions in improving outcomes for this population of young children, and factors that influence their success. What are the potential or actual clinical implications of this work? The high prevalence of Developmental Language Disorder, estimated at around 7% of the population, and the strong association with risk factors including low SES, prematurity and late language emergence, coupled with the low awareness of DLD and low diagnostic rate, mean that a strengthened evidence base and additional insights into the factors affecting success of family-based interventions can increase the effectiveness of service provision and care planning for this population. Timely and effective intervention in this group of young children has the potential to improve wellbeing and attainment outcomes across the lifespan. This review contributes to our understanding of how to implement cost-effective, socially valid and maximally engaging partnership working with families of young children at risk for DLD.

Humans

Characterizing Caregiver-Child Interactions Through a Transactional Lens: A Baseline Analysis of a Caregiver-Implemented Intervention.

PURPOSE: This study was motivated by the transactional model of development and examined the reciprocal influences that children and caregivers have on caregiver-child interactions (CCXs) prior to a caregiver-implemented intervention. We tested whether child communication characteristics were associated with caregiver strategy use and whether these strategies, in turn, influenced children's communication to understand how caregivers and children mutually shaped the language learning environment. METHOD: Caregiver-child dyads (N = 105) were participants in a randomized controlled trial. CCXs were collected when children were approximately 30 months of age, transcribed, and coded for four caregiver language facilitation strategies and child communication variables. RESULTS: Least Absolute Shrinkage and Selection Operator regression and postselection inference indicated that child communication characteristics in CCXs were associated with both the frequency and type of strategies caregivers used. Children's overall communication acts were significantly associated with caregiver use of vocabulary strategies, whereas children's vocabulary diversity was significantly associated with caregiver use of sentence strategies. Mixed-effects logistic regression demonstrated that all four caregiver strategies significantly increased the likelihood of spontaneous lexical overlap in subsequent child turns. CONCLUSIONS: Prior to the intervention, caregivers and children reciprocally shaped the language environment. This supports a transactional perspective and warrants further consideration of reciprocal influences when assessing the impact of caregiver-implemented interventions. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.32995796.

Humans

Functional Reading Activities to Motivate and Empower: Maintenance of Reading Outcomes for Young Adults With Intellectual and Developmental Disabilities Following a Randomized Controlled Trial.

PURPOSE: This study examined whether the effects of Functional Reading Activities to Motivate and Empower (FRAME), a functional, strategy-based reading comprehension intervention for young adults with intellectual and developmental disabilities (IDDs), were maintained 6 months following the completion of the intervention and explored participants' perceptions of the intervention's feasibility, relevance, and perceived impact. METHOD: Participants were 44 young adults with IDDs (ages 18-26 years) who participated in a previously reported randomized controlled trial (FRAME participants: n = 23; controls: n = 21). Trial outcomes were assessed via telepractice at pretest, posttest, and 6-month follow-up. Six-month maintenance analyses focused on outcomes that demonstrated significant posttest group differences: use of (a) reading comprehension strategies (proximal) and (b) reading comprehension questions (distal). Participant perceptions (social validity) were collected post-intervention from FRAME participants using a structured interview protocol with closed- and open-ended items. RESULTS: At the 6-month follow-up, FRAME participants demonstrated sustained but reduced improvements in strategy use relative to controls (p = .040). Between-groups differences were not maintained for reading comprehension questions (p = .091). Participants reported high acceptability and perceived relevance of FRAME, with qualitative themes reflecting perceived improvements in comprehension, self-improvement, and increased independence. CONCLUSION: Findings suggest that FRAME supports sustainable gains in reading comprehension strategy use and is perceived as meaningful and feasible for young adults with IDDs, although additional supports may be needed to promote sustained improvements in distal comprehension outcomes. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.33307218.

Humans

Characterizing Submental Neuromuscular Activity of Swallowing Rehabilitation: An Electromyographic Evaluation of Rehabilitative Maneuvers in Healthy Adults.

PURPOSE: The effortful swallow (ES), the Mendelsohn maneuver (MM), and isometric tongue presses (TPs) are widely used swallowing maneuvers/exercises to improve elements of swallowing, such as muscle strength and biomechanics. However, the underlying neuromuscular mechanisms of these exercises remain unclear, potentially limiting our ability to specify treatment targets and improve treatment efficacy. This study aimed to compare submental neuromuscular activation patterns during the ES, MM, TPs, and typical swallows in healthy young and older adults. METHOD: As part of a larger randomized crossover validation study, 60 healthy adults (30 young and 30 older) completed typical swallows and three maneuvers using a wearable submental surface electromyographic (sEMG) system (i-Phagia). Outcome variables included (a) normalized mean sEMG amplitude and (b) time to peak sEMG amplitude. Linear mixed models were used to examine effects of task, age, and sex on both outcomes. RESULTS: Normalized mean amplitude was significantly different across tasks. Post hoc pairwise comparisons confirmed that all three maneuvers produced higher normalized mean sEMG amplitude than typical swallows, with the ES eliciting the highest amplitude across age groups. Time to peak amplitude differed significantly across tasks, with typical swallows requiring the shortest time to reach peak amplitude, followed by the ES, MM, and TP. CONCLUSIONS: The ES required the highest neuromuscular effort with the shortest time to reach peak amplitude, suggesting its potential for targeting submental muscle power. Typical swallows required the least neuromuscular effort with the shortest time to reach peak amplitude, suggesting their potential for training submental muscle speed. The MM and TP may also improve submental muscle strength; however, given their temporal requirements (longer durations), they may be more beneficial for targeting coordination and endurance, though more research in this area is warranted. These findings underscore the importance of task-specific neuromuscular profiling to inform mechanism-based swallowing rehabilitation. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.32948549.

Humans

The voice clone intelligibility benefit in noise in middle-aged listeners.

Research with younger adults showed that cloned voices are more intelligible than human voices in noise, with a benefit of 13.4%. This study tested whether this benefit extends to 40 middle-aged listeners (45-65 years), as this population may show emerging difficulties with speech-in-noise. Participants recognised sentences by ten human voices and ten voice clones in four noise levels. Cloned voices were 11.8% more intelligible, with benefits enhanced at the two most severe noise levels (15.9% at -6 dB and 17.5% at -3 dB), suggesting cloned speech enhanced perception in middle-aged listeners, potentially by reducing listening effort and compensating for emerging age-related auditory-cognitive decline.

Humans

Effects of anodal transcranial direct current stimulation over the right primary motor cortex on a sequential motor finger tapping task in developmental stuttering.

INTRODUCTION: This study investigates the impact of anodal transcranial direct current stimulation (tDCS) on non-speech sequential motor practice in adults who stutter (AWS), compared to non-stuttering controls (ANS). Recent research has explored the effects of tDCS on speech fluency in stuttering. However, its effect on non-speech motor tasks has not yet been studied. METHODS: 20 AWS and 30 ANS right-handed participants were randomly assigned to anodal or sham tDCS conditions, performing a sequential finger tapping task. We targeted over the right primary motor cortex, stimulating at 2 mA for 20 min. Sequence duration and reaction time were analyzed. RESULTS: AWS analysis revealed that the anodal condition had significantly slower reaction times in the second half of the task compared to sham. For sequence durations, AWS in the anodal condition had slower overall sequence durations than the sham condition. However, there were no block-by-block differences in sequence duration. When comparing AWS and ANS, no significant differences were observed for sequence duration. However, there were significant differences in reaction time between AWS and ANS, specifically in earlier blocks. Additionally, there was no significant Group × Condition interaction. DISCUSSION: The findings suggest that anodal stimulation impeded finger sequencing in AWS, showing overall slower sequence durations and a diminishing effect on reaction times in the second half of the experiment, suggesting anodal tDCS may interact uniquely with the neural mechanisms in stuttering. Future studies should explore the effects of anodal tDCS on non-speech motor tasks to gain a broader understanding of its impact on motor control and motor learning.

Humans

Evaluating Patient Satisfaction and Oral Health Impact Profile-14 (OHIP-14): A Multicenter Crossover Study Comparing Selective Pressure Impression Conventional Dentures with Mucostatic Digital Dentures.

PURPOSE: To compare patient satisfaction and oral health impact between individuals receiving complete dentures made by digital methods and those using conventional techniques. MATERIALS AND METHODS: In this randomized crossover clinical trial, 23 patients aged 40 years and older with completely edentulous arches were enrolled at three treatment centers. Each participant received two sets of complete dentures: one set created using conventional methods (selective pressure impression) and the other through digital techniques (mucostatic digital impression). The order of denture placement was randomized, with each set used for 4 weeks. A trained specialist administered treatments alongside research tools, including a general information questionnaire, a denture satisfaction survey, and the OHIP-14 interview tool. Statistical analysis was conducted using Mann-Whitney U test. RESULTS: Participants with digital dentures reported significantly higher satisfaction regarding treatment duration, comfort, confidence, chewing ability, esthetics, and overall satisfaction compared to those with conventional dentures. There were no significant differences in satisfaction concerning speech and pronunciation. Overall, the oral health impact on quality of life was similar between denture types, but participants indicated improved quality of life while using dentures compared to being edentulous. CONCLUSIONS: Patients with digital dentures exhibited greater satisfaction across various domains compared to those with conventional dentures, despite similar satisfaction levels in speech and pronunciation. The impact on quality of life was comparable between both types, as measured by the OHIP-14.

Humans

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans

Systematic Review of Symptoms of Catatonia in Autism Spectrum Disorder.

Catatonia is a complex neuropsychiatric syndrome characterized by disturbances in mood, motor function, behavior and speech. It is increasingly recognized in individuals with autism spectrum disorder (ASD), although its identification remains challenging due to the overlapping clinical features of the two conditions. Shared characteristics, such as echophenomena, mannerisms, social indifference and repetitive behaviors can obscure accurate diagnosis. Although reports suggest a significant prevalence of catatonia among individuals with ASD, the condition remains poorly understood and frequently under recognized, leading to substantial diagnostic and treatment challenges. A systematic review was conducted to characterize the symptoms of catatonia in individuals with ASD. The literature search included peer-reviewed journal articles published in English from 1980 onward, focusing on studies examining co-occurring catatonia and ASD. A qualitative framework analysis was implemented to evaluate 45 peer-reviewed studies, with findings interpreted in relation to, and extending beyond, the diagnostic criteria for catatonia outlined in the International Classification of Diseases, 11th revision (ICD-11). The objective was to identify symptom patterns extending beyond current diagnostic frameworks and to support improved clinical recognition and diagnostic precision in ASD populations. The review identified six primary symptom clusters associated with catatonia in individuals with ASD: (1) psychomotor activity, (2) speech disturbances, (3) changes in behavior/skills/functions, (4) mental health symptoms, (5) physiological symptoms, and (6) symptoms related to arousal and awareness. Notably, several symptoms observed within these clusters are not currently included in the ICD-11 diagnostic criteria for catatonia. These additional symptoms include tics, motor compliance, incoherent speech, self-injury, impaired cognition, and appetite changes, suggesting a broader clinical presentation of catatonia in ASD populations than is presently captured in existing diagnostic frameworks. The findings of this review highlight the significance of enhancing clinicians' awareness and understanding of how catatonia manifests in individuals with ASD. Most notably, six symptom clusters, psychomotor changes, speech disturbances, behavioral and functional regression, affective and psychiatric symptoms, physiological symptoms, and arousal/awareness disturbances, were observed. Several symptoms identified in this review are not included in the current diagnostic criteria, and their recognition may facilitate in earlier identification and timely intervention, potentially preventing the severe consequences of untreated catatonia in this population.

Humans

Using the OPTIMAL Theory to Optimize Aerodynamics in Respiratory Training for Healthy Adults and Individuals With Parkinson's Disease.

BACKGROUND: The OPTIMAL (Optimizing Performance Through Intrinsic Motivation and Attention for Learning) theory is a motor learning framework proposing that optimizing intrinsic motivation enhances motor performance and learning. The theory identifies three key components-Enhanced Expectancies (EE), Autonomy Support (AS) and External Focus of Attention (EF)-which facilitate more efficient, goal-directed movement. These components have been shown to improve motor outcomes in limb-based tasks; however, their application to respiratory training, particularly in clinical contexts such as voice and swallowing therapy in patients with Parkinson's disease (pwPD), has not yet been systematically explored. AIMS: This study aimed to investigate whether implementing OPTIMAL theory strategies during a respiratory muscle strength training (RMST) task improves immediate respiratory motor performance in healthy adults and pwPD. Additionally, we aimed to examine the effects of these strategies on motivation and cognitive engagement. METHODS: This quasi-randomized, single-session trial included 47 participants: Healthy CONTROL (n = 17), Healthy OPTIMAL (n = 16) and PD OPTIMAL (n = 14). Healthy participants were quasi-randomly assigned to either intervention or control conditions, whereas pwPD completed the intervention only. All participants completed a single respiratory session that included baseline, practice and retention phases. Outcome measures included peak expiratory flow, cough peak expiratory flow, cognitive engagement (EEG-based Cognitive Engagement Index) and self-administered motivation questionnaire. OUTCOMES AND RESULTS: Exhalation force improved from baseline to retention in the Healthy OPTIMAL group (baseline: M = 296 L/min; retention: M = 338 L/min; p < 0.001) and the PD OPTIMAL group (baseline: M = 315 L/min; retention: M = 370 L/min; p < 0.0001), but not in the Healthy CONTROL group (p > 0.05). No significant changes in cough strength were observed in any group. No correlations were found between cognitive engagement and exhalation force or motivation scores. However, motivation increased more in the Healthy OPTIMAL group (Questionnaire 1: M = 57.2; Questionnaire 2: M = 60.7) and the PD OPTIMAL group (Questionnaire 1: M = 60.1; Questionnaire 2: M = 62.8) than in the Healthy CONTROL group (Questionnaire 1: M = 61.1; Questionnaire 2: M = 62.5). CONCLUSIONS AND IMPLICATIONS: Implementing the OPTIMAL theory enhances immediate respiratory motor performance in both healthy participants and pwPD. OPTIMAL theory has clinical value in voice and swallowing therapy, although further research is needed to establish long-term efficacy and clinical impact. WHAT THIS PAPER ADDS: What is already known on the subject Motivation is a critical factor in rehabilitation. The OPTIMAL theory has been shown to improve both motivation and motor performance in limb-based tasks. Its impact on respiratory training, however, has not been previously examined. What this paper adds to the existing knowledge This study shows that applying OPTIMAL strategies during a respiratory muscle strength training task significantly improved peak expiratory flow in both healthy adults and people with Parkinson's disease. What are the potential or clinical implications of this work? Integrating the OPTIMAL theory principles into respiratory therapy may enhance motor outcomes, supporting voice, swallowing and cough rehabilitation.

Humans

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

A bimodal large language model reduces misalignment in patient education: A double-blinded randomized trial.

BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.

Humans

BMT4me En Espa&#xf1;ol: Multisite Feasibility and Usability Testing of a Spanish-Language mHealth Adherence Support App for Spanish-Speaking Caregivers of Children After Hematopoietic Stem Cell Transplantation and Cancer Treatment.

BACKGROUND: Medication nonadherence during the first 100 days after pediatric hematopoietic stem cell transplantation (HSCT) and during oncology treatment increases risk for complications. BMT4me is a caregiver-facing mobile health (mHealth) application providing medication reminders, symptom tracking, and note-taking features to support medication management. Spanish-speaking caregivers are frequently excluded from digital adherence interventions due to the lack of language-accessible tools. PROCEDURE: We conducted a multisite, mixed-methods usability testing of a Spanish-language version of BMT4me ("BMT4me en Espa&#xf1;ol") with Spanish-speaking caregivers of children (ages 2-17 years) post-HSCT or with an oncology diagnosis on active treatment. Caregivers completed a facilitated, three-step usability session (unobtrusive observation, interactive observation, and debriefing), followed by a semi-structured interview, and then completed the system usability scale (SUS). Quantitative outcomes were summarized descriptively; qualitative data were analyzed using content analysis with constant comparison. RESULTS: Fifteen participants enrolled at each site for a total of 30 participants. Across both sites, the recruitment rate was 91%. All participants completed all parts of the study. The SUS score (M&#xa0;=&#xa0;80.09; SD&#xa0;=&#xa0;17.35) was above average (>68). Two key qualitative themes emerged: (1) the perceived positive impact of BMT4me on managing a serious illness and (2) the acceptance and sociocultural relevance of BMT4me for Spanish-speaking families. Caregivers also shared suggestions to add educational content and multiuser functionalities to BMT4me. CONCLUSIONS: The acceptance and perceived positive impact of the Spanish BMT4me app indicates that socioculturally relevant, Spanish mHealth interventions have strong potential to support Spanish-speaking caregivers in pediatric oncology and HSCT settings. CLINICAL TRIALS NCT: NCT06361173.

Adolescent

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Sports Medicine