Search PubMedSearch

SEARCH · Search PubMed

Results for “motor learning”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

217 records · Page 4Linked to original sources

Effects of extended problem-based learning interventions on undergraduate nursing education: A systematic review.

OBJECTIVE: Exploring the effects of long-term PBL (problem-based learning) intervention on undergraduate nursing students. METHODS: The article retrieved literature from CINAHL Complete, Academic Search Complete, Web of Science, PubMed, EMBASE, OVID, and Cochrane Library up to January 2025. Studies had to meet all of these criteria: (1) They used a randomized controlled trial (RCT) and quasi-experimental design. (2) The PBL pedagogy intervention lasted 4 weeks or longer. (3) The participants were undergraduate nursing students. (4) They reported primary outcomes. These included critical thinking, problem-solving skills and self-directed learning. Two researchers screened articles, extracted data, and assessed quality independently using blinding. They used Cochrane ROB2 for RCTs and ROBINS-I for quasi-experimental studies to judge bias risk. Meta-analysis was performed using RevMan 5.4 software. For continuous variables, standardized mean difference (SMD) and 95% confidence interval were calculated. Heterogeneity was assessed by I2 statistic. When I2 > 50%, sensitivity analysis was conducted. The source of heterogeneity was explored by excluding studies one by one. The primary outcomes included standardized critical thinking, problem-solving, and self-directed learning assessment results. RESULTS: A total of 11 randomized controlled trials and quasi-experimental studies were retrieved and included for meta-analysis. The experimental group significantly outperformed the control group in critical thinking, problem-solving, and self-directed learning, with differences being statistically significant (P ≤ 0.05). However, high heterogeneity was observed. After sensitivity analysis, the heterogeneity was reduced and the results remained statistically significant, indicating that the findings were not solely dependent on the excluded studies.

Problem-Based Learning

Comparative effectiveness of game-based learning modalities in nursing and medical education: a systematic review and Bayesian network meta-analysis.

BACKGROUND: Game-based learning (GBL) is increasingly used in healthcare education, but educators must choose among diverse modalities (e.g., quiz platforms, apps, serious games and metaverse environments). Comparative evidence on which modalities perform best across learning domains (knowledge, attitudes, and practice) remains limited. AIM: To compare the effects of distinct GBL modalities on knowledge, attitudes, and practice outcomes in nursing and medical education and to explore whether comparative effects differ by learner group (pre-licensure students and in-service professionals). DESIGN: PRISMA-NMA-aligned systematic review and Bayesian network meta-analysis. METHODS: We searched eight databases and trial registries through September 2, 2024, for randomized controlled trials comparing GBL with traditional teaching (TT). Outcomes were transformed to a 0-100 scale and analysed as change from baseline in Bayesian consistency models; random-effects models were selected using deviance information criterion (DIC). Risk of bias was assessed using RoB 2. We report mean differences (MDs) with 95% credible intervals (CrIs) versus TT, ranking probabilities, and subgroup NMAs by learner group. RESULTS: Thirty-one RCTs (n = 3439) were included; 15 contributed complete data to the network. Risk of bias was low in 15 trials and raised some concerns in 16. The network was modest for knowledge (11 trials) and sparse for attitudes (3) and practice (4). Compared with TT, metaverse-based learning showed improved attitudes (MD 15; 95% CrI 12 to 18), based on a single trial. For knowledge and practice, Kahoot-based quizzes (MD 9.1; 95% CrI -8.9 to 27) and app-based learning (MD 4.6; 95% CrI -4.4 to 14) had the highest estimated mean improvements, but credible intervals were wide and included the null for most comparisons. Subgroup rankings differed by learner group, but several comparisons were imprecise and uncertainty was substantial, particularly in sparse networks. CONCLUSIONS: GBL modalities may improve learning outcomes compared with TT, but relative effects appear domain-specific and the certainty of rankings is limited by sparse evidence and imprecision. Future trials should prioritise head-to-head comparisons, robust outcome measurement, and longer-term retention and transfer outcomes in both student and in-service populations.

Humans

Nondominant Hand Training in Laparoscopy for Surgical Interns: Feasibility and Impact.

OBJECTIVE: Laparoscopy requires bimanual proficiency, yet early trainees demonstrate underdeveloped nondominant hand (NDH) performance. Although deliberate practice of NDH skill contributes to overall performance, NDH training is rarely incorporated into residency simulation curricula and has not been formally evaluated in surgical trainees. We assessed feasibility and impact of integrating structured NDH training with established laparoscopic curriculum for surgery interns. DESIGN: Prospective, single-institution randomized pilot study. Interns were assigned the standard 4-week curriculum of laparoscopic dominant hand and bimanual tasks (Control) or completed assigned NDH tasks in addition to the standard curriculum (Intervention). Feasibility was determined by assigned task completion, daily standard and NDH-specific self-reported practice time, and improvement in bimanual task performance. Performance was video recorded weekly and assessed by blinded evaluators using MISTELS and GOALS scoring. Cognitive workload during laparoscopic tasks was measured via NASA-TLX. Exploratory analyses were conducted within a Bayesian framework. SETTING: A single academic institution with a surgical simulation training program. PARTICIPANTS: General surgery interns on their 4-week simulation rotation. RESULTS: Eleven general surgery interns (6 intervention, 5 controls; all right-hand dominant) completed the study with 100% task completion and practice log compliance. Both groups improved in bimanual performance and perceived cognitive load. Reduction in cognitive workload during bimanual task performance was greater in the NDH group. Time spent on NDH practice over 4 weeks was associated with improved bimanual performance, independent of time spent on standard curriculum tasks. CONCLUSIONS: Structured NDH training is feasible to implement within an existing curriculum and reduces perceived cognitive workload during bimanual laparoscopic tasks. NDH practice demonstrates a beneficial dose-response relationship with performance, supporting its integration into early laparoscopic training.

Laparoscopy

Does Inquiry-Based Learning Improve Students' Critical Thinking? A Meta-Analysis Accounting for Control Group Variations.

BACKGROUND: Inquiry learning is widely recognized, through empirical studies, as an appropriate instruction in enhancing students' critical thinking, yet the results were varied across context. The previous meta-analysis did not include the control group variations as a potential moderator and the studies subject domain was limited only to science subjects. Consequently, it is difficult to generalize the effectiveness of IBL in enhancing critical thinking. This meta-analysis aims to investigate whether inquiry learning is effective in improving the students' critical thinking skills and examine the moderating roles of each study characteristic. Methods The literature search applying the PRISMA protocol 2020 was conducted by utilizing SCOPUS, ERIC, and DOAJ databases. A total of 57 studies from 51 articles, published from 2015 to 2025, were synthesized using a random-effects model with standardized mean difference (SMD). RESULTS: The analysis revealed that IBL has a large and significant effect on enhancing students' critical thinking (g = 1.336; 95% CI [1.061, 1.611]). However, substantial heterogeneity was observed (I 2 = 92.09%), suggesting variability across contexts. Moderator analyses revealed that the main moderator, control group variations, was statistically significant in moderating the effectiveness of IBL (Qm = 5.21; p = .022). in contrast, subject domain (Qm = 1.43; p = .698), education level (Qm = 1.11; p = .774), and country ( Q m  = 3.33; p = .650), were insignificantly moderating the effectiveness of inquiry learning. CONCLUSIONS: The present meta-analysis highlighted that IBL is effective in improving students' critical thinking. However, the effectiveness of IBL was relative to the type of control group variations. Its effect on critical thinking was greater when compared with teacher-centered learning but smaller when compared with other student-centered learning.

Thinking

Revealing potential biomarkers and metabolic mechanisms of ovarian aging in hens during late laying period based on machine learning and metabolomics.

Ovarian function decline during the late laying period represents a major bottleneck for the economic efficiency of the global poultry industry. However, the underlying metabolic mechanisms and reliable early-warning biomarkers for ovarian aging remain poorly understood. In this study, we performed the first untargeted LC-MS/MS metabolomics analysis of ovarian tissues from Taihe silky fowls at peak laying (30 weeks) and late laying (50 weeks) stages, and employed an ensemble machine learning strategy integrating LASSO, random forest, and support vector machine (SVM) algorithms to identify high-confidence core biomarkers of ovarian aging. Gene expression analysis was further conducted to validate the potential molecular mechanisms. Our results showed that the metabolic profiles of ovarian tissues differed significantly between the two groups. A total of 6 core biomarkers were identified, 4 of which were long-chain acylcarnitines. Mechanistic analysis revealed that downregulation of key genes in the carnitine shuttle system led to impaired mitochondrial fatty acid β-oxidation, which in turn triggered excessive oxidative stress and compromised ovarian endocrine function. In conclusion, this study identifies long-chain acylcarnitines as potential metabolic biomarkers for ovarian aging in Taihe silky fowls. These findings provide novel insights into the metabolic basis of poultry ovarian aging and lay a theoretical foundation for the precise regulation of reproductive performance in indigenous poultry breeds.

Animals

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Predicting ACL injury risk in athletes: A systematic review of machine learning-based models.

BACKGROUND: Early ACL injury risk identification in athletes is essential. This systematic review examines machine learning (ML) models for predicting ACL injuries, evaluating their methodological quality, performance, and reliability. METHOD: A comprehensive electronic search was conducted across PubMed, Scopus, Web of Science, and IEEE Xplore databases, supplemented by Google Scholar for grey literature, covering articles published between January 1, 2015, and August 30, 2025. Eligible studies were appraised using the Prediction Model Study Risk of Bias Assessment Tool (PROBAST) for methodological quality and risk of bias, and the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) guidelines for quality of evidence. RESULTS: Ten studies were included. PROBAST showed eight studies had moderate risk of bias and two low risk. TRIPOD found only two studies met quality criteria. ML models included logistic regression (n = 5), support vector machines (n = 4), k-nearest neighbor (n = 3), decision trees (n = 3), random forests (n = 5), neural networks (n = 2), linear discriminant analysis (n = 1), and pre-trained CNNs (n = 1). AUC ranged from 0.63 to 0.98. Accuracy (reported in six studies) ranged from 26% to 95%; however, these values should be interpreted with caution due to the absence of confidence intervals, lack of class imbalance handling, and limited external validation across studies. Tree-based ensemble methods such as random forest achieved competitive accuracy (74-86%), while SVM, a non-ensemble classifier, reported accuracy ranging from 71% to 95%; however, the highest values were obtained in studies with notably small sample sizes (n = 12 to n = 39), raising concerns about overfitting and generalizability. CONCLUSION: Current ML algorithms show promise for identifying athletes at high ACL injury risk and detecting relevant risk factors. Although study quality was generally satisfactory, future research should prioritize external validation and model interpretability to support clinical translation.

Humans

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype‑dependent opioid consumption over 72 h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non‑carriers, despite reporting similar subjective pain scores. This consistent genotype‑dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

Predicting training outcomes for developmental dyslexia from EEG data.

Developmental dyslexia (DD) is characterised by lower-than-average reading abilities and is diagnosed in approximately 10% of individuals. The societal barriers may limit professional fulfilment and psychological wellbeing of individuals with DD, calling for the development of effective interventions to counteract them. As DD is associated with challenges in both phonological and visuo-attentional domains, different longitudinal training approaches were developed to strengthen them. However, they require a considerable amount of personal, social and economic resources and the outcomes may vary depending on individual differences in behavioural and neurophysiological functionality. Hence, predicting training outcomes might help in developing personalised treatment protocols and optimising the use of resources. In the present work we applied machine learning to resting-state EEG to predict longitudinal training outcomes in adults with DD enrolled in a randomized clinical trial. In particular, one group received a visuo-attentional training combined with transcranial alternating current stimulation (tACS), another group received visuo-attentional training with sham/placebo stimulation, and the third group received a phonological training with sham/placebo stimulation. The improvement in text reading speed was associated with spectral power in low-beta and individual frequencies in the alpha (IAF) and beta (IBF) bands, while the improvement in pseudoword reading was associated with IBF. The findings highlight the potential of capturing neural markers of treatment responsiveness in DD. Future studies should focus on the generalisability of predictive models to real-world settings, while investigating whether specific EEG markers predict responsiveness to distinct remediation protocols, thus supporting the development of personalised interventions.

Humans

The role of simulator immersion on learning and transfer of decision-making skill in sport.

Virtual reality has become popular in sport and other domains because it can immerse the user within a sporting context and solve logistical problems for additional off-field training. There is limited evidence, however, of whether immersion is crucial for learning and transfer. This study compared training of decision-making skill between 360-degree video virtual reality (360VR) and two-dimensional video. Twenty-eight Australian Rules Football players were randomly assigned to one of three training groups: 360VR, two-dimensional video, and control. Across four weeks, participants in the training groups were exposed to decision-making scenarios consisting of visual, contextual and auditory cues. Performance was assessed pre- and post-training with virtual reality and field-based decision-making tests. Results indicated that the two-dimensional video training group showed significantly superior decision-making in the field-based transfer test compared to 360VR and control groups post intervention. There was also indication that two-dimensional video training was superior to the control post intervention in the virtual reality test. Findings indicate that immersion created in virtual reality is not an underpinning mechanism for learning and transfer, rather the use of perceptual information is crucial. 360VR may facilitate uptake through engagement, but two-dimensional video is adequate for learning and transfer of decision-making to the field.

Humans

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Short-term psychodynamic psychotherapy for functional neurological disorder: A pilot randomized controlled trial.

BACKGROUND: Evidence-based psychotherapeutic treatments for Functional Neurological Disorder (FND) remain limited. This pilot trial evaluated the preliminary efficacy of Short-term Psychodynamic Psychotherapy (STPP) plus Standard Medical Care (SMC) compared with SMC alone in reducing FND symptom frequency. METHODS: Adults with FND were randomized (1:1) to receive either SMC alone or 12 weekly sessions of STPP plus SMC. The primary outcome was symptom frequency (days with symptoms in the last 4 weeks) assessed at the end of treatment (3 months) and at 6-month follow-up. Secondary outcomes included treatment response (&#x2265;50% reduction in symptom frequency) and scores on the Hamilton Depression Rating Scale (HAM-D), Hamilton Anxiety Rating Scale (HAM-A), and World Health Organization Disability Assessment Schedule 2.0 (WHODAS 2.0). RESULTS: Of 91 randomized patients (mean age 38.2 years, 75.8% female), 81.3% completed follow-up. Intention-to-treat analysis using Linear Mixed Models showed that STPP plus SMC significantly reduced symptom frequency compared with SMC alone (estimated mean difference -5.72 [95% CI -8.68 to -2.77]; Cohen's d = 0.77; p&#x202f;<&#x202f;0.001).Treatment response was achieved by 65.8% in the intervention group versus 16.7% in controls (OR 8.21 [95% CI 2.79-24.19]; p&#x202f;<&#x202f;0.001; NNT 2.0).Significant improvements were also observed for depression (HAM-D: estimated mean difference -10.80; d = 1.45), anxiety (HAM-A: -7.94; d = 1.06), and disability (WHODAS 2.0: -5.77; d = 0.74), all p&#x202f;<&#x202f;0.001. CONCLUSIONS: STPP was associated with clinically meaningful improvements in FND symptom frequency and all secondary outcomes, with large effect sizes and high treatment response rates. These findings support the preliminary efficacy of STPP for FND and justify larger, multicenter confirmatory trials.

Humans

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Clinical Efficacy and Learning Curve of Far-Lateral Approach (FLA) in Uni-Portal Non-Coaxial Spinal Endoscopic Surgery (UNSES) in the Treatment of Lumbar Degenerative Diseases: A Prospective Study.

BACKGROUND: Uniportal non-coaxial spinal endoscopic surgery (UNSES) via far-lateral approach (FLA) is an innovative minimally invasive procedure for lumbar degenerative diseases, particularly far-lateral disc herniation and foraminal stenosis. However, complex lateral lumbar anatomy and strict endoscope-instrument coordination create a distinct learning curve that may compromise early surgical efficiency and safety. This study aimed to evaluate the efficacy and safety, quantify the learning curve, and to provide clinical guidance for the standardized promotion and application of this technology. METHODS: A total of 40 consecutive patients with lumbar degenerative diseases who underwent UNSES via FLA by a single surgeon between January 2025 and December 2025 were included. All data were analyzed using SPSS 26.0 statistical software (IBM, USA). Primary outcomes included operation time, blood loss, fluoroscopy frequency, and intraoperative complication rate. Secondary outcomes were VAS, ODI, and modified Macnab criteria at 1, 3, and 6&#x2009;months postoperatively. The learning curve and the inflection point of the learning curve was determined using cumulative sum (CUSUM) analysis. The differences in clinical indicators between early and proficient stage were compared. RESULT: Operation time, blood loss, and fluoroscopy times decreased significantly with case accumulation (p&#x2009;<&#x2009;0.05). CUSUM identified an inflection point at the 16th case, after which operation time stabilized at (55.3&#x2009;&#xb1;&#x2009;8.6) min, much shorter than the early phase (89.5&#x2009;&#xb1;&#x2009;10.3) min (p&#x2009;<&#x2009;0.001). Before the 16th case, the curve was in an upward trend; after the 16th case, the curve tended to be flat, indicating the proficiency stage. Postoperative VAS and ODI improved significantly than those before surgery at each follow-up time (p&#x2009;<&#x2009;0.05). There was no significant difference in postoperative VAS score and ODI between the two groups at each follow-up time point (p&#x2009;>&#x2009;0.05). The total complication rate was 12.5% (5/40), were cured by conservative treatment. The total excellent-good rate was 90.0% (36/40). L5/S1 and Bertolotti's syndrome were independent factors affecting the learning curve. CONCLUSION: UNSES via FLA is a safe and effective minimally invasive technique for treating complex lumbar degenerative diseases. It has a certain learning curve, and the inflection point is about the 16th case. After mastering the key techniques such as anatomical positioning, endoscopic manipulation and hemostasis, the surgeon can gradually reach the proficiency stage, with significantly improved surgical efficiency and clinical efficacy, and controllable complications. This study provides a theoretical basis for the clinical training and technology promotion of UNSES via FLA.

Humans

Reliability-aware hierarchical learning for Chagas disease screening from 12-lead ECGs: tackling label uncertainty and class imbalance.

Objective.Chagas disease, a neglected tropical disease (NTD) with significant cardiovascular impact, remains underdiagnosed in resource-limited regions. Electrocardiogram (ECG) screening offers a low-cost tool for detecting cardiac involvement, yet algorithm development is challenged by label noise, data scarcity, and the latent nature of infection. This study proposes a robust ECG-based screening framework that explicitly addresses these constraints.Approach.We introduce aReliability-Aware Hierarchical Learningstrategy that calibrates supervision according to data provenance, prioritizing serology-confirmed labels over noisy self-reports. To mitigate data scarcity, we compare a specialized convolutional neural network (CNN) trained from scratch with a transfer learning approach based on a Spatio-Temporal ECG foundation Model (FM). Performance is evaluated across varying data scales, and the representation structure is analyzed to interpret model behavior.Main results.On the official hidden test set of the George B. Moody PhysioNet/Computing in Cardiology Challenge 2025, our approach achieved a Challenge Score of 0.163. We observe that while the specialized CNN performs competitively in data-rich regimes, the FM exhibits superior robustness in extreme low-resource settings. Furthermore, performance reaches a plateau imposed by underlying disease physiology. Bimodal score distributions suggest that models distinguish established cardiomyopathy from indeterminate infection, which remains electrophysiologically indistinguishable from healthy controls.Significance.These findings clarify both the potential and intrinsic limits of ECG-based AI screening for NTD-associated cardiac involvement. Reliability-aware supervision and data-efficient transfer learning provide a practical framework toward scalable and clinically meaningful ECG screening systems in resource-constrained environments.

Humans