Search PubMedSearch

SEARCH · Search PubMed

Results for “PROBAST+AI”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

102 records · Page 2Linked to original sources

Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis.

BACKGROUND: Emergency radiographic interpretation for fractures is prone to missed or misdiagnoses. Artificial intelligence (AI) is expected to become a powerful tool to assist clinicians in fracture detection. PURPOSE: A systematic review and meta-analysis was performed to assess whether AI improves clinicians' ability to detect fractures on radiographs. MATERIALS AND METHODS: A literature search was conducted in PubMed, Web of Science, and Cochrane Library for studies published between January 1, 2010, and October 10, 2025. A meta-analysis of diagnostic accuracy studies was performed using a Summary Receiver Operating Characteristic (SROC) curve. The quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Subgroup analysis and meta-regression were conducted to explore potential sources of heterogeneity. RESULTS: A total of 26 studies were included . The pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94). The corresponding AUC values were 0.90 (95% CI: 0.87-0.92) before and 0.95 (95% CI: 0.93-0.97) after AI assistance. Eight studies were rated as high risk of bias. Subgroup analysis and meta-regression identified potential sources of heterogeneity, including fracture location, AI model type, high risk of bias, and reference standards. CONCLUSION: AI assistance significantly improves clinicians' diagnostic performance in detecting fractures on radiographs for extremity and trunk fractures.

Humans

Challenges and future directions in AI-driven biomaterials for microbiome-associated oral infectious diseases: A systematic review.

Oral biofilm-induced antimicrobial resistance is the core pathogenic mechanism of microbiome-associated oral infectious diseases (dental caries, periodontitis, peri-implantitis, and endodontic infection). Traditional therapies and biomaterials are limited by poor biofilm penetration, drug resistance induction, single functionality, and inadequate adaptation to dynamic oral microenvironmental changes (e.g., pH fluctuations, salivary rinsing, masticatory stimulation). Artificial intelligence (AI) has transformed the field by integrating materials science, microbiology, and stomatology data. Via machine learning, deep learning, and multi-physics simulation, AI optimizes biomaterial physicochemical properties, decodes microenvironmental signals, constructs precise sensing-response loops, and supports the full chain of material design, performance prediction, and action simulation, advancing treatment from empirical intervention to precision regulation. This systematic review retrieved literature from PubMed, Embase, and Web of Science (January 2016-January 2026) using keywords across three dimensions: AI, biomaterials, and oral microbiome. Following inclusion/exclusion criteria, 99 articles were included. It elaborates on five core mechanisms of AI-driven oral biomaterials (precise oral microbiome analysis, targeted material design/optimization, performance prediction/simulation, targeted delivery/intervention, effect evaluation/dynamic regulation), analyzes their applications in microbiome-targeted biomaterial research and development (R&D) and clinical practice for the four major oral infectious diseases, addresses technical bottlenecks (insufficient targeting specificity and precision of biomaterials, poor stability and durability in complex oral microenvironments, inadequate biofilm disruption capacity, and clinical translation obstacles), and proposes future directions (multimodal design to enhance targeting specificity, structural and component optimization to improve stability/durability, development of multi-mechanism synergistic biofilm disruption strategies, strengthening translational research for clinical application, and deep integration of AI in the full chain of biomaterial R&D). This work provides comprehensive theoretical and practical support for the R&D, optimization, and clinical translation of AI-driven microbiome-targeted oral biomaterials.

Humans

AI-driven snapshot hyperspectral imaging for on-line sorting systems in food industry: From real-time sensing to intelligent decision-making.

High-throughput food sorting requires rapid, non-destructive detection of external defects, foreign materials, and internal quality attributes in heterogeneous food matrices. Conventional scanning hyperspectral imaging may suffer from motion-induced spatial-spectral mismatches, whereas snapshot hyperspectral imaging (S-HSI) captures spectral images within a single integration time. However, its advantage is limited by trade-offs in resolution, signal-to-noise ratio (SNR), reconstruction uncertainty, and calibration stability, which are further amplified by variable tissue structure, surface reflection, moisture, and fat distribution in foods. This review critically examines artificial intelligence (AI)-driven S-HSI for on-line food sorting within a sensing-representation-decision-execution framework. Compact architectures are compared according to their physical constraints, food-sorting suitability, and ability to support mapping between spectral responses and physicochemical quality attributes. AI strategies are reviewed for spectral reconstruction, image restoration, spatial-spectral representation, band selection, uncertainty-aware decision-making, and edge implementation. AI can partially compensate for snapshot-specific limitations, but current evidence remains largely limited to laboratory or prototype studies. Future work should link system performance to food safety and quality outcomes by reporting throughput, decision latency, calibration drift, missed-detection risk, false-rejection cost, and closed-loop sorting success.

Hyperspectral Imaging

An AI-assisted Clinical Decision Support System for Green Classification of Cystocele on Dynamic Transperineal Ultrasound.

Green classification of cystocele on dynamic transperineal ultrasound (TPUS) remains operator-dependent because it requires manual frame selection and landmark-based assessment of the Valsalva maneuver. We developed a workflow-oriented AI-assisted clinical decision support system for automated urethrovesical junction localization and dynamic Green classification and prospectively evaluated its standalone and reader-support performance. This diagnostic accuracy and reader study included 881 patients from a tertiary referral hospital, comprising a retrospective development cohort (n = 688) and an independent prospective test cohort (n = 193). A nested subset of 67 prospective patients was used for a reader study involving two junior and two intermediate radiologists under unaided and AI-assisted conditions. In the complete prospective test cohort, Green-AttGRU achieved a macro-averaged AUC of 0.939 (95% CI, 0.897-0.971) and an overall accuracy of 0.902 (95% CI, 0.860-0.943). In the reader study, overall accuracy increased from 0.761 to 0.821 without AI to 0.851-0.881 with AI, while macro-F1 increased from 0.660 to 0.777 to 0.820-0.860. Overall inter-reader agreement increased from a Fleiss' κ of 0.453 to 0.786, and pooled median interpretation time decreased from 26.7 s to 9.9 s. These findings support the preliminary feasibility of the system as a workflow-oriented decision-support tool for dynamic TPUS interpretation.

Humans

Manual, digital, and AI tumour-infiltrating lymphocyte scoring: a secondary analysis of the APHINITY randomised trial.

BACKGROUND: Stromal tumour-infiltrating lymphocytes (sTILs) are prognostic in early-stage HER2-positive breast cancer, but their role in the context of dual HER2 blockade remains undefined. We evaluated manual, digital, and artificial intelligence (AI)-based sTIL quantification, together with AI-derived spatial metrics, for prognostic and treatment-benefit stratification using tumour samples from the phase 3 APHINITY trial. METHODS: In the APHINITY trial, 4805 patients were randomly assigned to receive chemotherapy plus trastuzumab with pertuzumab or chemotherapy plus trastuzumab with placebo. Median follow-up was 74&#xb7;1 months (IQR 68&#xb7;3-75&#xb7;4). We analysed 4262 haematoxylin and eosin-stained images using manual assessment, an automated digital approach, AI-based lymphocyte quantification (AI percentage lymphocytes), and two AI-derived spatial features (AI-TIL and immune hotspot). Interobserver reproducibility was assessed in 262 randomly chosen tumour samples scored independently by five pathologists. Multivariable Cox models were used to assess associations between TIL levels and invasive disease-free survival (primary outcome in APHINITY), distant recurrence-free interval, and overall survival. The heterogeneity of pertuzumab benefit was evaluated using subgroup analyses, subpopulation treatment effect pattern plot analyses, and nested Cox models with treatment-by-biomarker interaction terms. FINDINGS: Manual scoring showed high interobserver reproducibility (intraclass correlation coefficient 0&#xb7;84 [95% CI 0&#xb7;79-0&#xb7;88]). Concordance between manual and automated methods was modest. AI-based scoring (AI percentage lymphocytes) reclassified 120 (11&#xb7;6%) of 1035 node-positive tumours from immune-low (by manual scoring) to immune-high; this subgroup of patients showed greater separation of 5-year invasive disease-free survival curves between pertuzumab and placebo groups compared with patients whose tumours were concordantly classified as immune-low by both manual and AI-based approaches. Higher levels of TILs were associated with improved invasive disease-free survival for all sTIL measurement approaches and spatial measurements (hazard ratios [HRs] 0&#xb7;41-0&#xb7;93). Pertuzumab was associated with improved invasive disease-free survival at higher sTIL levels across all measurement approaches (HRs 0&#xb7;36-0&#xb7;48), but was not associated with higher values of spatial measures. The largest 6-year absolute improvements with pertuzumab were observed in patients with node-positive disease whose tumours scored in the highest level of immune infiltration of manual sTIL scoring (&#x2265;70&#xb7;0%; mean absolute improvement 12&#xb7;1 percentage points [SD 2&#xb7;8]). In nested prognostic and predictive models, AI-based immune hotspot scores provided the most consistent additional information when combined with any sTIL measurement (all p<0&#xb7;010). INTERPRETATION: Standardised manual sTIL scoring was reproducible, and digital and AI-based methods showed consistent prognostic stratification and potential for treatment-benefit stratification despite only modest correlation between platforms. AI spatial metrics provided complementary information beyond sTIL density and could support more scalable immune assessment. Future studies are needed to validate these approaches in independent cohorts and to clarify their clinical utility for stratifying contemporary HER2-directed therapies. FUNDING: None.

Humans

Experimental validation of an AI-driven digital healthcare platform for oral health behavior and plaque assessment among vietnamese children.

BACKGROUND: Oral health among children in developing countries, including Vietnam, remains a significant public health concern. Innovative approaches leveraging artificial intelligence AI-based digital health platforms may offer effective strategies for managing dental plaque and promoting better oral hygiene behaviors among school-aged children. This study aimed to evaluate the effectiveness of an AI-driven oral healthcare platform (Denti-i Vietnam) in improving oral hygiene and behavioral outcomes among Vietnamese primary school students. METHODS: A total of 204 primary school students aged 8-10&#xa0;years in Hanoi, Vietnam, participated in this experimental study. Participants were randomly assigned to an intervention group (n&#xa0;=&#xa0;107), which used the AI-driven oral healthcare platform, and a comparison group (n&#xa0;=&#xa0;97), which received traditional oral health education via pamphlets. Oral health behaviors, dental plaque levels (Simplified Oral Hygiene Index; OHI-S), and caries indices (dft/DMFT) were assessed at baseline and after the intervention period. RESULTS: The intervention group demonstrated a significant reduction in the OHI-S score compared to baseline (2.49&#xa0;&#xb1;&#xa0;0.60 to 1.70&#xa0;&#xb1;&#xa0;0.76, p&#xa0;<&#xa0;0.001), particularly in the debris component, indicating enhanced plaque control. Notable improvements were also observed in oral hygiene behaviors, including increased frequency of toothbrushing before and after breakfast (p&#xa0;<&#xa0;0.01) and more frequent parental assistance during brushing (p&#xa0;=&#xa0;0.03). Furthermore, parental awareness of dental caries significantly increased in the intervention group (p&#xa0;=&#xa0;0.001). CONCLUSIONS: The AI-driven oral healthcare platform significantly improved both oral hygiene behaviors and plaque control among Vietnamese primary school children. These findings suggest that AI-driven digital health tools can serve as practical and scalable solutions for promoting oral health in developing countries.

Humans

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis.

BACKGROUND: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. OBJECTIVE: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. METHODS: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (&#x2265;18 years of age), reporting mean polyp detection counts stratified by size (&#x2264;5 mm, 6-9 mm, and &#x2265;10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and &#x3c4;2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. RESULTS: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (&#x2264;5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI -1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI -0.02 to 0.06, 95% PI -0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI -0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. CONCLUSIONS: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment-prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932.

Colonoscopy

User Engagement and Feature Preferences in an AI-Powered mHealth Intervention for Diabetes Prevention: Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Prediabetes is highly prevalent and increasing globally, yet lifestyle interventions remain underused. AI-driven mobile health (mHealth) tools can help scale diabetes prevention efforts, but the key factors driving their success are not well understood. OBJECTIVE: This post hoc secondary analysis of a randomized controlled trial (RCT) aimed to characterize the most valued features and the role of user engagement in outcomes of a fully automated mHealth intervention for diabetes prevention. METHODS: Data from 151 participants with prediabetes and overweight or obesity who were assigned to an AI-based diabetes prevention program (Sweetch) in a parent RCT (NCT05056376) were analyzed. Engagement (defined as the total number of days the app was used) was categorized into tertiles (low, medium, and high). Baseline characteristics were compared across engagement groups using ANOVA, Kruskal-Wallis, and chi-square tests, and regression models assessed the association between engagement and achievement of diabetes risk reduction outcomes (&#x2265;5% weight loss, &#x2265;4% weight loss with &#x2265;150 min/week of physical activity, or &#x2265;0.2 percentage point reduction in hemoglobin A1c [HbA1c] at 12 months). Perceived usefulness of intervention features was surveyed at 12 months. RESULTS: Median engagement was 98 (IQR 34-232) days. Older age (P<.001) and lower baseline BMI (P=.04) were significantly associated with higher engagement. Compared with low engagement, high engagement was associated with greater odds of achieving the composite diabetes risk reduction outcome (odds ratio [OR] 2.59, 95% CI 1.11-6.01; P=.03), &#x2265;5% weight loss (OR 3.31, 95% CI 1.16-9.42; P=.03), and &#x2265;0.2 percentage point reduction in HbA1c (OR 3.57, 95% CI 1.19-10.75; P=.02). Participants most frequently rated weight tracking, physical activity tracking, and the digital body weight scale as the features that were most helpful for achieving their health goals. CONCLUSIONS: Higher engagement with an AI-driven intervention requiring no human intervention was associated with improved diabetes risk reduction. Contrary to concerns about lower digital literacy, older adults engaged with the intervention more than younger adults. Features related to weight and physical activity tracking were most valued by patients in the program. TRIAL REGISTRATION: ClinicalTrials.gov NCT05056376; https://clinicaltrials.gov/study/NCT05056376.

Humans

Artificial intelligence (AI) uses in stereotactic radiosurgery (SRS): diagnosis with brain metastasis (BM) - A systematic review.

BACKGROUND: Brain metastases (BM) are the most common intracranial tumors in adults, and stereotactic radiosurgery (SRS) has become a mainstay of management. However, several diagnostic challenges persist in the SRS pathway, particularly the differentiation of radiation necrosis (RN) from true tumor progression, which conventional MRI and even advanced imaging techniques often cannot reliably resolve. Recent advances in artificial intelligence (AI) offer the potential to address these diagnostic limitations. This systematic review synthesizes current literature on AI applications for MRI-based diagnostic decision support in BM patients undergoing SRS, with a focus on radiomics and deep learning tools for distinguishing RN from progression, classifying molecular and histologic subtypes, and predicting treatment response. METHODS: A systematic review was performed in accordance with PRISMA guidelines. PubMed, Web of Science, and Scopus were searched using a targeted query combining terms related to AI, brain metastasis, diagnosis or imaging, and SRS. After screening 483 records and applying strict inclusion and exclusion criteria, 18 studies published between 2015 and 2025 were included. Data were extracted on study design, cohort characteristics, imaging modality, AI methodology, validation strategy, and reported diagnostic performance. RESULTS: Among the 18 included studies, AI models demonstrated strong performance across diagnostic tasks in the BM-SRS pathway. The differentiation of RN from true tumor progression was the most extensively studied application, addressed by 14 of 18 studies, with reported AUCs ranging from 0.71 to 0.94. Support vector machines, random-forest ensembles, convolutional neural networks, and transformer-based multimodal architectures were widely used. The literature evolved from single-sequence radiomic classifiers in 2018 to multimodal deep learning frameworks fusing imaging with clinical and genomic data in 2025. Contrast-enhanced T1-weighted MRI was the dominant imaging input, and texture-based radiomic features (GLCM, GLSZM, GLDM, and wavelet-derived features) were the most consistently predictive. The highest-performing models reached AUCs of 0.85-0.91 through multimodal integration of imaging with clinical and genomic features, and consistently outperformed expert neuroradiologist read on matched cases. Remaining studies addressed longitudinal segmentation-based detection of local failure and adverse radiation effects, BRAF mutation status in melanoma BM, early Gamma Knife treatment response, and primary tumor histology classification, with more variable performance. CONCLUSION: AI models, particularly those integrating MRI-derived radiomic features with clinical and genomic data, show high accuracy in supporting diagnostic decisions for BM patients treated with SRS. The post-SRS differentiation of radiation necrosis from true tumor progression has reached the greatest level of maturity and is closest to clinical translation, with potential to reduce unnecessary biopsies, personalize surveillance intervals, and rationalize treatment-pathway decisions. Other diagnostic applications, including molecular subtyping and primary tumor histology classification, remain exploratory and require further multicenter validation. Integration of AI tools into multidisciplinary tumor-board workflows, combined with prospective validation and standardized reporting, will be essential to realize the full clinical benefits of AI in SRS for brain metastases.

Humans

From prediction to mechanism: Explainable AI uncovers plasma and CSF proteomic signatures of Alzheimer's disease.

Alzheimer's disease (AD) plasma and cerebrospinal fluid (CSF) proteomics can distinguish AD from cognitively normal controls, but the generalizability of machine learning performance and the recurrence of biological signals across datasets require cautious interpretation. We developed an explainable artificial intelligence framework spanning two fluids and four ADNI proteomic datasets, covering 2082 modality specific samples, all analysed internally within ADNI. Phase 1 analysed plasma using a 119 analyte NULISA and targeted UPENN panel (n&#xa0;=&#xa0;727; 216&#xa0;CE, 511 controls). Phase 2 extended the analysis to CSF using SOMAscan7k, TMT-MS and targeted SET2, with Elecsys A&#x3b2;42, A&#x3b2;40, total tau and p-tau181 as anchor biomarkers. Only SOMAscan was subject-independent relative to Phase 1 plasma; TMT-MS and SET2 overlapped with Phase 1 for 96.0% and 97.7% of subjects and therefore are not independent replication cohorts. Under subject-level splits with fold internal preprocessing, we compared Elastic Net, Explainable Boosting Machines and gradient boosted trees with SHAP-based explanations. Among the candidate pipelines, we selected the pipeline with the highest held-out test ROC AUC for each platform; the selected values were 0.927 in plasma and 0.954-0.973 across the three CSF datasets. Because the same held out test performance was used for pipeline selection and headline reporting, these are optimistically selected single-holdout estimates, not unbiased estimates of generalizable or clinical performance. Explanations identified five recurring biological axes within ADNI: cholinergic (ACHE), tau/14-3-3 (YWHAG, YWHAZ, YWHAB, YWHAE), neuro-axonal (NEFL, NEFH), microglial/complement (CHIT1, SMOC1, CHI3L1, C7, CFH) and synaptic (NPTXR, NPTX2, DLG4, SYT5, VSNL1, ELAVL2). CSF analyses showed synaptic vesicle-cycle enrichment (q&#xa0;=&#xa0;2&#xa0;&#xd7;&#xa0;10-6), and CSF YWHAG correlated strongly with total tau (&#x3c1;&#xa0;=&#xa0;0.87). Cross-fluid directional concordance was modest overall (54-57%) but increased to 73-80% among mapped analyte/protein rows reaching q&#xa0;<&#xa0;0.05 in CSF. These findings provide hypothesis-generating, internally supported evidence within ADNI. Independent external cohorts with locked pipelines are required to evaluate generalizable performance and biological reproducibility; the overlapping TMT-MS and SET2 analyses should not be interpreted as independent replication.

Alzheimer Disease

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype&#x2011;dependent opioid consumption over 72&#xa0;h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non&#x2011;carriers, despite reporting similar subjective pain scores. This consistent genotype&#x2011;dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

Effectiveness of an AI-based home exercise app for rehabilitation of rotator cuff-related shoulder pain: A randomized controlled trial.

BACKGROUND: Rotator cuff-related shoulder pain contributes to disability and healthcare use. Although therapeutic exercise is first-line treatment, limited supervision and adherence may reduce its effectiveness; digital rehabilitation with real-time feedback may address these limitations. OBJECTIVES: To evaluate the effectiveness of adding a digital rehabilitation program to standard physiotherapy on pain, function, fear-avoidance beliefs, and healthcare utilization. DESIGN: Single-center, assessor-blinded, randomized controlled trial with two parallel groups. METHOD: Forty-six adults (mean age 59 years) with rotator cuff-related shoulder pain were randomized to 12 weeks of conventional physiotherapy or physiotherapy plus an AI-based digital rehabilitation program using computer vision for real-time feedback and performance monitoring. Outcomes were assessed at baseline and at 2, 4, and 12 weeks. Pain intensity (NPRS) was primary outcome; secondary outcomes included upper limb function (QuickDASH), fear-avoidance beliefs (FABQ), and post-intervention healthcare utilization. Analyses followed an intention-to-treat approach. RESULTS: Pain reduction exceeded the MCID (1.3) at 4 and 12 weeks. Between-group differences favoured the intervention at Weeks 2 and 4 (MD -0.7; 95% CI -1.13 to -0.14 and MD -1.01; 95% CI -1.8 to -0.2, respectively). Upper limb function improved more at Week 4 (MD -7.3; 95% CI -12.3 to -2.2). FABQ scores decreased more at Week 12 (MD -7.6; 95% CI -14 to -0.5). Fewer participants in the experimental group required post-intervention healthcare (3 vs 10; p&#x202f;=&#x202f;0.02). CONCLUSION: Adding AI-based home exercise app to conventional treatment improve pain and may improve function and reduce healthcare utilization in rotator cuff-related shoulder pain.

Humans

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Sports Medicine

Impact of Commercial Artificial Intelligence on Radiologist Reading Time for Pulmonary Nodule Evaluation at Chest CT.

Background Chest CT is a primary method for identifying pulmonary nodules, yet interpreting scans remains time-intensive and demanding. Currently, artificial intelligence (AI) is expected to reduce reading times, but the effect of AI on reporting times in this setting is unknown. Purpose To evaluate the impact of a commercial AI software on radiologists' reading time for pulmonary nodule assessment on chest CT scans within a real-world clinical setting. Materials and Methods This retrospective study included patients who underwent chest CT examinations at a tertiary medical center between September 2021 and May 2024. The study period was divided into pre- and post-AI phases. The primary outcome was radiology reporting time. The association between AI implementation and reporting time was evaluated using a multivariable parametric Weibull shared frailty survival model adjusted for reader function, examination type, patient location, and requesting specialty, with clustering at the radiologist level. Interaction analyses assessed heterogeneity across prespecified subgroups. An exploratory extrapolation estimated projected workforce and financial impact. Results This study included 19&#x2009;433 patients (mean age, 62 years &#xb1; 14.2 [SD]; 21&#x2009;814 men; 39&#x2009;323 chest CT examinations, 19&#x2009;190 pre-AI, and 20&#x2009;133 post-AI). AI implementation was associated with faster report completion (adjusted hazard ratio, 1.17; 95% CI: 1.14, 1.21; P < .001). The adjusted median reporting time decreased from 21.3 minutes pre-AI to 18.2 minutes post-AI (14.6% reduction; P < .001). Heterogeneity was observed across reader function (P < .001), examination type (P = .048), and requesting specialty (P = .03). The largest relative reductions were observed for CT thorax electrocardiogram-gated examinations (-41.1%; P < .001) and thoracic radiologists (-25.0%; P < .001), whereas emergency department examinations showed increased median reporting time (7.1%; P < .001). At institutional scan volumes (approximately 20&#x2009;000-22&#x2009;000 chest CT examinations annually), exploratory modeling suggested an approximate reduction of 0.5 full-time equivalent radiologist workload. Conclusion Implementation of commercial AI-assisted pulmonary nodule assessment on chest CT scans reduced radiologist reporting time in a real-world clinical setting. &#xa9; The Author(s) 2026. Published by the Radiological Society of North America under a CC BY 4.0 license. Supplemental material is available for this article. See also the editorial by Iwasawa in this issue.

Humans

Intervention Without Borders - an Automated Self-Guided AI-Enhanced Psychoeducation Intervention for Dementia Caregivers: Parallel-Group Randomized Waitlist-Controlled Trial.

OBJECTIVE: To examine whether a fully automated, self-guided intervention (PDC30) could improve caregiver well-being over a 1-month waitlist control in an international sample. DESIGN: Randomized waitlist-controlled trial. SETTING: Web-based platform accessible globally. PARTICIPANTS: 441 individuals responded to study promotion on the internet, of whom 274 from 43 countries met the study criteria and were randomized. Eligible participants were adults providing &#x2265;10 care hours weekly to community-dwelling relatives with dementia, scoring &#x2265;5 on Patient Health Questionnaire-9 (PHQ-9), and without recent caregiver intervention. INTERVENTION: Available 24/7, PDC30 is a self-guided, automated intervention consisting of a Guidebook, an AI-powered counseling chatbot, and interactive applications for cognitive-behavioral techniques, relaxation, and caregiver-recipient bonding. MEASUREMENTS: At baseline and follow-ups at 1, 2, and 3 months, depression was assessed by PHQ-9. Secondary outcomes were measured with validated brief versions of anxiety, burden, and positive gains. RESULTS: Intent-to-treat analysis using mixed-effects regression showed treatment x time2 effects on all outcomes except anxiety. At 1-month follow-up, coinciding with exclusive access to PDC30, intervention caregivers showed significant improvements in depression (d = -0.37), burden (d = -0.34), and positive gains (d = 0.42). The differences mostly disappeared after control participants received the intervention, while improvements in both groups were sustained thereafter. Participants reported using the website several times weekly, were generally satisfied with it, and found the chatbot most helpful. CONCLUSIONS: The effects on depression and other outcomes were consistent with those observed for in-person programs, suggesting the viability of well-designed automated intervention. The study demonstrates the feasibility, acceptability, and potential global health impact of PDC30.

Humans

Decoding the spatiotemporal patterns of food spoilage microbial communities: Integrating multi-omics and artificial intelligence to enable precision preservation.

In the global food supply chain, food wastage caused by spoilage has resulted in significant economic losses, food shortages, and environmental pressure. This process is fundamentally driven by the spatiotemporal dynamics of microbial communities. However, traditional research methods struggle to elucidate the complex mechanisms of spatial heterogeneity, interspecies interactions, and functional succession. This limits the development of effective preservation strategies. This review systematically reviews the cutting-edge progress of integrating multi-omics technologies and artificial intelligence (AI) to study food spoilage microbial communities, breaking through this bottleneck. We propose an intelligent theoretical framework that could potentially analyze microbial metabolic activities and predict dynamic shelf life if implemented. The conceptual framework integrates multidimensional data, including spatial metabolomics, temporal metatranscriptomics, single-cell transcriptomics, and longitudinal metagenomics. It can also be combined with AI models, such as graph neural networks. The article elaborates on the principles and applications of spatio-temporal monitoring technologies, such as nano secondary ion mass spectrometry, hyperspectral imaging, and the Internet of Things sensing. Through illustrative cases of typical perishable foods, it also explores how such a multi-omics - AI system might be applied to spoilage warning and precise intervention. Additionally, the article addresses the current challenges in data coverage, model generalization, and federated learning implementation. Then the research further explores emerging areas such as engineered probiotics, edge AI, and microfluidic sensing. These areas are targeted at transforming food preservation from an empirical control approach to a data-driven, precise regulatory framework. This transformation provides theoretical support and technical approaches for developing a smart, sustainable food preservation system.

Multiomics

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans

How Following Medical Artificial Intelligence Advice Can Mitigate Malpractice Liability: Cross-National Insights from a Randomized Trial.

Artificial intelligence (AI) increasingly influences clinical decision-making, yet its recommendations may diverge from standard care. Although malpractice concerns are thought to discourage physicians from following AI advice, experimental evidence from the United States suggests the opposite: lay jurors are more likely to hold physicians liable when they reject AI recommendations. Whether this pattern extends to systems in which court-appointed experts, not lay jurors, determine liability remains unknown. Methods: To examine how physicians and laypeople in expert-based and lay-juror legal systems evaluate physicians' acceptance or rejection of AI recommendations, particularly when those recommendations deviate from standard care, we designed a randomized vignette study: a 2 &#xd7; 2 factorial design varying the AI recommendation (standard vs. nonstandard care) and a fictional physician's decision (accept vs. reject). The study was conducted online in 2023 among nationally representative samples of U.S. and German adults and from 2023 to 2024 among German physicians. In total, 387 German physicians, 2291 U.S. adults, and 2283 German adults participated; those not completing the survey or failing attention checks were excluded per preregistered criteria. Participants were randomly assigned to 1 of 4 vignettes, varying the AI recommendation (standard vs. nonstandard care) and physician's decision (accept vs. reject). The reasonableness of the fictional physician's decision was measured, rated by participants on a Likert scale. Results: Analysis, following preregistered exclusion criteria, included 248 German physicians, 1202 U.S. adults, and 1358 German adults. Physicians accepting standard-care AI recommendations were rated more reasonable than those rejecting them (U.S. laypeople: t = 5.36; 95% CI, 0.45-0.97; P < 0.001; German physicians: t = 2.47; 95% CI, 0.14-1.30; P = 0.02; German laypeople: t = 4.14; 95% CI, 0.27-0.76; P < 0.001). Ratings of physicians accepting versus rejecting AI nonstandard-care recommendations were statistically equivalent. Equivalence was tested at an &#x3b1;-value of 0.05 using a two 1-sided tests procedure, reported with 90% CIs per standard convention (U.S. laypeople: t = -4.90; 90% CI, -0.1 to 0.36; P < 0.001; German physicians: t = -1.76; 90% CI, -0.12 to 0.67; P = 0.04; German laypeople: t = 5.35; 90% CI, -0.35 to 0.06; P < 0.001). Conclusion: Across the United States and Germany, samples representative of lay jurors and court-appointed experts viewed accepting standard-care AI advice as more reasonable, whereas accepting or rejecting nonstandard-care AI advice was judged similarly. Contrary to predictions, malpractice liability regimes do not necessarily pose a barrier to AI use in precision medicine.

Artificial Intelligence