Search PubMedSearch

SEARCH · Search PubMed

Results for “AI-supported systematic review”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

2,007 records · Page 2Linked to original sources

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis.

BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG-derived inputs. METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Assessing AI literacy and attitudes among medical students: implications for integration into healthcare practice.

PURPOSE: This study aims to assess AI literacy and attitudes among medical students and explore their implications for integrating AI into healthcare practice. DESIGN/METHODOLOGY/APPROACH: A quantitative research design was employed to comprehensively evaluate AI literacy and attitudes among 374 Lusaka Apex Medical University medical students. Data were collected from April 3, 2024, to April 30, 2024, using a closed-ended questionnaire. The questionnaire covered various aspects of AI literacy, perceived benefits of AI in healthcare, strategies for staying informed about AI, relevant AI applications for future practice, concerns related to AI algorithm training and AI-based chatbots in healthcare. FINDINGS: The study revealed varying levels of AI literacy among medical students with a basic understanding of AI principles. Perceptions regarding AI's role in healthcare varied, with recognition of key benefits such as improved diagnosis accuracy and enhanced treatment planning. Students relied predominantly on online resources to stay informed about AI. Concerns included bias reinforcement, data privacy and over-reliance on technology. ORIGINALITY/VALUE: This study contributes original insights into medical students' AI literacy and attitudes, highlighting the need for targeted educational interventions and ethical considerations in AI integration within medical education and practice.

Students, Medical

Artificial intelligence enabled social robotic interventions (PARO) in Australian dementia care: A systematic review and meta-analysis.

BACKGROUND: Although there is a growing body of research indicating that Personal Robot/Social Robot could be used in various aspects of care for individuals with dementia, little is known about how well these types of interventions work in an actual hospital setting in Australia. AIMS & OBJECTIVES: The objective of the present systematic review and meta-analysis is to assess the effectiveness of PARO-based socially assistive robotic intervention in terms of its effectiveness outcomes towards the reduction of dementia-related behavioural and psychological symptoms in Australian based healthcare settings. METHODS: A systematic search was conducted across five electronic databases, including MEDLINE (PubMed), EMBASE, CINAHL, PsycINFO, and the Cochrane Library, to identify randomised controlled trials (RCTs) investigating PARO-based socially assistive robotic interventions for dementia in Australian healthcare settings. This review was registered with PROSPERO (CRD420251251916) and followed the PRISMA 2020 guidelines. In addition, the Cochrane Risk of Bias tool (RoB 2) was used to evaluate the risk of bias across all studies. Pooled standardised mean differences (SMD) with 95 % confidence intervals (CI) were calculated for agitation, anxiety, and depression. Heterogeneity across studies was evaluated using the I2 statistic. RESULTS: Six RCTs involving 1444 participants were identified for inclusion in this review. AI-enabled socially assistive robotic interventions, specifically the PARO therapeutic robot, significantly reduced agitation and anxiety when compared to standard treatment or control conditions. The pooled analysis showed that agitation [SMD = -0.44 (95 % CI: -0.70, -0.18) p = 0.0008] and anxiety [SMD = -0.59 (95 % CI: -0.91, -0.27) p = 0.0003] were reduced significantly, while the decrease in depression [SMD = -0.44 (95 % CI: -0.95, -0.07) p = 0.09] scores was non-significant among dementia patients receiving PARO-based socially assistive robotic interventions as compared to the control. The overall risk of bias across all six studies was considered low to moderate. CONCLUSION: PARO-based socially assistive robotic interventions may provide preliminary evidence of effectiveness in reducing agitation and anxiety in individuals with dementia in Australian healthcare, but the evidence regarding the reduction of depression remains unclear. Therefore, additional high-quality trials with consistent methodology and extended follow-up will be necessary to determine both the short-term and long-term clinical efficacy and practicality of implementing these interventions into practice.

Humans

Telmisartan-based monotherapy and combination regimens for blood pressure control in adults with hypertension: a systematic review, meta-analysis, and GRADE assessment.

PURPOSE: To evaluate the efficacy, safety, and certainty of evidence for telmisartan-based antihypertensive regimens in adults with hypertension. METHODS: This systematic review and meta-analysis followed PRISMA 2020. PubMed/MEDLINE, Scopus, Web of Science, and Cochrane CENTRAL were searched from inception to 2026. Eligible studies enrolled adults with hypertension and compared telmisartan monotherapy or telmisartan-containing combinations with placebo, usual care, non-telmisartan antihypertensive agents, or alternative telmisartan-based regimens. Continuous outcomes were pooled as mean differences (MDs) and dichotomous outcomes as risk ratios (RRs), both with 95% confidence intervals (CIs), using random-effects models, with additional subgroup analyses conducted by comparator type. Risk of bias was assessed using RoB 2, and certainty of evidence was evaluated using GRADE. RESULTS: Twenty-five included reports (24 unique trials, since two reports present secondary outcomes from the same underlying trial) involving 6,521 participants were included, spanning placebo-controlled, usual-care-controlled, active-comparator, and telmisartan-combination-versus-telmisartan-monotherapy designs. Telmisartan-based therapy significantly reduced office systolic blood pressure (MD - 6.39 mm Hg; 95% CI - 7.86 to - 4.93; low certainty) and office diastolic blood pressure (MD - 4.88 mm Hg; 95% CI - 6.67 to - 3.09; low certainty), although the magnitude of effect was comparator-dependent. Based on only two trials, 24-h ambulatory systolic blood pressure (MD - 7.16 mm Hg; 95% CI - 10.61 to - 3.72) and ambulatory diastolic blood pressure (MD - 4.42 mm Hg; 95% CI - 6.36 to - 2.48) were reduced with moderate certainty. Telmisartan-based regimens improved blood pressure response (RR 1.68; 95% CI 1.31 to 2.16; moderate certainty) but not blood pressure control achievement (RR 1.44; 95% CI 0.92 to 2.24; very low certainty). Overall adverse events, dizziness, and headache were comparable (very low to low certainty), while edema was less frequent with telmisartan-based therapy (RR 0.33; 95% CI 0.15 to 0.73; moderate certainty). CONCLUSION: Telmisartan-based regimens, particularly fixed-dose and multidrug combinations, effectively reduce office and ambulatory blood pressure and improve blood pressure response, with broadly comparable short-term safety and less edema. These effect sizes are comparator-dependent, and certainty of evidence for absolute blood pressure control achievement and for major adverse events is very low; heterogeneity, limited long-term data, and a predominance of Asian-population trials warrant cautious interpretation pending larger, higher-quality, and more geographically diverse confirmatory studies.

Humans

Applications of artificial intelligence in robot-assisted surgery: a systematic review.

To characterize applications of artificial intelligence (AI) in robot-assisted surgery, summarize technical and clinical performance, and assess the quality of the available evidence. PubMed, Web of Science Core Collection, and Scopus were searched for English-language journal articles published from 1 January 2020 through 31 October 2025. Randomized, observational, model-development, validation, and feasibility studies evaluating AI in robot-assisted surgery or closely related image-guided minimally invasive workflows were eligible. Two reviewers independently performed study selection, data extraction, and risk-of-bias assessment. Owing to heterogeneity in surgical procedures, AI tasks, analytical units, validation strategies, and outcomes, findings were synthesized descriptively without statistical pooling. The review was registered in the International Prospective Register of Systematic Reviews (CRD420251175699). Seventeen studies were included: seven clinical prediction or decision-support studies, eight intraoperative recognition, segmentation, or image-guided studies, and two training or workflow studies. Five prediction studies reported area-under-the-curve values of 0.74-0.95. Technical studies reported F1 or Dice scores of 0.525-0.995 and task-specific accuracies of 0.840-0.998. Two randomized studies suggested benefits for personalized suturing feedback and automated camera control, but neither established improved patient outcomes. Only one study had low overall risk of bias; the remaining studies were at high or unclear risk or raised some concerns. AI applications in robot-assisted surgery show promise for prediction, intraoperative perception, training, and workflow support. Evidence primarily demonstrates technical feasibility rather than established clinical effectiveness. Independent multicenter validation and prospective evaluation of patient, educational, and workflow outcomes are required before widespread implementation.

Robotic Surgical Procedures

Experimental validation of an AI-driven digital healthcare platform for oral health behavior and plaque assessment among vietnamese children.

BACKGROUND: Oral health among children in developing countries, including Vietnam, remains a significant public health concern. Innovative approaches leveraging artificial intelligence AI-based digital health platforms may offer effective strategies for managing dental plaque and promoting better oral hygiene behaviors among school-aged children. This study aimed to evaluate the effectiveness of an AI-driven oral healthcare platform (Denti-i Vietnam) in improving oral hygiene and behavioral outcomes among Vietnamese primary school students. METHODS: A total of 204 primary school students aged 8-10&#xa0;years in Hanoi, Vietnam, participated in this experimental study. Participants were randomly assigned to an intervention group (n&#xa0;=&#xa0;107), which used the AI-driven oral healthcare platform, and a comparison group (n&#xa0;=&#xa0;97), which received traditional oral health education via pamphlets. Oral health behaviors, dental plaque levels (Simplified Oral Hygiene Index; OHI-S), and caries indices (dft/DMFT) were assessed at baseline and after the intervention period. RESULTS: The intervention group demonstrated a significant reduction in the OHI-S score compared to baseline (2.49&#xa0;&#xb1;&#xa0;0.60 to 1.70&#xa0;&#xb1;&#xa0;0.76, p&#xa0;<&#xa0;0.001), particularly in the debris component, indicating enhanced plaque control. Notable improvements were also observed in oral hygiene behaviors, including increased frequency of toothbrushing before and after breakfast (p&#xa0;<&#xa0;0.01) and more frequent parental assistance during brushing (p&#xa0;=&#xa0;0.03). Furthermore, parental awareness of dental caries significantly increased in the intervention group (p&#xa0;=&#xa0;0.001). CONCLUSIONS: The AI-driven oral healthcare platform significantly improved both oral hygiene behaviors and plaque control among Vietnamese primary school children. These findings suggest that AI-driven digital health tools can serve as practical and scalable solutions for promoting oral health in developing countries.

Humans

Artificial intelligence for anticancer drug discovery from natural products of macroalgae and sponges: A systematic review.

Marine natural products (MNPs) from macroalgae and marine sponges have inspired clinically important anticancer agents, including the cytarabine pharmacophore and the eribulin scaffold, while cyanobacterial dolastatin chemistry supplies the auristatin payloads of several marine-inspired antibody-drug conjugates (ADCs) such as brentuximab vedotin. Artificial intelligence (AI) methods, encompassing both classical machine learning (ML) with hand-engineered features and modern deep learning (DL) with many-layered neural networks, are increasingly supporting key decisions in natural-product anticancer drug discovery, including bioactivity prediction, target identification, absorption, distribution, metabolism, excretion and toxicity (ADMET) filtering, generative analogue design, and the selection of preclinical candidates. DL architectures relevant to this field include graph neural networks, transformer-based molecular generators, diffusion models for protein-ligand docking, and convolutional networks for mass spectrometry, while classical ML contributes interpretable fingerprint-based bioactivity models and molecular networking for dereplication. This review follows a systematic literature review methodology to organize the landscape of AI methods now applied to MNP anticancer discovery, distinguishing ML and DL approaches where relevant, situating them within the chemical context of macroalgal and sponge-derived oncology leads, and critically examining published case studies, including validation level (computational, in vitro, in vivo, clinical). The principal bottleneck for medical translation has shifted partly from algorithmic capability toward data infrastructure and experimental validation. Sparse, heterogeneous, and taxonomically biased bioactivity records limit what current models can learn and reduce the reliability of AI-prioritized candidates entering the preclinical pipeline. A roadmap is proposed that prioritizes open MNP-specific benchmarks, symbiont-aware modeling, and active learning loops with synthesizability and ADMET constraints. These AI workflows may accelerate the prioritization of marine-derived anticancer leads and support earlier, more evidence-based translational decisions in oncology drug development.

Biological Products

Comparative Efficacy of Different AI Systems for Polyp Detection by Size During Colonoscopy: Systematic Review and Network Meta-Analysis.

BACKGROUND: Colorectal cancer remains a leading cause of death despite being largely preventable through polypectomy. AI systems designed to enhance polyp detection during colonoscopy have shown promise, but the extent to which they improve detection of different-sized polyps remains unclear. OBJECTIVE: This study compared the size-stratified efficacy of AI-assisted colonoscopy vs standard colonoscopy using the Hartung-Knapp-Sidik-Jonkman (HKSJ) method, and generated exploratory rankings while acknowledging all cross-platform comparisons are indirect. METHODS: This systematic review and network meta-analysis (NMA) searched PubMed, Embase, Cochrane CENTRAL, and Web of Science from inception to July 25, 2026, supplemented by citation searching. We included randomized controlled trials (RCTs) comparing AI-assisted vs standard colonoscopy in adults (&#x2265;18 years of age), reporting mean polyp detection counts stratified by size (&#x2264;5 mm, 6-9 mm, and &#x2265;10 mm). Two reviewers screened studies, extracted data, and assessed risk of bias using the Cochrane Risk of Bias 2.0. We conducted frequentist NMA using the HKSJ method with restricted maximum likelihood estimation, calculated 95% prediction intervals (PIs), and assessed heterogeneity using I2 and &#x3c4;2. Certainty of evidence was rated using the GRADE (Grading of Recommendations Assessment, Development, and Evaluation) framework. RESULTS: A total of 13 RCTs (4156 participants) compared 8 AI systems to standard colonoscopy, forming a network without direct AI comparisons. For diminutive polyps (&#x2264;5 mm), AI showed a modest advantage (standardized mean difference [SMD] 0.21, 95% CI 0.07 to 0.35, 95% PI -1.12 to 1.54), but substantial heterogeneity (I2=86.6%) and wide PI crossing the null indicated high uncertainty. EndoScreener showed the most consistent evidence (SMD 0.36, 95% CI 0.18-0.54). For small and large polyps, effects were minimal (SMD 0.02, 95% CI -0.02 to 0.06, 95% PI -0.03 to 0.07; SMD 0.01, 95% CI 0.00-0.02, 95% PI -0.01 to 0.03). GRADE certainty was very low for diminutive polyps and low for small and large polyps. Sensitivity analysis excluding Tianjin YuJin did not materially change findings. CONCLUSIONS: AI may modestly enhance diminutive polyp detection, but effects on small and large polyps are minimal, with no platform superiority. Given very low to low certainty, findings are hypothesis-generating. This exploratory NMA provides size-stratified comparisons that can inform future head-to-head trial design. Unlike prior reviews aggregating all polyp sizes, we show the overall AI benefit is driven by diminutive polyp detection, providing a framework for targeted deployment-prioritizing AI for diminutive polyp screening, with limited value for larger lesions. Head-to-head trials are urgently needed. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD420251266932; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251266932.

Colonoscopy

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75&#x2009;161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et&#xa0;al., Nanda et&#xa0;al., Naylor et&#xa0;al., and Van Leeuwen et&#xa0;al., each showing fair discrimination. The Teede et&#xa0;al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et&#xa0;al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et&#xa0;al. and van Leeuwen et&#xa0;al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans

Machine learning-ready genomic biomarkers: ATF3 polymorphisms predict postoperative analgesic demand through AI-compatible phenotyping.

PURPOSE: To determine whether ATF3 polymorphisms can serve as genetic biomarkers for machine learning-based precision analgesia by establishing a genotype-phenotype association suitable for predictive modeling of postoperative opioid requirements. METHODS: In a prospective cohort of 167 adults undergoing abdominal surgery, ATF3 SNPs rs3122721 and rs3125293 were genotyped. A structured dataset architecture was developed to represent genetic profiles as input features for supervised learning models, enabling translational analysis of genotype&#x2011;dependent opioid consumption over 72&#xa0;h. RESULTS: Patients with homozygous genotypes of the ATF3 SNPs had significantly higher opioid requirements than non&#x2011;carriers, despite reporting similar subjective pain scores. This consistent genotype&#x2011;dependent pattern provided a clinically relevant phenotype suitable for integration into predictive algorithms. CONCLUSION: ATF3 genotyping offers a promising biomarker for computationally informed precision analgesia. By linking genomic variability to clinically meaningful outcomes within a structured clinical and genomic framework, this approach supports the future development of risk-stratified clinical decision-support systems to optimize postoperative pain management.Trial registration ChiCTR1900021991, registered 30 April 2019. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13755-026-00480-9.

ATF3

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans

The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis.

BACKGROUND: Current endoscopic retrograde cholangiopancreatography (ERCP) and cholangioscopic-based diagnostic sampling for indeterminant biliary strictures remain suboptimal. Artificial intelligence (AI)-based algorithms by means of computer vision in machine learning have been applied to cholangioscopy in an effort to improve diagnostic yield. The aim of this study was to perform a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based diagnostic performance of AI-associated cholangioscopic diagnosis of indeterminant or malignant biliary strictures. METHODS: Individualized searches were developed in accordance with PRISMA and MOOSE guidelines, and meta-analysis according to Cochrane Diagnostic Test Accuracy working group methodology. A bivariate model was used to compute pooled sensitivity and specificity, likelihood ratio, diagnostic odds ratio, and summary receiver operating characteristics curve (SROC). RESULTS: Five studies (n=675 lesions; a total of 2,685,674 cholangioscopic images) were included. All but one study analyzed a deep learning AI-based system using a convoluted neural network (CNN) with an average image processing speed of 30 to 60 frames per second. The pooled sensitivity and specificity were 95% (95% CI: 85-98) and 88% (95% CI: 76-94), with a diagnostic accuracy (SROC) of 97% (95% CI: 95-98). Sensitivity analysis of CNN studies (4 studies, 538 patients) demonstrated a pooled sensitivity, specificity, and accuracy (SROC) of 95% (95% CI: 82-99), 88% (95% CI: 72-95), and 97% (95% CI: 95-98), respectively. CONCLUSIONS: Artificial intelligence-based machine learning of cholangioscopy images appears to be a promising modality for the diagnosis of indeterminant and malignant biliary strictures.

Humans

Decoding the spatiotemporal patterns of food spoilage microbial communities: Integrating multi-omics and artificial intelligence to enable precision preservation.

In the global food supply chain, food wastage caused by spoilage has resulted in significant economic losses, food shortages, and environmental pressure. This process is fundamentally driven by the spatiotemporal dynamics of microbial communities. However, traditional research methods struggle to elucidate the complex mechanisms of spatial heterogeneity, interspecies interactions, and functional succession. This limits the development of effective preservation strategies. This review systematically reviews the cutting-edge progress of integrating multi-omics technologies and artificial intelligence (AI) to study food spoilage microbial communities, breaking through this bottleneck. We propose an intelligent theoretical framework that could potentially analyze microbial metabolic activities and predict dynamic shelf life if implemented. The conceptual framework integrates multidimensional data, including spatial metabolomics, temporal metatranscriptomics, single-cell transcriptomics, and longitudinal metagenomics. It can also be combined with AI models, such as graph neural networks. The article elaborates on the principles and applications of spatio-temporal monitoring technologies, such as nano secondary ion mass spectrometry, hyperspectral imaging, and the Internet of Things sensing. Through illustrative cases of typical perishable foods, it also explores how such a multi-omics - AI system might be applied to spoilage warning and precise intervention. Additionally, the article addresses the current challenges in data coverage, model generalization, and federated learning implementation. Then the research further explores emerging areas such as engineered probiotics, edge AI, and microfluidic sensing. These areas are targeted at transforming food preservation from an empirical control approach to a data-driven, precise regulatory framework. This transformation provides theoretical support and technical approaches for developing a smart, sustainable food preservation system.

Multiomics

Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the American Journal of Sports Medicine.

BACKGROUND: The peer review system faces increasing strain from rising manuscript volumes, reviewer fatigue, and well-documented interreviewer disagreement. Large language models (LLMs) have shown potential to support the peer review process, but their ability to replicate editorial decisions at high-impact medical journals and their utility as manuscript screening tools remain unknown. PURPOSE: To compare the agreement between an LLM and the final editorial decision on manuscripts submitted to the American Journal of Sports Medicine and to evaluate the potential of LLMs as a manuscript screening tool. STUDY DESIGN: Cross-sectional agreement study. METHODS: Fifty-four manuscripts randomly selected from submissions to the American Journal of Sports Medicine (September 2024-October 2024) were reviewed by a locally deployed LLM (Ministral 3 14B; Mistral AI) using a standardized prompt. The artificial intelligence (AI) produced a categorical recommendation (reject, cascade, revision, or accept) and a numerical score (0-100) for each manuscript. Agreement with the final editorial decision was assessed by Cohen kappa (4-category model) for pooled human reviewers (n = 139 reviews) and the AI (n = 54). Screening performance was evaluated by positive predictive value (PPV), sensitivity, and specificity. RESULTS: Pooled human reviewers demonstrated fair agreement with the final decision (&#x3ba; = 0.181 [P < .001]; 42.4% agreement), while the AI demonstrated slight, nonsignificant agreement (&#x3ba; = 0.126 [P = .099]; 37.0% agreement). The AI recommended revision for 61.1% of manuscripts, of which 72.7% were ultimately rejected or cascaded, demonstrating systematic "revision bias." When the AI recommended rejection, 54.5% of those manuscripts were ultimately rejected and 27.3% were cascaded; when the AI recommended cascade, 50% were rejected and 50% were cascaded. However, when the AI recommended rejection or cascade (n = 21), 90.5% received a final decision of rejection or cascade (PPV, 90.5%; specificity, 81.8%). Manuscripts with an AI score <70 were rejected or cascaded 88.0% of the time (PPV, 88.0%). CONCLUSION: AI cannot replicate the nuanced judgment of human peer reviewers at a high-impact sports medicine journal. When AI recommended rejection or cascade, 90.5% of manuscripts received that final decision (descriptive PPV, 90.5%; 95% CI, 71.1%-97.3%), suggesting potential utility as an exploratory first-pass screening tool warranting further validation in larger cohorts. However, AI could not reliably distinguish manuscripts destined for outright rejection from those that would be cascaded to a sister journal-an important limitation for editorial triage applications.

Sports Medicine

Artificial intelligence in genitourinary oncology: publication trends and systematic review.

OBJECTIVE: To conduct an analysis of publication trends and a systematic review of randomized controlled trials (RCTs) to characterize the current state of artificial intelligence (AI) use in genitourinary (GU) oncology, as AI has emerged as a transformative tool in healthcare with potential applications in diagnostics, treatment planning, and prognostication. METHODS: We searched the Medical Literature Analysis and Retrieval System Online (MEDLINE), Excerpta Medica dataBASE (EMBASE; Ovid), and Cumulative Index to Nursing and Allied Health Literature (CINAHL) Ultimate for studies related to AI and GU oncology, excluding non-English papers, non-human studies, review articles, and articles using AI solely for manuscript writing. Publication trends were analysed from 2013 to 2023 and categorized by study design and cancer type. RCTs were evaluated through systematic review using Covidence (Veritas Health Innovation Ltd, Melbourne, Victoria, Australia) for screening and data extraction. Two reviewers independently assessed all studies, with risk of bias (RoB) evaluated using the Cochrane RoB 2.0 tool. RESULTS: Of 2409 articles identified, 1220 met inclusion criteria. These included 962 retrospective articles, 175 prospective studies, 79 studies with combined retrospective/prospective methods, and four RCTs. Studies most commonly addressed prostate (n&#x2009;=&#x2009;923), renal (n&#x2009;=&#x2009;274), and urothelial (n&#x2009;=&#x2009;194) cancers. Publications grew from 14 in 2013 to 362 in 2023, with substantial acceleration in 2019. Four RCTs were identified - one in urothelial cancer and three in prostate cancer. Two RCTs evaluated AI-based diagnostics, demonstrating improved performance over conventional methods; the remaining two RCTs evaluated AI in prognostication and treatment planning, showing improved gains in imaging interpretation and operational efficiency. RoB varied across studies, primarily related to randomisation and deviations from intended interventions. CONCLUSIONS: Artificial intelligence research in GU oncology has grown, although high-level evidence from RCTs remains limited. Existing trials underscore AI's promise in diagnostics, prognostication, and treatment planning, and the rapidly evolving nature of this field warrants continued prospective investigation.

Humans