Search PubMedSearch

SEARCH · Search PubMed

Results for “learning curve”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

233 recordsLinked to original sources

Clinical Efficacy and Learning Curve of Far-Lateral Approach (FLA) in Uni-Portal Non-Coaxial Spinal Endoscopic Surgery (UNSES) in the Treatment of Lumbar Degenerative Diseases: A Prospective Study.

BACKGROUND: Uniportal non-coaxial spinal endoscopic surgery (UNSES) via far-lateral approach (FLA) is an innovative minimally invasive procedure for lumbar degenerative diseases, particularly far-lateral disc herniation and foraminal stenosis. However, complex lateral lumbar anatomy and strict endoscope-instrument coordination create a distinct learning curve that may compromise early surgical efficiency and safety. This study aimed to evaluate the efficacy and safety, quantify the learning curve, and to provide clinical guidance for the standardized promotion and application of this technology. METHODS: A total of 40 consecutive patients with lumbar degenerative diseases who underwent UNSES via FLA by a single surgeon between January 2025 and December 2025 were included. All data were analyzed using SPSS 26.0 statistical software (IBM, USA). Primary outcomes included operation time, blood loss, fluoroscopy frequency, and intraoperative complication rate. Secondary outcomes were VAS, ODI, and modified Macnab criteria at 1, 3, and 6&#x2009;months postoperatively. The learning curve and the inflection point of the learning curve was determined using cumulative sum (CUSUM) analysis. The differences in clinical indicators between early and proficient stage were compared. RESULT: Operation time, blood loss, and fluoroscopy times decreased significantly with case accumulation (p&#x2009;<&#x2009;0.05). CUSUM identified an inflection point at the 16th case, after which operation time stabilized at (55.3&#x2009;&#xb1;&#x2009;8.6) min, much shorter than the early phase (89.5&#x2009;&#xb1;&#x2009;10.3) min (p&#x2009;<&#x2009;0.001). Before the 16th case, the curve was in an upward trend; after the 16th case, the curve tended to be flat, indicating the proficiency stage. Postoperative VAS and ODI improved significantly than those before surgery at each follow-up time (p&#x2009;<&#x2009;0.05). There was no significant difference in postoperative VAS score and ODI between the two groups at each follow-up time point (p&#x2009;>&#x2009;0.05). The total complication rate was 12.5% (5/40), were cured by conservative treatment. The total excellent-good rate was 90.0% (36/40). L5/S1 and Bertolotti's syndrome were independent factors affecting the learning curve. CONCLUSION: UNSES via FLA is a safe and effective minimally invasive technique for treating complex lumbar degenerative diseases. It has a certain learning curve, and the inflection point is about the 16th case. After mastering the key techniques such as anatomical positioning, endoscopic manipulation and hemostasis, the surgeon can gradually reach the proficiency stage, with significantly improved surgical efficiency and clinical efficacy, and controllable complications. This study provides a theoretical basis for the clinical training and technology promotion of UNSES via FLA.

Humans

Distal versus proximal radial access for diagnostic cerebral angiography: comparative outcomes and learning curve analysis.

BACKGROUND AND PURPOSE: Distal transradial access (dTRA) is an alternative to proximal transradial access (pTRA) for neuroangiography, but comparative real-world data and evidence on its early learning curve remain limited. We compared procedural performance and access-site complications between dTRA and pTRA and evaluated the early learning curve of dTRA. METHODS: We retrospectively analyzed 470 diagnostic cerebral angiography procedures, representing 421 unique patients, performed via radial access at a single center between January 2025 and February 2026, including 237 dTRA and 233 pTRA procedures. Baseline characteristics, including age, sex, body mass index (BMI) category, aortic arch type, and antiplatelet/anticoagulant use, procedural performance, and clinically assessed access-site events were compared between groups. Radial artery occlusion (RAO) was assessed by postoperative bedside pulse examination and confirmed with Doppler ultrasound when clinical findings were uncertain. Multivariable logistic regression was used to evaluate predictors of RAO, persistent bleeding or repeated compression, hand edema, and a composite access-site event endpoint. Because repeated procedures occurred in a subset of patients and event counts were limited, first-procedure sensitivity analysis and analyses of infrequent outcomes were interpreted cautiously. The dTRA learning process was assessed in the first 100 dTRA cases performed by a single operator using multivariable regression, cumulative sum (CUSUM) analysis, segmented trend analysis, and phase-based comparisons. RESULTS: Baseline characteristics were comparable between groups, including age, male sex, BMI category, aortic arch type, and antiplatelet/anticoagulant use. Compared with pTRA, dTRA was associated with more puncture attempts (3.0 [2.0-4.0] vs 2.0 [1.0-3.0], P&#xa0;<&#xa0;0.001), longer puncture time (2.0 [1.0-5.0] vs 2.0 [1.0-3.0] min, P&#xa0;=&#xa0;0.003), lower first-pass success (19.4% vs 35.2%, P&#xa0;<&#xa0;0.001), and a higher crossover rate (11.4% vs 6.0%, P&#xa0;=&#xa0;0.037). However, dTRA was associated with a lower clinically assessed RAO rate (2.5% vs 7.7%, P&#xa0;=&#xa0;0.011). On multivariable analysis, pTRA was independently associated with higher odds of RAO (OR 3.27, 95% CI 1.26-8.49, P&#xa0;=&#xa0;0.015) and the composite access-site event endpoint (OR 3.12, 95% CI 1.55-6.28, P&#xa0;=&#xa0;0.001). Similar findings were observed in a sensitivity analysis restricted to the first procedure per patient. In the first 100 dTRA cases, cumulative dTRA experience was independently associated with shorter total procedure time (beta&#xa0;=&#xa0;-0.074&#xa0;min/case, P&#xa0;=&#xa0;0.009), while CUSUM and moving-average analyses suggested that the major learning effect occurred within approximately the first 10-15 cases. CONCLUSIONS: In this retrospective single-operator cohort, dTRA was associated with lower clinically assessed RAO than pTRA despite greater access difficulty. The early learning effect was mainly reflected in shorter total procedure time. These findings support the feasibility of dTRA but should be interpreted cautiously given the study's observational design and limited anatomical data.

Humans

Structured robotic colorectal training in a non-tertiary NHS hospital: a 502-case consecutive cohort implementation study.

Robotic-assisted colorectal surgery has expanded rapidly across NHS practice in the UK. Structured unit-wide training pathways are essential for safe technology adoption, yet published outcome data from non-tertiary hospitals remain limited. This study describes the implementation and feasibility of a unit-wide robotic colorectal program at a high-volume non-tertiary hospital, reporting outcomes across 502 consecutive resections performed by eight consultant surgeons and presenting these in the context of nationally published benchmarks. A retrospective cohort study of 502 consecutive robotic colorectal resections performed at York Teaching Hospital between May 2022 and December 2025. Eight consultant surgeons (A-H) participated in a structured four-phase training pathway incorporating simulation training, proctored cases, complexity-based case progression, and formal credentialing. Primary outcomes were 30-day mortality, unplanned return to theatre (RTT), and anastomotic leak (AL). Anastomotic leak was calculated using only patients who underwent anastomosis as the denominator. Procedure-stratified and individual surgeon outcomes with 95% confidence intervals were reported. Risk-adjusted cumulative sum (RA-CUSUM) analysis was performed to evaluate learning curves. Outcomes are presented descriptively alongside nationally published reference data; no formal statistical comparison against national benchmarks was performed. 502 robotic colorectal resections were performed. Mean patient age was 70.0 &#xb1; 11.3&#xa0;years; 58.4% were male. Median ASA grade was III. The indication was malignancy in 89.2% of cases. Length of stay was non-normally distributed and is therefore reported using median and interquartile range in the revised analysis. Key outcomes: - 30-day mortality: 1.0% (5/502; 95% CI 0.4-2.3%) - Unplanned return to theatre (RTT): 5.2% (26/502; 95% CI 3.6-7.5%) - Anastomotic leak (AL): 3.3% (15/450; 95% CI 2.0-5.5%; denominator = patients with anastomosis) - 30-day unplanned readmission: 5.0% (25/502; 95% CI 3.4-7.2%) - Conversion to open surgery: 3.6% (18/502; 95% CI 2.3-5.6%) - Lymph node yield &#x2265;12: 91.3% of cancer resections - R0 resection rate: 95.1% of cancer resections All primary outcomes fell within or below the published reference ranges used for descriptive context. RA-CUSUM trajectories were heterogeneous: no surgeon crossed the predefined upper control limit, but several curves showed later upward movement. Accordingly, the analysis is interpreted as safety surveillance rather than evidence of uniform performance improvement. RA-CUSUM monitoring showed that no surgeon crossed the predefined upper control limit; however, heterogeneous trajectories precluded a claim of uniform performance improvement.

Humans

Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis.

BACKGROUND: Emergency radiographic interpretation for fractures is prone to missed or misdiagnoses. Artificial intelligence (AI) is expected to become a powerful tool to assist clinicians in fracture detection. PURPOSE: A systematic review and meta-analysis was performed to assess whether AI improves clinicians' ability to detect fractures on radiographs. MATERIALS AND METHODS: A literature search was conducted in PubMed, Web of Science, and Cochrane Library for studies published between January 1, 2010, and October 10, 2025. A meta-analysis of diagnostic accuracy studies was performed using a Summary Receiver Operating Characteristic (SROC) curve. The quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Subgroup analysis and meta-regression were conducted to explore potential sources of heterogeneity. RESULTS: A total of 26 studies were included . The pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94). The corresponding AUC values were 0.90 (95% CI: 0.87-0.92) before and 0.95 (95% CI: 0.93-0.97) after AI assistance. Eight studies were rated as high risk of bias. Subgroup analysis and meta-regression identified potential sources of heterogeneity, including fracture location, AI model type, high risk of bias, and reference standards. CONCLUSION: AI assistance significantly improves clinicians' diagnostic performance in detecting fractures on radiographs for extremity and trunk fractures.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n&#xa0;=&#xa0;549) and a validation set (n&#xa0;=&#xa0;236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60&#xa0;mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60&#xa0;mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

Integrative machine learning and transcriptomic analysis reveals molecular mechanisms underlying low survival rate in larval Chinese Bahaba (Bahaba taipingensis).

Chinese Bahaba (Bahaba taipingensis) is a Class I protected marine fish endemic to China. Low larvae survival during artificial breeding severely hinder population recovery. To investigate the molecular mechanism of high mortality in larval fish, this study performed RNA-seq on liver from naturally deceased (ND) and mass-dead (MD) individuals, combined with least absolute shrinkage and selection operator (LASSO) regression and random forest (RF) algorithms to screen for core signature genes. A total of 873 differentially expressed genes (DEGs) were identified, including 112 upregulated and 761 downregulated genes. GO and KEGG enrichment analyses revealed significant enrichment in amino acid metabolism disorders, one&#x2011;carbon folate pool impairment, PPAR signaling abnormalities, ECM-receptor interaction, focal adhesion pathway, indicating widespread metabolic suppression accompanied by extracellular matrix remodeling and signaling disturbances in the livers of MD fish. MAD pre-filtering combined with dual machine learning algorithms yielded 18 robust core signature genes, among which SLC38A4, MMP1, FADD, FKBP5, and APOB were consistently identified as high-frequency core genes by both algorithms. SLC38A4 exhibited the highest importance score in the RF model and was significantly downregulated, making it the primary molecule distinguishing ND from MD phenotypes. ROC curve analysis showed that both models achieved an AUC of 1.000 (95% CI lower bound: 0.610), confirming the precise discriminatory ability of the core genes. GSEA further demonstrated significant enrichment of this core gene set in ND samples. This study provides the first systematic elucidation of the molecular mechanisms underlying liver dysfunction in low survival rate B. taipingensis, characterized by amino acid transport impairment, metabolic reprogramming, and structural remodeling, offering theoretical foundations for health assessment, early mortality risk warning, and artificial breeding conservation of this species.

Animals

The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis.

BACKGROUND: Current endoscopic retrograde cholangiopancreatography (ERCP) and cholangioscopic-based diagnostic sampling for indeterminant biliary strictures remain suboptimal. Artificial intelligence (AI)-based algorithms by means of computer vision in machine learning have been applied to cholangioscopy in an effort to improve diagnostic yield. The aim of this study was to perform a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based diagnostic performance of AI-associated cholangioscopic diagnosis of indeterminant or malignant biliary strictures. METHODS: Individualized searches were developed in accordance with PRISMA and MOOSE guidelines, and meta-analysis according to Cochrane Diagnostic Test Accuracy working group methodology. A bivariate model was used to compute pooled sensitivity and specificity, likelihood ratio, diagnostic odds ratio, and summary receiver operating characteristics curve (SROC). RESULTS: Five studies (n=675 lesions; a total of 2,685,674 cholangioscopic images) were included. All but one study analyzed a deep learning AI-based system using a convoluted neural network (CNN) with an average image processing speed of 30 to 60 frames per second. The pooled sensitivity and specificity were 95% (95% CI: 85-98) and 88% (95% CI: 76-94), with a diagnostic accuracy (SROC) of 97% (95% CI: 95-98). Sensitivity analysis of CNN studies (4 studies, 538 patients) demonstrated a pooled sensitivity, specificity, and accuracy (SROC) of 95% (95% CI: 82-99), 88% (95% CI: 72-95), and 97% (95% CI: 95-98), respectively. CONCLUSIONS: Artificial intelligence-based machine learning of cholangioscopy images appears to be a promising modality for the diagnosis of indeterminant and malignant biliary strictures.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans

Adverse Experiences in Brief Meditation Practices: Randomized Controlled Trial.

BACKGROUND: Meditation has become increasingly popular in recent decades. However, relatively little remains known about the prevalence of and risk factors for adverse experiences related to a single meditation practice. OBJECTIVE: The objective of our study was to examine adverse experiences associated with 3 brief, digitally delivered meditation practices (mindfulness, self-compassion, and gratitude) relative to using the internet as usual, as well as to investigate whether preintervention characteristics could predict such outcomes. METHODS: In a secondary analysis of a randomized controlled trial using samples that were representative of the US and UK adult populations with regard to ethnicity, sex, and age, we examined adverse experiences associated with 3 brief (ie, 5 or 10 minutes) meditation practices (ie, mindfulness, self-compassion, and gratitude) relative to using the internet as usual. We also investigated the potential of using preintervention characteristics to predict such outcomes. RESULTS: A total of 5049 participants completed all preintervention measures and were randomly assigned to meditation or control conditions. Across the sample, 4.1% (204/4925) of participants reported having a distressing experience during the intervention, and 7.1% (348/4908) of participants experienced an increase in negative affect from before to after the intervention. The results showed that participants who were randomized to a brief meditation intervention were no more likely to report a distressing experience than those who were randomized to use the internet as usual (odds ratio [OR] 1.05, 95% CI 0.76-1.47; P=.76). The results also showed that participants who were randomized to a brief meditation intervention were less likely to report clinically relevant increases in negative affect relative to using the internet as usual (OR 0.63, 95% CI 0.50-0.80; P<.001). Notably, participants in the 10-minute condition had a significantly higher likelihood of reporting a distressing experience than those in the 5-minute condition (OR 1.42, 95% CI 1.07-1.89; P=.02). Preintervention characteristics showed acceptable discrimination ability to predict a distressing experience (area under the curve=0.73) and slightly lower ability to predict increased negative affect (area under the curve=0.67). CONCLUSIONS: Taken together, we found that the brief, digitally delivered meditation practices tested in this study carry risks of adverse experiences that are comparable to or lower than those of typical activities on the internet; 10-minute condition was more likely to result in distressing experiences than 5-minute condition; and adverse responses to a brief meditation practice can, at least to a certain degree, be predicted using preintervention characteristics. TRIAL REGISTRATION: Open Science Framework 94HKS; https://osf.io/94hks/overview.

Humans

Applications of artificial intelligence in robot-assisted surgery: a systematic review.

To characterize applications of artificial intelligence (AI) in robot-assisted surgery, summarize technical and clinical performance, and assess the quality of the available evidence. PubMed, Web of Science Core Collection, and Scopus were searched for English-language journal articles published from 1 January 2020 through 31 October 2025. Randomized, observational, model-development, validation, and feasibility studies evaluating AI in robot-assisted surgery or closely related image-guided minimally invasive workflows were eligible. Two reviewers independently performed study selection, data extraction, and risk-of-bias assessment. Owing to heterogeneity in surgical procedures, AI tasks, analytical units, validation strategies, and outcomes, findings were synthesized descriptively without statistical pooling. The review was registered in the International Prospective Register of Systematic Reviews (CRD420251175699). Seventeen studies were included: seven clinical prediction or decision-support studies, eight intraoperative recognition, segmentation, or image-guided studies, and two training or workflow studies. Five prediction studies reported area-under-the-curve values of 0.74-0.95. Technical studies reported F1 or Dice scores of 0.525-0.995 and task-specific accuracies of 0.840-0.998. Two randomized studies suggested benefits for personalized suturing feedback and automated camera control, but neither established improved patient outcomes. Only one study had low overall risk of bias; the remaining studies were at high or unclear risk or raised some concerns. AI applications in robot-assisted surgery show promise for prediction, intraoperative perception, training, and workflow support. Evidence primarily demonstrates technical feasibility rather than established clinical effectiveness. Independent multicenter validation and prospective evaluation of patient, educational, and workflow outcomes are required before widespread implementation.

Robotic Surgical Procedures

Could the preoperative urethral curve be used to predict immediate urinary continence following Retzius-sparing robot-assisted radical prostatectomy? A retrospective multi-center study.

PURPOSE: Immediate urinary continence (UC) recovery following Retzius-sparing robot-assisted radical prostatectomy (RS-RARP) remains highly variable, highlighting the need for reliable preoperative prediction. We aimed to develop and validate models to identify patients likely to achieve immediate UC recovery following RS-RARP. MATERIALS AND METHODS: A total of 580 prostate cancer patients who underwent RS-RARP from four medical centers were assigned to a training set (n=348), an internal validation set (n=103) and an external validation set (n=129). Independent predictors were identified through univariate analysis and LASSO regression. A nomogram was constructed using multivariate logistic regression. Its performance was evaluated with receiver operating characteristic (ROC) curve, calibration curves, and decision curve analysis. RESULTS: Immediate UC recovery was observed in 84.5% (294/348) of patients in the training cohort, 80.6% (83/103) in the internal validation cohort, and 81.4% (105/129) in the external validation cohort, respectively. Multivariate analysis identified membranous urethral length (MUL) (OR=1.23, P=0.029) and urethral curvature (OR=2.84, P<0.001) as independent predictors, while prostate volume (PV) (OR=0.84, P <0.001) as a protective factor. The nomogram integrating MUL, PV, and urethral curvature demonstrated superior predictive accuracy, with an AUC of 0.87 (95% CI, 0.83-0.91) in the training cohort. The bootstrap-corrected calibration slope was 0.96, and the Brier score was 0.08.&#xa0;Calibration curves and decision curve analysis confirmed the predictive accuracy and clinical utility of the nomogram. CONCLUSIONS: Our study introduces a novel quantitative method for assessing urethral curvature. The mpMRI-based model, integrating urethral curvature and prostate spatial configuration, offers enhanced predictive accuracy for postoperative immediate UC recovery.

Humans

Pharmacokinetic Differences Between Fast-Acting, Standard, and Placebo Cannabis Edibles.

INTRODUCTION: Edibles have become the second-most used cannabis product in legal U.S. states, wherein 64% of cannabis consumers reported using edibles within the past year. Among expansions to the legal cannabis industry are the newly marketed "fast-acting" edible compounds, which may address many of the issues associated with edible use related to overdose and dose management. The study hypotheses were that fast-acting edibles would reach peak concentration significantly faster than standard edibles and placebo edibles. MATERIALS AND METHODS: Twenty participants completed three arms within-subjects designed study to test hypotheses. The three arms were ingestion of a (1) fast-acting edible, (2) a standard edible, and (3) a &#x394;9-tetrahydrocannabinol (THC) terpene-derived placebo edible that was indistinguishable from the two THC-containing edibles. Blood plasma was analyzed for the presence of THC and THC analytes. The pharmacokinetic parameters tested were time to max concentration (Tmax), maximum concentration (Cmax), terminal half-life (t1/2), and area under the curve (AUC). RESULTS: Results supported study hypotheses in that Tmax was significantly faster for the fast-acting edible, observed 30 min post-ingestion and, on average, 30 min earlier than the Tmax for the standard edible. There were no significant differences between the fast-acting and standard edibles on Cmax, t1/2, and AUC; however, both the fast-acting and standard edibles were significantly different compared with the placebo across all pharmacokinetic parameters. DISCUSSION: The results indicate that the microencapsulation technology used to create the fast-acting edible enabled analyte concentrations to peak significantly faster compared to the standard and placebo edibles.

Humans