Search PubMedSearch

SEARCH · Search PubMed

Results for “External beam radiotherapy”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

126 records · Page 3Linked to original sources

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation

Advanced/Novel Stenting for Pediatric Dynamic Airway Collapse.

Pediatric dynamic airway collapse is a complex condition that can impact all levels of the pediatric airway. These conditions can pose life threatening risk to pediatric patients and carry lasting impacts. While traditionally, tracheostomy has been used to address all levels of dynamic collapse, recent advances have allowed for more individualized, anatomy-specific stenting and splinting strategies for treatment. This article covers pathophysiology and the latest evidence on strategies to address nasopharyngeal, oropharyngeal, proximal trachea, and tracheobronchial dynamic collapse.

Humans

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning

Outcomes of Stage IVA Cervical Cancer Treated with Radiation Therapy: A Systematic Review and Meta-Analysis.

FIGO stage IVA cervical cancer, defined by bladder or rectal mucosal invasion without distant metastasis, is an uncommon but clinically challenging disease with limited high-quality evidence to guide management. We performed a systematic review and meta-analysis to evaluate survival outcomes, treatment-related morbidity, and prognostic factors in patients with stage IVA cervical cancer treated with definitive radiotherapy. Following PRISMA 2020 guidelines and PROSPERO registration (CRD42024602426), PubMed, Embase, and Web of Science were searched through February 2025. Eleven studies comprising 492 patients met eligibility criteria. Pooled random-effects analyses demonstrated 2-, 3-, and 5-year disease-free survival rates of 41.5% (95% CI, 28.1-55.0), 34.2% (95% CI, 21.1-47.3), and 30.9% (95% CI, 16.0-45.8), respectively. Corresponding overall survival rates were 56.0% (95% CI, 46.2-65.9), 45.9% (95% CI, 35.9-55.9), and 34.8% (95% CI, 26.4-43.3). Weighted median disease-free survival and overall survival were 15.6 and 33.6 months, respectively. The pooled incidence of vesicovaginal or rectovaginal fistula was 23.7% (95% CI, 12.9-34.6). Adverse prognostic factors included pelvic nodal involvement, hydronephrosis, rectal invasion, omission of brachytherapy, and total EQD2 below 80-85 Gy. Concurrent chemoradiation and completion of brachytherapy were consistently associated with improved outcomes. Despite curative-intent treatment, long-term survival remains poor and treatment-related morbidity substantial. Durable pelvic control remains the principal therapeutic challenge. Concurrent chemoradiation with adequate-dose brachytherapy appears essential for optimal outcomes, while future stage IVA-specific studies incorporating image-guided adaptive brachytherapy, advanced radiotherapy techniques, and systemic treatment intensification are needed to improve survival and reduce treatment-related morbidity.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95 % CI 0.85-0.94; 95 % prediction interval 0.62-0.98), with sensitivity of 0.80 (95 % CI 0.77-0.83) and specificity of 0.87 (95 % CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Online Risk Behavior in Adolescents: A Systematic Review.

Identifying and categorizing online risk behaviors is crucial for assessing their impact on adolescents. Despite extensive research, previous studies have not provided a clear classification of these behaviors. This systematic review synthesizes the quantitative literature on adolescent online risk behaviors from the inception of research to September 2023, aiming to: (a) offer a comprehensive overview of the types of online risk behaviors and the specific actions encompassed within each category among adolescents; (b) summarize the adverse outcomes associated with these behaviors; and (c) discuss the implications and future research directions. Utilizing key terms, this study sourced studies from four electronic databases (Scopus, PubMed, Web of Science, and EMBASE), ultimately including 22 English-language quantitative studies. The review reveals that online risk behaviors are primarily categorized into content risk behaviors, contact risk behaviors, and conduct risk behaviors. Adolescents engaging in these behaviors are at an increased risk of experiencing physical health issues, mental health problems, externalizing behaviors, and even self-harm and suicidal thoughts or actions. Further research is needed to develop and validate an online risk behavior scale and conduct longitudinal and experimental studies to establish causal relationships and examine the long-term effects of these behaviors on adolescent well-being. The review concludes with implications for future research and potential prevention, intervention, and policy strategies to mitigate online risk behaviors in adolescents.

Humans

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Radiographic assessment and orthodontic intervention effects on orthodontically induced root resorption: a systematic review and meta-analysis of clinical trials.

The relative contribution of radiographic methods and characteristics of the orthodontic intervention to orthodontically induced root resorption (OIRR) remains unknown. The aims of this systematic review and meta-analysis were to (1) estimate the pooled OIRR effect across orthodontic intervention versus comparator contrasts, (2) compare pooled estimates by radiographic method (2D [two-dimensional] vs. 3D/CBCT [three-dimensional/cone-beam computed tomography]), and (3) explore whether force mechanics (intrusive versus nonintrusive) modified OIRR magnitude. Seven randomized controlled trials and one prospective study (January 2010-October 2025) were included. Only OIRR was the outcome, reported as correlation coefficients (r). The primary analysis combined within-study intervention-versus-comparator estimates. Subgroup analysis of 2D versus 3D/CBCT imaging was prespecified, whereas post-hoc analysis of intrusive versus nonintrusive mechanics was performed. The pooled analysis for the primary outcome showed a small, nonsignificant OIRR effect (r = 0.07; 95% confidence interval [CI]: -0.12 to 0.27; p = 0.372) with high heterogeneity (I2 = 84.0%). Radiographic method did not change the pooled estimates significantly (p = 0.331). Force-mechanics analysis showed that intrusive mechanics was related to significantly higher root resorption than nonintrusive mechanics (r = 0.40; 95% CI = 0.15 to 0.65 versus r = -0.03; 95% CI = -0.16 to 0.10; p < 0.001). This accounted for 87.1% of the between-study variance. The average orthodontic intervention effect on OIRR was small and not significant; however, the OIRR magnitude was strongly affected by force mechanics, particularly by intrusive forces. There was no significant difference in pooled estimates by radiographic method; however, 3D/CBCT provides superior volumetric quantification and should be used judiciously according ALARA (as low as reasonably achievable) principles.

Root Resorption

Ventriculostomy-Related Infections by Country-Income Level: A Systematic Review and Bayesian Hierarchical Meta-analysis.

Our objective was to perform a systematic review and meta-analysis of published literature on ventriculostomy-related infection (VRI) and evaluate temporal and global trends. We conducted a systematic review and Bayesian hierarchical random-effects meta-analysis of VRI rates in adults, stratified by country-income level (high-income countries [HIC]; low- or middle-income countries [LMIC]), study design, sample size, enrollment period, VRI intervention, and VRI definition. We identified 159 articles published between 1989 and 2025 that included 523,704 patients with 7293 VRIs. The pooled VRI rate was 8.64% [95% CI: 7.44-9.97], with moderate heterogeneity and good model fit. The leave-one-out sensitivity analysis showed a mean absolute change of 0.06% and a maximum change of 0.2%, indicating robust analysis. Five of the 33 represented countries had VRI rates below the global pooled rate of 8.64%. Four were HICs: Singapore (VRI rate 3.3% [0.8-7]), the United States (VRI rate 4.6% [3.4-5.9]), Germany (VRI rate 6.1% [1.1-18.9]), Norway (8.3% [0.3-68.4]), with 1 LMIC: China (8.5% [5.4-12.4]). VRI was significantly higher in studies using definitions beyond CSF culture alone for VRI (+3.16% [0.11- 6.52]) and in those from Europe (+7.29% [4.62-10.10]) and the Western Pacific (+4.09% [1.55-6.98]). No other subgroup demonstrated significant differences. This Bayesian meta-analysis provides global estimates and factors associated with VRI. Standardization of VRI definitions is critical for future benchmarking of VRI rates.

Humans

The role of artificial intelligence in the diagnosis and prognosis of traumatic brain injury based on brain CT scans: a systematic review.

Traumatic brain injury (TBI) is a leading cause of emergency department visits and a major contributor to injury-related mortality and long-term neurological disability. Non-contrast computed tomography (CT) is the gold-standard imaging modality for the rapid diagnosis of TBI. Clinical outcomes depend strongly on early detection and prompt acute management. Artificial intelligence (AI)-based models may support faster automated identification of traumatic findings and early prediction of patient prognosis.&#xa0;A systematic literature search was conducted in PubMed/MEDLINE, Scopus, IEEE Xplore, ACM Digital Library, and the Cochrane Library in accordance with PRISMA 2020 guidelines to evaluate AI-based models for automated detection of TBI-related findings on CT and for prediction of clinical outcomes. Risk of bias and applicability were assessed using QUADAS-2 for diagnostic accuracy studies and PROBAST&#x2009;+&#x2009;AI for prediction model studies.&#xa0;Twenty-two studies were included. Sixteen studies evaluated diagnostic tasks and 10 evaluated prognostic outcomes, with four studies contributing to both categories. Diagnostic performance was generally high, with many studies reporting AUC values approaching or exceeding 0.90, particularly for larger lesion volumes.Prognostic performance was more variable, with moderate to high discrimination and substantial heterogeneity. Only 9 studies incorporated independent external validation, and performance was frequently lower in external cohorts. All prognostic model studies were judged to be at high overall risk of bias using PROBAST&#x2009;+&#x2009;AI, and most diagnostic accuracy studies also demonstrated high or unclear risk of bias in at least one QUADAS-2 domain, most frequently in patient selection.&#xa0;AI-based models applied to brain CT demonstrate strong technical performance for both diagnostic and prognostic tasks in TBI. However, most studies relied on retrospective designs and lacked independent external validation which limits models generalizability and raises concern for potential overfitting. Prospective, multicenter studies with standardized methodologies and rigorous external validation are required before widespread clinical implementation.

Humans

Quo vadis, BGA? A collaborative EDNAP exercise on the challenges and progress in forensic biogeographical ancestry inference.

There is a broad consensus that forensic tests for the prediction of externally visible characteristics (EVC) and analysis of biogeographic ancestry (BGA) of an individual are technically reliable. However, interpretation of the results and population-specific genotype distribution patterns remains challenging. EVC and BGA analyses provide valuable information for population genetics studies and as investigative leads for criminal cases, as well as for historical and contemporary identification tests. However, inaccurate or incorrect predictions, for example, from subjective bias in the interpretations made, have the potential to misdirect police investigations. The legal situation regarding EVC and BGA testing varies by country: ranging from countries where it is explicitly prohibited, to those without specific regulations on biogeographic ancestry prediction, and others that have already enacted laws governing its use. The reluctance to utilize these analyses is not only due to legal restrictions and data protection concerns, but also to initial limited sets of sufficiently comprehensive forensic DNA assays. Forensic BGA marker panels typically contain up to &#x223c;300 SNPs. This relatively small number of genetic markers, along with limited reference population data, complicates the interpretation of results from donors of unknown origin. This paper presents the results of a collaborative EDNAP study, which, for the first time, evaluated the approach to reporting EVC and BGA data between international laboratories. For the study, DNA from nine individuals with self-reported ancestry was collected and analysed using various forensic panels differing in the number and composition of ancestry-informative markers genotyped, comprising: the Precision ID mtDNA Whole Genome Panel, the VISAGE Basic Tool and the VISAGE Enhanced Tool for Appearance and Ancestry Prediction, and the Ion AmpliSeq&#x2122; PhenoTrivium Panel. To ensure full data protection, all SNP genotypes and uniparental marker haplotypes obtained were not shared with third parties. Instead, the genetic data were analysed using a range of commonly used population analysis software packages. These analysis outcomes were then distributed to twelve European forensic laboratories (both academic and law enforcement institutions), who were asked to prepare reports based on their interpretation of the phenotypes and ancestry they inferred from the analysis data. A questionnaire sent alongside the genetic information, aimed to evaluate which difficulties were encountered by the participants in processing the BGA analysis data they were given.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Ensemble DNA methylation clock demonstrates Immune-metabolic aging signatures associated with mortality.

Aging is a multifactorial process that is best described in terms of the progressive acquisition of multiple layers of phenotypic changes, such as epigenetic modifications, inflammation, and metabolic dysregulation. DNA methylation clocks have been extensively used to construct epigenetic clocks based on the DNAm profiles that can be used to estimate biological age and predict age-associated outcomes. Nevertheless, the vast majority of clocks constructed so far have been based on linear models, which are unlikely to fully account for the heterogeneity and non-linearity of survival-related DNAm signatures. In this work, we constructed a heterogeneous stacked ensemble survival model based on DNAm data obtained from the Framingham Heart Study. We first identified 190 CpG loci using elastic net Cox regression and subsequently constructed a survival prediction model based on the fusion of five complementary survival models by means of a neural network meta-learner. The prediction power of the survival model was evaluated in an external validation cohort, where we observed strong performance for predicting all-cause mortality that significantly exceeded PhenoAge and was statistically comparable to GrimAge. These performance estimates were derived in cohorts of European ancestry and externally validated in postmenopausal women aged 50-79 years, and should therefore be interpreted as applicable only to demographically similar populations.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

The Dual Role of Executive Functioning in the Association Between Family Socioeconomic Status and Children's Problem Behaviors.

Although prior research suggests that executive functioning may either mediate or moderate the association between family socioeconomic status (SES) and children's problem behaviors, examining these roles separately provides an incomplete account: mediation models may understate individual differences that are not attributable to the environment, whereas moderation models may understate the role of the environment in shaping personal characteristics. To integrate these perspectives, the present longitudinal study examined whether executive functioning simultaneously mediates and moderates the association between family SES and children's internalizing and externalizing problems. A total of 308 children in the early years of elementary school (MageT1 = 7.34 years; 138 girls) were assessed and followed up 39 months later. After controlling for children's gender, age, and grade, lower family SES at T1 significantly predicted higher levels of both internalizing and externalizing problems at T2. Executive functioning at T1 partially mediated these associations, indicating that differences in children's executive functioning partly accounted for socioeconomic disparities in problem behaviors. Executive functioning also moderated both associations: SES was negatively associated with problem behaviors among children with lower executive functioning but not among those with higher executive functioning. These findings highlight the dual role of executive functioning in the longitudinal association between SES and children's problem behaviors and suggest that executive functioning may be a promising target for efforts to reduce mental health disparities associated with socioeconomic disadvantage.

Humans

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans

Not just when, but how: An exploratory dual-control approach to video feedback in motor learning.

The present study provides exploratory evidence for a novel dual-control paradigm. It examines whether combining temporal over video feedback timing with learner-controlled interactive playback functions (pause, slow-motion, rewind) would enhance motor skill acquisition beyond temporal autonomy alone. Sixty-four novice adults were randomly assigned to one of four conditions: Full Control (self-controlled timing + interactive replay), Partial Control (self-controlled timing + non-interactive replay), Yoked Full Control (externally controlled timing + interactive replay), or Yoked Partial Control (externally controlled timing + non-interactive replay). Motor accuracy (Radial Error), movement consistency (Bivariate Variable Error), technical execution, and self-efficacy were assessed at pre-test, 24-h retention, and 72-h retention following two acquisition sessions on a dart-throwing task (120 trials total). The Full Control group demonstrated the greatest and most durable learning gains across all outcomes. The Group &#xd7; Time interaction was significant across all dependent variables (&#x3b7;2&#x209a; ranging from 0.140 to 0.234), with Full Control demonstrating superior retention at both 24 and 72&#xa0;h relative to other groups (though differences relative to Partial Control were more pronounced at 72-h retention). Critically, the Yoked Full Control group showed comparatively weaker outcomes despite access to the same interactive playback functions. These findings suggest that interactive video tools may be most useful when learners can regulate both when feedback is accessed and how it is inspected. Theoretical and practical implications for the design of learner-centered video feedback systems are discussed.

Humans