Search PubMedSearch

SEARCH · Search PubMed

Search Search PubMed

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

3 recordsLinked to original sources

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

Diagnostic accuracy of nuclear STAT6 immunohistochemistry for solitary fibrous tumour: a systematic review and meta-analysis.

Nuclear STAT6 immunohistochemistry is the diagnostic surrogate for the NAB2::STAT6 fusion of solitary fibrous tumour (SFT); its sensitivity is established, but specificity varies for unexamined reasons. This review quantified pooled accuracy and tested whether antibody clone and nuclear threshold govern specificity. PubMed, Scopus and Web of Science were searched to 29 June 2026 for studies reporting nuclear STAT6 immunohistochemistry against a reference standard (NAB2::STAT6 confirmation and/or expert consensus) in SFT and comparators, with extractable two-by-two data. Two reviewers screened, extracted data and applied QUADAS-2. A bivariate generalised linear mixed model gave summary sensitivity and specificity, and exploratory subgroup analysis and meta-regression tested antibody clone, anatomical site and reference-standard type. Twenty-three studies (1216 SFT and 4715 comparators) were included. Summary sensitivity was 98.7% (95% confidence interval 96.7-99.5) and specificity 99.1% (97.8-99.6); the diagnostic odds ratio was approximately 8656. The monoclonal YE361 subgroup (8 studies) reached specificity 99.9% (99.3-100), with one false positive among 861 comparators, versus 98.1% (96.0-99.1) for polyclonal and other antibodies. False positives concentrated in dedifferentiated liposarcoma and prostatic stromal tumours. Estimates were stable after removing studies at higher risk of bias (98.9%/99.1%) and on leave-one-out analysis; the Deeks test was non-significant (p = 0.08). Nuclear STAT6 immunohistochemistry is therefore highly sensitive and specific for SFT, and the residual specificity loss is structured and largely avoidable: the monoclonal YE361 read at a strict nuclear threshold is preferred, with MDM2 and CDK4 applied to exclude dedifferentiated liposarcoma when nuclear STAT6 is unexpectedly positive.

Humans

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans