Search PubMedSearch

SEARCH · Search PubMed

Results for “Diagnostic test accuracy”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

691 recordsLinked to original sources

Diagnostic accuracy of bronchoalveolar lavage fluid-based testing for pulmonary cryptococcosis: A systematic review and meta-analysis.

BACKGROUND: Pulmonary cryptococcosis(PC) presents diagnostic challenges because of its non-specific clinical and radiological manifestations. Bronchoalveolar lavage fluid (BALF)-based testing, which includes latex agglutination (LA) and lateral flow assay (LFA), offers a minimally invasive diagnostic method, yet its pooled diagnostic accuracy remains unclear. METHODS: We systematically searched PubMed, Embase, Cochrane Library, and Scopus from inception to May 2026. Studies evaluating BALF-based testing for PC with extractable 2 × 2 data were included. The methodological quality of relevant studies was assessed by the QUADAS-2 tool. Pooled sensitivity, specificity, likelihood ratios, and diagnostic odds ratio (DOR) were estimated using a bivariate random-effects model. Subgroup analyses were performed by testing method and reference standard type. Heterogeneity was evaluated through paired forest plots, HSROC visualization, and exploratory bivariate meta-regression. RESULTS: The pooled sensitivity was 0.87 (95% CI: 0.81-0.91), and the specificity was 0.99 (95% CI: 0.982 - 0.995). The pooled positive likelihood ratio (PLR) was 88.00 (95% CI: 47.39 - 163.42), the negative likelihood ratio (NLR) was 0.13 (95% CI: 0.09 -0.20), and the DOR was 658.50 (95% CI: 285.36-1519.55). No significant threshold effect or publication bias was detected. Exploratory meta-regression suggested a possible assay-method effect in the joint model (P = 0.03), mainly driven by specificity (P = 0.01). CONCLUSIONS: The study demonstrates the high accuracy of CrAg in BALF for the diagnosis of pulmonary cryptococcosis, supporting its role as an important adjunctive diagnostic tool, particularly when tissue biopsy is not feasible or rapid results are needed. Larger prospective studies with standardized protocols are needed to validate these estimates.

Humans

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Diagnostic accuracy of nuclear STAT6 immunohistochemistry for solitary fibrous tumour: a systematic review and meta-analysis.

Nuclear STAT6 immunohistochemistry is the diagnostic surrogate for the NAB2::STAT6 fusion of solitary fibrous tumour (SFT); its sensitivity is established, but specificity varies for unexamined reasons. This review quantified pooled accuracy and tested whether antibody clone and nuclear threshold govern specificity. PubMed, Scopus and Web of Science were searched to 29 June 2026 for studies reporting nuclear STAT6 immunohistochemistry against a reference standard (NAB2::STAT6 confirmation and/or expert consensus) in SFT and comparators, with extractable two-by-two data. Two reviewers screened, extracted data and applied QUADAS-2. A bivariate generalised linear mixed model gave summary sensitivity and specificity, and exploratory subgroup analysis and meta-regression tested antibody clone, anatomical site and reference-standard type. Twenty-three studies (1216 SFT and 4715 comparators) were included. Summary sensitivity was 98.7% (95% confidence interval 96.7-99.5) and specificity 99.1% (97.8-99.6); the diagnostic odds ratio was approximately 8656. The monoclonal YE361 subgroup (8 studies) reached specificity 99.9% (99.3-100), with one false positive among 861 comparators, versus 98.1% (96.0-99.1) for polyclonal and other antibodies. False positives concentrated in dedifferentiated liposarcoma and prostatic stromal tumours. Estimates were stable after removing studies at higher risk of bias (98.9%/99.1%) and on leave-one-out analysis; the Deeks test was non-significant (p = 0.08). Nuclear STAT6 immunohistochemistry is therefore highly sensitive and specific for SFT, and the residual specificity loss is structured and largely avoidable: the monoclonal YE361 read at a strict nuclear threshold is preferred, with MDM2 and CDK4 applied to exclude dedifferentiated liposarcoma when nuclear STAT6 is unexpectedly positive.

Humans

Evaluation of three Aspergillus antibody assays for screening of chronic pulmonary aspergillosis: prospective diagnostic accuracy study.

OBJECTIVES: Chronic pulmonary aspergillosis (CPA) is a frequent complication of pulmonary tuberculosis (PTB), particularly in high-burden settings where access to reliable serological diagnostics remains limited. We evaluated the diagnostic performance of two immunochromatographic technology (ICT) lateral flow assays (LFAs) and an ELISA for CPA screening among patients with active or previously treated PTB. METHODS: In this two-year prospective multicentre diagnostic evaluation, serum from adults with prior or active PTB was tested using the Era Biology Aspergillus IgG ICT LFA, LDBio Aspergillus IgG/IgM ICT LFA, and Bordier Aspergillus fumigatus IgG ELISA. CPA diagnosis was established using a consensus composite reference standard incorporating clinical, immunological, radiological, and microbiological criteria. The Bordier ELISA was used as part of the immunological component of the consensus CPA diagnosis, with a cutoff optical density of ≥1.0. Diagnostic accuracy, agreement statistics, receiver operating characteristic analysis, and latent class analysis (LCA) were performed. RESULTS: Among 340 participants, 24 (7.06%) had CPA. Proportion of participants with positive antibody tests among all tested individuals were 6.76% for LDBio ICT LFA, 20.0% for Era Biology ICT LFA, and 11.47% for Bordier ELISA. Against consensus CPA diagnosis, Bordier ELISA showed 87.50% sensitivity and 94.30% specificity, LDBio ICT LFA 58.33% sensitivity and 97.15% specificity, and Era Biology LFA 66.67% sensitivity and 83.54% specificity. LCA estimated CPA prevalence at 7.72%. LCA-derived sensitivities and specificities were 86.58% and 99.92% for LDBio ICT LFA, 83.39% and 85.31% for Era Biology LFA, and 79.10% and 94.19% for Bordier ELISA. CONCLUSIONS: The Bordier ELISA showed high sensitivity and specificity, while the LDBio ICT LFA demonstrated very high specificity with strong LCA-derived performance. These findings support the use of ELISA for laboratory diagnosis and ICT as a point-of-care screening tool for CPA in resource-limited settings. Era Biology Aspergillus IgG LFA demonstrated moderate sensitivity and acceptable diagnostic performance, indicating its potential utility as a supplementary screening assay for CPA in settings where rapid, point-of-care testing is required.

Humans

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95&#xa0;% CI 0.85-0.94; 95&#xa0;% prediction interval 0.62-0.98), with sensitivity of 0.80 (95&#xa0;% CI 0.77-0.83) and specificity of 0.87 (95&#xa0;% CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

The Role of Artificial Intelligence Combined With Digital Cholangioscopy for Indeterminant and Malignant Biliary Strictures: A Systematic Review and Meta-analysis.

BACKGROUND: Current endoscopic retrograde cholangiopancreatography (ERCP) and cholangioscopic-based diagnostic sampling for indeterminant biliary strictures remain suboptimal. Artificial intelligence (AI)-based algorithms by means of computer vision in machine learning have been applied to cholangioscopy in an effort to improve diagnostic yield. The aim of this study was to perform a systematic review and meta-analysis to evaluate the diagnostic performance of AI-based diagnostic performance of AI-associated cholangioscopic diagnosis of indeterminant or malignant biliary strictures. METHODS: Individualized searches were developed in accordance with PRISMA and MOOSE guidelines, and meta-analysis according to Cochrane Diagnostic Test Accuracy working group methodology. A bivariate model was used to compute pooled sensitivity and specificity, likelihood ratio, diagnostic odds ratio, and summary receiver operating characteristics curve (SROC). RESULTS: Five studies (n=675 lesions; a total of 2,685,674 cholangioscopic images) were included. All but one study analyzed a deep learning AI-based system using a convoluted neural network (CNN) with an average image processing speed of 30 to 60 frames per second. The pooled sensitivity and specificity were 95% (95% CI: 85-98) and 88% (95% CI: 76-94), with a diagnostic accuracy (SROC) of 97% (95% CI: 95-98). Sensitivity analysis of CNN studies (4 studies, 538 patients) demonstrated a pooled sensitivity, specificity, and accuracy (SROC) of 95% (95% CI: 82-99), 88% (95% CI: 72-95), and 97% (95% CI: 95-98), respectively. CONCLUSIONS: Artificial intelligence-based machine learning of cholangioscopy images appears to be a promising modality for the diagnosis of indeterminant and malignant biliary strictures.

Humans

Artificial Intelligence in Diagnosing Depression Through Behavioural Cues: A Diagnostic Accuracy Systematic Review and Meta-Analysis.

AIM: To synthesise existing evidence concerning the application of AI methods in detecting depression through behavioural cues among adults in healthcare and community settings. DESIGN: This is a diagnostic accuracy systematic review. METHODS: This review included studies examining different AI methods in detecting depression among adults. Two independent reviewers screened, appraised and extracted data. Data were analysed by meta-analysis, narrative synthesis and subgroup analysis. DATA SOURCES: Published studies and grey literature were sought in 11 electronic databases. Hand search was conducted on reference lists and two journals. RESULTS: In total, 30 studies were included in this review. Twenty of which demonstrated that AI models had the potential to detect depression. Speech and facial expression showed better sensitivity, reflecting the ability to detect people with depression. Text and movement had better specificity, indicating the ability to rule out non-depressed individuals. Heterogeneity was initially high. Less heterogeneity was observed within each modality subgroup. CONCLUSIONS: This is the first systematic review examining AI models in detecting depression using all four behavioural cues: speech, texts, movement and facial expressions. IMPLICATIONS: A collaborative effort among healthcare professionals can be initiated to develop an AI-assisted depression detection system in general healthcare or community settings. IMPACT: It is challenging for general healthcare professionals to detect depressive symptoms among people in non-psychiatric settings. Our findings suggested the need for objective screening tools, such as an AI-assisted system, for screening depression. Therefore, people could receive accurate diagnosis and proper treatments for depression. REPORTING METHOD: This review followed the PRISMA checklist. PATIENTS OR PUBLIC CONTRIBUTION: No patients or public contribution.

Humans

AI echo INSIGHT study: A prospective blinded randomized trial of artificial intelligence echocardiogram interpretation.

BACKGROUND: Transthoracic echocardiography (TTE) is the most commonly performed cardiac imaging modality with over 30 million studies annually. Demand for timely expert interpretation continues to outpace capacity, creating diagnostic delays and inter-observer variability that impact patient care. Recent research has suggested computer vision artificial intelligence (AI) models can generate accurate preliminary comprehensive TTE reports, however, prospective evaluation is needed to determine whether AI-assisted TTE interpretation can improve clinician efficiency while preserving diagnostic accuracy. METHODS: AI ECHO INSIGHT is a prospective randomized blinded clinical trial conducted at Kaiser Permanente Northern California that will evaluate 1200 historical TTE studies (1000 consecutive unselected studies plus 200 with moderate or greater valvular disease) interpreted using three workflows: (1) AI-generated preliminary report finalized by a blinded cardiologist (AI-assisted); (2) cardiologist-generated preliminary report finalized by a blinded cardiologist (cardiologist-assisted); and (3) sonographer-generated preliminary report finalized by a blinded cardiologist (sonographer-assisted). The primary outcome is the rate of substantial change between preliminary and final reports, comparing the AI-assisted workflow to the pooled cardiologist-assisted and sonographer-assisted workflows. Secondary outcomes include cardiologist interpretation time for report finalization, superiority testing for diagnostic accuracy, and reporting consistency. CONCLUSION: AI ECHO INSIGHT is a prospective randomized blinded clinical trial evaluating the clinical impact of AI-assisted TTE interpretation on diagnostic accuracy, cardiologist efficiency, and reporting consistency in real-world echocardiography workflows. TRIAL REGISTRATION: ClinicalTrials.gov registration number NCT07229300.

Humans

Diagnostic criteria and severity assessment for syndesmosis injury using magnetic resonance imaging: A systematic review.

High ankle sprains involving syndesmosis injury present challenges in both diagnosis and severity assessment. Magnetic resonance imaging is widely regarded as the preferred modality for evaluating syndesmosis injury and related structural damage. This systematic review primarily examined the diagnostic utility of magnetic resonance imaging. Secondarily, it explores grading and prognostics of syndesmosis injuries with magnetic resonance imaging and identified possible imaging parameters predictive of injury severity. A comprehensive search of MEDLINE, Embase, CINAHL Complete, and Scopus was performed through February 12, 2025, following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Peer-reviewed human studies in English that used magnetic resonance imaging to assess syndesmosis injury were included. Excluded were review articles, case reports, abstract-only studies, and biomechanical or cadaveric investigations. Twenty-seven studies comprising 1931 ankles met inclusion criteria. Magnetic resonance imaging demonstrated high diagnostic accuracy for complete tears of the anterior and posterior inferior tibiofibular ligaments. Ancillary signs such as the ring-of-fire edema pattern, distal tibiofibular joint effusion, and widening of the distal joint space exhibited high specificity with variable sensitivity and may assist in grading injury severity. Magnetic resonance imaging in chronic syndesmosis injury primarily detects fibrotic scarring and post-injury changes. Evidence gaps remain regarding the parameters that best determine injury severity and indicate early surgical intervention in competitive athletes. Consolidating multiple magnetic resonance imaging findings into standardized diagnostic criteria may improve reliability and clinical decision-making.

Humans

Feasibility of implementation, diagnostic accuracy, and end-user impact of an electronic health record (EHR)-based ureteral stent tracking tool in a pediatric population.

INTRODUCTION & OBJECTIVES: Ureteral stent tracking systems have reduced stent retention in adults, but their accuracy and impact in pediatrics have been minimally explored. With low event rates in children, such tools may yield high false positives, raising questions on balancing event prevention with provider burden. We aimed to evaluate the feasibility, diagnostic accuracy, and end-user impact of an Electronic Surveillance Tool for Evaluating Nephroureteral stent Tracking (eSTENT) at our institution. STUDY DESIGN: eSTENT, implemented in 1/2024, flags ureteral stents at risk for retention based on implant documentation, expected explant date, and explant documentation. Monthly reports are generated for stents missing explant documentation. We retrospectively evaluated the diagnostic performance of eSTENT from 1/2024-8/2025 at our pediatric hospital. A usability survey including a validated 1-7 implementation score (higher = easier implementation) was distributed to pediatric urologists and operating room nurses. RESULTS: Of 172 cases with ureteral stent placement, eSTENT flagged 28 events (16%) in 24 patients. Of these, 26 represented documentation gaps where explant had been appropriate. Two flags had no documentation of explant, representing near miss events that were identified. No retained stents occurred, consistent with high sensitivity and modest specificity. There were no flags in the last 6 months of the study period. Survey response rate was 100% for surgeons and 55% for nurses. Before eSTENT, stents were not routinely tracked. All surgeons and 93% of nurses reported no added burden, despite occasional misidentification of retained stents. Three surgeons found eSTENT beneficial, four were neutral, and free-text responses generally cited eSTENT's "fail safe" nature as positive. Nurses suggested improvements, including user support and integrated documentation reminders. The average implementation score among both groups was 6/7, indicating easy adoption. DISCUSSION: While the impact of stent tracking tools in adult literature has been positive, our study emphasizes the feasibility of broader adoption at a pediatric hospital. Integration of eSTENT may avoid the potentially devastating consequences of a retained stent. Prioritizing sensitivity over specificity appears acceptable for a "never event" in patient safety. Our study is limited by the retrospective nature of data collection and survey bias. CONCLUSIONS: Though no stents were retained in the study period, eSTENT appropriately flagged two cases without added burden to most end-users. Further optimization is warranted, but adoption in pediatric centers may enhance care reliability.

Humans

Characteristics of p53 and Smad4 immunohistochemistry in pancreatic ductal adenocarcinoma and validation by next-generation sequencing.

BACKGROUND: Mutations in four major driver genes -KRAS, CDKN2A, TP53, and SMAD4- are central to the pathogenesis of pancreatic ductal adenocarcinoma (PDAC) and critically inform diagnosis, therapeutic decision-making, and prognostic assessment. Although next-generation sequencing (NGS) is widely regarded as the gold standard for detecting these mutations, its clinical application is often limited by suboptimal analytical efficiency and substantial economic cost. Among these genes, immunohistochemical (IHC) staining for the proteins encoded by TP53 and SMAD4 has been extensively adopted in routine pathology practice. However, standardized IHC pattern classification schemes and rigorous validation of their predictive accuracy for underlying genomic alterations remain lacking in PDAC. METHODS: We retrospectively enrolled 63 PDAC patients and systematically characterized the typical IHC expression patterns of p53 and Smad4. Targeted NGS was subsequently performed on all available tumor specimens, and the resulting mutational profiles were correlated with corresponding IHC findings. Diagnostic performance including sensitivity, specificity and accuracy of p53 IHC for predicting TP53 mutations and of Smad4 IHC for predicting SMAD4 mutations was rigorously evaluated. RESULTS: Among the four canonical driver genes, co-occurring double- or triple-gene mutations were prevalent; within TP53 and SMAD4, missense mutations constituted the most frequent variant type. Using NGS as the reference standard, we validated the diagnostic utility of a three-tiered p53 IHC classification system, particularly in fine-needle biopsy (FNB) specimens. Furthermore, we proposed a novel, refined Smad4 IHC pattern classification that incorporates an "intermediate" category, thereby expanding upon conventional binary interpretation. This new scheme achieved markedly improved mutation prediction accuracy (0.76) compared with traditional approaches (0.57). CONCLUSION: Our study highlights the complementary diagnostic value of p53 and Smad4 IHC relative to molecular testing in PDAC, especially when tissue is limited, as commonly encountered in FNB specimens. The newly established Smad4 IHC classification system, which integrates an intermediate expression category into the conventional two-tier framework, demonstrates superior clinical utility and enhances predictive accuracy for SMAD4 genomic alterations.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29&#x2009;709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85)&#x2009;and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Comparative evaluation of molecular technologies for the identification of prevalent non-tuberculous mycobacteria in pulmonary infections: a systematic review and meta-analysis.

BACKGROUND: The increasing prevalence of non-tuberculous mycobacteria pulmonary disease (NTM PD) is a burden to public health. Successful management of NTM PD critically depends on accurate species identification and reliable drug susceptibility testing to guide appropriate antibiotic therapy. Emerging molecular technologies offer rapid diagnostic solutions compared to conventional methods, but their performance varies. This study aims to provide a comprehensive evaluation of current molecular techniques for NTM identification and to present a global antibiotic resistance profile. METHODS: A systematic literature search was conducted in PubMed and Web of Science for studies published between 2005 and 2024. Studies applying molecular methods for NTM identification and resistance detection in humans were included. Data on study characteristics, diagnostic methods, sample types, sample sizes, identification sensitivity, and drug susceptibility results were extracted. Meta-analysis was performed using R with the meta4diag package. The quality of included studies was assessed using the QUADAS-2 tool. RESULTS: The analysis included 49 studies on NTM identification and 33 studies on antibiotic resistance. For species identification, all evaluated molecular technologies (MALDI-TOF MS, PCR-based methods, Sequencing, DNA chip, and DNA strip) demonstrated high pooled sensitivities (>0.92). Subgroup analysis revealed that sample type significantly affected performance for MALDI-TOF MS. Preliminary analysis of antibiotic resistance rates revealed varying patterns. For slowly growing mycobacteria, a significantly high Ethambutol resistance rate was observed in M. avium (69.20%). Among rapidly growing mycobacteria, resistance to Imipenem was notable (54.22%), and Clarithromycin resistance varied significantly within the Mycobacterium abscessus complex. CONCLUSION: Emerging molecular technologies have revolutionized the methodology for NTM identification with excellent performance. However, their performance can be influenced by sample type, particularly for MALDI-TOF MS. The alarming and heterogeneous antibiotic resistance patterns also highlight the critical need for rapid and accurate species identification and drug susceptibility testing to inform effective therapeutic strategies. Key messagesMolecular technologies demonstrate high accuracy for NTM identification.Antibiotic resistance is a serious concern with variations among NTM species and subspecies.Rapid and accurate species identification and drug susceptibility testing are crucial for guiding effective clinical management of NTM PD.

Humans

A child and young person focused systematic review of scrotal ultrasound with Mitigants to identify missed torsion.

BACKGROUND: Diagnostic accuracy of ultrasound for adults is stated as excellent. Paediatric and adult acute scrotal diseases are different. Findings in adults may not apply in children. The authors have seen overconfidence and interpretation with adult disease in mind in children with scrotal pain and have seen testicular loss as a result. OBJECTIVES: To interrogate the literature methodologically to report diagnostic accuracy of ultrasound (US) for the acute scrotum specifically in CYP with a follow-up methodology which allows identification of missed torsion. STUDY DESIGN: We performed systematic review and meta-analysis of studies reporting true diagnostic accuracy of ultrasound for testicular torsion (TT). PubMed, MEDLINE, Scopus, and Web of Science databases were searched from platform start until September 2023 using MeSH terms with protocol as detailed on PROSPERO. Multiple author search was undertaken. Two authors extracted data for included studies. Study quality was reported through QUADAS-2 methodology. Surgical findings or 12 months follow up in non-operative cases would be the gold standard to identify late atrophy due to missed torsion. Studies with methodology to identify missed torsion outwith the initial consult were included for meta-analysis. RESULTS: A total of 6790 cases were reported from the 38 papers identified. Studies were often flawed with poor definition of outcomes and short follow-up duration. Meta-analysis of the 10 high-quality studies revealed pooled sensitivity and specificity of 0.93 and 0.99. Doppler US by trained individuals demonstrated excellent diagnostic accuracy. Nonetheless, on extrapolation of the data, especially where flow was preserved, 5.6% of torsions were missed. DISCUSSION/CONCLUSION: Our review confirms a high diagnostic accuracy of US for TT in CYP. However, a 5.6% missed torsion rate was found if flow alone was considered diagnostic. The authors suggest US should be used as an adjunct in ambiguous cases, ensuring that clinically obvious torsion goes to theatre immediately and is not delayed by ultrasound. TRIAL REGISTRATION: PROSPERO: CRD42023412619.

Humans

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

Premature closure underlies bias in medical diagnosis in students: A randomised controlled experiment.

OBJECTIVE: The purpose of the study reported in this article was to shed light on the cognitive mechanism mediating between biasing information and diagnostic error. The literature suggests at least two different hypotheses: premature closure leading biased participants to spend less time on diagnosis or increased competition between diagnostic hypotheses. The latter hypothesis predicts that biased participants would spend more time reaching a diagnosis. METHOD: Using the salient distracting findings (SDF) experimental paradigm, we biased 58 fourth-year medical students while diagnosing 12 clinical vignettes in a within-group incomplete block design under three conditions: cases presented without SDF, with SDF at the beginning and with SDF at the end. For each of these conditions, diagnostic accuracy, the number of SDF-related mistakes and time per word needed to process the case were recorded. The data were analysed using linear mixed modelling. Estimated marginal mean scores were reported. RESULTS: Participants confronted with salient distracting features (SDFs) at the beginning of a clinical case demonstrated significantly lower diagnostic accuracy (mean 0.11) compared with the No-SDF condition (0.27), representing a 61% reduction (F2,693&#x2009;=&#x2009;11.995, p&#x2009;<&#x2009;0.001), and made more SDF-related mistakes (F2, 693&#x2009;=&#x2009;16.395, p&#x2009;<&#x2009;0.001). When SDFs were presented at the end of the case, diagnostic accuracy was also reduced (mean 0.17; 36% reduction), but processing time did not differ from the No-SDF condition. Only early presentation of SDFs was associated with reduced processing time per word (F2,636&#x2009;=&#x2009;4.799, p&#x2009;<&#x2009;0.01), consistent with premature closure. CONCLUSION: These findings demonstrate that biasing information increases diagnostic error in medical students and that only early bias is associated with reduced information processing. The data do not support the competition hypothesis for early bias, as processing time did not increase under biasing conditions. Premature closure can therefore be directly observed rather than inferred, inviting further research.

Humans