Search PubMedSearch

SEARCH · Search PubMed

Results for “High sensitivity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,055 records · Page 2Linked to original sources

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Sex Differences in Postoperative Recovery and Mortality After High-Risk Cardiac Surgery: A Propensity Score-Matched Post Hoc Analysis of the SUSTAIN-CSX Trial.

BACKGROUND: Sex-related differences after cardiac surgery remain controversial because women often present with higher baseline risk and complexity than men. We performed a post hoc propensity score-matched analysis of the SUSTAIN-CSX (Sodium Selenite Administration in Cardiac Surgery) trial to evaluate sex differences in mortality, postoperative complications, and recovery after high-risk cardiac surgery. METHODS: Of 1394 trial participants, 1386 had complete data. Women were matched 1:1 to men using nearest-neighbor propensity score matching based on age and European System for Cardiac Operative Risk Evaluation II (EuroSCORE II), with exact matching on surgical category, yielding 327 female-male pairs. Prespecified sensitivity analyses adjusted for frailty, baseline hemoglobin, renal disease, left ventricular ejection fraction, previous myocardial infarction, preoperative medications, and baseline creatinine. RESULTS: In the primary matched analysis, 180-day survival did not differ between women and men (log-rank P=0.086; unadjusted hazard ratio, 1.80 [95% CI, 0.91-3.55]; P=0.091). In descriptive matched comparisons, women had numerically longer intensive care unit stay (median, 3 days [quartile 1, quartile 3 (Q1, Q3)=1, 6 days] versus 2 days [Q1, Q3=1, 5 days]) and hospital stay (median, 10 days [Q1, Q3=7, 18 days] versus 9 days [Q1, Q3=6, 16 days]; P=0.292), whereas major postoperative complications were similar. In adjusted sensitivity analyses accounting for the matched design and residual imbalance, female sex remained associated with longer intensive care unit stay (adjusted incidence rate ratio [IRR], 1.8 [95% CI, 1.2-2.9]; P=0.009) and hospital stay (adjusted IRR, 1.4 [95% CI, 1.0-1.9]; P=0.031). Mortality sensitivity analyses were model-dependent. CONCLUSIONS: In this propensity score-matched cohort of high-risk cardiac surgery patients, women showed a longer postoperative recovery trajectory in adjusted analyses, whereas mortality findings were sensitive to model specification and should be interpreted cautiously. REGISTRATION: URL: https://www.clinicaltrials.gov; Unique identifier: NCT02002247.

Aged

Diagnostic and prognostic value of fibroblast growth factor 23 in acute kidney injury: systematic review and meta-analysis.

Background: Acute kidney injury (AKI) is associated with high mortality and adverse outcomes. Fibroblast growth factor 23 (FGF23) has emerged as a potential biomarker for AKI; however, its diagnostic and prognostic utility remains inconsistent.Methods: We conducted a systematic review and meta-analysis of studies evaluating circulating intact FGF23 (iFGF23) or C-terminal FGF23 (cFGF23) (PROSPERO: CRD42022302659). PubMed, EMBASE, CNKI, and Wanfang databases were searched through June 9, 2026. QUADAS-2 was used for quality assessment. A random-effects bivariate model pooled sensitivity, specificity, positive/negative likelihood ratio (PLR/NLR), diagnostic odds ratio (DOR), and area under the summary receiver operating characteristic curve (SROC AUC).Results: Twenty-three studies were included: 17 diagnostic, 6 prognostic (one addressing both). For AKI diagnosis, the pooled sensitivity was 0.79 (95% CI 0.73-0.86), specificity 0.82 (95% CI 0.75-0.89), PLR 4.40 (95% CI 2.59-6.21), NLR 0.25 (95% CI 0.16-0.34), DOR 17.49 (95% CI 8.67-35.16), and SROC AUC 0.87 (95% CI 0.81-0.92). Substantial heterogeneity was observed (I2 = 67%), with iFGF23 demonstrating higher accuracy than cFGF23 (AUC 0.91 vs 0.81). For AKI mortality, pooled sensitivity was 0.77 (95% CI 0.69-0.84), specificity 0.76 (95% CI 0.70-0.82), DOR 10.89 (95% CI 6.86-17.30), and SROC AUC 0.77 (95% CI 0.70-0.83). Significant heterogeneity was noted (I2 = 86.2% for sensitivity, 80.4% for specificity). No significant publication bias was detected.Conclusions: Circulating FGF23 exhibits moderate-to-high diagnostic and moderate prognostic performance in AKI, though interpretation is limited by substantial heterogeneity. It may serve as a complementary biomarker for risk stratification, pending further validation with standardized protocols.

Humans

Accelerated Diagnostic Pathways for Suspected Acute Coronary Syndrome in Practice: A Randomized Trial of 0/1-Hour vs 0/3-Hour Troponin Testing.

BACKGROUND: For suspected acute coronary syndrome (ACS), guidelines recommend using high-sensitivity troponins (hs-cTn) in accelerated diagnostic pathways (ADPs) with 0/1-hour recommended over 0/3-hour ADP. However, implementation of these ADPs, with universal use of hs-cTns, has not been directly compared in randomized trials OBJECTIVES: This study sought to compare the efficiency and safety of the European Society of Cardiology (ESC) 0/1-hour and a 0/3-hour ADP when implemented in real-world clinical practice. METHODS: This pragmatic, randomized, noninferiority implementation trial compared the safety and efficiency of clinician decision making using these 2 pathways. To prevent incorporation bias, an independent hs-cTnI was used for formal adjudication using the fourth universal definition of myocardial infarction (MI). Efficiency was judged by the proportion of patients discharged within 4 hours. The safety endpoint was major adverse cardiac events (MACE) within 30 days (adjudicated index or representation type 1 MI, cardiovascular death, and urgent coronary revascularization) for those who were considered not to have ACS and discharged. The noninferiority margin, for absolute difference in sensitivity, between the ESC 0/1-hour and the 0/3-hour ADP was set at 3%, assessed with a 1-sided 97.5% CI. RESULTS: From December 2021 to July 2024, of 13,983 screened 3,543 individual patients with suspected ACS were recruited and consented from 2 major emergency departments in North-West England, with 100% follow-up achieved for all representations to any national hospital. The median age was 60 years (IQR: 49.5-70.5 years), 53% were men, 6.7%, and 7.6% had adjudicated index type 1 MI and MACE within 30 days, respectively. The turnaround time from sample to result for central laboratory hs-cTnT was 81 minutes (IQR: 69-101 minutes). The proportion of patients discharged within 4 hours was relatively low and did not differ substantially (21.8% vs 19.2%, P = 0.07). In addition, the 0/1-hour pathway was noninferior for safety, in patients discharged, compared with the 0/3-hour pathway, absolute difference in sensitivity was +4.2% (1-sided 97.5% CI: -2.5) in favor of the 0/1-hour pathway. The calculated sensitivities were 93.7% (95% CI: 88.4%-97.1%) vs 89.5% (95% CI: 82.7%-94.3%), respectively. CONCLUSIONS: Implementation of the ESC 0/1-hour pathway failed to discharge significantly more patients within 4 hours of presentation compared with the 0/3-hour ADP. In addition, The ESC 0/1-hour was noninferior to the 0/3-hour hs-cTn pathway for safety of discharge, although safety for both pathways was less than that imputed by observational studies. This trial demonstrates that perceived benefits to emergency department efficiency of a reduced sampling interval are mitigated by central laboratory turnaround times as well as system constraints. (Pragmatic Randomised Trial of the ESC 0/​1 Versus 0/​3 Hour Troponin Pathway [MACROS2]; NCT05322395).

Acute Coronary Syndrome

Ultra-high-frequency ECG quantifies residual electrical dyssynchrony during left bundle branch area pacing in patients with wide QRS: a paired within-patient study.

BACKGROUND: Left bundle branch area pacing (LBBAP) may restore a more physiological pattern of ventricular activation in patients with conduction delay; however, QRS narrowing alone may incompletely characterize electrical resynchronization. Ultra-high-frequency ECG (UHF-ECG) provides quantitative markers of ventricular activation timing and dyssynchrony. OBJECTIVE: To quantify paired OFF-to-ON changes in conventional ECG and UHF-ECG metrics during LBBAP in patients with baseline wide QRS and to assess the relationship between paced R-wave peak time (RWPT) and residual UHF-ECG dyssynchrony. METHODS: In this prospective single-center paired study, 21 patients with bradycardia and baseline wide QRS underwent standard ECG and UHF-ECG assessment during intrinsic rhythm (pacing OFF) and during LBBAP (pacing ON). Endpoints included QRS duration, signed VED16, absolute VED16 (|VED16|), mean ventricular delay (meanVD), and a clinically interpretable distance-to-normal metric defined as dist&#xa0;=&#xa0;max(|VED16|-20, 0). Paired changes were summarized as medians with bootstrap 95% confidence intervals and tested using the Wilcoxon signed-rank test. Associations between paced RWPT and residual dyssynchrony during pacing were evaluated using Pearson and Spearman correlation coefficients. RESULTS: LBBAP significantly narrowed QRS duration from 136.8 [130.2-153.6] ms during intrinsic rhythm to 116.0 [107.8-125.6] ms during pacing (median &#x394; -21.0&#xa0;ms; 95% CI -33.9 to -18.6; p&#xa0;<&#xa0;0.001). Signed VED16 did not change significantly (median &#x394; 0.4&#xa0;ms; p&#xa0;=&#xa0;1.000), consistent with the mixed conduction-phenotype composition of the cohort. In contrast, severity-oriented UHF-ECG endpoints improved: |VED16| decreased numerically (median &#x394; -5.2&#xa0;ms; p&#xa0;=&#xa0;0.070), whereas dist decreased significantly (median &#x394; -0.7&#xa0;ms; 95% CI -14.4 to 0.0; p&#xa0;=&#xa0;0.015). The proportion of patients within the normal dyssynchrony band (|VED16|&#xa0;&#x2264;&#xa0;20&#xa0;ms) increased from 7/21 (33.3%) to 12/21 (57.1%). Median paced RWPT was 66.6 [58.6-74.6] ms, and shorter RWPT correlated with lower residual |VED16| during pacing (Pearson r&#xa0;=&#xa0;-0.45, p&#xa0;=&#xa0;0.038). CONCLUSIONS: In patients with baseline wide QRS, LBBAP produces marked QRS narrowing, whereas UHF-ECG provides complementary quantification of residual electrical dyssynchrony. Severity-oriented UHF-ECG endpoints, particularly a distance-to-normal metric, may offer an interpretable mechanistic framework beyond conventional ECG alone. Shorter paced RWPT was associated with lower residual dyssynchrony during pacing, supporting physiological coherence between procedural and high-resolution electrocardiographic markers.

Humans

Engineering bubble structures as Cas12a activators for highly sensitive monitoring of WRN helicase function.

The Werner syndrome helicase (WRN) is a critical synthetic lethal target in microsatellite instability cancers, essential for resolving complex genomic structures like replication bubbles and R-loops. However, strategies to simultaneously discriminate WRN activity on DNA versus DNA-RNA substrates in living cells are lacking. Here, we developed a structure-specific CRISPR/Cas12a biosensing strategy to visualize WRN functional activity by engineering bubble-structure probes. These probes were rationally designed to structurally mimic DNA replication bubbles and R-loop associated DNA-RNA hybrids. Upon specific unwinding by WRN, the probes release a sequestered activator strand that triggers Cas12a trans-cleavage, effectively converting the unwinding event into an amplified fluorescent signal. This assay achieves low picomolar sensitivity (LODs: 5.6-6.0 pM) and exceptional selectivity against homologous RecQ helicases. Uniquely, this strategy enables the parallel quantification of WRN activity on both substrate types, providing insights into distinct WRN-mediated pathways for resolving genomic stress. We further demonstrated the strategy's utility by visualizing endogenous WRN dynamics in living cells and profiling the efficacy of small-molecule inhibitors. This work offers a powerful molecular toolkit for dissecting WRN biology and facilitating high-throughput drug screening in targeted cancer therapy.

Werner Syndrome Helicase

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

Enhancing Hemoglobin Bart's hydrops fetalis syndrome prevention: a single-tube multiplex real-time PCR assay for the comprehensive detection of four significant &#x3b1;0-thalassemia deletions (--SEA, --THAI, --CR, and --SA) found in Thailand.

BACKGROUND: Hemoglobin (Hb) Bart's hydrops fetalis is a major public health concern in Southeast Asia, particularly in Thailand. Current screening strategies target the two most common &#x3b1;0 -thalassemia deletions (--SEA and --THAI). METHOD: In this study, we developed a single-tube multiplex real-time PCR assay for the simultaneous detection of four clinically relevant &#x3b1;0-thalassemia deletions (--SEA, --THAI, --CR, and --SA). The assay was validated using 538 clinical samples with diverse thalassemia genotypes and compared against conventional gap-PCR as the reference method. Analytical performance, including sensitivity, specificity, and limit of detection (LOD), was evaluated. In addition, clinical utility was assessed in 22 prenatal diagnosis cases at risk of Hb Bart's hydrops fetalis. RESULTS: The study cohort demonstrated substantial genetic heterogeneity, comprising 43 distinct genotypes. The developed assay achieved 100% sensitivity and specificity for all targeted deletions, with complete concordance with gap-PCR results. No cross-reactivity was observed with &#x3b1;+-thalassemia. The assay demonstrated a high analytical sensitivity with a LOD of 9.76&#x2009;&#xd7;&#x2009;10-3&#x2009;ng per reaction. Whereas in prenatal diagnosis, all 22 fetal genotypes were accurately identified, including five cases of homozygous --SEA and one rare compound heterozygous --SEA/--CR fetus. CONCLUSIONS: This study presents a rapid, accurate, and cost-effective multiplex real-time PCR assay capable of detecting both common and rare &#x3b1;0-thalassemia deletions in a single reaction. The assay demonstrates strong potential for implementation in routine clinical laboratories and large-scale population screening, contributing to improved prevention and control of severe thalassemia syndromes in high-prevalence regions.

Humans

Ecotoxicological responses of aquatic macrophytes to 2,4-D: A global synthesis of species sensitivity and ecological risk.

The widespread use of 2,4-dichlorophenoxyacetic acid (2,4-D) has raised concern about its persistence, mobility, and effects on non-target aquatic vegetation in freshwater ecosystems. Here, we provide a global synthesis of the ecotoxicological responses of aquatic macrophytes to 2,4-D based on a PRISMA-guided systematic review of 86 peer-reviewed studies published between 1947 and 2025. A consistent gradient of species-specific sensitivity was observed across macrophyte growth forms. The submerged species Myriophyllum spicatum showed high susceptibility, with EC&#x2085;&#x2080; values of 0.04-0.182 mg/L and marked growth inhibition at low concentrations, whereas floating species such as Lemna minor and Pontederia crassipes were more tolerant, requiring higher concentrations (7.08 to >100 and 8.1 mg/L, respectively) to produce comparable effects. Importantly, this sensitivity ranking was consistent across laboratory and field experimental settings. These interspecific differences likely reflect variation in herbicide uptake, translocation, and detoxification capacity associated with growth form. The overlap between EC&#x2085;&#x2080; values for M. spicatum and regulatory thresholds for 2,4-D in surface waters suggests that current limits may be insufficient to protect sensitive submerged macrophyte communities. Regarding remediation, L. minor and Salvinia natans emerged as the most promising candidates for phytoremediation, while P. crassipes showed limited capacity to reduce herbicide concentrations in water. Despite advances, no study directly compared oxidative stress biomarkers between submerged and floating species, representing a critical gap in understanding the biochemical basis of the sensitivity gradient. Overall, this synthesis highlights the need to account for taxon-dependent sensitivity when evaluating the ecological risks of 2,4-D and provides a basis for improving regulatory frameworks and management of herbicide contamination in freshwater ecosystems.

2,4-Dichlorophenoxyacetic Acid

Enhanced fracture detection on radiographs with AI assistance for clinicians: a systematic review and meta-analysis.

BACKGROUND: Emergency radiographic interpretation for fractures is prone to missed or misdiagnoses. Artificial intelligence (AI) is expected to become a powerful tool to assist clinicians in fracture detection. PURPOSE: A systematic review and meta-analysis was performed to assess whether AI improves clinicians' ability to detect fractures on radiographs. MATERIALS AND METHODS: A literature search was conducted in PubMed, Web of Science, and Cochrane Library for studies published between January 1, 2010, and October 10, 2025. A meta-analysis of diagnostic accuracy studies was performed using a Summary Receiver Operating Characteristic (SROC) curve. The quality of included studies was assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool. Subgroup analysis and meta-regression were conducted to explore potential sources of heterogeneity. RESULTS: A total of 26 studies were included . The pooled sensitivity of clinicians increased from 77% (95% CI: 72-81) to 87% (95% CI: 83-90) with AI assistance, while the pooled specificity improved from 88% (95% CI: 85-90) to 92% (95% CI: 89-94). The corresponding AUC values were 0.90 (95% CI: 0.87-0.92) before and 0.95 (95% CI: 0.93-0.97) after AI assistance. Eight studies were rated as high risk of bias. Subgroup analysis and meta-regression identified potential sources of heterogeneity, including fracture location, AI model type, high risk of bias, and reference standards. CONCLUSION: AI assistance significantly improves clinicians' diagnostic performance in detecting fractures on radiographs for extremity and trunk fractures.

Humans

A mechanism-guided framework for prioritizing membrane-interaction anti-Vibrio peptides from peptidomics data.

A mechanism-guided framework for prioritizing membrane-interaction antimicrobial peptide candidates from proteomics-derived peptide mixtures is presented. The framework integrates conservative machine-learning-based antimicrobial peptide (AMP) screening with a literature-derived membrane-interaction plausibility (MAP) assessment and a data-driven membrane-interaction ranking function (AIPx), followed by structural visualization for interpretability. MAP encodes physicochemical characteristics commonly associated with peptide-membrane interaction and provides a graded plausibility assessment. Building upon this physicochemically interpretable framework, AIPx ranks peptides using feature weights calibrated from experimentally characterized anti-Vibrio peptides, where minimum inhibitory concentration (MIC) values are used as a coarse-grained ranking reference rather than a direct prediction target. In a peptidomics-based peptide fractionation study targeting Vibrio spp., AIPx exhibited a consistent relationship with experimentally observed antibacterial activity. Distributional analysis revealed that peptide fractions exhibiting high anti-Vibrio activity are characterized by enrichment of high-ranking peptides rather than by AMP abundance alone. By structuring AMP identification and prioritization as sequential stages, the MAP&#xa0;+&#xa0;AIPx framework enables interpretable and experimentally actionable candidate selection by reducing biologically implausible candidates. The framework facilitates species-oriented prioritization of AMP candidates, addressing a key challenge in antimicrobial peptide discovery where activity may depend on target-specific membrane characteristics. Moreover, the approach is extensible through species-specific calibration and supports interpretable, mechanism-informed prioritization in antimicrobial peptide discovery.

Proteomics

Tigecycline-resistant Staphylococcus in waiting pens of a pig slaughterhouse: genomic insights into a food safety alert.

BACKGROUND: The waiting pens of slaughterhouses represent a critical control point in the 'farm-to-fork' continuum, yet their role in the emergence and dissemination of antimicrobial resistance remains understudied. This study investigated tigecycline-resistant Staphylococcus (TRS) in these high-risk zones to assess their prevalence, resistance mechanisms, and transmission dynamics. METHODS: 400 samples were collected from the waiting pens of a pig slaughterhouse in Guangzhou, China. Antimicrobial susceptibility testing, whole-genome sequencing, phylogenetic analysis, and molecular cloning were employed to characterize resistance mechanisms and transmission patterns. RESULTS: 78 TRS strains were isolated and classified into three species, including S. borealis, S. ureilyticus, and S. pasteuri. These isolates exhibited multidrug-resistant phenotypes and carried new mutations in rpsJ and tet(M), which were functionally confirmed to reduce tigecycline susceptibility. Phylogenetic evidence demonstrated clonal transmission between pig farms and the slaughterhouse. The tet(M) gene was located within Staphylococcal cassette chromosome mec elements mediated by IS257, while tet(L) was carried by plasmids formed through IS256/IS257-mediated recombination. CONCLUSIONS: Waiting pens serve as crucial reservoirs for the amplification and dissemination of antimicrobial resistance. Our findings underscore the urgent need for enhanced biosecurity measures, improved waste management, and routine molecular surveillance in these high-risk zones to mitigate the spread of resistance along the food production chain.

Animals

Can ChatGPT Replace Human Clinical Coders? A Comparative Study in Otology Billing.

OBJECTIVE: Evaluate the utility of the large language model (LLM), ChatGPT, for the analysis of operative notes and the generation of Current Procedural Terminology (CPT) codes in comparison to human clinical coders. STUDY DESIGN: CPT billing codes assigned by ChatGPT were compared to existing billing data. Otology practice within a tertiary academic center. METHODS: About 191 operative notes from a single surgeon (9/2022-10/2023) were analyzed. ChatGPT-3.5 and 4 models were prompted for CPT codes based on operative notes. Assessment included determining exact and partial match rates, sensitivity and specificity for targeted procedures, and work Relative Value Units (wRVU) differences between ChatGPT-generated and human-assigned codes. RESULTS: ChatGPT-3.5 achieved exact matches in 22% of cases and partial matches in 32%, while ChatGPT-4 achieved 14% exact and 33% partial matches. When cochlear implantation (CI) was excluded, performance dropped significantly. For CI, ChatGPT-3.5 demonstrated a sensitivity of 94% and specificity of 90%, while ChatGPT-4 showed a sensitivity of 96% and specificity of 92%. In contrast, performance on cartilage grafting was poor, with sensitivities of 4.2% for ChatGPT-3.5 and 0% for ChatGPT-4. ChatGPT-3.5 and 4 showed moderate CPT code matching accuracy among themselves, with slight agreement to human coders. Both models tended to underbill for wRVUs compared to human coders, with significant differences in the values generated. CONCLUSION: This study assessed ChatGPT's effectiveness in automating CPT code assignment for otologic surgeries. While the models achieved high sensitivity values for assigning codes related to cochlear implantation, both models struggled with complex cases, failed to apply modifiers, and often assigned fewer wRVUs. The findings highlight ChatGPT's potential in medical billing but indicate a need for further refinement.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

Longitudinal Prediction of Retinal Sensitivity Based on Disease Progression Quantified From Optical Coherence Tomography in Geographic Atrophy.

PURPOSE: The purpose of this study was to analyze the association between disease progression of geographic atrophy (GA) from optical coherence tomography (OCT) with retinal sensitivity (RS) in microperimetry (MP) over a 2-year follow-up period. METHODS: This is a longitudinal analysis of the OAKS Phase-III clinical trial. Both study and fellow eyes with GA that underwent imaging with the Spectralis OCT and consecutive MP examination were eligible. Pointwise quantification of ellipsoid zone (EZ) thickness, EZ and retinal pigment epithelium (RPE) loss from OCT volumes was correlated with localized RS. A longitudinal predictive model using a Markov Chain framework was implemented to predict RS change over time based on OCT biomarkers. The modeling of morphological and functional progression was based on the fellow-eye cohort. RESULTS: A total of 39,681 MP points from 406 patients were analyzed. In the fellow eye cohort, baseline (BSL) EZ thickness was positively associated with RS (0.3 decibel [dB]/&#xb5;m, P < 0.001). Decrease in EZ thickness between visits during follow-up was significantly associated with decrease in RS (0.1 dB / 1&#xa0;&#xb5;m change). RS was significantly lower in MP points within EZ loss during follow-up compared with MP points within the retina with measurable EZ (P < 0.001). The largest functional decline was observed within RPE loss, also associated with the highest probability of absolute scotoma (P < 0.001). Morphological progression to EZ and RPE loss was influenced by EZ thickness and the morphology of adjacent MP points (P < 0.001). CONCLUSIONS: Two exploratory endpoints were developed, namely quantification of EZ thickness and loss, and localized RS within high-risk OCT areas. RS decline during follow-up is associated with automatically quantified disease progression in OCT.

Humans

Online Social Anxiety in the Digital Age: Transitions, Predictors, and Mental Health Associations in Emerging Adulthood.

BACKGROUND: Online social anxiety (OSA), a multidimensional form of social evaluative anxiety in online social contexts, disproportionately affects emerging adults who constitute the largest active group of media users and face heightened psychological sensitivity due to growing pressures and immature sociocognitive regulation during the transition to adulthood. However, its heterogeneity, transitions, and longitudinal associations with mental health outcomes remain underexplored. METHODS: This study utilized data from two waves of a three-wave longitudinal survey, with 849 Chinese participants (Meanage = 21.6 years; 50.4 percent female) assessed at 4-month intervals. Individuals were classified using latent profile analysis and the stability and changes of profiles were assessed via latent transition analysis (LTA). Multinomial logistic regressions were conducted separately at baseline and follow-up to identify correlates of profile membership. Predictors of profile transitions were examined using manual three-step LTA models, and associations between latent transition patterns and follow-up mental health outcomes were examined using BCH-LTA distal outcome analyses controlling for the corresponding baseline symptom level. RESULTS: Four profiles of OSA were identified: low, privacy-sensitive, moderate-high, and high OSA. Extreme profiles (low/high OSA) showed high stability (80.4 percent and 79.1 percent), while privacy-sensitive OSA exhibited the lowest stability (55.1 percent). Profile memberships were influenced by social-cognitive biases and digital interaction, particularly fear of negative evaluation and online interpersonal trust, whereas profile transitions were mainly associated with anxiety. Transitions toward less severe OSA profiles were generally associated with better subsequent mental health, whereas transitions toward more severe profiles corresponded to poorer outcomes, particularly for offline social anxiety. CONCLUSION: OSA was heterogeneous in its manifestation, severity and transitions. Personalized and early interventions targeting profile-specific vulnerabilities are critical to prevent the worsening of OSA and mitigate its psychological burden.

Humans

Assessment of atypical glandular cell interpretation in Pap tests using the Hologic Genius Digital Diagnostics System.

Atypical glandular cells (AGC) are a diagnostic challenge. The aim of this study was to evaluate the efficacy and diagnostic performance of AGC detection on the Hologic Genius Digital Diagnostics System (HGDDS). A retrospective analysis of 451 ThinPrep Pap cases was conducted, including 207 cases of AGC, 27 cases of high-grade squamous intraepithelial lesion (HSIL), 25 cases of low-grade squamous intraepithelial lesion (LSIL), and 192 benign cases. All AGC cases had follow-up histologic diagnoses, with 66 cases subsequently diagnosed as adenocarcinoma. The slides were randomized, scanned, and analyzed by the HGDDS. Patient age and HPV test results were provided to reviewers, an experienced cytologist, who screened the cases, followed by two cytopathologists who independently examined the cases on the HGDDS. Diagnostic concordance between the two cytopathologists indicated strong agreement (&#x3ba; = 0.829). Sensitivity of AGC on Papanicolaou (Pap) tests for adenocarcinoma detection on HGDDS was 98.5% and 95.5%, respectively, comparable to the original ThinPrep interpretation (OTPI). Specificity for adenocarcinoma detection was significantly higher (84.6% and 85.6%) with the HGDDS than 27.7% with OTPI. Overall, the diagnostic performance for AGC/HSIL interpretation to detect CIN2/3/adenocarcinoma appeared to have improved with HGDDS compared with OTPI, particularly for specificity and positive predictive value (PPV). This is the first study evaluating AGC diagnosis using the HGDDS. The findings demonstrate that the sensitivity of adenocarcinoma detection as AGC on HGDDS is comparable to the ThinPrep Imaging System, but the specificity and PPV are improved. This suggests the potential of artificial intelligence to augment the performance of cervical cancer screening.

Humans

Diagnostic accuracy of bronchoalveolar lavage fluid-based testing for pulmonary cryptococcosis: A systematic review and meta-analysis.

BACKGROUND: Pulmonary cryptococcosis(PC) presents diagnostic challenges because of its non-specific clinical and radiological manifestations. Bronchoalveolar lavage fluid (BALF)-based testing, which includes latex agglutination (LA) and lateral flow assay (LFA), offers a minimally invasive diagnostic method, yet its pooled diagnostic accuracy remains unclear. METHODS: We systematically searched PubMed, Embase, Cochrane Library, and Scopus from inception to May 2026. Studies evaluating BALF-based testing for PC with extractable 2 &#xd7; 2 data were included. The methodological quality of relevant studies was assessed by the QUADAS-2 tool. Pooled sensitivity, specificity, likelihood ratios, and diagnostic odds ratio (DOR) were estimated using a bivariate random-effects model. Subgroup analyses were performed by testing method and reference standard type. Heterogeneity was evaluated through paired forest plots, HSROC visualization, and exploratory bivariate meta-regression. RESULTS: The pooled sensitivity was 0.87 (95% CI: 0.81-0.91), and the specificity was 0.99 (95% CI: 0.982 - 0.995). The pooled positive likelihood ratio (PLR) was 88.00 (95% CI: 47.39 - 163.42), the negative likelihood ratio (NLR) was 0.13 (95% CI: 0.09 -0.20), and the DOR was 658.50 (95% CI: 285.36-1519.55). No significant threshold effect or publication bias was detected. Exploratory meta-regression suggested a possible assay-method effect in the joint model (P = 0.03), mainly driven by specificity (P = 0.01). CONCLUSIONS: The study demonstrates the high accuracy of CrAg in BALF for the diagnosis of pulmonary cryptococcosis, supporting its role as an important adjunctive diagnostic tool, particularly when tissue biopsy is not feasible or rapid results are needed. Larger prospective studies with standardized protocols are needed to validate these estimates.

Humans