Search PubMedSearch

SEARCH · Search PubMed

Results for “Reference Standards”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Xpert MTB/RIF Ultra assay for tuberculosis disease and rifampicin resistance in children.

BACKGROUND: In 2023, an estimated 1.3 million children (aged 0-14 years) became ill with tuberculosis, and 166,000 children (aged 0-15 years) died from the disease. Xpert MTB/RIF Ultra (Xpert Ultra) is a molecular World Health Organization (WHO)-recommended rapid diagnostic test that detects Mycobacterium tuberculosis complex and rifampicin resistance. This is an update of a Cochrane review first published in 2020 and last updated in 2022. Parts of the current update informed the 2024 WHO updated guidance for the diagnosis of tuberculosis. OBJECTIVES: To assess the diagnostic accuracy of Xpert Ultra for detecting pulmonary tuberculosis, tuberculous meningitis, lymph node tuberculosis, and rifampicin resistance in children (aged 0-9 years) with presumed tuberculosis. SEARCH METHODS: We searched the Cochrane Central Register of Controlled Trials (CENTRAL), MEDLINE, Embase, three other databases, and three trial registers without language restrictions to 6 October 2023. SELECTION CRITERIA: For study design, we included cross-sectional and cohort studies and randomized trials that evaluated Xpert Ultra in HIV-positive and HIV-negative children aged birth to nine years. Regarding specimen type, we included studies evaluating sputum, gastric, stool, or nasopharyngeal specimens (pulmonary tuberculosis); cerebrospinal fluid (tuberculous meningitis); and fine needle aspirate or surgical biopsy tissue (lymph node tuberculosis). Reference standards for detection of tuberculosis were microbiological reference standard (MRS; including culture) or composite reference standard (CRS); for stool, we considered Xpert Ultra in sputum or gastric aspirates in addition to culture. Reference standards for detection of rifampicin resistance in sputum were phenotypic drug susceptibility testing or targeted or whole genome sequencing. DATA COLLECTION AND ANALYSIS: Two review authors independently extracted data and assessed methodological quality using the tailored QUADAS-2 tool, judging risk of bias separately for each target condition and sample type. We conducted separate meta-analyses for detection of pulmonary tuberculosis, tuberculous meningitis, lymph node tuberculosis, and rifampicin resistance. We used a bivariate model to estimate summary sensitivity and specificity with 95% confidence intervals (CIs). We assessed certainty of evidence using the GRADE approach. MAIN RESULTS: This update included 23 studies (including 9 new studies since the previous review) that evaluated detection of pulmonary tuberculosis (21 studies, 9223 children), tuberculous meningitis (3 studies, 215 children), lymph node tuberculosis (2 studies, 58 children), and rifampicin resistance (3 studies, 130 children). Seventeen studies (74%) took place in countries with a high tuberculosis burden. Overall, risk of bias and applicability concerns were low. Detection of pulmonary tuberculosis (microbiological reference standard) Sputum (11 studies) Xpert Ultra summary sensitivity was 75.3% (95% CI 68.9% to 80.8%; 345 children; moderate-certainty evidence), and specificity was 95.9% (95% CI 92.3% to 97.9%; 2645 children; high-certainty evidence). Gastric aspirate (12 studies) Xpert Ultra summary sensitivity was 69.6% (95% CI 60.3% to 77.6%; 167 children; moderate-certainty evidence), and specificity was 91.0% (95% CI 82.5% to 95.6%; 1792 children; moderate-certainty evidence). Stool (10 studies) Xpert Ultra summary sensitivity was 68.0% (95% CI 50.3% to 81.7%; 255 children; moderate-certainty evidence), and specificity was 98.2% (95% CI 96.3% to 99.1%; 2630 children; high-certainty evidence). Nasopharyngeal aspirate (6 studies) Xpert Ultra summary sensitivity was 46.2% (95% CI 34.9% to 57.9%; 94 children; moderate-certainty evidence), and specificity was 97.5% (95% CI 95.1% to 98.7%; 1259 children; high-certainty evidence). Xpert Ultra sensitivity was lower against CRS than against MRS for all specimen types, while the specificities were similar. Extrapulmonary tuberculosis Meta-analysis was not possible for lymph node tuberculosis and tuberculous meningitis due to low study numbers. Interpretation of results For a population of 1000 children, where 100 have pulmonary tuberculosis: In sputum: • 112 would be Xpert Ultra positive, of whom 75 would have pulmonary tuberculosis (true positives) and 37 would not (false positives). • 888 would be Xpert Ultra negative, of whom 863 would not have pulmonary tuberculosis (true negatives) and 25 would have pulmonary tuberculosis (false negatives). In gastric aspirate: • 151 would be Xpert Ultra positive, of whom 70 would have pulmonary tuberculosis (true positives) and 81 would not (false positives). • 849 would be Xpert Ultra negative, of whom 819 would not have pulmonary tuberculosis (true negatives) and 30 would have pulmonary tuberculosis (false negatives). In stool: • 85 would be Xpert Ultra positive, of whom 68 would have pulmonary tuberculosis (true positives) and 17 would not (false positives). • 915 would be Xpert Ultra negative, of whom 883 would not have pulmonary tuberculosis (true negatives) and 32 would have pulmonary tuberculosis (false negatives). In nasopharyngeal aspirate: • 68 would be Xpert Ultra positive, of whom 46 would have pulmonary tuberculosis (true positives) and 22 would not (false positives). • 932 would be Xpert Ultra negative, of whom 878 would not have pulmonary tuberculosis (true negatives), and 54 would have pulmonary tuberculosis (false negatives). Detection of rifampicin resistance Three studies with 76 children evaluated detection of rifampicin resistance (sputum only); two of these studies reported no cases and one reported rifampicin resistance in two children. AUTHORS' CONCLUSIONS: Xpert Ultra sensitivity was moderate in sputum, gastric aspirate, and stool specimens. Nasopharyngeal aspirate had the lowest sensitivity. Xpert Ultra specificity was high against both MRS and CRS. We were unable to determine the accuracy of Xpert Ultra for detecting tuberculous meningitis, lymph node tuberculosis, and rifampicin resistance due to a paucity of data. FUNDING: This update was funded through WHO. REGISTRATION: The protocol for this review was originally published through Cochrane in 2019. The protocol for this update was a generic protocol that consolidated previously published Cochrane protocols of Xpert Ultra for tuberculosis detection and can be accessed at https://osf.io/26wg7/. Protocol (2019) DOI: 10.1002/14651858.CD013359 Original review (2020) DOI: 10.1002/14651858.CD013359.pub2 Review update (2022) DOI: 10.1002/14651858.CD013359.pub3.

Adolescent

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

Standardisation in the Analytical Characterization of Adeno-Associated Virus (AAV) Vectors.

Adeno-associated virus (AAV) has become a leading vector for in vivo gene therapy, with eight products currently holding marketing authorization. As the field rapidly evolves, the need for robust analytical methods to characterize critical quality attributes (CQAs)-including capsid titer, genome titer, capsid content (empty/full ratio), identity, and purity-continues to grow. Reference Standard Materials (RSMs) play a pivotal role by providing well-characterized, standardized AAV batches that serve as universal benchmarks. RSMs facilitate the validation of emerging analytical technologies, ensure the accuracy and reproducibility of routine assays, and enable inter-laboratory comparability. However, developing universal AAV RSMs is fundamentally constrained by the complex biology, diversity of serotypes, vector genomes, and engineered capsid variants, necessitating serotype-specific and application-specific standards. Recent advances, including the release of pharmacopeial AAV8 reference standards characterized by multiple orthogonal methods, represent meaningful progress toward measurement harmonisation. This review addresses the critical need for RSMs in AAV gene therapy, evaluates the currently available pharmacopeial and commercial standards, and outlines practical strategies for in-house RSM development. Establishing robust, serotype-specific AAV RSMs and harmonised standard operating protocols (SOPs) are essential for advancing AAV gene therapy and ensuring accuracy, reproducibility, and safety across research, development, and clinical manufacturing.

Dependovirus

Diagnostic performance of the Sanity 2.0 assay to detect resistance to rifampicin, isoniazid, and fluoroquinolones in tuberculosis.

UNLABELLED: Effective tuberculosis (TB) management relies on prompt diagnosis of Mycobacterium tuberculosis complex (MTBC) and associated drug resistance. The Sanity 2.0 assay is a high-resolution melting assay designed for direct respiratory sample testing, enabling simultaneous detection of MTBC and resistance to rifampicin (RIF), isoniazid (INH), and fluoroquinolones (FQ) in a single step. This study evaluated its diagnostic performance in two registered multicenter trials among bacteriologically confirmed TB patients. Diagnostic performance was evaluated for MTBC detection, as well as for the identification of resistance to RIF, INH, and FQ, using phenotypic drug susceptibility testing, whole-genome sequencing, and a composite reference standard. Agreement analyses were conducted between the Sanity 2.0 assay and Xpert MTB/RIF and Xpert MTB/XDR. Among 611 patients, the Sanity 2.0 assay detected MTBC in 563 patients, exhibiting a sensitivity of 92.1% (95% CI: 89.7-94.0). For detecting resistance to RIF, INH, and FQ, sensitivities exceeded 90%, with specificities of 95.8% (95% CI: 88.5-98.6), 100.0% (95% CI: 96.4-100.0), and 97.8% (95% CI: 93.8-99.3) against the composite reference standard, respectively. The agreement with Xpert MTB/RIF for RIF detection was 98.6% (95% CI: 96.9-99.3). For INH and FQ resistance, the agreement with Xpert MTB/XDR was 92.0% (95% CI: 88.5-94.5) and 94.3% (95% CI: 91.2-96.3), respectively. The Sanity 2.0 assay is a rapid and user-friendly platform capable of detecting both MTBC and key drug resistance. It demonstrated good diagnostic performance and could potentially be an effective alternative to guide individualized anti-TB treatment, especially in resource-limited settings. IMPORTANCE: Rapid and accurate detection of both Mycobacterium tuberculosis complex (MTBC) and key drug resistance is critical to improving tuberculosis treatment outcomes and reducing transmission. However, current molecular diagnostic workflows often require sequential testing, which can delay the initiation of effective and individualized therapy. We evaluated the Sanity 2.0 assay, an integrated high-resolution melting test that simultaneously detects MTBC and resistance to rifampicin, isoniazid, and fluoroquinolone resistance directly from respiratory samples in about 2-3 hours. The assay demonstrated excellent performance, with MTBC detection sensitivity of 92.1% and drug resistance sensitivities exceeding 90% and specificities over 95% against a composite reference standard, as well as strong concordance with World Health Organization-endorsed molecular assays. Implementation of the Sanity 2.0 assay could streamline TB diagnostic workflows; enable rapid, single-step resistance profiling; and facilitate timely, individualized treatment-particularly in resource-limited settings where rapid and comprehensive resistance testing remains a critical unmet need.

Humans

Diagnostic accuracy of nuclear STAT6 immunohistochemistry for solitary fibrous tumour: a systematic review and meta-analysis.

Nuclear STAT6 immunohistochemistry is the diagnostic surrogate for the NAB2::STAT6 fusion of solitary fibrous tumour (SFT); its sensitivity is established, but specificity varies for unexamined reasons. This review quantified pooled accuracy and tested whether antibody clone and nuclear threshold govern specificity. PubMed, Scopus and Web of Science were searched to 29 June 2026 for studies reporting nuclear STAT6 immunohistochemistry against a reference standard (NAB2::STAT6 confirmation and/or expert consensus) in SFT and comparators, with extractable two-by-two data. Two reviewers screened, extracted data and applied QUADAS-2. A bivariate generalised linear mixed model gave summary sensitivity and specificity, and exploratory subgroup analysis and meta-regression tested antibody clone, anatomical site and reference-standard type. Twenty-three studies (1216 SFT and 4715 comparators) were included. Summary sensitivity was 98.7% (95% confidence interval 96.7-99.5) and specificity 99.1% (97.8-99.6); the diagnostic odds ratio was approximately 8656. The monoclonal YE361 subgroup (8 studies) reached specificity 99.9% (99.3-100), with one false positive among 861 comparators, versus 98.1% (96.0-99.1) for polyclonal and other antibodies. False positives concentrated in dedifferentiated liposarcoma and prostatic stromal tumours. Estimates were stable after removing studies at higher risk of bias (98.9%/99.1%) and on leave-one-out analysis; the Deeks test was non-significant (p = 0.08). Nuclear STAT6 immunohistochemistry is therefore highly sensitive and specific for SFT, and the residual specificity loss is structured and largely avoidable: the monoclonal YE361 read at a strict nuclear threshold is preferred, with MDM2 and CDK4 applied to exclude dedifferentiated liposarcoma when nuclear STAT6 is unexpectedly positive.

Humans

Evaluating culture-free targeted next-generation sequencing for diagnosing drug-resistant tuberculosis: a multicentre clinical study of two end-to-end commercial workflows.

BACKGROUND: Drug-resistant tuberculosis remains a major obstacle in ending the global tuberculosis epidemic. Deployment of molecular tools for comprehensive drug resistance profiling is imperative for successful detection and characterisation of tuberculosis drug resistance. We aimed to assess the diagnostic accuracy of a new class of molecular diagnostics for drug-resistant tuberculosis. METHODS: We conducted a prospective, cross-sectional, multicentre clinical evaluation of the performance of two targeted next-generation sequencing (tNGS) assays for drug-resistant tuberculosis at reference laboratories in three countries (Georgia, India, and South Africa) to assess diagnostic accuracy and index test failure rates. Eligible participants were aged 18 years or older, with molecularly confirmed pulmonary tuberculosis, and at risk for rifampicin-resistant tuberculosis. Sensitivity and specificity for both tNGS index tests (GenoScreen Deeplex Myc-TB and Oxford Nanopore Technologies [ONT] Tuberculosis Drug Resistance Test) were calculated for rifampicin, isoniazid, fluoroquinolones (moxifloxacin, levofloxacin), second line-injectables (amikacin, kanamycin, capreomycin), pyrazinamide, bedaquiline, linezolid, clofazimine, ethambutol, and streptomycin against a composite reference standard of phenotypic drug susceptibility testing and whole-genome sequencing. FINDINGS: Between April 1, 2021, and June 30, 2022, 832 individuals were invited to participate in the study, of whom 720 were included in the final analysis (212, 376, and 132 participants in Georgia, India, and South Africa, respectively). Of 720 clinical sediment samples evaluated, 658 (91%) and 684 (95%) produced complete or partial results on the GenoScreen and ONT tNGS workflows, respectively, with 593 (96%) and 603 (98%) of 616 smear-positive samples producing tNGS sequence data. Both workflows had sensitivities and specificities of more than 95% for rifampicin and isoniazid, and high accuracy for fluoroquinolones (sensitivity approximately ≥94%) and second line-injectables (sensitivity 80%) compared with the composite reference standard. Importantly, these assays also detected mutations associated with resistance to critical new and repurposed drugs (bedaquiline, linezolid) not currently detectable by any other WHO-recommended rapid diagnostics on the market. We note that the current format of assays have low sensitivity (≤50%) for linezolid and more work on mutations associated with drug resistance is needed. INTERPRETATION: This multicentre evaluation demonstrates that culture-free tNGS can provide accurate sequencing results for detection and characterisation of drug resistance from Mycobacterium tuberculosis clinical sediment samples for timely, comprehensive profiling of drug-resistant tuberculosis. FUNDING: Unitaid.

Humans

Performance of AI-Based Screening Tools for Obstructive Sleep Apnea Across Apnea-Hypopnea Index Thresholds: Systematic Review and Meta-Analysis.

BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain. OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG-derived inputs. METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method. CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.

Humans

Which radiographic plane should be used to quantify the distal tibia angle on weightbearing CT images?

BACKGROUND: Precise quantification of distal tibial alignment is essential for planning corrective osteotomies and ankle joint replacement surgery. The lateral distal tibial angle (LDTA) is the principal radiographic parameter used for this purpose. While LDTA is increasingly measured on weightbearing cone-beam CT (WBCT) using two-dimensional coronal slices, the optimal measurement plane remains unclear. METHODS: In this retrospective comparative study, full-leg WBCT scans of patients scheduled for supramalleolar osteotomy (n&#x202f;=&#x202f;20; mean age 47&#x202f;&#xb1;&#x202f;12.8 years) were analyzed. LDTA was measured on three coronal planes of the distal tibial plafond (anterior edge, mid-dome, posterior edge) and compared with semi-automated three-dimensional (3D) tibial alignment measurements as the reference standard. RESULTS: Mid-dome LDTA showed no significant difference from the 3D reference (p&#x202f;>&#x202f;0.05) and demonstrated excellent agreement. Anterior measurements significantly overestimated LDTA, while posterior measurements underestimated it (both p&#x202f;<&#x202f;0.05), with only fair agreement. CONCLUSION: LDTA should be measured at the mid-dome of the distal tibial plafond on WBCT to ensure accurate and reproducible alignment assessment. LEVEL OF EVIDENCE: Level III - Retrospective Comparative Study.

Humans

Use of indocyanine green fluorescence versus patent blue V dye for sentinel lymph node biopsy in early breast cancer, a randomized controlled trial.

BACKGROUND: Sentinel lymph node biopsy (SLNB) is standard for axillary staging in early breast cancer. While the combination of radioisotope and blue dye (e.g., patent blue V, PBV) remains the standard, it has limitations including logistics, variable identification rate (IR), and allergic potential. Indocyanine green (ICG) fluorescence is a promising alternative, but high-quality comparative evidence is needed. METHODS: This was a single-center, prospective, randomized controlled trial. Forty patients with early-stage, node-negative breast cancer were allocated to SLNB using either ICG (n&#x2009;=&#x2009;20) or PBV (n&#x2009;=&#x2009;20). All patients subsequently underwent completion level I-II axillary lymph node dissection (ALND) as the pathological reference standard for diagnostic performance assessment. Primary outcome was sentinel lymph node (SLN) IR. Secondary outcomes included detection time, number of SLNs retrieved, false-negative rate (FNR), and safety. RESULTS: Baseline characteristics were comparable between groups. The SLN IR was significantly higher with ICG (100% [20/20]) than with PBV (75% [15/20], p&#x2009;=&#x2009;0.047). ICG was associated with a significantly shorter median detection time (14.5 vs. 24.0&#xa0;min, p&#x2009;<&#x2009;0.001) and retrieved more SLNs (mean: 3.6 vs. 2.4, p&#x2009;=&#x2009;0.002). Most critically, ICG demonstrated 100% sensitivity, specificity, negative predictive value (NPV), and overall diagnostic accuracy, with a 0% FNR. In contrast, PBV achieved a sensitivity of 75%, an overall diagnostic accuracy of 90%, and an FNR of 25%. No ICG-related adverse events occurred. PBV caused skin discoloration in 75% of patients and one (5%) allergic reaction. CONCLUSION: ICG fluorescence achieved a higher SLN IR, shorter detection time, higher sensitivity, lower FNR, and fewer tracer-related adverse events than PBV as a single tracer for SLNB in patients with early-stage breast cancer. These findings suggest that ICG is a promising standalone tracer when radioisotope mapping is unavailable. Larger multicenter studies are required before widespread adoption can be recommended.

Humans

Characteristics of p53 and Smad4 immunohistochemistry in pancreatic ductal adenocarcinoma and validation by next-generation sequencing.

BACKGROUND: Mutations in four major driver genes -KRAS, CDKN2A, TP53, and SMAD4- are central to the pathogenesis of pancreatic ductal adenocarcinoma (PDAC) and critically inform diagnosis, therapeutic decision-making, and prognostic assessment. Although next-generation sequencing (NGS) is widely regarded as the gold standard for detecting these mutations, its clinical application is often limited by suboptimal analytical efficiency and substantial economic cost. Among these genes, immunohistochemical (IHC) staining for the proteins encoded by TP53 and SMAD4 has been extensively adopted in routine pathology practice. However, standardized IHC pattern classification schemes and rigorous validation of their predictive accuracy for underlying genomic alterations remain lacking in PDAC. METHODS: We retrospectively enrolled 63 PDAC patients and systematically characterized the typical IHC expression patterns of p53 and Smad4. Targeted NGS was subsequently performed on all available tumor specimens, and the resulting mutational profiles were correlated with corresponding IHC findings. Diagnostic performance including sensitivity, specificity and accuracy of p53 IHC for predicting TP53 mutations and of Smad4 IHC for predicting SMAD4 mutations was rigorously evaluated. RESULTS: Among the four canonical driver genes, co-occurring double- or triple-gene mutations were prevalent; within TP53 and SMAD4, missense mutations constituted the most frequent variant type. Using NGS as the reference standard, we validated the diagnostic utility of a three-tiered p53 IHC classification system, particularly in fine-needle biopsy (FNB) specimens. Furthermore, we proposed a novel, refined Smad4 IHC pattern classification that incorporates an "intermediate" category, thereby expanding upon conventional binary interpretation. This new scheme achieved markedly improved mutation prediction accuracy (0.76) compared with traditional approaches (0.57). CONCLUSION: Our study highlights the complementary diagnostic value of p53 and Smad4 IHC relative to molecular testing in PDAC, especially when tissue is limited, as commonly encountered in FNB specimens. The newly established Smad4 IHC classification system, which integrates an intermediate expression category into the conventional two-tier framework, demonstrates superior clinical utility and enhances predictive accuracy for SMAD4 genomic alterations.

Humans

Stretched penile length in boys with hypospadias: Population-based analysis using validated nomogram.

BACKGROUND: Hypospadias affects 1 in 200-300 male births. Parents are often concerned about penile adequacy beyond the urethral defect itself, yet few studies have systematically compared stretched penile length (SPL) in hypospadias against population-based reference standards. OBJECTIVE: To evaluate SPL distribution patterns in boys with Types I and II hypospadias and compare them with established normative data. METHODS: The authors studied 876 consecutive boys aged 1-14 years with unoperated Types I (distal) and II (mid-shaft) hypospadias. Two observers independently measured SPL using the validated SPLINT technique. The SPL measurements were compared against age-matched normative data from 1276 Indian children. Exact binomial probability tests were used for percentile distributions, chi-square tests for subtype comparisons and t-tests for mean deviations. RESULTS: The cohort included 479 Type I and 397 Type II cases. SPL distribution showed a marked leftward shift: 71% fell below the 50th percentile (expected 50%, p < 0.001) and 41.5% below the 25th percentile. Lower percentiles were overrepresented, 20.7% were below the 10th percentile and 20.8% in the 10th-25th range. Upper percentiles were depleted: only 7.4% in the 75th-90th range and 1.7% above the 90th percentile (all p < 0.001). Mean SPL was reduced by 6.8% (95% CI: -8.18 to -5.42%) in Type I and 7.5% (95% CI: -9.05 to -5.92%) in Type II. The two subtypes showed no significant distributional difference (&#x3c7;2 = 6.22, p = 0.18), suggesting that meatal position does not predict SPL reduction. CONCLUSIONS: Boys with distal and mid-shaft hypospadias show clinically meaningful SPL reduction that follows a continuous distribution rather than an all-or-none pattern. SPL reduction appears independent of meatal position. These findings support routine SPL assessment using population-specific references and can guide preoperative counselling.

Humans

Soft Tissue Volume Augmentation at Single Implant Sites Applying Collagen Matrices or Connective Tissue Grafts: 10-Year Follow-Up of a Randomized Controlled Trial.

AIM: To compare up to 10&#x2009;years clinical, profilometric and patient-reported outcomes of implant sites previously augmented using a volume-stable collagen matrix (VCMX) or connective tissue graft (SCTG) in the aesthetic zone. METHODS: The original non-inferiority randomized controlled trial (RCT) enrolled 20 patients who received soft tissue volume augmentation with VCMX or SCTG at single implant sites. Clinical assessments and standardized measurements were performed at baseline after crown insertion and at 6&#x2009;months, 1, 3, 5, 7.5, and 10&#x2009;years. The primary outcome was mucosal thickness. Secondary outcomes included marginal bone levels (MBL), probing depth (PD), bleeding on probing (BOP), plaque control record, Pink Aesthetic Score (PES), OHIP-14 and buccal profilometric changes. Group comparisons were performed using mixed-effects and generalized estimating equation (GEE) models, which account for within-patient correlations due to repeated measurements and allow inclusion of all available data without requiring imputation for missing observations. RESULTS: Of the 20 originally enrolled patients, 10 (5 in the SCTG group and 5 in the VCMX group) were available for re-examination at 10&#x2009;years. The adjusted between-group difference in mucosal thickness was -0.02&#x2009;mm (95% CI -0.99 to 0.96). As the lower bound of the confidence interval remained above the prespecified non-inferiority margin of -1&#x2009;mm, non-inferiority of VCMX was shown. Buccal contour changes were comparable during the early follow-up, while a trend toward a greater long-term contour decrease was observed in group VCMX (-0.31&#x2009;mm [95% CI, -0.65 to 0.03]; p&#x2009;=&#x2009;0.07). Mean PES values were 10.6 in the SCTG group and 9.6 in the VCMX group, with no significant between-group differences (p&#x2009;=&#x2009;0.45). Both groups revealed high levels of oral health-related quality of life, with low median OHIP-14 scores (SCTG, 0.0; VCMX, 1.0; p&#x2009;=&#x2009;0.26). CONCLUSION: These preliminary long-term findings showed no clinically relevant differences between SCTG and VCMX in terms of clinical, profilometric and patient-reported outcomes. While SCTG remains the reference standard, VCMX represents a less invasive alternative but with a slight tendency toward greater long-term contour reduction. CLINICAL SIGNIFICANCE: Volume-stable collagen matrices serve as a viable alternative to autogenous connective tissue grafts for peri-implant soft tissue volume augmentation, particularly in patients seeking a reduced morbidity, without compromising long-term clinical or aesthetic outcomes. TRIAL REGISTRATION: German Clinical Trials Register: DRKS00017484.

Humans

Diagnostic Accuracy of a CRISPR-Based Assay in Smear- and Culture-Negative Fungal Keratitis.

IMPORTANCE: Diagnosing fungal keratitis (FK) in patients with negative smear and culture results remains clinically challenging, highlighting the need for alternative diagnostic approaches. OBJECTIVE: To determine the diagnostic accuracy of the clustered regularly interspaced short palindromic repeats (CRISPR)-based Rapid Identification of Mycoses using CRISPR (RID-MyC) assay for detecting FK in patients with negative smear and culture results using in vivo confocal microscopy (IVCM) as the reference standard. DESIGN, SETTING, AND PARTICIPANTS: This prospective diagnostic accuracy study was conducted from December 2024 to March 2025 at Aravind Eye Hospital, a tertiary ophthalmology referral hospital in Coimbatore, India. Consecutive patients clinically suspected to have microbial keratitis with negative smear and culture results were eligible for inclusion. Data were analyzed from March 2025 to June 2025. INTERVENTIONS: All included participants underwent corneal scraping for RID-MyC assay and imaging by IVCM. MAIN OUTCOMES AND MEASURES: The primary outcomes were sensitivity, specificity, positive predictive value, negative predictive value, and diagnostic concordance of the RID-MyC assay compared with IVCM results. RESULTS: Of 245 consecutive patients clinically suspected to have microbial keratitis, 82 were smear negative. After exclusions due to contraindications or positive subsequent cultures, 41 patients with smear- and culture-negative results were ultimately included in the final analysis. Of these 41 patients (mean [SD] age, 51.0 [14.6] years; 21 [51.2%] women), RID-MyC demonstrated sensitivity of 82.1% (95% CI, 63%-94%) and specificity of 76.9% (95% CI, 46%-95%). Positive predictive value was 88.5% (95% CI, 74%-95%) and negative predictive value was 66.7% (95% CI, 46%-82%). Concordance between RID-MyC and IVCM was observed in 33 cases (80.5%). Notably, prior antifungal treatment was most frequent (4 of 5 [80%]) among patients with positive IVCM but negative RID-MyC results. Conversely, all patients (3 of 3 [100%]) with negative IVCM but positive RID-MyC findings had smaller, peripheral, or paracentral lesions. CONCLUSIONS AND RELEVANCE: In this diagnostic study, in patients with smear- and culture-negative FK, the RID-MyC assay showed good diagnostic accuracy comparable with IVCM and was feasible in all cases, including those in whom imaging was not possible. With its rapid turnaround and minimal equipment needs, RID-MyC may serve as a practical adjunct to conventional diagnostics, particularly in high-burden, resource-limited settings where IVCM is unavailable or contraindicated.

Humans

Holmium laser enucleation of the prostate for the treatment of lower urinary tract symptoms in men with benign prostatic hyperplasia.

RATIONALE: A range of surgical options is available for the treatment of benign prostatic hyperplasia (BPH), including holmium laser enucleation of the prostate (HoLEP). The evidence is unclear regarding differences in functional, perioperative, and morbidity outcomes between these modalities. OBJECTIVES: To assess the effects of holmium laser enucleation of the prostate compared with other surgical treatments for lower urinary tract symptoms in men with benign prostatic hyperplasia. SEARCH METHODS: We searched multiple databases (including MEDLINE, Embase, CENTRAL, Web of Science, LILACS, and the International HTA database), trial registries, and conference abstracts through April 08, 2026. ELIGIBILITY CRITERIA: We only included randomized trials of men over 40 years of age with a prostate volume of at least 20 mL (assessed by digital rectal examination, ultrasound, or conventional imaging) who exhibited lower urinary tract symptoms (LUTS) defined by an International Prostate Symptom Score (IPSS) of eight or greater undergoing surgical interventions for BPH. OUTCOMES: The critical outcomes measured were the urologic symptoms score, the quality-of-life score, and major adverse events. The important outcomes measured were: re-treatment, erectile function, ejaculatory function, transfusions, acute urinary retention, indwelling urinary catheter duration, and hospital stay duration. RISK OF BIAS: We used the Cochrane risk of bias tool (RoB 1) to assess for potential sources of bias on a study and outcome level basis. SYNTHESIS METHODS: We pooled outcome data using the random-effects model and performed meta-analyses using the Mantel-Haenszel method. We assessed statistical heterogeneity in the pooled data by visually inspecting forest plots and using the I2 statistic to quantify it. We used the GRADE framework to assess the certainty of evidence. INCLUDED STUDIES: We included 52 trials that included 6242 participants that compared HoLEP to other surgical interventions for benign prostatic hyperplasia. The median age of participants across the studies ranged from 65 to 74 years. The baseline prostate volume ranged from 30 cc to 142 cc. Baseline IPSS scores ranged from 19.6 to 28.6 (range 0-35). SYNTHESIS OF RESULTS: We prioritized comparing HoLEP with transurethral resection of the prostate (TURP) at short-term follow-up (up to 12 months), because TURP is the long-standing reference standard and the predominant comparator in randomized surgical trials. Findings for the four remaining comparisons (laser ablation, alternative energy source enucleation, other minimally invasive therapies, and simple prostatectomy), for long-term follow-up, and for all remaining outcomes are reported in full in the review. Compared to TURP, at short-term follow-up: Critical outcomes - HoLEP may result in little to no difference in short-term urologic symptom scores measured using the IPSS (range 0 to 35; lower values reflect fewer symptoms) (MD -0.67, 95% CI -1.20 to -0.14; I&#xb2; = 93%; 14 studies, 1666 participants, low-certainty evidence). - HoLEP may result in little to no difference in short-term quality of life (range 0 to 6; lower values reflect better quality of life) (MD -0.04, 95% CI -0.23 to 0.15; I&#xb2; = 73%; 6 studies, 876 participants, low-certainty evidence). - HoLEP may result in little to no difference in short-term major adverse events (RR 0.75, 95% CI 0.35 to 1.58; I&#xb2; = 0%; 10 studies, 1147 participants, low-certainty evidence). Important outcomes - HoLEP likely results in little to no difference in re-treatment (RR 0.45, 95% CI 0.14 to 1.50; I&#xb2; = 0%; 8 studies, 813 participants, moderate-certainty evidence). - HoLEP likely results in little to no difference in erectile function (MD -0.03, 95% CI -0.47 to 0.42; I&#xb2; = 0%; 3 studies, 518 participants, moderate-certainty evidence). - Ejaculatory function: we did not find any data for this outcome. - HoLEP likely reduces the need for blood transfusion (RR 0.19, 95% CI 0.09 to 0.42; I&#xb2; = 0%; 15 studies, 1755 participants, moderate-certainty evidence). AUTHORS' CONCLUSIONS: Compared with TURP, HoLEP may achieve similar relief of urologic symptoms, similar quality of life, and similar rates of major adverse events in the first 12 months after surgery, and probably similar re-treatment rates and erectile function. HoLEP likely reduces the need for blood transfusion; this is the only advantage of HoLEP that the randomized evidence, as summarized here, supports as clinically important. There was insufficient evidence to assess outcomes in the subset of individuals with larger prostates or on anticoagulation. Future research should prioritize long-term trials reporting sexual function and urinary incontinence outcomes, recruit men with very large prostates (&#x2265; 150 cc) or on anticoagulation therapy, and evaluate cost-effectiveness and training requirements. FUNDING: No external funding was received for this review. REGISTRATION: The protocol for this review was published in the Cochrane Database 2019 (https://doi.org/10.1002/14651858.CD013291).

Humans

Absolute Quantification of Cellular and Cell-Free Mitochondrial DNA Copy Number from Human Blood and Urinary Samples Using Real Time Quantitative PCR.

Mitochondrial DNA copy number (mtDNA-CN) in human body fluids is widely used as a biomarker of mitochondrial dysfunction in common metabolic diseases. Here we describe protocols to measure cellular and/or cell free (cf)-mtDNA-CN in human peripheral blood and urine. Cellular mtDNA is located inside the mitochondria where it encodes key subunits of the respiratory complexes in mitochondria and is usually normalized with reference to the nuclear genome as the mitochondrial genome to nuclear genome ratio (Mt/N) in either whole blood, peripheral blood mononuclear cells (PBMCs), or whole urine. Cf -mtDNA is usually found outside of the mitochondria, often released following mitochondrial damage, can trigger inflammatory pathways, and is usually measured as mtDNA-CN per volume of the starting material. Here we describe how to (1) separate whole blood into PBMCs, plasma, and serum fractions and whole urine into urinary supernatant and pellet, (2) prepare DNA from each of these fractions, (3) prepare reference&#xa0;standards&#xa0;for absolute quantification, (4) carry out qPCR for either relative or absolute quantification from test samples, (5) analyze qPCR data, and (6) calculate the sample size to adequately power studies. The protocol presented here is suitable for high throughput use and can be modified to quantify mtDNA from other body fluids, human cells, and tissues.

Humans

Repeated scoring with the adult appendicitis score improves the sensitivity and the specificity of appendicitis diagnosis in patients with early equivocal signs of appendicitis: a secondary analysis.

PURPOSE: The utilization of computed tomography in the early stage of acute appendicitis may result in overdiagnosis and unnecessarily expose patients to ionising radiation. The Adult Appendicitis Score (AAS) can be used to select patients for imaging. Observation and re-scoring in the DIAMOND trial reduced the need for imaging. Now, we wanted to determine if the change in AAS (&#x2206;AAS) can serve as a diagnostic tool to select patients for imaging even more precisely. METHODS: Eighty-eight patients with early equivocal appendicitis participated in the observation arm of the DIAMOND trial. The data for these patients were reanalysed, and &#x2206;AAS during the observation was calculated. The baseline AAS, final AAS, and the change in C-reactive protein (&#x2206;CRP) were selected as reference standards. RESULTS: Eighty-three patients with complete data were included in the analysis. The AUROC (Area Under the Receiver Operating Characteristic) values are as follows: &#x2206;AAS, 0.932 (95% CI 0.868-0.996); baseline AAS, 0.629 (95% CI 0.498-0.760); final AAS, 0.936 (95% CI 0.886-0.987); and &#x2206;CRP, 0.796 (95% CI 0.696-0.897). Using receiver operating characteristic curves, we established the thresholds for low (AAS&#x2009;&#x2264;&#x2009;-2), intermediate (AAS -1 to 0), and high (AAS&#x2009;&#x2265;&#x2009;1) probability of appendicitis. The negative predictive value for the low-probability group and the positive predictive value for the high-probability group concerning acute appendicitis were 97% and 94%, respectively. CONCLUSION: Patients with equivocal signs of appendicitis may benefit from short observation and the calculation of &#x2206;AAS to reduce overdiagnosis and exposure to excessive imaging. REGISTRATION: The DIAMOND trial was officially registered on ClinicalTrials.gov (NCT02742402) on April 13, 2016.

Adult

Clinical performance of the Abbott RealTime Mycobacterium tuberculosis (MTB) PCR on bronchoscopic specimens for diagnosing pulmonary tuberculosis.

PURPOSE: We evaluated the performance of the Abbott RealTime Mycobacterium tuberculosis (MTB) PCR (RT MTB) on bronchoscopic specimens using conventional culture as the reference standard in a low-Tuberculosis (TB) prevalence setting. METHODS: A total of 6,988 specimens (4,682 bronchial aspirates [BAS] and 2,306 bronchoalveolar lavages [BAL]) from 4,118 patients with suspected pulmonary TB were included. When BAS and BAL specimens from the same bronchoscopy procedure were available, these were mixed 1:1 prior to culture inoculation and PCR testing. Following processing, specimens were inoculated into a L&#xf6;wenstein-Jensen and a Bactec MGIT 960 tube and incubated at 37&#xa0;&#xb0;C for 3 months and at 35&#xa0;&#xb0;C for 8 weeks, respectively. RT MTB was performed as indicated by the manufacturer. RT MTB targets the insertion sequence IS6110 and the protein antigen B (PAB) gene, both highly conserved within the Mycobacterium tuberculosis complex. Whole genome Next generation sequencing of clinical MTBC isolates was performed when indicated. RESULTS: Among the 104 culture-positive specimens, 84 (1.2%) were detected by PCR. Additionally, 16 specimens (0.3% of all samples), from 16 patients, were PCR-positive despite negative culture results. Conversely, 20 specimens (0.4% of all samples), from 19 patients, were culture-positive but not detected by PCR. No significant differences were found between PCR-positive and PCR-negative specimens with respect to the number of IS6110 copies per isolate or PAB gene sequences (P&#x2009;=&#x2009;0.69). Finally, there were 4,683 specimens (97.4%) from the remaining 4,014 patients tested PCR-negative/Culture-negative. Following the resolution of discrepancies based on clinical grounds the sensitivity and specificity of RT MTB were 83.9% (CI 95%, 76.0-90.0) and 99.9% (CI 95%, 99.9-99.9), respectively. These results exceed the minimum performance requirements defined in the WHO Target Product Profiles for molecular TB diagnostics. Although RT MTB is not a point-of-care test but rather a moderate-complexity automated NAAT, it is recommended by WHO as part of the Abbott RealTime MTB/MTB RIF-INH testing algorithm, in which MTBC detection by RT MTB is followed by reflex testing with the MTB RIF/INH assay for detection of rifampicin and isoniazid resistance. CONCLUSION: RT MTB shows a good performance on bronchoscopic specimens.

Mycobacterium tuberculosis

Evaluation of a cornea-specialized large language model for diagnostic and management accuracy in complex corneal cases.

PURPOSE: To evaluate whether a cornea-specialized large language model (LLM) enhanced with retrieval-augmented generation (RAG) improves clinicians' diagnostic and management accuracy in complex corneal cases compared to a general-purpose GPT-4o model and unaided clinician performance. METHODS: This prospective, randomized, masked evaluation study involved three cornea trainees who each independently reviewed 39 real-world corneal cases under three experimental conditions: unaided, GPT-4o-assisted, and assisted by a cornea-specialized GPT-4o model. The cornea-specialized model was constructed by embedding over 200 publicly available Wikipedia articles into GPT-4o's RAG framework. Participants provided open-ended diagnoses and selected the next-step management options (multiple choice). They were allowed up to three GPT-4o queries per case, and the AI-assisted arms were randomized to minimize bias. Accuracy for both tasks was compared against expert reference standards using McNemar's test. RESULTS: Diagnostic accuracy was 48.7%, 20.5%, and 38.5% unaided, improving to 69.2%, 46.2%, and 59.0% with general GPT-4o (p<0.04). The cornea-specialized GPT-4o further improved accuracy to 71.8%, 48.7%, and 74.4%, with improvements over unaided performance for all clinicians (p<0.01). For next-step decisions, unaided accuracy was 76.9%, 87.2%, and 59.0%. With the specialized model, Ophthalmologist 3 improved to 71.8% (p<0.05), Ophthalmologist 1 remained high at 82.1%, and Ophthalmologist 2 declined to 64.1% (p<0.05). CONCLUSIONS: A cornea-specialized LLM enhanced with RAG improved diagnostic accuracy in complex corneal cases, particularly among clinicians with lower baseline performance. Effects on management accuracy were inconsistent. Future studies should explore the use of open-ended management tasks and examine whether smaller, curated retrieval corpora yield better model performance.

Humans