Search PubMedSearch

PubMed · 40024070

Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.

Abstract

OBJECTIVE: To quantify and improve the performance of standard rheumatoid arthritis (RA) algorithms in a biobank setting. METHODS: This retrospective cohort study within the Mayo Clinic (MC) Biobank and MC Tapestry Study identified RA cases by presence of at least two RA codes OR positive anti-cyclic citrullinated peptide antibodies (CCP) plus disease-modifying anti-rheumatic drug (DMARD) prescription as of 7/18/2022. Rheumatology physicians manually verified all RA cases using RA criteria and/or rheumatology physician diagnosis plus DMARD use. All other biobank participants served as non-RA controls. We defined seropositivity as rheumatoid factor and/or anti-CCP positivity. We assessed rules-based and Electronic Medical Records and Genomics (eMERGE) RA algorithms using positive predictive value (PPV). Finally, we developed a novel RA algorithm using a LASSO-based machine learning approach with five-fold cross validation. RESULTS: We identified 1,316 confirmed RA cases (968 MC Biobank, 348 Tapestry, 70 % seropositive) and 82,123 non-RA controls (mean age 65, 61 % female). The PPV of 3 RA codes was 43 %, codes plus DMARD was 54 %, and codes plus DMARD plus seropositivity was 85 %. The PPV of eMERGE was 77 %. Available in the MC Biobank, self-reported RA (PPV 10 %) only minimally improved algorithm performance (PPV from 83 % to 85 %), whereas family history of RA (PPV 3 %) worsened performance. At 90 % PPV, the novel RA algorithm incorporating key variables such as anti-CCP and DMARD use increased sensitivity by 4-11 % compared to eMERGE. CONCLUSION: Rules-based and eMERGE RA algorithms had worse performance in biobank than administrative settings. Our novel RA algorithm outperformed these standard algorithms.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vanessa L Kronzer, Katrina A Williamson, Andrew C Hanson, Jennifer A Sletten, Jeffrey A Sparks, John M Davis, Cynthia S Crowson. 2025-02-22. Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.. https://doi.org/10.1016/j.semarthrit.2025.152668

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and environmental information within a One Health framework, while addressing critical challenges in governance, equity, and interoperability. The discussion covers the entire genomic surveillance workflow, from sample collection and sequencing to bioinformatic analysis and phylogenetic inference, and highlights the transformative role of artificial intelligence (AI) in predictive surveillance. By analyzing global initiatives, operational barriers, and emerging technologies, this chapter underscores the necessity of sustainable, equitable, and interoperable genomic systems to proactively address current and future infectious disease threats.

Humans

Systematic Dissection of Key Driver Perturbation Signatures in Single Cells via ECCITE-seq.

CRISPR screens, such as expanded CRISPR-compatible cellular indexing of transcriptomes and epitopes by sequencing (ECCITE-seq), enable the simultaneous measurement of transcriptomes, gRNA identity, and cell-surface protein expression at single-cell resolution to systematically interrogate gene function. This platform provides a powerful and scalable experimental approach for validating disease-associated regulators identified by large-scale association studies and other computational methods, including network-based analyses of multi-omics data. Here, as an example application, we describe an ECCITE-seq framework to characterize the transcriptomic consequences of perturbing multiple neuronal key driver genes associated with Alzheimer's disease (AD) in human-induced pluripotent stem cell (hiPSC)-derived neurons. More broadly, by integrating customized pooled gRNA libraries with different CRISPR effectors across multiple cell types, this approach allows for the assessment of the regulatory impact of candidate genes implicated in development and disease processes.

Humans

Identification of Genome-Wide Chromatin Structural Aberration in Cancer by Hi-C Analysis.

Aberrant three-dimensional genome organization is a hallmark of cancer, often driving oncogene activation through mechanisms such as enhancer hijacking. High-throughput chromosome conformation capture (Hi-C) maps these interactions on a genome-wide scale. Unlike earlier dilution-based methods, in situ Hi-C performs proximity ligation within intact nuclei, minimizing random ligation noise and enabling fine-scale structure detection. This chapter describes an optimized in situ Hi-C protocol tailored for cancer cell lines using MboI digestion and biotin-mediated pull-down to generate high-complexity libraries. We further outline a computational workflow that extends beyond standard topological mapping of compartments and topologically associating domains to identify cancer-specific aberrations. Specifically, we focus on detecting chromosomal rearrangements (structural variants) and characterizing the distinct circular topology of extrachromosomal DNA. This integrated experimental and analytical framework provides the necessary tools to dissect the spatial dysregulation underlying tumor evolution.

Humans