Search PubMedSearch

PubMed · 41000838

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Abstract

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Caralyn Reisle, Cameron J Grisdale, Kilannin Krysiak, Arpad M Danos, Mariam Khanfar, Erin Pleasance, Jason Saliba, Melika Hanos, Nilan V Patel, Asmita Jain, Joshua F McMichael, Ajay C Venigalla, Malachi Griffith, Obi L Griffith, Steven J M Jones. 2025-09-15. Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.. https://doi.org/10.1101/2025.09.10.675443

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Effectiveness of high-dose versus standard-dose influenza vaccines against hospitalisation according to frailty risk: a prespecified analysis of the randomised trial DANFLU-2.

BACKGROUND: Frailty is a major risk factor for influenza-related complications and can influence vaccine effectiveness. We aimed to assess the relative vaccine effectiveness (rVE) of high-dose (HD-IIV) versus standard-dose inactivated influenza vaccine (SD-IIV) in older adults aged 65 years or older according to frailty risk. METHODS: This study was a prespecified analysis of DANFLU-2, an open-label, individually randomised trial, conducted in Denmark during three consecutive influenza seasons (2022-23, 2023-24, and 2024-25). Adults aged 65 years or older were randomised (1:1) to the HD-IIV or SD-IIV group. The primary endpoint was hospitalisation for influenza or pneumonia. Frailty was defined according to the validated Hospital Frailty Risk Score (HFRS) based on ICD-10 codes within 10 years before randomisation. Participants were stratified into three HFRS categories, namely low (<5 points), intermediate (5-15 points), and high (>15 points) frailty risk. The rVE of HD-IIV versus SD-IIV against the primary endpoint was assessed across prespecified HFRS categories and treating HFRS as a continuous variable. Pearson's chi-square test was used to compare safety events across frailty risk groups and randomisation groups. FINDINGS: Among 332&#x2009;438 randomised participants (mean age 73&#xb7;7 years [SD 5&#xb7;8]; 161&#x2009;538 [48&#xb7;6%] were female), 276&#x2009;173 (83&#xb7;1%) had low frailty risk, 52&#x2009;395 (15&#xb7;8%) had intermediate frailty risk, and 3861 (1&#xb7;2%) had high frailty risk. The primary endpoint of hospitalisation for influenza or pneumonia occurred in 1424 (0&#xb7;5%) of 276&#x2009;173 participants with low frailty risk, 761 (1&#xb7;5%) of 52&#x2009;395 with intermediate frailty risk, and 163 (4&#xb7;2%) of 3861 with high frailty risk (relative risk [RR] for intermediate vs low frailty risk 2&#xb7;8 [95% CI 2&#xb7;6-3&#xb7;1]; RR for high vs low frailty risk 8&#xb7;2 [7&#xb7;0-9&#xb7;6]). HFRS as a continuous variable significantly modified the effect of HD-IIV versus SD-IIV against the primary endpoint with higher rVE estimates with increasing HFRS (pinteraction=0&#xb7;020). The rVE was 0&#xb7;2% (95% CI -10&#xb7;8 to 10&#xb7;2) among those with low frailty risk, 13&#xb7;1% (-0&#xb7;4 to 24&#xb7;8) among those with intermediate frailty risk, and 19&#xb7;9% (-10&#xb7;3 to 42&#xb7;1) among those with high frailty risk. No significant interaction was observed when HFRS was assessed according to the prespecified categorical frailty groups (pinteraction=0&#xb7;17). The proportion of participants with at least one serious adverse event increased across frailty risk groups (13&#x2009;366 [4&#xb7;8%] of 275&#x2009;795 for low frailty risk, 5475 [10&#xb7;5%] of 52&#x2009;315 for intermediate frailty risk, and 777 [20&#xb7;2%] of 3850 for high frailty risk; p<0&#xb7;0001), with similar proportions of serious adverse events in the HD-IIV and SD-IIV groups for each frailty risk group. INTERPRETATION: Among adults aged 65 years or older in Denmark, frailty risk might modify the effects of HD-IIV versus SD-IIV against hospitalisation for influenza or pneumonia, with higher rVE estimates with increasing frailty risk. These findings might support considering high-dose influenza vaccines for frail older adults. However, effect modification was not evident when frailty was assessed using prespecified categorical subgroups, and subgroup-specific estimates were imprecise, with 95% CIs crossing the null. These results should be considered exploratory, warranting further investigation. FUNDING: The DANFLU-2 trial was funded by Sanofi.

Journal Article