Search PubMed⌕ Search

PubMed · 9744903

Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms.

Abstract

This article reviews five approximate statistical tests for determining whether one learning algorithm outperforms another on a particular learning task. These tests are compared experimentally to determine their probability of incorrectly detecting a difference when no difference exists (type I error). Two widely used statistical tests are shown to have high probability of type I error in certain situations and should never be used: a test for difference of two proportions and a paired-differences t test based on taking several random train-test splits. A third test, a paired-differences t test based on 10-fold cross-validation, exhibits somewhat elevated probability of type I error. A fourth test, McNemar's test, is shown to have low type I error. The fifth test is a new test, 5 x 2 cv, based on five iterations of twofold cross-validation. Experiments show that this test also has acceptable type I error. The article also measures the power (ability to detect algorithm differences when they do exist) of these tests. The cross-validated t test is the most powerful. The 5 x 2 cv test is shown to be slightly more powerful than McNemar's test. The choice of the best test is determined by the computational cost of running the learning algorithm. For algorithms that can be executed only once, McNemar's test is the only test with acceptable type I error. For algorithms that can be executed 10 times, the 5 x 2 cv test is recommended, because it is slightly more powerful and because it directly measures variation due to the choice of training set.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

TG Dietterich. 1998-09-15. Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms.. https://doi.org/10.1162/089976698300017197

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Effectiveness of high-dose versus standard-dose influenza vaccines against hospitalisation according to frailty risk: a prespecified analysis of the randomised trial DANFLU-2.

BACKGROUND: Frailty is a major risk factor for influenza-related complications and can influence vaccine effectiveness. We aimed to assess the relative vaccine effectiveness (rVE) of high-dose (HD-IIV) versus standard-dose inactivated influenza vaccine (SD-IIV) in older adults aged 65 years or older according to frailty risk. METHODS: This study was a prespecified analysis of DANFLU-2, an open-label, individually randomised trial, conducted in Denmark during three consecutive influenza seasons (2022-23, 2023-24, and 2024-25). Adults aged 65 years or older were randomised (1:1) to the HD-IIV or SD-IIV group. The primary endpoint was hospitalisation for influenza or pneumonia. Frailty was defined according to the validated Hospital Frailty Risk Score (HFRS) based on ICD-10 codes within 10 years before randomisation. Participants were stratified into three HFRS categories, namely low (<5 points), intermediate (5-15 points), and high (>15 points) frailty risk. The rVE of HD-IIV versus SD-IIV against the primary endpoint was assessed across prespecified HFRS categories and treating HFRS as a continuous variable. Pearson's chi-square test was used to compare safety events across frailty risk groups and randomisation groups. FINDINGS: Among 332&#x2009;438 randomised participants (mean age 73&#xb7;7 years [SD 5&#xb7;8]; 161&#x2009;538 [48&#xb7;6%] were female), 276&#x2009;173 (83&#xb7;1%) had low frailty risk, 52&#x2009;395 (15&#xb7;8%) had intermediate frailty risk, and 3861 (1&#xb7;2%) had high frailty risk. The primary endpoint of hospitalisation for influenza or pneumonia occurred in 1424 (0&#xb7;5%) of 276&#x2009;173 participants with low frailty risk, 761 (1&#xb7;5%) of 52&#x2009;395 with intermediate frailty risk, and 163 (4&#xb7;2%) of 3861 with high frailty risk (relative risk [RR] for intermediate vs low frailty risk 2&#xb7;8 [95% CI 2&#xb7;6-3&#xb7;1]; RR for high vs low frailty risk 8&#xb7;2 [7&#xb7;0-9&#xb7;6]). HFRS as a continuous variable significantly modified the effect of HD-IIV versus SD-IIV against the primary endpoint with higher rVE estimates with increasing HFRS (pinteraction=0&#xb7;020). The rVE was 0&#xb7;2% (95% CI -10&#xb7;8 to 10&#xb7;2) among those with low frailty risk, 13&#xb7;1% (-0&#xb7;4 to 24&#xb7;8) among those with intermediate frailty risk, and 19&#xb7;9% (-10&#xb7;3 to 42&#xb7;1) among those with high frailty risk. No significant interaction was observed when HFRS was assessed according to the prespecified categorical frailty groups (pinteraction=0&#xb7;17). The proportion of participants with at least one serious adverse event increased across frailty risk groups (13&#x2009;366 [4&#xb7;8%] of 275&#x2009;795 for low frailty risk, 5475 [10&#xb7;5%] of 52&#x2009;315 for intermediate frailty risk, and 777 [20&#xb7;2%] of 3850 for high frailty risk; p<0&#xb7;0001), with similar proportions of serious adverse events in the HD-IIV and SD-IIV groups for each frailty risk group. INTERPRETATION: Among adults aged 65 years or older in Denmark, frailty risk might modify the effects of HD-IIV versus SD-IIV against hospitalisation for influenza or pneumonia, with higher rVE estimates with increasing frailty risk. These findings might support considering high-dose influenza vaccines for frail older adults. However, effect modification was not evident when frailty was assessed using prespecified categorical subgroups, and subgroup-specific estimates were imprecise, with 95% CIs crossing the null. These results should be considered exploratory, warranting further investigation. FUNDING: The DANFLU-2 trial was funded by Sanofi.

Journal Article↗