Search PubMed⌕ Search

PubMed · 8896134

Explained variation for logistic regression.

Abstract

Different measures of the proportion of variation in a dependent variable explained by covariates are reported by different standard programs for logistic regression. We review twelve measures that have been suggested or might be useful to measure explained variation in logistic regression models. The definitions and properties of these measures are discussed and their performance is compared in an empirical study. Two of the measures (squared Pearson correlation between the binary outcome and the predictor, and the proportional reduction of squared Pearson residuals by the use of covariates) give almost identical results, agree very well with the multiple R2 of the general linear model, have an intuitively clear interpretation and perform satisfactorily in our study. For all measures the explained variation for the given sample and also the one expected in future samples can be obtained easily. For small samples an adjustment analogous to Radj2 in the general linear model is suggested. We discuss some aspects of application and recommend the routine use of a suitable measure of explained variation for logistic models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M Mittlböck, M Schemper. 1996-10-15. Explained variation for logistic regression.. https://doi.org/10.1002/(sici)1097-0258(19961015)15%3A19%3C1987%3A%3Aaid-sim318%3E3.0.co%3B2-9

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Complexity of the simplest phylogenetic estimation problem.

The maximum-likelihood (ML) solution to a simple phylogenetic estimation problem is obtained analytically The problem is estimation of the rooted tree for three species using binary characters with a symmetrical rate of substitution under the molecular clock. ML estimates of branch lengths and log-likelihood scores are obtained analytically for each of the three rooted binary trees. Estimation of the tree topology is equivalent to partitioning the sample space (space of possible data outcomes) into subspaces, within each of which one of the three binary trees is the ML tree. Distance-based least squares and parsimony-like methods produce essentially the same estimate of the tree topology, although differences exist among methods even under this simple model. This seems to be the simplest case, but has many of the conceptual and statistical complexities involved in phylogeny estimation. The solution to this real phylogeny estimation problem will be useful for studying the problem of significance evaluation.

Likelihood Functions↗

Optimal tests for no contamination in reliability models.

Inferences on mixtures of probability distributions, in general, and of life distributions, in particular, are receiving considerable importance in recent years. The likelihood ratio procedure of testing for the null hypothesis of no contamination is often very cumbersome and lacks its usual asymptotic properties. Recently, SenGupta (1991) has introduced the notion of an 'L-optimal' test for such testing problems. The idea is to recast the original several parametric hypotheses representation of the null hypothesis in terms of only a single hypothesis involving an appropriately chosen parametric function. This approach is shown to be both mathematically elegant and operationally simple for a quite general class of mixture distributions which contains, in particular, all mixtures of the one-parameter exponential family and also a very rich subclass of mixtures useful in life-testing and reliability analysis. It is also illustrated through two examples--one based on real-life data and the other on a simulated sample.

Likelihood Functions↗