Search PubMed⌕ Search

PubMed · 1789885

Cost-efficient study designs for binary response data with Gaussian covariate measurement error.

Abstract

When mismeasurement of the exposure variable is anticipated, epidemiologic cohort studies may be augmented to include a validation study, where a small sample of data relating the imperfect exposure measurement method to the better method is collected. Optimal study designs (i.e., least expensive subject to specified power constraints) are developed that give the overall sample size and proportion of the overall sample size allocated to the validation study. If better exposure measurements can be collected on a sample of subjects, an optimal design can be suggested that conforms to realistic budgetary constraints. The properties of three designs--those that include an internal validation study, those where the validated subsample is derived from subjects external to the primary investigation, and those that use the better method of exposure assessment on all subjects--are compared. The proportion of overall study resources allocated to the validation substudy increases with increasing sample disease frequency, decreasing unit cost of the superior exposure measurement relative to the imperfect one, increasing unit cost of outcome ascertainment, increasing distance between two alternative values of the relative risk between which the study is designed to discriminate, and increasing magnitude of hypothesized values. This proportion also depends in a nonlinear fashion on the severity of measurement error, and when the validation study is internal, measurement error reaches a point after which the optimal design is the smaller, fully validated one.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D Spiegelman, R Gray. 1991. Cost-efficient study designs for binary response data with Gaussian covariate measurement error.. https://pubmed.ncbi.nlm.nih.gov/1789885/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A boosting approach to flexible semiparametric mixed models.

In linear mixed models the influence of covariates is restricted to a strictly parametric form. With the rise of semi- and non-parametric regression also the mixed model has been expanded to allow for additive predictors. The common approach uses the representation of additive models as mixed models. An alternative approach that is proposed in the present paper is likelihood based boosting. Boosting originates in the machine learning community where it has been proposed as a technique to improve classification procedures by combining estimates with reweighted observations. Likelihood based boosting is a general method which may be seen as an extension of L2 boost. In additive mixed models the advantage of boosting techniques in the form of componentwise boosting is that it is suitable for high dimensional settings where many explanatory variables are present. It allows to fit additive models for many covariates with implicit selection of relevant variables and automatic selection of smoothing parameters. Moreover, boosting techniques may be used to incorporate the subject-specific variation of smooth influence functions by specifying 'random slopes' on smooth effects. This results in flexible semiparametric mixed models which are appropriate in cases where a simple random intercept is unable to capture the variation of effects across subjects.

Cohort Studies↗

Estimation of attributable number of deaths and standard errors from simple and complex sampled cohorts.

Estimates of the attributable number of deaths (AD) from all causes can be obtained by first estimating population attributable risk (AR) adjusted for confounding covariates, and then multiplying the AR by the number of deaths determined from vital mortality statistics that occurred in the population for a specific time period. Proportional hazard regression estimates of adjusted relative hazards obtained from mortality follow-up data from a cohort is combined with a joint distribution of risk factor and confounders to compute an adjusted AR. Two estimators of adjusted AR are examined. These estimators differ according to which reference population is used to obtain the joint distribution of risk factor and confounders. Two types of reference populations were considered: (i) the population represented by the baseline cohort and (ii) a population that is external to the cohort. Methods used in survey sampling are applied to obtain estimates of the variance of the AD estimator. These variances can be applied to data that range from simple random samples to multistage stratified cluster samples, which are used in national household surveys. The variance estimation of AD is illustrated in an analysis of excess deaths due to having a non-ideal body mass index using the second National Health and Examination Survey (NHANES) Mortality Study and the 1999-2002 NHANES. These methods can also be used to estimate the attributable number of cause-specific deaths and their standard errors when the time period for the accrual of deaths is short.

Cohort Studies↗

Longitudinal variable selection by cross-validation in the case of many covariates.

Longitudinal models are commonly used for studying data collected on individuals repeatedly through time. While there are now a variety of such models available (marginal models, mixed effects models, etc.), far fewer options exist for the closely related issue of variable selection. In addition, longitudinal data typically derive from medical or other large-scale studies where often large numbers of potential explanatory variables and hence even larger numbers of candidate models must be considered. Cross-validation is a popular method for variable selection based on the predictive ability of the model. Here, we propose a cross-validation Markov chain Monte Carlo procedure as a general variable selection tool which avoids the need to visit all candidate models. Inclusion of a 'one-standard error' rule provides users with a collection of good models as is often desired. We demonstrate the effectiveness of our procedure both in a simulation setting and in a real application.

Cohort Studies↗