Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Strategies for the analysis of imputed data from a sample survey. The National Medical Care Utilization and Expenditure Survey.

Missing data in sample surveys is virtually unavoidable, whether it is an entire unit that is missing or only an item for a responding unit. Compensation for unit nonresponse is usually made through the assignments of weights to responding units; for item nonresponse, the compensation often is by an imputation procedure. This paper reviews the extent of missing data in a large federal survey, the National Medical Care Utilization and Expenditure Survey, and the imputation procedures used to compensate for item missing data. The effects of imputation on several types of estimates from the survey are examined. In addition, several methods for analyzing survey data with imputed values are reviewed, and recommendations about preferred strategies are made for selected circumstances.

Data Collection↗

A multiple imputation approach to regression analysis for doubly censored data with application to AIDS studies.

Sun, Liao, and Pagano (1999) proposed an interesting estimating equation approach to Cox regression with doubly censored data. Here we point out that a modification of their proposal leads to a multiple imputation approach, where the double censoring is reduced to single censoring by imputing for the censored initiating times. For each imputed data set one can take advantage of many existing techniques and software for singly censored data. Under the general framework of multiple imputation, the proposed method is simple to implement and can accommodate modeling issues such as model checking, which has not been adequately discussed previously in the literature for doubly censored data. Here we illustrate our method with an application to a formal goodness-of-fit test and a graphical check for the proportional hazards model for doubly censored data. We reanalyze a well-known AIDS data set.

Acquired Immunodeficiency Syndrome↗

Marginal analysis of incomplete longitudinal binary data: a cautionary note on LOCF imputation.

In recent years there has been considerable research devoted to the development of methods for the analysis of incomplete data in longitudinal studies. Despite these advances, the methods used in practice have changed relatively little, particularly in the reporting of pharmaceutical trials. In this setting, perhaps the most widely adopted strategy for dealing with incomplete longitudinal data is imputation by the "last observation carried forward" (LOCF) approach, in which values for missing responses are imputed using observations from the most recently completed assessment. We examine the asymptotic and empirical bias, the empirical type I error rate, and the empirical coverage probability associated with estimators and tests of treatment effect based on the LOCF imputation strategy. We consider a setting involving longitudinal binary data with longitudinal analyses based on generalized estimating equations, and an analysis based simply on the response at the end of the scheduled follow-up. We find that for both of these approaches, imputation by LOCF can lead to substantial biases in estimators of treatment effects, the type I error rates of associated tests can be greatly inflated, and the coverage probability can be far from the nominal level. Alternative analyses based on all available data lead to estimators with comparatively small bias, and inverse probability weighted analyses yield consistent estimators subject to correct specification of the missing data process. We illustrate the differences between various methods of dealing with drop-outs using data from a study of smoking behavior.

Adolescent↗

Multiple imputation for model checking: completed-data plots with missing and latent data.

In problems with missing or latent data, a standard approach is to first impute the unobserved data, then perform all statistical analyses on the completed dataset--corresponding to the observed data and imputed unobserved data--using standard procedures for complete-data inference. Here, we extend this approach to model checking by demonstrating the advantages of the use of completed-data model diagnostics on imputed completed datasets. The approach is set in the theoretical framework of Bayesian posterior predictive checks (but, as with missing-data imputation, our methods of missing-data model checking can also be interpreted as "predictive inference" in a non-Bayesian context). We consider the graphical diagnostics within this framework. Advantages of the completed-data approach include: (1) One can often check model fit in terms of quantities that are of key substantive interest in a natural way, which is not always possible using observed data alone. (2) In problems with missing data, checks may be devised that do not require to model the missingness or inclusion mechanism; the latter is useful for the analysis of ignorable but unknown data collection mechanisms, such as are often assumed in the analysis of sample surveys and observational studies. (3) In many problems with latent data, it is possible to check qualitative features of the model (for example, independence of two variables) that can be naturally formalized with the help of the latent data. We illustrate with several applied examples.

Animals↗

The New Zealand Socio-economic Index of Occupational Status: methodological revision and imputation for missing data.

OBJECTIVES: To revise and update the New Zealand Socio-economic Index (NZSEI) in the light of methodological issues in its construction, and to develop an imputation method for use where occupational information is not available. METHODS: Data were drawn from the following New Zealand national surveys: 1996 Population Census; 1996/97 and 1997/98 Household Economic Surveys; 1996/97 Household Health Survey. Three sets of statistical analyses were applied: alternating least squares to generate socio-economic scores; cluster and discriminant function analyses to identify cut-points; and regression and logistic regression to develop and test imputation methods. RESULTS: Socio-economic scores for the full-time workforce in 1996 showed a different distribution, but much the same occupational ordering, as in 1991. The introduction of part-time workers and income adjustment multipliers for self-employed workers significantly affected scores for management and agricultural titles. The application of cluster and discriminant function analyses generated six groupings that were relatively distinct occupationally. An imputation method based on an averaging of scores within age/qualification categories was found to achieve acceptable results. CONCLUSIONS: Methodological improvements in the construction of the NZSEI have enhanced its empirical robustness, while a simple imputation technique has widened the potential application of the scale.

Data Collection↗

Differences in mail and telephone responses to self-rated health: use of multiple imputation in correcting for response bias.

OBJECTIVES: To estimate differences in self-rated health by mode of administration and to assess the value of multiple imputation to make self-rated health comparable for telephone and mail. METHODS: In 1996, Survey 1 of the Australian Longitudinal Study on Women's Health was answered by mail. In 1998, 706 and 11,595 mid-age women answered Survey 2 by telephone and mail respectively. Self-rated health was measured by the physical and mental health scores of the SF-36. Mean change in SF-36 scores between Surveys 1 and 2 were compared for telephone and mail respondents to Survey 2, before and after adjustment for sociodemographic and health characteristics. Missing values and SF-36 scores for telephone respondents at Survey 2 were imputed from SF-36 mail responses and telephone and mail responses to sociodemographic and health questions. RESULTS: At Survey 2, self-rated health improved for telephone respondents but not mail respondents. After adjustment, mean changes in physical health and mental health scores remained higher (0.4 and 1.6 respectively) for telephone respondents compared with mail respondents (-1.2 and 0.1 respectively). Multiple imputation yielded adjusted changes in SF-36 scores that were similar for telephone and mail respondents. CONCLUSIONS AND IMPLICATIONS: The effect of mode of administration on the change in mental health is important given that a difference of two points in SF-36 scores is accepted as clinically meaningful. Health evaluators should be aware of and adjust for the effects of mode of administration on self-rated health. Multiple imputation is one method that may be used to adjust SF-36 scores for mode of administration bias.

Attitude to Health↗

[Imputability of drug side effects].

By imputability it is meant the assessment of the probable responsability of a drug in the development of undesirable effect. Its principle is based on the evaluation of some criteria derived from observation (intrinsic imputability) or relevant literature (extrinsic imputability). Around ten imputability methods have so far been reported; although they are not comparable, personnel differences in evaluation may be avoided.

Drug-Related Side Effects and Adverse Reactions↗

Sample design, sampling weights, imputation, and variance estimation in the 1995 National Survey of Family Growth.

OBJECTIVES: Cycle 5 of the National Survey of Family Growth (NSFG) was conducted by the National Center for Health Statistics (NCHS) in 1995. The NSFG collects data on pregnancy, childbearing, and women's health from a national sample of women 15-44 years of age. This report describes how the sample was designed, shows response rates for various subgroups of women, describes how the sampling weights were computed to make national estimates possible, shows how missing data were imputed for a limited set of key variables, and describes the proper ways to estimate sampling errors from the NSFG. The report includes both nontechnical summaries for readers who need only general information and more technical detail for readers who need an in-depth understanding of these topics. METHODS: The 1995 NSFG was based on a national probability sample of women 15-44 years of age in the United States and was drawn from 14,000 households interviewed in the 1993 National Health Interview Survey (NHIS). Of the 13,795 women eligible for the NSFG, 10,847 (79 percent) gave complete interviews. RESULTS: This report recommends using weighted data for analysis and a software package that will estimate sampling errors from complex samples (for example, SUDAAN or comparable software). The rate of missing data in the 1995 NSFG was very low. However, missing data were imputed for 315 key variables, called "recodes." Of the 315 recodes defined for Cycle 5, 271 variables had missing data on less than 1 percent of the cases; only 44 had 1 percent or more with missing data. These missing values were imputed for all of these 315 variables. The imputation procedures are described in this report.

Adolescent↗

The relationship between hot-deck multiple imputation and weighted likelihood.

Hot-deck imputation is an intuitively simple and popular method of accommodating incomplete data. Users of the method will often use the usual multiple imputation variance estimator which is not appropriate in this case. However, no variance expression has yet been derived for this easily implemented method applied to missing covariates in regression models. The simple hot-deck method is in fact asymptotically equivalent to the mean-score method for the estimation of a regression model parameter, so that hot-deck can be understood in the context of likelihood methods. Both of these methods accommodate data where missingness may depend on the observed variables but not on the unobserved value of the incomplete covariate, that is, missing at random (MAR). The asymptotic properties of hot-deck are derived here for the case where the fully observed variables are categorical, though the incomplete covariate(s) may be continuous. Simulation studies indicate that the two methods compare well in small samples and for small numbers of imputations. Current users of hot-deck may now conduct their analysis using mean-score, which is a weighted likelihood method and can thus be implemented by a single pass through the data using any standard package which accommodates weighted regression models. Valid inference is now straightforward using the variance expression provided here. The equivalence of mean-score and hot-deck is illustrated using three clinical data sets where an important covariate is missing for a large number of study subjects.

Angioplasty, Balloon, Coronary↗

Multiple imputation of missing blood pressure covariates in survival analysis.

This paper studies a non-response problem in survival analysis where the occurrence of missing data in the risk factor is related to mortality. In a study to determine the influence of blood pressure on survival in the very old (85+ years), blood pressure measurements are missing in about 12.5 per cent of the sample. The available data suggest that the process that created the missing data depends jointly on survival and the unknown blood pressure, thereby distorting the relation of interest. Multiple imputation is used to impute missing blood pressure and then analyse the data under a variety of non-response models. One special modelling problem is treated in detail; the construction of a predictive model for drawing imputations if the number of variables is large. Risk estimates for these data appear robust to even large departures from the simplest non-response model, and are similar to those derived under deletion of the incomplete records.

Aged↗

Linear regression for bivariate censored data via multiple imputation.

Bivariate survival data arise, for example, in twin studies and studies of both eyes or ears of the same individual. Often it is of interest to regress the survival times on a set of predictors. In this paper we extend Wei and Tanner's multiple imputation approach for linear regression with univariate censored data to bivariate censored data. We formulate a class of censored bivariate linear regression methods by iterating between the following two steps: 1. the data is augmented by imputing survival times for censored observations; 2. a linear model is fit to the imputed complete data. We consider three different methods to implement these two steps. In particular, the marginal (independence) approach ignores the possible correlation between two survival times when estimating the regression coefficient. To improve the efficiency, we propose two methods that account for the correlation between the survival times. First, we improve the efficiency by using generalized least squares regression in step 2. Second, instead of generating data from an estimate of the marginal distribution we generate data from a bivariate log-spline density estimate in step 1. Through simulation studies we find that the performance of the two methods that take the dependence into account is close and that they are both more efficient than the marginal approach. The methods are applied to a data set from an otitis media clinical trial.

Anti-Bacterial Agents↗

A two-sample test with interval censored data via multiple imputation.

Interval censored data arise naturally in large scale panel studies where subjects can only be followed periodically and the event of interest can only be recorded as having occurred between two examination times. In this paper we consider the problem of comparing two interval-censored samples. We propose to impute exact failure times from interval-censored observations to obtain right censored data, then apply existing techniques, such as Harrington and Fleming's G(rho) tests to imputed right censored data. To appropriately account for variability, a multiple imputation algorithm based on the approximate Bayesian bootstrap (ABB) is discussed. Through simulation studies we find that it performs well. The advantage of our proposal is its simplicity to implement and adaptability to incorporate many existing two-sample comparison techniques for right censored data. The method is illustrated by reanalysing the Breast Cosmesis Study data set.

Algorithms↗

Imputation strategies for missing data in a school-based multi-centre study: the Pathways study.

Pathways is a multi-centre school-based trial sponsored by the National Heart, Lung, and Blood Institute testing the efficacy of an obesity prevention intervention in American Indian children. During the study's protocol development, we prepared an analysis plan that accounted for missing data. In this paper, we present a case study of the process we used to decide upon the final analysis plan. The primary endpoint of the Pathways study is a comparison of per cent body fat between treatment and usual care groups at the end of a three-year intervention. Other studies on children and Native Americans have had moderate to large amounts of missing data. As a result we were concerned that missing data in Pathways would affect the type I error rate and power of the test of our primary endpoint. We present results from our evaluation of three alternative procedures in this paper. The first is a multiple imputation procedure in which we replace missing values with resampled values from the observed data. The second is based on the Wilcoxon rank sum test; missing data in the intervention group receive the worst ranks. In the third, we use a multiple imputation procedure and replace missing values with predicted values from a regression equation with the coefficients estimated from observed follow-up data and baseline values. We found that the multiple imputation procedure that replaces missing values with predicted values had the best properties of the procedures we considered. The results from our simulation study showed that, for missing data patterns that are relevant to the Pathways study, this procedure has high power and maintains the type I error rate. Published in 2001 by John Wiley & Sons, Ltd.

Analysis of Variance↗

Multiple imputation to estimate the association between eyes in disease progression with interval-censored data.

In many ophthalmologic studies, progression of diseases such as diabetic retinopathy, age-related maculopathy, cataract, and glaucoma is only noted when each eye is examined at intervals that commonly vary between subjects. Such data are often analysed using continuous time survival methods with observed progression assumed to occur at the end of the interval. Tied times of progression can lead to substantial bias in estimation of the association between progression in right and left eyes. We describe a multiple imputation strategy to create multiple data sets without ties, based on drawing interval-censored progression times from a parametric gamma frailty model that accounts for continuous and discrete covariates. We illustrate the method with data from 478 patients with insulin-dependent diabetes mellitus who were followed for progression of diabetic retinopathy in the Sorbinil Retinopathy Trial. Resolution of tied failure times allows for valid estimation of the hazard of progression in one eye given the progression status of the other eye. A simulation study suggests that the method performs well. Results highlight the advantage of multiple imputation that data imputed under one model can be analysed under several alternative models.

Adolescent↗

A new classification rule for incomplete doubly multivariate data using mixed effects model with performance comparisons on the imputed data.

A mixed effects model, enhanced by a Kronecker product structure for the residual variance-covariance matrix, is used in conjunction with a discriminant analysis technique, to devise a new statistical classification method on incomplete doubly multivariate data. The proposed method is efficient in small scale clinical trials that use relatively few patients. The new classification method is also applied to multiply imputed data sets. The misclassification error rates (MERs) are compared in order to investigate the effectiveness of the new classification rule on an incomplete data set. The classification method is applied to a real data set. The error rates on the incomplete data set are found to be much less than the median error rate on the multiply imputed data sets. Non-parametric methods, such as kernel method and k-nearest neighbourhood method, are also applied to multiply imputed data sets. Results illustrating the advantages of the new classification method over classic non-parametric classification methods are presented.

Bone Density↗

Multiple imputation for correcting verification bias.

In the case in which all subjects are screened using a common test and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, we are actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt multiple imputation framework for verification bias problem. We propose several imputation procedures for this problem and compare different methods of estimation. We show that our imputation methods are better than the existing methods with regard to nominal coverage and confidence interval length.

Bias↗

Multiple imputation for the comparison of two screening tests in two-phase Alzheimer studies.

Two-phase designs are common in epidemiological studies of dementia, and especially in Alzheimer research. In the first phase, all subjects are screened using a common screening test(s), while in the second phase, only a subset of these subjects is tested using a more definitive verification assessment, i.e. golden standard test. When comparing the accuracy of two screening tests in a two-phase study of dementia, inferences are commonly made using only the verified sample. It is well documented that in that case, there is a risk for bias, called verification bias. When the two screening tests have only two values (e.g. positive and negative) and we are trying to estimate the differences in sensitivities and specificities of the tests, one is actually estimating a confidence interval for differences of binomial proportions. Estimating this difference is not trivial even with complete data. It is well documented that it is a tricky task. In this paper, we suggest ways to apply imputation procedures in order to correct the verification bias. This procedure allows us to use well-established complete-data methods to deal with the difficulty of the estimation of the difference of two binomial proportions in addition to dealing with incomplete data. We compare different methods of estimation and evaluate the use of multiple imputation in this case. Our simulation results show that the use of multiple imputation is superior to other commonly used methods. We demonstrate our finding using Alzheimer data.

Aged↗

Partial imputation approach to analysis of repeated measurements with dependent drop-outs.

In clinical trials repeated measurements of a response variable are usually taken at prespecified time-points to compare the treatment effects. However, the comparison of treatment effects is often complicated by missing data caused by the withdrawal of some patients before the end of the study (that is, drop-outs). When the drop-out process depends on the response variable of interest, ignoring missing data may lead to biased comparison of the treatment effect. In this paper, conditions for ignoring the dependent missingness are investigated and a new approach using the usual testing procedure based on data with partial carrying-forward imputation is proposed. The proposed approach is conceptually and practically simple, and is motivated by making incremental improvement on the familiar 'all available data' (AAD) approach and the 'last value carrying forward' (LVCF) approach, which are commonly used in data analysis with drop-outs by practitioners. It is also compared favourably to the mixed-effect model approach with dependent drop-outs. Simulations and real data are used to evaluate and illustrate statistical properties of the proposed approach. The principle of the proposed approach can also be extended to using other imputation methods such as the multiple imputation.

Analysis of Variance↗