Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Robustness of a multivariate normal approximation for imputation of incomplete binary data.

Multiple imputation has become easier to perform with the advent of several software packages that provide imputations under a multivariate normal model, but imputation of missing binary data remains an important practical problem. Here, we explore three alternative methods for converting a multivariate normal imputed value into a binary imputed value: (1) simple rounding of the imputed value to the nearer of 0 or 1, (2) a Bernoulli draw based on a 'coin flip' where an imputed value between 0 and 1 is treated as the probability of drawing a 1, and (3) an adaptive rounding scheme where the cut-off value for determining whether to round to 0 or 1 is based on a normal approximation to the binomial distribution, making use of the marginal proportions of 0's and 1's on the variable. We perform simulation studies on a data set of 206,802 respondents to the California Healthy Kids Survey, where the fully observed data on 198,262 individuals defines the population, from which we repeatedly draw samples with missing data, impute, calculate statistics and confidence intervals, and compare bias and coverage against the true values. Frequently, we found satisfactory bias and coverage properties, suggesting that approaches such as these that are based on statistical approximations are preferable in applied research to either avoiding settings where missing data occur or relying on complete-case analyses. Considering both the occurrence and extent of deficits in coverage, we found that adaptive rounding provided the best performance.

Adolescent↗

A multiple imputation strategy for incomplete longitudinal data.

Longitudinal studies are commonly used to study processes of change. Because data are collected over time, missing data are pervasive in longitudinal studies, and complete ascertainment of all variables is rare. In this paper a new imputation strategy for completing longitudinal data sets is proposed. The proposed methodology makes use of shrinkage estimators for pooling information across geographic entities, and of model averaging for pooling predictions across different statistical models. Bayes factors are used to compute weights (probabilities) for a set of models considered to be reasonable for at least some of the units for which imputations must be produced, imputations are produced by draws from the predictive distributions of the missing data, and multiple imputations are used to better reflect selected sources of uncertainty in the imputation process. The imputation strategy is developed within the context of an application to completing incomplete longitudinal variables in the so-called Area Resource File. The proposed procedure is compared with several other imputation procedures in terms of inferences derived with the imputations, and the proposed methodology is demonstrated to provide valid estimates of model parameters when the completed data are analysed. Extensions to other missing data problems in longitudinal studies are straightforward so long as the missing data mechanism can be assumed to be ignorable.

Bayes Theorem↗

Imputation of missing longitudinal data: a comparison of methods.

BACKGROUND AND OBJECTIVES: Missing information is inevitable in longitudinal studies, and can result in biased estimates and a loss of power. One approach to this problem is to impute the missing data to yield a more complete data set. Our goal was to compare the performance of 14 methods of imputing missing data on depression, weight, cognitive functioning, and self-rated health in a longitudinal cohort of older adults. METHODS: We identified situations where a person had a known value following one or more missing values, and treated the known value as a "missing value." This "missing value" was imputed using each method and compared to the observed value. Methods were compared on the root mean square error, mean absolute deviation, bias, and relative variance of the estimates. RESULTS: Most imputation methods were biased toward estimating the "missing value" as too healthy, and most estimates had a variance that was too low. Imputed values based on a person's values before and after the "missing value" were superior to other methods, followed by imputations based on a person's values before the "missing value." Imputations that used no information specific to the person, such as using the sample mean, had the worst performance. CONCLUSIONS: We conclude that, in longitudinal studies where the overall trend is for worse health over time and where missing data can be assumed to be primarily related to worse health, missing data in a longitudinal sequence should be imputed from the available longitudinal data for that person.

Aged↗

The effect of imputation of exposure estimates on the association between fine particulate matter and mortality.

PURPOSE: The Harvard Six Cities Study (HSCS) found a small but significant association between daily PM2.5 and daily mortality count. The HSCS findings have been used as the basis for new EPA regulations, requiring lower levels of PM2.5. We feel that there are unresolved issues regarding the HSCS that should be fully evaluated prior to its findings being used as the basis of new regulation, including how the extent and method of imputing exposure data affect the association with daily mortality counts.METHODS: We examined the association between PM2.5 levels and daily mortality count, comparing the results from the HSCS methods with results based on an alternate imputation method, and with non-missing data.RESULTS: Overall, approximately 30% of the data points used in the HSCS were imputed. The method of imputation affected the association between particulate matter and mortality to a substantial degree in most of the cities. When the model using the HSCS method was compared to the model using the alternate method, in two areas the coefficients decreased substantially and lost significance. In two areas they changed little; in one area it rose substantially and became significant; and in one area it declined substantially but remained significant. When compared to the model based on the non-missing data, somewhat different patterns were observed. In both comparisons there were some large changes in the magnitude of the effect, but these were not consistent with the model used.CONCLUSIONS: We are concerned about the degree of data imputation and the effect that the method of imputation has on the association between particulate matter levels and mortality. In the case of the HSCS it appears that the imputed data are more strongly associated with the outcome than other methods of imputation and than the non-missing data. The reasons for these observations are not readily apparent, but the differences in effect should be explored and explained.

Journal Article↗

Usefulness of imputation for the analysis of incomplete otoneurologic data.

The usefulness of imputation in the treatment of missing values of an otoneurologic database for the discriminant analysis was evaluated on the basis of the agreement of imputed values and the analysis results. The data consisted of six patient groups with vertigo (N=564). There were 38 variables and 11% of the data was missing. Missing values were filled in with the means, regression and Expectation-Maximisation (EM) imputation methods and a random imputation method provided the baseline results. Means, regression and EM methods agreed on 41-42% of the imputed missing values. The level of agreement between these and the random method was 20-22%. Despite the moderate agreement between the means, regression and EM methods, the discriminant functions were similar and accurate (prediction accuracy 83-99%). The discriminant functions obtained from the randomly imputed data were also accurate having prediction accuracy 88-97%. Imputation seems to be a useful method for treating the missing data in this database. However, a lot of data was missing in otoneurologic tests, which are likely to be of less importance in the diagnosis of vertiginous patients. Consequently, the disagreement of the methods did not affect clearly the discriminant analysis, and, therefore, future research requires more complete data and advanced imputation methods.

Data Collection↗

Multiple imputation of baseline data in the cardiovascular health study.

Most epidemiologic studies will encounter missing covariate data. Software packages typically used for analyzing data delete any cases with a missing covariate to perform a complete case analysis. The deletion of cases complicates variable selection when different variables are missing on different cases, reduces power, and creates the potential for bias in the resulting estimates. Recently, software has become available for producing multiple imputations of missing data that account for the between-imputation variability. The implementation of the software to impute missing baseline data in the setting of the Cardiovascular Health Study, a large, observational study, is described. Results of exploratory analyses using the imputed data were largely consistent with results using only complete cases, even in a situation where one third of the cases were excluded from the complete case analysis. There were few differences in the exploratory results across three imputations, and the combined results from the multiple imputations were very similar to results from a single imputation. An increase in power was evident and variable selection simplified when using the imputed data sets.

Aged↗

Comparison of imputation and modelling methods in the analysis of a physical activity trial with missing outcomes.

BACKGROUND: Longitudinal studies almost always have some individuals with missing outcomes. Inappropriate handling of the missing data in the analysis can result in misleading conclusions. Here we review a wide range of methods to handle missing outcomes in single and repeated measures data and discuss which methods are most appropriate. METHODS: Using data from a randomized controlled trial to compare two interventions for increasing physical activity, we compare complete-case analysis; ad hoc imputation techniques such as last observation carried forward and worst-case; model-based imputation; longitudinal models with random effects; and recently proposed joint models for repeated measures data and non-ignorable dropout. RESULTS: Estimated intervention effects from ad hoc imputation methods vary widely. Standard multiple imputation and longitudinal modelling agree closely, as they should. Modifying the modelling method to allow for non-ignorable dropout had little effect on estimated intervention effects, but imputing using a common imputation model in both groups gave more conservative results. CONCLUSIONS: Results from ad hoc imputation methods should be avoided in favour of methods with more plausible assumptions although they may be computationally more complex. Although standard multiple imputation methods and longitudinal modelling methods are equivalent for estimating the treatment effect, the two approaches suggest different ways of relaxing the assumptions, and the choice between them depends on contextual knowledge.

Bias↗

Survival analysis using auxiliary variables via multiple imputation, with application to AIDS clinical trial data.

We develop an approach, based on multiple imputation, to using auxiliary variables to recover information from censored observations in survival analysis. We apply the approach to data from an AIDS clinical trial comparing ZDV and placebo, in which CD4 count is the time-dependent auxiliary variable. To facilitate imputation, a joint model is developed for the data, which includes a hierarchical change-point model for CD4 counts and a time-dependent proportional hazards model for the time to AIDS. Markov chain Monte Carlo methods are used to multiply impute event times for censored cases. The augmented data are then analyzed and the results combined using standard multiple-imputation techniques. A comparison of our multiple-imputation approach to simply analyzing the observed data indicates that multiple imputation leads to a small change in the estimated effect of ZDV and smaller estimated standard errors. A sensitivity analysis suggests that the qualitative findings are reproducible under a variety of imputation models. A simulation study indicates that improved efficiency over standard analyses and partial corrections for dependent censoring can result. An issue that arises with our approach, however, is whether the analysis of primary interest and the imputation model are compatible.

Acquired Immunodeficiency Syndrome↗

The validity of using multiple imputation for missing out-of-hospital data in a state trauma registry.

OBJECTIVES: To assess 1) the agreement of multiply imputed out-of-hospital values previously missing in a state trauma registry compared with known ambulance values and 2) the potential impact of using multiple imputation versus a commonly used method for handling missing data (i.e., complete case analysis) in a typical multivariable injury analysis. METHODS: This was a retrospective cohort analysis. Multiply imputed out-of-hospital data from 1998 to 2003 for four variables (intubation attempt, Glasgow Coma Scale score, systolic blood pressure, and respiratory rate) were compared with known values from probabilistically linked ambulance records using measures of agreement (kappa, weighted kappa, and Bland-Altman plots). Ambulance values were assumed to represent the "true" values for all analyses. A hypothetical multivariable regression model was used to demonstrate the impact (i.e., bias and precision of model results) of handling missing out-of-hospital data with multiple imputation versus complete case analysis. RESULTS: A total of 6,150 matched ambulance and trauma registry records were available for comparison. Multiply imputed values for the four out-of-hospital variables demonstrated fair to good agreement with known ambulance values. When included in typical multivariable analyses, multiple imputation increased precision and reduced bias compared with using complete case analysis for the same data set. CONCLUSIONS: Multiply imputed out-of-hospital values for intubation attempt, Glasgow Coma Scale score, systolic blood pressure, and respiratory rate have fair to good agreement with known ambulance values. Multiple imputation also increased precision and reduced bias compared with complete case analysis in a typical multivariable injury model, and it should be considered for studies using out-of-hospital data from a trauma registry, particularly when substantial portions of data are missing.

Ambulances↗

Imputation of missing data when measuring physical activity by accelerometry.

PURPOSE: We consider the issue of summarizing accelerometer activity count data accumulated over multiple days when the time interval in which the monitor is worn is not uniform for every subject on every day. The fact that counts are not being recorded during periods in which the monitor is not worn means that many common estimators of daily physical activity are biased downward. METHODS: Data from the Trial for Activity in Adolescent Girls (TAAG), a multicenter group-randomized trial to reduce the decline in physical activity among middle-school girls, were used to illustrate the problem of bias in estimation of physical activity due to missing accelerometer data. The effectiveness of two imputation procedures to reduce bias was investigated in a simulation experiment. Count data for an entire day, or a segment of the day were deleted at random or in an informative way with higher probability of missingness at upper levels of body mass index (BMI) and lower levels of physical activity. RESULTS: When data were deleted at random, estimates of activity computed from the observed data and those based on a data set in which the missing data have been imputed were equally unbiased; however, imputation estimates were more precise. When the data were deleted in a systematic fashion, the bias in estimated activity was lower using imputation procedures. Both imputation techniques, single imputation using the EM algorithm and multiple imputation (MI), performed similarly, with no significant differences in bias or precision. CONCLUSIONS: Researchers are encouraged to take advantage of software to implement missing value imputation, as estimates of activity are more precise and less biased in the presence of intermittent missing accelerometer data than those derived from an observed data analysis approach.

Acceleration↗

[Imputation of the date of HIV seroconversion in cohorts of haemophiliacs].

OBJECTIVES: To describe the methods used to impute HIV seroconversion date in the haemophiliac cohorts from GEMES project and to validate its use. METHOD: 632 haemophiliacs coming from three hemophilia units identified as HIV+ and 1.092 individuals coming from 5 project GEMES cohorts with a seroconversion window (time among test HIV and HIV+) less than 3 years where mid point (PM) was assumed as seroconversion date. For both groups, seroconversion date was imputed after estimating the probability distribution of seroconversion by means of the EM algorithm. Two imputation methods are used: one obtained from the expected value and the other from the geometric mean of 5 random samples. from the estimated distribution. Imputations have been validated in the non haemophiliacs cohorts comparing with the PM seroconversion date. Also AIDS free time and survival from the different seroconversion imputed dates were compared. RESULTS: Median seroconversion date is located in May of 1993 for the non haemophiliacs and in 1982 for the haemophiliacs. Not big differences are observed among the imputed seroconversion dates and the mid-point seroconversion date in the non-haemophiliac cohorts. Similar results are found for the haemophiliac cohorts. Also no differences are observed in the estimated AIDS-free time for both groups of cohorts. CONCLUSIONS: Geometric mean imputation from several random samples provides a good estimate of the HIV seroconversion date that can be used to estimate AIDS-free time and survival in haemophiliac cohorts where seroconversion date is ignored.

Cohort Studies↗

Missing data imputation in two phase III trials treating HIV1 infection.

In most longitudinal clinical trials, some patients drop out before the end of the planned follow-up, and, in order to allow an all-patient intent-to-treat analysis to be performed, it is common practice to use some method of imputation to estimate values for missing data. However, different imputation methods may provide different results, and it is essential to investigate the sensitivity of the analysis using different imputation rules. In our analysis of two trials of the new HIV1 fusion inhibitor enfuvirtide, we compared some standard methods of imputing and analyzing HIV1-RNA data with two novel alternatives, to check the robustness of the primary endpoint results. The standard methods were: (1) last-observation-carried-forward, (2) baseline carried forward, and (3) multiple imputation. These were compared with a nearest-neighbour hot-deck method, specifically proposed for imputation of missing HIV1-RNA data, and with a heuristic approach: censored regression analysis of the last-observation-carried-forward. To supplement this analysis of real clinical trial data, we investigated the performance of the same imputation methods on simulated datasets designed to cover a broader range of missing data patterns.

Algorithms↗

The influence of missing value imputation on detection of differentially expressed genes from microarray data.

MOTIVATION: Missing values are problematic for the analysis of microarray data. Imputation methods have been compared in terms of the similarity between imputed and true values in simulation experiments and not of their influence on the final analysis. The focus has been on missing at random, while entries are missing also not at random. RESULTS: We investigate the influence of imputation on the detection of differentially expressed genes from cDNA microarray data. We apply ANOVA for microarrays and SAM and look to the differentially expressed genes that are lost because of imputation. We show that this new measure provides useful information that the traditional root mean squared error cannot capture. We also show that the type of missingness matters: imputing 5% missing not at random has the same effect as imputing 10-30% missing at random. We propose a new method for imputation (LinImp), fitting a simple linear model for each channel separately, and compare it with the widely used KNNimpute method. For 10% missing at random, KNNimpute leads to twice as many lost differentially expressed genes as LinImp. AVAILABILITY: The R package for LinImp is available at http://folk.uio.no/idasch/imp.

Algorithms↗

Imputation of SF-12 health scores for respondents with partially missing data.

OBJECTIVE: To create an efficient imputation algorithm for imputing the SF-12 physical component summary (PCS) and mental component summary (MCS) scores when patients have one to eleven SF-12 items missing. STUDY SETTING: Primary data collection was performed between 1996 and 1998. STUDY DESIGN: Multi-pattern regression was conducted to impute the scores using only available SF-12 items (simple model), and then supplemented by demographics, smoking status and comorbidity (enhanced model) to increase the accuracy. A cut point of missing SF-12 items was determined for using the simple or the enhanced model. The algorithm was validated through simulation. DATA COLLECTION: Thirty-thousand-three-hundred and eight patients from 63 physician groups were surveyed for a quality of care study in 1996, which collected the SF-12 and other information. The patients were classified as "chronic" patients if they reported that they had diabetes, heart disease, asthma/chronic obstructive pulmonary disease, or low back pain. A follow-up survey was conducted in 1998. PRINCIPAL FINDINGS: Thirty-one percent of the patients missed at least one SF-12 item. Means of variance of prediction and standard errors of the mean imputed scores increased with the number of missing SF-12 items. Correlations between the observed and the imputed scores derived from the enhanced models were consistently higher than those derived from the simple model and the increments were significant for patients with > or =6 missing SF-12 items (p<.03). CONCLUSION: Missing SF-12 items are prevalent and lead to reduced analytical power. Regression-based multi-pattern imputation using the available SF-12 items is efficient and can produce good estimates of the scores. The enhancement from the additional patient information can significantly improve the accuracy of the imputed scores for patients with > or =6 items missing, leading to estimated scores that are as accurate as that of patients with <6 missing items.

Adult↗

Multiple imputation of missing genotype data for unrelated individuals.

The objective of this study was to investigate the performance of multiple imputation of missing genotype data for unrelated individuals using the polytomous logistic regression model, focusing on different missingness mechanisms, percentages of missing data, and imputation models. A complete dataset of 581 individuals, each analysed for eight biallelic polymorphisms and the quantitative phenotype HDL-C, was used. From this dataset one hundred replicates with missing data were created, in different ways for different scenarios. The performance was assessed by comparing the mean bias in parameter estimates, the root mean squared standard errors, and the genotype-imputation error rates. Overall, the mean bias was small in all scenarios, and in most scenarios the mean did not differ significantly from 'no bias'. Including polymorphisms that are highly correlated in the imputation model reduced the genotype-imputation error rate and increased precision of the parameter estimates. The method works well for data that are missing completely at random, and for data that are missing at random. In conclusion, our results indicate that multiple imputation with the polytomous logistic regression model can be used for association studies to deal with the problem of missing genotype data, when attention is paid to the imputation model and the percentage of missing data.

Cholesterol, HDL↗

Improving missing value imputation of microarray data by using spot quality weights.

BACKGROUND: Microarray technology has become popular for gene expression profiling, and many analysis tools have been developed for data interpretation. Most of these tools require complete data, but measurement values are often missing A way to overcome the problem of incomplete data is to impute the missing data before analysis. Many imputation methods have been suggested, some naïve and other more sophisticated taking into account correlation in data. However, these methods are binary in the sense that each spot is considered either missing or present. Hence, they are depending on a cutoff separating poor spots from good spots. We suggest a different approach in which a continuous spot quality weight is built into the imputation methods, allowing for smooth imputations of all spots to larger or lesser degree. RESULTS: We assessed several imputation methods on three data sets containing replicate measurements, and found that weighted methods performed better than non-weighted methods. Of the compared methods, best performance and robustness were achieved with the weighted nearest neighbours method (WeNNI), in which both spot quality and correlations between genes were included in the imputation. CONCLUSION: Including a measure of spot quality improves the accuracy of the missing value imputation. WeNNI, the proposed method is more accurate and less sensitive to parameters than the widely used kNNimpute and LSimpute algorithms.

Algorithms↗

Multiple imputation technique applied to appropriateness ratings in cataract surgery.

Missing data such as appropriateness ratings in clinical research are a common problem and this often yields a biased result. This paper aims to introduce the multiple imputation method to handle missing data in clinical research and to suggest that the multiple imputation technique can give more accurate estimates than those of a complete-case analysis. The idea of multiple imputation is that each missing value is replaced with more than one plausible value. The appropriateness method was developed as a pragmatic solution to problem of trying to assess "appropriate" surgical and medical procedures for patients. Cataract surgery was selected as one of four procedures that were evaluated as a part of the Clinical Appropriateness Initiative. We created mild to high missing rates of 10%, 30% and 50% and compared the performance of logistic regression in cataract surgery. We treated the coefficients in the original data as true parameters and compared them with the other results. In the mild missing rate (10%), the deviation from the true coefficients was quite small and ignorable. After removing the missing data, the complete-case analysis did not reveal any serious bias. However, as the missing rate increased, the bias was not ignorable and it distorted the result. This simulation study suggests that a multiple imputation technique can give more accurate estimates than those of a complete-case analysis, especially for moderate to high missing rates (30 - 50%). In addition, the multiple imputation technique yields better accuracy than a single imputation technique. Therefore, multiple imputation is useful and efficient for a situation in clinical research where there is large amounts of missing data.

Cataract Extraction↗

Multivariate outlier detection applied to multiply imputed laboratory data.

In clinical laboratory safety data, multivariate outlier detection methods may highlight a patient whose laboratory measurements do not follow the same pattern of relationships as the majority of patients, although their individual measurements are not found to be outlying when considered one at a time. Missing data problems are often dealt with by imputing a single value as an estimate of the missing value. The completed data set may then be analysed using traditional methods. A disadvantage of using single imputation is the underestimation of variability, with a corresponding distortion of power in hypothesis testing. Multiple imputation methods attempt to overcome this problem, and in this paper a study is described which considers the application of multivariate outlier detection methods to multiply imputed clinical laboratory safety data sets. Three different proportions of missing data are generated in laboratory data sets of dimensions 4, 7, 12 and 30, and a comparison of eight multiple imputation methods is carried out. Two outlier detection techniques, Mahalanobis distance and generalized principal component analysis, are applied to the multiply imputed data sets, and their performances are discussed. Measures are introduced for assessing the accuracy of the missing data results, depending on which method of analysis is used.

Algorithms↗