Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Multiple imputation for threshold-crossing data with interval censoring.

Medical statistics often involve measurements of the time when a variable crosses a threshold value. The time to threshold crossing may be the outcome variable in a survival analysis, or a time-dependent covariate in the analysis of a subsequent event. This paper presents new methods for analysing threshold-crossing data that are interval censored in that the time of threshold crossing is known only within a specified interval. Such data typically arise in event-history studies when the threshold is crossed at some time between data-collection points, such as visits to a clinic. We propose methods based on multiple imputation of the threshold-crossing time with use of models that take into account values recorded at the times of visits. We apply the methods to two real data sets, one involving hip replacements and the other on the prostate specific antigen (PSA) assay for prostate cancer. In addition, we compare our methods with the common practice of imputing the threshold-crossing time as the right endpoint of the interval. The two examples require different imputation models, but both lead to simple analyses of the multiply imputed data that automatically take into account variability due to imputation.

Data Interpretation, Statistical↗

A multiple imputation method for missing covariates in non-linear mixed-effects models with application to HIV dynamics.

We propose a three-step multiple imputation method, implemented by Gibbs sampler, for estimating parameters in non-linear mixed-effects models with missing covariates. Estimates obtained by the proposed multiple imputation method are compared to those obtained by the mean-value imputation method and the complete-case method through simulations. We find that the proposed multiple imputation method offers smaller biases and smaller mean-squared errors for the estimates of covariate coefficients compared to other two methods. We apply the three missing data methods to modelling HIV viral dynamics from an AIDS clinical trial. We believe that the results from the proposed multiple imputation method are more reliable than that from the other two commonly used methods.

Antiretroviral Therapy, Highly Active↗

Imputation of missing values in the case of a multiple item instrument measuring alcohol consumption.

Missing values in survey instruments are a common problem for survey researchers. It is aggravated in the case of instruments used to measure alcohol consumption: they usually consist of item batteries from which summary measures, such as grams of pure alcohol per day, are constructed, and a missing value (for example, quantity or frequency) in regard to a single item for only one of several beverages results in a missing summary measure across all of the beverages, though the values for the remaining items are known. The present paper examines different approaches to imputation of missing values, feasible with standard statistical software packages. Hot-deck imputation is shown to have certain advantages, but even single-value imputation (for example, median imputation) results in values that are comparable to those of the other three imputation methods.

Adult↗

Response and non-response to a quality-of-life question on sexual life: a case study of the simple mean imputation method.

We investigated the non-response rates to the question "I am satisfied with my sex life" in the Functional Assessment of Cancer Therapy-General questionnaire in Chinese (n = 769), Malay (n = 41) and Indian (n = 33) patients in Singapore, a multi-ethnic society whose residents are said to have a conservative sexual attitude. Non-response rates to the question were 44%, 22% and 24% in the three groups respectively. The rates were much higher than that reported previously in a US study (7%) and used in the associated simulation study of the simple mean imputation method. We further examined the Chinese respondents in detail. The odds of non-response and the scores among the responders were associated with several demographic and clinical characteristics. Using the checklist proposed by Fayers et al. [Stat Med 1998; 17: 679-696] to assess the data patterns, we found that the application of the simple mean imputation is questionable. We employed an alternative (multiple) imputation procedure that took into account covariates that predicted the odds of non-response and the observed response scores. We compared the analytic results based on different approaches to handling missing values, and found that analysis based on the simple mean imputation gave results similar to that based on multiply imputed data even in this quite extreme example.

Asia, Southeastern↗

Imputation of missing values is superior to complete case analysis and the missing-indicator method in multivariable diagnostic research: a clinical example.

BACKGROUND AND OBJECTIVES: To illustrate the effects of different methods for handling missing data--complete case analysis, missing-indicator method, single imputation of unconditional and conditional mean, and multiple imputation (MI)--in the context of multivariable diagnostic research aiming to identify potential predictors (test results) that independently contribute to the prediction of disease presence or absence. METHODS: We used data from 398 subjects from a prospective study on the diagnosis of pulmonary embolism. Various diagnostic predictors or tests had (varying percentages of) missing values. Per method of handling these missing values, we fitted a diagnostic prediction model using multivariable logistic regression analysis. RESULTS: The receiver operating characteristic curve area for all diagnostic models was above 0.75. The predictors in the final models based on the complete case analysis, and after using the missing-indicator method, were very different compared to the other models. The models based on MI did not differ much from the models derived after using single conditional and unconditional mean imputation. CONCLUSION: In multivariable diagnostic research complete case analysis and the use of the missing-indicator method should be avoided, even when data are missing completely at random. MI methods are known to be superior to single imputation methods. For our example study, the single imputation methods performed equally well, but this was most likely because of the low overall number of missing values.

Adult↗

The Soifua Manuia reference panel with 2,570 Samoan haplotypes improves genotype imputation quality among Samoans.

Genotype imputation is fundamental to association studies, and yet even gold standard panels like TOPMed are limited in the populations for which they yield good imputation. Specifically, Pacific Islanders are poorly represented in extant panels. To address this, we used whole-genome sequencing from 1,285 Samoan individuals combined with 1000 Genomes Project (1KGP) individuals to construct an imputation reference panel that better represents Pacific Islander, specifically Samoan, genetic variation. Here we show that this panel yielded up to two times more well-imputed (r2 ≥ 0.80) variants than TOPMed-R3 and 1KGP and was enriched for moderate and high impact variants. There was improved imputation accuracy across the minor allele frequency (MAF) spectrum; accuracy (r2) was greater for population-specific variants (high fixation index, FST) and those from larger haplotypes (high LD score). However, the gain in accuracy over TOPMed-R3 was largest for small haplotypes, reflecting the Samoan panel's ability to capture variation not well tagged by other panels.

Haplotypes↗

Treating missing data in a clinical neuropsychological dataset--data imputation.

Missing data frequently reduce the applicability of clinically collected data in research requiring multivariate statistics. In data imputation, missing values are replaced by predicted values obtained from models based on auxiliary information. Our aim was to complete a clinical child neuropsychological data set containing 5.2% of missing observations. This was to be used in research requiring multivariate statistics. We compared four data imputation methods by artificially deleting some data. A real-donor imputation method which preserved the parameter estimates and which predicted the observed values with acceptable accuracy was used to complete the data set. In addressing the lack of studies with regard to treatment of missing data in neuropsychological data sets, this study presents information on the outcomes of applying data imputation methods to such data. The imputation modeling described can be applied to a variety of clinical neuropsychological data sets.

Child↗

Imputing cross-sectional missing data: comparison of common techniques.

OBJECTIVE: Increasing awareness of how missing data affects the analysis of clinical and public health interventions has led to increasing numbers of missing data procedures. There is little advice regarding which procedures should be selected under different circumstances. This paper compares six popular procedures: listwise deletion, item mean substitution, person mean substitution at two levels, regression imputation and hot deck imputation. METHOD: Using a complete dataset, each was examined under a variety of sample sizes and differing levels of missing data. The criteria were the true t-values for the entire sample. RESULTS: The results suggest important differences. If missing data are from a scale where about half the items are present, hot deck imputation or person mean substitution are best. Because person mean substitution is computationally simpler, similar in its efficiency, advocated by other researchers and more likely to be an option on statistical software packages, it is the method of choice. If the missing data are from a scale where more than half the items are missing, or with single-item measures, then hot deck imputation is recommended. The findings also showed that listwise deletion and item mean substitution performed poorly. CONCLUSIONS: Person mean and hot deck imputation are preferred. Since listwise deletion and item mean substitution performed poorly, yet are the most widely reported methods, the findings have broad implications.

Anxiety↗

The Parkinson's Disease Questionnaire (PDQ-39): evidence for a method of imputing missing data.

BACKGROUND: The Parkinson's Disease Questionnaire (PDQ-39) is the most widely used Parkinson's specific measure of health status. It is increasingly used in treatment trials, sometimes as a primary end-point, where any missing data can potentially cause difficulties in analyses. OBJECTIVES: The purpose of this article is to evaluate the Expectation Maximisation (EM) algorithm for the imputation of missing dimension scores on the 39-item PDQ-39. METHODS: A postal survey of patients diagnosed with Parkinson's disease (PD). A total of 1,372 patients were surveyed and 839 (61.15%) questionnaires returned completed or partially completed. Of these, complete PDQ data were available in 715 (85.22%) cases. Data were deleted from this complete dataset and a sub-set of 200 respondents from this dataset and then imputed using the EM algorithm; results were then compared to the dataset before data deletion. RESULTS: Results gained from imputation of data closely mirrored that of the complete dataset in each case. Descriptive statistics, mean scores and spread of scores were almost identical between original and imputed datasets. Furthermore, original and imputed datasets were highly correlated [intra-class correlation coefficient (ICC) = 0.93 or greater], and mean differences were small (+/-1.00). CONCLUSIONS: The results suggest that the use of EM for the PDQ-39 provides data that closely mirrors the original when this has been deliberately removed. Consequently, EM is likely to be appropriate for trials using the PDQ that contains missing data points.

Adult↗

Collateral missing value imputation: a new robust missing value estimation algorithm for microarray data.

MOTIVATION: Microarray data are used in a range of application areas in biology, although often it contains considerable numbers of missing values. These missing values can significantly affect subsequent statistical analysis and machine learning algorithms so there is a strong motivation to estimate these values as accurately as possible before using these algorithms. While many imputation algorithms have been proposed, more robust techniques need to be developed so that further analysis of biological data can be accurately undertaken. In this paper, an innovative missing value imputation algorithm called collateral missing value estimation (CMVE) is presented which uses multiple covariance-based imputation matrices for the final prediction of missing values. The matrices are computed and optimized using least square regression and linear programming methods. RESULTS: The new CMVE algorithm has been compared with existing estimation techniques including Bayesian principal component analysis imputation (BPCA), least square impute (LSImpute) and K-nearest neighbour (KNN). All these methods were rigorously tested to estimate missing values in three separate non-time series (ovarian cancer based) and one time series (yeast sporulation) dataset. Each method was quantitatively analyzed using the normalized root mean square (NRMS) error measure, covering a wide range of randomly introduced missing value probabilities from 0.01 to 0.2. Experiments were also undertaken on the yeast dataset, which comprised 1.7% actual missing values, to test the hypothesis that CMVE performed better not only for randomly occurring but also for a real distribution of missing values. The results confirmed that CMVE consistently demonstrated superior and robust estimation capability of missing values compared with other methods for both series types of data, for the same order of computational complexity. A concise theoretical framework has also been formulated to validate the improved performance of the CMVE algorithm. AVAILABILITY: The CMVE software is available upon request from the authors.

Algorithms↗

Imputing physical health status scores missing owing to mortality: results of a simulation comparing multiple techniques.

BACKGROUND: Having missing data complicates the statistical analysis of health-related quality-of-life (HRQOL) data and, depending on the extent and nature of missing data, can introduce significant bias in treatment comparisons. OBJECTIVE: We evaluated the bias associated with 4 different imputation methods for estimating physical health status (PHS) scores missing as a result of mortality. METHODS: A simulation study was conducted in which we systematically varied mortality rates from 0% to 30% and change in PHS scores from -20 to 20 on a 100-point scale for a 2-group clinical trial with follow-up over 18 months. The 4 imputation methods were last value carried forward (LVCF), arbitrary substitution (ARBSUB), empirical Bayes (BAYES), and within-subject modeling (WSMOD). Pseudo-root mean square residuals (RMSRs) and differences between true and estimated slopes were used to evaluate how well the imputation methods reproduced the true characteristics of the simulated population data. RESULTS: ARBSUB and BAYES methods have the smallest RMSRs compared with LVCF and WSMOD across all mortality rates. As the rate of missing data resulting from mortality increased, all imputation techniques deviated more from population data. The BAYES technique was best at reproducing group slopes in cases with differential mortality rates or when mortality rates exceeded 15%. WSMOD and LVCF significantly underestimated changes in PHS. CONCLUSIONS: The different imputation methods produced comparable results when there were few missing data. The BAYES approach most closely estimated true population differences and change in PHS regardless of missing data rates. These findings are limited to physical health and functioning measures.

Analysis of Variance↗

Handling missing data in nursing research with multiple imputation.

BACKGROUND: In the data analysis phase of research, missing values present a challenge to nurse investigators. Common approaches for addressing missing data generally include complete-case analysis, available-case analysis, and single-value imputation methods. These methods have been the subject of increasing criticism with respect to their tendency to underestimate standard errors, overstate statistical significance, and introduce bias. OBJECTIVES: This article reviews the limitations of standard approaches for handling missing data, and suggests multiple imputation is a useful method for nursing research. METHOD: Secondary analysis was conducted to examine the effect of a public policy on the health of women using a data set that had a large degree and complex patterns of missing data. DISCUSSION: In the example, accommodation of the incomplete data was critical to making valid inferences; however, complete-case, available-case, or single imputation could not be defended as an adequate method for dealing with the missing data patterns. Alternative methods for dealing with incomplete data were sought, and a multiple imputation approach was selected given the missing data pattern. Nurse researchers confronting similar complex patterns of missing data may find multiple imputation a useful procedure for conducting data analysis and avoiding the bias associated with other methods of handling missing data.

Adult↗

Using multiple imputation for analysis of incomplete data in clinical research.

BACKGROUND: Sample loss and missing data are inevitable in multivariate and longitudinal research. Ad hoc approaches such as analysis of incomplete data or substituting the group mean for missing data, while common, may unnecessarily reduce statistical power and threaten study validity. Multiple imputation for missing data is a newly accessible, methodologically rigorous approach to dealing with the problem of missing data. APPROACH: To (a) discuss the problem of missing data in clinical research, and (b) describe the technique of multiple imputation. A case of analysis of multivariate psychosocial data is presented to illustrate the practice of multiple imputation. RESULTS: The advantages of multiple imputation are it (a) results in unbiased estimates, providing more validity than ad hoc approaches to missing data; (b) uses all available data, preserving sample size and statistical power; (c) may be used with standard statistical software; and, (d) results are readily interpreted. DISCUSSION: Accessible, user-friendly computer programs are available to perform multiple imputation for missing data making ad hoc approaches to missing data obsolete.

Data Interpretation, Statistical↗

Imputation strategies for blood pressure data nonignorably missing due to medication use.

BACKGROUND: Underlying or untreated blood pressure (BP) is often an outcome of interest, but is unobservable when study participants are on anti-hypertensive medications. Untreated levels are not missing at random but would be higher among those on such medication. In such cases, standard methods of analysis may lead to bias. PURPOSE: BPs obtained at the private physician's office (out-of-study BPs) at the time of prescription of anti-hypertensive medications were available from Phase II of the Trials of Hypertension Prevention (TOHP) and were used to adjust for the potential bias. METHODS: Observed out-of-study BPs were used to estimate the conditional expectation and variance of the unobserved unmedicated study BPs. For those with no physician data, imputation from bootstrap samples of out-of-study BPs was used. An iterative method based on the EM algorithm was used to estimate the unknown study parameters in a random-effects model using multiple imputations. This was compared to an alternative model for the out-of-study BPs based on a theoretical truncated normal distribution, and to standard analyses, including both multivariate repeated measures and last-observation-carried-forward (LOCF) analyses, using data from Phase II of TOHP. RESULTS: Differences between methods were seen in the decline in BP over time in the reference group, where the changes from baseline to 36 months were 3.0 in univariate analyses, 2.4 using LOCF, and 2.6 in the multivariate analysis, compared to 2.0 or 1.7 in the imputation analyses, depending on the number of physician visits. Estimated intervention effects tended to be slightly larger using the imputation methods. LIMITATIONS: out-of-study measures may not be available for other studies. CONCLUSIONS: Because the proposed strategy was based on an empirically observed distribution for out-of-study BP, fewer assumptions about the missing data were made. These data may be useful in suggesting imputation strategies for other studies.

Blood Pressure↗

Treatment of missing values with imputation for the analysis of otologic data.

Usefulness of imputation in the treatment of missing values in an otologic database was studied. Missing values were filled in with means (ME), regression (LR) and Expectation-Maximization (EM) imputation methods. A random imputation method (RA) provided baseline results. ME, LR and EM methods agreed on 41-42% of the imputed missing values. The level of agreement between these and RA method was 20-22%. Despite the moderate agreement, discriminant functions were similar and accurate (prediction accuracy 83-99%) for each diagnosis. A lot of data were missing in otoneurotologic tests which have less weight in the diagnosis of vertiginous patients. Consequently, the disagreement of the methods did not affect discriminant analysis. Inputation seems to be a useful method to treat missing data in this database, but future research requires more complete data and advanced imputation methods.

Data Collection↗

Analyses of public use decennial census data with multiply imputed industry and occupation codes.

"This paper gives a brief introduction to multiple imputation for handling non-response in surveys. We then describe a recently completed project in which multiple imputation was used to recalibrate industry and occupation codes in 1970 U.S. census public use samples to the 1980 standard. Using analyses of data from the project, we examine the utility of analysing a large data set having imputed values compared with analysing a small data set having true values, and we provide examples of the amount by which variability is underestimated by using just one imputation rather than multiple imputations."

Americas↗

Multiple imputation as a missing data machine.

This paper deals with problems concerning missing data in clinical databases. After signalling some shortcomings of popular solutions to incomplete data problems, we outline the concepts behind multiple imputation. Multiple imputation is a statistically sound method for handling incomplete data. Application of multiple imputation requires a lot of work and not every user is able to do this. A transparent implementation of multiple imputation is necessary. Such an implementation is possible in the HERMES medical workstation. A remaining problem is to find proper imputations.

Algorithms↗

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r = 0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans↗