Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Imputation of individual cancer cases to occupational causes.

OBJECTIVES: Many potential occupational causes of cancer have been documented. Imputation of an individual cancer to occupational or other causes is, however, difficult. A method based on the Bayes theorem is proposed for assessing causal relationships at the individual level. METHODS: Causality assessment, dealing with four types of persons defined by exposure and the occurrence of cancer, was linked with imputation, only dealing with persons who have cancer and were exposed. Imputation was then formulated using the Bayes theorem, relating epidemiologic information regarding causes, a patient's exposure history, and the posterior odds that the cancer was caused by a suspected occupational exposure. Data needed to apply a Bayesian method were defined in terms of relative risks, proportion of people exposed in populations, and the frequency of a positive relevant characteristic for persons without cancer. A relevant characteristic was defined using a formal consensus between experts. The method was then illustrated with cases of mesothelioma and lung cancer in possible relation to asbestos. RESULTS: Experts defined the relevant characteristics as being qualification of occupational exposure, intensity of exposure, latency, disease characteristics, and presence of causal agent in the body. Application to mesothelioma and lung cancer cases illustrated the potential usefulness of the method. CONCLUSIONS: The importance of occupational exposure in the formulation of imputation underscores the need for available and reliable data sources on occupational exposures. The proposed method could become a powerful tool for the expert assessment of causes of cancer cases, provided data become available in individual files and the literature.

Aged↗

Multiple imputation for missing data.

Missing data occur frequently in survey and longitudinal research. Incomplete data are problematic, particularly in the presence of substantial absent information or systematic nonresponse patterns. Listwise deletion and mean imputation are the most common techniques to reconcile missing data. However, more recent techniques may improve parameter estimates, standard errors, and test statistics. The purpose of this article is to review the problems associated with missing data, options for handling missing data, and recent multiple imputation methods. It informs researchers' decisions about whether to delete or impute missing responses and the method best suited to doing so. An empirical investigation of AIDS care data outcomes illustrates the process of multiple imputation.

Acquired Immunodeficiency Syndrome↗

From single-race reporting to multiple-race reporting: using imputation methods to bridge the transition.

In 1997, the Office of Management and Budget issued revised standards for the collection of race information within the Federal statistical system. One revision allows individuals to choose more than one race group when responding to Federal surveys and other Federal data collections. This paper explores methods that impute single-race categories for those who have given multiple-race responses. Such imputations would be useful when it is desired to conduct analyses involving only single-race categories, such as when trends over time are being examined by race group so that data collected under the old and new standards are being combined. The National Health Interview Survey has allowed multiple-race responses for several years, while also asking respondents to specify one race as their primary race. Exploratory analyses of data from the survey suggest that imputation methods that use demographic and contextual covariate information to predict primary race can have advantages with respect to lower bias and improved variance estimation compared to simpler methods discussed by the Office of Management and Budget. It also appears, however, that the relationships between primary race and covariates might be changing over time. Thus, caution should be exercised if an imputation model fitted to data from one time period is to be applied to data from another time period. Published in 2003 by John Wiley & Sons, Ltd.

Adolescent↗

Estimating the distribution of times from HIV seroconversion to AIDS using multiple imputation. Multicentre AIDS Cohort Study.

Multiple imputation is a model based technique for handling missing data problems. In this application we use the technique to estimate the distribution of times from HIV seroconversion to AIDS diagnosis with data from a cohort study of 4954 homosexual men with 4 years of follow-up. In this example the missing data are the dates of diagnosis with AIDS. The imputation procedure is performed in two stages. In the first stage, we estimate the residual AIDS-free time distribution as a function of covariates measured on the study participants with data provided by the participants who were seropositive at study entry. Specifically, we assume the residual AIDS-free times follow a log-normal regression model that depends on the covariates measured at enrolment on the seropositive participants. In the second stage we impute the date of AIDS diagnosis for the participants who seroconverted during the course of the study and are AIDS-free with use of the log-normal distribution estimated in the first stage and the covariates from each seroconverter's latest visit. The estimated proportions developing AIDS within 4 and within 7 years of seroconversion are 15 and 36 per cent respectively, with associated 95 per cent confidence intervals of (10, 21) and (26, 47) per cent. We discuss the Bayesian foundations of the multiple imputation technique and the statistical and scientific assumptions.

AIDS Serodiagnosis↗

Effects of mid-point imputation on the analysis of doubly censored data.

Doubly censored data arise in some cohort studies of the AIDS incubation period because the time of infection may be known only up to an interval defined by two successive screening tests for HIV antibody. A simple analytic approach is to impute the infection time by the mid-point of the interval and then apply standard survival techniques for right censored data. The objective of this paper is to investigate the statistical properties of such a mid-point imputation approach. We investigated the asymptotic bias of the Kaplan-Meier estimate, coverage probabilities of associated confidence intervals, bias in hazard ratio, and the size of the logrank test. We show that the statistical properties of mid-point imputation depend strongly on the underlying distributions of infection times and the incubation periods, and the width of the interval between screening tests. In the absence of treatment, the median incubation period of HIV infection is approximately 10 years, and we conclude that, for this situation, mid-point imputation is a reasonable procedure for interval widths of 2 years or less.

Bias↗

Analysis of the benefits of a Mediterranean diet in the GISSI-Prevenzione study: a case study in imputation of missing values from repeated measurements.

The problem of missing values has increasingly being recognized in epidemiology. New methods allow for the analysis of missing data that can provide valid estimates of epidemiological quantities of interest. The GISSI-Prevenzione study was aimed to reliably assess the long-term relationship between the consumption of foods typical of the Mediterranean diet and the risk of mortality amongst 11,323 Italians with prior myocardial infarction. Food intake frequencies were recorded repeatedly over the 4.5 years of follow-up and missing values affected each food variable at increasing rates over the course of the study. Comparisons were made between the results obtained from the analysis of the complete data and those obtained after imputing the missing data with simple imputation methods and with various implementations of the multiple imputation (MI) method. MI appeared to best address the issue of missing data on the food intake frequencies, preserving the observed distributions and relationships between variables whilst producing plausible estimates of variability. Given its theoretical properties and flexibility to different types of data, MI is more likely to provide valid estimates, compared to complete data analysis and imputation by simple methods, and is thus worthy of wider consideration amongst epidemiological researchers.

Confidence Intervals↗

Imputing missing standard deviations in meta-analyses can provide accurate results.

BACKGROUND AND OBJECTIVES: Many reports of randomized controlled trials (RCTs) fail to provide standard deviations (SDs) of their continuous outcome measures. Some meta-analysts substitute them by those reported in other studies, either from another meta-analysis or from other studies in the same meta-analysis. But the validity of such practices has never been empirically examined. METHODS: We compared the actual standardized mean difference (SMD) of individual RCTs and the meta-analytically pooled SMD of all RCTs against those based on the above-mentioned two imputation methods in two meta-analyses of antidepressants. RESULTS: Two meta-analyses included 39 RCTs of fluoxetine (n = 3,681) and 25 RCTs of amitriptyline (n = 1,832), which had actually reported means and SDs of the Hamilton Rating Scale for Depression. According to either of the two proposed imputation methods, the agreement between actual SMDs and imputed SMDs for individual RCTs was very good with ANOVA intraclass correlation coefficients between 0.61 and 0.97. The agreement between the actual pooled SMD and the imputed one was even better, with minimal differences in both their point estimates and 95% confidence intervals. CONCLUSION: For a systematic review where some of the identified trials do not report SDs, it appears safe to borrow SDs from other studies.

Amitriptyline↗

Using the outcome for imputation of missing predictor values was preferred.

BACKGROUND AND OBJECTIVE: Epidemiologic studies commonly estimate associations between predictors (risk factors) and outcome. Most software automatically exclude subjects with missing values. This commonly causes bias because missing values seldom occur completely at random (MCAR) but rather selectively based on other (observed) variables, missing at random (MAR). Multiple imputation (MI) of missing predictor values using all observed information including outcome is advocated to deal with selective missing values. This seems a self-fulfilling prophecy. METHODS: We tested this hypothesis using data from a study on diagnosis of pulmonary embolism. We selected five predictors of pulmonary embolism without missing values. Their regression coefficients and standard errors (SEs) estimated from the original sample were considered as "true" values. We assigned missing values to these predictors--both MCAR and MAR--and repeated this 1,000 times using simulations. Per simulation we multiple imputed the missing values without and with the outcome, and compared the regression coefficients and SEs to the truth. RESULTS: Regression coefficients based on MI including outcome were close to the truth. MI without outcome yielded very biased--underestimated--coefficients. SEs and coverage of the 90% confidence intervals were not different between MI with and without outcome. Results were the same for MCAR and MAR. CONCLUSION: For all types of missing values, imputation of missing predictor values using the outcome is preferred over imputation without outcome and is no self-fulfilling prophecy.

Adult↗

Analysis of the imputed female urinary incontinence data for the evaluation of expert system parameters.

We evaluated parameters for an expert system which will be designed to aid the differential diagnosis of female urinary incontinence by using knowledge discovered from data. To allow the statistical analysis, we applied means, regression and Expectation-Maximization (EM) imputation methods to fill in missing values. In addition, complete-case analysis was performed. Logistic regression results from the imputed data were reasonable. The significant parameters were mostly those that are important in the diagnostic work-up. Moreover, directions of relations between the parameters and the stress, mixed and sensory urge diagnoses were as expected. Analysis with the complete reduced data set gave clearly insufficient results. Imputed values had a moderate agreement, but odds ratios and classification accuracies of logistic regression equations were similar. Results suggest that with these data, simpler methods may be used to allow multivariate analysis and knowledge discovery, when better methods, such as EM imputation, are unavailable. Cluster analysis detected clusters corresponding to the small normal class, but was unable to clearly separate the larger incontinence classes.

Adult↗

The influence of mortality on twin models of change: addressing missingness through multiple imputation.

Twin analyses of phenotypes that are associated with mortality may provide biased heritability estimates if the models require that data from both members of a pair are available. This is particularly true when longitudinal analyses are applied to measures of cognition or biomarkers of aging. The effect of applying imputational techniques that include information on age at death was tested on longitudinal data from two twin studies of aging, each with up to four occasions of measurement. Measures of twin similarity for intercepts and slopes from three latent growth curve models were compared: without imputed data, including imputed data but without information on age at death, and including imputed data with information on age at death. Results indicated that twin similarity for slopes decreases when mortality is accounted for, but that considerable age-related covariation remains.

Humans↗

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents↗

Archaic ancestry inference in imputed ancient human genomes.

When modern humans expanded from Africa into Eurasia, they interbred with archaic hominins such as Neanderthals and Denisovans. This introgression shaped human evolution, yet most insights have been gained from present-day genomes, leaving little known about how archaic variants evolved after interbreeding. Ancient genomes offer a direct view of this process, but low coverage and poor quality have limited their use. Recent advances in genotype imputation offer a way to overcome these challenges by reconstructing missing information from reference panels and recovering evolutionary signals from low-coverage data. Here, we show that imputation enables accurate detection and quantification of archaic introgression in ancient genomes, improves local archaic ancestry inference, and that regions of archaic ancestry are imputed with especially high accuracy. We further demonstrate that imputed genomes can reconstruct the trajectories of introgressed haplotypes, distinguish populations across time and geography, and identify both known and additional candidates for adaptive introgression.

Humans↗

Use of multiple imputation to correct for nonresponse bias in a survey of urologic symptoms among African-American men.

The Flint Men's Health Study is an ongoing population-based study of African-American men designed to address questions related to prostate cancer and urologic symptoms. The initial phase of the study was conducted in 1996-1997 in two stages: an interviewer-administered survey followed by a clinical examination. The response rate in the clinical examination phase was 52%. Thus, some data were missing for clinical examination variables, diminishing the generalizability of the results to the general population. This paper is a case study demonstrating the application of multiple imputation to address important questions related to prostate cancer and urologic symptoms in a data set with missing values. On the basis of the observed clinical examination data, the American Urological Association Symptoms Score showed a surprising reduction in symptoms in the oldest age group, but after multiple imputation there was a monotonically increasing trend with age. It appeared that multiple imputation corrected for nonresponse bias associated with the observed data. For other outcome measures-namely, the age-adjusted 95th percentile of prostate-specific antigen level and the association between urologic symptoms and prostate volume-results from the observed data and the multiply imputed data were similar.

Adult↗

slideimp: efficient imputation of DNA methylation data.

SUMMARY: We developed slideimp, an R package that extends and optimizes K-nearest neighbor (K-NN) and Principal Component Analysis (PCA) imputation with grouped and sliding-window modes for accurate and efficient imputation of microarray and whole-genome DNA methylation (DNAm) data, respectively. Under a realistic scenario, slideimp achieved ≈12-28× faster runtime and ≈3-6× peak memory usage reduction for DNAm microarray imputation (GSE286313, EPICv2, N = 72) and achieved high imputation accuracy in a whole-genome DNAm dataset (N = 41). AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/hhp94/slideimp_paper. The R package slideimp is available on CRAN (DOI: 10.32614/CRAN.package.slideimp). Version 1.0.0 of slideimp, which was used in this study, is archived on Zenodo (DOI: 10.5281/zenodo.20029382).

DNA Methylation↗

Towards clustering of incomplete microarray data without the use of imputation.

MOTIVATION: Clustering technique is used to find groups of genes that show similar expression patterns under multiple experimental conditions. Nonetheless, the results obtained by cluster analysis are influenced by the existence of missing values that commonly arise in microarray experiments. Because a clustering method requires a complete data matrix as an input, previous studies have estimated the missing values using an imputation method in the preprocessing step of clustering. However, a common limitation of these conventional approaches is that once the estimates of missing values are fixed in the preprocessing step, they are not changed during subsequent processes of clustering; badly estimated missing values obtained in data preprocessing are likely to deteriorate the quality and reliability of clustering results. Thus, a new clustering method is required for improving missing values during iterative clustering process. RESULTS: We present a method for Clustering Incomplete data using Alternating Optimization (CIAO) in which a prior imputation method is not required. To reduce the influence of imputation in preprocessing, we take an alternative optimization approach to find better estimates during iterative clustering process. This method improves the estimates of missing values by exploiting the cluster information such as cluster centroids and all available non-missing values in each iteration. To test the performance of the CIAO, we applied the CIAO and conventional imputation-based clustering methods, e.g. k-means based on KNNimpute, for clustering two yeast incomplete data sets, and compared the clustering result of each method using the Saccharomyces Genome Database annotations. The clustering results of the CIAO method are more significantly relevant to the biological gene annotations than those of other methods, indicating its effectiveness and potential for clustering incomplete gene expression data. AVAILABILITY: The software was developed using Java language, and can be executed on the platforms that JVM (Java Virtual Machine) is running. It is available from the authors upon request.

Algorithms↗

Microarray missing data imputation based on a set theoretic framework and biological knowledge.

Gene expressions measured using microarrays usually suffer from the missing value problem. However, in many data analysis methods, a complete data matrix is required. Although existing missing value imputation algorithms have shown good performance to deal with missing values, they also have their limitations. For example, some algorithms have good performance only when strong local correlation exists in data while some provide the best estimate when data is dominated by global structure. In addition, these algorithms do not take into account any biological constraint in their imputation. In this paper, we propose a set theoretic framework based on projection onto convex sets (POCS) for missing data imputation. POCS allows us to incorporate different types of a priori knowledge about missing values into the estimation process. The main idea of POCS is to formulate every piece of prior knowledge into a corresponding convex set and then use a convergence-guaranteed iterative procedure to obtain a solution in the intersection of all these sets. In this work, we design several convex sets, taking into consideration the biological characteristic of the data: the first set mainly exploit the local correlation structure among genes in microarray data, while the second set captures the global correlation structure among arrays. The third set (actually a series of sets) exploits the biological phenomenon of synchronization loss in microarray experiments. In cyclic systems, synchronization loss is a common phenomenon and we construct a series of sets based on this phenomenon for our POCS imputation algorithm. Experiments show that our algorithm can achieve a significant reduction of error compared to the KNNimpute, SVDimpute and LSimpute methods.

Algorithms↗

Multiple imputation to account for missing data in a survey: estimating the prevalence of osteoporosis.

BACKGROUND: Nonresponse bias is a concern in any epidemiologic survey in which a subset of selected individuals declines to participate. METHODS: We reviewed multiple imputation, a widely applicable and easy to implement Bayesian methodology to adjust for nonresponse bias. To illustrate the method, we used data from the Canadian Multicentre Osteoporosis Study, a large cohort study of 9423 randomly selected Canadians, designed in part to estimate the prevalence of osteoporosis. Although subjects were randomly selected, only 42% of individuals who were contacted agreed to participate fully in the study. The study design included a brief questionnaire for those invitees who declined further participation in order to collect information on the major risk factors for osteoporosis. These risk factors (which included age, sex, previous fractures, family history of osteoporosis, and current smoking status) were then used to estimate the missing osteoporosis status for nonparticipants using multiple imputation. Both ignorable and nonignorable imputation models are considered. RESULTS: Our results suggest that selection bias in the study is of concern, but only slightly, in very elderly (age 80+ years), both women and men. CONCLUSIONS: Epidemiologists should consider using multiple imputation more often than is current practice.

Aged↗

Imputing response rates from means and standard deviations in meta-analyses.

The principle of intention-to-treat analysis must be strictly applied to both individual randomized controlled trial and meta-analysis but, in doing so, would involve imputation of some missing data. There is little literature on how to perform this in the case of meta-analysis. For dichotomous outcome measures, one possible strategy is to carry out a sensitivity analysis based on the so-called best case/worst case analyses. For continuous outcomes, it may be possible to achieve this if we can dichotomise the continuous outcomes. Here, we empirically examined the appropriateness of converting continuous outcomes (expressed as mean+/-SD) into dichotomous outcomes (expressed as response rates) in four completed meta-analyses of depression and anxiety, assuming normal distribution of the continuous outcome measures. The agreement between the actually observed versus the imputed raw numbers of responders was indicated by an intraclass correlation coefficient of 0.97 (95% confidence interval 0.95-0.98). The pooled relative risks of the four meta-analyses based on the imputed values were virtually identical to those based on the actually observed values. When individual trials report the means+/-SDs of their outcome measures but fail to report response rates, it may therefore be possible to impute the response rates based on the means+/-SDs, and then submit the meta-analysis to worst case/best case analyses. This would allow a more robust and clinically interpretable estimation of the true, underlying treatment effect to be made.

Data Interpretation, Statistical↗