Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Linkage analysis with sequential imputation.

Multilocus calculations, using all available information on all pedigree members, are important for linkage analysis. Exact calculation methods in linkage analysis are limited in either the number of loci or the number of pedigree members they can handle. In this article, we propose a Monte Carlo method for linkage analysis based on sequential imputation. Unlike exact methods, sequential imputation can handle large pedigrees with a moderate number of loci in its current implementation. This Monte Carlo method is an application of importance sampling, in which we sequentially impute ordered genotypes locus by locus, and then impute inheritance vectors conditioned on these genotypes. The resulting inheritance vectors, together with the importance sampling weights, are used to derive a consistent estimator of any linkage statistic of interest. The linkage statistic can be parametric or nonparametric; we focus on nonparametric linkage statistics. We demonstrate that accurate estimates can be achieved within a reasonable computing time. A simulation study illustrates the potential gain in power using our method for multilocus linkage analysis with large pedigrees. We simulated data at six markers under three models. We analyzed them using both sequential imputation and GENEHUNTER. GENEHUNTER had to drop between 38-54% of pedigree members, whereas our method was able to use all pedigree members. The power gains of using all pedigree members were substantial under 2 of the 3 models. We implemented sequential imputation for multilocus linkage analysis in a user-friendly software package called SIMPLE.

Genetic Linkage↗

Imputation methods to improve inference in SNP association studies.

Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we develop two haplotype-based imputation approaches and one tree-based imputation approach for association studies. The emphasis is to evaluate the impact of imputation on parameter estimation, compared to the standard practice of ignoring missing data. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to iteratively sample from a full conditional distribution, which is obtained from the classification and regression tree (CART) algorithm. We employ a standard multiple imputation procedure to account for the uncertainty of imputation. We apply the methods to simulated data as well as a case-control study on developmental dyslexia. Our results suggest that imputation generally improves efficiency over the standard practice of ignoring missing data. The tree-based approach performs comparably well as haplotype-based approaches, but the former has a computational advantage. The WEM approach yields the smallest bias at a price of increased variance.

Algorithms↗

Multiple imputation in health-care databases: an overview and some applications.

Multiple imputation for non-response replaces each missing value by two or more plausible values. The values can be chosen to represent both uncertainty about the reasons for non-response and uncertainty about which values to impute assuming the reasons for non-response are known. This paper provides an overview of methods for creating and analysing multiply-imputed data sets, and illustrates the dramatic improvements possible when using multiple rather than single imputation. A major application of multiple imputation to public-use files from the 1970 census is discussed, and several exploratory studies related to health care that have used multiple imputation are described.

Databases, Factual↗

A simple imputation algorithm reduced missing data in SF-12 health surveys.

OBJECTIVE: The SF-12 Health Survey is a 12-item questionnaire that yields two summary scores (physical and mental health). Neither score can be computed when an item is missing. We explored imputation methods for missing scores for this instrument. STUDY DESIGN AND SETTING: Using data from a population-based survey, we tested several ways of imputing simulated missing data. RESULTS: Among 1250 participants, 118 (9.6%) had at last one missing SF-12 item. Missing data were more common among women, older respondents, non-Swiss nationals, and health service users. Among the 1132 respondents with complete data, replacement of any item with the mean population item weight yielded good results: the mean correlation between imputed and true score was 0.979 for both the physical and mental score. Results remained satisfactory when up to three of the six key items for each score (items that contribute predominantly to a given score), and any number of non-key items, were replaced by the mean. Application of this imputation algorithm to the original survey reduced the proportion of missing scores to <1%. Respondents with incomplete surveys, hence imputed scores, had lower scores than respondents with complete data (physical score: 44.9 vs. 49.8, p < 0.001, mental score: 44.4 vs. 46.3, p=0.064). CONCLUSIONS: A simple imputation algorithm can substantially reduce the proportion of missing scores for the SF-12 health survey, and consequently reduce non-response bias.

Algorithms↗

Survival estimates of a prognostic classification depended more on year of treatment than on imputation of missing values.

BACKGROUND AND OBJECTIVE: The International Germ Cell Consensus (IGCC) classification defines good, intermediate, and poor prognosis groups among patients with nonseminomatous germ cell cancer. In the database used to develop the IGCC classification (n = 5,202), >40% of patients were excluded because of missing values (n = 2,154). We looked for effects of this exclusion on survival estimates in the three IGCC prognosis groups. STUDY DESIGN AND SETTING: We imputed missing values using a multiple imputation procedure. The IGCC classification was applied to patients with complete data (n = 3,048) and with imputed data (n = 2,154), and 5-year survival was calculated for each prognosis group. RESULTS: Patients with missing values had a lower 5-year survival than those without missing values: 76% vs. 82%. Five-year survival in the complete and imputed data samples was 92% and 87% for the good prognosis groups and 80% and 70% for the intermediate prognosis groups, whereas 5-year survival for the poor prognosis groups in both samples was similar (50% and 47%, respectively). This difference in survival was largely explained by a higher proportion of missing values among patients treated before 1985, who had a worse survival than patients treated after 1985. CONCLUSION: Multiple imputation of the missing values led to lower survival estimates across the IGCC prognosis groups, compared with estimates based on the complete data. Although imputation of missing values gives statistically better survival estimates, adjustments for year of treatment are necessary to make the estimates applicable to currently diagnosed patients with testicular cancer.

Classification↗

Gaussian mixture clustering and imputation of microarray data.

MOTIVATION: In microarray experiments, missing entries arise from blemishes on the chips. In large-scale studies, virtually every chip contains some missing entries and more than 90% of the genes are affected. Many analysis methods require a full set of data. Either those genes with missing entries are excluded, or the missing entries are filled with estimates prior to the analyses. This study compares methods of missing value estimation. RESULTS: Two evaluation metrics of imputation accuracy are employed. First, the root mean squared error measures the difference between the true values and the imputed values. Second, the number of mis-clustered genes measures the difference between clustering with true values and that with imputed values; it examines the bias introduced by imputation to clustering. The Gaussian mixture clustering with model averaging imputation is superior to all other imputation methods, according to both evaluation metrics, on both time-series (correlated) and non-time series (uncorrelated) data sets.

Algorithms↗

Imputation and variable selection in linear regression models with missing covariates.

Across multiply imputed data sets, variable selection methods such as stepwise regression and other criterion-based strategies that include or exclude particular variables typically result in models with different selected predictors, thus presenting a problem for combining the results from separate complete-data analyses. Here, drawing on a Bayesian framework, we propose two alternative strategies to address the problem of choosing among linear regression models when there are missing covariates. One approach, which we call "impute, then select" (ITS) involves initially performing multiple imputation and then applying Bayesian variable selection to the multiply imputed data sets. A second strategy is to conduct Bayesian variable selection and missing data imputation simultaneously within one Gibbs sampling process, which we call "simultaneously impute and select" (SIAS). The methods are implemented and evaluated using the Bayesian procedure known as stochastic search variable selection for multivariate normal data sets, but both strategies offer general frameworks within which different Bayesian variable selection algorithms could be used for other types of data sets. A study of mental health services utilization among children in foster care programs is used to illustrate the techniques. Simulation studies show that both ITS and SIAS outperform complete-case analysis with stepwise variable selection and that SIAS slightly outperforms ITS.

Algorithms↗

Applications of multiple imputation in medical studies: from AIDS to NHANES.

Rubin's multiple imputation is a three-step method for handling complex missing data, or more generally, incomplete-data problems, which arise frequently in medical studies. At the first step, m (> 1) completed-data sets are created by imputing the unobserved data m times using m independent draws from an imputation model, which is constructed to reasonably approximate the true distributional relationship between the unobserved data and the available information, and thus reduce potentially very serious nonresponse bias due to systematic difference between the observed data and the unobserved ones. At the second step, m complete-data analyses are performed by treating each completed-data set as a real complete-data set, and thus standard complete-data procedures and software can be utilized directly. At the third step, the results from the m complete-data analyses are combined in a simple, appropriate way to obtain the so-called repeated-imputation inference, which properly takes into account the uncertainty in the imputed values. This paper reviews three applications of Rubin's method that are directly relevant for medical studies. The first is about estimating the reporting delay in acquired immune deficiency syndrome (AIDS) surveillance systems for the purpose of estimating survival time after AIDS diagnosis. The second focuses on the issue of missing data and noncompliance in randomized experiments, where a school choice experiment is used as an illustration. The third looks at handling nonresponse in United States National Health and Nutrition Examination Surveys (NHANES). The emphasis of our review is on the building of imputation models (i.e. the first step), which is the most fundamental aspect of the method.

Acquired Immunodeficiency Syndrome↗

Missing value estimation for DNA microarray gene expression data by Support Vector Regression imputation and orthogonal coding scheme.

BACKGROUND: Gene expression profiling has become a useful biological resource in recent years, and it plays an important role in a broad range of areas in biology. The raw gene expression data, usually in the form of large matrix, may contain missing values. The downstream analysis methods that postulate complete matrix input are thus not applicable. Several methods have been developed to solve this problem, such as K nearest neighbor impute method, Bayesian principal components analysis impute method, etc. In this paper, we introduce a novel imputing approach based on the Support Vector Regression (SVR) method. The proposed approach utilizes an orthogonal coding input scheme, which makes use of multi-missing values in one row of a certain gene expression profile and imputes the missing value into a much higher dimensional space, to obtain better performance. RESULTS: A comparative study of our method with the previously developed methods has been presented for the estimation of the missing values on six gene expression data sets. Among the three different input-vector coding schemes we tried, the orthogonal input coding scheme obtains the best estimation results with the minimum Normalized Root Mean Squared Error (NRMSE). The results also demonstrate that the SVR method has powerful estimation ability on different kinds of data sets with relatively small NRMSE. CONCLUSION: The SVR impute method shows better performance than, or at least comparable with, the previously developed methods in present research. The outstanding estimation ability of this impute method is partly due to the use of the most missing value information by incorporating orthogonal input coding scheme. In addition, the solid theoretical foundation of SVR method also helps in estimation of performance together with orthogonal input coding scheme. The promising estimation ability demonstrated in the results section suggests that the proposed approach provides a proper solution to the missing value estimation problem. The source code of the SVR method is available from http://202.38.78.189/downloads/svrimpute.html for non-commercial use.

Algorithms↗

Imputation for exposure histories with gaps, under an excess relative risk model.

In reconstructing exposure histories needed to calculate cumulative exposures, gaps often occur. Our investigation was motivated by case-control studies of residential radon exposure and lung cancer, where half or more of the targeted homes may not be measurable. Investigators have adopted various schemes for imputing exposures for such gaps. We first undertook simulations to assess the performance of five such methods under an excess relative risk model, in the presence of random missingness and under assumed independence among the true exposure levels for different epochs of exposure (houses). Assuming no other source of measurement error, one of the methods performed without bias and with coverage of nominally 95% confidence intervals that was close to 95%. This method assigns to the missing residences the arithmetic mean across all measured control residences. We show that its good properties can be explained by the fact that this approach produces approximate "Berkson errors." To take advantage of predictive information that might exist about the missing epochs of exposure, one might prefer to carry out the imputations within strata. In further simulations, we asked whether the method would still perform well if imputations were carried out within many strata. It does, and much of the lost statistical power/precision can be recovered if the stratification system is moderately predictive of the missing exposures. Thus, observed control mean imputation provides a way to impute missing exposures without corrupting the study's validity; and stratifying the imputations can enhance precision. The technique is applicable in other settings where exposure histories contain gaps.

Environmental Exposure↗

Selphi, a tool for improving genotype imputation accuracy.

Genotype imputation is a powerful tool for inferring missing genotype data in large-scale genetic studies. Over the last two decades, multiple imputation algorithms have been developed, steadily improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge, largely because existing methods rely on local haplotype matching within genomic windows and do not fully exploit the extended patterns of haplotype sharing that span entire chromosomes. Here we present Selphi, a new genotype imputation algorithm that combines the Positional Burrows-Wheeler Transform (PBWT) with a multi-stage haplotype selection heuristic operating across entire chromosomes. When compared to state-of-the-art methods Beagle 5.4, IMPUTE5, and Minimac4, Selphi showed higher accuracy on the 1000 Genomes Project and TOPMed datasets, across all super-populations and allele frequencies. Similarly, Selphi achieved higher accuracy than Beagle 5.4 on the UK Biobank dataset, which translated into improved concordance with hc-WGS GWAS summary statistics at known trait-associated loci and more accurate polygenic risk scores (PRS). Selphi outputs standard VCF files with genotype dosages (DS), haplotype-specific allele probabilities (AP1, AP2), and a per-variant dosage R-squared quality score (DR2), enabling direct integration with downstream analytical pipelines including standard post-imputation quality filtering.

Genome-Wide Association Study↗

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics↗

Evaluating the Antigen and Eplet Accuracy of DQA1 Imputations With the HaploSFHI Two-Field HLA Typing Inference Tool.

Donor/recipient mismatched HLA antigens can lead to the production of Donor-Specific Antibodies by the recipient, which are deleterious to organ transplants. The HLA-DQ locus is the most frequent target, with both the DQ beta and alpha chains involved. For deceased donors in particular, while HLA-DQB1 has been typed in emergencies for a long time, HLA-DQA1 has only recently been included. No imputation algorithmic tool was available to impute HLA-DQA1 until the development of HaploSFHI, trained on 61,393 two-field typings by NGS methods. We evaluated the accuracy of two-field HLA-DQA1 imputation from serological and two-field level HLA-A, B, DRB1, and DQB1 typings. We report a highly accurate two-field HLA-DQA1 prediction using a French test cohort of 7696 individuals, respectively reaching 92.30% and 96.45% accuracy. The average 'False Positive eplet load' stood at 0.19 and 0.07, respectively, and the average 'False Negative eplet load' at 0.18 and 0.08, respectively. A similar performance was obtained on three independent test cohorts of European ancestry (from the USA, the UK, and Portugal). Interestingly, performance was only slightly inferior on five independent test cohorts of other ethnicities (from Hong Kong and the USA) whereas it was significantly lower for two-field DRB1 imputation from its serological level. These results suggest that DQA1 can reliably be imputed even when information is totally missing, with low error risk at both antigen and eplet levels, even if the reference population is not matched. Similar additional initiatives would be welcome to confirm these findings.

Humans↗

Estimation and testing of genotype and haplotype effects in case-control studies: comparison of weighted regression and multiple imputation procedures.

A popular approach for testing and estimating genotype and haplotype effects associated with a disease outcome is to conduct a population-based case/control study, in which haplotypes are not directly observed but may be inferred probabilistically from unphased genotype data. A variety of methods exist to analyse the resulting data while accounting for the uncertainty in haplotype assignment, but most focus on the issue of testing the global null hypothesis that no genotype or haplotype effects exist. A more interesting question, once a region of disease association has been identified, is to estimate the relevant genotypic or haplotypic effects and to perform tests of complex null hypotheses such as the hypothesis that some loci, but not others, are associated with disease. Here I examine the assumptions behind, and the performance of, two classes of methods for addressing this question. The first is a weighted regression approach in which posterior probabilities of haplotype assignments are used as weights in a logistic regression analysis, generating a test based on either a weighted pseudo-likelihood, or a weighted log-likelihood. The second is a multiple imputation approach using either an improper procedure in which the posterior probabilities are used to generate replicate imputed data sets, or a proper data augmentation procedure. I compare these approaches to a simple expectation substitution (haplotype trend regression) approach. In simulations, all methods gave unbiased parameter estimation but the weighted pseudo-likelihood, expectation substitution and multiple imputation methods had superior confidence interval coverage. For the weighted pseudo-likelihood and expectation substitution methods it was necessary to estimate posterior haplotype assignment probabilities using the combined case/control data, whereas for the multiple imputation approaches it was necessary to estimate these probabilities in the case and control groups separately. Overall, multiple imputation was easiest approach to implement in standard statistical software and to extend to more complex models such as those that include gene-gene or gene-environment interactions.

Case-Control Studies↗

Use of the mean, hot deck and multiple imputation techniques to predict outcome in intensive care unit patients in Colombia.

A cohort of intensive care unit (ICU) patients in 20 Colombian ICUs is used to describe the application of three imputation techniques: single, hot deck and multiple imputation. These strategies were used to impute the missing data in the variables used to construct APACHE II scores, a scoring system for the ICU patients that provides an unbiased standardized estimate of the probability of hospital death. Imputed APACHE II scores were then used in the APACHE II model to estimate adjusted hospital mortality rates. The area under the receiver operating characteristic (ROC) curve was used to compare imputation strategies with respect to predictive power. While statistically significant differences were found for the area under the ROC curve, these differences were not clinically significant.

APACHE↗

Multiple imputation for body mass index: lessons from the Australian Longitudinal Study on Women's Health.

In large epidemiological studies missing data can be a problem, especially if information is sought on a sensitive topic or when a composite measure is calculated from several variables each affected by missing values. Multiple imputation is the method of choice for 'filling in' missing data based on associations among variables. Using an example about body mass index from the Australian Longitudinal Study on Women's Health, we identify a subset of variables that are particularly useful for imputing values for the target variables. Then we illustrate two uses of multiple imputation. The first is to examine and correct for bias when data are not missing completely at random. The second is to impute missing values for an important covariate; in this case omission from the imputation process of variables to be used in the analysis may introduce bias. We conclude with several recommendations for handling issues of missing data.

Australia↗

Multiple imputation techniques in small sample clinical trials.

Clinical trials allow researchers to draw conclusions about the effectiveness of a treatment. However, the statistical analysis used to draw these conclusions will inevitably be complicated by the common problem of attrition. Resorting to ad hoc methods such as case deletion or mean imputation can lead to biased results, especially if the amount of missing data is high. Multiple imputation, on the other hand, provides the researcher with an approximate solution that can be generalized to a number of different data sets and statistical problems. Multiple imputation is known to be statistically valid when n is large. However, questions still remain about the validity of multiple imputation for small samples in clinical trials. In this paper we investigate the small-sample performance of several multiple imputation methods, as well as the last observation carried forward method.

Bayes Theorem↗

Gaussianization-based quasi-imputation and expansion strategies for incomplete correlated binary responses.

New quasi-imputation and expansion strategies for correlated binary responses are proposed by borrowing ideas from random number generation. The core idea is to convert correlated binary outcomes to multivariate normal outcomes in a sensible way so that re-conversion to the binary scale, after performing multiple imputation, yields the original specified marginal expectations and correlations. This conversion process ensures that the correlations are transformed reasonably which in turn allows us to take advantage of well-developed imputation techniques for Gaussian outcomes. We use the phrase 'quasi' because the original observations are not guaranteed to be preserved. We argue that if the inferential goals are well-defined, it is not necessary to strictly adhere to the established definition of multiple imputation. Our expansion scheme employs a similar strategy where imputation is used as an intermediate step. It leads to proportionally inflated observed patterns, forcing the data set to a complete rectangular format. The plausibility of the proposed methodology is examined by applying it to a wide range of simulated data sets that reflect alternative assumptions on complete data populations and missing-data mechanisms. We also present an application using a data set from obesity research. We conclude that the proposed method is a promising tool for handling incomplete longitudinal or clustered binary outcomes under ignorable non-response mechanisms.

Adolescent↗