Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Probability imputation revisited for prognostic factor studies.

The analysis of prognostic factor studies by Cox or logistic regression models is often impeded by missing covariate values. In 1990 Schemper and Smith recommended a conditional probability imputation technique (PIT) for the analysis of treatment studies which can be easily applied using standard software and which has been demonstrated to outperform the complete case and omission of covariates strategies. Recent research, however, showed that PIT cannot universally be recommended and it was concluded that model-based methods should be preferred. We agree with these conclusions but also think that there is enough empirical evidence to judge the performance of PIT to be satisfactory in typical prognostic factor studies. Furthermore, comparisons of PIT with multiple imputation in the same context did not indicate an advantage of the latter more involved technique. By means of an analysis of a prostate cancer data set various aspects of application of PIT are discussed, in particular that PIT permits direct comparability of marginal and partial effects analyses. We conclude that PIT continues to be an appropriate and attractive choice for analyses of prognostic factor studies.

Data Interpretation, Statistical↗

Smoking imputation and lung cancer in railroad workers exposed to diesel exhaust.

BACKGROUND: An association between diesel exhaust exposure and lung cancer mortality in a large retrospective cohort study of US railroad workers has previously been reported. However, specific information regarding cigarette smoking was unavailable. METHODS: Birth cohort, age, job, and cause of death specific smoking histories from a companion case-control study were used to impute smoking behavior for 39,388 railroad workers who died 1959-1996. Mortality analyses incorporated the effect of smoking on lung cancer risk. RESULTS: The smoking adjusted relative risk of lung cancer in railroad workers exposed to diesel exhaust compared to unexposed workers was 1.22 (95% CI = 1.12-1.32), and unadjusted for smoking the relative risk was 1.35 (95% CI = 1.24-1.46). CONCLUSIONS: These analyses illustrate the use of imputation in record-based occupational health studies to assess potential confounding due to smoking. In this cohort, small differences in smoking behavior between diesel exposed and unexposed workers did not explain the elevated lung cancer risk.

Adult↗

Prediction of survival and opportunistic infections in HIV-infected patients: a comparison of imputation methods of incomplete CD4 counts.

In evaluating the risk of mortality or development of opportunistic infections in HIV-infected patients, the number of CD4 lymphocyte cells per cubic millimetre of blood is widely recognized as one of the best available predictors of such future events. However, its usefulness is limited by the incompleteness and variability of such CD4 measurements during follow-up. Because of these limitations, analysis of such data requires the missing measurements to be 'filled in' or the patients without them to be excluded. We consider multiple imputation of CD4 values based partly on information from other health status measures such as haemoglobin, as well as on the event status of interest. These alternative health status measures are also considered as possible independent predictors of survival endpoints. Our work is motivated by a cohort of 1530 patients enrolled in two AIDS clinical trials. We compare our approach to other strategies such as basing evaluation of risk on baseline CD4, the last measured CD4 before an event, or a time-dependent covariate based on carrying the last CD4 value forward; we conclude with a strong recommendation for multiple imputation.

AIDS-Related Opportunistic Infections↗

A comparison of imputation methods in a longitudinal randomized clinical trial.

It is common for longitudinal clinical trials to face problems of item non-response, unit non-response, and drop-out. In this paper, we compare two alternative methods of handling multivariate incomplete data across a baseline assessment and three follow-up time points in a multi-centre randomized controlled trial of a disease management programme for late-life depression. One approach combines hot-deck (HD) multiple imputation using a predictive mean matching method for item non-response and the approximate Bayesian bootstrap for unit non-response. A second method is based on a multivariate normal (MVN) model using PROC MI in SAS software V8.2. These two methods are contrasted with a last observation carried forward (LOCF) technique and available-case (AC) analysis in a simulation study where replicate analyses are performed on subsets of the originally complete cases. Missing-data patterns were simulated to be consistent with missing-data patterns found in the originally incomplete cases, and observed complete data means were taken to be the targets of estimation. Not surprisingly, the LOCF and AC methods had poor coverage properties for many of the variables evaluated. Multiple imputation under the MVN model performed well for most variables but produced less than nominal coverage for variables with highly skewed distributions. The HD method consistently produced close to nominal coverage, with interval widths that were roughly 7 per cent larger on average than those produced from the MVN model.

Aged↗

Multiple imputation under Bayesianly smoothed pattern-mixture models for non-ignorable drop-out.

Conventional pattern-mixture models can be highly sensitive to model misspecification. In many longitudinal studies, where the nature of the drop-out and the form of the population model are unknown, interval estimates from any single pattern-mixture model may suffer from undercoverage, because uncertainty about model misspecification is not taken into account. In this article, a new class of Bayesian random coefficient pattern-mixture models is developed to address potentially non-ignorable drop-out. Instead of imposing hard equality constraints to overcome inherent inestimability problems in pattern-mixture models, we propose to smooth the polynomial coefficient estimates across patterns using a hierarchical Bayesian model that allows random variation across groups. Using real and simulated data, we show that multiple imputation under a three-level linear mixed-effects model which accommodates a random level due to drop-out groups can be an effective method to deal with non-ignorable drop-out by allowing model uncertainty to be incorporated into the imputation process.

Antipsychotic Agents↗

Multiple imputation methods for modelling relative survival data.

In population-based cancer survival studies, the cause-specific survival measures the net survival (excess mortality) due to cancer when the cause of death information is available and reliable. In contrast, when the cause of death is uncertain or unavailable, relative survival, the ratio of the survival rate due to all causes to the expected survival rate, is more appropriate. There is a large body of work on the modelling and hypothesis testing of cause-specific survival, but many of these methods are not directly applicable to relative survival. In this paper, we extend the multiple imputation (MI) methods (Stat. Methods Med. Res. 1999; 8:3-15) to the case of relative survival data. The MI methodology is combined with relative survival to estimate the net survival by changing relative survival data to cause-specific data. This facilitates the direct application to statistical methods developed for the cause-specific survival to the special situation of relative survival. The parameter estimates and the log-rank statistics are obtained by combining the results from multiple imputed cause-specific data. The likelihood-based methods for modelling relative survival data have been implemented by a Windows application called CANSURV (Comput. Meth. Prog. Biomed 2005). Although these methods produce accurate parameter estimates, the choice for models and diagnostic tools is limited. The MI method is presented as a simpler alternative. The relative survival data for the colorectal cancer patients from Surveillance, Epidemiology, and End Results (SEER) program (SEER Cancer Statistics Review, 1973-1999. National Cancer Institute: Bethesda, 2002) is used as an illustration. The results are compared with those obtained from the likelihood-based relative survival analysis methods. A sample SAS macro for the MI method is provided.

Black People↗

A model-based approach to the imputation of missing data: home injury incidences.

Missing or incomplete data cases are a problem in all types of statistical analyses. In disease surveillance, this problem inhibits determining the actual incidence of a disease event and monitoring the disease occurrence. Several statistical techniques have been developed to impute values for incomplete data cases. We present a model-based approach to the imputation of missing data elements as applied to determining the incidence of home injury deaths.

Accidents, Home↗

A multiple imputation strategy for clinical trials with truncation of patient data.

Clinical trials of drug treatments for psychiatric disorders commonly employ the parallel groups, placebo-controlled, repeated measure randomized comparison. When patients stop adhering to their originally assigned treatment, investigators often abandon data collection. Thus, non-adherence produces a monotone pattern of unit-level missing data, disabling the analysis by intent-to-treat. We propose an approach based on multiple imputation of the missing responses, using the approximate Bayesian bootstrap to draw ignorable repeated imputations from the posterior predictive distribution of the missing data, stratifying by a balancing score for the observed responses prior to withdrawal. We apply the method and some variations to data from a large randomized trial of treatments for panic disorder, and compare the results to those obtained by the original analysis that used the standard (endpoint) method.

Adult↗

Analysis of incomplete quality of life data in advanced stage cancer: a practical application of multiple imputation.

This paper presents a practical approach to analyzing incomplete quality of life (QOL) data that contains non-ignorable dropouts in patients with advanced non-small-cell lung cancer (NSCLC). QOL scores for the physical domain at baseline and at the end of the first and second courses of chemotherapy were compared between two treatment groups in a phase III trial. One hundred and 103 eligible patients were randomized to receive cisplatin and irinotecan (CPT-P) or cisplatin and vindesine, respectively; of those two groups, 83 and 85, respectively, completed a QOL questionnaire at least at baseline. A multiple imputation incorporating auxiliary QOL variables was implemented as one of alternatives of sensitivity analyses; these were complete case, available case, and pattern mixture analyses. Although larger sensitivity to missing data was found for CPT-P treatment, none of the alternative analyses demonstrated a significant difference in estimated slopes over time between the groups. This study presents an analytical approach for dealing with the complex problem of missing QOL data. It must be noted, however, that the validity of the multiple imputation method we present is not certain unless we can specify sufficiently informative auxiliary variables to ensure the conversion of non-ignorable missingness to ignorable.

Aged↗

Variance imputation for overviews of clinical trials with continuous response.

Overviews of clinical trials are an efficient and important means of summarizing information about a particular scientific area. When the outcome is a continuous variable, both treatment effect and variance estimates are required to construct a confidence interval for the overall treatment effect. Often, only partial information about the variance is provided in the publication of the clinical trial. This paper provides heuristic suggestions for variance imputation based on partial variance information. Both pretest-posttest (parallel groups) and crossover designs are considered. A key idea is to use separate sources of incomplete information to help choose a better variance estimate. The imputation suggestions are illustrated with a data set.

Analysis of Variance↗

On the statistical analysis of the GS-NS0 cell proteome: imputation, clustering and variability testing.

We have undertaken two-dimensional gel electrophoresis proteomic profiling on a series of cell lines with different recombinant antibody production rates. Due to the nature of gel-based experiments not all protein spots are detected across all samples in an experiment, and hence datasets are invariably incomplete. New approaches are therefore required for the analysis of such graduated datasets. We approached this problem in two ways. Firstly, we applied a missing value imputation technique to calculate missing data points. Secondly, we combined a singular value decomposition based hierarchical clustering with the expression variability test to identify protein spots whose expression correlates with increased antibody production. The results have shown that while imputation of missing data was a useful method to improve the statistical analysis of such data sets, this was of limited use in differentiating between the samples investigated, and highlighted a small number of candidate proteins for further investigation.

Algorithms↗

Including multiple imputation in a sensitivity analysis for clinical trials with treatment failures.

When treatment failures occur during the course of a clinical trial, the treatment regimen following failure may be changed. This change in therapy complicates comparisons among the original treatment arms. As in some clinical trials with dropouts, intent-to-treat analysis can yield a large bias. We examine the use of multiple imputation to replace observations after treatment failure has occurred. As a sensitivity analysis, this approach is compared to existing methods for handling treatment failures - removing treatment failure subjects, removing data after the onset of treatment failure, and imputation the last observation prior to treatment failure for all subsequent observations - in addition to an analysis of all collected data based on randomized treatment assignment. A data set from the Asthma Clinical Research Network is used to demonstrate the methods.

Algorithms↗

Missing data imputation through GTM as a mixture of t-distributions.

The Generative Topographic Mapping (GTM) was originally conceived as a probabilistic alternative to the well-known, neural network-inspired, Self-Organizing Maps. The GTM can also be interpreted as a constrained mixture of distribution models. In recent years, much attention has been directed towards Student t-distributions as an alternative to Gaussians in mixture models due to their robustness towards outliers. In this paper, the GTM is redefined as a constrained mixture of t-distributions: the t-GTM, and the Expectation-Maximization algorithm that is used to fit the model to the data is modified to carry out missing data imputation. Several experiments show that the t-GTM successfully detects outliers, while minimizing their impact on the estimation of the model parameters. It is also shown that the t-GTM provides an overall more accurate imputation of missing values than the standard Gaussian GTM.

Algorithms↗

Multiple imputation versus data enhancement for dealing with missing data in observational health care outcome analyses.

The problem of missing data is frequently encountered in observational studies. We compared approaches to dealing with missing data. Three multiple imputation methods were compared with a method of enhancing a clinical database through merging with administrative data. The clinical database used for comparison contained information collected from 6,065 cardiac care patients in 1995 in the province of Alberta, Canada. The effectiveness of the different strategies was evaluated using measures of discrimination and goodness of fit for the 1995 data. The strategies were further evaluated by examining how well the models predicted outcomes in data collected from patients in 1996. In general, the different methods produced similar results, with one of the multiple imputation methods demonstrating a slight advantage. It is concluded that the choice of missing data strategy should be guided by statistical expertise and data resources.

Alberta↗

Multiple imputation for the Cox proportional hazards model with missing covariates.

We present three multiple imputation estimates for the Cox model with missing covariates. Two of the suggested estimates are asymptotically equivalent to estimates in the literature when the number of multiple imputations approaches infinity. The third estimate can be implemented using standard software that could handle time-varying covariates.

Analysis of Variance↗

Rapid N-acetyltransferase 2 imputed phenotype and smoking may increase risk of colorectal cancer in women (Netherlands).

OBJECTIVE: The relationship between smoking and colorectal cancer risk and whether such effect is modified by variations in the NAT2 genotype is investigated. METHODS: In the prospective DOM (Diagnostisch Onderzoek Mammacarcinoom; 27,722 women) cohort follow-up from 1976 until 1987 revealed 54 deaths due to colon or rectal cancer, and follow-up from 1987 to 01-01-1996 revealed 204 incident colorectal cancer cases. A random sample (n = 857) from the baseline cohort was used as controls. Four NAT2 restriction fragment length polymorphisms (RFLPs) were analysed using DNA extracted from urine samples. Rapid or slow acetylator phenotype status was attributed to individuals. RESULTS: Smoking may increase the risk for colon cancer (RR = 1.36, 95% CI 0.97-1.92) as well as for rectal cancer (RR = 1.31, 95% CI 0.76-2.25), although not statistically significant. Rapid NAT2 acetylation did not increase colorectal cancer risk, but in combination with smoking the risk was statistically significant increased, compared to women who had a slow NAT2 imputed phenotype and never smoked (RR = 1.56, 95% CI 1.03-2.37). For colon cancer, but not for rectal cancer the increased risk was statistically significant (RR = 1.67, 95% CI, 1.05-2.67 versus RR = 1.30 95% CI 0.63-2.68). CONCLUSIONS: Our study points to smoking as a risk factor for colon and rectal cancer and, in addition, especially in women with rapid NAT2 imputed phenotype.

Arylamine N-Acetyltransferase↗

A SAS macro for a simulation study of imputation methods for missing values--an application of Bebbington's algorithm.

This paper presents a SAS macro for a simulation study of comparing a new variant of hot-deck imputation with mean imputation for missing values, in which a simple algorithm proposed by Bebbington (Applied Statistics, 1975) for carrying out simple random sampling without replacement was employed to draw repeated random samples efficiently. A simulated example of drawing repeated random samples from a regional survey of obesity in school children was used to demonstrate the SAS macro.

Algorithms↗

Causality assessment of adverse drug reactions: comparison of the results obtained from published decisional algorithms and from the evaluations of an expert panel, according to different levels of imputability.

OBJECTIVES: To evaluate agreement between causality assessments of reported adverse drug reactions (ADRs) obtained from decisional algorithms, with those obtained from an expert panel using the WHO global introspection method (GI), according to different levels of imputability and to evaluate the influence of confounding variables. METHOD: Two hundred reports were included in this study. An independent researcher used decisional algorithms, while an expert panel assessed the same ADR reports using the GI, both aimed at evaluating causality. Reports were divided according to the presence, absence or lack of information on confounding variables. RESULTS: The rates of concordance between assessments made using the algorithms and GI according to levels of imputability were: 45% for 'certain', 61% for 'probable', 46% for 'possible' and 17% for drug unrelated terms. When confounding variables were taken into account, the rates of concordance for the 'absence of information', 'lack of information' and 'presence of confounding variables' in the 'certain' group were 49, 69 and 7%, respectively. The corresponding values for the 'probable' group were 80, 68 and 24% and 30, 51 and 51%, respectively for the 'possible' group. CONCLUSION: Full agreement with global introspection was not found for any level of causality assessment. Confounding variables were found to be associated with low levels of agreement between decision algorithms and the GI method compromising the algorithms' sensitivity and specificity.

Adverse Drug Reaction Reporting Systems↗