Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A model of contextual effect on reproduced extents in recall tasks: the issue of the imputed motion hypothesis.

In this article the fundamental question of space and time dependencies in the reproduction of spatial or temporal extents is studied. The functional dependence of spatial responses on the temporal context and the corresponding dependence of temporal responses on spatial context are reported as the tau and kappa effects, respectively. A common explanation suggested that the participant imputes motion to discontinuous displays. Using a mathematical model we explore the imputed velocity hypothesis and provide a globally fit model that addresses the question of sequences modelling. Our model accounts for observed data in the tau experiment. The accuracy of the model is improved introducing a new hypothesis based on small velocity variations. On the other hand, results show that the imputed velocity hypothesis fails to reproduce the kappa effect. This result definitively shows that both effects are not symmetric.

Adult↗

A multiple imputation approach to linear regression with clustered censored data.

We extend Wei and Tanner's (1991) multiple imputation approach in semi-parametric linear regression for univariate censored data to clustered censored data. The main idea is to iterate the following two steps: 1) using the data augmentation to impute for censored failure times; 2) fitting a linear model with imputed complete data, which takes into consideration of clustering among failure times. In particular, we propose using the generalized estimating equations (GEE) or a linear mixed-effects model to implement the second step. Through simulation studies our proposal compares favorably to the independence approach (Lee et al., 1993), which ignores the within-cluster correlation in estimating the regression coefficient. Our proposal is easy to implement by using existing softwares.

Algorithms↗

Imputing missing repeated measures data: how should we proceed?

OBJECTIVE: This paper compares six missing data methods that can be used for carrying out statistical tests on repeated measures data: listwise deletion, last value carried forward (LVCF), standardized score imputation, regression and two versions of a closest match method. METHOD: The efficacy of each was investigated under a variety of sample sizes and with differing levels of missingness. Randomly selected samples from a dataset (n = 804) were used to compare the methods using t-tests. Efficacy was defined as the closeness of the estimated t-values to the true t-values from the complete dataset. RESULTS: The results suggest a reliable and efficacious basis for imputation method for repeated measures data is to substitute a missing datum with a value from another individual who has the closest scores on the same variable measured at other timepoints, or the average value of four individuals who have the closest scores on the same variable at other timepoints. The LVCF and standardized score methods performed relatively poorly, which is of concern since these are often recommended. Listwise deletion was also an inefficient missing data method. CONCLUSIONS: Researchers should consider using closest match missing data imputation. Since listwise deletion performed poorly, is widely reported and is the default method in many statistical software packages, the findings have broad implications.

Alcoholism↗

Missing value estimation for DNA microarray gene expression data: local least squares imputation.

MOTIVATION: Gene expression data often contain missing expression values. Effective missing value estimation methods are needed since many algorithms for gene expression data analysis require a complete matrix of gene array values. In this paper, imputation methods based on the least squares formulation are proposed to estimate missing values in the gene expression data, which exploit local similarity structures in the data as well as least squares optimization process. RESULTS: The proposed local least squares imputation method (LLSimpute) represents a target gene that has missing values as a linear combination of similar genes. The similar genes are chosen by k-nearest neighbors or k coherent genes that have large absolute values of Pearson correlation coefficients. Non-parametric missing values estimation method of LLSimpute are designed by introducing an automatic k-value estimator. In our experiments, the proposed LLSimpute method shows competitive results when compared with other imputation methods for missing value estimation on various datasets and percentages of missing values in the data. AVAILABILITY: The software is available at http://www.cs.umn.edu/~hskim/tools.html CONTACT: hpark@cs.umn.edu

Algorithms↗

Multiple imputation method for estimating incidence of HIV infection. The Multicenter Prospective HIV Study.

BACKGROUND: CD4+ T-lymphocyte (CD4) and platelet counts are good predictors of the 'maturity' of HIV infection and can be used to impute the date of infection/seroconversion in individuals for whom this date is unknown. METHODS: Data from the Italian Seroconversion Study were used to develop a Weibull regression model for time since seroconversion as a function of the haematologic markers. The model was used to impute time since HIV infection/seroconversion in individuals from a prevalent cohort, recruited through the Lazio regional HIV surveillance system. RESULTS: The range of the imputed calendar times of infection/seroconversion in 2599 HIV prevalent individuals was 1972-1992; the earliest seroconversions occurred among injecting drug users (IDU). The peak of incidence was reached in 1986 with 340 seroconversions. Among males, the estimated median time from seroconversion to HIV diagnosis was shorter in IDU (30 months) as compared to non-IDU (36 months). This difference was smaller for females (26.6 versus 28.4 in IDU and non-IDU, respectively). CONCLUSIONS: This method permits the estimation of population-based curves of HIV incidence, using data from surveillance. The results support the hypotheses of an early spread of the epidemic among IDU in the Lazio region, and of shorter lead times in this population.

AIDS Serodiagnosis↗

A multiple imputation approach to Cox regression with interval-censored data.

We propose a general semiparametric method based on multiple imputation for Cox regression with interval-censored data. The method consists of iterating the following two steps. First, from finite-interval-censored (but not right-censored) data, exact failure times are imputed using Tanner and Wei's poor man's or asymptotic normal data augmentation scheme based on the current estimates of the regression coefficient and the baseline survival curve. Second, a standard statistical procedure for right-censored data, such as the Cox partial likelihood method, is applied to imputed data to update the estimates. Through simulation, we demonstrate that the resulting estimate of the regression coefficient and its associated standard error provide a promising alternative to the nonparametric maximum likelihood estimate. Our proposal is easily implemented by taking advantage of existing computer programs for right-censored data.

Biometry↗

Multiple imputation and posterior simulation for multivariate missing data in longitudinal studies.

This paper outlines a multiple imputation method for handling missing data in designed longitudinal studies. A random coefficients model is developed to accommodate incomplete multivariate continuous longitudinal data. Multivariate repeated measures are jointly modeled; specifically, an i.i.d. normal model is assumed for time-independent variables and a hierarchical random coefficients model is assumed for time-dependent variables in a regression model conditional on the time-independent variables and time, with heterogeneous error variances across variables and time points. Gibbs sampling is used to draw model parameters and for imputations of missing observations. An application to data from a study of startle reactions illustrates the model. A simulation study compares the multiple imputation procedure to the weighting approach of Robins, Rotnitzky, and Zhao (1995, Journal of the American Statistical Association 90, 106-121) that can be used to address similar data structures.

Acoustic Stimulation↗

Multiple imputation for multivariate data with missing and below-threshold measurements: time-series concentrations of pollutants in the Arctic.

Many chemical and environmental data sets are complicated by the existence of fully missing values or censored values known to lie below detection thresholds. For example, week-long samples of airborne particulate matter were obtained at Alert, NWT, Canada, between 1980 and 1991, where some of the concentrations of 24 particulate constituents were coarsened in the sense of being either fully missing or below detection limits. To facilitate scientific analysis, it is appealing to create complete data by filling in missing values so that standard complete-data methods can be applied. We briefly review commonly used strategies for handling missing values and focus on the multiple-imputation approach, which generally leads to valid inferences when faced with missing data. Three statistical models are developed for multiply imputing the missing values of airborne particulate matter. We expect that these models are useful for creating multiple imputations in a variety of incomplete multivariate time series data sets.

Air Pollutants↗

Multiple imputation methods for estimating regression coefficients in the competing risks model with missing cause of failure.

We propose a method to estimate the regression coefficients in a competing risks model where the cause-specific hazard for the cause of interest is related to covariates through a proportional hazards relationship and when cause of failure is missing for some individuals. We use multiple imputation procedures to impute missing cause of failure, where the probability that a missing cause is the cause of interest may depend on auxiliary covariates, and combine the maximum partial likelihood estimators computed from several imputed data sets into an estimator that is consistent and asymptotically normal. A consistent estimator for the asymptotic variance is also derived. Simulation results suggest the relevance of the theory in finite samples. Results are also illustrated with data from a breast cancer study.

Breast Neoplasms↗

Iterated local least squares microarray missing value imputation.

Microarray gene expression data often contains multiple missing values due to various reasons. However, most of gene expression data analysis algorithms require complete expression data. Therefore, accurate estimation of the missing values is critical to further data analysis. In this paper, an Iterated Local Least Squares Imputation (ILLSimpute) method is proposed for estimating missing values. Two unique features of ILLSimpute method are: ILLSimpute method does not fix a common number of coherent genes for target genes for estimation purpose, but defines coherent genes as those within a distance threshold to the target genes. Secondly, in ILLSimpute method, estimated values in one iteration are used for missing value estimation in the next iteration and the method terminates after certain iterations or the imputed values converge. Experimental results on six real microarray datasets showed that ILLSimpute method performed at least as well as, and most of the time much better than, five most recent imputation methods.

Algorithms↗

Multiple imputation methods for longitudinal blood pressure measurements from the Framingham Heart Study.

Missing data are a great concern in longitudinal studies, because few subjects will have complete data and missingness could be an indicator of an adverse outcome. Analyses that exclude potentially informative observations due to missing data can be inefficient or biased. To assess the extent of these problems in the context of genetic analyses, we compared case-wise deletion to two multiple imputation methods available in the popular SAS package, the propensity score and regression methods. For both the real and simulated data sets, the propensity score and regression methods produced results similar to case-wise deletion. However, for the simulated data, the estimates of heritability for case-wise deletion and the two multiple imputation methods were much lower than for the complete data. This suggests that if missingness patterns are correlated within families, then imputation methods that do not allow this correlation can yield biased results.

Adult↗

Measurement of vaccination coverage at age 24 and 19-35 months: a case study of multiple imputation in public health.

AIM: Childhood immunization coverage in the United States (US) is often measured at age 24 months or, in the National Immunization Survey (NIS) at age of interview, which is between 19 and 35 months. This paper compares these standards. METHODS: Data from the NIS is used to compare immunization coverage at time of interview, retrospectively among all children aged 24 or more months at time of interview, and obtained via multiple imputation (with 10 imputations) for all children, both nationally, by state, and by demographic groups. RESULTS: At the national level, the difference between the 19-35 month estimate and the 24 month complete-case estimate was 1.9 percentage points. For most but not all states and subgroups, the 19-35 month estimate was higher than the 24 month complete-case estimate. The difference between vaccination coverage measured at 19-35 months and 24 months ranged from -2.3 to 7.5 percentage points among states. For three states, the difference between the 19-35 month and 24 month complete-case estimate was more than 6 percentage points, in twelve states there was a 4-6 percentage point difference, and in sixteen states a 2-4 percentage point difference. Conversely, five states had higher 24 month complete-case estimates than 19-35 month estimates. CONCLUSION: We found that the coverages at 19-35 and 24 months differ such that they would rarely be adequate surrogates for one another, particularly at a state level. Multiple imputation, which is easily implemented, increases precision of estimates of coverage at age 24 months.

Journal Article↗

Multiple imputation procedures allow the rescue of missing data: an application to determine serum tumor necrosis factor (TNF) concentration values during the treatment of rheumatoid arthritis patients with anti-TNF therapy.

Longitudinal studies aimed at evaluating patients clinical response to specific therapeutic treatments are frequently summarized in incomplete datasets due to missing data. Multivariate statistical procedures use only complete cases, deleting any case with missing data. MI and MIANALYZE procedures of the SAS software perform multiple imputations based on the Markov Chain Monte Carlo method to replace each missing value with a plausible value and to evaluate the efficiency of such missing data treatment. The objective of this work was to compare the evaluation of differences in the increase of serum TNF concentrations depending on the -308 TNF promoter genotype of rheumatoid arthritis (RA) patients receiving anti-TNF therapy with and without multiple imputations of missing data based on mixed models for repeated measures. Our results indicate that the relative efficiency of our multiple imputation model is greater than 98% and that the related inference was significant (p-value < 0.001). We established that under both approaches serum TNF levels in RA patients bearing the G/A -308 TNF promoter genotype displayed a significantly (p-value < 0.0001) increased ability to produce TNF over time than the G/G patient group, as they received successively doses of anti-TNF therapy.

Antibodies, Monoclonal↗

[Markov Chain Monte Carlo Method of multiple imputation for longitudinal data with missing values in the survey of maternal and children health].

OBJECTIVE: To deal with arbitrary missing pattern in longitudinal data of the Survey of Maternal and Child Health and make the most appropriate inferences with multiple imputation (MI) for further analysis. METHODS: SAS 9.0 was used for Markov Chain Monte Carlo (MCMC) method of MI procedure to impute missing values and combine inferences. RESULTS: The result is acceptable as the data set was imputed 5 times. CONCLUSION: MI is able to solve a variety of problems in missing data sets and to improve the statistical power, especially with the use of MCMC method, for complicated missing data sets.

Bias↗

The effect of imputation procedures on first birth intervals: evidence from five African fertility surveys.

In most African societies there is little motivation to remember dates of demographic events with the level of precision required in demographic surveys. Consequently it is common that the large majority of survey respondents can provide only the calendar year of occurrence or their age at the time of the event. The World Fertility Survey Group decided to handle the problem of poor date reporting by using a computer program to impute the missing information. This article illustrates the effect of these imputation procedures on cross-national differentials in the proportion of premarital first births in Benin, Cameroon, Côte d'Ivoire, Ghana, and Nigeria. The analysis demonstrates that the exceptionally low proportion of premarital first births in Ghana is an artifact of the imputation procedures.

Adolescent↗

A case study on the use of multiple imputation.

Multiple imputation is a relatively new technique for dealing with missing values on items from survey data. Rather than deleting observations for which a value is missing, or assigning a single value to incomplete observations, one replaces each missing item with two or more values. Inferences then can be made with the complete data set. This paper presents an application of multiple imputation using the 1987-1988 National Survey of Families and Households. We impute several binary indicators of whether the respondent's elderly mother/mother-in-law is married. Descriptive statistics are then presented for the sample of adult children with an unmarried mother or mother-in-law.

Adult↗

Imputability of heparin in heparin induced thrombocytopenia: a 10 year experience of the Centre Regional de Pharmacovigilance of Reims-Champagne Ardenne.

The imputability of heparin in heparin induced thrombocytopenia (HIT) was analysed retrospectively in the chart records of 86 cases documented by the Centre Regional de Pharmacovigilance (CRPV) of Reims-Champagne Ardenne over a period of 10 years. Considerable difficulties are encountered in evaluating the degree of imputability. Chronological criteria seem to be determinant in the final imputability score, whereas semiological criteria are particularly difficult to interpret, especially as it is not yet clearly established whether biological tests should be taken into account. The method of assessment requires more precise adaptation to the specific case of HIT and could be improved by redefinition of the criteria in collaboration between pharmacologists and haematologists.

Adult↗

Random regression with imputed values for dropouts.

The random regression model (RRM) has been advocated as a potential solution to problems of statistical analysis posed by dropouts in clinical trials. However, the power of the RRM tests for differences in rates of change can be seriously attenuated by presence of dropouts. The use of imputed scores and other modifications are examined in an attempt to render a simple growth-curve form of the RRM analysis more robust against dropouts. Methods that extrapolate from an individual's own performance were found effective, although inclusion of time-in-treatment as a covariate was documented to be important under identifiable conditions. Of the methods evaluated, those that used group data to impute missing values for dropouts produced nonconservative bias. The results suggest the importance of careful evaluation of potential bias when integrating any group-based imputation procedure into the RRM analyses.

Clinical Trials as Topic↗