Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “imputation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Multiple imputation compared with some informative dropout procedures in the estimation and comparison of rates of change in longitudinal clinical trials with dropouts.

Statistical analysis based on multiple imputation (MI) of missing data when analyzing data with missing observations is gaining popularity among statisticians because of availability of computing softwares; it might be tempting to use MI whenever data is missing. An important assumption behind MI is the "ignorability of missingness." In this paper, we demonstrate the use of MI in conjunction with random effects models and several other methods that are devised to handle nonignorable missingness (informative dropouts). We then compare the results to assess sensitivity to underlying assumptions. Our focus is primarily to estimate and compare rates of change (of a primary variable). The application dataset has a high dropout rate and has features to suggest informativeness of the dropout process. The estimates obtained under random effects modeling with multiple imputation were found to differ substantially from those obtained by methods devised to handle informative dropouts.

Algorithms↗

Numerical equivalence of imputing scores and weighted estimators in regression analysis with missing covariates.

Imputation, weighting, direct likelihood, and direct Bayesian inference (Rubin, 1976) are important approaches for missing data regression. Many useful semiparametric estimators have been developed for regression analysis of data with missing covariates or outcomes. It has been established that some semiparametric estimators are asymptotically equivalent, but it has not been shown that many are numerically the same. We applied some existing methods to a bladder cancer case-control study and noted that they were the same numerically when the observed covariates and outcomes are categorical. To understand the analytical background of this finding, we further show that when observed covariates and outcomes are categorical, some estimators are not only asymptotically equivalent but also actually numerically identical. That is, although their estimating equations are different, they lead numerically to exactly the same root. This includes a simple weighted estimator, an augmented weighted estimator, and a mean-score estimator. The numerical equivalence may elucidate the relationship between imputing scores and weighted estimation procedures.

Case-Control Studies↗

Imputing nonresponses to mail-back questionnaires.

Many mail-back questionnaires are expected at the outset to elicit poor response rates, perhaps as low as 15-30%. Corrections can be designed into such a survey by using either two or three mailouts of the questionnaire at regular intervals. Assuming a trend in responses as a function of the number of mailouts a person receives before filling out and mailing back the questionnaire, responses are imputed for those who do not mail back the questionnaire after the final mailout. Standard errors are derived, and an example is included. The imputation is easily programmed. A validation of this method is also included.

Analysis of Variance↗

On imputing function to structure from the behavioural effects of brain lesions.

What is the link, if any, between the patterns of connections in the brain and the behavioural effects of localized brain lesions? We explored this question in four related ways. First, we investigated the distribution of activity decrements that followed simulated damage to elements of the thalamocortical network, using integrative mechanisms that have recently been used to successfully relate connection data to information on the spread of activation, and to account simultaneously for a variety of lesion effects. Second, we examined the consequences of the patterns of decrement seen in the simulation for each type of inference that has been employed to impute function to structure on the basis of the effects of brain lesions. Every variety of conventional inference, including double dissociation, readily misattributed function to structure. Third, we tried to derive a more reliable framework of inference for imputing function to structure, by clarifying concepts of function, and exploring a more formal framework, in which knowledge of connectivity is necessary but insufficient, based on concepts capable of mathematical specification. Fourth, we applied this framework to inferences about function relating to a simple network that reproduces intact, lesioned and paradoxically restored orientating behaviour. Lesion effects could be used to recover detailed and reliable information on which structures contributed to particular functions in this simple network. Finally, we explored how the effects of brain lesions and this formal approach could be used in conjunction with information from multiple neuroscience methodologies to develop a practical and reliable approach to inferring the functional roles of brain structures.

Behavior↗

A comparison of imputation techniques for handling missing data.

Researchers are commonly faced with the problem of missing data. This article presents theoretical and empirical information for the selection and application of approaches for handling missing data on a single variable. An actual data set of 492 cases with no missing values was used to create a simulated yet realistic data set with missing at random (MAR) data. The authors compare and contrast five approaches (listwise deletion, mean substitution, simple regression, regression with an error term, and the expectation maximization [EM] algorithm) for dealing with missing data, and compare the effects of each method on descriptive statistics and correlation coefficients for the imputed data (n = 96) and the entire sample (n = 492) when imputed data are inculded. All methods had limitations, although our findings suggest that mean substitution was the least effective and that regression with an error term and the EM algorithm produced estimates closest to those of the original variables.

Algorithms↗

Multiple imputation: a primer.

In recent years, multiple imputation has emerged as a convenient and flexible paradigm for analysing data with missing values. Essential features of multiple imputation are reviewed, with answers to frequently asked questions about using the method in practice.

Bayes Theorem↗

Imputation of a true endpoint from a surrogate: application to a cluster randomized controlled trial with partial information on the true endpoint.

BACKGROUND: The Anglia Menorrhagia Education Study (AMES) is a randomized controlled trial testing the effectiveness of an education package applied to general practices. Binary data are available from two sources; general practitioner reported referrals to hospital, and referrals to hospital determined by independent audit of the general practices. The former may be regarded as a surrogate for the latter, which is regarded as the true endpoint. Data are only available for the true end point on a sub set of the practices, but there are surrogate data for almost all of the audited practices and for most of the remaining practices. METHODS: The aim of this paper was to estimate the treatment effect using data from every practice in the study. Where the true endpoint was not available, it was estimated by three approaches, a regression method, multiple imputation and a full likelihood model. RESULTS: Including the surrogate data in the analysis yielded an estimate of the treatment effect which was more precise than an estimate gained from using the true end point data alone. CONCLUSIONS: The full likelihood method provides a new imputation tool at the disposal of trials with surrogate data.

Bayes Theorem↗

Direct likelihood analysis versus simple forms of imputation for missing data in randomized clinical trials.

BACKGROUND: In many clinical trials, data are collected longitudinally over time. In such studies, missingness, in particular dropout, is an often encountered phenomenon. METHODS: We discuss commonly used but often problematic methods such as complete case analysis and last observation carried forward and contrast them with broadly valid and easy to implement direct-likelihood methods. We comment on alternatives such as multiple imputation and the expectation-maximization algorithm. RESULTS: We apply these methods in particular to data from a study with continuous outcomes. The outcomes are modelled using a general linear mixed-effects model. The bias with CC and LOCF is established in the case study and the advantages of the direct-likelihood approach shown. CONCLUSIONS: We have established formal but easy to understand arguments for a shift towards a direct-likelihood paradigm when analysing incomplete data from longitudinal clinical trials, necessitating neither imputation nor deletion.

Data Collection↗

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans↗

Value of monitoring pulse oximetry for imputability of patent foramen ovale in transient dyspnoea.

Diagnosis of patent foramen ovale (PFO) is commonly made by echocardiography with contrast injection. PFO can be responsible for a transient right-to-left shunting with paroxysmal dyspnoea but punctual measurements of oxygen saturation may fail to detect arterial desaturations. Thus, claiming the imputability of PFO in dyspnoeic symptoms remains difficult. We report on the case of a 64-year-old man presenting an intermittent disabilitating dyspnoea, for which the pulse oximetry monitoring allowed to impute symptoms to the right-to-left shunting through the PFO and influenced the decision of percutaneous closure.

Dyspnea↗

An analysis of alternative imputation strategies for individuals with partial data in the National Medical Care Expenditure Survey.

Data collection in the National Medical Care Expenditure Survey was applied to the same panel of sample households in six rounds of interviewing, with 1977 as the reference period. Approximately 11 percent of all survey participants provided data for only part of the time they were eligible to respond. To allow for national estimates of relevant health parameters, the data for the partial participants must be adjusted for the entire time frame for which they were eligible. Consequently, three alternative imputation strategies were considered for implementation: a weighted adjustment to the partial data, a substitution of data from complete participants who matched the partial respondents on relevant demographic characteristics, and use of only the data from participants with complete information to characterize the nation. To determine the optimal strategy, a controlled experiment was conducted by artificially creating partial data for participants with complete information, then adjusting the synthetically produced partial data by the three imputation strategies.

Analysis of Variance↗

The use of multiple imputation for the analysis of missing data.

This article provides a comprehensive review of multiple imputation (MI), a technique for analyzing data sets with missing values. Formally, MI is the process of replacing each missing data point with a set of m > 1 plausible values to generate m complete data sets. These complete data sets are then analyzed by standard statistical software, and the results combined, to give parameter estimates and standard errors that take into account the uncertainty due to the missing data values. This article introduces the idea behind MI, discusses the advantages of MI over existing techniques for addressing missing data, describes how to do MI for real problems, reviews the software available to implement MI, and discusses the results of a simulation study aimed at finding out how assumptions regarding the imputation model affect the parameter estimates provided by MI.

Bias↗

[Trauma and imputability].

The imputability of a bodily damage to a traumatism and the bond of causality are at the basis of the mission of the medical expert in the procedure that leads to indemnification in reparation of a suffered prejudice. The aim of this paper is to enumerate the bases on which the expert grounds to prove the imputability and the bond of causality. He will also insist on the final role of the jurist who, grounding his opinion on the elements of the expert appraisement, will establish the causality and hence the responsibility and the level of reparation.

Belgium↗

[Jurisdiction and imputability].

Validity, efficacy and responsibility of acts depend on the intelligence and will of the acting subject; therefore when they are reduced or debilitated, these acts may be declared as non-valid and the author, not-responsible for the acts. Some neurological pathologies may generate physical and/or psychic permanent deficiencies, which prevent subjects from acting on their own. For these cases, the law establishes the incapacity state, in order to protect the disabled and complete the reduced ability, guaranteeing their rights and security. The disabled state will be determined by a legal sentence, which states the lack of ability to manage. In that sentence extension and limits of the disability will be determined; disability level will be proportional to the insight degree.Similarly, a subject suffering a pathological condition that invalidates his/her will and intelligence will be considered non-responsible and not imputable, since there is no culpability ability. The Penal Code establishes the criteria that will determine the possibility of imputability or its absence, as well as modifying circumstances.

Persons with Disabilities↗

Multiple imputation methods for the missing covariates in generalized estimating equation.

This paper discusses the missing covariates problem in the generalized estimating equation (GEE) model. Estimates by various multiple imputation techniques (MI) are examined and compared to the sample average imputation method (SA) through simulations and an example. The simulation results show that, under the correct model specification, the MI estimators have negligible bias and have fairly similar efficiencies as the SA estimator. A practical advantage of the MI estimates is that the standard errors can be more easily computed than the SA estimates.

Aged↗

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability↗

An imputation method for non-ignorable missing data in studies of blood pressure.

In studies with repeated measures of blood pressure (BP), particularly in trials of hypertension prevention, BP measurements often become censored once a participant commences antihypertensive medication. When prescribed by non-study physicians under uncontrolled conditions, the missing data mechanism is non-ignorable and may bias the BP effects of interest. I propose a method that models the distribution of BPs measured by non-study physicians and their relation to study BPs using random effects models. If treated for hypertension, I assume that BP measured outside the study is greater than a clinical cutpoint, such as diastolic BP > or = 90 mmHg. I then compute estimates for the missing study BPs conditional on previously observed study BPs and treatment for hypertension. Multiple imputation is used to model the variability of the BP values and adjust the standard error estimates of the parameters. Examples are given using simulated data and data from the weight loss intervention of phase I of the Trials of Hypertension Prevention.

Antihypertensive Agents↗

Multiple imputation for simple estimation of the hazard function based on interval censored data.

A data augmentation algorithm is presented for estimating the hazard function and pointwise variability intervals based on interval censored data. The algorithm extends that proposed by Tanner and Wong for grouped right censored data to interval censored data. It applies multiple imputation and local likelihood methods to obtain smooth non-parametric estimates for the hazard function. This approach considerably simplifies the problem of estimation for interval censored data as it transforms it into the more tractable problem of estimation for right censored data. The method is illustrated for two real data sets: times to breast cosmesis deterioration and times to HIV-1 infection for individuals with haemophilia. Simulations are presented to assess the effects of various parameters on the estimates and their variances.

Algorithms↗