Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Analysis of change in the presence of informative censoring: application to a longitudinal clinical trial of progressive renal disease.

The rate of change in a continuous variable, measured serially over time, is often used as an outcome in longitudinal studies or clinical trials. When patients terminate the study before the scheduled end of the study, there is a potential for bias in estimation of rate of change using standard methods which ignore the missing data mechanism. These methods include the use of unweighted generalized estimating equations methods and likelihood-based methods assuming an ignorable missing data mechanism. We present a model for analysis of informatively censored data, based on an extension of the two-stage linear random effects model, where each subject's random intercept and slope are allowed to be associated with an underlying time to event. The joint distribution of the continuous responses and the time-to-event variable are then estimated via maximum likelihood using the EM algorithm, and using the bootstrap to calculate standard errors. We illustrate this methodology and compare it to simpler approaches and usual maximum likelihood using data from a multi-centre study of the effects of diet and blood pressure control on progression of renal disease, the Modification of Diet in Renal Disease (MDRD) Study. Sensitivity analyses and simulations are used to evaluate the performance of this methodology in the context of the MDRD data, under various scenarios where the drop-out mechanism is ignorable as well as non-ignorable.

Algorithms↗

Efficacy and safety of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with active psoriatic arthritis: 52-week results from the randomised, double-blind, placebo-controlled phase 3 POETYK PsA-1 trial.

OBJECTIVES: The randomised, double-blind, placebo-controlled, phase 3 Program fOr Evaluation of TYK2 inhibitor Psoriatic Arthritis-1 (POETYK PsA-1) trial evaluated the efficacy, safety, and tolerability of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with PsA na&#xef;ve to biologic disease-modifying antirheumatic drugs. METHODS: Adults with active PsA, high-sensitivity C-reactive protein concentration &#x2265; 3 mg/L, and &#x2265; 1 PsA-related hand and/or foot erosion detectable via radiograph were randomised 1:1 to oral deucravacitinib 6 mg once daily or placebo through week (W) 16. At W16, patients continued receiving deucravacitinib or switched from placebo to deucravacitinib through W52. The primary endpoint was American College of Rheumatology 20% improvement in response (ACR20) at W16. Nonresponder imputation was used for missing data. Efficacy and safety were evaluated through W52. Post hoc rank analysis of covariance was used to evaluate structural damage with no missing data imputation. RESULTS: In 670 patients, a significantly greater proportion of those receiving deucravacitinib vs placebo achieved ACR20 at W16 (54.2% vs 34.1%, P < .001). Responses with deucravacitinib were increased at W52. Patients who switched from placebo to deucravacitinib achieved improvements similar to those in patients who received continuous deucravacitinib. Inhibition of structural damage was observed at W16 and W52. At W16, incidences of serious adverse events (AEs) (deucravacitinib, 1.8%; placebo, 2.4%) and discontinuations due to AEs (2.4%; 1.8%) were low and remained low through W52, without imbalances in cardiovascular events, malignancies, or opportunistic infections. No new safety signals were detected; no deaths occurred. CONCLUSIONS: Deucravacitinib demonstrated superiority vs placebo for clinical responses, patient-reported outcomes, and structural damage inhibition in patients with PsA, with favourable tolerability and safety.

Humans↗

Sensitivity analysis for pattern mixture models.

Incomplete series of data is a common feature in quality-of-life studies, in particular in chronic diseases where attrition of patients is high. Two alternative approaches to modeling longitudinal data with incomplete measurements have frequently been proposed in the literature, selection models and pattern-mixture models. In this paper we focus on, by way of sensitivity analysis, extrapolating incomplete patterns using identifying restrictions. Perhaps the best known ones are so-called complete case missing value restrictions (CCMV), where for a given pattern, the conditional distribution of the missing data, given the observed data, is equated to its counterpart in the completers. Available case missing value (ACMV) restrictions equate this conditional density to the one calculated from the subgroup of all patterns for which all required components have been observed. Neighboring case missing value restrictions (NCMV) equate this conditional density to the one calculated from the the pattern with one additional measurement obtained. In this paper, these three identifying restriction strategies are used to multiply impute missing data in a study in metastatic prostate cancer. Multiple imputation is employed to reduce the uncertainty of single imputation. It is shown how hypothesis testing and sensitivity analyses are carried out in this setting.

Humans↗

Catquest questionnaire for use in cataract surgery care: assessment of surgical outcomes.

PURPOSE: To demonstrate the outcome for patients after cataract extraction using the Catquest cataract questionnaire and discuss the models validity in assessing outcome. SETTING: Thirty-five Swedish departments of ophthalmology. METHODS: Patients having cataract extraction performed by surgeons from 35 Swedish departments of opthalmology participated in the study. The questionnaire was given to 2970 consecutive patients having surgery during March 1995 at the participating surgical units. The questionnaire was sent by mail to patients and completed on a voluntary basis. It focuses on visual disabilities in daily life, activity level, cataract symptoms, and degree of independence. The results form the questionnaire are interpreted using a benefit matrix that credits not only a decrease in visual disabilities and cataract symptoms but also an improvement in or maintenance of a preoperative activity level. RESULTS: Complete surgical outcome data and completed preoperative and postoperative questionnaires were available in 1933 cases (65.1%). Benefit from surgery according to the model was achieved by 90.9% of the patients. Patients having their second cataract extraction had the highest frequency of the greatest benefit form surgery. There was good agreement between the different levels of benefit from surgery according to the model and the patient's global rating of his or her vision or achieved visual acuity after surgery, respectively. Patients with missing data (did not return postoperative questionnaire or had missing surgical result variables) were older and had a higher frequency of other diseases and handicaps. CONCLUSION: The Catquest cataract questionnaire allowed the outcome of cataract surgery to be graded by different levels of benefit. There seemed to be good agreement between this model of assessment and the patient's global rating of his or her vision. Missing data may be a problem when a postal questionnaire is used.

Activities of Daily Living↗

Life satisfaction following spinal cord injury: long-term follow-up.

OBJECTIVE: To determine the course of self-reported life satisfaction in a spinal cord injury (SCI) cohort. DESIGN: Prospective study using longitudinal data from the Injury Control Research Center. PARTICIPANTS: Adult persons with traumatic-onset SCI (n = 207) evaluated at 1, 2, 4, and 5 years postinjury using the Life Satisfaction Index-A. RESULTS: A nonsignificant (P > 0.05) main effect of time was found using a repeated-measures analysis controlling for education and employment status. Several methods were used that provided a range of liberal to conservative estimates for missing data (ie, 38% retention rate at year 5). Subsequent missing data analyses tended to corroborate the finding of a nonsignificant effect of time, although the most conservative methods showed a significant decrease in life satisfaction between year 1 and year 5 postinjury (P < 0.05). Examination of numerous demographic, injury, and treatment-related characteristics at each follow-up time point suggested that the main findings of the study were not merely the result of differential dropout rates. CONCLUSION: Life satisfaction after the first year of injury remains largely the same over the next 4 years. Methodologic and analytic recommendations are discussed.

Adolescent↗

Ranking the 'balance' of state long-term care systems: a comparative exposition of the SMARTER and CaRBS Techniques.

The use of attribute sets to rank units of health provision (e.g., states, organizations) against policy goals is an essential task within decision-making and analysis. This paper elucidates and compares two techniques, SMARTER (Simple Multiattribute Rating Technique Exploiting Ranks) and CaRBS (Classification and Ranking Belief Simplex), within an expositional ranking of US states' long-term care (LTC) systems against the policy goal of providing a balance between (traditionally dominant) institutional care, and alternative home and community-based services (HCBS). While the (more established) SMARTER technique is used primarily for comparative purposes, greater emphasis is placed on elucidating CaRBS which is based on the Dempster-Shafer theory of evidence. It is shown that CaRBS offers four appealing features for health policy analysis: (1) the capacity to rank using either of two confidence measures (DST-related belief and plausibility values), (2) a systematic approach to managing missing data, (3) the production of stable rankings, and (4) the simplex plot method of data representation. In addition to discussing the LTC policy implications of the study findings, the issues of rank order stability and the management of missing data are discussed with respect to the two techniques employed.

Health Services Accessibility↗

A simulation study of the effects of assignment of prior identity-by-descent probabilities to unselected sib pairs, in covariance-structure modeling of a quantitative-trait locus.

Sib pair-selection strategies, designed to identify the most informative sib pairs in order to detect a quantitative-trait locus (QTL), give rise to a missing-data problem in genetic covariance-structure modeling of QTL effects. After selection, phenotypic data are available for all sibs, but marker data-and, consequently, the identity-by-descent (IBD) probabilities-are available only in selected sib pairs. One possible solution to this missing-data problem is to assign prior IBD probabilities (i.e., expected values) to the unselected sib pairs. The effect of this assignment in genetic covariance-structure modeling is investigated in the present paper. Two maximum-likelihood approaches to estimation are considered, the pi-hat approach and the IBD-mixture approach. In the simulations, sample size, selection criteria, QTL-increaser allele frequency, and gene action are manipulated. The results indicate that the assignment of prior IBD probabilities results in serious estimation bias in the pi-hat approach. Bias is also present in the IBD-mixture approach, although here the bias is generally much smaller. The null distribution of the log-likelihood ratio (i.e., in absence of any QTL effect) does not follow the expected null distribution in the pi-hat approach after selection. In the IBD-mixture approach, the null distribution does agree with expectation.

Alleles↗

An eigenvector method for estimating item parameters of the dichotomous and polytomous Rasch models.

The purpose of this paper is to describe a technique for obtaining item parameters of the Rasch model, a technique in which the item parameters are extracted from the eigenvector of a matrix derived from comparisons between pairs of items. The technique can be applied to both dichotomous and polytomous data. In application to a previously published data set, it is shown that the technique provides item parameter estimates comparable to those produced by joint maximum likelihood estimation, and for the most difficult items, the technique appears to produce superior estimates. This method has several advantages. It easily accommodates missing data, and makes transparent the basis for item parameter estimation in the presence of missing data. Furthermore, the method provides a link to other methods in the social sciences and, in particular, provides the framework for application of graph theory to the analysis of assessment networks. Finally, it exploits several characteristics that are unique to the Rasch model.

Algorithms↗

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis↗

An illness-death stochastic model in the analysis of longitudinal dementia data.

A significant source of missing data in longitudinal epidemiological studies on elderly individuals is death. Subjects in large scale community-based longitudinal dementia studies are usually evaluated for disease status in study waves, not under continuous surveillance as in traditional cohort studies. Therefore, for the deceased subjects, disease status prior to death cannot be ascertained. Statistical methods assuming deceased subjects to be missing at random may not be realistic in dementia studies and may lead to biased results. We propose a stochastic model approach to simultaneously estimate disease incidence and mortality rates. We set up a Markov chain model consisting of three states, non-diseased, diseased and dead, and estimate the transition hazard parameters using the maximum likelihood approach. Simulation results are presented indicating adequate performance of the proposed approach.

Aged↗

Imputation methods to improve inference in SNP association studies.

Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we develop two haplotype-based imputation approaches and one tree-based imputation approach for association studies. The emphasis is to evaluate the impact of imputation on parameter estimation, compared to the standard practice of ignoring missing data. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to iteratively sample from a full conditional distribution, which is obtained from the classification and regression tree (CART) algorithm. We employ a standard multiple imputation procedure to account for the uncertainty of imputation. We apply the methods to simulated data as well as a case-control study on developmental dyslexia. Our results suggest that imputation generally improves efficiency over the standard practice of ignoring missing data. The tree-based approach performs comparably well as haplotype-based approaches, but the former has a computational advantage. The WEM approach yields the smallest bias at a price of increased variance.

Algorithms↗

Clinical significance of antiviral therapy for episodic treatment of herpes labialis: exploratory analyses of the combined data from two valaciclovir trials.

Valaciclovir (Valtrex) 2 g twice daily for 1 day was recently approved in the United States for treatment of cold sores. In order to apply more clinically relevant assumptions to the analysis, we examined the effect of different missing data and endpoint assumptions on apparent valaciclovir efficacy. Results of each analysis demonstrate statistically significant increases in the proportion of subjects whose cold sores were aborted with valaciclovir compared with placebo, and significant decreases in healing times for subjects with cold sore lesions who were treated with valaciclovir compared with placebo. These exploratory analyses provide evidence of the robustness of the results to differing missing data assumptions and show that use of more clinically relevant endpoint assumptions increases the magnitude of some therapeutic responses. We also introduce a new measure that combines the two observed drug effects (reduced lesion duration, increased aborted lesions) into a single endpoint that captures the global benefit of the drug to the patient.

Acyclovir↗

A comparison of imputation methods in a longitudinal randomized clinical trial.

It is common for longitudinal clinical trials to face problems of item non-response, unit non-response, and drop-out. In this paper, we compare two alternative methods of handling multivariate incomplete data across a baseline assessment and three follow-up time points in a multi-centre randomized controlled trial of a disease management programme for late-life depression. One approach combines hot-deck (HD) multiple imputation using a predictive mean matching method for item non-response and the approximate Bayesian bootstrap for unit non-response. A second method is based on a multivariate normal (MVN) model using PROC MI in SAS software V8.2. These two methods are contrasted with a last observation carried forward (LOCF) technique and available-case (AC) analysis in a simulation study where replicate analyses are performed on subsets of the originally complete cases. Missing-data patterns were simulated to be consistent with missing-data patterns found in the originally incomplete cases, and observed complete data means were taken to be the targets of estimation. Not surprisingly, the LOCF and AC methods had poor coverage properties for many of the variables evaluated. Multiple imputation under the MVN model performed well for most variables but produced less than nominal coverage for variables with highly skewed distributions. The HD method consistently produced close to nominal coverage, with interval widths that were roughly 7 per cent larger on average than those produced from the MVN model.

Aged↗

Imputation of missing values is superior to complete case analysis and the missing-indicator method in multivariable diagnostic research: a clinical example.

BACKGROUND AND OBJECTIVES: To illustrate the effects of different methods for handling missing data--complete case analysis, missing-indicator method, single imputation of unconditional and conditional mean, and multiple imputation (MI)--in the context of multivariable diagnostic research aiming to identify potential predictors (test results) that independently contribute to the prediction of disease presence or absence. METHODS: We used data from 398 subjects from a prospective study on the diagnosis of pulmonary embolism. Various diagnostic predictors or tests had (varying percentages of) missing values. Per method of handling these missing values, we fitted a diagnostic prediction model using multivariable logistic regression analysis. RESULTS: The receiver operating characteristic curve area for all diagnostic models was above 0.75. The predictors in the final models based on the complete case analysis, and after using the missing-indicator method, were very different compared to the other models. The models based on MI did not differ much from the models derived after using single conditional and unconditional mean imputation. CONCLUSION: In multivariable diagnostic research complete case analysis and the use of the missing-indicator method should be avoided, even when data are missing completely at random. MI methods are known to be superior to single imputation methods. For our example study, the single imputation methods performed equally well, but this was most likely because of the low overall number of missing values.

Adult↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗

The effect of question structure on self-reports of heavy drinking: closed-ended versus open-ended questions.

OBJECTIVE: We compared open-ended versus closed-ended questions on the frequency of consuming five or more drinks in a single sitting. METHOD: From a general population survey of Ontario adults (N = 2,022, 62% male), we analyzed a subsample of 649 respondents who reported drinking five or more drinks in a single sitting at least once in the past year. Differences in agreement between the two questions and rates of missing data were evaluated. RESULTS: For the most part, the two measures were not consistent, with the closed-ended question eliciting higher rates of heavier drinking. Rates of missing data were also higher for the open-ended question. CONCLUSIONS: Open-ended question may not necessarily be more suitable than closed-ended questions for estimating the frequency of heavy alcohol use.

Adult↗

Parametric and nonparametric linkage analysis: a unified multipoint approach.

In complex disease studies, it is crucial to perform multipoint linkage analysis with many markers and to use robust nonparametric methods that take account of all pedigree information. Currently available methods fall short in both regards. In this paper, we describe how to extract complete multipoint inheritance information from general pedigrees of moderate size. This information is captured in the multipoint inheritance distribution, which provides a framework for a unified approach to both parametric and nonparametric methods of linkage analysis. Specifically, the approach includes the following: (1) Rapid exact computation of multipoint LOD scores involving dozens of highly polymorphic markers, even in the presence of loops and missing data. (2) Non-parametric linkage (NPL) analysis, a powerful new approach to pedigree analysis. We show that NPL is robust to uncertainty about mode of inheritance, is much more powerful than commonly used nonparametric methods, and loses little power relative to parametric linkage analysis. NPL thus appears to be the method of choice for pedigree studies of complex traits. (3) Information-content mapping, which measures the fraction of the total inheritance information extracted by the available marker data and points out the regions in which typing additional markers is most useful. (4) Maximum-likelihood reconstruction of many-marker haplotypes, even in pedigrees with missing data. We have implemented NPL analysis, LOD-score computation, information-content mapping, and haplotype reconstruction in a new computer package, GENEHUNTER. The package allows efficient multipoint analysis of pedigree data to be performed rapidly in a single user-friendly environment.

Algorithms↗

Active life expectancy from annual follow-up data with missing responses.

Active life expectancy (ALE) at a given age is defined as the expected remaining years free of disability. In this study, three categories of health status are defined according to the ability to perform activities of daily living independently. Several studies have used increment-decrement life tables to estimate ALE, without error analysis, from only a baseline and one follow-up interview. The present work conducts an individual-level covariate analysis using a three-state Markov chain model for multiple follow-up data. Using a logistic link, the model estimates single-year transition probabilities among states of health, accounting for missing interviews. This approach has the advantages of smoothing subsequent estimates and increased power by using all follow-ups. We compute ALE and total life expectancy from these estimated single-year transition probabilities. Variance estimates are computed using the delta method. Data from the Iowa Established Population for the Epidemiologic Study of the Elderly are used to test the effects of smoking on ALE on all 5-year age groups past 65 years, controlling for sex and education.

Aged↗