Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Item response models for longitudinal quality of life data in clinical trials.

Assessment of quality of life is becoming standard in clinical trials. A popular method for measuring quality of life is with instruments which utilize multiple-item subscales, in which each item is scored on a Likert scale. Most statistical methods for the analysis of quality of life data in clinical trials do not explicity consider the properties and psychometric features which were of interest in scale development. In this regard, the measurement and statistical summarization of quality of life data, along with the clinical interpretation, can be somewhat disjoint from the psychometric concerns of the development process. The aim of this paper is to address the complicated issues present in analysing multiple-item ordinal quality of life data in clinical trials while maintaining fidelity to the psychometrical foundations upon which quality of life instruments are built. Accomplishing this will require the development of item response models which recognize the longitudinal aspects of clinical trial designs as well as the potential problem of informatively missing data. A general item response modeling approach is presented for longitudinal multiple-item quality of life data measured on ordinal scales with model components for missing data mechanisms and latent trait regression on treatment indicators and other covariates.

Clinical Trials as Topic↗

Single-channel data and missed events: analysis of a two-state Markov model.

Patch-clamp recording permits investigation of the gating kinetics of single ion channels. Careful statistical analysis of kinetic data can yield clues as to the molecular events underlying channel gating. However, it is important that such analysis should take full account of the limitations that arise from the finite time resolution of patch-clamp recording techniques. Single-ion-channel data are generally interpreted in terms of Markov process models of channel gating mechanisms. Experimental channel records suffer from time interval omission, i.e. failure to detect brief channel openings and closings. This leads to an identifiability problem when analysing single-channel data, i.e. different gating mechanisms provide equally convincing descriptions of the same experimental data. We consider a two-state Markov model of receptor-channel gating in which the channel opening rate is proportional to the agonist concentration, C in equilibrium with OA. By using computer-simulated data, the approximate likelihood of the data is maximized to yield parameter estimates for the model. At a single agonist concentration there is an identifiability problem in that two pairs of parameter estimates are obtained. The 'true' parameter estimates cannot be distinguished from the 'false' ones. By considering data corresponding to a range of agonist concentrations one may identify the 'true' parameter estimates as those that do not change as the agonist concentration is increased. Alternatively, one may identify the 'true' parameter estimates directly by maximizing a global likelihood, the latter being obtained by simultaneous consideration of data obtained at several different agonist concentrations.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Estimating equations with nonignorably missing response data.

Troxel, Lipsitz, and Brennan (1997, Biometrics 53, 857-869) considered parameter estimation from survey data with nonignorable nonresponse and proposed weighted estimating equations to remove the biases in the complete-case analysis that ignores missing observations. This paper suggests two alternative modifications for unbiased estimation of regression parameters when a binary outcome is potentially observed at successive time points. The weighting approach of Robins, Rotnitzky, and Zhao (1995, Journal of the American Statistical Association 90, 106-121) is also modified to obtain unbiased estimating functions. The suggested estimating functions are unbiased only when the missingness probability is correctly specified, and misspecification of the missingness model will result in biases in the estimates. Simulation studies are carried out to assess the performance of different methods when the covariate is binary or normal. For the simulation models used, the relative efficiency of the two new methods to the weighting methods is about 3.0 for the slope parameter and about 2.0 for the intercept parameter when the covariate is continuous and the missingness probability is correctly specified. All methods produce substantial biases in the estimates when the missingness model is misspecified or underspecified. Analysis of data from a medical survey illustrates the use and possible differences of these estimating functions.

Biometry↗

Multivariate outlier detection applied to multiply imputed laboratory data.

In clinical laboratory safety data, multivariate outlier detection methods may highlight a patient whose laboratory measurements do not follow the same pattern of relationships as the majority of patients, although their individual measurements are not found to be outlying when considered one at a time. Missing data problems are often dealt with by imputing a single value as an estimate of the missing value. The completed data set may then be analysed using traditional methods. A disadvantage of using single imputation is the underestimation of variability, with a corresponding distortion of power in hypothesis testing. Multiple imputation methods attempt to overcome this problem, and in this paper a study is described which considers the application of multivariate outlier detection methods to multiply imputed clinical laboratory safety data sets. Three different proportions of missing data are generated in laboratory data sets of dimensions 4, 7, 12 and 30, and a comparison of eight multiple imputation methods is carried out. Two outlier detection techniques, Mahalanobis distance and generalized principal component analysis, are applied to the multiply imputed data sets, and their performances are discussed. Measures are introduced for assessing the accuracy of the missing data results, depending on which method of analysis is used.

Algorithms↗

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗

Analysis of change in the presence of informative censoring: application to a longitudinal clinical trial of progressive renal disease.

The rate of change in a continuous variable, measured serially over time, is often used as an outcome in longitudinal studies or clinical trials. When patients terminate the study before the scheduled end of the study, there is a potential for bias in estimation of rate of change using standard methods which ignore the missing data mechanism. These methods include the use of unweighted generalized estimating equations methods and likelihood-based methods assuming an ignorable missing data mechanism. We present a model for analysis of informatively censored data, based on an extension of the two-stage linear random effects model, where each subject's random intercept and slope are allowed to be associated with an underlying time to event. The joint distribution of the continuous responses and the time-to-event variable are then estimated via maximum likelihood using the EM algorithm, and using the bootstrap to calculate standard errors. We illustrate this methodology and compare it to simpler approaches and usual maximum likelihood using data from a multi-centre study of the effects of diet and blood pressure control on progression of renal disease, the Modification of Diet in Renal Disease (MDRD) Study. Sensitivity analyses and simulations are used to evaluate the performance of this methodology in the context of the MDRD data, under various scenarios where the drop-out mechanism is ignorable as well as non-ignorable.

Algorithms↗

Efficacy and safety of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with active psoriatic arthritis: 52-week results from the randomised, double-blind, placebo-controlled phase 3 POETYK PsA-1 trial.

OBJECTIVES: The randomised, double-blind, placebo-controlled, phase 3 Program fOr Evaluation of TYK2 inhibitor Psoriatic Arthritis-1 (POETYK PsA-1) trial evaluated the efficacy, safety, and tolerability of deucravacitinib, an oral, selective tyrosine kinase 2 inhibitor, in patients with PsA na&#xef;ve to biologic disease-modifying antirheumatic drugs. METHODS: Adults with active PsA, high-sensitivity C-reactive protein concentration &#x2265; 3 mg/L, and &#x2265; 1 PsA-related hand and/or foot erosion detectable via radiograph were randomised 1:1 to oral deucravacitinib 6 mg once daily or placebo through week (W) 16. At W16, patients continued receiving deucravacitinib or switched from placebo to deucravacitinib through W52. The primary endpoint was American College of Rheumatology 20% improvement in response (ACR20) at W16. Nonresponder imputation was used for missing data. Efficacy and safety were evaluated through W52. Post hoc rank analysis of covariance was used to evaluate structural damage with no missing data imputation. RESULTS: In 670 patients, a significantly greater proportion of those receiving deucravacitinib vs placebo achieved ACR20 at W16 (54.2% vs 34.1%, P < .001). Responses with deucravacitinib were increased at W52. Patients who switched from placebo to deucravacitinib achieved improvements similar to those in patients who received continuous deucravacitinib. Inhibition of structural damage was observed at W16 and W52. At W16, incidences of serious adverse events (AEs) (deucravacitinib, 1.8%; placebo, 2.4%) and discontinuations due to AEs (2.4%; 1.8%) were low and remained low through W52, without imbalances in cardiovascular events, malignancies, or opportunistic infections. No new safety signals were detected; no deaths occurred. CONCLUSIONS: Deucravacitinib demonstrated superiority vs placebo for clinical responses, patient-reported outcomes, and structural damage inhibition in patients with PsA, with favourable tolerability and safety.

Humans↗

Sensitivity analysis for pattern mixture models.

Incomplete series of data is a common feature in quality-of-life studies, in particular in chronic diseases where attrition of patients is high. Two alternative approaches to modeling longitudinal data with incomplete measurements have frequently been proposed in the literature, selection models and pattern-mixture models. In this paper we focus on, by way of sensitivity analysis, extrapolating incomplete patterns using identifying restrictions. Perhaps the best known ones are so-called complete case missing value restrictions (CCMV), where for a given pattern, the conditional distribution of the missing data, given the observed data, is equated to its counterpart in the completers. Available case missing value (ACMV) restrictions equate this conditional density to the one calculated from the subgroup of all patterns for which all required components have been observed. Neighboring case missing value restrictions (NCMV) equate this conditional density to the one calculated from the the pattern with one additional measurement obtained. In this paper, these three identifying restriction strategies are used to multiply impute missing data in a study in metastatic prostate cancer. Multiple imputation is employed to reduce the uncertainty of single imputation. It is shown how hypothesis testing and sensitivity analyses are carried out in this setting.

Humans↗

Catquest questionnaire for use in cataract surgery care: assessment of surgical outcomes.

PURPOSE: To demonstrate the outcome for patients after cataract extraction using the Catquest cataract questionnaire and discuss the models validity in assessing outcome. SETTING: Thirty-five Swedish departments of ophthalmology. METHODS: Patients having cataract extraction performed by surgeons from 35 Swedish departments of opthalmology participated in the study. The questionnaire was given to 2970 consecutive patients having surgery during March 1995 at the participating surgical units. The questionnaire was sent by mail to patients and completed on a voluntary basis. It focuses on visual disabilities in daily life, activity level, cataract symptoms, and degree of independence. The results form the questionnaire are interpreted using a benefit matrix that credits not only a decrease in visual disabilities and cataract symptoms but also an improvement in or maintenance of a preoperative activity level. RESULTS: Complete surgical outcome data and completed preoperative and postoperative questionnaires were available in 1933 cases (65.1%). Benefit from surgery according to the model was achieved by 90.9% of the patients. Patients having their second cataract extraction had the highest frequency of the greatest benefit form surgery. There was good agreement between the different levels of benefit from surgery according to the model and the patient's global rating of his or her vision or achieved visual acuity after surgery, respectively. Patients with missing data (did not return postoperative questionnaire or had missing surgical result variables) were older and had a higher frequency of other diseases and handicaps. CONCLUSION: The Catquest cataract questionnaire allowed the outcome of cataract surgery to be graded by different levels of benefit. There seemed to be good agreement between this model of assessment and the patient's global rating of his or her vision. Missing data may be a problem when a postal questionnaire is used.

Activities of Daily Living↗

Life satisfaction following spinal cord injury: long-term follow-up.

OBJECTIVE: To determine the course of self-reported life satisfaction in a spinal cord injury (SCI) cohort. DESIGN: Prospective study using longitudinal data from the Injury Control Research Center. PARTICIPANTS: Adult persons with traumatic-onset SCI (n = 207) evaluated at 1, 2, 4, and 5 years postinjury using the Life Satisfaction Index-A. RESULTS: A nonsignificant (P > 0.05) main effect of time was found using a repeated-measures analysis controlling for education and employment status. Several methods were used that provided a range of liberal to conservative estimates for missing data (ie, 38% retention rate at year 5). Subsequent missing data analyses tended to corroborate the finding of a nonsignificant effect of time, although the most conservative methods showed a significant decrease in life satisfaction between year 1 and year 5 postinjury (P < 0.05). Examination of numerous demographic, injury, and treatment-related characteristics at each follow-up time point suggested that the main findings of the study were not merely the result of differential dropout rates. CONCLUSION: Life satisfaction after the first year of injury remains largely the same over the next 4 years. Methodologic and analytic recommendations are discussed.

Adolescent↗

A simulation study of the effects of assignment of prior identity-by-descent probabilities to unselected sib pairs, in covariance-structure modeling of a quantitative-trait locus.

Sib pair-selection strategies, designed to identify the most informative sib pairs in order to detect a quantitative-trait locus (QTL), give rise to a missing-data problem in genetic covariance-structure modeling of QTL effects. After selection, phenotypic data are available for all sibs, but marker data-and, consequently, the identity-by-descent (IBD) probabilities-are available only in selected sib pairs. One possible solution to this missing-data problem is to assign prior IBD probabilities (i.e., expected values) to the unselected sib pairs. The effect of this assignment in genetic covariance-structure modeling is investigated in the present paper. Two maximum-likelihood approaches to estimation are considered, the pi-hat approach and the IBD-mixture approach. In the simulations, sample size, selection criteria, QTL-increaser allele frequency, and gene action are manipulated. The results indicate that the assignment of prior IBD probabilities results in serious estimation bias in the pi-hat approach. Bias is also present in the IBD-mixture approach, although here the bias is generally much smaller. The null distribution of the log-likelihood ratio (i.e., in absence of any QTL effect) does not follow the expected null distribution in the pi-hat approach after selection. In the IBD-mixture approach, the null distribution does agree with expectation.

Alleles↗

An eigenvector method for estimating item parameters of the dichotomous and polytomous Rasch models.

The purpose of this paper is to describe a technique for obtaining item parameters of the Rasch model, a technique in which the item parameters are extracted from the eigenvector of a matrix derived from comparisons between pairs of items. The technique can be applied to both dichotomous and polytomous data. In application to a previously published data set, it is shown that the technique provides item parameter estimates comparable to those produced by joint maximum likelihood estimation, and for the most difficult items, the technique appears to produce superior estimates. This method has several advantages. It easily accommodates missing data, and makes transparent the basis for item parameter estimation in the presence of missing data. Furthermore, the method provides a link to other methods in the social sciences and, in particular, provides the framework for application of graph theory to the analysis of assessment networks. Finally, it exploits several characteristics that are unique to the Rasch model.

Algorithms↗

Analysis strategies for serial multivariate ultrasonographic data that are incomplete.

Ultrasonographic measurement of intima-media thickness in the carotid artery has emerged as an important non-invasive means of assessing atherosclerosis, and has served to define primary outcome measures related to progression of arterial lesions in several large clinical trials and epidemiologic studies. It is characteristic that measurements often cannot be obtained from all sites during repeated examinations. This leads to incomplete multivariate serial data, for which the set and number of visualized sites may vary across time. We have contrasted several conditional and unconditional maximum likelihood analytical approaches, and have evaluated these with a simulation experiment based on characteristics of ultrasound measurements collected during the course of the Asymptomatic Carotid Artery Plaque Study. We examined analyses based on unweighted and generalized least squares regression in which we estimated cross-sectional summary statistics using raw means, unconditional maximum likelihood estimates and full maximum likelihood estimates. Since the genesis of missing data is not fully clear, and since the approaches we examined are based, to some degree, on the assumption that data are missing at random, we also examined the relative impact of deviations from such an assumption on each of the approaches considered. We found that maximum likelihood based approaches increased the expected efficiency of the analysis of serial ultrasound data over ignoring missing data by up to 21 per cent.

Arteriosclerosis↗

An illness-death stochastic model in the analysis of longitudinal dementia data.

A significant source of missing data in longitudinal epidemiological studies on elderly individuals is death. Subjects in large scale community-based longitudinal dementia studies are usually evaluated for disease status in study waves, not under continuous surveillance as in traditional cohort studies. Therefore, for the deceased subjects, disease status prior to death cannot be ascertained. Statistical methods assuming deceased subjects to be missing at random may not be realistic in dementia studies and may lead to biased results. We propose a stochastic model approach to simultaneously estimate disease incidence and mortality rates. We set up a Markov chain model consisting of three states, non-diseased, diseased and dead, and estimate the transition hazard parameters using the maximum likelihood approach. Simulation results are presented indicating adequate performance of the proposed approach.

Aged↗

Clinical significance of antiviral therapy for episodic treatment of herpes labialis: exploratory analyses of the combined data from two valaciclovir trials.

Valaciclovir (Valtrex) 2 g twice daily for 1 day was recently approved in the United States for treatment of cold sores. In order to apply more clinically relevant assumptions to the analysis, we examined the effect of different missing data and endpoint assumptions on apparent valaciclovir efficacy. Results of each analysis demonstrate statistically significant increases in the proportion of subjects whose cold sores were aborted with valaciclovir compared with placebo, and significant decreases in healing times for subjects with cold sore lesions who were treated with valaciclovir compared with placebo. These exploratory analyses provide evidence of the robustness of the results to differing missing data assumptions and show that use of more clinically relevant endpoint assumptions increases the magnitude of some therapeutic responses. We also introduce a new measure that combines the two observed drug effects (reduced lesion duration, increased aborted lesions) into a single endpoint that captures the global benefit of the drug to the patient.

Acyclovir↗

Applications of computer-intensive statistical methods to environmental research.

Conventional statistical approaches rely heavily on the properties of the central limit theorem to bridge the gap between the characteristics of a sample and some theoretical sampling distribution. Problems associated with nonrandom sampling, unknown population distributions, heterogeneous variances, small sample sizes, and missing data jeopardize the assumptions of such approaches and cast skepticism on conclusions. Conventional nonparametric alternatives offer freedom from distribution assumptions, but design limitations and loss of power can be serious drawbacks. With the data-processing capacity of today's computers, a new dimension of distribution-free statistical methods has evolved that addresses many of the limitations of conventional parametric and nonparametric methods. Computer-intensive statistical methods involve reshuffling, resampling, or simulating a data set thousands of times to empirically define a sampling distribution for a chosen test statistic. The only assumption necessary for valid results is the random assignment of experimental units to the test groups or treatments. Application to a real data set illustrates the advantages of these methods, including freedom from distribution assumptions without loss of power, complete choice over test statistics, easy adaptation to design complexities and missing data, and considerable intuitive appeal. The illustrations also reveal that computer-intensive methods can be more time consuming than conventional methods and the amount of computer code required to orchestrate reshuffling, resampling, or simulation procedures can be appreciable.

Analysis of Variance↗

The effect of question structure on self-reports of heavy drinking: closed-ended versus open-ended questions.

OBJECTIVE: We compared open-ended versus closed-ended questions on the frequency of consuming five or more drinks in a single sitting. METHOD: From a general population survey of Ontario adults (N = 2,022, 62% male), we analyzed a subsample of 649 respondents who reported drinking five or more drinks in a single sitting at least once in the past year. Differences in agreement between the two questions and rates of missing data were evaluated. RESULTS: For the most part, the two measures were not consistent, with the closed-ended question eliciting higher rates of heavier drinking. Rates of missing data were also higher for the open-ended question. CONCLUSIONS: Open-ended question may not necessarily be more suitable than closed-ended questions for estimating the frequency of heavy alcohol use.

Adult↗

Parametric and nonparametric linkage analysis: a unified multipoint approach.

In complex disease studies, it is crucial to perform multipoint linkage analysis with many markers and to use robust nonparametric methods that take account of all pedigree information. Currently available methods fall short in both regards. In this paper, we describe how to extract complete multipoint inheritance information from general pedigrees of moderate size. This information is captured in the multipoint inheritance distribution, which provides a framework for a unified approach to both parametric and nonparametric methods of linkage analysis. Specifically, the approach includes the following: (1) Rapid exact computation of multipoint LOD scores involving dozens of highly polymorphic markers, even in the presence of loops and missing data. (2) Non-parametric linkage (NPL) analysis, a powerful new approach to pedigree analysis. We show that NPL is robust to uncertainty about mode of inheritance, is much more powerful than commonly used nonparametric methods, and loses little power relative to parametric linkage analysis. NPL thus appears to be the method of choice for pedigree studies of complex traits. (3) Information-content mapping, which measures the fraction of the total inheritance information extracted by the available marker data and points out the regions in which typing additional markers is most useful. (4) Maximum-likelihood reconstruction of many-marker haplotypes, even in pedigrees with missing data. We have implemented NPL analysis, LOD-score computation, information-content mapping, and haplotype reconstruction in a new computer package, GENEHUNTER. The package allows efficient multipoint analysis of pedigree data to be performed rapidly in a single user-friendly environment.

Algorithms↗