Search PubMedSearch

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Missing quality of life data in cancer clinical trials: serious problems and challenges.

Measurement of quality of life (QOL) in cancer clinical trials has increased in recent years as more groups realize the importance of such endpoints. A key problem has been missing data. Some QOL data may unavoidably be missing, as for example when patients are too ill to complete forms. Other important sources are potentially avoidable and can broadly be divided into three categories: (i) methodological factors; (ii) logistic and administrative factors; (iii) patient-related factors. Logistic and administrative factors, for example, staff oversights, have proven to be most important. Since most QOL measurements require patient self-report, it is usually not possible to rectify the failure to collect baseline data or any follow-up assessments. There is strong evidence that such data are not 'missing at random', and cannot be ignored without introducing bias. Although several approaches to the analysis of partly missing data have been described, none is entirely satisfactory. Prevention of avoidable missing data is better than attempted cure. In July 1996, an international conference on missing QOL data in cancer clinical trials reported the experience of most major groups involved. This paper will serve as an introduction to the problem and provide an estimation of its magnitude, and approaches to its prevention and solution.

Bias

Quality of life as an endpoint in EORTC clinical trials. European Organization for Research and Treatment for Cancer.

For more than 30 years the European Organization for Research and Treatment for Cancer (EORTC) has conducted, co-ordinated, and stimulated research on the experimental and clinical bases of treatment of cancer and related problems. For more than a decade the EORTC has included quality of life as an outcome measure in some of its trials. The number of clinical studies that include QOL as an evaluation endpoint has increased rapidly in the last few years, and is still increasing steadily. This necessitated a careful and critical evaluation of procedures and results so far in order to generate appropriate guidelines and procedures for incorporating QOL issues in all stages of the clinical trial process, including protocol writing, data collection, data analysis, and reporting of results. This paper provides an overview of the types and the design of studies, data management of quality of life assessment, compliance, missing data and lessons learned during the past years with respect to QOL assessments in the EORTC studies.

Adult

A systematic overview of the use of diary cards, quality-of-life questionnaires, and psychometric tests in treatment trials of Helicobacter pylori-positive and -negative non-ulcer dyspepsia.

BACKGROUND: Our aim was to evaluate the use of diary cards, quality-of-life questionnaires, and psychometric tests in treatment trials of non-ulcer dyspepsia. METHODS: Data sources were a Medline search (up to 1966) and a manual search of five gastrointestinal journals (up to 1980) for original, randomized, double-blind, placebo-controlled trials with at least 20 patients which evaluated treatment regimens for non-ulcer dyspepsia. RESULTS: Of the 67 eligible studies, 31 used diary cards. Diary cards were used alone in 15 of the 31 studies (48%), whereas the others (52%) also used a physician assessment. The symptoms assessed by diary cards were epigastric pain (100%), nausea/vomiting (65%), heartburn (52%), belching (39%), regurgitation (29%), fullness (29%), and bloating (23%). Forty-five per cent also recorded antacid use. Severity of outcome measures was assessed by a visual analogue scale in 5 of the 31 studies (16%), Likert scales in 17 studies (55%), and unclear methods in 3 studies (10%). For statistical analysis daily averages of symptoms were used in 5 of the 31 studies (16%), weekly averages in 11 studies (35%), and 2-week intervals during the treatment period in the rest, with some studies using a combination (such as daily and weekly averages). Only 3 of the 31 studies (8%) checked for compliance with diary card data. None of the studies mention anything about missing data and how this was handled. One study evaluated quality of life questionnaires and one evaluated a psychometric test. CONCLUSIONS: Non-ulcer dyspepsia treatment trials frequently use diary cards but need to be much clearer about how information was obtained and how it was used in the statistical analysis. Not much information is available to comment on the use of quality-of-life questionnaires or psychometric tests for evaluation of outcome measures.

Dyspepsia

Race and delayed kidney allograft function.

BACKGROUND: Allograft survival among black recipients is poorer than among whites. Delayed allograft function is associated with a significant reduction in renal allograft survival. The relationship between delayed allograft function and black race is incompletely specified and was the focus of this investigation. METHODS: A non-concurrent study of 325 recipients of cadaveric allografts followed for the occurrence of delayed allograft function defined as dialysis during the first week following transplantation for the principal analysis. A secondary definition of delayed allograft function was formulated based on the serum creatinine 2 weeks after transplantation. Unadjusted and adjusted logistic regression analysis were used to examine the unconfounded relationship between race and delayed allograft function. RESULTS: Fifty-seven of 91 (62.6%) black recipients experienced delayed allograft function compared to 113 of 234 (48.3%) whites. The odds ratio for black race as a predictor of delayed allograft function was 1.80, P=0.02, (95% CI, 1.09, 2.95). This finding was stable despite adjustment for other predictors of delayed allograft function in a multivariate model, but the precision of this estimate was less (P=0.10) because of missing data. Additionally, adjusted models with imputed values for missing covariates, models using a secondary definition of delayed allograft function, and models excluding patients whose cyclosporin therapy was delayed, all consistently demonstrated a similar association between black race and delayed allograft function. CONCLUSIONS: This study demonstrated an increased risk of delayed allograft function among black recipients. This relationship may play a role in the poorer allograft outcomes experienced by black recipients. Given the negative effect of delayed allograft function on allograft survival, efforts to identify its modifiable risk factors should be a high priority.

Adult

The impact of changing methods of data collection on the reliability of self-reported drug use of adolescents.

The purpose of this study is to determine the impact of different modes of data collection on the reliability of self-reported drug use of adolescents in a panel study. Adolescents were assigned to four groups based upon the ways they chose to respond to the survey instruments: 1) mailed questionnaires in both years, 2) survey interview in one year and mailed questionnaire in the next year, 3) mailed questionnaire in one year and survey interview in the following year, and 4) survey interview in both years. The quality of the self-reported data was examined in terms of return rates, missing data, internal consistency, and consistency of reported information over time. No significant differences were found between groups, suggesting that the mode of data collection does not affect the reliability of adolescents' self-reports of substance use.

Adolescent

Applications of multiple imputation to the analysis of censored regression data.

The first part of the article reviews the Data Augmentation algorithm and presents two approximations to the Data Augmentation algorithm for the analysis of missing-data problems: the Poor Man's Data Augmentation algorithm and the Asymptotic Data Augmentation algorithm. These two algorithms are then implemented in the context of censored regression data to obtain semiparametric methodology. The performances of the censored regression algorithms are examined in a simulation study. It is found, up to the precision of the study, that the bias of both the Poor Man's and Asymptotic Data Augmentation estimators, as well as the Buckley-James estimator, does not appear to differ from zero. However, with regard to mean squared error, over a wide range of settings examined in this simulation study, the two Data Augmentation estimators have a smaller mean squared error than does the Buckley-James estimator. In addition, associated with the two Data Augmentation estimators is a natural device for estimating the standard error of the estimated regression parameters. It is shown how this device can be used to estimate the standard error of either Data Augmentation estimate of any parameter (e.g., the correlation coefficient) associated with the model. In the simulation study, the estimated standard error of the Asymptotic Data Augmentation estimate of the regression parameter is found to be congruent with the Monte Carlo standard deviation of the corresponding parameter estimate. The algorithms are illustrated using the updated Stanford heart transplant data set.

Algorithms

Ensuring data quality in a multicenter clinical trial: remote site data entry, central coordination and feedback.

In an ongoing multicenter clinical trial, "Treatment Strategies in Schizophrenia," the five participating sites have the capacity to perform a variety of tasks or study functions independently. These tasks include (a) verification of diagnostic eligibility through the use of computerized decision algorithms; (b) assignment of patients to treatment based on prognostic indicators using a computerized randomization algorithm; (c) entry of data into a microcomputer using a clinical trial data management system that performs simple range and missing data item checks; and (d) regular transfer of all data to the central coordinating team. The clinical trial data management system employed allows for both independent site functioning and assurance of consistency across sites. The integration of a variety of software outside the main data management system provides the central coordinators with the tools to monitor critical data as it is collected, as well as the capacity to assess the flow, quality, and uniformity of the ongoing trial.

Clinical Trials as Topic

Developing an outcomes infrastructure for nursing. The Outcomes Taskforce.

An infrastructure to support the evaluation of patient care sensitive to the intervention of nursing personnel is being developed within a major health maintenance organization. In addition to traditional administrative measures of care, the database infrastructure will include measures of the patient's functional status, knowledge and engagement in care and psychosocial well-being. These measures are believed to be particularly sensitive to the independent intervention of the nurse. Reported here are the structures in place to monitor and support the reliability and validity of the administrative data elements; algorithm elements created to account for missing data; the model for the first generation of successful practice reports and the results of a study establishing the content validity of the clinical data elements.

Costs and Cost Analysis

Application of random-effects regression models in relapse research.

This article describes and illustrates use of random-effects regression models (RRM) in relapse research. RRM are useful in longitudinal analysis of relapse data since they allow for the presence of missing data, time-varying or invariant covariates, and subjects measured at different timepoints. Thus, RRM can deal with "unbalanced" longitudinal relapse data, where a sample of subjects are not all measured at each and every timepoint. Also, recent work has extended RRM to handle dichotomous and ordinal outcomes, which are common in relapse research. Two examples are presented from a smoking cessation study to illustrate analysis using RRM. The first illustrates use of a random-effects ordinal logistic regression model, examining longitudinal changes in smoking status, treating status as an ordinal outcome. The second example focuses on changes in motivation scores prior to and following a first relapse to smoking. This latter example illustrates how RRM can be used to examine predictors and consequences of relapse, where relapse can occur at any study timepoint.

Alcoholism

Non-ignorable missing covariates in generalized linear models.

We propose a likelihood method for estimating parameters in generalized linear models with missing covariates and a non-ignorable missing data mechanism. In this paper, we focus on one missing covariate. We use a logistic model for the probability that the covariate is missing, and allow this probability to depend on the incomplete covariate. We allow the covariates, including the incomplete covariate, to be either categorical or continuous. We propose an EM algorithm in this case. For a missing categorical covariate, we derive a closed form expression for the E- and M-steps of the EM algorithm for obtaining the maximum likelihood estimates (MLEs). For a missing continuous covariate, we use a Monte Carlo version of the EM algorithm to obtain the MLEs via the Gibbs sampler. The methodology is illustrated using an example from a breast cancer clinical trial in which time to disease progression is the outcome, and the incomplete covariate is a quality of life physical well-being score taken after the start of therapy. This score may be missing because the patients are sicker, so this covariate could be non-ignorably missing.

Algorithms

Evaluating interventions with differential attrition: the importance of nonresponse mechanisms and use of follow-up data.

Evaluations of psychological interventions are often criticized because of differential attrition, which is cited as a severe threat to validity. The present study shows that differential attrition is not a problem unless the mechanism causing the attrition is inaccessible (unavailable for analysis). With a simulation study, we show that conclusions about program effects (a) are unbiased when there is no differential attrition, even with usual complete cases analysis; (b) may be severely biased when based on usual complete cases analyses and there is differential attrition; (c) are unbiased when based on the expectation-maximization (EM) algorithm, even when there is differential attrition, as long as the attrition mechanism is accessible; and (d) are biased, even with the EM algorithm, when the attrition mechanism is inaccessible. Following Little and Rubin (1987), we advocate the collection of new data from a random sample of subjects with initially missing data. On the basis of these data, we propose a simple correction to the EM algorithm estimates. In our study, the correction produced unbiased estimates of program effects parameters, even with an inaccessible attrition mechanism and substantial differential attrition.

Bias

Free-living energy expenditure of adult men assessed by continuous heart-rate monitoring and doubly-labelled water.

Free-living energy expenditure was estimated by doubly-labelled water (DLW) and continuous heart-rate (HR) monitoring over nine consecutive days in nine healthy men with sedentary occupations but different levels of leisure-time physical activity. Individual calibrations of the HR-energy expenditure (EE) relationship were obtained for each subject using 30 min average values of HR and EE obtained during 24 h whole-body calorimetry with a defined exercise protocol, and additional data points for individual leisure activities measured with an Oxylog portable O2 consumption meter. The HR data were processed to remove spurious values and insert missing data before the calculation of EE from second-order polynomial equations relating EE to HR. After data processing, the HR-derived EE for this group of subjects was on average 0.8 (SEM 0.6) MJ/d, or 6.0 (SEM 4.2) % higher than that estimated by DLW. The diary-respirometer method, used over the same 9 d, gave values which were 1.9 (SEM 0.7) MJ/d, or -12.1 (SEM 4.0) % lower than the DLW method. The results suggest that HR monitoring can provide a better estimate of 24 h EE of groups than the diary-respirometer method, but show that both methods can introduce errors of 20% or more in individuals.

Adult

Issues in incorporation semantic integrity in molecular biological object-oriented databases.

Issues critical to ensuring semantic integrity in molecular biological data collections have been identified and include complexity, exceptions, missing data, changing models, holism and integration, delocalized data, interoperability and nomenclature. This combination is peculiar to biology and presents some interesting problems as a result. Little is known about semantic checking in object-oriented databases in general, but because such technology appears highly suitable for modeling biological data, it is appropriate to examine the ways in which object-oriented technology can support this functionality. It is concluded that object-oriented technology will support semantic checking even in a complex domain like biology. We propose 10 guidelines for future work including ways of treating exceptional cases and 'positioning' of constraints in a schema.

Biotechnology

Admission base deficit predicts transfusion requirements and risk of complications.

BACKGROUND: Trauma center resource management could be facilitated by a readily available indicator of resource consumption. This marker should identify patients more likely to require transfusion and intensive care services and to develop complications. Base deficit (BD) has been shown to be a valuable indicator of shock, abdominal injury, fluid requirements, efficacy of resuscitation, and to be predictive of mortality after trauma. This study was performed to determine whether BD could be used to identify which patients were likely to require blood transfusion in the first 24 hours of hospitalization, and to develop shock-related complications and increased intensive care unit (ICU) and hospital stays. METHODS: A retrospective review of 2,954 patients admitted to the Valley Medical Center Level I trauma service from July 1990 through August 1995 was done using the trauma registry and blood bank data bases. Medical record review was done to supplement missing data. RESULTS: Transfusion requirements increased as the BD category became more severe (p < 0.001). Transfusions were required within 24 hours of admission in 72% of patients with a BD < or = -6 versus 18% of patients with a BD > -6 (p < 0.001, chi 2). Both ICU and hospital length of stay increased with worsening BD (p < 0.015 and p < 0.05, respectively). The frequency of adult respiratory distress syndrome (ARDS) (p < 0.01), renal failure (p = 0.015), coagulopathy (p < 0.001), and multiorgan system failure (MOF) (p = 0.002) all increased with increasingly severe BD. Discriminate analysis using Injury Severity Score (ISS) and BD category demonstrated predictive accuracy of 81%, 77%, and 77% for coagulopathy, ARDS, and MOF, respectively. Mortality also increased with worsening BD. When stratified by BD category, there was no difference between observed and predicted survival. CONCLUSIONS: Admission BD identifies patients likely to require early transfusion and increased ICU and hospital stays, and be at increased risk for shock-related complications. Patients with BD < or = -6 should undergo type and cross-match rather than type and screen. The use of ISS and BD category probability curves may identify candidates for early invasive monitoring.

Acid-Base Imbalance

On summary measures analysis of the linear mixed effects model for repeated measures when data are not missing completely at random.

Subjects often drop out of longitudinal studies prematurely, yielding unbalanced data with unequal numbers of measures for each subject. A simple and convenient approach to analysis is to develop summary measures for each individual and then regress the summary measures on between-subject covariates. We examine properties of this approach in the context of the linear mixed effects model when the data are not missing completely at random, in the sense that drop-out depends on the values of the repeated measures after conditioning on fixed covariates. The approach is compared with likelihood-based approaches that model the vector of repeated measures for each individual. Methods are compared by simulation for the case where repeated measures over time are linear and can be summarized by a slope and intercept for each individual. Our simulations suggest that summary measures analysis based on the slopes alone is comparable to full maximum likelihood when the data are missing completely at random but is markedly inferior when the data are not missing completely at random. Analysis discarding the incomplete cases is even worse, with large biases and very poor confidence coverage.

Computer Simulation

Individual subject random assignment is the preferred means of evaluating behavioral lifestyle modification.

Of the three most important approaches to evaluating lifestyle and health outcomes--observational studies, individual subject random assignment clinical trials, and community random assignment clinical trials--individual subject random assignment clinical trials provide the most useful information and the most certain inferences. Observational studies are limited by the collection of information on association of lifestyle instead of change in lifestyle with health outcomes, as well as by individual lifestyle selection that may be associated with particular outcomes. Community randomized trials may be the best way to decide such public health policy issues as whether or not to add fluoride to a water supply or to use a community-wide anti-smoking program. The limited amount of individual-specific data collected in community randomized trials, difficulties in accounting for missing data, and problems in data analysis because people move into or out of communities that are under study limit the value of community randomized trials for advising individuals whether or not to embark on a lifestyle modification program. Clinical trials of pharmacologic agents have successfully addressed challenges to individual subject random assignment clinical trials of lifestyle modification, such as long duration of study, access to study intervention(s) by individuals not assigned them, and cost.

Community Health Services

Nicardipine and propranolol in the treatment of essential hypertension.

Two hundred thirty-four patients with supine diastolic blood pressure of between 95 and 114 mm Hg were enrolled into a double-blind, randomized, parallel, multicenter trial. The patients were randomized to either nicardipine 30 mg tid, propranolol 40 mg tid, or nicardipine 30 mg tid and propranolol 40 mg tid for six weeks. Two hundred six patients yielded data for analyses. Of the 28 not included, seven had missing data, whereas the remaining 21 were excluded because they either failed to meet inclusion criteria or were noncompliant at endpoint. Both nicardipine and propranolol as monotherapies and in combination achieved statistically significant, (P less than .01), supine diastolic blood pressure reduction relative to baseline. The combination of nicardipine and propranolol showed a greater reduction in supine diastolic and systolic measurements than either of the monotherapies. Nicardipine produced greater blood pressure reductions one hour after dosing, whereas the propranolol treatment tended to produce slightly greater blood pressure decreases eight hours after dose. The combination always resulted in the greatest blood pressure reduction, independent of time after dose. Adverse experiences were reported by 26% of patients in the nicardipine-treated group, most often transient vasodilatory effects, by 17% of the propranolol-treated patients, and by 18% of the combination-treated group. This study demonstrated at the doses studied that nicardipine alone produced equivalent blood pressure reductions to those obtained by propranolol alone, but that the combination of these two drugs produced greater reductions in blood pressures than either of the monotherapies.

Adult

Molecular phylogeny of the genus Hypochaeris using internal transcribed spacers of nuclear rDNA: inference for chromosomal evolution.

Sequences of the internal transcribed spacers (ITSs) of 18S-26S nuclear ribosomal DNA were used to resolve phylogenetic relationships and chromosomal evolution among 14 species of the genus Hypochaeris (Asteraceae). Parsimony analysis was performed for phylogenetic reconstruction, and sequence divergence between species was estimated. Pairwise sequence divergence within Hypochaeris genus ranged from 0% to 25.68% in ITS1 and from 0% to 17.08% in ITS2. A highly resolved strict-consensus tree was obtained that showed the phylogenetically useful information of ITS sequences within the genus Hypochaeris. Four clades could be well distinguished, one of them formed by the single species H. robertia, which appeared to be the most related to the ancestral species of the genus. The results agree with taxonomic classification based on morphological data, and the tree obtained, when indels are coded as missing data, aggregates the species having the same chromosome number, except in one clade. According to the ITS phylogenetic tree, the chromosomal evolution within the genus Hypochaeris conflicts with the previous hypothesis and suggests that karyotype evolution in Hypochaeris was accompanied with both decreasing and increasing dysploidy, probably with several chromosomal rearrangements, and from an ancestral basic chromosome number of 4 or 5.

Asteraceae