Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Comparison of several model-based methods for analysing incomplete quality of life data in cancer clinical trials.

This paper considers five methods of analysis of longitudinal assessment of health related quality of life (QOL) in two clinical trials of cancer therapy. The primary difference in the two trials is the proportion of participants who experience disease progression or death during the period of QOL assessments. The sensitivity of estimation of parameters and hypothesis tests to the potential bias as a consequence of the assumptions of missing completely at random (MCAR), missing at random (MAR) and non-ignorable mechanisms are examined. The methods include complete case analysis (MCAR), mixed-effects models (MAR), a joint mixed-effects and survival model and a pattern-mixture model. Complete case analysis overestimated QOL in both trials. In the adjuvant breast cancer trial, with 15 per cent disease progression, estimates were consistent across the remaining four methods. In the advanced non-small-cell lung cancer trial, with 35 per cent mortality, estimates were sensitive to the missing data assumptions and methods of analysis.

Breast Neoplasms↗

Bayesian infinite mixture model based clustering of gene expression profiles.

MOTIVATION: The biologic significance of results obtained through cluster analyses of gene expression data generated in microarray experiments have been demonstrated in many studies. In this article we focus on the development of a clustering procedure based on the concept of Bayesian model-averaging and a precise statistical model of expression data. RESULTS: We developed a clustering procedure based on the Bayesian infinite mixture model and applied it to clustering gene expression profiles. Clusters of genes with similar expression patterns are identified from the posterior distribution of clusterings defined implicitly by the stochastic data-generation model. The posterior distribution of clusterings is estimated by a Gibbs sampler. We summarized the posterior distribution of clusterings by calculating posterior pairwise probabilities of co-expression and used the complete linkage principle to create clusters. This approach has several advantages over usual clustering procedures. The analysis allows for incorporation of a reasonable probabilistic model for generating data. The method does not require specifying the number of clusters and resulting optimal clustering is obtained by averaging over models with all possible numbers of clusters. Expression profiles that are not similar to any other profile are automatically detected, the method incorporates experimental replicates, and it can be extended to accommodate missing data. This approach represents a qualitative shift in the model-based cluster analysis of expression data because it allows for incorporation of uncertainties involved in the model selection in the final assessment of confidence in similarities of expression profiles. We also demonstrated the importance of incorporating the information on experimental variability into the clustering model. AVAILABILITY: The MS Windows(TM) based program implementing the Gibbs sampler and supplemental material is available at http://homepages.uc.edu/~medvedm/BioinformaticsSupplement.htm CONTACT: medvedm@email.uc.edu

Bayes Theorem↗

Epidemiologic evaluation of measurement data in the presence of detection limits.

Quantitative measurements of environmental factors greatly improve the quality of epidemiologic studies but can pose challenges because of the presence of upper or lower detection limits or interfering compounds, which do not allow for precise measured values. We consider the regression of an environmental measurement (dependent variable) on several covariates (independent variables). Various strategies are commonly employed to impute values for interval-measured data, including assignment of one-half the detection limit to nondetected values or of "fill-in" values randomly selected from an appropriate distribution. On the basis of a limited simulation study, we found that the former approach can be biased unless the percentage of measurements below detection limits is small (5-10%). The fill-in approach generally produces unbiased parameter estimates but may produce biased variance estimates and thereby distort inference when 30% or more of the data are below detection limits. Truncated data methods (e.g., Tobit regression) and multiple imputation offer two unbiased approaches for analyzing measurement data with detection limits. If interest resides solely on regression parameters, then Tobit regression can be used. If individualized values for measurements below detection limits are needed for additional analysis, such as relative risk regression or graphical display, then multiple imputation produces unbiased estimates and nominal confidence intervals unless the proportion of missing data is extreme. We illustrate various approaches using measurements of pesticide residues in carpet dust in control subjects from a case-control study of non-Hodgkin lymphoma.

Bias↗

Prevalence of rheumatoid arthritis and hepatitis C in those age 60 and older in a US population based study.

OBJECTIVE: A positive association between rheumatoid arthritis (RA) and hepatitis C virus (HCV) infection has been reported in clinic based cross sectional studies. We investigated if RA and HCV are associated in a population based survey. METHODS: Using data from the National Health and Nutrition Examination Survey III, hepatitis C and RA status were determined for subjects > or = 60 years of age. RA was defined to be present when 3 of 6 American College of Rheumatology (ACR) criteria were met. RESULTS: Of 6596 subjects, 1827 (27.7%) were excluded due to missing data. Of the remaining 4769, 196 subjects (4.1%) met our modified ACR criteria for probable RA: 63 tested positive for anti-HCV antibodies (1.3%) while 35 were HCV RNA positive (0.7%). Two subjects had both HCV antibodies and RA, while one subject was both HCV RNA positive and had RA. HCV antibody positivity was not associated with RA (OR 0.44, 95% CI 0.07-2.80). Similarly, HCV positivity by polymerase chain reaction was not associated with RA (OR 0.77, 95% CI 0.10-6.19). CONCLUSION: These results argue against a potential role for HCV in the etiology of RA in the US population aged 60 years and over.

Age Distribution↗

Size and power of two-sample tests of repeated measures data.

One method of using repeated measures data to compare treatment groups in a clinical trial is to summarize each subject's outcomes with a single summary statistic, and then perform a distribution-free comparison based on the resulting statistics. We examine extensions of this approach and conditions under which they retain proper size in the presence of missing data. The asymptotic relative efficiencies of several summary statistic tests are calculated to show which perform best in a variety of situations. The techniques are illustrated using data from an AIDS clinical trial.

Analysis of Variance↗

Measuring the impact of MS on walking ability: the 12-Item MS Walking Scale (MSWS-12).

OBJECTIVE: To develop a patient-based measure of walking ability in MS. METHODS: Twelve items describing the impact of MS on walking (12-Item MS Walking Scale [MSWS-12]) were generated from 30 patient interviews, expert opinion, and literature review. Preliminary psychometric evaluation (data quality, scaling assumptions, acceptability, reliability, validity) was undertaken in the data generated by 602 people from the MS Society membership database. Further psychometric evaluation (including comprehensive validity assessment, responsiveness, and relative efficiency) was conducted in two hospital-based samples: people with primary progressive MS (PPMS; n = 78) and people with relapses admitted for IV steroid treatment (n = 54). RESULTS: In all samples, missing data were low (< or =3.8%), item test-retest reproducibility was high (> or =0.78), scaling assumptions were satisfied, and reliability was high (> or =0.94). Correlations between the MSWS-12 and other scales were consistent with a priori hypotheses. The MSWS-12 (relative efficiency = 1.0) was more responsive than the Functional Assessment of Multiple Sclerosis mobility scale (0.72), the 36-Item Short Form Health Survey physical functioning scale (0.33), the Expanded Disability Status Scale (0.03), the 25-ft Timed Walk Test (0.44), and Guy's Neurologic Disability Scale lower limb disability item (0.10). CONCLUSIONS: The MSWS-12 satisfies standard criteria as a reliable and valid patient-based measure of the impact of MS on walking. In these samples, the MSWS-12 was more responsive than other walking-based scales.

Adult↗

The Psychiatric Out-Patient Experiences Questionnaire (POPEQ): data quality, reliability and validity in patients attending 90 Norwegian clinics.

The aim of this study was to develop and evaluate the Psychiatric Out-Patient Experiences Questionnaire (POPEQ). The instrument was developed following a literature review, patient interviews and pre-testing of questionnaire items. The POPEQ was administered as part of a postal survey of 15,422 adult outpatients attending Norwegian clinics; 6677 (43.3%) patients responded to the questionnaire. Items had low levels of missing data. Factor analysis showed that 11 widely applicable items contribute to a measure of overall experiences. Sub-dimensions include clinician interaction (six items) information (two items) and outcomes (three items). Item-total correlations ranged from 0.5 to 0.8. Cronbach's alpha and test-retest reliability estimates exceeded the criterion of 0.7; the majority were over 0.8 and total scores over 0.9. Construct validity was supported by the results of 128 tests. The POPEQ includes important aspects of patient experience for psychiatric outpatients and has excellent evidence for reliability and construct validity. The instrument is recommended for the measurement of psychiatric outpatient experiences.

Adult↗

Method for cohort and nested case-control studies: the prevalence, timing and effectiveness of obstetric ultrasound, Victoria 1991-1992.

The study was designed to assess the effectiveness of obstetric ultrasound in the diagnosis of congenital malformations and to establish its prevalence of use and timing. Statewide data were collected in 138 of the 141 obstetric hospitals in Victoria over a 12-month period during 1991-1992. Within the final cohort of 55,226 mothers providing responses, a nested case-control study group was formed. This group comprised 719 cases (infants born with one or more malformations potentially diagnosable at 16-20 weeks) and 703 controls (non-malformed infants). The cases for the group were extracted from the Victorian Congenital Malformations Register; controls were randomly selected from the Victorian Perinatal Data Collection Unit's database. Of the 1422 in the study group, 1328 medical records were validated in 100 hospitals. The design, method, procedure, sample and outcome are described for the cohort and nested case-control studies. At the conclusion of the study it was established that major variations between the cohort and missing data were confined to mothers less likely to have had a spontaneous vaginal delivery or to those with poor perinatal outcome. There was no significant selective loss of cases or controls in the nested case-control study group.

Case-Control Studies↗

Causal effects in clinical and epidemiological studies via potential outcomes: concepts and analytical approaches.

A central problem in public health studies is how to make inferences about the causal effects of treatments or agents. In this article we review an approach to making such inferences via potential outcomes. In this approach, the causal effect is defined as a comparison of results from two or more alternative treatments, with only one of the results actually observed. We discuss the application of this approach to a number of data collection designs and associated problems commonly encountered in clinical research and epidemiology. Topics considered include the fundamental role of the assignment mechanism, in particular the importance of randomization as an unconfounded method of assignment; randomization-based and model-based methods of statistical inference for causal effects; methods for handling noncompliance and missing data; and methods for limiting bias in the analysis of observational data, including propensity score matching and sensitivity analysis.

Analysis of Variance↗

Urinary albumin excretion and its relation with C-reactive protein and the metabolic syndrome in the prediction of type 2 diabetes.

OBJECTIVE: To investigate urinary albumin excretion (UAE) and its relation with C-reactive protein (CRP) and the metabolic syndrome in the prediction of the development of type 2 diabetes. RESEARCH DESIGN AND METHODS: We used data from the Prevention of Renal and Vascular End Stage Disease (PREVEND) study, an ongoing, community-based, prospective cohort study initiated in 1997 in the Netherlands. The initial cohort consisted of 8,592 subjects. After 4 years, 6,894 subjects participated in a follow-up survey. Subjects with diabetes at baseline or missing data on fasting glucose were excluded, leaving 5,654 subjects for analysis. The development of type 2 diabetes, defined as a fasting glucose > or = 7.0 mmol/l and/or the use of antidiabetic medication, was used as the outcome measure. UAE was calculated as the mean UAE from two consecutive 24-h urine collections. Logistic regression models were used, with the development of type 2 diabetes as the dependent variable. RESULTS: Of the 5,654 subjects for whom data were analyzed, 185 (3.3%) developed type 2 diabetes during a mean follow-up period of 4.2 years. UAE, CRP, and the presence of the metabolic syndrome at baseline were significantly associated with the incidence of type 2 diabetes (P < 0.001 for all variables). In a univariate model, the odds ratio (OR) for UAE was 1.59 (95% CI 1.42-1.79). In our full model, adjusted for age, sex, number of criteria of metabolic syndrome, and other known risk factors for the development of type 2 diabetes (including fasting insulin), the association between UAE and type 2 diabetes remained significant (OR 1.53, 95% CI 1.25-1.88, P < 0.001). There was a significant interaction between UAE and CRP (P = 0.002). After CRP was stratified into tertiles, the ORs for the association between baseline UAE and the development of type 2 diabetes were 2.2 (1.47-3.3), 1.33 (0.96-1.84), and 1.04 (0.83-1.31) for the lowest to highest tertiles, respectively. CONCLUSIONS: UAE predicts type 2 diabetes independent of the metabolic syndrome and other known risk markers of development of type 2 diabetes. The predictive value of UAE was modified by the level of CRP.

Aged↗

[Delivery logbook: pertinence of the collected data for scientific exploitation].

OBJECTIVES: To determine if the data recorded in the delivery logbook are relevant and complete compared to the data registered in the patient's medical chart. MATERIAL AND METHODS: Prospective study during one month recording all the data registered by the midwives or the medical doctors in the delivery room immediately after the birth. To compare them to the information collected in the patient's medical chart. Our delivery logbook has 55 headings for each patient. During October 1998, we had 156 births. RESULTS: We had 5.3% of errors divided by missing data and erroneous entries. More precisely, we recorded 9.5% of errors in the antepartum data, 3.2% in intrapartum or postpartum events and 5.8% in the newborn information. CONCLUSION: This study demonstrates that medical data used for scientific projects and recorded in the delivery logbook, have to be interpreted very cautiously. It seems to be reasonable that each obstetrical unit should standardize its registration of data and carry out internal audits.

Bias↗

Handling missing values in population data: consequences for maximum likelihood estimation of haplotype frequencies.

Haplotype frequency estimation in population data is an important problem in genetics and different methods including expectation maximisation (EM) methods have been proposed. The statistical properties of EM methods have been extensively assessed for data sets with no missing values. When numerous markers and/or individuals are tested, however, it is likely that some genotypes will be missing. Thus, it is of interest to investigate the behaviour of the method in the presence of incomplete genotype observations. We propose an extension of the EM method to handle missing genotypes, and we compare it with commonly used methods (such as ignoring individuals with incomplete genotype information or treating a missing allele as any other allele). Simulations were performed, starting from data sets of haematopoietic stem cell donors genotyped at three HLA loci. We deleted some data to create incomplete genotype observations in various proportions. We then compared the haplotype frequencies obtained on these incomplete data sets using the different methods to those obtained on the complete data. We found that the method proposed here provides better estimations, both qualitatively and quantitatively, but increases the computation time required. We discuss the influence of missing values on the algorithm's efficiency and the advantages and disadvantages of deleting incomplete genotypes. We propose guidelines for missing data handling in routine analysis.

Computational Biology↗

A new approach to the analysis of analgesic drug trials, illustrated with bromfenac data.

A clinical trial of an analgesic agent compares pain relief scores (ordered categorical responses) over time among groups of patients, each subject to a painful procedure and given various doses of active agent (including zero, i.e., placebo) on demand. Patients may elect to remedicate with an active agent if their pain relief is insufficient, so the sample of patients at any given time is biased toward those with better relief. Standard analyses usually (1) fill in the missing data but make no correction for so doing and (2) treat the ordered categorical variable as continuous. Both of these create problems in interpretation and inference, but the former is more serious than the latter. An alternative analysis has been recently proposed that deals with these problems. This article presents that method for a nonstatistical audience and illustrates its use on some data from the analgesic bromfenac.

Analgesics↗

An application of hierarchical linear models to longitudinal studies.

Nursing researchers are increasingly interested in studying changes in patients' outcomes, such as physiologic and psychological status, across time. The most frequently used approaches, univariate repeated measures, multivariate repeated measures, and pre- and posttest differences, have restrictive assumptions and unrealistic data requirements. Therefore, a more flexible approach is needed. Hierarchical linear models (HLM) can be used to solve these problems. The advantages of HLM are (a) it describes each individual's growth trajectory and its relationship with initial status, (b) it is not restricted by unrealistic assumptions, (c) if solves the commonly observed problems of missing data, (d) it does not require fixed time intervals, and (e) it provides more precise estimation.

Analysis of Variance↗

A computer program for regression analysis of ordered categorical repeated measurements.

RMORD is an easy-to-use FORTRAN program for the analysis of clustered ordinal data using the method of Stram, Wei, and Ware. This method constitutes an extension of the proportional-odds model to the situation in which groups of responses are correlated. At each measurement occasion, a proportional-odds regression model is fit to the data by maximizing the occasion-specific likelihood function. The joint asymptotic distribution of the occasion-specific regression parameter estimators is obtained along with a consistent estimator of their asymptotic covariance matrix. RMORD may be used when ordinal measurements are obtained at a common set of observation times for multiple subjects or clusters. Both missing data and covariates which vary within clusters can be accommodated. The program can be run on microcomputers, workstations, and mainframe computers. Two examples illustrating the usage and features of RMORD are provided.

Age Distribution↗

A survey of laboratory and statistical issues related to farmworker exposure studies.

Developing internally valid, and perhaps generalizable, farmworker exposure studies is a complex process that involves many statistical and laboratory considerations. Statistics are an integral component of each study beginning with the design stage and continuing to the final data analysis and interpretation. Similarly, data quality plays a significant role in the overall value of the study. Data quality can be derived from several experimental parameters including statistical design of the study and quality of environmental and biological analytical measurements. We discuss statistical and analytic issues that should be addressed in every farmworker study. These issues include study design and sample size determination, analytical methods and quality control and assurance, treatment of missing data or data below the method's limits of detection, and post-hoc analyses of data from multiple studies. Key words: analytical methodology, biomarkers, laboratory, limit of detection, omics, quality control, sample size, statistics.

Agriculture↗

Drinking and driving: explaining beverage-specific risks.

OBJECTIVE: This study tested whether the association of beer drinking with drinking and driving is due to cultural norms or is an artifact arising from the demographic profile of beer drinkers (young and male), the drinking patterns of this subpopulation (frequent and heavy), and the venues in which they prefer to drink (bars and restaurants). METHOD: Telephone survey data from six U.S. communities were used to establish the demographic characteristics of drinkers, their consumption patterns, beverage preferences, preferred drinking venues and self-reported drinking and driving rates. The survey completion rate was 64.6%. A total sample of 5,231 drinkers was divided into test and validity samples. After deletion of cases with missing data, the test sample included 2,275 drinkers, of whom 985 had driven after drinking. RESULTS: Controlling for a broad set of covariates, the analyses showed that frequent consumers were more likely to drink outside the home, preferred beer and spirits to wine, and were more likely than others to drink and drive. Beverage preferences were not directly associated with drinking and driving. Beer drinkers, however, were from the subpopulation most likely to drink and drive: heavier drinking younger men, who prefer to drink at bars and restaurants. CONCLUSIONS: These results suggest that the association of beer consumption with drinking-driving arises from the circumstances in which the subpopulation of beer drinkers more commonly find themselves (as a result of their efforts to maximize, within economic constraints, the social and amenity value of drinking), as opposed to any culturally induced disposition beer drinkers may have to drink and drive.

Age Factors↗

Intention-to-treat analyses for incomplete repeated measures data.

In a randomized longitudinal clinical trial designed to evaluate two or more rival treatments, an intent-to-treat analysis requires inclusion of all randomized patients, regardless of whether they remain on protocol for the duration of the study. We propose a piecewise linear random effects model for analyzing longitudinal data where the multivariate outcome can depend upon time spent on treatment. The model assumes that data are available on a random sample of subjects after treatment is terminated, and allows either a pragmatic or explanatory analysis (as defined by Schwartz and Lellouch, 1967, Journal of Chronic Diseases 20, 637-648). Full maximum likelihood estimation of the model parameters is carried out using widely available statistical software for repeated measures with missing data and for nonparametric survival curve estimation. Data from a national, multicenter pediatric AIDS clinical trial are analyzed to illustrate implementation and interpretation of the model.

Acquired Immunodeficiency Syndrome↗