Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Culling before testing in swine: identification of culling strategy and estimation of culling precision.

The aim of this simulation study was to identify culling strategy and to estimate culling precision based on various characteristics available in field data in order to evaluate the ability to detect situations in which adjustment for missing data should be applied in genetic evaluation. Data were simulated for age at 100 kg of live weight (AGE) measured on the farm. Culling was done within (C-W/IN) or over (C-OVER) litters by deleting records from the simulated datasets with culling intensities of .33 and .67. The culling variate (CVAR) used indicated the culling precision and had genetic and phenotypic correlations of 1.00, .75, .50, .25, or .00 with AGE (r(CVAR,AGE)). We were able to distinguish between culling strategies C-OVER and C-W/IN by means of decision rules based on proportion of tested animals per litter. Estimates of r(CVAR,AGE) were obtained from calibration curves for linear regression coefficients of litter average or within-litter variance for AGE on proportion of tested animals, and within- and between-litter variance (V(W) and V(B)) for AGE. Moderate to high r(CVAR,AGE) could be identified with little error by using V(W) or V(B) in C-W/IN and V(W) in C-OVER. Within-litter variance and the weighted average of the estimates from all four characteristics were well able to detect r(CVAR,AGE) values of .50 and higher in both C-W/IN and C-OVER. In conclusion, characteristics of swine field data with missing observations contain information that makes it possible to determine culling strategy, intensity, and precision. This information can be used to decide whether missing data should be replaced by their expected values in genetic evaluation.

Animal Husbandry↗

Evaluation of the emergency department logbook for population-based surveillance of firearm-related injury.

STUDY OBJECTIVE: To evaluate existing emergency department logbooks as a source of population-based data on firearm-related injuries. METHODS: We examined the logbooks of the 24 acute care and specialty-hospital EDs in Allegheny County, Pennsylvania, to determine the number and type of data variables each contained and the completeness of reporting of each variable for selected firearm-related cases. The amount of missing data for certain variables was determined and the cause for the missing data described. RESULTS: Logbooks from 18 of the 24 eligible hospitals were reviewed. We identified 785 cases of firearm-related injury recorded between January 1, 1992, and December 31, 1993. Of the variables we selected for analysis, only date (100%), chief complaint or diagnosis (100%), name (98%), and time of admission (97%) were consistently documented. In 37% of cases the patient's county of residence could not be determined. Similarly incomplete data were found for body part injured (31%), race (28%), age (26%), sex (22%), and mode of arrival (21%). The factor most responsible for the high percentage of incomplete data was the considerable variation in the data elements contained in the different hospitals' logbooks. CONCLUSION: Missing data resulting from inconsistencies in the variables contained in different EDs' logbooks and errors of omission prevent ED logbooks, in their current state, from providing population-based data for surveillance of firearm-related injury. Standardization of such variables in ED logbooks would yield a more useful source of information for injury and disease surveillance. In lieu of standardized logbooks, multiple sources of data are necessary to establish a more comprehensive and useful system of surveillance of firearm-related injury.

Emergency Service, Hospital↗

Summary measures and statistics for comparison of quality of life in a clinical trial of cancer therapy.

Assessment of health related quality of life (QOL) has become an important endpoint in many clinical trials of cancer therapy. Most of these studies entail multiple QOL scales that are assessed repeatedly over time. As a result, the problem of multiple comparisons is a primary analytic challenge with these trials. The use of summary measures and statistics both reduces the number of hypotheses tested and facilitates the interpretation of trial results where the primary question is 'Does the overall QOL differ between treatment arms?' I present two classes of summary measures that are sensitive to consistent trends in the same direction across multiple assessment times or multiple QOL scales. Missing data strongly influences the choice between the two classes, where one class handles missing data on an individual basis, while the other class uses model-based strategies. I present the results from a clinical trial of adjuvant therapy for breast cancer that use summary measures with a focus on the practical issues that affect these analysis strategies, such as missing data and integration of QOL with efficacy endpoints such as survival.

Analysis of Variance↗

Incorporation of potential for multimedia exposure into chemical hazard scores for pollution prevention.

We are reporting a chemical hazard score for pollution prevention, called the Purdue score. The Purdue score provides a relative quantitative measure combining a variety of chemical hazards into a single quantitative hazard weighting factor for the non-expert to use. The main expected uses are to design safer products, assist in implementing and measuring achievement in pollution prevention, and as an adjunct for reporting Toxic Release Inventory data to the U.S. Government. Scoring results are presented for 200 Superfund chemicals, rank ordered by the worker hazard part of the score, by the environmental hazard part, and by combined worker and environmental hazard scores. We have reviewed the extent to which the Purdue score presently incorporates potential for multimedia pathway and multiroute absorption exposure. Until other possible uses have been carefully tested, peer-reviewed and published, users are advised to limit use of this system to planning, implementing and measuring pollution prevention and to enhancing the interpretation of Toxic Release Inventory data. The objective of this report is to look at how the structure of this score handles exposure to chemicals, both via multi-compartment pathways and multi-routes for contact or absorption health damage, as well as how it handles habitat degradation by chemicals. For all of these, the approach is built on inherent properties of each chemical, which are true for all sites and scenarios. The biggest obstacle to scoring is lack of measured chemical property data needed for scoring. We handle missing data by regression, quantitative structure activity relationship estimations, and a missing data default rule. The limitations of chemical hazard scoring are reviewed. At present, there is no widely accepted single measure of relative chemical hazard, against which to calibrate this hazard score for accuracy, except experience from industrial use. However, despite limitations, we suggest there is a strong value added for industry and society in availability of a concise, simple-to-use measure of relative chemical hazard. The Purdue score enables separate or combined consideration of chemical hazard to workers and to the natural environment. The Purdue score has potential for major cost savings in relative hazard ranking and business decision making regarding little-studied organic chemicals, because of the extensive use of advanced property estimation software. We conclude that there is societal need to warrant advanced development of this risk management tool, which is now ready for pilot use by industry. The Purdue score is mainly intended to assist and encourage businesses to implement and measure pollution prevention-especially small businesses--in a cost-effective way. The Purdue score relies strongly on sublethal toxicity, and there is practical potential for it to be used with thousands of chemicals.

Animals↗

Riluzole for amyotrophic lateral sclerosis (ALS)/motor neuron disease (MND).

BACKGROUND: Riluzole has been approved for treatment of patients with amyotrophic lateral sclerosis (ALS) in some countries but not others. Questions persist about its clinical utility because of high cost, modest efficacy and concern over adverse effects. OBJECTIVES: To examine the efficacy of riluzole in prolonging survival, and in delaying the use of surrogates (tracheostomy and mechanical ventilation) to sustain survival. SEARCH STRATEGY: Search of the Cochrane Neuromuscular Disease Group Register for randomized trials and enquiry from authors of trials and other experts in the field. The most recent search was conducted in June 1999. SELECTION CRITERIA: Types of studies: randomized trials TYPES OF PARTICIPANTS: adults with a diagnosis of ALS Types of interventions: treatment with riluzole or placebo Types of outcome measures: Primary: per cent mortality at 12 months with riluzole 100 mg Secondary: per cent mortality as a function of time with 100 mg and with all doses of riluzole, scales of neurologic function, quality of life, muscle strength and adverse events. DATA COLLECTION AND ANALYSIS: We identified two randomized trials. Each reviewer graded them for methodological quality. Data extraction was performed by a single reviewer and checked by the other two. We obtained some missing data from investigators. We performed meta-analyses with RevMan software using a fixed effects model. MAIN RESULTS: The two eligible trials included a total of 794 riluzole treated patients and 320 placebo treated patients. The methodological quality was acceptable and the trials were easily comparable. There were significant differences between the riluzole and placebo groups of both trials, in terms of the primary outcome measure, which was per cent mortality at 12 months with the 100 mg dose of riluzole. The odds ratio for the combined studies was 0.57 (95%CI 0.41 to 0.80) at 12 months. In the secondary outcome measures, there was a survival advantage with riluzole 100 mg at six, nine, 12 and 15 months, but not at three or 18 months. Pooled data from the 50, 100 and 200mg dose groups in the larger trial showed a lower per cent mortality with riluzole compared to placebo only at 12 months (odds ratio (OR) 0.64, 95% CI 0.47 to 0.88). There was no beneficial effect on bulbar function, or muscle strength. There were scant data on quality of life, but patients treated with riluzole remained in a more moderately affected health state significantly longer than placebo-treated patients (weighted mean difference (WMD) 35.5 days, 95% CI 5.9 to 65. 0). A threefold increase in serum alanine transferase was more frequent in riluzole treated patients than controls (WMD 2.65, 95% CI 1.51 to 4.65). REVIEWER'S CONCLUSIONS: Riluzole 100 mg per day appears to be modestly effective in prolonging survival for patients with ALS.

Amyotrophic Lateral Sclerosis↗

Analysis of pregnancy and other factors on detection of human papilloma virus (HPV) infection using weighted estimating equations for follow-up data.

Generalized estimating equations have been well established to draw inference for the marginal mean from follow-up data. Many studies suffer from missing data that may result in biased parameter estimates if the data are not missing completely at random. Robins and co-workers proposed using weighted estimating equations (WEE) in estimating the mean structure if drop-out occurs missing at random. We illustrate the differences between the WEE and the commonly applied available case analysis in a simulation study. We apply the WEE and reanalyse data of a longitudinal study of pregnancy and human papilloma virus (HPV) infection. We estimate the response probabilities and demonstrate that the data are not missing completely at random. Upon use of the WEE, we are able to show that pregnant women have an increased odds for an HPV infection compared with non-pregnant women after delivery (p=0.027). We conclude that the WEE are useful for dealing with monotone missing data due to drop-outs in follow-up data.

Cohort Studies↗

Early use of inhaled corticosteroids in the emergency department treatment of acute asthma.

BACKGROUND: Systemic corticosteroids therapy is central to the management of acute asthma The use of inhaled steroids may also be beneficial in this setting. OBJECTIVES: To determine the benefit of ICS for the treatment of patients with acute asthma managed in the emergency department (ED). SEARCH STRATEGY: Randomised controlled trials (RCTs) were identified from the Cochrane Airways Review Group register. Bibliographies from included studies, known reviews, and texts also were searched. SELECTION CRITERIA: Only RCTs or quasi-randomised trials were eligible for inclusion. Studies were included if patients presented with acute asthma to the ED or its equivalent, and were treated with ICS or placebo, in addition to standard therapy. Two reviewers independently selected potentially relevant articles, and then independently selected articles for inclusion. Methodological quality was independently assessed by two reviewers. DATA COLLECTION AND ANALYSIS: Data were extracted independently by two reviewers if the authors were unable to verify the validity of extracted information. Missing data were obtained from the authors or calculated from other data presented in the paper. MAIN RESULTS: Seven trials were selected for inclusion, but data were not available for one of them. In the six usable rials, (4 adult, 2 paediatric), a total of 352 patients were studied (179 ICS, 173 non-ICS treated). Patients treated with ICS were less likely to be admitted to hospital (OR: 0.30; 95% CI: 0.16, 0.57). This benefit was confined to patients not receiving concomitant systemic steroids. Such patients showed the same, but non-significant, trend towards reduced admissions compared to placebo treatment (OR 0.46; 95% CI: 0. 19, 1.11). In children, ICS appeared to be at least as effective as systemic steroids (OR 0.5; 95% CI: 0.24, 1.06). Patients receiving ICS demonstrated small, significant improvements in peak expiratory flows (PEFR WMD: 7%; 95% CI: 3, 13) and forced expiratory volumes (FEV-FEV1 WMD: 5.0%; 95% CI: 0.4, 9.7). The treatment was well tolerated, with few reported adverse side effects. REVIEWER'S CONCLUSIONS: Inhaled steroids reduced admission rates in patients with acute asthma who were not receiving concomitant systemic steroids. In children, inhaled steroids appear to be at least as effective as systemic steroids. Further research is needed to clarify the effect of ICS when used in addition to systemic corticosteroids, and to determine the optimal dose, agent, and frequency of ICS administration.

Acute Disease↗

Conditional pairwise estimation in the Rasch model for ordered response categories using principal components.

In the Rasch model for items with more than two ordered response categories, the thresholds that define the successive categories are an integral part of the structure of each item in that the probability of the response in any category is a function of all thresholds, not just the thresholds between any two categories. This paper describes a method of estimation for the Rasch model that takes advantage of this structure. In particular, instead of estimating the thresholds directly, it estimates the principal components of the thresholds, from which threshold estimates are then recovered. The principal components are estimated using a pairwise maximum likelihood algorithm which specialises to the well known algorithm for dichotomous items. The method of estimation has three advantageous properties. First, by considering items in all possible pairs, sufficiency in the Rasch model is exploited with the person parameter conditioned out in estimating the item parameters, and by analogy to the pairwise algorithm for dichotomous items, the estimates appear to be consistent, though unlike for the dichotomous case, no formal proof has yet been provided. Second, the estimates of each item parameter is a function of frequencies in all categories of the item rather than just a function of frequencies of two adjacent categories. This stabilizes estimates in the presence of low frequency data. Third, the procedure accounts readily for missing data. All of these properties are important when the model is used for constructing variables from large scale data sets which must account for structurally missing data. A simulation study shows that the quality of the estimates is excellent.

Algorithms↗

Genotype relative-risks and association tests for nuclear families with missing parental data.

The development of a new method for testing the association of genetic markers with disease is presented. This approach is applicable when sampling nuclear families with one or more affected siblings and when neither, one, or both parents are missing marker genotype data. All siblings, affected and not affected, are used to probabilistically infer the missing parental marker data. A likelihood ratio statistic, which treats marker allele frequencies as nuisance parameters, is presented to test whether all marker relative risks are equal to one (i.e., no marker association). This approach offers a solution to test for marker associations when parents are difficult to obtain.

Algorithms↗

[SF-36 Health Survey in Rehabilitation Research. Findings from the North German Network for Rehabilitation Research, NVRF, within the rehabilitation research funding program].

The SF-36 Health Survey and its 12-item abridged form is an instrument for the assessment of health related quality of life that can be used with healthy persons and patient populations. Its use has been recommended within a large German multicentre rehabilitation research programme. The paper examines missing data across all five study projects of the North German Network for Rehabilitation Research (NVRF) as well as psychometric properties of the instrument. In addition, data were compared to representative norm data using the SF-36 (SF-12) in the German National Health Survey. Results showed that there were few missing data in the SF-36. Examining the impact of age, gender and health status yielded effects of higher age and female gender on missing data. Psychometric analyses showed good to excellent results of the instrument in terms of scale fit and reliability. In terms of convergent validity, medium to high correlation of the SF-36 subscales with comparable instruments (e. g. SCL-90-R) could be found. Summarizing, the SF-36/SF-12 can be recommended for use in rehabilitation research. Analyses regarding sensitivity should be conducted in future studies.

Activities of Daily Living↗

The analysis of incomplete data in the three-period two-treatment cross-over design for clinical trials.

The additional time to complete a three-period two-treatment (3P2T) cross-over trial may cause a greater number of patient dropouts than with a two-period trial. This paper develops maximum likelihood (ML), single imputation and multiple imputation missing data analysis methods for the 3P2T cross-over designs. We use a simulation study to compare and contrast these methods with one another and with the benchmark method of missing data analysis for cross-over trials, the complete case (CC) method. Data patterns examined include those where the missingness differs between the drug types and depends on the unobserved data. Depending on the missing data mechanism and the rate of missingness of the data, one can realize substantial improvements in information recovery by using data from the partially completed patients. We recommend these approaches for the 3P2T cross-over designs.

Analysis of Variance↗

National Survey of Family Growth: design, estimation, and inference.

The purpose of this report is to document the procedures used in the 1988 National Survey of Family Growth (NSFG) to select the sample, weight the data to produce national estimates, impute missing data, and estimate sampling errors. Therefore, this report necessarily contains a great deal of technical detail. For readers who do not need this level of detail, this summary briefly describes the procedures used. The National Survey of Family Growth is conducted every few years by the National Center for Health Statistics (NCHS), a part of the U.S. Department of Health and Human Services. The purpose of the survey is to collect and publish data from a national sample of women on childbearing, factors affecting childbearing (such as contraception, sterilization, and infertility), and related aspects of maternal and infant health. Interviewing for Cycle IV of the survey was done in 1988 by Westat, Inc., under a contract with NCHS. Personal interviews were conducted between January and August of 1988 with a national sample of 8,450 women in the civilian noninstitutionalized population of the United States. Interviews were conducted in person by trained female interviewers and lasted an average of 70 minutes. The interview focused on the woman's pregnancies, if any; her use of contraception; her ability to bear children (fecundity and infertility); her use of medical services for family planning, infertility, and prenatal care; her marriage and cohabitation history, if any; and a wide range of demographic and economic characteristics. This report describes some of the main methodological aspects of the survey, including the sample design, weighting, sampling errors, and imputation of missing data. These topics will be described briefly and less technically in this summary. Each topic is discussed in more detail in the rest of the report.

Adolescent↗

Identification of significant host factors for HIV dynamics modelled by non-linear mixed-effects models.

Non-linear mixed-effects models are powerful tools for modelling HIV viral dynamics. In AIDS clinical trials, the viral load measurements for each subject are often sparse. In such cases, linearization procedures are usually used for inferences. Under such linearization procedures, however, standard covariate selection methods based on the approximate likelihood, such as the likelihood ratio test, may not be reliable. In order to identify significant host factors for HIV dynamics, in this paper we consider two alternative approaches for covariate selection: one is based on individual non-linear least square estimates and the other is based on individual empirical Bayes estimates. Our simulation study shows that, if the within-individual data are sparse and the between-individual variation is large, the two alternative covariate selection methods are more reliable than the likelihood ratio test, and the more powerful method based on individual empirical Bayes estimates is especially preferable. We also consider the missing data in covariates. The commonly used missing data methods may lead to misleading results. We recommend a multiple imputation method to handle missing covariates. A real data set from an AIDS clinical trial is analysed based on various covariate selection methods and missing data methods.

Acquired Immunodeficiency Syndrome↗

Multicenter trial of fluoxetine as an adjunct to behavioral smoking cessation treatment.

The authors evaluated the efficacy of fluoxetine hydrochloride (Prozac; Eli Lilly and Company, Indianapolis, IN) as an adjunct to behavioral treatment for smoking cessation. Sixteen sites randomized 989 smokers to 3 dose conditions: 10 weeks of placebo, 30 mg, or 60 mg fluoxetine per day. Smokers received 9 sessions of individualized cognitive-behavioral therapy, and biologically verified 7-day self-reported abstinence follow-ups were conducted at 1, 3, and 6 months posttreatment. Analyses assuming missing data counted as smoking observed no treatment difference in outcomes. Pattern-mixture analysis that estimates treatment effects in the presence of missing data observed enhanced quit rates associated with both the 60-mg and 30-mg doses. Results support a modest, short-term effect of fluoxetine on smoking cessation and consideration of alternative models for handling missing data.

Adult↗

Sensitivity analysis for the estimation of rates of change with non-ignorable drop-out: an application to a randomized clinical trial of the vitamin D3.

The vitamin D(3) trial was a repeated measures randomized clinical trial for secondary hyperparathyroidism in haemodialysis patients where the efficacy of the vitamin D(3) infusions for suppressing the secretion of parathyroid hormone (PTH) was compared among four dose groups over 12 weeks. In this trial, patients terminated the study before the scheduled end of the study due to their elevated serum calcium (Ca) level, that is, the administration of the vitamin D(3) was expected to cause hypercalcaemia as an adverse event. In this setting of monotone missingness, there is a potential for bias in estimation of mean rates of decline in PTH for each treatment group using the standard methods such as the generalized estimating equations (GEE) which ignore the observed past Ca histories. We estimated the treatment-group-specific mean rates of decline in PTH by the inverse probability of censoring weighted (IPCW) methods which account for the observed past histories of time-dependent factors that are both a predictor of drop-out and are correlated with the outcomes. The IPCW estimator can be viewed as an extension of the GEE estimator that allows for the data to be MAR but not MCAR. With missing data, it is rarely appropriate to analyse the data solely under the assumption that the missing data process is ignorable, because the assumption of ignorable missingness cannot be guaranteed to hold and is untestable from the observed data. We proposed a sensitivity analysis that examines how inference about the IPCW estimates of the treatment-group-specific mean rates of decline in PTH changes as we vary the non-ignorable selection bias parameter over a range of plausible values.

Bias↗

A global sensitivity analysis of performance of a medical diagnostic test when verification bias is present.

Current advances in technology provide less invasive or less expensive diagnostic tests for identifying disease status. When a diagnostic test is evaluated against an invasive or expensive gold standard test, one often finds that not all patients undergo the gold standard test. The sensitivity and specificity estimates based only on the patients with verified disease are often biased. This bias is called verification bias. Many authors have examined the consequences of verification bias and have proposed bias correction methods based on the assumption of independence between disease status and election for verification conditionally on the test result, or equivalently on the assumption that the disease status is missing at random using missing data terminology. This assumption may not be valid and one may need to consider adjustment for a possible non-ignorable verification bias resulting from the non-ignorable missing data mechanism. Such an adjustment involves ultimately uncheckable assumptions and requires sensitivity analysis. The sensitivity analysis is most often accomplished by perturbing parameters in the chosen model for the missing data mechanism, and it has a local flavour because perturbations are around the fitted model. In this paper we propose a global sensitivity analysis for assessing performance of a diagnostic test in the presence of verification bias. We derive a region of all sensitivity and specificity values consistent with the observed data and call this region a test ignorance region (TIR). The term 'ignorance' refers to the lack of knowledge due to the missing disease status for the not verified patients. The methodology is illustrated with two clinical examples.

Diagnostic Tests, Routine↗

Do quality of life assessments make a difference in the evaluation of cancer treatments?

The question posed by this set of quality of life papers is whether or not quality of life assessments in cancer clinical trials help evaluate the effects of cancer treatment on patient functioning. In this discussion, missing data problems, particularly those commonly found in advanced stage disease trials, are highlighted. Researchers are encouraged to investigate the extent of bias associated with missing data and to select analysis approaches accordingly. In the worst case, it may not be possible to analyze data longitudinally; descriptive or graphical portrayals of the data may be more appropriate. The importance of instrument reliability (minimizing measurement error) is emphasized for clinical trials research, particularly with respect to enhancing a trial's ability to detect quality of life differences by treatment arm. One strategy for addressing missing data is evaluated with respect to its impact on the measurement properties of the quality of life questionnaire. Clinical trials groups have been successful in obtaining quality of life data in multi-site settings and patients, by and large, appreciate the effort to include a systematic and standardized report of the effects of treatment on their functioning.

Bias↗

Using the national registry of HIV-infected veterans in research: lessons for the development of disease registries.

Disease-specific registries have many important applications in epidemiologic, clinical and health services research. Since 1989 the Department of Veterans Affairs has maintained a national HIV registry. VA's HIV registry is national in scope, it contains longitudinal data and detailed resource utilization and clinical information. To describe the structure, function, and limitations of VA's national HIV registry, and to test its accuracy and completeness. The VA's national HIV registry contains data that are electronically extracted from VA's computerized comprehensive clinical and administrative databases, called Veterans Integrated Health Systems Technology and Architecture (VISTA). We examined the number of AIDS patients and the number of new patients identified to the registry, by year, through December 1996. We verified data elements against information obtained from the medical records at five VA sites. By December 1996, 40,000 HIV-infected patients had been identified to the registry. We encountered missing data and problems with data classification. Missing data occurred for some elements related to the computer programming that creates the registry (e.g., pharmacy files), and for other elements because manual entry is required (e.g., ethnicity). Lack of a standardized data classification system was a problem, especially for the pharmacy and laboratory files. In using VA's national HIV registry we have learned important lessons, which, if taken into account in the future, could lead to the creation of model disease-specific registries.

HIV Infections↗