Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Mapping quantitative trait loci by an extension of the Haley-Knott regression method using estimating equations.

The Haley-Knott (HK) regression method continues to be a popular approximation to standard interval mapping (IM) of quantitative trait loci (QTL) in experimental crosses. The HK method is favored for its dramatic reduction in computation time compared to the IM method, something that is particularly important in simultaneous searches for multiple interacting QTL. While the HK method often approximates the IM method well in estimating QTL effects and in power to detect QTL, it may perform poorly if, for example, there is strong epistasis between QTL or if QTL are linked. Also, it is well known that the estimation of the residual variance by the HK method is biased. Here, we present an extension of the HK method that uses estimating equations based on both means and variances. For normally distributed phenotypes this estimating equation (EE) method is more efficient than the HK method. Furthermore, computer simulations show that the EE method performs well for very different genetic models and data set structures, including nonnormal phenotype distributions, nonrandom missing data patterns, varying degrees of epistasis, and varying degrees of linkage between QTL. The EE method retains key qualities of the HK method such as computational speed and robustness against nonnormal phenotype distributions, while approximating the IM method better in terms of accuracy and precision of parameter estimates and power to detect QTL.

Chromosome Mapping↗

The impact of changing methods of data collection on the reliability of self-reported drug use of adolescents.

The purpose of this study is to determine the impact of different modes of data collection on the reliability of self-reported drug use of adolescents in a panel study. Adolescents were assigned to four groups based upon the ways they chose to respond to the survey instruments: 1) mailed questionnaires in both years, 2) survey interview in one year and mailed questionnaire in the next year, 3) mailed questionnaire in one year and survey interview in the following year, and 4) survey interview in both years. The quality of the self-reported data was examined in terms of return rates, missing data, internal consistency, and consistency of reported information over time. No significant differences were found between groups, suggesting that the mode of data collection does not affect the reliability of adolescents' self-reports of substance use.

Adolescent↗

The life history calendar: a technique for collecting retrospective data.

"This paper details the authors' selection, design, and use of a life history calendar (LHC) to collect retrospective life course data. A sample of nine hundred [U.S.] 23-year-olds, originally interviewed in 1980, were asked about the incidence and timing of various life events in the nine years since their 15th birthday.... The following aspects of the LHC are described: (a) the concept, uses, and advantages of the LHC, (b) the time units and domains used, (c) the mode of recording the responses and the decisions and problems involved, (d) interviewer training, and (e) coding. The following results attest to the accuracy of the LHC retrospective data: (a) only four of the calendars had missing data in any month; (b) the data obtained in 1980 about current work, school attendance, marriage, and children showed a remarkable correspondence to the retrospective 1985 LHC reports of these events; (c) the interviewers were positive about the LHC's ability to increase respondent recall."

Americas↗

[Longitudinal studies: concepts and particularities].

In this review the definition of "longitudinal study" is analysed. Most current textbooks on epidemiology do not define a longitudinal study, whereas statistical textbooks do. It is more common to talk about longitudinal data than about longitudinal studies. A longitudinal study implies the existence of repeated measurements (more than two) across follow-up. According to these ideas, a longitudinal study can be considered a subtype of cohort study that, in contrast with life-table cohort studies, allows inference to the subject level, to analyze changes in variables (exposures and outcomes) and transitions among different health states. The characteristics of this design force to paid special attention to quality control during data collection, losses during follow-up, and missing data in some measurements. The statistical analysis should take repeated measures into account, and it is what finally gives the longitudinal character to a study with repeated measurements.

Biomedical Research↗

Applications of multiple imputation to the analysis of censored regression data.

The first part of the article reviews the Data Augmentation algorithm and presents two approximations to the Data Augmentation algorithm for the analysis of missing-data problems: the Poor Man's Data Augmentation algorithm and the Asymptotic Data Augmentation algorithm. These two algorithms are then implemented in the context of censored regression data to obtain semiparametric methodology. The performances of the censored regression algorithms are examined in a simulation study. It is found, up to the precision of the study, that the bias of both the Poor Man's and Asymptotic Data Augmentation estimators, as well as the Buckley-James estimator, does not appear to differ from zero. However, with regard to mean squared error, over a wide range of settings examined in this simulation study, the two Data Augmentation estimators have a smaller mean squared error than does the Buckley-James estimator. In addition, associated with the two Data Augmentation estimators is a natural device for estimating the standard error of the estimated regression parameters. It is shown how this device can be used to estimate the standard error of either Data Augmentation estimate of any parameter (e.g., the correlation coefficient) associated with the model. In the simulation study, the estimated standard error of the Asymptotic Data Augmentation estimate of the regression parameter is found to be congruent with the Monte Carlo standard deviation of the corresponding parameter estimate. The algorithms are illustrated using the updated Stanford heart transplant data set.

Algorithms↗

Ensuring data quality in a multicenter clinical trial: remote site data entry, central coordination and feedback.

In an ongoing multicenter clinical trial, "Treatment Strategies in Schizophrenia," the five participating sites have the capacity to perform a variety of tasks or study functions independently. These tasks include (a) verification of diagnostic eligibility through the use of computerized decision algorithms; (b) assignment of patients to treatment based on prognostic indicators using a computerized randomization algorithm; (c) entry of data into a microcomputer using a clinical trial data management system that performs simple range and missing data item checks; and (d) regular transfer of all data to the central coordinating team. The clinical trial data management system employed allows for both independent site functioning and assurance of consistency across sites. The integration of a variety of software outside the main data management system provides the central coordinators with the tools to monitor critical data as it is collected, as well as the capacity to assess the flow, quality, and uniformity of the ongoing trial.

Clinical Trials as Topic↗

Developing an outcomes infrastructure for nursing. The Outcomes Taskforce.

An infrastructure to support the evaluation of patient care sensitive to the intervention of nursing personnel is being developed within a major health maintenance organization. In addition to traditional administrative measures of care, the database infrastructure will include measures of the patient's functional status, knowledge and engagement in care and psychosocial well-being. These measures are believed to be particularly sensitive to the independent intervention of the nurse. Reported here are the structures in place to monitor and support the reliability and validity of the administrative data elements; algorithm elements created to account for missing data; the model for the first generation of successful practice reports and the results of a study establishing the content validity of the clinical data elements.

Costs and Cost Analysis↗

Application of random-effects regression models in relapse research.

This article describes and illustrates use of random-effects regression models (RRM) in relapse research. RRM are useful in longitudinal analysis of relapse data since they allow for the presence of missing data, time-varying or invariant covariates, and subjects measured at different timepoints. Thus, RRM can deal with "unbalanced" longitudinal relapse data, where a sample of subjects are not all measured at each and every timepoint. Also, recent work has extended RRM to handle dichotomous and ordinal outcomes, which are common in relapse research. Two examples are presented from a smoking cessation study to illustrate analysis using RRM. The first illustrates use of a random-effects ordinal logistic regression model, examining longitudinal changes in smoking status, treating status as an ordinal outcome. The second example focuses on changes in motivation scores prior to and following a first relapse to smoking. This latter example illustrates how RRM can be used to examine predictors and consequences of relapse, where relapse can occur at any study timepoint.

Alcoholism↗

Efficient inference of haplotypes from genotypes on a large animal pedigree.

We present a simple algorithm for reconstruction of haplotypes from a sample of multilocus genotypes. The algorithm is aimed specifically for analysis of very large pedigrees for small chromosomal segments, where recombination frequency within the chromosomal segment can be assumed to be zero. The algorithm was tested both on simulated pedigrees of 155 individuals in a family structure of three generations and on real data of 1149 animals from the Israeli Holstein dairy cattle population, including 406 bulls with genotypes, but no females with genotypes. The rate of haplotype resolution for the simulated data was >91% with a standard deviation of 2%. With 20% missing data, the rate of haplotype resolution was 67.5% with a standard deviation of 1.3%. In both cases all recovered haplotypes were correct. In the real data, allele origin was resolved for 22% of the heterozygous genotypes, even though 70% of the genotypes were missing. Haplotypes were resolved for 36% of the males. Computing time was insignificant for both data sets. Despite the intricacy of large-scale real pedigree genotypes, the proposed algorithm provides a practical rule-based solution for resolving haplotypes for small chromosomal segments in commercial animal populations.

Algorithms↗

Non-ignorable missing covariates in generalized linear models.

We propose a likelihood method for estimating parameters in generalized linear models with missing covariates and a non-ignorable missing data mechanism. In this paper, we focus on one missing covariate. We use a logistic model for the probability that the covariate is missing, and allow this probability to depend on the incomplete covariate. We allow the covariates, including the incomplete covariate, to be either categorical or continuous. We propose an EM algorithm in this case. For a missing categorical covariate, we derive a closed form expression for the E- and M-steps of the EM algorithm for obtaining the maximum likelihood estimates (MLEs). For a missing continuous covariate, we use a Monte Carlo version of the EM algorithm to obtain the MLEs via the Gibbs sampler. The methodology is illustrated using an example from a breast cancer clinical trial in which time to disease progression is the outcome, and the incomplete covariate is a quality of life physical well-being score taken after the start of therapy. This score may be missing because the patients are sicker, so this covariate could be non-ignorably missing.

Algorithms↗

Numerical algorithms for spatial registration of line fiducials from cross-sectional images.

We present several numerical algorithms for six-degree-of-freedom rigid-body registration of line fiducial objects to their marks in cross-sectional planar images, such as those obtained in CT and MRI, given the correspondence between the marks and line fiducials. The area of immediate application is frame-based stereotactic procedures, such as radiosurgery and functional neurosurgery. The algorithms are also suitable to problems where the fiducial pattern moves inside the imager, as is the case in robot-assisted image-guided surgical applications. We demonstrate the numerical methods on clinical CT images and computer-generated data and compare their performance in terms of robustness to missing data, robustness to noise, and speed. The methods show two unique strengths: (1) They provide reliable registration of incomplete fiducial patterns when up to two-thirds of the total fiducials are missing from the image; and (2) they are applicable to an arbitrary combination of line fiducials without algorithmic modification. The average speed of the fastest algorithm is 0.3236 s for six fiducial lines in real CT data in a Matlab implementation.

Algorithms↗

Antibiotics for acute bronchitis.

BACKGROUND: Antibiotic treatment of acute bronchitis, which is one of the most common illnesses seen in primary care, is controversial. Most clinicians prescribe antibiotics in spite of expert recommendations against this practice. OBJECTIVES: People with acute bronchitis may show little evidence of bacterial infection. If effective, antibiotics could shorten the course of the disease. However if they are not effective, the risk of antibiotic resistance may be increased. The objective of this review was to assess the effects of antibiotic treatment for patients with a clinical diagnosis of acute bronchitis. SEARCH STRATEGY: We searched Medline, Embase, reference lists of articles and the authors' personal collections up to 1996, and Scisearch from 1989 to 1996. SELECTION CRITERIA: Randomised trials comparing any antibiotic therapy with placebo in acute bronchitis. DATA COLLECTION AND ANALYSIS: At least two reviewers extracted data and assessed trial quality. MAIN RESULTS: Eight trials involving 750 patients aged eight to over 65 and including smokers and non-smokers were included. The quality of the trials was variable. A variety of outcome measures were assessed. In many cases, only outcomes that showed a statistically significant difference between groups were reported. Overall, patients receiving antibiotics had slightly better outcomes than did those receiving placebo. They were less likely to report feeling unwell at a follow up visit (odds ratio 0.42, 95% confidence interval 0.22 to 0.82), to show no improvement on physician assessment (odds ratio 0.43; 0.23 to 0.79), or to have abnormal lung findings (odds ratio 0.33, 95% confidence interval 0.13 to 0.86), and had a more rapid return to work or usual activities (weighted mean difference 0.7 days earlier, 95% confidence interval 0.2 to 1. 3). Antibiotic-treated patients reported significantly more adverse effects (odds ratio 1.64; 1.05 to 2.57) such as nausea, vomiting, headache, skin rash or vaginitis. REVIEWER'S CONCLUSIONS: Antibiotics appear to have a modest beneficial effect in the treatment of acute bronchitis, with a corresponding small risk of adverse effects. The benefits of antibiotics may be overestimated in this analysis because of the tendency of published reports to include complete data on only the outcomes found to be statistically significant.

Acute Disease↗

Antibiotics for acute bronchitis.

BACKGROUND: Antibiotic treatment of acute bronchitis, which is one of the most common illnesses seen in primary care, is controversial. Most clinicians prescribe antibiotics in spite of expert recommendations against this practice. OBJECTIVES: The objective of this review was to assess the effects of antibiotic treatment for patients with a clinical diagnosis of acute bronchitis. SEARCH STRATEGY: In this updated review, we searched the Cochrane Central Register of Controlled trials (CENTRAL) (The Cochrane Library Issue 2, 2004); MEDLINE (January 1966 to March 2004); EMBASE (January 2000 to December 2003); SciSearch from 1989 to 2004; reference lists of articles and the authors' personal collections up to 1996, and also wrote to study authors and drug manufacturers. EMBASE has previously been searched from 1974 to 2000). SELECTION CRITERIA: Randomised controlled trials comparing any antibiotic therapy with placebo in acute bronchitis or acute productive cough without other obvious cause in patients without underlying pulmonary disease. DATA COLLECTION AND ANALYSIS: At least two reviewers extracted data and assessed trial quality. Authors were contacted for missing data. MAIN RESULTS: Nine trials involving over 750 patients aged eight to over 65 and including smokers and non-smokers were included in the primary analysis. The quality of the trials was variable. A variety of outcome measures were assessed. Overall, patients receiving antibiotics had better outcomes than did those receiving placebo. At a follow-up visit, they were less likely to have a cough (relative risk (RR) 0.64, 95% confidence interval (CI) 0.49 to 0.85; number-needed-to-treat (NNT) 5; 95% CI 3 to 14), show no improvement on physician assessment (RR 0.52; 95% CI 0.31 to 0.87; NNT 14; 95% CI 8 to 50), or have abnormal lung findings (RR 0.48; 95% CI 0.26 to 0.89; NNT 11; 95% CI 6 to 50); and had shorter durations of cough (weighted mean difference 0.58 days; 95% CI 0.01 to 1.16 days), productive cough (weighted mean difference (WMD) 0.52 days; 95% CI 0.01 to 1.03 days), and feeling ill (WMD 0.58 days; 95% CI 0.00 to 1.16 days). There were no significant differences regarding the presence of night cough, productive cough, or activity limitations at follow up, or in the mean duration of activity limitations. The benefits of antibiotics were less apparent in a sensitivity analysis that included data from two other studies of patients with upper respiratory tract infections with productive cough. There was a non significant trend towards an increase in adverse effects in the antibiotic group, relative risk (RR) 1.22 (95% CI 0.94 to 1.58). REVIEWERS' CONCLUSIONS: Overall, antibiotics appear to have a modest beneficial effect in patients who are diagnosed with acute bronchitis. The magnitude of this benefit, however, is similar to that of the detriment from potential adverse effects.

Acute Disease↗

CHIP: Defining a dimension of the vulnerability to attention deficit hyperactivity disorder (ADHD) using sibling and individual data of children in a community-based sample.

We are taking a quantitative trait approach to the molecular genetic study of attention deficit hyperactivity disorder (ADHD) using a truncated case-control association design. An epidemiological sample of children aged 5 to 15 years was evaluated for symptoms of ADHD using a parent rating scale. Individuals scoring high or low on this scale were selected for further investigation with additional questionnaires and DNA analysis. Data in studies like this are typically complicated. In the study reported on here, individuals have from 1 to 4 questionnaires completed on them and the sample is composed of a mixture of singletons and siblings. In this paper, we describe how we used a genetic hierarchical model to fit our data, together with a twin dataset, in order to estimate genetic factor loadings. Correlation matrices were estimated for our data using a maximum likelihood approach to account for missing data. We describe how we used these results to create a composite score, the heritability of which was estimated to be acceptably high using the twin dataset. This score measures a quantitative dimension onto which molecular genetic data will be mapped.

Adolescent↗

Use of OSWALD for analyzing longitudinal data with informative dropout.

OSWALD (Object-oriented Software for the Analysis of Longitudinal Data) is flexible and powerful software written for S-PLUS for the analysis of longitudinal data with dropout for which there is little other software available in the public domain. The implementation of OSWALD is described through analysis of a psychiatric clinical trial that compares antidepressant effects in an elderly depressed sample and a simulation study. In the simulation study, three different dropout mechanisms: completely random dropout (CRD), random dropout (RD) and informative dropout (ID), are considered and the results from using OSWALD are compared across mechanisms. The parameter estimates for ID-simulated data show less bias with OSWALD under the ID missing data assumption than under the CRD or RD assumptions. Under an ID mechanism, OSWALD does not provide standard error estimates. We supplement OSWALD with a bootstrap procedure to derive the standard errors. This report illustrates the usage of OSWALD for analyzing longitudinal data with dropouts and how to draw appropriate conclusions based on the analytic results under different assumptions regarding the dropout mechanism.

Humans↗

Statistical analysis of data from studies on experimental autoimmune encephalomyelitis.

Research in multiple sclerosis often employs animal models of the disease, especially experimental autoimmune encephalomyelitis (EAE) in rodents. The statistical analysis procedures chosen for these studies are often suboptimal, either because of violations of the assumptions of the procedure or because the analysis selected is inappropriate for the research question. In this paper, we discuss the types of research questions frequently asked in EAE studies and suggest appropriate and useful research designs and statistical methods that will optimize the information contained within the data. We also discuss other troublesome issues such as missing data, atypical disease profiles, and power analysis.

Animals↗

A case-control study on scrapie in Norwegian sheep flocks.

Scrapie first was detected in indigenous sheep in Norway in 1981, and from 1995 to 1997 an increase in the number of flocks with scrapie cases was recorded. These flocks were mainly in one geographical region. A study to identify risk factors for scrapie was conducted. The study had three frequency-matched controls selected for every case within the same Veterinary District. A questionnaire was submitted to 176 sheep flocks (42 had been scrapie flocks). The data obtained by the questionnaire were linked to data collected from governmental and industry registers. After imputing missing data using single random imputation, the statistical analysis was performed using multivariable conditional logistic regression. Purchase of female sheep from scrapie flocks, sharing of rams, or sharing of pastures between different flocks were the risk factors associated with the occurrence of scrapie. Of factors potentially sustaining and promoting the infection in the flock, number of winter-fed sheep, number of buildings for housing sheep, rams and ewes shared room during mating period and increase in the flock size were associated with scrapie. We interpret these findings to show that factors involving transfer of sheep between flocks or direct contact between sheep of different flocks are important for the spread of scrapie. Management factors are important for the development of scrapie. However, it was not possible to discriminate between the different management factors in this study at the flock level. Also, factors indicating awareness and interest of the farmer (as well as willingness to contact a veterinarian for diseased sheep) were related to the detection of scrapie in the flock.

Animal Husbandry↗

Resolving the long-term trends of polycyclic aromatic hydrocarbons in the Canadian Arctic atmosphere.

Polycyclic aromatic hydrocarbon (PAH) air concentrations measured over the period 1992-2000 at the Canadian High Arctic station of Alert were subject to time-series analysis using dynamic harmonic regression (DHR). For most of the PAHs, the DHR model fit to the observed data was good, with DHR capable of interpolating over missing data points during periods when air concentrations were below detection limits. As expected, DHR identified seasonal increases in PAH air concentrations. However, it has also identified additional, subtler "seasonal" patterns as a series of harmonics with varying periodicity. For example, a regular summer high in air concentrations was apparent for many PAHs, particularly the lower molecular weight (two- to three-ringed) compounds, which may be attributed to summertime regional combustion events such as forestfires and/or revolatilization from surfaces (e.g., soil and oceans, as well as arctic surfaces). Comparison of wintertime PAH concentrations (where sigmaPAH ranged from 260 to 516 pg m(-3)) with an earlier arctic study did not reveal a reduction in PAH levels. However, removal of the seasonal components by DHR revealed a declining trend in PAH concentrations over the 1992-2000 period. For many lighter PAHs, this was typified by a linear decrease over the whole time series, although, for the higher molecular weight PAHs, a marked reduction was apparent in the first few years of sampling followed by a leveling off in concentrations by the mid/late-1990s. This behavior is similar to reported trends of other air pollutants in the Arctic, may be attributed to the decline in Soviet industry during the early 1990s, and has implications regarding the major PAH sources affecting the Arctic.

Air Pollutants↗