Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Precision and type I error rate in the presence of genotype errors and missing parental data: a comparison between the original transmission disequilibrium test (TDT) and TDTae statistics.

BACKGROUND: Two factors impacting robustness of the original transmission disequilibrium test (TDT) are: i) missing parental genotypes and ii) undetected genotype errors. While it is known that independently these factors can inflate false-positive rates for the original TDT, no study has considered either the joint impact of these factors on false-positive rates or the precision score of TDT statistics regarding these factors. By precision score, we mean the absolute difference between disease gene position and the position of markers whose TDT statistic exceeds some threshold. METHODS: We apply our transmission disequilibrium test allowing for errors (TDTae) and the original TDT to phenotype and modified single-nucleotide polymorphism genotype simulation data from Genetic Analysis Workshop. We modify genotype data by randomly introducing genotype errors and removing a percentage of parental genotype data. We compute empirical distributions of each statistic's precision score for a chromosome harboring a simulated disease locus. We also consider inflation in type I error by studying markers on a chromosome harboring no disease locus. RESULTS: The TDTae shows median precision scores of approximately 13 cM, 2 cM, 0 cM, and 0 cM at the 5%, 1%, 0.1%, and 0.01% significance levels, respectively. By contrast, the original TDT shows median precision scores of approximately 23 cM, 21 cM, 15 cM, and 7 cM at the corresponding significance levels, respectively. For null chromosomes, the original TDT falsely rejects the null hypothesis for 28.8%, 14.8%, 5.4%, and 1.7% at the 5%, 1%, 0.1% and 0.01%, significance levels, respectively, while TDTae maintains the correct false-positive rate. CONCLUSION: Because missing parental genotypes and undetected genotype errors are unknown to the investigator, but are expected to be increasingly prevalent in multilocus datasets, we strongly recommend TDTae methods as a standard procedure, particularly where stricter significance levels are required.

Chromosomes, Human, Pair 3↗

A PC program for diagnosing abnormal growth, growth velocity and acceleration from longitudinal observations.

A PC program, written in GAUSS386i, implementing Zerbe's (Growth, 43 (1979) 263-272) procedure for diagnosis on the basis of longitudinal data is described, illustrated and made available to interested readers. Given longitudinal observations on N normal individuals, this technique can be used to characterize normal growth, velocity and acceleration, and to determine whether or not a new individual can be considered normal with respect to any or all of these parameters. Missing data are allowed, and there is no requirement that the variable whose growth is being monitored has a normal distribution. The method and program are illustrated using a data set with a substantial amount of missing data. Information on obtaining a copy of the program and hardware requirements are given in the Appendix.

Algorithms↗

Source of bias in prenatal care utilization indices: implications for evaluating the Medicaid expansion.

BACKGROUND: Recent expansions in eligibility for coverage of prenatal care services by the Medicaid program reflect national initiatives to improve pregnancy outcomes. This study investigates the potential impact that completeness of reporting of prenatal care and gestational age variables and strategies to impute missing data may have on evaluations of the Medicaid expansion. METHODS: This study, examining 15 years of vital record data from a single state and comparing 1 year of data from four mid-Atlantic states, selected single live births to resident mothers for analyses. The "day 15" and the "preceding case" methods were used to impute missing gestational age data. RESULTS: Considerable temporal and geographic variation was detected in completeness of reporting of variables used to construct prenatal care indices. After imputing values for cases with missing data, the proportion of cases for which adequacy of prenatal care utilization could not be determined ranged from 3% to 24% among the states investigated. For those cases where gestational age data could be imputed, the distribution of prenatal care utilization was not markedly disparate from those cases with complete reporting of gestational age. CONCLUSIONS: The results indicate that variations in reporting, decisions regarding the treatment of missing data, and the choice of the denominator can alter prenatal care utilization percentages and have implications for evaluations of the impact of the recent Medicaid expansion on prenatal care utilization.

Bias↗

Surrogate endpoints in clinical trials: cardiovascular diseases.

A surrogate endpoint in a cardiovascular clinical trial is defined as endpoint measured in lieu of some other so-called 'true' endpoint. A surrogate is especially useful if it is easily measured and highly correlated with the true endpoint. Often the 'true' endpoint is one with clinical importance to the patient, for example, mortality or a major clinical outcome, while a surrogate is one biologically closer to the process of disease, for example, ejection fraction. Use of the surrogate can often lead to dramatic reductions in sample size and much shorter studies than use of the true endpoint. We discuss several problems common in trials with surrogate endpoints. Most important is the effect of missing data, especially in the face of informative censoring. Possible solutions are the assignment of scores or formal penalties to missing data.

Cardiovascular Diseases↗

Multiple imputation methods for longitudinal blood pressure measurements from the Framingham Heart Study.

Missing data are a great concern in longitudinal studies, because few subjects will have complete data and missingness could be an indicator of an adverse outcome. Analyses that exclude potentially informative observations due to missing data can be inefficient or biased. To assess the extent of these problems in the context of genetic analyses, we compared case-wise deletion to two multiple imputation methods available in the popular SAS package, the propensity score and regression methods. For both the real and simulated data sets, the propensity score and regression methods produced results similar to case-wise deletion. However, for the simulated data, the estimates of heritability for case-wise deletion and the two multiple imputation methods were much lower than for the complete data. This suggests that if missingness patterns are correlated within families, then imputation methods that do not allow this correlation can yield biased results.

Adult↗

Using technology to improve longitudinal studies: self-reporting with ChronoRecord in bipolar disorder.

OBJECTIVES: Longitudinal studies are an optimal approach to investigating the highly variable course and outcome associated with bipolar disorder, but are expensive and often have missing data. This study validates patient self-reported mood ratings using a home computer-based system (ChronoRecord) with clinician mood ratings on the Hamilton Depression Rating scale (HAMD) and Young Mania Rating scale (YMRS), and investigates the patient acceptance of the technology. METHODS: After brief training, outpatients with bipolar disorder were given the software version of an established paper based self-reporting form (ChronoSheet) to install on a home computer. Every day for 3 months, patients entered mood, medications, sleep, life events, and menstrual data. Weight was entered weekly. RESULTS: Eighty of 96 (83%) patients returned 8662 days of data. The mean days of data returned was 114.7 +/- 32.3 SD The mean percentage of days missing for mood data was 6.1% +/- 9.3 SD, equivalent to missing 7.3 day of the 114.7 days. Self-reported ratings were strongly correlated with clinician HAMD ratings (-0.683, p < 0.001). CONCLUSIONS: This study demonstrates concurrent validity between ChronoRecord and HAMD. Patients with bipolar disorder showed high acceptance of a computer-based system for self-reporting of daily data. Automation of data collection can reduce missing data and eliminate errors associated with data entry. This technology also enables on-going feedback for both patient and researcher during a long-term study.

Adult↗

Evaluation of most frequent errors in daily compilation and use of a radiation treatment chart.

Between 1 March and 30 April (1994) we recorded the errors detected by the physician, the radiographer or the physicist during prescription, preparation and execution phases of 227 treatment plans. The radiation treatment modalities used were the following: (i) single or opposed fields, moulded or not; and (ii) multiple fields or kinetic techniques. The total number of sessions performed is 1613 with the cobalt unit and 2131 with the linear accelerator (total, 3744). The total number of wrong data is 155, consisting of 24/227 (10.5%) in compilation, 22/3744 (0.58%) in execution and 109/3744 (2.9%) in registration phases. The number of missing data is 140, consisting of 10/227 (4.4%) in compilation, 9/3744 (0.2%) in execution and 121/3744 (3.2%) in registration phases. Wrong data of compilation, even if in high rate (10.5%), were all found during the same compilation phase or at the first treatment, so that they did not alter the exactness of the treatment plan. Wrong and missing data, found in the registration phase (2.9% and 3.2%, respectively), depend on the repetition of daily treatment and on the registration of data on the chart after having digitized them on the display.

Cobalt Radioisotopes↗

An appraisal of echocardiography as an epidemiological tool. The Strong Heart Study.

PURPOSE: Despite the prognostic importance of left ventricular (LV) mass (LVM) by M-mode echocardiography, concern exists about bias introduced by missing data. The American Society of Echocardiography has made recommendations for linear measurements of LV wall thickness and internal dimension used to calculate LVM, but it is unknown whether their substitution for suboptimal M-modes improves measurement yield and reduces bias. METHODS: LVM measurement yield and associations of missing data with risk factors were assessed in 3487 American Indian participants in Strong Heart Study (SHS) Phase II and compared to data from other large-scale studies. RESULTS: In SHS, LVM was measurable in 3188 (91%) participants compared to 4947/6148 (80%) Framingham participants studied by classic M-mode technique, with less decrease in measurement yield with age in SHS. In univariate SHS analyses, missing LVM was significantly associated with male gender, older age, greater height, body mass index, fat-free mass, waist/hip ratio, fibrinogen and, marginally, diabetes but not smoking, blood pressure, or lipids. In logistic regression analysis, missing LVM was independently associated with male gender, older age, greater body mass index and lower forced expiratory volume (FEV(1)) (with a low multiple R(2) [.04]), but not other risk factors. Doppler stroke volume, a measure of hemodynamic volume load, was measurable in 96% of SHS participants; missing values were weakly associated with older age, higher creatinine and lower FEV(1). During 48 +/- 11 months of follow-up, inability to measures LV mass or stroke volume was not associated with higher rates of cardiovascular events or death (p = 0.25 to 0.96). CONCLUSIONS: Improvements in echocardiographic methods have increased the yield of LVM in middle-aged and older adults and allow even more consistent assessment of cardiac volume load. Despite small persistent biases, due to associations of missing LVM and Doppler stroke volume data with male gender, greater obesity, lower FEV(1) and (for LVM only) older age, individuals with missing measurement are not at higher risk of cardiovascular events.

Aged↗

Medical image registration with partial data.

We have developed a general-purpose registration algorithm for medical images and volumes. The transformation between images is modeled as locally affine but globally smooth, and explicitly accounts for local and global variations in image intensities. An explicit model of missing data is also incorporated, allowing us to simultaneously segment and register images with partial or missing data. The algorithm is built upon a differential multiscale framework and incorporates the expectation maximization algorithm. We show that this approach is highly effective in registering a range of synthetic and clinical medical images.

Algorithms↗

Long acting beta-agonists versus theophylline for maintenance treatment of asthma.

BACKGROUND: Theophylline and long acting beta2-agonists are bronchodilators used for the management of persistent asthma symptoms, especially nocturnal asthma. They represent different classes of drug with differing side-effect profiles. OBJECTIVES: To assess the comparative efficacy, safety and side-effects of long-acting beta-agonists and theophylline in the maintenance treatment of asthma. SEARCH STRATEGY: Randomised, controlled trials (RCTs) were identified using the Cochrane Airways Group register. The register was searched using the following terms: asthma and theophylline and long acting beta-agonist or formoterol or foradile or eformoterol or salmeterol or bambuterol or bitolterol. Titles and abstracts were then screened to identify potentially relevant studies. The bibliography of each RCT was searched for additional RCTs. Authors of identified RCTs were contacted for other relevant published and unpublished studies. SELECTION CRITERIA: All included studies were RCTs involving adults and children with clinical evidence of asthma. These studies must have compared oral sustained release and/or dose adjusted theophylline with an inhaled long-acting beta-agonist. DATA COLLECTION AND ANALYSIS: Potentially relevant trials, identified by screening titles and/or abstracts, were obtained. Two reviewers independently assessed full text versions of these trials to decided whether the trial should be included in the review, and assessed its methodological quality. Where there was disagreement between reviewers, this was resolved by consensus, or reference to a third party. Data were extracted by two independent reviewers. Inter-rater reliability was assessed by simple agreement. Study authors were contacted to clarify randomisation methods, provide missing data, verify the data extracted and identify unpublished studies. Relevant pharmaceutical manufacturers were also contacted. MAIN RESULTS: Six trials met the inclusion criteria. Five used salmeterol and one, biltoterol. They were of varying quality. There was a trend for salmeterol to improve FEV1 more than theophylline in three studies and salmeterol use was associated with more symptom free nights. Bitolterol, used in only one study, was reported to be less effective than theophylline. Subjects taking salmeterol experienced fewer adverse events than those using theophylline (Relative Risk 0.38; 95%Confidence Intervals 0.25, 0.57). Significant reductions were reported for central nervous system adverse events (Relative Risk 0.51; 95%Confidence Intervals 0.30, 0.88) and gastrointestinal adverse events (Relative Risk 0.32; 95%Confidence Intervals 0.17, 0.59). REVIEWER'S CONCLUSIONS: Salmeterol may be more effective than theophylline in reducing asthma symptoms including night waking and improving lung function. More adverse events occurred in subjects using theophylline when compared to salmeterol.

Adrenergic beta-Agonists↗

Estimating the distribution of times from HIV seroconversion to AIDS using multiple imputation. Multicentre AIDS Cohort Study.

Multiple imputation is a model based technique for handling missing data problems. In this application we use the technique to estimate the distribution of times from HIV seroconversion to AIDS diagnosis with data from a cohort study of 4954 homosexual men with 4 years of follow-up. In this example the missing data are the dates of diagnosis with AIDS. The imputation procedure is performed in two stages. In the first stage, we estimate the residual AIDS-free time distribution as a function of covariates measured on the study participants with data provided by the participants who were seropositive at study entry. Specifically, we assume the residual AIDS-free times follow a log-normal regression model that depends on the covariates measured at enrolment on the seropositive participants. In the second stage we impute the date of AIDS diagnosis for the participants who seroconverted during the course of the study and are AIDS-free with use of the log-normal distribution estimated in the first stage and the covariates from each seroconverter's latest visit. The estimated proportions developing AIDS within 4 and within 7 years of seroconversion are 15 and 36 per cent respectively, with associated 95 per cent confidence intervals of (10, 21) and (26, 47) per cent. We discuss the Bayesian foundations of the multiple imputation technique and the statistical and scientific assumptions.

AIDS Serodiagnosis↗

Calculation of multipoint likelihoods using flanking marker data: a simulation study.

The calculation of multipoint likelihoods is computationally challenging, with the exact calculation of multipoint probabilities only possible on small pedigrees with many markers or large pedigrees with few markers. This paper explores the utility of calculating multipoint likelihoods using data on markers flanking a hypothesized position of the trait locus. The calculation of such likelihoods is often feasible, even on large pedigrees with missing data and complex structures. Performance characteristics of the flanking marker procedure are assessed through the calculation of multipoint heterogeneity LOD scores on data simulated for Genetic Analysis Workshop 14 (GAW14). Analysis is restricted to data on the Aipotu population on chromosomes 1, 3, and 4, where chromosomes 1 and 3 are known to contain disease loci. The flanking marker procedure performs well, even when missing data and genotyping errors are introduced.

Chromosome Mapping↗

Gametic phase estimation over large genomic regions using an adaptive window approach.

The authors present ELB, an easy to programme and computationally fast algorithm for inferring gametic phase in population samples of multilocus genotypes. Phase updates are made on the basis of a window of neighbouring loci, and the window size varies according to the local level of linkage disequilibrium. Thus, ELB is particularly well suited to problems involving many loci and/or relatively large genomic regions, including those with variable recombination rate. The authors have simulated population samples of single nucleotide polymorphism genotypes with varying levels of recombination and marker density, and find that ELB provides better local estimation of gametic phase than the PHASE or HTYPER programs, while its global accuracy is broadly similar. The relative improvement in local accuracy increases both with increasing recombination and with increasing marker density. Short tandem repeat (STR, or microsatellite) simulation studies demonstrate ELB's superiority over PHASE both globally and locally. Missing data are handled by ELB; simulations show that phase recovery is virtually unaffected by up to 2 per cent of missing data, but that phase estimation is noticeably impaired beyond this amount. The authors also applied ELB to datasets obtained from random pairings of 42 human X chromosomes typed at 97 diallelic markers in a 200 kb low-recombination region. Once again, they found ELB to have consistently better local accuracy than PHASE or HTYPER, while its global accuracy was close to the best.

Algorithms↗

Item response models for longitudinal quality of life data in clinical trials.

Assessment of quality of life is becoming standard in clinical trials. A popular method for measuring quality of life is with instruments which utilize multiple-item subscales, in which each item is scored on a Likert scale. Most statistical methods for the analysis of quality of life data in clinical trials do not explicity consider the properties and psychometric features which were of interest in scale development. In this regard, the measurement and statistical summarization of quality of life data, along with the clinical interpretation, can be somewhat disjoint from the psychometric concerns of the development process. The aim of this paper is to address the complicated issues present in analysing multiple-item ordinal quality of life data in clinical trials while maintaining fidelity to the psychometrical foundations upon which quality of life instruments are built. Accomplishing this will require the development of item response models which recognize the longitudinal aspects of clinical trial designs as well as the potential problem of informatively missing data. A general item response modeling approach is presented for longitudinal multiple-item quality of life data measured on ordinal scales with model components for missing data mechanisms and latent trait regression on treatment indicators and other covariates.

Clinical Trials as Topic↗

Single-channel data and missed events: analysis of a two-state Markov model.

Patch-clamp recording permits investigation of the gating kinetics of single ion channels. Careful statistical analysis of kinetic data can yield clues as to the molecular events underlying channel gating. However, it is important that such analysis should take full account of the limitations that arise from the finite time resolution of patch-clamp recording techniques. Single-ion-channel data are generally interpreted in terms of Markov process models of channel gating mechanisms. Experimental channel records suffer from time interval omission, i.e. failure to detect brief channel openings and closings. This leads to an identifiability problem when analysing single-channel data, i.e. different gating mechanisms provide equally convincing descriptions of the same experimental data. We consider a two-state Markov model of receptor-channel gating in which the channel opening rate is proportional to the agonist concentration, C in equilibrium with OA. By using computer-simulated data, the approximate likelihood of the data is maximized to yield parameter estimates for the model. At a single agonist concentration there is an identifiability problem in that two pairs of parameter estimates are obtained. The 'true' parameter estimates cannot be distinguished from the 'false' ones. By considering data corresponding to a range of agonist concentrations one may identify the 'true' parameter estimates as those that do not change as the agonist concentration is increased. Alternatively, one may identify the 'true' parameter estimates directly by maximizing a global likelihood, the latter being obtained by simultaneous consideration of data obtained at several different agonist concentrations.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Estimating equations with nonignorably missing response data.

Troxel, Lipsitz, and Brennan (1997, Biometrics 53, 857-869) considered parameter estimation from survey data with nonignorable nonresponse and proposed weighted estimating equations to remove the biases in the complete-case analysis that ignores missing observations. This paper suggests two alternative modifications for unbiased estimation of regression parameters when a binary outcome is potentially observed at successive time points. The weighting approach of Robins, Rotnitzky, and Zhao (1995, Journal of the American Statistical Association 90, 106-121) is also modified to obtain unbiased estimating functions. The suggested estimating functions are unbiased only when the missingness probability is correctly specified, and misspecification of the missingness model will result in biases in the estimates. Simulation studies are carried out to assess the performance of different methods when the covariate is binary or normal. For the simulation models used, the relative efficiency of the two new methods to the weighting methods is about 3.0 for the slope parameter and about 2.0 for the intercept parameter when the covariate is continuous and the missingness probability is correctly specified. All methods produce substantial biases in the estimates when the missingness model is misspecified or underspecified. Analysis of data from a medical survey illustrates the use and possible differences of these estimating functions.

Biometry↗

Multivariate outlier detection applied to multiply imputed laboratory data.

In clinical laboratory safety data, multivariate outlier detection methods may highlight a patient whose laboratory measurements do not follow the same pattern of relationships as the majority of patients, although their individual measurements are not found to be outlying when considered one at a time. Missing data problems are often dealt with by imputing a single value as an estimate of the missing value. The completed data set may then be analysed using traditional methods. A disadvantage of using single imputation is the underestimation of variability, with a corresponding distortion of power in hypothesis testing. Multiple imputation methods attempt to overcome this problem, and in this paper a study is described which considers the application of multivariate outlier detection methods to multiply imputed clinical laboratory safety data sets. Three different proportions of missing data are generated in laboratory data sets of dimensions 4, 7, 12 and 30, and a comparison of eight multiple imputation methods is carried out. Two outlier detection techniques, Mahalanobis distance and generalized principal component analysis, are applied to the multiply imputed data sets, and their performances are discussed. Measures are introduced for assessing the accuracy of the missing data results, depending on which method of analysis is used.

Algorithms↗

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance↗