Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Inferences concerning exponential distributions in the presence of randomly right censored data with missing censored values.

Inferences concerning exponential distributions are considered from a sampling theory viewpoint when the data are randomly right censored and the censored values are missing. Both one-sample and m-sample (m > or = 2) problems are considered. Likelihood functions are obtained for situations in which the censoring mechanism is informative which leads to natural and intuitively appealing estimators of the unknown proportions of censored observations. For testing hypotheses about the unknown parameters, three well-known test statistics, namely, likelihood ratio test, score test, and Wald-type test are considered.

Data Interpretation, Statistical↗

[Methods for handling incomplete data in health research: a critical look].

OBJECTIVE: To illustrate methods for handling incomplete data in health research. METHODS: Two strategies for handling missing data are presented: complete-case analysis and imputations. The imputations used were mean imputations, regression imputations, and multiple imputations. These strategies are illustrated in the context of logistic regression through an example using data from the "Second Cuban national survey on risk factors and non communicable disease", carried out in 2001. RESULTS: The results obtained via mean and regression imputation were similar. The odds ratios were overestimated by 10%. The results of complete-case analysis showed the greatest difference from the reference odds ratios, with a variation of between 2 and 65%. The three methods distorted the relationship between age and hypertension. Multiple imputations produced estimates closest to those of the reference estimates with a variation of less than 16%. This was the only procedure preserving the relationship between age and hypertension. CONCLUSIONS: Selecting methods for handling missing data is difficult, since the same procedure can give precise estimations in certain circumstances and not in others. Complete-case analysis should be used with caution due to the substantial loss of information it produces. Mean and regression imputations produce unreliable estimates under missing at random (MAR) mechanisms.

Cuba↗

Computerized information-gathering in specialist rheumatology clinics: an initial evaluation of an electronic version of the Short Form 36.

OBJECTIVES: Longitudinal outcome data are important for research and are becoming part of routine clinical practice. We assessed an initial version of an electronic Short Form 36 (SF-36), a well-established health assessment questionnaire, in comparison with standard paper forms, in two specialist rheumatology clinics. METHODS: Out-patients (20 with systemic lupus erythematosus and 31 with vasculitis) were randomly selected to complete either paper (n=29) or electronic and paper SF-36 versions (n=51) before and after consultation (paper vs paper comparison). Data were evaluated as the response correlation, internal consistency, missing data, patient satisfaction and preference. RESULTS: There were very good correlations in SF-36 responses (P<0.001) between the paper and electronic forms and the paper and paper forms. Internal reliability coefficients (Cronbach's alpha) showed good internal consistency for all reported responses in either computer or paper forms. There were no missing data in the computerized version but 24% of patients failed to answer all of the paper form questions. Ease of use of the computer version was rated highly by 71% of all the respondents, and 69% would prefer to use the computer version in future. DISCUSSION: Computerized data collection is acceptable to patients and feasible in clinical settings. It provides responses that are at least comparable to those to the paper form, improves data capture and is available immediately.

Ambulatory Care↗

Eliminating missing race/ethnicity data from a sexually transmitted disease case registry.

Data regarding race and ethnicity are usually requested when conducting public health surveillance. However, such data are frequently not included in case reports by providers. This report describes efforts to reduce the extent of missing race and/or ethnicity data in reports of sexually transmitted diseases in Massachusetts. A list of cases reported to the Department of Public Health between March 1 and May 31 1999 lacking race and/or ethnicity data was generated. A student intern tried contacting the providers with a request for complete information. Of the 2,954 cases of syphilis, gonorrhea, and chlamydia infection reported during the study period, 34.8% (1,028 cases) lacked race/ethnicity data. Despite an average of 2.27 calls and 1.5 transfers per call, data was successfully added to only 143 cases, increasing the percent of reported cases with complete data from 65.2% to 70.0%. The telephone calls, while inefficient for collecting this data, had some advantages. For example, they offered opportunities for communication between the STD Division and providers regarding other provider needs or services that the Division might meet. Consideration can also be given to using surnmame lists, ethnic marketing lists, birth records, and matching the case's address with census block data to infer race and/or ethnicity.

Data Collection↗

Maximum likelihood estimation in random effects cure rate models with nonignorable missing covariates.

We introduce a method of parameter estimation for a random effects cure rate model. We also propose a methodology that allows us to account for nonignorable missing covariates in this class of models. The proposed method corrects for possible bias introduced by complete case analysis when missing data are not missing completely at random and is motivated by data from a pair of melanoma studies conducted by the Eastern Cooperative Oncology Group in which clustering by cohort or time of study entry was suspected. In addition, these models allow estimation of cure rates, which is desirable when we do not wish to assume that all subjects remain at risk of death or relapse from disease after sufficient follow-up. We develop an EM algorithm for the model and provide an efficient Gibbs sampling scheme for carrying out the E-step of the algorithm.

Journal Article↗

A comparison of measures of socioeconomic status for adolescents in a Canadian national health survey.

The purpose of this study was to explore and compare measures of socioeconomic status (SES) in a national sample of Canadian adolescents. Issues of missing data and interrelationships among the measures were addressed. Measures of SES included household income, parental education, two parental occupation-based measures, and four neighbourhood proxy indicators. The proportion of adolescents with missing data was largest for household income (21.1 percent). Data were not missing at random, as adolescents missing household income information were less likely to reside in a high income neighbourhood. Pair-wise Spearman correlations ranged from: 0.40-0.79 between neighbourhood SES measures; 0.12-0.37 between household/parental and neighbourhood indicators; and 0.36-0.87 between household/parental measures. Correlations were lower among rural adolescents, particularly for the neighbourhood SES measures. The results highlight both measurement and conceptual challenges for researchers who wish to gain insight into SES-health relationships for adolescents. In particular, the findings emphasize the importance of incorporating multiple measures of SES and suggest a need to further explore the meaning of socioeconomic position for this population.

Adolescent↗

Estimating single-channel kinetic parameters from idealized patch-clamp data containing missed events.

We present here a maximal likelihood algorithm for estimating single-channel kinetic parameters from idealized patch-clamp data. The algorithm takes into account missed events caused by limited time resolution of the recording system. Assuming a fixed dead time, we derive an explicit expression for the corrected transition rate matrix by generalizing the theory of Roux and Sauve (1985, Biophys. J. 48:149-158) to the case of multiple conductance levels. We use a variable metric optimizer with analytical derivatives for rapidly maximizing the likelihood. The algorithm is applicable to data containing substates and multiple identical or nonidentical channels. It allows multiple data sets obtained under different experimental conditions, e.g., concentration, voltage, and force, to be fit simultaneously. It also permits a variety of constraints on rate constants and provides standard errors for all estimates of model parameters. The algorithm has been tested extensively on a variety of kinetic models with both simulated and experimental data. It is very efficient and robust; rate constants for a multistate model can often be extracted in a processing time of approximately 1 min, largely independent of the starting values.

Algorithms↗

Impact of missing genotype data on Monte-Carlo simulation based haplotype analysis.

In the context of haplotype association analysis of unphased genotype data, methods based on Monte-Carlo simulations are often used to compensate for missing or inappropriate asymptotic theory. Moreover, such methods are an indispensable means to deal with multiple testing problems. We want to call attention to a potential trap in this usually useful approach: The simulation approach may lead to strongly inflated type I errors in the presence of different missing rates between cases and controls, depending on the chosen test statistic. Here, we consider four different testing strategies for haplotype analysis of case-control data. We recommend to interpret results for data sets with non-comparable distributions of missing genotypes with special caution, in case the test statistic is based on inferred haplotypes per individual. Moreover, our results are important for the conduction and interpretation of genome-wide association studies.

Case-Control Studies↗

Genetic association tests for family data with missing parental genotypes: a comparison.

We consider three tests for genetic association in data from nuclear families (the Family-Based Association Test (FBAT) test proposed by Rabinowitz and Laird ([2000] Hum. Hered. 50:211-223), a second test proposed by Rabinowitz ([2002] J. Am. Stat. Assoc. 97:742-758), and the Family Genotype Analysis Program (FGAP) nonfounder or partial score test proposed by Clayton ([1999] Am. J. Hum. Genet. 65:1170-1177) and Whittemore and Tu ([2000] Am. J. Hum. Genet. 66:1329-1340)). We show that each test statistic arises from the efficient score of the family data as the solution to a set of constraints on its null expectation. Moreover, the FBAT and Rabinowitz tests (but not the FGAP test) are locally the most powerful among all tests satisfying their constraints. We used simulations to examine how the three tests perform in situations when their assumptions are violated and the number of families is not huge. We found that the FBAT test tended to have less power than the other two tests, particularly when applied to families in whom all offspring were affected. The Rabinowitz and FGAP tests performed similarly, although the latter tended to extract more information from families containing one typed parent. While none of the tests showed good power to detect rare, recessively acting genes, the Rabinowitz test with a sample variance estimate performed particularly poorly in this case. However, the Rabinowitz test with a model-based variance had power comparable to that of the FGAP test, and more accurate type I error rates. We conclude that for the situations we considered, the Rabinowitz test with model-based variance has good power without forfeiting robustness against misspecification of parental genotype probabilities. However, its utility is limited by the lack of a simple algorithm to apply it to families with varying structures and phenotypes.

Family↗

Population one-compartment pharmacokinetic analysis with missing dosage data.

OBJECTIVE: Our objective was to develop a population 1-compartment pharmacokinetic (PK) method of analysis to deal with suspect or missing prior dosage history. METHODS: Population PK data from a 1-compartment model with first-order elimination and absorption, described by PK parameters clearance, volume of distribution, and absorption rate constant, are simulated. A PK sample is drawn just before a test dose (Dt), followed by a (varying) number of additional samples over 1 interdose interval (tau). For 60% of the subjects, the true history of the scheduled dose (Ds) preceding Dt differs from that prescribed, whereas doses taken before Ds do not. Two settings are evaluated: considerable accumulation of drug in the body (typical drug half-life t1/2 approximately equal to tau) and very little such accumulation (t1/2 approximately equal to tau/5). Precision and bias of several PK analysis methods--Missing Dose Method (MDM), Missing Dose Mixture Method (MDMM) and Extrapolation-Subtraction Method (ESM), all of which essentially do not use prior dose history--are compared with those of the Prescribed Dose Method (PDM), which assumes nominal dosage, and an Ideal Method (IDM), which uses true (but unknown) pre-test dose history. RESULTS: At t1/2 approximately equal to tau, MDM and MDMM are the most precise methods. The accuracy of ESM and PDM is poor. At t1/2 approximately equal to tau/5, no significant differences, in terms of precision or bias, are observed between methods. Misspecification of the structural or statistical model seems not to influence these results. The results of analysis of a real (caffeine) data set are compatible with the findings from the simulations. CONCLUSION: When a test dose is given and a predose baseline observation is taken as part of an "intensive" PK study during outpatient therapy of a 1-compartment drug, an analysis that assumes that the nominal dose history is correct is not robust to past dosage history misspecification, whereas methods that do not do this are robust and reliable.

Absorption↗

Analysis of proportions from clustered data with missing observations in a matched-pair design.

OBJECTIVES: Typically, methods for the estimation of differences in proportions from clustered data are based on complete cases with no missing information. In this paper we propose an extension to the method of Rao and Scott and Obuchowski to allow for the explicit computation of the variance of the estimator for the difference in presence of incomplete cases. METHODS: We divided the full analysis set into a set of complete cases and a set of incomplete cases. The differences in proportions of correct diagnoses were estimated for each set by taking into consideration the clustering effect for both sets and the correlation between the procedures in the set with complete cases. Then the estimates of the two parts were combined by appropriate weights, which then allowed the explicit calculation of the variance. The performance of the extension as compared to the original method and generalized estimation equations model (GEEs) was examined by simulations. RESULTS: The results of the examples suggest that the extended approach is superior to the complete-case method and is therefore appropriate when all data are to be used. In comparison to GEEs, the extended method appears to be slightly inferior, when the number of observations per patient is high, but of similar efficiency with a low number of observations per patient. CONCLUSIONS: With the extension of the method by Rao and Scott and Obuchowski we make use of all available data. Therefore, we follow the intent-to-treat principle as close as possible.

Germany↗

Smoking cessation studies: a methodological comparison.

A wide variety in outcome criteria hinders comparison of results between smoking cessation studies. Three important methodological issues are discussed: analysis of data of participants who drop out of therapy, treatment of missing data, and repeated use of significance tests. These issues determine to a great extent the results of evaluation studies. In general, they are of interest to all researchers of addiction who study the effects of interventions. Several ways to decide on these issues and the consequences of these decisions are considered. Little consensus exists about the criterion for dropout. It is concluded that a dropout criterion is a burden rather than a help. A better criterion would be the number of sessions present. Few satisfying techniques exist to handle the problem of missing data. Evaluation studies need to set a priori standards to counter the increased risk of a Type I error, caused by the repeated use of significance tests. Reviewers need to be aware of the variety in data treatment before comparing results.

Behavior Therapy↗

A simulation study of estimators for rates of change in longitudinal studies with attrition.

Many longitudinal studies and clinical trials are designed to compare rates of change over time in one or more outcome variables in several groups. Most such studies have incomplete data because some patients drop out before completing the study. The missing data may induce bias and inefficiency in naive estimates of important parameters. This paper uses Monte Carlo methods to compare the bias and efficiency of several two-stage estimators of the effect of treatment on the mean rate of change when the missing data arise from one of four processes. We also study the validity of confidence intervals and the power of hypothesis tests based on these estimates and their standard errors. In general, the weighted least squares estimator does relatively well, as does an analysis of covariance type estimator proposed by Wu et al. The best estimates of variance components are based on complete cases or maximum likelihood.

Analysis of Variance↗

Plasma exchange for Guillain-Barré syndrome.

BACKGROUND: Guillain-Barré syndrome is an acute symmetric usually ascending and usually paralysing illness due to inflammation of peripheral nerves. It is thought to be caused by autoimmune factors, such as antibodies. Plasma exchange removes antibodies and other potentially injurious factors from the blood stream. It involves connecting the patient's blood circulation to a machine which exchanges the plasma for a substitute solution, usually albumin. Several studies have evaluated plasma exchange for Guillain-Barré syndrome. OBJECTIVES: To systematically review the evidence concerning the efficacy of plasma exchange for treating Guillain-Barré syndrome. SEARCH STRATEGY: Search of the Cochrane Neuromuscular Disease Trial Register for randomised trials concerning plasma exchange in Guillain-Barré syndrome, search of the bibliographies of identified papers and enquiry from the authors of the papers. SELECTION CRITERIA: Randomised and quasi-randomised trials of plasma exchange versus sham exchange or supportive treatment. DATA COLLECTION AND ANALYSIS: Potentially relevant papers were scrutinised by two reviewers and the selection of eligible studies was agreed by them and a third reviewer. Data were extracted by one reviewer and checked by a second reviewer. Some missing data were obtained from the authors of studies. MAIN RESULTS: Six eligible trials concerning 649 patients were identified, all comparing plasma exchange versus supportive treatment alone. Primary outcome measures ~Bullet~Time to recover walking with aid In the only two trials for which this measure was reported the median time to recover this ability was faster in the plasma exchange than the control group. ~Bullet~Time to onset of motor recovery in mildly affected patients In the one trial for which this measure was available the time was significantly shortened in the plasma exchange group. Secondary outcome measures ~Bullet~Improvement in disability grade at 4 weeks In five trials, there were significantly more patients who had improved by one disability grade or more in the plasma exchange group as compared to the control group. Patients treated with plasma exchange fared significantly better in the following secondary outcome measures: time to recover walking without aid, percentage of patients requiring artificial ventilation, duration of ventilation, full muscle strength recovery after one year, and severe sequelae after one year. There were less patients with infectious events and cardiac arrhythmias in the plasma exchange than the control group. Subgroup analyses Plasma exchange was beneficial in patients with mild, moderate and severe (needing ventilation) Guillain-Barré syndrome. It was beneficial in patients with a disease duration of seven or less days and also in those with disease lasting more than seven days. However, in the only trial that enrolled patients up to 30 days from disease onset, the benefit of plasma exchange in patients treated after seven days was less apparent. Type of treatment Single studies showed that two plasma exchanges were significantly superior to none for mild Guillain-Barré syndrome and four to two for moderate Guillain-Barré syndrome but that six were not superior to four for severe Guillain-Barré syndrome requiring ventilation. One study suggested that continuous flow plasma exchange was significantly superior to intermittent flow. Another study found no significant difference between the two techniques. The same study found a significantly higher rate of adverse events with fresh frozen plasma as the replacement fluid than albumin. REVIEWER'S CONCLUSIONS: Plasma exchange is the first and only treatment that has been proven to be superior to supportive treatment alone in Guillain-Barré syndrome. Consequently, plasma exchange should be regarded as the treatment against which new treatments, such as intravenous immunoglobulin, should be judged. In mild Guillain-Barré syndrome two sessions of plasma exchange are superior to none. In moderate Guillain-Barré syndrome four sessions are superior to two. In severe Guillain-Barré syndrome six sessions are no better than four. Continuous flow plasma exchange machines may be superior to intermittent flow machines and albumin to fresh frozen plasma as the exchange fluid. Plasma exchange is more beneficial when started within seven days after disease onset rather than later, but was still beneficial in patients treated up to 30 days after disease onset. The value of plasma exchange in children less than 12 years old is not known.

Guillain-Barre Syndrome↗

The phylogeny of acorn weevils (genus Curculio) from mitochondrial and nuclear DNA sequences: the problem of incomplete data.

We considered the contribution of two mitochondrial and two nuclear data sets for the phylogenetic reconstruction of 22 species of seed beetles in the genus Curculio (Coleoptera: Cuculionidae). A phylogenetic tree from representatives found on various hosts was inferred from a combined data set of mitochondrial DNA cytochrome oxidase subunit I, mitochondrial cytochrome b, nuclear elongation factor 1alpha, and nuclear phosphoglycerate mutase, used for the first time as a molecular marker. Separate parsimony analyses of each data set showed that individual gene trees were mainly congruent and often complementary in the support of clades but the analysis was complicated by failure of PCR amplification of nuclear genes for many taxa and hence missing data entries. When the four gene partitions were combined in a simultaneous analysis despite the missing data, this increased the resolution and taxonomic coverage compared to the individual source trees. Alternative approaches of combining the information via supertree methodology produced a comparatively less resolved tree, and hence seem inferior to combining data matrices even in cases where numerous taxa are missing. The molecular data suggest a classification of the European species into two species groups that are in accordance with morphological characteristics but the data do no support any of the previously recognised American species groups.

Animals↗

Complete imputation of missing repeated categorical data: one-sample applications.

Longitudinal studies with repeated measures are often subject to non-response. Methods currently employed to alleviate the difficulties caused by missing data are typically unsatisfactory, especially when the cause of the missingness is related to the outcomes. We present an approach for incomplete categorical data in the repeated measures setting that allows missing data to depend on other observed outcomes for a study subject. The proposed methodology also allows a broader examination of study findings through interpretation of results in the framework of the set of all possible test statistics that might have been observed had no data been missing. The proposed approach consists of the following general steps. First, we generate all possible sets of missing values and form a set of possible complete data sets. We then weight each data set according to clearly defined assumptions and apply an appropriate statistical test procedure to each data set, combining the results to give an overall indication of significance. We make use of the EM algorithm and a Bayesian prior in this approach. While not restricted to the one-sample case, the proposed methodology is illustrated for one-sample data and compared to the common complete-case and available-case analysis methods.

Algorithms↗

Geographic bias related to geocoding in epidemiologic studies.

BACKGROUND: This article describes geographic bias in GIS analyses with unrepresentative data owing to missing geocodes, using as an example a spatial analysis of prostate cancer incidence among whites and African Americans in Virginia, 1990-1999. Statistical tests for clustering were performed and such clusters mapped. The patterns of missing census tract identifiers for the cases were examined by generalized linear regression models. RESULTS: The county of residency for all cases was known, and 26,338 (74%) of these cases were geocoded successfully to census tracts. Cluster maps showed patterns that appeared markedly different, depending upon whether one used all cases or those geocoded to the census tract. Multivariate regression analysis showed that, in the most rural counties (where the missing data were concentrated), the percent of a county's population over age 64 and with less than a high school education were both independently associated with a higher percent of missing geocodes. CONCLUSION: We found statistically significant pattern differences resulting from spatially non-random differences in geocoding completeness across Virginia. Appropriate interpretation of maps, therefore, requires an understanding of this phenomenon, which we call "cartographic confounding."

Journal Article↗

Performance of multi-layer feedforward neural networks to predict liver transplantation outcome.

A novel multisolutional clustering and quantization (MCQ) algorithm has been developed that provides a flexible way to preprocess data. It was tested whether it would impact the neural network's performance favorably and whether the employment of the proposed algorithm would enable neural networks to handle missing data. This was assessed by comparing the performance of neural networks using a well-documented data set to predict outcome following liver transplantation. This new approach to data preprocessing leads to a statistically significant improvement in network performance when compared to simple linear scaling. The obtained results also showed that coding missing data as zeroes in combination with the MCQ algorithm, leads to a significant improvement in neural network performance on a data set containing missing values in 59.4% of cases when compared to replacement of missing values with either series means or medians.

Algorithms↗