Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Simplifying the assessment of rural emergency medical services trauma transport.

OBJECTIVES: The authors determine whether assessments of effects of rural emergency medical service (EMS) system characteristics on trauma outcomes (using patient-level data) are significantly biased if the Injury Severity Score (ISS) is not available. METHODS: The data are from ambulance trip reports merged with the trauma registry data for the Georgia EMS region VI trauma center hospital, located in Augusta. All 294 trauma patients for the rural counties surrounding Richmond County for the calendar year 1991 who were not dead at the scene and were treated at the trauma center are included. A 20% random sample of trauma patients from Richmond county from May to September 1991 not dead at the scene and treated at the trauma center yielded an additional 96 cases. Excluding 43 patients with missing data yields 347 trauma cases with 18 trauma deaths. A logistic regression model for trauma mortality is estimated using the Revised Trauma Score (RTS), ISS, type of trauma, and patient age (analogous to the standard Trauma Related Injury Severity Score model). The predicted probability of patient mortality from this model is compared with the predicted probability of mortality when the logistic regression model omits ISS. Correlations between the difference in predicted probability (i.e., the error in predicted probability associated with the omitted ISS variable) and EMS system characteristics are determined. RESULTS: Although ISS adds to the predictive power of the trauma outcome model, the errors in predicted probabilities associated with the omission of ISS generally are small and uncorrelated with patient or EMS system characteristics, with the exception of patient gender. CONCLUSIONS: In rural settings, where a patient's ISS generally is not available, studies of rural EMS system characteristics and trauma outcomes may use RTS, patient age, and type of trauma to control for expected survival. The patient's ISS does not appear to be essential, at least for the rural area analyzed in this study.

Emergency Medical Services↗

Addressing an idiosyncrasy in estimating survival curves using double sampling in the presence of self-selected right censoring.

We investigate the use of follow-up samples of individuals to estimate survival curves from studies that are subject to right censoring from two sources: (i) early termination of the study, namely, administrative censoring, or (ii) censoring due to lost data prior to administrative censoring, so-called dropout. We assume that, for the full cohort of individuals, administrative censoring times are independent of the subjects' inherent characteristics, including survival time. To address the loss to censoring due to dropout, which we allow to be possibly selective, we consider an intensive second phase of the study where a representative sample of the originally lost subjects is subsequently followed and their data recorded. As with double-sampling designs in survey methodology, the objective is to provide data on a representative subset of the dropouts. Despite assumed full response from the follow-up sample, we show that, in general in our setting, administrative censoring times are not independent of survival times within the two subgroups, nondropouts and sampled dropouts. As a result, the stratified Kaplan-Meier estimator is not appropriate for the cohort survival curve. Moreover, using the concept of potential outcomes, as opposed to observed outcomes, and thereby explicitly formulating the problem as a missing data problem, reveals and addresses these complications. We present an estimation method based on the likelihood of an easily observed subset of the data and study its properties analytically for large samples. We evaluate our method in a realistic situation by simulating data that match published margins on survival and dropout from an actual hip-replacement study. Limitations and extensions of our design and analytic method are discussed.

Arthroplasty, Replacement, Hip↗

Affiliation with deviant peers among children of substance dependent fathers from pre-adolescence into adolescence: associations with problem behaviors.

OBJECTIVE: Affiliation with delinquent peers has been shown to be a major risk factor for the development of antisocial and substance abuse behaviors in adolescence. However, little data are available concerning the developmental trajectories of deviant peer affiliation. METHOD: In this study, we have prospectively examined the density of deviant peers among the social networks of children of drug dependent fathers at age 10, and at 2 and 5 year follow-ups, and compared them with those of controls. Measures of internalizing and externalizing psychopathology were employed as time varying covariates, while socioeconomic status (SES) was used as a time invariant covariate. A pattern mixture analysis of missing data was conducted. RESULTS: Using mixed effects models, we found significant main effects of time, group, externalizing psychopathology, and to lesser extent, SES on the magnitude of affiliation with deviant peers. Greater deviant peer affiliation among the high-risk children was found at each time point. Externalizing psychopathology augmented the magnitude of deviant peer affiliation in both high-risk and comparison children. CONCLUSION: Offspring of drug dependent fathers have heightened affiliation with deviant peers from pre-adolescence through mid-adolescence. This social developmental process may be a component of the familial risk for substance abuse and antisocial behaviors.

Adolescent↗

Is the Addiction Severity Index a reliable and valid assessment instrument among clients with severe and persistent mental illness and substance abuse disorders?

OBJECTIVE: This study examined aspects of reliability, validity and utility of Addiction Severity Index (ASI) data as administered to clients with severe and persistent mental illness (SMI) and concurrent substance abuse disorders enrolled in a publicly-funded community mental health center. METHODS: A total of 62 clients with SMI volunteered to participate in an interobserver and test-retest reliability study of the ASI. Spearman-Brown and Pearson correlation coefficients were calculated to examine the extent of agreement among client responses. RESULTS: Overall 16% of the composite scores could not be calculated due to missing data and 31% of the clients misunderstood or confused items in at least one of the seven ASI domains. As a whole, the interobserver reliability of the ASI composite scores for those subjects where sufficient data were available was satisfactory. However, there was more variance in the stability of client responses, with four composite scores producing test-retest reliability coefficients below .65. CONCLUSION: Evidence from this study suggests that the ASI has a number of limitations in assessing the problems of clients with severe and persistent mental illness, and it is likely that other similar instruments based on the self-reports of persons with severe and persistent mental illness would also encounter these limitations.

Adolescent↗

Assessing consistency of responses to questions on cocaine use.

This study examines consistency of self-reported responses to items within the questionnaire of a multi-site, prospective study of drug abuse treatment in the United States (DATOS). The analyses use data from 2842 interviewer-administered intake interviews. Questions that were logically related are paired and responses compared. The questions cover three topics: (1) age at which different types of cocaine was used, (2) reports on most recent use and (3) frequency of cocaine use during period of "heaviest" use. Responses are coded as consistent, inconsistent, or as survey administration error. The latter is related to interviewer errors such as erroneous skip pattern, out-of-range responses, "don't know" responses, missing data, or illegible responses. Contrary to expectations inconsistent responses were relatively rare in this study, with fewer than 5% (0.5-4.6%) of respondents reporting inconsistent answers for pairs of logically related questions. A careful review of responses also found few survey administration errors (0.2-1.3%).

Adult↗

Simplifying the assessment of rural emergency medical service trauma transport.

OBJECTIVES: The authors determine whether assessments of effects of rural emergency medical services (EMS) system characteristics on trauma outcomes using patient-level data are biased significantly if the Injury Severity Score (ISS) is not available. METHODS: Data are taken from ambulance trip reports merged with the trauma registry data for the Georgia EMS region VI trauma center hospital, located in Augusta. All 294 trauma patients for the rural counties surrounding Richmond County for the calendar year 1991 who were not dead at the scene and who were treated at the trauma center are included. A 20% random sample of trauma patients from Richmond county from May 1991 to September 1991 not dead at the scene and treated at the trauma center yielded an additional 96 cases. Excluding 43 patients with missing data yields 347 trauma cases with 18 trauma deaths. A logistic regression model for trauma mortality is estimated using the Revised Trauma Score, ISS, type of trauma, and patient age (analogous to the standard Trauma Related Injury Severity Score model). The predicted probability of patient mortality from this model is compared with the predicted probability of mortality when the logistic regression model omits ISS. Correlations between the difference in predicted probability (ie, the error in predicted probability associated with the omitted ISS variable) and EMS system characteristics are determined. RESULTS: Although ISS adds to the predictive power of the trauma outcome model, the errors in predicted probabilities associated with the omission of ISS generally are small and uncorrelated with patient or EMS system characteristics, with the exception of patient gender. CONCLUSIONS: In rural settings, where a patient's ISS generally is not available, studies of rural EMS system characteristics and trauma outcomes may use Revised Trauma Score, patient age, and type of trauma to control for expected survival. The patient's ISS does not appear to be essential, at least for the rural area analyzed in this study.

Emergency Medical Services↗

A new clustering method for microarray data analysis.

A novel clustering approach is introduced to overcome data missing and inconsistency of gene expression levels under different conditions in the stage of clustering. It is based on the so-called smooth score, which is defined for measuring the deviation of the expression level of a gene and the average expression level of all the genes involved under a condition. We present an efficient greedy algorithm for finding clusters with smooth score below a threshold after studying its computational complexity. The algorithm was tested intensively on random matrixes and a yeast data. It was shown to perform well in finding co-regulation patterns in a test with the yeast data.

Algorithms↗

The urge to merge: linking vital statistics records and Medicaid claims.

This paper describes a procedure used to link Medicaid claims data to California vital statistics records for very low birthweight infants. The linkage involved about 53,000 infants born from 1980 to 1987 and 1.46 million claims for delivery/birth-related hospital admissions during the same period. Because the two data files did not share a unique identifier, record linkage required combining evidence across several linking variables: delivery hospital, delivery/birth date or hospitalization period, names, mother's age, and zip code. To combine the various pieces of evidence, we used record linkage theory to compute scores that measure the likelihood of a match, i.e., that two records correspond to the same delivery. These scores appropriately weight the various pieces of evidence for or against a match. Implementation required dealing with large amounts of missing data in one of the files, errors and variations in reported names, and the need to minimize the number of incorrect links. The approach applies to a wide range of linkage problems. The ability to combine existing datasets to form new datasets containing analysis variables from each facilitates analyses that would otherwise be impossible, or prohibitively expensive.

Bias↗

Algorithms for association study design using a generalized model of haplotype conservation.

There is considerable interest in computational methods to assist in the use of genetic polymorphism data for locating disease-related genes. Haplotypes, contiguous sets of correlated variants, may provide a means of reducing the difficulty of the data analysis problems involved. The field to date has been dominated by methods based on the "haplotype block" hypothesis, which assumes discrete population-wide boundaries between conserved genetic segments, but there is strong reason to believe that haplotype blocks do not fully capture true haplotype conservation patterns. In this paper, we address the computational challenges of using a more flexible, block-free representation of haplotype structure called the "haplotype motif" model for downstream analysis problems. We develop algorithms for htSNP selection and missing data inference using this more generalized model of sequence conservation. Application to a dataset from the literature demonstrates the practical value of these block-free methods.

Algorithms↗

Analyzing laboratory marker changes in AIDS clinical trials.

Repeated measurements of laboratory markers of immunologic or disease status, such as CD4 lymphocyte counts and HIV p24 antigen levels, can be important end points in comparative clinical trials. In this report, we consider comparison of treatment groups with respect to such markers, focusing on a distribution-free approach in which each participant's data are characterized by a single summary statistic. The summary statistics examined are (a) the slope of the least-squares regression of the marker, (b) the average of the last r measurements, and (c) the difference between the averages of the last r and the first s measurements. Under various models of marker time trends, these methods are compared with regard to statistical power. It is found that the slope is usually more efficient than the other two types of summaries. Adaptations for missing data are discussed and illustrated in an analysis of CD4 counts from a recent AIDS clinical trial.

Acquired Immunodeficiency Syndrome↗

Analysis of longitudinal data in an Alzheimer's disease clinical trial.

Evidence of delayed progression is the primary mechanism for demonstrating therapeutic efficacy in clinical trials in Alzheimer's disease. In the major trials of therapeutic treatment of AD, to date, measures based on clinical judgement and cognitive performance, instead of mortality, have been used as the primary response measures. There is good reason for this since the course of the disease is quite long, and AD trials designed around mortality would require either very large sample sizes or very long follow-up in order to have adequate power. However, the evaluation of progression in AD using clinical markers is subject to a number of challenges often found in longitudinal databases, for example, missing data, floor and ceiling effects and non-linearity. Unfortunately, few of these issues are being addressed in the typical analysis of progression data. This paper explores these analytic issues in the context of the recently completed Alzheimer's Disease Cooperative Study trial of vitamin E and Selegeline in moderate AD patients.

Aged↗

Advancing paternal age and autism.

CONTEXT: Maternal and paternal ages are associated with neurodevelopmental disorders. OBJECTIVE: To examine the relationship between advancing paternal age at birth of offspring and their risk of autism spectrum disorder (ASD). DESIGN: Historical population-based cohort study. SETTING: Identification of ASD cases from the Israeli draft board medical registry. PARTICIPANTS: We conducted a study of Jewish persons born in Israel during 6 consecutive years. Virtually all men and about three quarters of women in this cohort underwent draft board assessment at age 17 years. Paternal age at birth was obtained for most of the cohort; maternal age was obtained for a smaller subset. We used the smaller subset (n = 132 271) with data on both paternal and maternal age for the primary analysis and the larger subset (n = 318 506) with data on paternal but not maternal age for sensitivity analyses. MAIN OUTCOME MEASURES: Information on persons coded as having International Classification of Diseases, 10th Revision ASD was obtained from the registry. The registry identified 110 cases of ASD (incidence, 8.3 cases per 10 000 persons), mainly autism, in the smaller subset with complete parental age data. RESULTS: There was a significant monotonic association between advancing paternal age and risk of ASD. Offspring of men 40 years or older were 5.75 times (95% confidence interval, 2.65-12.46; P<.001) more likely to have ASD compared with offspring of men younger than 30 years, after controlling for year of birth, socioeconomic status, and maternal age. Advancing maternal age showed no association with ASD after adjusting for paternal age. Sensitivity analyses indicated that these findings were not the result of bias due to missing data on maternal age. CONCLUSIONS: Advanced paternal age was associated with increased risk of ASD. Possible biological mechanisms include de novo mutations associated with advancing age or alterations in genetic imprinting.

Adult↗

A probabilistic treatment of phylogeny and sequence alignment.

Carrying out simultaneous tree-building and alignment of sequence data is a difficult computational task, and the methods currently available are either limited to a few sequences or restricted to highly simplified models of alignment and phylogeny. A method is given here for overcoming these limitations by Bayesian sampling of trees and alignments simultaneously. The method uses a standard substitution matrix model for residues together with a hidden Markov model structure that allows affine gap penalties. It escapes the heavy computational burdens of other models by using an approximation called the "*" rule, which replaces missing data by a sum over all possible values of variables. The behavior of the model is demonstrated on test sets of globins.

Bayes Theorem↗

A computerized method for accurately predicting fetal macrosomia up to 11 weeks before delivery.

OBJECTIVE: To improve the prediction of birth weight and fetal macrosomia by combining sonographically derived fetal biometric data with routinely recorded pregnancy-specific information. STUDY DESIGN: Retrospective data were obtained for 218 normal gravidas who had obstetrical ultrasonography performed within 11 weeks of delivery. Multiple regression was employed to derive a set of equations for predicting birth weight that used different combinations of ultrasonographic and pregnancy-specific variables. RESULTS: A set of 38 unique combination equations was derived to accurately predict birth weight up to 11 weeks before delivery. The equations use different combinations of ultrasonographic and pregnancy-specific variables, so that predictions are still possible in the face of missing data. When ultrasonographic measurements are taken within 3 weeks of delivery, fetal macrosomia is predicted with 75% sensitivity, 93% specificity, and 67% and 95% positive and negative predictive value, respectively. The equations are equally as accurate for primiparous and multiparous women from all racial groups. A jackknifing procedure was used to validate the predictive accuracy of the equations for use with new subjects. CONCLUSION: The combined approach of predicting fetal macrosomia using ultrasonographic fetal measurements and pregnancy-specific characteristics is superior to pre-existing approaches that rely on either method alone. The method can be used up to 11 weeks before delivery, allowing fetal macrosomia to be predicted reliably in low-risk populations sufficiently early for prospective clinical intervention to be undertaken.

Female↗

The Short-Form Headache Impact Test (HIT-6) was psychometrically equivalent in nine languages.

BACKGROUND AND OBJECTIVE: This study examined the psychometric properties and equivalence of the six-item Headache Impact Test (HIT-6) across 11 languages in 14 countries. METHODS: A multicenter, international cross-sectional study conducted in a primary care setting. Data obtained from 1,171 adults from 14 countries who consulted their primary care physician for headache completed the HIT-6 questionnaire and a headache survey were included in this analysis. Item-level statistics (e.g., range of response choices used by participants), item-scale statistics (e.g., item-total correlations), scale level statistics (e.g., internal consistency reliability), and tests of differential item functioning were conducted to examine the psychometric properties of all HIT-6 translations and their comparability across translations. RESULTS: Across languages, missing data were low, item-scale correlations were high, reliability was adequate, and item-level statistics were generally comparable. We found only minor differential item functioning, suggesting that the HIT-6 translations are equivalent to the U.S. English form. CONCLUSIONS: Psychometric analyses indicate that most HIT-6 translations (Canadian English, French, Greek, Hungarian, UK English, Hebrew, Portuguese, German, Spanish, and Dutch) are comparable to U.S. English. Improvements may be needed in the Finnish and Slovakian translations and the appropriateness of using the HIT-6 in South Africa should be explored further.

Cross-Sectional Studies↗

Estimating cumulative probabilities from incomplete longitudinal binary responses with application to HIV vaccine trials.

When describing longitudinal binary response data, it may be desirable to estimate the cumulative probability of at least one positive response by some time point. For example, in phase I and II human immunodeficiency virus (HIV) vaccine trials, investigators are often interested in the probability of at least one vaccine-induced CD8+ cytotoxic T-lymphocyte (CTL) response to HIV proteins at different times over the course of the trial. In this setting, traditional estimates of the cumulative probabilities have been based on observed proportions. We show that if the missing data mechanism is ignorable, the traditional estimator of the cumulative success probabilities is biased and tends to underestimate a candidate vaccine's ability to induce CTL responses. As an alternative, we propose applying standard optimization techniques to obtain maximum likelihood estimates of the response profiles and, in turn, the cumulative probabilities of interest. Comparisons of the empirical and maximum likelihood estimates are investigated using data from simulations and HIV vaccine trials. We conclude that maximum likelihood offers a more accurate method of estimation, which is especially important in the HIV vaccine setting as cumulative CTL responses will likely be used as a key criterion for large scale efficacy trial qualification.

AIDS Vaccines↗

A review of selected patient-generated outcome measures and their application in clinical trials.

BACKGROUND: Patient-generated outcome measures have been developed in an effort to capture the individualistic nature of health-related quality of life (HRQoL). These measures differ from traditional HRQoL instruments in that they allow patients to individually define HRQoL domains or weights. Nevertheless, application of these measures may be challenging, particularly in a clinical trial setting. OBJECTIVE: The objective of this study was to provide a critical review of the following patient-generated outcome measures: Patient-Generated Index (PGI), Schedule for the Evaluation of Individual Quality of Life (SEIQoL), Repertory Grid, and Asthma Quality of Life Questionnaire (AQLQ). METHODS: We conducted a systematic literature review of the Medline and Olga databases and the journal Quality of Life Research and consulted experts. We abstracted data from eligible studies on the instruments' content, psychometric properties, and applicability in a clinical trial setting. RESULTS: The SEIQoL has shown to be reliable, valid, and responsive, but the PGI has not. Both instruments have poor practicality and have not been used in a clinical trial. The Repertory Grid's psychometric properties have not been well studied. The AQLQ is in part patient-generated and has been used in clinical trials. Nevertheless, the practicality of the individualized section is poor owing to the issue of missing data. All four instruments fail to provide a form of standardization needed for estimating population effects in a clinical trial. CONCLUSION: The applicability of patient-generated outcome measures in a clinical trial setting remains questionable. Patient-generated outcome measures appear to be useful primarily in complementing traditional HRQoL measures, guiding individual patient treatment decisions, and assisting the design of new measures.

Clinical Trials as Topic↗

An investigation of nonresponse to self-assessment of health by older persons. Associations with mortality.

This study examined the association between mortality and nonresponse to questions about health status (both refusals and "don't know" responses) using a national sample of persons aged 70 and over. Data were drawn from the 1984-1990 Longitudinal Study of Aging. Three time points of vital status were used as the outcome indicators (1984-1986, 1984-1988, 1984-1990). Five self-assessment questions were examined; three of the five questions had bivariate odds ratios that indicated significant associations between a nonresponse and all three mortality indexes. Results of the study suggest that nonresponses by older persons can convey meaningful information. Research on self-assessments of health in later life should not routinely exclude nonresponses as missing data, even if they are an infrequent response.

Aged↗