Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Multipoint linkage analysis with many multiallelic or dense diallelic markers: Markov chain-Monte Carlo provides practical approaches for genome scans on general pedigrees.

Computations for genome scans need to adapt to the increasing use of dense diallelic markers as well as of full-chromosome multipoint linkage analysis with either diallelic or multiallelic markers. Whereas suitable exact-computation tools are available for use with small pedigrees, equivalent exact computation for larger pedigrees remains infeasible. Markov chain-Monte Carlo (MCMC)-based methods currently provide the only computationally practical option. To date, no systematic comparison of the performance of MCMC-based programs is available, nor have these programs been systematically evaluated for use with dense diallelic markers. Using simulated data, we evaluate the performance of two MCMC-based linkage-analysis programs--lm_markers from the MORGAN package and SimWalk2--under a variety of analysis conditions. Pedigrees consisted of 14, 52, or 98 individuals in 3, 5, or 6 generations, respectively, with increasing amounts of missing data in larger pedigrees. One hundred replicates of markers and trait data were simulated on a 100-cM chromosome, with up to 10 multiallelic and up to 200 diallelic markers used simultaneously for computation of multipoint LOD scores. Exact computation was available for comparison in most situations, and comparison with a perfectly informative marker or interprogram comparison was available in the remaining situations. Our results confirm the accuracy of both programs in multipoint analysis with multiallelic markers on pedigrees of varied sizes and missing-data patterns, but there are some computational differences. In contrast, for large numbers of dense diallelic markers, only the lm_markers program was able to provide accurate results within a computationally practical time. Thus, programs in the MORGAN package are the first available to provide a computationally practical option for accurate linkage analyses in genome scans with both large numbers of diallelic markers and large pedigrees.

Alleles↗

Blood pressure and symptoms of depression and anxiety: a prospective study.

This study investigated whether symptoms of depression and anxiety were related to the development of elevated blood pressure in initially normotensive adults. The study's hypothesis was addressed with an existing set of prospective data gathered from an age-, sex-, and weight-stratified sample of 508 adults. Four years of follow-up data were analyzed both with logistic analysis, which used hypertension (blood pressure > or =140 mm Hg systolic or 90 mm Hg diastolic) as the dependent variable, and with multiple regression analysis, which used change in blood pressure as the dependent variable. Five physical risk factors for hypertension (age, sex, baseline body mass index, family history of hypertension, and baseline blood pressure levels) were controlled for in the regression analyses. Use of antidepressant/antianxiety and antihypertensive medications were controlled for in the study. Of the 433 normotensive participants who were eligible for our study, 15% had missing data in the logistic regression analysis focusing on depression (n = 371); similarly, 15% of the eligible sample had missing data in the logistic regression using anxiety as the psychological variable of interest (n = 370). Both logistic regression analyses showed no significant relationship for either depression or anxiety in the development of hypertension. The multiple regression analyses (n = 369 for the depression analysis; n = 361 for the anxiety analysis) similarly showed no relationship between either depression or anxiety in changes in blood pressure during the 4-year follow-up. Thus, our results do not support the role of depressive or anxiety symptoms in the development of hypertension in our sample of initially normotensive adults.

Adult↗

The SF-36 in multiple sclerosis: why basic assumptions must be tested.

OBJECTIVES: To evaluate, in people with multiple sclerosis, two psychometric assumptions that must be satisfied for valid use of the medical outcomes study 36-item short form health survey (SF-36): the data are of high quality and, it is legitimate to generate scores for eight scales and two summary measures using the standard algorithms. METHODS: SF-36 data from 438 people representing the full range of multiple sclerosis were examined (mean age 48; 70% women). Data quality (per cent missing data and computable scale and summary scores) were determined, six scaling criteria were tested to determine the legitimacy of generating the eight SF-36 scale scores using Likert's method of summed ratings, and two scaling criteria were tested to determine the appropriateness of the standard SF-36 algorithms for weighting scale scores to generate two summary measures. RESULTS: Data quality was excellent except in the most disabled subgroup where missing responses reached a maximum of 16.5% and summary scores could only be computed for 72%. There was clear support for the generation of SF-36 scale scores. Item response distributions were symmetric, item mean scores and variances were equivalent, corrected item-total correlations were high (range 0.46-0.85) and similar, and definite scaling success rates exceeded 96%. Nevertheless, there were notable floor or ceiling effects in four of the eight scales. Assumptions for generating two SF-36 summary measures were only partially satisfied. Although principal components analysis suggested a two component model, these components explained less than 60% of the total variance in SF-36 scales, and less than 75% of the variance in five of the eight scales. Moreover, scale to component correlations did not support the use of scale weights derived from United States population data. CONCLUSIONS: When using the SF-36 as a health measure in multiple sclerosis summary scores should be reported with caution.

Activities of Daily Living↗

HAPLOFREQ--estimating haplotype frequencies efficiently.

A commonly used tool in disease association studies is the search for discrepancies between the haplotype distribution in the case and control populations. In order to find this discrepancy, the haplotypes frequency in each of the populations is estimated from the genotypes. We present a new method HAPLOFREQ to estimate haplotype frequencies over a short genomic region given the genotypes or haplotypes with missing data or sequencing errors. Our approach incorporates a maximum likelihood model based on a simple random generative model which assumes that the genotypes are independently sampled from the population. We first show that if the phased haplotypes are given, possibly with missing data, we can estimate the frequency of the haplotypes in the population by finding the global optimum of the likelihood function in polynomial time. If the haplotypes are not phased, finding the maximum value of the likelihood function is NP-hard. In this case, we define an alternative likelihood function which can be thought of as a relaxed likelihood function. We show that the maximum relaxed likelihood can be found in polynomial time and that the optimal solution of the relaxed likelihood approaches asymptotically to the haplotype frequencies in the population. In contrast to previous approaches, our algorithms are guaranteed to converge in polynomial time to a global maximum of the different likelihood functions. We compared the performance of our algorithm to the widely used program PHASE, and we found that our estimates are at least 10% more accurate than PHASE and about ten times faster than PHASE. Our techniques involve new algorithms in convex optimization. These algorithms may be of independent interest. Particularly, they may be helpful in other maximum likelihood problems arising from survey sampling.

Algorithms↗

Improving estimates of exposures for epidemiologic studies of plutonium workers.

Epidemiologic studies of nuclear facilities usually focus on relations between cancer and doses from external penetrating radiation, and describe these exposures with little detail on measurement error and missing data. We demonstrate ways to document complex exposures to nuclear workers with data on external and internal exposures to ionizing radiation and toxic chemicals. We describe methods for assessing internal exposures to plutonium and external doses from neutrons; the use of a job exposure matrix for estimating chemical exposures; and methods for imputing missing data for exposures and doses. For plutonium workers at Rocky Flats, errors in estimating neutron doses resulted in underestimating the total external dose for production workers by about 16%. Estimates of systemic deposition do not correlate well with estimates of organ doses. Only a small percentage of workers had exposures to toxic chemicals, making epidemiologic assessments of risk difficult.

Colorado↗

Interpreting treatment differences when patients drop out of a clinical trial.

Clinical trials are the standard for identifying new drugs for the treatment of disease, but results are dependent on patient compliance. The success of treatments for HIV disease in particular may be judged in part by their effect on immunologic, virologic, or clinical measures collected on patients at regular predefined intervals. If patients drop out of a trial before study completion, the analysis of the repeatedly collected parameters needs to be undertaken and interpreted with care. The authors recommend using graphic techniques to assess the impact of the missing data on the profiles of the parameters over time. To assess treatment differences, a variety of simple tests are proposed that allow different assumptions to be made regarding the reasons for the incomplete data. A case study is presented providing an analysis of CD4 data from the Pediatric Aids Clinical Trials Group (PACTG) Protocol 051, in which only 52% of the patients completed the study while remaining on treatment; younger patients with lower CD4 counts were more likely to stop treatment earlier. This type of systematic missing data can lead to incorrect conclusions regarding different treatment effects on CD4 counts. With the data of PACTG 051, however, regardless of the methodology used, no treatment differences were found. Inconsistent conclusions would have indicated the need for more sophisticated statistical techniques to adequately test for treatment differences.

Age Factors↗

Comparison of alternative strategies for analysis of longitudinal trials with dropouts.

PROBLEM: Patients may withdraw from longitudinal clinical trials for many reasons. Methods for handling the problem presented by missing data of patients who withdraw before reaching the time point of the primary measurement include carrying the last observation forward (LOCF), data as observed analysis (DAO), mixed model approaches, and pattern mixture models. METHOD: We evaluate a multiple imputation (MI) approach that has the flexibility to adjust inferences about the treatment effect for the withdrawn patients relative to currently used alternatives. Sensitivity analyses are performed under a collection of scenarios that include many circumstances that may arise in practice, including different assumptions about treatment effects post-withdrawal and about the missing data mechanism. Simulations are used to compare the results of analyses based on the MI approach with those based on the LOCF, DAO, and the mixed model approaches. RESULTS: The LOCF and DAO approaches cannot be recommended as strategies for handling missing responses, at least for these scenarios, because they provide biased estimates of treatment effects and biased tests of the null hypothesis of no treatment effect. Application of the various approaches to the analysis of clinical data from a longitudinal trial confirms the underestimation of the variability when the LOCF approach is used.

Analysis of Variance↗

Quality of life in advanced non-small-cell lung cancer: results of a Southwest Oncology Group randomized trial.

PURPOSE: The main purpose of this paper is to present the results of a randomized trial comparing the effects of two chemotherapy regimens on the Quality of life (QOL) of patients with advanced non-small-cell lung cancer (NSCLC). Trials in advanced stage disease represent an important treatment context for QOL assessment. A second purpose of this paper is to examine methods for handling the level of missing data commonly observed in the advanced stage disease context. METHODS: Patients were randomized to receive cisplatin plus vinorelbine or carboplatin plus paclitaxel. The QOL of 222 patients was assessed with the Functional Assessment of Cancer Therapy-Lung (FACT-L) prior to randomization; follow-up assessments occurred at 13 and 25 weeks. Three methods were used to analyze the QOL data: (1) cross-sectional analysis of four patient categories (improved, stable, missing, and declined) based on changes in the FACT-L score, (2) a mixed linear model, and (3) a pattern mixture model. The longitudinal analyses addressed two potential data biases. RESULTS: Questionnaire submission rates were 91% at baseline, 68% at 13 weeks, and 47% at 25 weeks. The cross-sectional and mixed linear model analyses did not show significant differences by treatment arm in patient-reported QOL. The pattern mixture model analysis, more appropriate given non-ignorable missing data, also found no statistically significant effect of treatment on patient QOL. CONCLUSION: We present a sensitivity analysis approach with multiple methods for analyzing treatment effects on patient QOL in the presence of substantial, non-ignorable missing data in an advanced stage disease clinical trial. We conclude that the two treatment arms did not differ statistically in their effects on patient QOL over a 25-week treatment period.

Antineoplastic Combined Chemotherapy Protocols↗

Comparison of open and closed questionnaire formats in obtaining demographic information from Canadian general internists.

The objective of this study was to compare the impact of closed- versus open-ended question formats on the completeness and accuracy of demographic data collected in a mailed survey questionnaire. We surveyed general internists in five Canadian provinces to determine their career satisfaction. We randomized respondents to receive versions of the questionnaire in which 16 demographic questions were presented in a closed-ended or open-ended format. Two questions required respondents to make a relatively simple computation (ensuring that three or four categories of response added to 100%). The response rate was 1007/1192 physicians (80.0%). The proportion of respondents with no missing data for all 16 questions was 44.7% for open-ended and 67.0% for closed-ended formats (P < 0.001). The odds of having missing items remained higher for open-ended response options after adjusting for a number of respondent characteristics (2.67, 95% confidence interval 2.01 to 3.55). For the two questions requiring computations focused on professional activity and income, there were more missing data (P = 0.02, 0.02, respectively) but fewer inaccurate responses (P = 0.009, 0.20, respectively) for the open-ended compared to the closed-ended format. Investigators can achieve higher response rates for demographic items using closed format response options, but at the risk of increasing inaccuracy in response to questions requiring computation.

Adult↗

Two methods for recommending bat weights.

Baseball players swung very light and very heavy bats through our instrument and the speed of the bat was recorded. These data were used to make mathematical models for each person. Then these models were coupled with equations of physics for bat-ball collisions to compute the Ideal Bat Weight for each individual. However, these calculations required the use of a sophisticated instrument that is not conveniently available to most people. So, we tried to find items in our database that correlated with Ideal Bat Weight. However, because many cells in the database were empty, we could not use traditional statistical techniques or even neural networks. Therefore, three new methods were used to estimate the missing data: (i) a neural network was trained using subjects that had no empty cells, then that neural network was used to predict the missing data, (ii) the data patching facility of a commercial software package was used, and (iii) the empty cells were filled with random numbers. Then, using these fully populated databases, several simple models were derived for recommending bat weights.

Adolescent↗

Prevalence and patterns of same-gender sexual contact among men.

The prevalence and patterns of same-gender sexual contact among men are key components of models of the spread of HIV infection and AIDS in the U.S. population. Previous estimates by Kinsey et al. from data collected between 1938 and 1948 have been widely criticized for inadequacies of sample design. New lower-bound estimates of prevalence developed from data from a national sample survey conducted in 1970 indicate that minimums of 20.3 percent of adult men in the United States in 1970 had sexual contact to orgasm with another man at some time in life; 6.7 percent had such contact after age 19; and between 1.6 and 2.0 percent had such contact within the previous year. Although these estimates incorporate adjustments for missing data, the likelihood of underreporting suggests that these estimates might be lower bounds on the prevalence of same-gender sex among men. Two sets of alternative estimates are derived to assess the sensitivity of these estimates to the assumptions made in imputing values to missing data. Detailed estimates are presented by frequency of contact, age, education, and marital status; and supporting estimates are derived from a 1988 national survey. Data from both the 1970 and 1988 surveys indicate that never-married men are more likely than other men to have had same-gender sexual contacts within the last year. The 1970 survey also indicates, however, that approximately half the men estimated to have such contacts are found among the more numerous population of currently or previously married men.

Adult↗

Surveying minorities with limited-English proficiency: does data collection method affect data quality among Asian Americans?

BACKGROUND: Little is known about how modes of survey administration affect response rates and data quality among populations with limited-English proficiency (LEP). Asian Americans are a rapidly growing minority group with large numbers of LEP immigrants. OBJECTIVE: We sought to compare the response rates and data quality of interviewer-administered telephone and self-administered mail surveys among LEP Asian Americans. DESIGN: This was a randomized, cross-sectional study using a 78-item survey about quality of medical care that was given to Vietnamese, Mandarin, or Cantonese Chinese patients in their native language. MEASURES: We examined response rates and missing data by mode of survey and language groups. To examine nonresponse bias, we compared the sociodemographic characteristics of respondents and nonrespondents. To assess response patterns, we compared the internal-consistency reliability coefficients across modes and language groups. RESULTS: We achieved an overall response rate of 67% (322 responses of 479 patients surveyed). A higher response rate was achieved by phone interviews (75%) as compared with mail surveys with telephone reminder calls (59%). There were no significant differences in response rates by language group. The mean number of missing item for the mail mode was 4.14 versus 1.67 for the phone mode (P< or =0.000). There were no significant differences in missing data among the language groups and no significant differences in scale reliability coefficients by modes or language groups. CONCLUSIONS: Telephone interviews and mail surveys with phone reminder calls are feasible options to survey LEP Chinese and Vietnamese Americans. These methods may be less costly and labor-intensive ways to include LEP minorities in research.

Adult↗

Analysis of incomplete longitudinal binary data using multiple imputation.

We propose a propensity score-based multiple imputation (MI) method to tackle incomplete missing data resulting from drop-outs and/or intermittent skipped visits in longitudinal clinical trials with binary responses. The estimation and inferential properties of the proposed method are contrasted via simulation with those of the commonly used complete-case (CC) and generalized estimating equations (GEE) methods. Three key results are noted. First, if data are missing completely at random, MI can be notably more efficient than the CC and GEE methods. Second, with small samples, GEE often fails due to 'convergence problems', but MI is free of that problem. Finally, if the data are missing at random, while the CC and GEE methods yield results with moderate to large bias, MI generally yields results with negligible bias. A numerical example with real data is provided for illustration.

Clinical Trials as Topic↗

Trends in trauma care in England and Wales 1989-97. UK Trauma Audit and Research Network.

BACKGROUND: In 1988, the Royal College of Surgeons reported major deficiencies in trauma care in UK hospitals. We investigated whether and how that care has changed in the last decade by use of data collected by the UK Trauma Audit and Research Network. METHODS: We analysed injury-severity, process, and outcome variables from 91602 patients' records on the database at the end of 1997, collected from 97 (49% of trauma-receiving) hospitals in England, Wales, and two in Ireland. We did longitudinal analyses of odds of death, process variables, and individual hospitals' performance. We took account of potential selection bias from missing data and recruitment of new hospitals. FINDINGS: The severity-adjusted odds of death after trauma declined gradually from 1989 (odds ratio 1997/1989 0.63 [95% CI [0.49-0.82]). In 1997, the reduction in odds of death was significant even after adjustment for missing data (ratio 1997/1989 0.72 [0.55-0.92]) and recruitment of new hospitals (0.64 [0.44-0.93]). There was significant variability in the proportion of survivors (adjusted for severity of injury and age) between the highest and lowest 10% of UK hospitals. The time between the call to the emergency services and arrival at hospital increased from 32 min in 1989 to 45 min in 1997, irrespective of injury severity. The proportion of severely injured patients seen first by senior doctors increased from 32% to 60%. INTERPRETATION: Hospital care has made a valuable but variable contribution to reductions in case fatality after injury in the UK in the past 10 years, though further improvement is possible.

Aged↗

Estimating vaccine efficacy using auxiliary outcome data and a small validation sample.

In vaccine studies, a specific diagnosis of a suspected case by culture or serology of the infectious agent is expensive and difficult. Implementing validation sets in the study is less expensive and is easier to carry out. In studies using validation sets, the non-specific or auxiliary outcome is measured on each participant while the specific outcome is measured only for a small proportion of the participants. Vaccine efficacy, defined as one minus some measure of relative risk, could be severely attenuated if based only on the auxiliary outcome. Applying missing data analysis techniques could thus correct the bias while maintaining statistical efficiency. However, when the sample size in the validation sets is small and the vaccine is highly efficacious, all specific outcomes are likely to be negative in the validation set in the vaccinated group. Two commonly used missing data analysis methods, the mean score method and multiple imputation, depend on the ad hoc continuity correction when none of the specific outcomes are positive and the normality or log-normality assumption of relative risk, which may not hold when the relative risk is highly skewed, to estimate the confidence interval. In this paper, we propose a Bayesian method to estimate vaccine efficacy and its highest probability density (HPD) credible set using Monte Carlo (MC) methods when using auxiliary outcome data and a small validation sample. Comparing the performance of these approaches using data from a field study of influenza vaccine and simulations, we recommend to use the Bayesian method in this situation.

Adolescent↗

Epidemiology of notified campylobacteriosis in Western Australia.

Campylobacteriosis is one of the most common causes of gastroenteritis in Australia and the rates are thought to be increasing. This study has included all cases of campylobacteriosis that were notified in Western Australia between 1991 and 2001. The data for the study were received from Western Australian Notifiable Infectious Diseases Database located at the Communicable Disease Control Directorate of Western Australia. Rates of notification were calculated using the census data from 1991 for the general population and 1996 census data for the Aboriginal population. The notification rate of campylobacteriosis 89 per 100000 (95.0% confidence interval (CI) 87.6-91.4) for males and for females it was 78 per 100 000 (95.0%CI 87.6-91.4). Increased notification rates were seen in the very young, in males, in non-metropolitan areas and in the spring season. Aboriginal people had a much higher incidence than the rest of the population. Rates increased when laboratory notification was introduced. This study concludes that the rate of campylobacteriosis notification in Western Australia is increasing and is affecting younger children and young adults. The rate is higher in the Aboriginal population. As there were missing data from some cases the study faced some difficulties in interpreting the results. Recommendations for an improved surveillance system are made in order to minimise missing data.

Adolescent↗

Evaluation of predictors of mortality in frontotemporal dementia-methodological aspects.

OBJECTIVES: To retrospectively evaluate pre-diagnostic clinical features (predictors) of mortality in frontotemporal dementia (FTD). The main aim was to investigate if there were indications against interpreting missing data as signs of absence. MATERIAL AND METHODS: 96 cases with FTD, here defined as Dementia in Pick's disease according to ICD-10. The predictors were behavioural/psychiatric features, language impairment and neurological deficits up to the date of diagnosis. Each predictor was rated as present (Yes), absent (No) or not recorded (Missing), and evaluated according to its distribution and mortality pattern: if a feature was not recorded because it was absent, the mortality of the Missing and the No-category should hypothetically be close. Statistical methods included Kaplan-Meier survival curves and Cox regression analyses. RESULTS: Neurological deficits and language impairments were frequently recorded as present or absent, while non-recordings were more prevalent among the behavioural/psychiatric features. Some features were excluded as predictors because they showed too little variation. Analyses of the survival pattern indicated that in some features, the observations of the Missing-category could be interpreted as absence of the symptoms. In other features these observations had to be regarded as truly missing. CONCLUSIONS: In the retrospective evaluation of predictors of mortality a method for treating missing data was applied. The interpretation of non-recordings as signs of absence was supported by the analyses of the survival patterns in some of the studied features. However, the study underscores the importance of systematic estimations of pre-diagnostic clinical features in dementia.

Aged↗

Treating asthma by the guidelines: developing a medication management information system for use in primary care.

The aim of this study was to develop, implement, and assess an automated asthma medication management information system (MMIS) that provides patient-specific evaluative guidance based on 1997 NAEPP clinical consensus guidelines. MMIS was developed and implemented in primary care settings within a pediatric asthma disease management program. MMIS infrastructure featured a centralized database with Internet access. MMIS collects detailed patient asthma medication data, evaluates pharmacotherapy relative to practitioner-reported disease severity, symptom control and model of guideline-recommended severity-appropriate medications and produces a patient-specific "curbside consult" feedback report. A system algorithm translates actual detailed medication data into actual severity-specific medication-class combinations. A table-driven computer program compares actual medication-class combinations to a guideline-based medication-class combinations model. Methodology determines whether the patient was prescribed a "severity-appropriate" amount or an amount "more" or "less" medication than indicated for patient's reported severity. Feedback messages comment on comparison. Missing data, unrecognized amounts of controller medication or unrecognized medication combinations create error cases. Post hoc review analyzed error cases to determine prevalence of non-guideline medicating practices among these practitioners. Proportion of valid and error cases across two clinical visits before and after post hoc clinical review were measured, as well as proportion of severity-appropriate, out-of-severity and non-guideline medications. MMIS produced a valid feedback report for 83% of patient visits. Missing data accounted for 60% of error cases. Practitioners used severity-appropriate medications for 60% of cases. When non-severity-appropriate medications were used they tended to be "too much" rather than "too little" (22%, 5%), suggesting appropriate use of guideline-recommended "step down" therapy by these practitioners.

Anti-Asthmatic Agents↗