Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “missing data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

An application of a weighting method to adjust for nonresponse in standardized incidence ratio analysis of cohort studies.

PURPOSE: Cohort studies often conduct periodic follow-up interviews (or waves) to determine disease incidence since the previous follow-up and to update measures of exposure and confounders. The common practice of excluding nonrespondents from standardized incidence ratio (SIR) analyses of these cohorts can bias the estimates of interest if nonrespondents and respondents differ on important characteristics related to outcomes of interest. We propose an analytic approach to reduce the impact of nonresponse in the analyses of SIRs. METHODS: Logistic regression models controlling baseline information are used to estimate the propensity, or the probability of response; the reciprocals of these propensities are used as weights in the analysis of risk. This is illustrated in the analysis of 15 years of follow-up of a cohort of US radiologic technologists after an initial interview to assess the risk at several cancer sites from occupational radiation exposure. We use information from the baseline survey and certification records to compute the propensity of responding to the second survey. SIRs are computed using Surveillance, Epidemiology, and End Results (SEER) cancer incidence rates. Variances of the SIRs are estimated by a jackknife method that accounts for additional variability resulting from estimation of the weights. RESULTS: We find that, in this application, weighting alters point estimates and confidence limits only to a small degree, thus providing reassurance that the results are robust to nonresponse. This indicates that results from the analyses excluding the missing data may be slightly biased and weighting helps in reducing the nonresponse bias. CONCLUSION: This method is flexible, practical, easy to use with existing software, and is applicable to missing data from cohorts with baseline information on all subjects.

Bias↗

Youth recall and TriTrac accelerometer estimates of physical activity levels.

PURPOSE: To examine significance of missing data and describe physical activity patterns using recall and accelerometer measures among youth in a nonlaboratory setting. METHODS: Fifty-four middle-school students wore TriTrac-R3D monitors (TTM) and completed an interviewer-prompted 24-h recall during two, 5-d monitoring sessions. We coded 2860 30-min recall intervals to a standard MET compendium. Complete TTM data were gathered for 43 students. Ordinal multinomial models tested for bias in TTM estimates of activity levels due to: 1) exclusion of subjects with incomplete TTM data, and 2) exclusion of intervals within days due to missing TTM data. RESULTS: Students with complete monitor data had an average 12.5 +/- 0.9 monitored hours per day over 5.5 +/- 2.1 d. Compared with students with incomplete monitoring data, they reported similar proportions of recall 30-min intervals at sedentary (68% vs 69%), light (14% vs 15%), moderate (11% vs 10%), and vigorous (7% vs 6%) intensity levels (P = 0.63). The proportion of recall intervals (within days) with and without simultaneous monitoring data did not differ by activity intensity (P = 0.64) across sedentary (69% vs 67%), light (14% vs 12%), moderate (11% vs 10%), and vigorous (6% vs 9%) categories. Recalls overestimated percent time per day in moderate and vigorous activity relative to TTM (22.8% vs 8.9%, P < 0.0001). Boys reported higher percent of time than girls in vigorous activity (10.9% vs 3.9%, P < 0.05). Girls reported more time than boys (9.5% vs 6.4%, P < 0.05) in light activities. No significant sex differences were observed using TTM. CONCLUSIONS: Missing TTM data did not bias estimates of activity levels. Self-reported activity measures overestimated moderate and vigorous activity relative to the TTM and varied by sex.

Adolescent↗

The effect of correlation structure on treatment contrasts estimated from incomplete clinical trial data with likelihood-based repeated measures compared with last observation carried forward ANOVA.

Valid analyses of longitudinal data can be problematic, particularly when subjects dropout prior to completing the trial for reasons related to the outcome. Regulatory agencies often favor the last observation carried forward (LOCF) approach for imputing missing values in the primary analysis of clinical trials. However, recent evidence suggests that likelihood-based analyses developed under the missing at random framework provide viable alternatives. The within-subject error correlation structure is often the means by which such methods account for the bias from missing data. The objective of this study was to extend previous work that used only one correlation structure by including several common correlation structures in order to assess the effect of the correlation structure in the data, and how it is modeled, on Type I error rates and power from a likelihood-based repeated measures analysis (MMRM), using LOCF for comparison. Data from four realistic clinical trial scenarios were simulated using autoregressive, compound symmetric and unstructured correlation structures. When the correct correlation structure was fit, MMRM provided better control of Type I error and power than LOCF. Although misfitting the correlation structure in MMRM inflated Type I error and altered power, misfitting the structure was typically less deleterious than using LOCF. In fact, simply specifying an unstructured matrix for use in MMRM, regardless of the true correlation structure, yielded superior control of Type I error than LOCF in every scenario. The present and previous investigations have shown that the bias in LOCF is influenced by several factors and interactions between them. Hence, it is difficult to precisely anticipate the direction and magnitude of bias from LOCF in practical situations. However, in scenarios where the overall tendency is for patient improvement, LOCF tends to: 1) overestimate a drug's advantage when dropout is higher in the comparator and underestimate the advantage when dropout is lower in the comparator; 2) overestimate a drug's advantage when the advantage is maximum at intermediate time points and underestimate the advantage when the advantage increases over time; and 3) have a greater likelihood of overestimating a drug's advantage when the advantage is small. In scenarios in which the overall tendency is for patient worsening, the above biases are reversed. In the simulation scenarios considered in this study, which were patterned after acute phase neuropsychiatric clinical trials, the likelihood-based repeated measures approach, implemented with standard software, was more robust to the bias from missing data than LOCF, and choice of correlation structure was not an impediment to its implementation.

Analysis of Variance↗

A comparison of drugs versus placebo for the treatment of dysthymia.

OBJECTIVES: Dysthymia is a depressive disorder of chronic nature but of less severity than major depression, which depressive symptoms are more or less continuous for at least two years. The aim of this review was to conduct a systematic review of all RCTs comparing drugs and placebo for dysthymia. SEARCH STRATEGY: Electronic searches of Cochrane Library, EMBASE, MEDLINE, PsycLIT, Biological Abstracts and LILACS; reference searching; personal communication; conference abstracts; unpublished trials from the pharmaceutical industry; book chapters on the treatment of depression. SELECTION CRITERIA: The inclusion criteria for all randomised controlled trials were that they should focus on the use of drugs versus placebo for dysthymic patients. Exclusion criteria were: non randomised, mixed major depression/ dysthymia (trials not providing separate data) and depression secondary to other disorders (e.g. substance abuse). DATA COLLECTION AND ANALYSIS: The reviewers extracted the data independently. In order to achieve an intention-to-treat analysis, when trials failed to report it was assumed that people who died or dropped out had no improvement. Authors of relevant trials were contacted for additional and missing data. Absence of treatment response as defined by authors was the main measure of outcome used. Relative Risks (RR) and 95% confidence intervals (CI) of dichotomous data were calculated with the Random Effects Model. Where possible, number needed to treat (NNT) and number needed to harm (NNH) were estimated, taking the reciprocal of the absolute risk reduction. MAIN RESULTS: Currently the review includes 15 trials. Similar results were obtained in terms of efficacy for different groups of drugs, such as tricyclic (TCA), selective serotonin reuptake inhibitors (SSRI), monoamine oxidase inhibitors (MAOI) and other drugs (sulpiride, amineptine, and ritanserin). The pooled RR for absence of treatment response was 0. 68 (95% CI 0.59-0.78) for TCA and the NNT was 4.3 (95% CI 3.2-6.5). SSRIs showed similar RR for this outcome: 0.64 (95% CI 0.55-0.74), the NNT being 4.7 (95% CI 3.5-6.9). Concerning MAOIs, the RR was 0. 59 (95% CI 0.48-0.71) and the NNT was 2.9 (95% CI 2.2-4.3). Other drugs (amisulpride, amineptine and ritanserin) showed similar results in terms of absence of treatment response. Using more stringent criteria for improvement - full remission - the results were unchanged. Patients treated on TCA were more likely to report adverse events, compared with placebo. REVIEWER'S CONCLUSIONS: Drugs are effective in the treatment of dysthymia with no differences between and within class of drugs. Tricyclic antidepressants are more likely to cause adverse events and dropouts. As dysthymia is a chronic condition, there remains little information on quality of life and medium or long-term outcome.

Antidepressive Agents↗

Correcting for numerator/denominator bias when assessing changing inequalities in occupational class mortality, Australia 1981 -2002.

OBJECTIVE: Comparisons of the changing patterns of inequalities in occupational mortality provide one way to monitor the achievement of equity goals. However, previous comparisons have not corrected for numerator/denominator bias, which is a consequence of the different ways in which occupational details are recorded on death certificates and on census forms. The objective of this study was to measure the impact of this bias on mortality rates and ratios over time. METHODS: Using data provided by the Australian Bureau of Statistics, we examined the evidence for bias over the period 1981 -2002, and used imputation methods to adjust for this bias. We compared unadjusted with imputed rates of mortality for manual/non-manual workers. FINDINGS: Unadjusted data indicate increasing inequality in the age-adjusted rates of mortality for manual/non-manual workers during 1981 -2002. Imputed data suggest that there have been modest fluctuations in the ratios of mortality for manual/non-manual workers during this time, but with evidence that inequalities have increased only in recent years and are now at historic highs. CONCLUSION: We found that imputation for missing data leads to changes in estimates of inequalities related to social class in mortality for some years but not for others. Occupational class comparisons should be imputed or otherwise adjusted for missing data on census or death certificates.

Adolescent↗

Analysis of incomplete multivariate data from repeated measurement experiments.

This paper analyses two sets of data that consist of repeated measurements with missing data. The missing observations always occur at the end of the series of repeated measurements. The score test for multivariate normal data is used to compare treatment groups; if the original data are not multivariate normal they are replaced by expected normal scores.

Animals↗

Partial imputation approach to analysis of repeated measurements with dependent drop-outs.

In clinical trials repeated measurements of a response variable are usually taken at prespecified time-points to compare the treatment effects. However, the comparison of treatment effects is often complicated by missing data caused by the withdrawal of some patients before the end of the study (that is, drop-outs). When the drop-out process depends on the response variable of interest, ignoring missing data may lead to biased comparison of the treatment effect. In this paper, conditions for ignoring the dependent missingness are investigated and a new approach using the usual testing procedure based on data with partial carrying-forward imputation is proposed. The proposed approach is conceptually and practically simple, and is motivated by making incremental improvement on the familiar 'all available data' (AAD) approach and the 'last value carrying forward' (LVCF) approach, which are commonly used in data analysis with drop-outs by practitioners. It is also compared favourably to the mixed-effect model approach with dependent drop-outs. Simulations and real data are used to evaluate and illustrate statistical properties of the proposed approach. The principle of the proposed approach can also be extended to using other imputation methods such as the multiple imputation.

Analysis of Variance↗

Simple adjustments for randomized trials with nonrandomly missing or censored outcomes arising from informative covariates.

In randomized trials with missing or censored outcomes, standard maximum likelihood estimates of the effect of intervention on outcome are based on the assumption that the missing-data mechanism is ignorable. This assumption is violated if there is an unobserved baseline covariate that is informative, namely a baseline covariate associated with both outcome and the probability that the outcome is missing or censored. Incorporating informative covariates in the analysis has the desirable result of ameliorating the violation of this assumption. Although this idea of including informative covariates is recognized in the statistics literature, it is not appreciated in the literature on randomized trials. Moreover, to our knowledge, there has been no discussion on how to incorporate informative covariates into a general likelihood-based analysis with partially missing outcomes to estimate the quantities of interest. Our contribution is a simple likelihood-based approach for using informative covariates to estimate the effect of intervention on a partially missing outcome in a randomized trial. The first step is to create a propensity-to-be-missing score for each randomization group and divide the scores into a small number of strata based on quantiles. The second step is to compute stratum-specific estimates of outcome derived from a likelihood-analysis conditional on the informative covariates, so that the missing-data mechanism is ignorable. The third step is to average the stratum-specific estimates and compute the estimated effect of intervention on outcome. We discuss the computations for univariate, survival, and longitudinal outcomes, and present an application involving a randomized study of dual versus triple combinations of HIV-1 reverse transcriptase inhibitors.

Acquired Immunodeficiency Syndrome↗

Substitution between Formal and Informal Care for Persons with Severe Mental Illness and Substance Use Disorders.

BACKGROUND: Persons with severe mental illness (SMI) often get extensive informal care from family members and friends as well as substantial amounts of formal treatment from paid professionals. Both sources of care are well documented, but very little is known about how one affects the other. AIMS OF THE STUDY: This analysis estimates the extent of substitution between direct care provided by family and friends and formal treatment for people with severe mental illness and substance use disorders. Separate estimates are generated for short-term and long-term effects. METHODS: Data are from a randomized clinical trial conducted at seven mental health centers in New Hampshire between 1989 and 1995. The study includes detailed data for 193 persons with dual disorders measured at study entry and every six months for three years. Hours of informal care were compared with total treatment costs within each six-month period to measure short-term effects. Average amount of informal care over three years represented long-term caregiving practices. Measures of informal care are from interviews with informal caregivers. Treatment costs are based on combined data from management information systems, Medicaid claims, hospital records, and self reports. We used mixed effects repeated measures regression to estimate longitudinal effects and a multiple imputation technique to test the sensitivity of results to missing data. RESULTS: In the short-term, persons with bipolar disorder used significantly more formal care as informal care increased (complementarity). The relationship between short-term informal and formal care was significantly weaker for persons with schizophrenia. For both diagnostic groups there was a long-term substitution effect; a 4-6% increase in informal care hours was associated with an approximate 1% decrease in formal care costs. DISCUSSION: Although they must be confirmed by further research, these findings suggest that there is a significant and strong relationship between care given by family and friends and that supplied by formal treatment providers. The analysis indicates that the short-term relationship between informal care and formal treatment tends to be complementary, but differs according to diagnosis. Long-term effects, which are possibly related to changing role perceptions, show substitution between the two forms of care. Missing data for family care hours in some time periods was a concern in this study. However, the consistency in results between the analyses that used imputed data and the model using only original data increase our confidence in the findings. Although there may be some endogeneity between formal and informal care in other treatment settings we believe the unique characteristics of the service-rich environment in which this study was conducted limit that concern here. IMPLICATIONS FOR HEALTH CARE PROVISION AND USE: The amount of care provided by informal caregivers has a significant impact on formal treatment costs. Models of care that explicitly acknowledge the interplay between the two types of care are needed to ensure efficient combinations of formal and informal care. IMPLICATIONS FOR HEALTH POLICY FORMULATION: How to best to encourage informal support, without overburdening caregivers, is a key challenge facing policy makers and providers of mental health services. The merits of various approaches to reducing caregiver burden is a subject that needs more attention from researchers. In the interim, the demands on informal caregivers may mount as efforts to reduce health care spending continue. IMPLICATIONS FOR FURTHER RESEARCH: Informal care is not often included in economic evaluations of mental health treatment. Although additional research is needed to understand better the mechanisms by which informal care and formal treatment are related, we believe our results offer a strong argument for including measures of informal care in future economic evaluations.

Journal Article↗

[Prevalence and management of patients with a prior history of atherothrombotic disease in primary care in France. Results of the ECLAT1 survey].

This cross-sectional study assessed the prevalence of subjects with a previous history of atherothrombotic disease (myocardial infarction, ischemic stroke and/or lower limb arterial disease) among patients treated in general medicine. A random sample of 3,009 French general practitioners was recruited. Patients who consulted one of these general practitioners on December 7th 2000 were included. Those with a previous history of atherothrombotic disease were identified and further data on their cardiovascular risk factors and drug use were collected. The prevalence of patients with a previous history of atherothrombotic disease was 2% [95% confidence interval: 1.9-2-1] in subjects younger than 65, 13.4% [12.7-14.2] between 65 and 74 and 17.0% [16.2-17.8] in subjects older than 74. Arterial hypertension was found in 62.2% of the patients with a previous history of atherothrombotic disease, overweight or obesity in 59.4%, hypercholesterolaemia in 55%, current or past smoking in 48.3%, and diabetes mellitus in 20.1%. The last blood pressure and LDL-cholesterol measurements were respectively higher than or equal to 140/90 mmHg and 3 mmol/l in 70.6% of the patients suffering from arterial hypertension (missing data in 2.2%) and in 48.2% of the patients suffering from hypercholesterolaemia (missing data in 31.4%). Atherothrombosis represents a significant part of the primary care activity in France. Despite a widespread antihypertensive and hypocholesterolaemic drug prescription, the control of cardiovascular risk factors is insufficient. The high prevalence of overweight may contribute to this poor control.

Adult↗

Using meta-regression in performing indirect-comparisons: comparing escitalopram with venlafaxine XR.

BACKGROUND: In the absence of well-powered, randomised, direct-comparison trials, indirect comparisons are the only option for comparing treatment strategies. Several methodologies have been developed and each has sparked criticism. Using direct comparisons of escitalopram versus venlafaxine extended release (XR), we explore the differences between the two compounds through indirect comparisons. METHODS: The CENTRAL, Medline and Embase databases were interrogated, focusing on randomized placebo-controlled clinical trials involving adult patients treated for major depressive disorder in the acute phase. Corresponding authors were contacted to reduce missing data. Effect sizes were derived from each study's primary outcome. For indirect comparisons, a global effect size was computed through meta-regression. For direct comparisons, the studies were considered separately due to missing data. Non-inferiority assessments were employed. The conclusion of the meta-regression was then compared with the conclusions made in direct comparison trials. RESULTS: Ten placebo-controlled studies--six assessing escitalopram and four assessing venlafaxine XR--and two direct comparison studies were retrieved. Escitalopram was found to be non-inferior to venlafaxine XR in both indirect and direct comparisons with results of mean -0.02 (unilateral 95% confidence interval [CI] -0.16 to infinity) and 0.23 (95% CI -0.01 to infinity), respectively. Results obtained by both indirect and direct comparisons were similar. Investigating the influence of age, gender repartition and severity at baseline suggests that results are consistent. Results were also considered robust against publication bias. CONCLUSIONS: This empirical finding suggests that escitalopram is non-inferior to venlafaxine XR. This reinforces the evidence found in direct comparisons trials. Indirect comparisons through meta-regression may be suitable to support decision-making. To fully assess its potential, further evaluation of this methodology, using other examples, is needed.

Adult↗

Measuring depression in nursing home residents with the MDS and GDS: an observational psychometric study.

BACKGROUND: The objective of this study was to examine the Minimum Data Set (MDS) and Geriatric Depression Scale (GDS) as measures of depression among nursing home residents. METHODS: The data for this study were baseline, pre-intervention assessment data from a research study involving nine nursing homes and 704 residents in Massachusetts. Trained research nurses assessed residents using the MDS and the GDS 15-item version. Demographic, psychiatric, and cognitive data were obtained using the MDS. Level of depression was operationalized as: (1) a sum of the MDS Depression items; (2) the MDS Depression Rating Scale; (3) the 15-item GDS; and (4) the five-item GDS. We compared missing data, floor effects, means, internal consistency reliability, scale score correlation, and ability to identify residents with conspicuous depression (chart diagnosis or use of antidepressant) across cognitive impairment strata. RESULTS: The GDS and MDS Depression scales were uncorrelated. Nevertheless, both MDS and GDS measures demonstrated adequate internal consistency reliability. The MDS suggested greater depression among those with cognitive impairment, whereas the GDS suggested a more severe depression among those with better cognitive functioning. The GDS was limited by missing data; the DRS by a larger floor effect. The DRS was more strongly correlated with conspicuous depression, but only among those with cognitive impairment. CONCLUSIONS: The MDS Depression items and GDS identify different elements of depression. This may be due to differences in the manifest symptom content and/or the self-report nature of the GDS versus the observer-rated MDS. Our findings suggest that the GDS and the MDS are not interchangeable measures of depression.

Aged↗

On the statistical analysis of the GS-NS0 cell proteome: imputation, clustering and variability testing.

We have undertaken two-dimensional gel electrophoresis proteomic profiling on a series of cell lines with different recombinant antibody production rates. Due to the nature of gel-based experiments not all protein spots are detected across all samples in an experiment, and hence datasets are invariably incomplete. New approaches are therefore required for the analysis of such graduated datasets. We approached this problem in two ways. Firstly, we applied a missing value imputation technique to calculate missing data points. Secondly, we combined a singular value decomposition based hierarchical clustering with the expression variability test to identify protein spots whose expression correlates with increased antibody production. The results have shown that while imputation of missing data was a useful method to improve the statistical analysis of such data sets, this was of limited use in differentiating between the samples investigated, and highlighted a small number of candidate proteins for further investigation.

Algorithms↗

Determinants of HIV infection among female commercial sex workers in northeastern Thailand: results from a longitudinal study.

Our objective was to estimate HIV seroconversion rates among commercial sex workers (CSWs) between 1990 and 1991 and to identify the behavioral, demographic, and reproductive determinants of these rates. This study has a prospective (n = 240 with 15 cases) and a cross-sectional component (n = 271 with 34 cases). In November 1990, HIV-negative female CSWs from 24 brothels in Khon Kaen city were interviewed and were followed prospectively for up to 1 year. In March, June, and September 1991, additional HIV-negative CSWs were enrolled and prospectively followed. HIV seroconversion rates were calculated, and the Cox regression model was used to estimate the relative risks of HIV seroconversion from demographic, sexual practice, and reproductive factors, adjusted for the effects of the others, among 232 of the 240 without missing data. Seroprevalence rates were also calculated for the 271 participants enrolled between March and December 1991, and relative risks of HIV seroprevalence were calculated for demographic, sexual practice, and reproductive risk factors among 184 of the 271 without missing data. The average seroprevalence was 12.5% (95% confidence interval 9.6-15.4%). With 1,947 person-months of observation obtained from 240 participants who were uninfected at baseline and seen at least twice during the course of the study, the cumulative incidence of HIV seroconversion between November 1990 and December 1991 was 9.4% (95% confidence interval 5.4-13.4%), and the average incidence rate of HIV seroconversion was 9.2 per 100 person-years (95% confidence interval 4.6-13.9 per 100 person-years). In the multivariate analysis, later date of enrollment into the study, having < 3 months experience as a CSW, and use of injectable contraceptives were the only risk factors that remained significant, with relative risks of 2.1 (95% confidence interval 1.2-3.7) for enrollment 3 months later, 3.8 (95% confidence interval 1.0-14.4) for < 3 months experience as a CSW versus > 3 months experience, and 3.9 (95% confidence interval 1.3-11.8) [corrected] for use of injectable contraceptives. In multivariate analysis of the cross-sectional data with 184 participants, of whom 21 were HIV seropositive, risk of HIV seropositivity increased significantly with current syphilis infection (odds ratio 5.8, 95% confidence interval 1.1-31.0). The results of this study will contribute to a better understanding of the risk factors of infection with HIV and thus allow for better targeting of group-specific interventions, particularly for CSWs and their clients. Further investigation of a possible association between injectable contraceptive use and HIV infection is needed.

Adolescent↗

On locating multiple interacting quantitative trait loci in intercross designs.

A modified version (mBIC) of the Bayesian Information Criterion (BIC) has been previously proposed for backcross designs to locate multiple interacting quantitative trait loci. In this article, we extend the method to intercross designs. We also propose two modifications of the mBIC. First we investigate a two-stage procedure in the spirit of empirical Bayes methods involving an adaptive (i.e., data-based) choice of the penalty. The purpose of the second modification is to increase the power of detecting epistasis effects at loci where main effects have already been detected. We investigate the proposed methods by computer simulations under a wide range of realistic genetic models, with nonequidistant marker spacings and missing data. In the case of large intermarker distances we use imputations according to Haley and Knott regression to reduce the distance between searched positions to not more than 10 cM. Haley and Knott regression is also used to handle missing data. The simulation study as well as real data analyses demonstrates good properties of the proposed method of QTL detection.

Algorithms↗

What is meant by intention to treat analysis? Survey of published randomised controlled trials.

OBJECTIVES: To assess the methodological quality of intention to treat analysis as reported in randomised controlled trials in four large medical journals. DESIGN: Survey of all reports of randomised controlled trials published in 1997 in the BMJ, Lancet, JAMA, and New England Journal of Medicine. MAIN OUTCOME MEASURES: Methods of dealing with deviations from random allocation and missing data. RESULTS: 119 (48%) of the reports mentioned intention to treat analysis. Of these, 12 excluded any patients who did not start the allocated intervention and three did not analyse all randomised subjects as allocated. Five reports explicitly stated that there were no deviations from random allocation. The remaining 99 reports seemed to analyse according to random allocation, but only 34 of these explicitly stated this. 89 (75%) trials had some missing data on the primary outcome variable. The methods used to deal with this were generally inadequate, potentially leading to a biased treatment effect. 29 (24%) trials had more than 10% of responses missing for the primary outcome, the methods of handling the missing responses were similar in this subset. CONCLUSIONS: The intention to treat approach is often inadequately described and inadequately applied. Authors should explicitly describe the handling of deviations from randomised allocation and missing responses and discuss the potential effect of any missing response. Readers should critically assess the validity of reported intention to treat analyses.

Data Collection↗

Accuracy of haplotype reconstruction from haplotype-tagging single-nucleotide polymorphisms.

Many investigators are now using haplotype-tagging single-nucleotide polymorphism (htSNPs) as a way of screening regions of the genome for association with disease. A common approach is to genotype htSNPs in a study population and to use this information to draw inferences about each individual's haplotypic makeup, including SNPs that were not directly genotyped. To test the validity of this approach, we simulated the exercise of typing htSNPs in a large sample of individuals and compared the true and inferred haplotypes. The accuracy of haplotype inference varied, depending on the method of selecting htSNPs, the linkage-disequilibrium structure of the region, and the amount of missing data. At the stage of selection of htSNPs, haplotype-block-based methods required a larger number of htSNPs than did unstructured methods but gave lower levels of error in haplotype inference, particularly when there was a significant amount of missing data. We present a Web-based utility that allows investigators to compare the likely error rates of different sets of htSNPs and to arrive at an economical set of htSNPs that provides acceptable levels of accuracy in haplotype inference.

Biometry↗