Search PubMed⌕ Search

Biomedical subjects

D J Spiegelhalter

Publications and source records attributed to D J Spiegelhalter.

At least 19 recordsLinked to original sources

Bayesian random effects meta-analysis of trials with binary outcomes: methods for the absolute risk difference and relative risk scales.

When conducting a meta-analysis of clinical trials with binary outcomes, a normal approximation for the summary treatment effect measure in each trial is inappropriate in the common situation where some of the trials in the meta-analysis are small, or the observed risks are close to 0 or 1. This problem can be avoided by making direct use of the binomial distribution within trials. A fully Bayesian method has already been developed for random effects meta-analysis on the log-odds scale using the BUGS implementation of Gibbs sampling. In this paper we demonstrate how this method can be extended to perform analyses on both the absolute and relative risk scales. Within each approach we exemplify how trial-level covariates, including underlying risk, can be considered. Data from 46 trials of the effect of single-dose ibuprofen on post-operative pain are analysed and the results contrasted with those derived from classical and Bayesian summary statistic methods. The clinical interpretation of the odds ratio scale is not straightforward. The advantages and flexibility of a fully Bayesian approach to meta-analysis of binary outcome data, considered on an absolute risk or relative risk scale, are now available.

Anti-Inflammatory Agents, Non-Steroidal↗

Prospective application of Bayesian monitoring and analysis in an "open" randomized clinical trial.

We describe the prospective application of Bayesian monitoring and analysis in an ongoing large multi-centre, randomized trial in which interim results are released to investigators. Substantial variability in prior opinion led us to reject the use of elicited clinical priors for monitoring, in favour of archetypal prior distributions representing reasonable scepticism and enthusiasm. Likelihoods for odds ratios for different covariate values are derived from a logistic regression model, which allows us to incorporate information from prognostic factors without resorting to specialized software. Priors, likelihoods and posterior distributions are regularly reported to both an independent Data Monitoring Committee and the trial investigators.

Bayes Theorem↗

Monitoring of large randomised clinical trials: a new approach with Bayesian methods.

BACKGROUND: In judging whether or not to continue enrolling patients into a randomised clinical trial, most data-monitoring and ethics committees (DMECs) rely on the p value for the difference in effect between the study groups. In the 1990s, two randomised controlled trials-one in patients with lung cancer and one in those with head and neck cancer-were instead monitored by Bayesian methods. We assessed the value of this approach in the monitoring of these clinical trials. METHODS: Before the trials opened, participating clinicians were asked their opinions on the expected difference between the study treatment (continuous hyperfractionated accelerated radiotherapy [CHART]) and conventional radiotherapy. These opinions were used to form an "enthusiastic" and a "sceptical" prior distribution. These prior distributions were combined with the trial data at each of the annual DMEC meetings. If, during monitoring, a result in favour of CHART was seen, the DMEC was to decide whether the results were sufficiently convincing to persuade a sceptic that CHART was worthwhile. Conversely, if there was apparently no or little difference, the DMEC was asked whether they thought the results sufficiently convincing to persuade an enthusiast that CHART was not worthwhile. FINDINGS: At each of the annual meetings, the DMEC concluded that there was insufficient evidence to convert either sceptics or enthusiasts, and that the trials should therefore remain open to recruitment. Neither trial was closed to recruitment earlier than planned. However if a conventional (p-value-based) stopping rule had been used, the lung-cancer trial would probably have been stopped. INTERPRETATION: This Bayesian approach to monitoring is simple to implement and straightforward for members of the DMEC to understand. In our opinion, it is more intuitively appealing than conventional approaches.

Bayes Theorem↗

Bayesian methods for cluster randomized trials with continuous responses.

Bayesian methods for cluster randomized trials extend the random-effects formulation by allowing both the use of external evidence on parameters and straightforward relaxation of the standard normality and constant variance assumptions. Care is required in specifying prior distributions on variance components, and a number of different options are explored with implied prior distributions for other parameters given in closed form. Markov chain Monte Carlo (MCMC) methods permit the fitting of very general models and the introduction of parameter uncertainty into power calculations. We illustrate these ideas using a published example in which general practices were randomized to intervention or control, and show that different choices of supposedly 'non-informative' prior distributions can have substantial influence on conclusions. We also illustrate the use of forward simulation methods in power calculations with uncertainty on multiple inputs. Bayesian methods have the potential to be very useful but guidance is required as to appropriate strategies for robust analysis. Our current experience leads us to recommend a standard 'non-informative' prior distribution for the within-cluster sampling variance, and an independent prior on the intraclass correlation coefficient (ICC). The latter may exploit background evidence or, as a reference analysis, be a uniform ICC or a 'uniform shrinkage' prior.

Bayes Theorem↗

Bayesian methods in health technology assessment: a review.

BACKGROUND: Bayesian methods may be defined as the explicit quantitative use of external evidence in the design, monitoring, analysis, interpretation and reporting of a health technology assessment. In outline, the methods involve formal combination through the use of Bayes's theorem of: 1. a prior distribution or belief about the value of a quantity of interest (for example, a treatment effect) based on evidence not derived from the study under analysis, with 2. a summary of the information concerning the same quantity available from the data collected in the study (known as the likelihood), to yield 3. an updated or posterior distribution of the quantity of interest. These methods thus directly address the question of how new evidence should change what we currently believe. They extend naturally into making predictions, synthesising evidence from multiple sources, and designing studies: in addition, if we are willing to quantify the value of different consequences as a 'loss function', Bayesian methods extend into a full decision-theoretic approach to study design, monitoring and eventual policy decision-making. Nonetheless, Bayesian methods are a controversial topic in that they may involve the explicit use of subjective judgements in what is conventionally supposed to be a rigorous scientific exercise. OBJECTIVES: This report is intended to provide: 1. a brief review of the essential ideas of Bayesian analysis 2. a full structured review of applications of Bayesian methods to randomised controlled trials, observational studies, and the synthesis of evidence, in a form which should be reasonably straightforward to update 3. a critical commentary on similarities and differences between Bayesian and conventional approaches 4. criteria for assessing the reporting of a Bayesian analysis 5. a comprehensive list of published 'three-star' examples, in which a proper prior distribution has been used for the quantity of primary interest 6. tutorial case studies of a variety of types 7. recommendations on how Bayesian methods and approaches may be assimilated into health technology assessments in a variety of contexts and by a variety of participants in the research process. METHODS: The BIDS ISI database was searched using the terms 'Bayes' or 'Bayesian'. This yielded almost 4000 papers published in the period 1990-98. All resultant abstracts were reviewed for relevance to health technology assessment; about 250 were so identified, and used as the basis for forward and backward searches. In addition EMBASE and MEDLINE databases were searched, along with websites of prominent authors, and available personal collections of references, finally yielding nearly 500 relevant references. A comprehensive review of all references describing use of 'proper' Bayesian methods in health technology assessment (those which update an informative prior distribution through the use of Bayes's theorem) has been attempted, and around 30 such papers are reported in structured form. There has been very limited use of proper Bayesian methods in practice, and relevant studies appear to be relatively easily identified. RESULTS: Bayesian methods in the health technology assessment context 1. Different contexts may demand different statistical approaches. Prior opinions are most valuable when the assessment forms part of a series of similar studies. A decision-theoretic approach may be appropriate where the consequences of a study are reasonably predictable. 2. The prior distribution is important and not unique, and so a range of options should be examined in a sensitivity analysis. Bayesian methods are best seen as a transformation from initial to final opinion, rather than providing a single 'correct' inference. 3. The use of a prior is based on judgement, and hence a degree of subjectivity cannot be avoided. However, subjective priors tend to show predictable biases, and archetypal priors may be useful for identifying a reasonable range of prior opinion.

Bayes Theorem↗

Estimating the true extent of cognitive decline in the old old.

OBJECTIVE: To measure cognitive change using a brief measure over a period of 9 years and to adjust for attrition in the sample. DESIGN: The Cambridge City over 75 Cohort (CC75C), a complete sample of the 75 years and older age group from five group general practices in the city of Cambridge with a systematic one-third of a further practice, all followed on four occasions. SETTING: Cambridge city, UK, the respondents' place of residence. PARTICIPANTS: A total of 2106 subjects were included at study entry. MEASUREMENTS: A brief interview, administered by a trained interviewer, containing a short cognitive scale and the Mini-Mental State Examination (MMSE) at baseline, 2.4 years, 6 years, and 9 years. RESULTS: Decline in MMSE scores occurred across the population and was greater in the oldest age groups. Attrition at later stages of the follow-up was associated with greater decline at earlier stages. Adjusting the results for loss to the sample leads to considerably higher estimates of decline, with the older age groups declining faster from lower levels. CONCLUSIONS: To date, cognitive decline in the very old has been considerably underestimated by longitudinal studies. If studies of population samples are to reflect the health and social needs of this frail group accurately, adjustments for the effect of attrition must be included before true decline can be estimated.

Age Factors↗

Use of person-years of CIN III exposure as a surrogate outcome measure in cervical cancer screening trials.

BACKGROUND: The reluctance to perform randomised trials on still unresolved issues of the cervical screening programme is largely due to the view that using invasive cancer as an end-point would require huge and lengthy studies. However, by assuming that the incidence of invasive cancer is related to the number of women with CIN III, an estimate of person-years of CIN III can act as a surrogate outcome. METHODS: By having a reliable model for the development of CIN III and the errors involved in taking the smears and biopsies, the person-years of CIN III can be estimated from a smear and biopsy history. Methods for comparing resulting distributions of person-years of CIN III are discussed. Sensitivity analyses on the error rates of smear and biopsy results, and on the incidence of onset and regression of CIN III are performed. RESULTS: Estimates of person-years of CIN III were calculated for women with mildly abnormal smears in two screening programmes. 11% of women in the Cambridge programme and 21% in the Aberdeen programme were estimated to have been exposed to CIN III for more than 12 months. The greater estimated person-years of CIN III in the Aberdeen study reflects the more conservative treatment policy which was operating there. DISCUSSION: The use of person-years of CIN III as a surrogate outcome can provide a practical and meaningful assessment of strategies for cervical cancer screening. Using CIN III, in place of invasive disease, considerably reduces the study duration and sample size required.

Biopsy↗

Reliability of league tables of in vitro fertilisation clinics: retrospective analysis of live birth rates.

OBJECTIVE: To determine to what extent institutions carrying out in vitro fertilisation can reasonably be ranked according to their live birth rates. DESIGN: Retrospective analysis of prospectively collected data on live birth rate after in vitro fertilisation. SETTING: 52 clinics in the United Kingdom carrying out in vitro fertilisation over the period April 1994 to March 1995. MAIN OUTCOME MEASURE: Estimated adjusted live birth rate for each clinic; their rank and its associated uncertainty. RESULTS: There were substantial and significant differences between the live birth rates of the clinics. There was great uncertainty, however, concerning the true ranks, particularly for the smaller clinics. Only one clinic could be confidently ranked in the bottom quarter according to this measure of performance. Many centres had substantial changes in rank between years, even though their live birth rate did not change significantly. CONCLUSIONS: Even when there are substantial differences between institutions, ranks are extremely unreliable statistical summaries of performance and change in performance, particularly for smaller institutions. Any performance indicator should always be associated with a measure of sampling variability.

Birth Rate↗

Comparison between two districts of the effects of an air pollution intervention on bronchial responsiveness in primary school children in Hong Kong.

STUDY OBJECTIVE: This study examined the impact on children's respiratory health of a government air quality intervention that restricted the sulphur content of fuels to 0.5% from July 1990 onwards. DESIGN/SETTING/PARTICIPANTS: This study examined the changes, one and two years after the introduction of the intervention, in airway hyperreactivity of non-asthmatic and non-wheezing, primary 4, 5, and 6, school children aged 9-12 years living in a polluted district compared with those in a less polluted district. Bronchial hyperreactivity (BHR)(a 20% decrease in FEV1 provoked by a cumulative dose of histamine less than 7.8 mumol) and bronchial reactivity slope (BR slope) (percentage change in logarithmic scale in FEV1 per unit dose of histamine) were used to estimate responses to a histamine challenge. The between districts differences after the intervention were studied to assess the effectiveness of the intervention. MAIN RESULTS: In cohorts, comparing measurements made before the intervention and one year afterwards, both BHR and BR slope declined from 29% to 16% (p = 0.026) and from 48 to 39 (p = 0.075) respectively in the polluted district; and from 21% to 10% (p = 0.001) and 42 to 36 (p > 0.100) in the less polluted district. Comparing measurements made in 1991 (one year after intervention) with those in 1992 (two years after intervention), only the polluted district showed a significant decline from 28% to 12% (p = 0.016) and from 46 to 35 (p = 0.014), for BHR and BR slope respectively, with a greater decline in both responses (p = 0.018 and 0.073) than in the less polluted district. CONCLUSION: Bronchial hyperresponsiveness tests can be used to support the evaluation of an air quality intervention. The demonstrated reduction in bronchial hyperresponsiveness is an indication of the effectiveness of the intervention.

Air Pollutants↗

Survival analysis in observational studies.

Multi-centre databases are making an increasing contribution to medical understanding. While the statistical handling of randomized experimental studies is well documented in the medical literature, the analysis of observational studies requires the addressing of additional important issues relating to the timing of entry to the study and the effect of potential explanatory variables not introduced until after that time. A series of analyses is illustrated on a small data set. The influence of single and multiple explanatory variables on the outcome after a fixed time interval and on survival time until a specific event are examined. The analysis of the effect on survival of factors that only come into play during follow-up is then considered. The aim of each analysis, the choice of data used, the essentials of the methodology, the interpretation of the results and the limitations and underlying assumptions are discussed. It is emphasized that, in contrast to randomized studies, the basis for selection and timing of interventions in observational studies is not precisely specified so that attribution of a survival effect to an intervention must be tentative. A glossary of terms is provided.

Algorithms↗

Setting the minimal metrically detectable change on disability rating scales.

OBJECTIVE: To determine the minimal metrically detectable change (MMDC) on rating scales by comparing two methods: (1) using the reliability coefficient derived from an external study in the calculation of the standard error of measurement (psychometric method); and (2) examining the variability of scores in a stable subsample from a longitudinal study (empirical method). DESIGN: Longitudinal survey. SETTING: General community. PARTICIPANTS: Population-based representative sample of community-dwelling people older than 75 (n = 572). MAIN OUTCOME MEASURE: Disability as measured by the Functional Autonomy Measuring System (SMAF). RESULTS: Using the psychometric method, a change in score of 3.7 on the SMAF was obtained using a reliability coefficient derived from an external test-retest study, and 5.2 using the reliability coefficient measured in a stable subsample of the longitudinal study. With the empirical method, a change of 5 points was established as the MMDC. CONCLUSION: The setting of the MMDC on a disability scale could be useful for calculating sample size or interpreting results from clinical trials because it helps to establish the minimal clinically important difference, which should be equal to or larger than the MMDC.

Aged↗

Trends in invasive cervical cancer incidence in East Anglia from 1971 to 1993.

OBJECTIVE: To study the trends in the incidence of invasive cervical cancer in East Anglia. DESIGN: Statistical analysis of age specific incidence rate for the period 1971-93 using East Anglian Cancer Registry data. SUBJECTS: All cases of invasive cervical cancer registered with the East Anglian Cancer registry, diagnosed in the period 1971-93. MAIN OUTCOME MEASURES: Changing incidence of cervical cancer. RESULTS: For the 20 years 1971-90, trends varied widely by district and by age group, with little discernible overall effect of the increasing screening activity. Since 1990, rates have fallen sharply in the age groups targeted for screening, with a reduction of 34% (95% confidence interval 26% to 42%) from that expected based on 1971-90 trends. This fall was preceded by a rapid rise in the national uptake of screening. A shift to more favourable stage at diagnosis has also occurred. CONCLUSION: Changes in the organisation and management of the national screening programme introduced in 1988 and 1989 seem to have led to substantial improvements in effectiveness.

Adult↗

Pharmacodynamics of cyclosporine in heart and heart-lung transplant recipients. I: Blood cyclosporine concentrations and other risk factors for cardiac allograft rejection.

We have attempted to determine the optimal clinical use of cyclosporine during the first 3 months after heart transplantation. We used multiple logistic regression to quantify how blood cyclosporine concentrations and other potential risk factors influence the risk of histologically confirmed acute rejection in 111 heart transplant recipients. A 50% increase in cyclosporine concentration was associated with a 15% reduction in risk of rejection in the subsequent 5 days (P=0.002). Increasing oral corticosteroid dose also protected against rejection (P=0.01). Rejection was over 2.5 times more likely during the first 20 postoperative days, and patients with 2 HLA-DR mismatches who were transplanted for cardiomyopathy or who had multiple previous rejection episodes were predisposed to further rejection (P<0.01). High short-term variability in cyclosporine concentrations was weakly associated with risk of rejection (P=0.1). Investigation of threshold levels for the cyclosporine concentration-effect relationship suggested that concentrations above 375 microgram L(-1) provide optimal protection against acute cardiac allograft rejection. This result yields an objectively defined therapeutic threshold for targeting early cyclosporine concentrations following heart transplantation, although the upper end of the range will depend on the individual's susceptibility to nephrotoxicity and infection.

Analysis of Variance↗

Pharmacodynamics of cyclosporine in heart and heart-lung transplant recipients. II: Blood cyclosporine concentrations and other risk factors for lung allograft rejection.

We have attempted to quantify the optimal clinical use of cyclosporine during the first 3 months after heart-lung transplantation. We used multiple logistic regression to investigate the influence of blood cyclosporine concentrations and other potential risk factors on histologically confirmed acute lung rejection in 50 heart-lung transplant recipients. A 50% increase in cyclosporine concentration was associated with a 25% reduction in risk of rejection in the subsequent 5 days (P=0.008). Increasing oral corticosteroid dose also protected against rejection (P=0.006). Rejection was over 4 times more likely to occur during the first 20 postoperative days (P=0.002). After 20 days, an FEV1 < or = 70% of the age-, sex-, and height-adjusted expected score was associated with a 4-fold increase in risk of rejection (P=0.01). Patients who had multiple previous rejection episodes were also predisposed to further rejection (P=0.005). An investigation of threshold levels for the cyclosporine concentration-effect relationship suggested that cyclosporine concentrations above 500 microg L(-1) provide optimal protection against acute lung allograft rejection. This result provides an objectively defined therapeutic threshold for targeting early cyclosporine concentrations following heart-lung transplantation.

Adult↗

An empirical comparison of expert-derived and data-derived classification trees.

Classification trees provide an attractively transparent discrimination technique, and may be derived from both expert opinion and from data analysis. We consider a real and complex problem concerning the diagnosis of babies with suspected critical congenital heart disease into one of 27 classes. A full loss matrix for all possible misclassifications was obtained from clinical assessments. A tree derived from expert opinion was compared with those derived from analysis of 571 past cases, both for the full problem and for a subset of 6 diseases. Automatic methods for tree creation and pruning were found to have problems for rare diseases, and hand-pruning was carried out. Inclusion of costs led to much improved clinical performance, even for trees that had originally been constructed to minimize classification errors. The expert tree showed a specific building strategy that could not be reproduced automatically. The expert tree generally outperformed those derived from data, particularly in the ability to identify important composite features.

Algorithms↗

Effects of an ambient air pollution intervention and environmental tobacco smoke on children's respiratory health in Hong Kong.

BACKGROUND: Two-thirds of complaints received by the Hong Kong Environmental Protection Department in 1988 were related to poor air quality. In July 1990 legislation was implemented to reduce fuel sulphur levels. The objective of this study was to measure the impact of the intervention on respiratory health in primary school children. METHODS: In all, 3521 children, mean age 9.51 years (SD = 0.78), from two districts with good and poor air quality respectively before intervention were followed yearly from 1989 to 1991. Children and parents reported the children's respiratory symptoms using self-completed questionnaires. Factor analysis was used to derive independent scores from 12 symptoms. Four groups of related symptoms were identified and binary variables (presence of any symptom in each group) were treated as dependent variables in modelling using generalized estimating equations procedures. RESULTS: In 1989 and 1990 an excess of respiratory symptoms was observed in the polluted compared with unpolluted district. The significant effects (odds ratio [OR], 95% confidence interval [CI], P value) associated with living in the polluted district were: cough and sore throat (OR = 1.22, 95% CI: 1.04-1.43, P < 0.01) and wheezing (OR = 1.35, 95% CI: 1.10-1.66, P < 0.01). After the intervention, in the polluted district only, sulphur dioxide levels fell by up to 80% and sulphate concentrations in respirable particulates by 38%. Between 1989 and 1990-1991 there was a greater decline in the polluted compared with the unpolluted district for reported symptoms of cough or sore throat, phlegm, and wheezing. The risks to respiratory health for children exposed to tobacco smoke in the home were higher than those for air pollution in both 1989 and 1990 and remained unchanged in 1991. CONCLUSIONS: Air quality can be improved by fuel controls but an effective intersectoral approach is required if other risks from environmental tobacco smoke are to be avoided.

Air Pollution↗

Bayesian approaches to random-effects meta-analysis: a comparative study.

Current methods for meta-analysis still leave a number of unresolved issues, such as the choice between fixed- and random-effects models, the choice of population distribution in a random-effects analysis, the treatment of small studies and extreme results, and incorporation of study-specific covariates. We describe how a full Bayesian analysis can deal with these and other issues in a natural way, illustrated by a recent published example that displays a number of problems. Such analyses are now generally available using the BUGS implementation of Markov chain Monte Carlo numerical integration techniques. Appropriate proper prior distributions are derived, and sensitivity analysis to a variety of prior assumptions carried out. Current methods are briefly summarized and compared to the full Bayes analysis.

Bayes Theorem↗