A statistics primer. Power analysis and sample size determination.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Meta-analyses of clinical trials are increasingly seeking to go beyond estimating the effect of a treatment and may also aim to investigate the effect of other covariates and how they alter treatment effectiveness. This requires the estimation of treatment-covariate interactions. Meta-regression can be used to estimate such interactions using published data, but it is known to lack statistical power, and is prone to bias.The use of individual patient data can improve estimation of such interactions, among other benefits, but it can be difficult and time-consuming to collect and analyse. This paper derives, under certain conditions, the power of meta-regression and IPD methods to detect treatment-covariate interactions. These power formulae are shown to depend on heterogeneity in the covariate distributions across studies. This allows the derivation of simple tests, based on heterogeneity statistics, for comparing the statistical power of the analysis methods.
Most analysts agree that screening women for breast cancer can reduce the death rate by at least 25% to 30%. In 1993 the National Cancer Institute created a great deal of confusion among women and their physicians by withdrawing support for screening women aged 40 to 49 years. The controversy arose as a result of the inappropriate and scientifically insupportable analysis of data from the early results from the randomized controlled trials of screening. Despite that the trials were not designed to evaluate women aged 40 to 49 years as a separate subgroup and that they lacked the statistical power to permit legitimate analysis of this subgroup, women aged 40 to 49 years were analyzed separately, and erroneous conclusions were drawn. In an effort to justify these conclusions, other data have been analyzed improperly to try to suggest that significant screening parameters change at menopause or at the age of 50, whereas the facts do not support any abrupt change. There are no parameters such as breast density, cancer detection rate, or positive predictive value that change abruptly at age 50 or any other age. With longer follow-up of the screening trial data and the commensurately greater statistical power, the most recent meta-analyses provide statistically significant proof that screening can reduce the death rate from breast cancer for women aged 40 to 49 years by at least 24%.
We review the evidence examining the relation of reproductive function and exposure to vinyl chloride and selected structural analogs. Investigation of these compounds for possible reproductive effects has focused on paternal exposure, a much less well studied route than maternal exposure. Drawing on animal models, we discuss what is known about the possible reproductive consequences of exposure to the father as well as to the mother. In evaluating the studies of reproductive outcome in relation to vinyl chloride or analogs, we consider what biologic model may have been tested and whether there was statistical power to detect moderate increases in risk. Parameters influencing statistical power are reviewed, and recommended sample sizes are set out which would insure sufficient power, in future studies, to detect adverse effects.
BACKGROUND: In method comparison studies, it is of importance to assure that the presence of a difference of medical importance is detected. For a given difference, the necessary number of samples depends on the range of values and the analytical standard deviations of the methods involved. For typical examples, the present study evaluates the statistical power of least-squares and Deming regression analyses applied to the method comparison data. METHODS: Theoretical calculations and simulations were used to consider the statistical power for detection of slope deviations from unity and intercept deviations from zero. For situations with proportional analytical standard deviations, weighted forms of regression analysis were evaluated. RESULTS: In general, sample sizes of 40-100 samples conventionally used in method comparison studies often must be reconsidered. A main factor is the range of values, which should be as wide as possible for the given analyte. For a range ratio (maximum value divided by minimum value) of 2, 544 samples are required to detect one standardized slope deviation; the number of required samples decreases to 64 at a range ratio of 10 (proportional analytical error). For electrolytes having very narrow ranges of values, very large sample sizes usually are necessary. In case of proportional analytical error, application of a weighted approach is important to assure an efficient analysis; e.g., for a range ratio of 10, the weighted approach reduces the requirement of samples by >50%. CONCLUSIONS: Estimation of the necessary sample size for a method comparison study assures a valid result; either no difference is found or the existence of a relevant difference is confirmed.
Independent contrasts are widely used to incorporate phylogenetic information into studies of continuous traits, particularly analyses of evolutionary trait correlations, but the effects of taxon sampling on these analyses have received little attention. In this paper, simulations were used to investigate the effects of taxon sampling patterns and alternative branch length assignments on the statistical performance of correlation coefficients and sign tests; "full-tree" analyses based on contrasts at all nodes and "paired-comparisons" based only on contrasts of terminal taxon pairs were also compared. The simulations showed that random samples, with respect to the traits under consideration, provide statistically robust estimates of trait correlations. However, exact significance tests are highly dependent on appropriate branch length information; equal branch lengths maintain lower Type I error than alternative topological approaches, and adjusted critical values of the independent contrast correlation coefficient are provided for use with equal branch lengths. Nonrandom samples, with respect to univariate or bivariate trait distributions, introduce discrepancies between interspecific and phylogenetically structured analyses and bias estimates of underlying evolutionary correlations. Examples of nonrandom sampling processes may include community assembly processes, convergent evolution under local adaptive pressures, selection of a nonrandom sample of species from a habitat or life-history group, or investigator bias. Correlation analyses based on species pairs comparisons, while ignoring deeper relationships, entail significant loss of statistical power and as a result provide a conservative test of trait associations. Paired comparisons in which species differ by a large amount in one trait, a method introduced in comparative plant ecology, have appropriate Type I error rates and high statistical power, but do not correctly estimate the magnitude of trait correlations. Sign tests, based on full-tree or paired-comparison approaches, are highly reliable across a wide range of sampling scenarios, in terms of Type I error rates, but have very low power. These results provide guidance for selecting species and applying comparative methods to optimize the performance of statistical tests of trait associations.
BACKGROUND: In many randomized and non-randomized comparative trials, researchers measure a continuous endpoint repeatedly in order to decrease intra-patient variability and thus increase statistical power. There has been little guidance in the literature as to selecting the optimal number of repeated measures. METHODS: The degree to which adding a further measure increases statistical power can be derived from simple formulae. This "marginal benefit" can be used to inform the optimal number of repeat assessments. RESULTS: Although repeating assessments can have dramatic effects on power, marginal benefit of an additional measure rapidly decreases as the number of measures rises. There is little value in increasing the number of either baseline or post-treatment assessments beyond four, or seven where baseline assessments are taken. An exception is when correlations between measures are low, for instance, episodic conditions such as headache. CONCLUSIONS: The proposed method offers a rational basis for determining the number of repeat measures in repeat measures designs.
BACKGROUND: Propagated tissue degeneration, especially during aging, has been shown to be enhanced through potentiation of innate immune responses. Neurodegenerative diseases and a wide variety of inflammatory conditions are linked together and several anti-inflammatory compounds considered as having therapeutic potential for example in Alzheimer's disease (AD). In vitro brain slice techniques have been widely used to unravel the complexity of neuroinflammation, but rarely, has the power of the model itself been reported. Our aim was to gain a more detailed insight and understanding of the behaviour of hippocampus tissue slices in serum-free, interface culture per se and after exposure to different pro- and anti-inflammatory compounds. METHODS: The responses of the slices to pro- and anti-inflammatory stimuli were monitored at various time points by measuring the leakage of lactate dehydrogenase (LDH) and the release of cytokines interleukin 6 (IL-6) and tumour necrosis factor alpha (TNF-alpha) and nitric oxide (NO) from the culture media. Histological methods were applied to reveal the morphological status after exposure to stimuli and during the time course of the culture period. Statistical power analysis were made with nQuery Advisor, version 5.0, (Statistical Solutions, Saugus, MA) computer program for Wilcoxon (Mann-Whitney) rank-sum test. RESULTS: By using the interface membrane culture technique, the hippocampal slices largely recover from the trauma caused by cutting after 4-5 days in vitro. Furthermore, the cultures remain stable and retain their responsiveness to inflammatory stimuli for at least 3 weeks. During this time period, cultures are susceptible to modification by inflammatory stimuli as assessed by quantitative biochemical assays and morphological characterizations. CONCLUSION: The present report outlines the techniques for studying immune responses using a serum-free slice culture model. Statistically powerful data under controlled culture conditions and with ethically justified use of animals can be obtained as soon as after 4-5 DIV. The model is most probably suitable also for studies of chronic inflammation.
A no-observed-effect level (NOEL) of 0.1 mg/kg/day was reported for inhibition of red blood cell (RBC) acetylcholinesterase (AChE) in two groups of Beagle dogs fed chlorpyrifos (0, 0.01, 0.03, 0.1, 1 or 3 mg/kg/day) in the diet for 1 or 2 years (McCollister et al., Food Cosmet. Toxicol. 12 (1974) 45-61). The statistical analyses were by t-test that had low statistical power due to small sample sizes. Common time points for blood samples in both phases allowed a reanalysis of the grouped data over a 1-year time period. The reanalysis increased statistical power by increasing the sample size to n=14 from n=3 or 4, and decreasing the variance, by statistical step-by-step aggregation of the data from both phases, both sexes, and four sample periods. Factors retained in the ANOVA were dose, sex, and phase (sex-by-dose was not significant). Contrasts with one-sided t-tests indicated the 1 and 3 mg/kg/day groups had significantly inhibited RBC AChE (P<0.0001). At alpha=0.05, the uncorrected one-sided model had 80% power to detect a 12% decrease, 93% power for a 15% decrease, and 99.5% power for a 20% decrease in AChE activity. Overall, the reanalysis had high power to detect a clinically significant decrease in RBC AChE activity, and substantiated the original NOEL for chronic treatment of dogs to dietary chlorpyrifos at 0.1 mg/kg/day.
The levels of PCBs, HCB, HCHs, DDTs, Cd, Pb, Hg and Se, and especially the variability in biota obtained during Phase 1 of the Greenland AMAP-programme have been used to illustrate the ability of the programme to detect differences in contaminant levels over time. The statistical power of t-tests of contaminant levels are illustrated according to various scenarios of magnitude of change, significance level and sample size. The statistical power of various time series of contaminant levels to detect linear trends in mean log-concentrations, including a random between-year variation component, is illustrated. We conclude that the ability to detect differences is rather poor for many combinations of contaminants and media, and that long time series are needed before temporal trends are likely to be detected.
This article is a primer on issues in designing, testing, and interpreting interaction or moderator effects in research on family psychology. The first section focuses on procedures for testing and interpreting simple effects and interactions, as well as common errors in testing moderators (e.g., testing differences among subgroup correlations, omitting components of products, and using median splits). The second section, devoted to difficulties in detecting interactions, covers such topics as statistical power, measurement error, distribution of variables, and mathematical constraints of ordinal interactions. The third section, devoted to design issues, focuses on recommendations such as including reliable measures, enhancing statistical power, and oversampling extreme scores. The topics covered should aid understanding of existing moderator research as well as improve future research on interaction effects.
Review articles have focused attention on and cited possible reasons for the nonreplication of genetic association studies. Herein, we illustrate how one might work through these possible reasons to make a judgment about the most plausible reason(s) when faced with two or more studies which yield seemingly inconsistent results. In the first study, 342 treatment-seeking smokers were genotyped for the Val108Met polymorphism in the functional catechol-O-methyl-transferase (COMT) locus. Alleles coding Val at codon 108 are denoted as H and those coding Met are denoted as L. An association between presence of the "H" (high activity) allele and pretreatment level of nicotine dependence level using the Fagerstrom Test for Nicotine Dependence was detected (P = 0.0072), after controlling for baseline body mass index (BMI, kg/m2), depression symptoms, and age. To validate this initial finding, 443 treatment-seeking smokers from an independent smoking cessation clinical trial were genotyped for the COMT polymorphism. Within the second study, no association between presence of the "H" allele and nicotine dependence was detected (P = 0.6418) after controlling for baseline BMI, depression symptoms, and age. We critically reviewed both studies with regard to often cited reasons for nonreplication, including type I error, population stratification, low statistical power, and imprecise measures of phenotype. Although in our opinion the failure to replicate the initial association in the second study is likely either the result of low statistical power to detect a small effect or effect heterogeneity, thorough analyses failed to definitively identify the reason for nonreplication.
Design of clinical trials to establish drug efficacy in chronic pain is a complicated issue, and numerous factors must be considered in identifying an optimal study design. Investigators should begin by identifying a focused and testable research question, with the outcome variables operationalized in a way that allows appropriate quantitative analysis. The prospective, randomized, placebo-controlled, doubled-blind study using validated quantitative measures is considered the optimal clinical trial design. Within this general study type, the between-subjects design has less statistical power than does the crossover design, in which the patient serves as his or her own control. However, potential problems with the drug effects from the first condition carrying over into, and confounding, the second drug condition present a noteworthy limitation that must be addressed through adequate washout periods and statistical control if crossover designs are used. Retrospective designs may be useful primarily to take advantage of samples of convenience for development of pilot data that provide the basis for conducting better-controlled prospective studies. Sample selection issues must be considered during study design, including the sample size required (based on statistical power analysis), appropriate inclusion and exclusion criteria, and likely availability of qualifying patients. There are numerous statistical options for analyzing data that must be selected based on whether data are parametric or nonparametric, whether within-subject (crossover) or between-subject comparisons are used, and whether baseline values of outcome measures affect the degree of change in these measures over the course of the study. Involvement of a biostatistical consultant is recommended during all phases of clinical trials.
Various surgical procedures may cause temporary interruption of spinal cord blood supply and may result in irreversible ischemic injury and neurological deficits. The cascade of events that leads to neuronal death following ischemia may be amenable to pharmacological manipulations that aim to increase the tolerable duration of ischemia. Many agents have been evaluated in experimental spinal cord ischemia (SCI). In order to investigate whether an agent is available that justifies clinical evaluation, the literature on pharmacological neuroprotection in experimental SCI was systematically reviewed to assess the neuroprotective efficacy of the various agents. In addition, the strength of the evidence for neuroprotection was investigated by analyzing the methodology. The authors used a systematic review to conduct this evaluation. The included studies were analyzed for neuroprotection and methodology. In order to be able to compare the various agents for neuroprotective efficacy, relative risks and confidence intervals were calculated from the data in the results sections. A total of 103 studies were included. Seventy-nine different agents were tested. Only 14 of the agents tested did not afford protection at all. A large variation was observed in the experimental models to produce SCI. This variation limited comparison of the individual agents. In 48 studies involving 31 single agents, the relative risks and confidence intervals could be calculated. An analysis of the methodology revealed poor temperature management and lack of statistical power in the majority of the 103 studies. The results suggest that numerous agents may protect the spinal cord from transient ischemia. However, poor temperature management and lack of statistical power severely weakened the evidence. Consequently, clinical evaluation of pharmacological neuroprotection in surgical procedures that carry a risk of ischemic spinal cord damage is not justified on the basis of this study.
BACKGROUND AND PURPOSE: Eligibility criteria determine the external validity (generalizability) of the results of randomized controlled trials. To increase the number of outcome events, and hence statistical power, some recent stroke prevention trials have required additional vascular risk factors for eligibility. METHODS: To assess the merits of additional eligibility criteria in stroke prevention trials, we analyzed data from 3 trials and 1 hospital-referred series of patients with a transient ischemic attack or minor ischemic stroke. Patients were stratified according to 2 sets of additional risk factors similar to those used in recent trials (MATCH, SPORTIF and PRoFESS); risk of stroke, myocardial infarction, or vascular death was calculated in relation to the number of risk factors. RESULTS: Although the observed risk during follow-up did increase with the number of risk factors present (P<0.01 for both sets), the risks in patients with > or =1 risk factors were not substantially greater than those in all patients. Consequently, although the proportions of patients with no risk factors in the 4 cohorts differed substantially between the 2 sets of eligibility criteria (21% to 28% versus 56% to 73%), in neither case could their exclusion be justified on statistical grounds. CONCLUSIONS: The degree of patient selection introduced by use of additional vascular risk factors as eligibility criteria for trials can differ substantially between apparently similar sets of risk factors. Given that the potential for additional eligibility criteria to undermine generalizability and prolong recruitment outweighs any benefits in terms of statistical power, the exclusion of patients with no risk factors is difficult to justify.
It has been claimed that protein-protein interaction (PPI) networks are scale-free, and that identifying high-degree "hub" proteins reveals important features of PPI networks. In this paper, we evaluate the claims that PPI node degree sequences follow a power law, a necessary condition for networks to be scale-free. We provide two PPI network examples which clearly do not have power laws when analyzed correctly, and thus at least these PPI networks are not scale-free. We also show that these PPI networks do appear to have power laws according to methods that have become standard in the existing literature. We explain the source of this error using numerically generated data from analytic formulas, where there are no sampling or noise ambiguities.
PURPOSE: A major goal of epidemiology is to discover the causes of disease in populations. The aim of this study was to develop a method for establishing research priorities within this very broad area of scientific inquiry. METHODS: While the approach is applicable to many diseases, cancer was used here, in part because of its importance to both individuals and governments. Measures of disease were estimated from data in the Ontario Cancer Registry, and combined to yield a single assessment of impact for each cancer site. Measures of exposure prevalence were identified from recent population health surveys. Cross-classification by disease and exposure rankings yielded a matrix of scores of "relative importance" for each cancer-exposure combination. Onto this matrix was overlaid: 1) estimates of statistical power for examining each association; 2) biological plausibility of each association; and 3) strength of the epidemiological evidence supporting each association. RESULTS: The disease-exposure matrix, viewed in light of statistical power, biological plausibility, and current epidemiological evidence, yielded, in the examples shown, some potentially interesting yet understudied associations (e.g., non-Hodgkin lymphoma (NHL) and certain aspects of dietary intake). CONCLUSIONS: The associations identified within the research hierarchy suggest not only new avenues for etiologic research, but also priorities for research focus.
Fifty-one randomized trials of neurosurgical topics have been identified. These have been analyzed for content, quality of reporting, and quality of statistical design and analysis. In general, there were major omissions in the reporting of critical portions of the studies, the statistical analysis was highly variable in quality, and the statistical power was low. Few, if any, major neurosurgical questions have been answered by clinical trials. This may reflect the generally low quality of these trials. To improve the quality of such trials, neurosurgeons must seek biostatistical advice in planning and conducting trials and must cooperate in multi-institutional trials to increase statistical power.