Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Sample size and the probability of a successful trial.

This paper describes the distinction between the concept of statistical power and the probability of getting a successful trial. While one can choose a very high statistical power to detect a certain treatment effect, the high statistical power does not necessarily translate to a high success probability if the treatment effect to detect is based on the perceived ability of the drug candidate. The crucial factor hinges on our knowledge of the drug's ability to deliver the effect used to power the study. The paper discusses a framework to calculate the 'average success probability' and demonstrates how uncertainty about the treatment effect could affect the average success probability for a confirmatory trial. It complements an earlier work by O'Hagan et al. (Pharmaceutical Statistics 2005; 4:187-201) published in this journal. Computer codes to calculate the average success probability are included.

Algorithms↗

The fallacy of enrolling only high-risk subjects in cancer prevention trials: is there a "free lunch"?

BACKGROUND: There is a common belief that most cancer prevention trials should be restricted to high-risk subjects in order to increase statistical power. This strategy is appropriate if the ultimate target population is subjects at the same high-risk. However if the target population is the general population, three assumptions may underlie the decision to enroll high-risk subject instead of average-risk subjects from the general population: higher statistical power for the same sample size, lower costs for the same power and type I error, and a correct ratio of benefits to harms. We critically investigate the plausibility of these assumptions. METHODS: We considered each assumption in the context of a simple example. We investigated statistical power for fixed sample size when the investigators assume that relative risk is invariant over risk group, but when, in reality, risk difference is invariant over risk groups. We investigated possible costs when a trial of high-risk subjects has the same power and type I error as a larger trial of average-risk subjects from the general population. We investigated the ratios of benefit to harms when extrapolating from high-risk to average-risk subjects. RESULTS: Appearances here are misleading. First, the increase in statistical power with a trial of high-risk subjects rather than the same number of average-risk subjects from the general population assumes that the relative risk is the same for high-risk and average-risk subjects. However, if the absolute risk difference rather than the relative risk were the same, the power can be less with the high-risk subjects. In the analysis of data from a cancer prevention trial, we found that invariance of absolute risk difference over risk groups was nearly as plausible as invariance of relative risk over risk groups. Therefore a priori assumptions of constant relative risk across risk groups are not robust, limiting extrapolation of estimates of benefit to the general population. Second, a trial of high-risk subjects may cost more than a larger trial of average risk subjects with the same power and type I error because of additional recruitment and diagnostic testing to identify high-risk subjects. Third, the ratio of benefits to harms may be more favorable in high-risk persons than in average-risk persons in the general population, which means that extrapolating this ratio to the general population would be misleading. Thus there is no free lunch when using a trial of high-risk subjects to extrapolate results to the general population. CONCLUSION: Unless the intervention is targeted to only high-risk subjects, cancer prevention trials should be implemented in the general population.

Confidence Intervals↗

Improving trial power through use of prognosis-adjusted end points.

BACKGROUND AND PURPOSE: The stroke patient population is heterogeneous, leading to wide variation in outcome caused by differences in age, initial severity, and presence of concomitant disease. Setting an identical recovery target for all patients in intervention trials may conceal individually important therapeutic treatment effects. Instead, a variable end point that takes severity or likely prognosis into account may be more informative. METHODS: We used data from the Glycine Antagonist in Neuroprotection (GAIN) International trial to assess statistical power of various primary end points for intervention trials. We selected prognosis-adjusted cut points based on Barthel Index (BI) or Rankin Scale (RS) using a prognostic model, or assigned a fixed end point within subgroups of patients defined by their Oxford category or National Institutes of Health Stroke Scale (NIHSS) score. We simulated a treatment effect and estimated statistical power with standard formulae. RESULTS: Assignment of end points using a prognostic model for individual patients increased statistical power, when compared with assigning end points using only the Oxford classification. For the BI, power was increased from 60% to 88% (equivalent to a 49% reduction in sample size if power remains unchanged). With the RS end points, power was increased from 84% to 92% (or a 24% reduction in sample size). Versus a fixed end point for all patients, model-based methods increased power by 22 percentage points for BI> or =95 and 14 percentage points for RS< or =1 (effective sample size reductions 43% and 34%). CONCLUSIONS: Prognosis-adjusted end points can increase statistical power compared with fixed end points. Assessment is based on realistic goals for individual patients and yet trial results remain generalizable.

Age Factors↗

Attributions and depression: why is the literature so inconsistent?

A large body of literature examining the relations between depression and causal attributions has produced inconsistent findings. Many studies have clearly had inadequate statistical power, however, so that negative findings cannot be readily interpreted. In this review, statistical power was computed for all published analyses relating depression to attributions to any of the following: internal, stable, or global causes, or their composite, ability/character, effort/behavior, luck, or task difficulty. On average, the power of these analyses was very poor. For example, only 8 of the 87 analyses had a probability of .80 or better of detecting a small-medium true population effect (e.g., r = .20). Separating studies by levels of power helped to clarify the inconsistencies in the literature. Whereas across all published studies depression was fairly consistently related only to the composite of internal, stable, and global attributions, those few studies with fairly high power all reported significant relations of depression to stable and global attributions as well as to the composite. It is suggested that increased attention be paid to the power of statistical analyses in planning studies and in drawing conclusions from completed studies.

Cognition↗

The power and statistical behaviour of allele-sharing statistics when applied to models with two disease loci.

We have evaluated the power for detecting a common trait determined by two loci, using seven statistics, of which five are implemented in the computer program SimWalk2, and two are implemented in GENEHUNTER. Unlike most previous reports which involve evaluations of the power of allele-sharing statistics for a single disease locus, we have used a simulated data set of general pedigrees in which a two-locus disease is segregating and evaluated several nonparametric linkage statistics implemented in the two programs. We found that the power for detecting linkage using the S(all) statistic in GENEHUNTER (GH, version 2.1), implemented as statistic E in SimWalk2 (version 2.82), is different in the two. The P values associated with statistic E output by SimWalk2 are consistently more conservative than those from GENEHUNTER except when the underlying model includes heterogeneity at a level of 50% where the P values output are very comparable. On the other hand, when the thresholds are determined empirically under the null hypothesis, S(all) in GENEHUNTER and statistic E have similar power.

Alleles↗

First measured plasma concentration value as C(max); impact on the C(max) confidence interval in bioequivalence studies.

In bioequivalence studies, the first blood or plasma sample taken after dosing sometimes yields a higher assayed drug concentration than any samples drawn thereafter. This circumstance ('first C(max)' or 'FCM'), is usually considered undesirable, since a 'true C(max)' requires that the sampled concentrations immediately preceding and immediately after the 'true C(max)' concentration should be lower than the 'true C(max)' concentration. Therefore, a question arises whether the presence of FCM in a bioequivalence study affects the power and accuracy of the computed statistical confidence interval (CI) for C(max). This study examines what effect, if any, the inclusion or exclusion of FCM data has on the statistical power and accuracy of the 90% CI computed for C(max) in the analysis of results for in vivo bioequivalence studies. Actual experimental study data as well as data from simulated studies were evaluated. In the simulated studies, up to half of the study subjects exhibited FCM, and various levels of intrasubject variability were incorporated into the absorption rate constant. The two one-sided tests procedure was used to assess equivalence of C(max) for the test versus reference products when either a complete set of C(max) data was analysed (designated 'CC(max)'), which included subjects with FCM profiles; or a truncated set of data was analysed (designated 'TC(max)'), that excluded all subjects with FCM. The results showed that the CC(max) metric had greater statistical power and comparable or greater statistical accuracy compared to TC(max) for both bioequivalent and non-bioequivalent drug product formulations. Even when up to 50% of the study subjects had FCM, the power and accuracy of the 90% CI for rate of absorption (i.e. C(max)) was not significantly affected. Consequently, this study shows that, in the analysis of data from conventional in vivo bioequivalence studies, the inclusion of 50% of the subjects exhibiting FCM does not greatly impact the statistical results obtained for C(max). Published in 2000 by John Wiley & Sons, Ltd.

Absorption↗

Observer studies involving detection and localization: modeling, analysis, and validation.

Although the receiver operating characteristic (ROC) paradigm is the accepted method for evaluation of diagnostic imaging systems, it has some serious shortcomings inasmuch as it is restricted to one observer report per image. By contrast the free-response ROC (FROC) paradigm and associated analysis method allows the observer to report multiple abnormalities within each imaging study, and uses the location of reported abnormalities to improve the measurement. Because the ROC method cannot accommodate multiple responses or use location information, its statistical power will suffer. The FROC paradigm/analysis has not enjoyed widespread acceptance because of concern about whether responses made to the same diagnostic study can be treated as independent. We propose a new jackknife FROC analysis method (JAFROC) that does not make the independence assumption. The new analysis method combines elements of FROC and the Dorfman-Berbaum-Metz (DBM) methods. To compare JAFROC to an earlier free-response analysis method (specifically the alternative free-response, or AFROC method), and to the DBM method, which uses conventional ROC scoring, we developed a model for generating simulated FROC data. The simulation model is based on an eye-movement model of how experts evaluate images. It allowed us to examine null hypothesis (NH) behavior and statistical power of the different methods. We found that AFROC analysis did not pass the NH test, being unduly conservative. Both the JAFROC method and the DBM method passed the NH test, but JAFROC had more statistical power than the DBM method. The results of this comparison suggest that future studies of diagnostic performance may enjoy improved statistical power or reduced sample size requirements through the use of the JAFROC method.

Algorithms↗

An overview of variance inflation factors for sample-size calculation.

For power and sample-size calculations, most practicing researchers rely on power and sample-size software programs to design their studies. There are many factors that affect the statistical power that, in many situations, go beyond the coverage of commercial software programs. Factors commonly known as design effects influence statistical power by inflating the variance of the test statistics. The authors quantify how these factors affect the variances so that researchers can adjust the statistical power or sample size accordingly. The authors review design effects for factorial design, crossover design, cluster randomization, unequal sample-size design, multiarm design, logistic regression, Cox regression, and the linear mixed model, as well as missing data in various designs. To design a study, researchers can apply these design effects, also known as variance inflation factors to adjust the power or sample size calculated from a two-group parallel design using standard formulas and software.

Analysis of Variance↗

Overweight and pregnancy complications.

The association between increased prepregnancy weight for height and seven pregnancy complications was studied in a multi-racial sample of more than 4100 recent deliveries. Body mass indices were calculated and used to classify women as average weight (90-119 percent of ideal or BMI 19.21-25.60), moderately overweight (120-135 percent ideal or BMI 25.61-28.90), and very overweight (greater than 135 percent ideal or BMI greater than 28.91) prior to pregnancy. Compared to women of average weight for height, very overweight women had a higher risk of diabetes, hypertension, pregnancy-induced hypertension and primary cesarean section delivery. Moderately overweight women were also at higher risk than average for diabetes, pregnancy-induced hypertension and primary cesarean deliveries but the relative risks were of a smaller magnitude than for very overweight women. With women of average prepregnancy body mass as reference, moderately elevated, but not significant relative risks were found for perinatal mortality in the very overweight group, for urinary tract infections in both overweight groups, and a decreased risk for anemia was found in the very overweight group. However, post-hoc power analyses indicated that the number of overweight women in the sample did not allow adequate statistical power to detect these small differences in risk. To overcome limitations associated with low statistical power, the results of three recent studies of these outcomes in very overweight pregnant women were combined and summarized using Mantel-Haenzel techniques. This second, larger analysis suggested that very overweight women are at significantly higher risk for all seven outcomes studied. Summary results for moderately overweight women could not be calculated, since only two of the studies had evaluated moderately overweight women separately. These latter results support other findings that both moderate overweight and very overweight are risk factors during pregnancy, with the highest risk occurring in the heaviest group. Although these results indicate that moderate overweight is a risk factor during pregnancy, additional studies are needed to confirm the impact of being 20-35 percent above ideal weight prior to pregnancy. The results of this analysis also imply that since the baseline incidence of many perinatal complications is low, studies relating overweight and pregnancy complications should include large enough samples of overweight women so that there is adequate statistical power to reliably detect differences in complication rates.

Adolescent↗

Negative clinical trials in cystic fibrosis research.

The statistical power of 61 negative clinical trials of therapeutic regimens in patients with cystic fibrosis published from 1977 through 1988 was reviewed and the ability of the investigations to detect small, medium, and large standardized differences was calculated. Small, medium, and large standardized differences were defined as ratios of 0.2, 0.5, and 0.8, respectively, of the observed difference compared with the standard deviation. The average numbers (+/- SD) of patients in the treatment and control groups were 14.3 +/- 6.9 and 14.5 +/- 7.9, respectively. None of the studies had 80% power to detect a small or medium standardized difference and only 4 of the reports had 80% statistical power to detect a large standardized difference. The variability of cystic fibrosis causes a decrease in the standardized difference, making it more difficult to demonstrate statistical significance. Statistical power of negative clinical trials reported in the literature deserves more attention from investigators as well as physicians who treat patients with cystic fibrosis.

Clinical Trials as Topic↗

The power of tests for bioequivalence in feed experiments with poultry.

Several studies have compared the feeding of genetically modified (GM) grains and conventional grains to poultry. The general conclusion has been that there were no significant differences detected in the biological performance of the birds (i.e., the grains were bioequivalent). However, the question has been posed whether the experimental designs used in the studies had sufficient statistical power to detect treatment differences. The power of tests can be used to determine the ability of an experimental design to detect treatment differences. The definition of statistical power is the probability of rejecting the null hypothesis when it is false and should be rejected. The complement of statistical power is the Type II error (beta). That is, accepting the null hypothesis that there is no difference in treatments when there is one. A priori power analysis can indicate the probability at which the sampling regimen or experiment can actually detect an effect if a difference exists. Post hoc power analysis indicates the sufficiency or the sample size needed for an experiment that has already been conducted. In the current study, the power of tests for experiments published in the literature where significant and nonsignificant differences were reported between control birds and birds fed new feed grains was examined. With some exceptions, the power of tests is rarely formally considered or mentioned in poultry research. The results of the survey of the literature showed, in general, low power of statistical tests for feeding experiments involving non-GM grains or in those cases when GM and non-GM grains were compared in poultry feeding experiments. These results suggest that care needs to be taken when designing experiments for bioequivalence of grains fed to poultry.

Animal Feed↗

An alternative method for sample size determination in substance misuse prevention research.

There is considerable evidence that social science researchers often fail to adequately address statistical power and related sample size issues. This tendency has been particularly salient in substance misuse prevention research. Although failure to carefully address statistical power and sample size issues most frequently results in studies lacking statistical power, other problems can also occur. For example, sample size determination in controlled studies which focus on rates of substance initiation and similar measures require close consideration of baseline rates indicated by these measures. Otherwise, when conventional procedures for sample size determination are utilized, samples in these studies can exceed the size necessary for predesignated levels of power. This article presents a simple alternative to conventional procedures for sample size determination which can be applied to controlled study of substance initiation and similar outcomes.

Alcoholism↗

The use of percentage change from baseline as an outcome in a controlled trial is statistically inefficient: a simulation study.

BACKGROUND: Many randomized trials involve measuring a continuous outcome - such as pain, body weight or blood pressure - at baseline and after treatment. In this paper, I compare four possibilities for how such trials can be analyzed: post-treatment; change between baseline and post-treatment; percentage change between baseline and post-treatment and analysis of covariance (ANCOVA) with baseline score as a covariate. The statistical power of each method was determined for a hypothetical randomized trial under a range of correlations between baseline and post-treatment scores. RESULTS: ANCOVA has the highest statistical power. Change from baseline has acceptable power when correlation between baseline and post-treatment scores is high;when correlation is low, analyzing only post-treatment scores has reasonable power. Percentage change from baseline has the lowest statistical power and was highly sensitive to changes in variance. Theoretical considerations suggest that percentage change from baseline will also fail to protect from bias in the case of baseline imbalance and will lead to an excess of trials with non-normally distributed outcome data. CONCLUSIONS: Percentage change from baseline should not be used in statistical analysis. Trialists wishing to report this statistic should use another method, such as ANCOVA, and convert the results to a percentage change by using mean baseline scores.

Computer Simulation↗

Large trials vs meta-analysis of smaller trials: how do their results compare?

OBJECTIVE: To evaluate the results of large clinical trials vs the pooled results of smaller trials. DATA IDENTIFICATION: Meta-analyses with at least 1 "large" study were identified from the Cochrane Pregnancy and Childbirth Database and from MEDLINE (1966-1995). STUDY SELECTION: We used a sample size approach to select 79 meta-analyses with at least 1 large study of 1000 or more patients. We used a statistical power approach to select 61 meta-analyses with at least 1 large study based on statistical power considerations. DATA EXTRACTION: The outcome of interest for each meta-analysis was the primary one stated in the original publication or, when not clearly specified, was decided on clinically. DATA SYNTHESIS: By random effects calculations, we found agreement between large and smaller trials in 90% of the meta-analyses selected by the sample size approach and in 82% of the meta-analyses selected by the statistical power approach. Twice as many disagreements appeared when the variability among large studies and among smaller studies was not considered (ie, fixed effects calculations). Of the 15 disagreements between results of large and smaller trials using the random effects model, plausible explanations were identified in 10 meta-analyses: 5 with differences in the control rate of events between large and smaller trials, 4 with specific protocol or study differences, and 1 with potential publication bias. Two other disagreements were not clinically important, and tentative reasons could be identified for 2 of the remaining 3 disagreements. CONCLUSIONS: Results of smaller studies are usually compatible with the results of large studies, but discrepancies do occur even when the diversity among both large studies and smaller studies is considered. Clinically important differences without a potential explanation are extremely uncommon. Future research should further examine sources of heterogeneity between the results of large and smaller trials.

Bias↗

Analysis of statistical tests to compare visual analog scale measurements among groups.

BACKGROUND: A common type of study performed by anesthesiologists determines the effect of an intervention on pain reported by groups of patients. The goal of this study was to evaluate the effectiveness of t, analysis of variance (ANOVA), Mann-Whitney, and Kruskal-Wallis tests to compare visual analog scale (VAS) measurements between two or among three groups of patients. These results may be particularly helpful during the design of studies that measure pain with a VAS. METHODS: One VAS measurement was obtained from each of 480 nulliparous women in labor who were receiving oxytocin (149), nalbuphine (159), or epidural bupivacaine (172). Multiple simulated samples were then drawn from these data. These simulated samples were used in computer simulations of clinical trials comparing VAS measurements among groups. t and ANOVA tests were performed before and after an arcsin transformation was used, to make the data closer to a normal distribution. VAS measurements were also compared after they were divided into five ranked categories. RESULTS: The statistical distributions of VAS measurements were not normal (P < 10(-7)). Arcsin transformation made the distributions closer to normal distributions. Nevertheless, no statistical test incorrectly suggested that a difference existed among groups, when there was no difference, more often than the expected rate. t or ANOVA tests had a slightly greater statistical power than the other tests to detect differences among groups. Because arcsin transformation both decreased differences among means and reduced the variance to a lesser extent, it decreased power to detect differences among groups. Statistical power to detect differences among groups was not less for a five-category VAS than for a continuous VAS. CONCLUSIONS: We conclude that t and ANOVA, without an accompanying arcsin transformation, are good tests to find differences in VAS measurements among groups.

Analgesia, Epidural↗

Tobacco-specific nitrosamines in tobacco from U.S. brand and non-U.S. brand cigarettes.

Tobacco-specific nitrosamines (TSNAs) are one of the major classes of carcinogens found in tobacco products. As part of collaborative efforts to reduce tobacco use and resulting disease, the U.S. Centers for Disease Control and Prevention (CDC) carried out a two-phase investigation into the worldwide variation of the levels of TSNAs in cigarette tobacco. In the first phase, representatives of the World Health Organization (WHO) purchased cigarettes; scientists from the CDC subsequently measured the levels of TSNAs in tobacco from 21 different countries. Although the data collected from this initial survey suggested that globally marketed U.S.-brand cigarettes typically had higher TSNA levels than locally popular non-U.S. cigarettes in many countries, the number of samples limited the statistical power of the study. To improve statistical power and to ensure adequate sampling, the CDC conducted a second survey of 14 countries. In addition to the United States, the CDC selected the world's 10 most populous countries and three additional countries, so that at least two countries from each of the six WHO regions were represented. For each country, the CDC compared 15 packs of Marlboro cigarettes, which is the world's most popular brand of cigarettes, with 15 packs of a locally popular non-U.S. brand in the study country. Marlboro cigarettes purchased in 11/13 foreign countries had significantly higher tobacco TSNA levels than the locally popular non-U.S. brands purchased in the same country. The findings suggest that TSNA levels in tobacco can be substantially reduced in some cigarettes.

Carcinogens↗

Strengthening acute stroke trials through optimal use of disability end points.

BACKGROUND AND PURPOSE: Suboptimal choices of primary end point for acute stroke trials may have contributed to inconclusive results. The Barthel Index (BI) and Rankin Scale (RS) have been widely used and analyzed in various ways. We sought to investigate the most powerful end point for use in acute stroke trials. METHODS: Data from the Glycine Antagonist in Neuroprotection (GAIN) International Trial were used to simulate 24 000 clinical trials exploring various patterns and magnitudes of treatment effect and thus to estimate the statistical power for a range of end points based on the BI or RS. RESULTS: RS end points were more powerful than BI end points. End points dichotomized toward the favorable extreme of either scale or adjusted according to baseline prognosis ("patient-specific" end point) were among the most powerful. Combining RS and BI in a "global" end point was also successful. Improvements in statistical power indicated that using a RS end point instead of BI > or =60 could reduce the sample size by up to 84% (95% CI, 80% to 87%), 73% (95% CI, 68% to 79%) for a patient-specific BI end point, or 81% (95% CI, 76% to 85%) for a global end point. CONCLUSIONS: The RS and global end points are preferable to BI end points; the position of the cut point is also important. Better choices of end point substantially strengthen trial power for a given trial size or allow reduced sample sizes without loss of statistical power.

Acute Disease↗

Correction for intracranial volume in analysis of whole brain atrophy in multiple sclerosis: the proportion vs. residual method.

Two techniques that correct (normalize) regional and whole brain volumes according to head size-the proportion method (tissue-to-intracranial volume ratio) and the residual method (regression-based predicted brain tissue volumes)-are used pervasively in neuroimaging research, but have received little critical evaluation or direct comparison. Using a quantitatively derived MRI data set of patients with multiple sclerosis (n = 18) and age-/sex-matched normal controls (n = 18), we introduced various types of error into estimates of intracranial volume (ICV) and absolute parenchymal volume (APV) to observe how this error affected the final outcome of normalized brain measures and their ability to detect group differences, as computed by a proportion (brain parenchymal fraction [BPF]) and residual method (predicted parenchymal volume [PPV]). The results indicated that systemic error in ICV and APV values considerably affected BPF means based on the proportion method, except with dependent-related systematic APV error, but essentially did not change statistical power associated with group differences in BPF. Random error altered BPF means to a much smaller extent, but was associated with moderate reductions in statistical power. On the other hand, PPV estimates based on the residual method were unaffected by these same ICV and APV errors, except with dependent-related systematic APV error, and were not associated with reductions in statistical power. Our findings suggest that head size correction of brain regions with the residual method generally may provide advantages over the proportion method.

Adult↗