Statistical power analysis of health, physical education, and recreation research.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The NCCOM3-R6 monitor continuously monitors cardiac output and five other cardiovascular variables from the thoracic electrical bioimpedance signal. We averaged data over 5-min intervals for 130 min in 100 control studies in 40 pediatric ICU patients, age 0.04 to 20.39 yr (median 1.39) and weighing 2.0 to 59.5 kg (median 8.8). For individual studies, 99% of the 5-min averages of cardiac output fell within +/- 44% of the baseline cardiac output for that study. Normal ranges were somewhat narrower for the other five variables. When we averaged data for 100 studies, 5-min interval observations for each variable did not deviate from baseline over a 2-h period (p greater than .70). With a sample size of 100 studies, we could detect a change in cardiac output of +/- 5% at the p less than .005 level with a power of 0.95. We conclude that with a sufficiently large sample size, studies employing the NCCOM3 can detect clinically significant cardiovascular changes due to pharmacologic or procedural stressors.
Failure to consider statistical power when achieving apparently "negative" results prevents accurate interpretation of the results. A nonsignificant result can be obtained when one includes an insufficient number of subjects to permit observation of a true effect (low power to detect an effect), or when one has an adequate number of subjects, but a meaningful effect does not exist (high power, no effect); one can also have a situation of lower power and no real effect. Without considering power, one is unable to distinguish a "negative" experiment from an inadequate one. This article examines 154 published nonsignificant t-test results. When power is calculated with an effect size equal to a standardized difference of unity, over 50% of the tests have inadequate power.
A discussion of the importance of statistical power in research is presented accompanied by nomograms for determining sample size and statistical power for the Student's paired and unpaired t tests with a Type I error of 5%. A brief review of statistical inference is presented. Some findings from Part I are reviewed.
Six different statistical methods for comparing limiting dilution assays were evaluated, using both real data and a power analysis of simulated data. Simulated data consisted of a series of 12 dilutions for two treatment groups with 24 cultures per dilution and 1,000 independent replications of each experiment. Data within each replication were generated by Monte Carlo simulation, based on a probability model of the experiment. Analyses of the simulated data revealed that the type I error rates for the six methods differed substantially, with only likelihood ratio and Taswell's weighted mean methods approximating the nominal 5% significance level. Of the six methods, likelihood ratio and Taswell's minimum Chi-square exhibited the best power (least probability of type II errors). Taswell's weighted mean test yielded acceptable type I and type II error rates, whereas the regression method was judged unacceptable for scientific work.
Regression analysis is often used to demonstrate associations among variables believed to be biologically related. Failure to demonstrate a "significant" relationship may be due to two factors: 1) the variables are truly unrelated, or 2) a relationship exists but goes undetected due to inadequate statistical power. Investigators must consider the second possibility since failure to detect a statistically significant relationship is often taken as evidence for no biological relationship. These issues are addressed in the context of the interrelationship between four features common to all statistical methods: the size of effect or relationship worth detecting, the Type I (alpha) error, the sample size, and the Type II (beta) error. An example derived from published data relating morphological characteristics of muscle fiber type and isokinetic strength performance illustrates the practical significance of this dilemma.
A survey of basic ideas in statistical power analysis demonstrates the advantages and ease of using power analysis throughout the design, analysis, and interpretation of research. The power of a statistical test is the probability of rejecting the null hypothesis of the test. The traditional approach to power involves computation of only a single power value. The more general power curve allows examining the range of power determinants, which are sample size, population difference, and error variance, in traditional ANOVA. Power analysis can be useful not only in study planning, but also in the evaluation of existing research. An important application is in concluding that no scientifically important treatment difference exists. Choosing an appropriate power depends on: a) opportunity costs, b) ethical trade-offs, c) the size of effect considered important, d) the uncertainty of parameter estimates, and e) the analyst's preferences. Although precise rules seem inappropriate, several guidelines are defensible. First, the sensitivity of the power curve to particular characteristics of the study, such as the error variance, should be examined in any power analysis. Second, just as a small type I error rate should be demonstrated in order to declare a difference nonzero, a small type II error should be demonstrated in order to declare a difference zero. Third, when ethical and opportunity costs do not preclude it, power should be at least .84, and preferably greater than .90.
The phenothiazines have exhibited a history of problems associated with the bioequivalence of solid oral dosage forms. The more recent availability of chemically equivalent forms of thioridazine has raised new and interesting questions about the appropriateness of generic product interchange, even among brands that have been designated "therapeutically equivalent" by the Food and Drug Administration. The scrutiny that has accompanied the consideration of thioridazine products for inclusion into various state generic substitution formularies has offered an opportunity to examine issues involving bioequivalency in considerable detail. Specific bioequivalency concerns relate to: correct analysis of drug in biological fluids; the importance of evaluating active metabolites: single-dose vs. multiple-dose crossover studies; appropriate statistical power analysis; the "70/70" rule, and comparison of product variabilities. Examples of problems are cited to illustrate that significant questions still remain about the appropriate factors that should be used to establish bioequivalency.
To assess the effects of drawing blood specimens from a site proximal to an intravenous (IV) infusion line, 24 volunteers were infused with approximately 30 mL of 5% (w/v) dextrose in 0.9% (w/v) NaCl solution. Specimens were drawn proximal to the IV line, while the solution was infusing, and at 1, 2, and 3 minutes after discontinuance of the infusion and assayed for glucose, sodium, chloride, and red blood cell (RBC) count. Control samples were obtained simultaneously from the opposite arm to determine any residual or dilutional effects from the infusing solution. Statistical analysis using the paired sample t-test showed no significant difference of RBCs between arms at or beyond 1 minute. Statistical power analysis reveals that there is a 95% level of certainty that there is less than 1% dilution of the test arm specimen. Analysis of sodium and chloride levels showed no contamination of the test arm specimen at 1 minute, but glucose concentrations still showed an average elevation over control of 5% at 3 minutes. The authors concluded that the drawing of blood specimens proximal to an IV infusion, 3 minutes after its discontinuance, has a clinically negligable dilutional effect but substances present at relatively high levels in the infused solution may still be detected.
There is a growing debate among researchers and practitioners concerning the validity of the Food and Drug Administration (FDA) standards for establishing generic interchange through assignment of "therapeutically equivalent" designations for noninnovator or generic drugs. This debate has particular significance for psychotropic drugs which are used in the management of patients with severely debilitating mental diseases. Thus, the controversy primarily focuses on the appropriateness of using the FDA's therapeutic equivalence designation as a criterion for interchange. This is particularly important in light of the lack of well-defined standards for bioequivalence studies and in view of the effect of mandatory substitution laws and financial incentives that encourage generic dispensing that can lead to frequent and indiscriminate interchange among multiple generic brands. Specific concerns include the validity of assay techniques for drug in biologic fluids, statistical power analysis, the appropriateness of the "70/70 rule," and the relevance of studies carried out in healthy normal volunteers in the determination of bioequivalence.
A statistical power analysis of The American Journal of Physical Anthropology (Volume 44, 1976) was conducted. Twenty-five articles, which included 3,304 major significance tests, constituted the final sample. Resultant power estimates of 0.38, 0.62, and 0.81, corresponding to small, medium, and large population effects respectively, were obtained. Although the medium effect size estimate falls short of the recommended 0.80 level, the statistical power of physical anthropological research fares well relative to several of the social scientific fields of inquiry.
The pharmacokinetics of intravenously administered theophylline were studied in five healthy nonsmokers. Each subject received 5 mg/kg of theophylline as aminophylline after an overnight fast and again after a standard high-fat meal. Although there was wide between-day variation in the elimination rate constant in three of the five subjects, no statistically significant differences were observed in area under the time-versus-concentration curve, maximum serum theophylline concentration, elimination rate constant, or apparent volume of distribution between the two treatments. A statistical power analysis indicated that if differences in volume of distribution and maximum serum theophylline concentration occur in the general population, the mean differences are less less than 15% and 20%, respectively. This suggests that alterations in intravascular drug distribution resulting from eating a high-fat meal do not contribute importantly to previously reported effects of food on serum theophylline concentrations after oral dosing.
This article describes a 3-year experience with focal neocortical ischemia in three rat strains. Multiple groups of adult Wistar (n = 50), Fisher 344 (n = 31), and spontaneously hypertensive (n = 72) rats were subjected to permanent occlusion of the distal middle cerebral (MCA) and ipsilateral common carotid arteries (CCA). Twenty-four hours later the animals were killed, and frozen brain sections were stained with hematoxylin and eosin to demarcate infarcted tissue. The infarct volume for each section was quantified with an image analyzer, and the total infarct volume was calculated with an iterative program that summed all interval volumes. Neocortical infarct volume was the largest and most reproducible in the spontaneously hypertensive rats (SHR). Statistical power analysis to project the numbers of animals necessary to detect a 25 or 50% change in infarct volume with alpha = 0.05 and beta = 0.2 revealed that only the SHR model was practical in terms of requisite animals: i.e., less than 10 animals per group. Tandem occlusion of the distal MCA and ipsilateral CCA in the SHR strain provides a surgically simple method for causing large neocortical infarcts with reproducible topography and volume. The interanimal variability in infarct volume that occurs even in the SHR strain dictates that randomized, concomitant controls are necessary in each study to ensure the accurate assessment of experimental manipulations or pharmacologic therapies.
This study evaluates infarct size measurement as an indicator of cerebral ischaemia outcome in a placebo-controlled trial of potential cerebral protection in the unilateral carotid artery ligation in the Mongolian gerbil. Ibuprofen was used in an effort to manipulate infarct size as this agent has been shown to reduce ischaemia in myocardial infarction. Using measurements obtained through an infarct-sizing technique and a statistical power analysis of the method, the sample sizes needed to obtain significant results were projected for this model. In this case, it was not possible to demonstrate an effect of ibuprofen on infarct size although a tendency towards larger infarct size in ibuprofen-treated compared with placebo-treated gerbils was observed (36.1 +/- 10.1% versus 30.0 +/- 17.5%). The sample sizes needed to find significant changes in infarct size indicate that this model finds a practical use in studying therapies which will alter infarct size by at least 50%. For example, to detect a 30% change in infarct size, 33 successfully infarcted gerbils per group would be needed, but a 50% change would require a more tenable 13 infarcted gerbils per group. However, given the 40% infarction rate of occluded gerbils seen in this study, almost 33 gerbils per group would be required to detect a 50% change. In addition, somatosensory evoked potential was compared with neurological examination as a predictor of infarction. It would be helpful to be able to pre-screen for infarcted gerbils immediately after occlusion in order to direct infarcted gerbils into control and treated groups. Somatosensory evoked potential successfully predicted infarction with a 90% accuracy in 21 gerbils compared with neurological evaluation which was 100% accurate. But the somatosensory evoked potential prediction was made within 15 min of occlusion as opposed to the 6 h of observation during which the neurological evaluation was made.
Although neuroendocrine adverse effects (NAEs) of antipsychotic agents in women have been widely reported, it has been generally assumed that either (1) tolerance to these effects develops with chronic use, (2) patients adjust to the effects, or (3) a trial of dopamine-agonist treatment is effective. We have begun to examine the prevalence of chronic adverse effects and their effect on compliance using a pilot study of self-reported NAEs, antipsychotic drugs, and compliance patterns in a naturalistic setting. Twenty chronic psychiatric outpatients who had been continuously prescribed antipsychotic agents for a minimum of six months were interviewed. The major finding is the greater antipsychotic dose exposure among those with self-reported NAEs compared with those without NAEs (781 +/- 606 chlorpromazine-equivalent mg/d vs. 125 +/- 117, p less than 0.001). High-potency agents were prescribed for all of the patients reporting amenorrhea and/or galactorrhea, although the relationship between potency group (high vs. low) and total neuroendocrine effects was not significant. Self-reported compliance was not significantly related to neuroendocrine adverse effects. However, a trend toward the association of self-reported galactorrhea and noncompliance (p = 0.08) is noted. The implications of these findings and a suggested approach for their replication in a more powerful statistical analysis is discussed.
Postural sway was measured in 12-14-month-old human infants and in adults while they were standing in the light and dark. Spectral density analyses conducted on all frequencies, at specific frequencies, and for individual subjects showed that infants generally did not sway significantly more in the dark than in the light, whereas adults did. For example, infants' dark/light sway proportions were 1.12 and 1.21 for the anterior-posterior and lateral dimensions, respectively, compared to adult values of 2.23 and 3.43 for one-footed stance, and 1.43 and 2.13 for two-footed stance. A statistical power analysis indicated that if the dark/light proportions for infants had been comparable to those for adults, significant differences could have been detected. These findings indicate that the early regulation of standing posture does not depend on the continuous availability of visual information.
The power of the association between oral contraceptives and breast cancer was analysed in all the papers published up to date. Seventy-seven publications (from 44 studies) were collected and graded as to quality using meta-analytical methods. Power achieved a figure of greater than or equal to 0.8 in a 10.8% of the associations studied. It showed a significant relationship with the existence of a significant relative risk of the oral contraceptives for breast cancer. The relationship with the sample size of a study was not linear. Power did not show any significant relationship to other variables related to the design of a study (apart from matching, being the power higher in unmatched studies), or to the biases detected, although studies considered as unbiased yielded a higher power. Logistic regression analysis included as predictors of a power greater than or equal to 0.80 the existence of a significant relative risk and the lack of biases in a research.
A novel multivariate statistical approach is presented for extracting and exploiting intrinsic information present in our ever-growing sequence data banks. The information extraction from the sequences avoids the pitfalls of intersequence alignment by analyzing secondary invariant functions derived from the sequences in the data bank rather than the sequences themselves. Such typical invariant function is a 20 x 20 histogram of occurrences of amino acid pairs in a given sequence or fragment thereof. To illustrate the potential of the approach an analysis of 10,000 protein sequences from the National Biomedical Research Foundation Protein Identification Resource is presented, whose analysis already reveals great biological detail. For example, zeta-hemoglobin is found to lie close to amphibian and fish chi-hemoglobin which, in turn, is an important clue to the physiological function of this mammalian early embryonic hemoglobin. The multivariate statistical framework presented unifies such apparently unrelated issues as phylogenetic comparisons between a set of sequences and distance matrices between the constituents of the biological sequences. The Multivariate Statistical Sequence Analysis (MSSA) principles can be used for a wide spectrum of sequence analysis problems such as: assignment of family memberships to new sequences, validation of new incoming sequences to be entered into the database, prediction of structure from sequence, discrimination of coding from non-coding DNA regions, and automatic generation of an atlas of protein or DNA sequences. The MSSA techniques represent a self-contained approach to learning continuously and automatically from the growing stream of new sequences. The MSSA approach is particularly likely to play a significant role in major sequencing efforts such as the human genome project.