Search PubMedSearch

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Increasing scientific power with statistical power.

A survey of basic ideas in statistical power analysis demonstrates the advantages and ease of using power analysis throughout the design, analysis, and interpretation of research. The power of a statistical test is the probability of rejecting the null hypothesis of the test. The traditional approach to power involves computation of only a single power value. The more general power curve allows examining the range of power determinants, which are sample size, population difference, and error variance, in traditional ANOVA. Power analysis can be useful not only in study planning, but also in the evaluation of existing research. An important application is in concluding that no scientifically important treatment difference exists. Choosing an appropriate power depends on: a) opportunity costs, b) ethical trade-offs, c) the size of effect considered important, d) the uncertainty of parameter estimates, and e) the analyst's preferences. Although precise rules seem inappropriate, several guidelines are defensible. First, the sensitivity of the power curve to particular characteristics of the study, such as the error variance, should be examined in any power analysis. Second, just as a small type I error rate should be demonstrated in order to declare a difference nonzero, a small type II error should be demonstrated in order to declare a difference zero. Third, when ethical and opportunity costs do not preclude it, power should be at least .84, and preferably greater than .90.

Data Interpretation, Statistical

Statistical power in psychiatric research.

Statistical power is neglected in much psychiatric research, with the consequence that many studies do not provide a reasonable chance of detecting differences between groups if they exist in the population. This paper attempts to improve current practice by providing an introduction to the essential quantities required for performing a power analysis (sample size, effect size, type 1 and type 2 error rates). We provide simplified tables for estimating the sample size required to detect a specified size of effect with a type 1 error rate of alpha and a type 2 error rate of beta, and for estimating the power provided by a given sample size for detecting a specified size of effect with a type 1 error rate of alpha. We show how to modify these tables to perform power analyses for multiple comparisons in univariate and some multivariate designs. Power analyses for each of these types of design are illustrated by examples.

Clinical Trials as Topic

Effect of crossover on the statistical power of randomized studies.

Randomized studies involving long-term follow-up are vulnerable to the effects of unplanned crossover. In surgical studies, such crossover usually occurs when control patients become more symptomatic and undergo operation. In several large studies of coronary bypass grafting, crossover ranged from 25% to 38%. The most common way of dealing with this problem is to apply the "intention-to-treat" principle, which analyzes such crossovers with their originally assigned groups. Besides the logical problem of counting a control patient who actually undergoes operation as "nonsurgical," a more subtle problem arises in terms of statistical power. When statistical power is low, a truly effective treatment may be mistakenly labeled as no better than control, causing a potentially valuable form of therapy to be ignored or discarded. This analysis demonstrates that crossover may have a profound effect on the statistical power of randomized studies and presents a method for predicting the effect of such crossover on statistical power.

Coronary Artery Bypass

Statistical power in physical anthropology: a technical report.

A statistical power analysis of The American Journal of Physical Anthropology (Volume 44, 1976) was conducted. Twenty-five articles, which included 3,304 major significance tests, constituted the final sample. Resultant power estimates of 0.38, 0.62, and 0.81, corresponding to small, medium, and large population effects respectively, were obtained. Although the medium effect size estimate falls short of the recommended 0.80 level, the statistical power of physical anthropological research fares well relative to several of the social scientific fields of inquiry.

Anthropology, Physical

Dichotomizing continuous outcome variables: dependence of the magnitude of association and statistical power on the cutpoint.

Dichotomizing a continuous outcome variable casts that variable in traditional epidemiologic terms (that is, disease, no disease). One consequence is overall reduced statistical power. A more fundamental concern is that the magnitude of various measures of association (for example, prevalence ratio, odds ratio) and statistical power depend on the cutpoint used to dichotomize the variable. The phenomenon is illustrated with a hypothetical situation assuming a two-level predictor variable and a normally distributed outcome variable. As the cutpoint is increased from lower to higher values, the prevalence ratio increases steadily, the odds ratio is described by a U-shaped curve, and statistical power is described by an inverted U-shaped curve. Furthermore, the extent of these effects depends on the difference between the means of the continuous outcome variable for the two levels of the predictor variable. An empirical example is given using data on education and blood pressure (dichotomized to create a high blood pressure vs low blood pressure variable). Except at each end of the distribution, the results follow the hypothetical example. The observation has implications for public health and medical treatment; different cutpoints should be examined to determine the optimal cutpoint in terms of policy and/or treatment decisions. The observation described here also has implications for statistical interpretation; statements about the magnitude of association or statistical significance have limited meaning unless both the cutpoint and the distribution of the outcome variable are specified.

Bias

The effect of trial size on statistical power.

Many research studies produce results that falsely support a null hypothesis due to a lack of statistical power. The purpose of this research was to demonstrate selected relationships between single subject (SS) and group analyses and the importance of data reliability (trial size) on results. A computer model was developed and used in conjunction with Monte Carlo procedures to study the effects of sample size (subjects and trials), within- and between-subject variability, and subject performance strategies on selected statistical evaluation procedures. The inherent advantages of the approach are control and replication. Selected results are presented in this paper. Group analyses on subjects using similar performance strategies identified 10, 5, and 3 trials for sample sizes of 5, 10, and 20, respectively, as necessary to achieve statistical power values greater than 90% for effect sizes equal to one standard deviation of the condition distribution. SS analyses produced results exhibiting considerably less power than the group results for corresponding trial sizes, indicating how much more difficult it is to detect significant differences using a SS design. These results should be of concern to all investigators especially when interpreting nonsignificant findings.

Data Interpretation, Statistical

Costs and statistical power associated with five methods of collecting occupation exposure information for population-based case-control studies.

The ascertainment of information on past occupational exposure of study subjects is perhaps the main problem in case-control studies of occupational risk factors. Several methods have been proposed and used but little is known of their relative merits. The present study, undertaken in the context of a large ongoing case-control study of occupational cancer in Montreal, was designed to compare the costs of and statistical power to be derived from five plausible methods of data collection: 1) job titles abstracted from routine records, 2) job titles abstracted from routine records and processed through a job exposure matrix to derive exposure data, 3) job titles obtained by interview, 4) job titles obtained by interview and processed through a job exposure matrix to derive exposure data, and 5) job descriptions obtained by interview and processed by a team of experts to derive exposure data. Statistical power of the five methods was derived for 160 hypothetical risk factors, partly on the basis of empirical data from the data set and partly on the basis of some theoretical constructs. The design based on interview and expert evaluation was used as a reference, and the degree of misclassification of other methods was estimated in relation to this reference. For fixed sample size the interview and expert evaluation design was estimated to be much more costly than the others, but it provides much greater statistical power for detecting risks. Under the conditions of this investigation, this design was the most cost-effective. However, it is not clear to what extent this finding is generalizable.

Case-Control Studies

Issue of statistical power in comparative evaluations of minimal and intensive controlled drinking interventions.

An analysis of recent studies of minimal and intensive cognitive-behavioural treatments for problem drinking was undertaken to decide to whether a lack of statistical power explains the failure of the majority of studies to find a difference in outcome between these two types of treatment. Although the sample sizes have typically been small (n = 12-21), the analysis suggests that low statistical power is unlikely to be the explanation for the majority of null findings. It seems more likely that the difference in outcome between one positive study and the majority of null results reflects some combination of differences in the type of clients who were treated, the therapists' experience, and the type of intensive therapy that was provided. The low power of these studies demonstrates the desirability of researchers calculating the sample size required to an effect before commencing an outcome study. If they continue to undertake studies with small sample sizes, then they should refrain from inferring that the failure to reject a null hypothesis means that there is no difference between treatments.

Alcohol Drinking

Considerations of statistical power in the shared-haplotypes test.

This paper considers statistical power properties of a test of whether a disease is caused by a recessive gene, using data on HLA sharing properties of affected sibs. It is found that the test has very poor power characteristics, and in particular we are sometimes just as likely to accept the hypothesis that the gene is recessive when it is dominant as when it is truly recessive.

Child

Familial and sporadic schizophrenia. A simulation study of statistical power.

The importance of genetic factors in schizophrenia is clear but the mechanism involved remains obscure. Etiological heterogeneity may be responsible. Recently there has been interest in a putative distinction between genetic and environmental forms of the illness based on a positive or negative family history of the disorder. Those with a positive family history are classified as 'familial' and are considered to be more likely to have the genetic form of the illness. Those with a negative family history are classified as 'sporadic' and considered more likely to have an environmental form of the illness. This paper reports the results of a Monte Carlo simulation study with varying rates of misclassification to determine the statistical power of comparisons between familial and sporadic groups. For a large sample (n = 175) statistical power was moderate to good for effect sizes greater than or equal to 1.0 standard deviation unit and positive predictive value of 0.3 or greater.

Humans

Prognosis in valvular heart disease. I. Description of purpose, organization, data collection techniques, estimates of statistical power, and criteria for termination of patient entry. VA Cooperative Study Group on Valvular Heart Disease.

This report describes the design of a multicenter study with two major goals: the identification of valvular heart disease patients at risk for death or a serious complication, and the comparison of hemodynamic function and late outcome of a mechanical prosthetic valve (Björk-Shiley) with a bioprosthesis (Hancock porcine heterograft valve). Strengths of the study design are quantitative assessment of valvular and left ventricular function before and 6 months after valve replacement, measurement in a central laboratory of critical data items such as valve orifice area and left ventricular volumes, and assignment of cause of death and valve-related complications by a committee blinded to valve type. Statistical power calculated by a simulation technique shows only modest loss of power with frequent examination of outcome compared to infrequent examination. Guidelines for premature termination of the study because of superiority of one valve type are described; this includes a critical region with a sloping boundary, which allows for greater chance variation early in the study when the number of patients and events is small, gives the greatest statistical power, and yet appears to provide adequate protection for the subjects in the study.

Bioprosthesis

Semen analysis and fertility assessment in rabbits: statistical power and design considerations for toxicology studies.

Semen analysis is commonly used in evaluating human response to reproductive toxicants. Serial semen samples can be collected from rabbits and fertility assessed by artificial insemination, hence this species is potentially well suited for male reproductive toxicity studies that might be extrapolated to humans. However, the size and cost of rabbits often restricts the number of animals used, reducing the sensitivity of such studies. Therefore, it was of interest to optimize study design for semen analysis and fertility assessment in rabbits. Semen samples were collected weekly from sexually mature New Zealand white rabbits and a range of parameters was analyzed (Semen--pH, volume, osmolality; Sperm--number and concentration, morphology, viability, percentage motility, motion characteristics; Seminal plasma--fructose, citric acid, carnitine and protein concentrations, acid phosphatase activity). Male fertility was assessed by inseminating female rabbits with the minimum number of motile sperm required for normal fertility, determined to be one million. The within- and between-buck variabilities were determined for all parameters and used to calculate the statistical power of different study designs. The variability of sperm number and concentration was decreased when measured in four ejaculates collected within a short period of time rather than in a single ejaculate; this was not true of other endpoints measured. In addition, use of preexposure observations further increased the statistical power for all of the parameters. These data can be used to determine the optimum design for studies of male reproductive toxicity using rabbits, with particular regard to cost and the number of animals used.

Acid Phosphatase

Detecting disease clusters: the importance of statistical power.

A variety of methods and models have been proposed for the statistical analysis of disease excesses, yet rarely are these methods compared with respect to their ability to detect possible clusters. Evaluation of statistical power is one approach for comparing different methods. In this paper, the authors study the probability that a test will reject the null hypothesis, given that the null hypothesis is indeed false. They present a discussion of some considerations involved in power studies of cluster methods and review two methods for detecting space-time clusters of disease, one based on cell occupancy models and the other based on interevent distance comparisons. The authors compare these approaches with respect to: 1) the sensitivity to detect disease excesses (false negatives); 2) the likelihood of detecting clusters that do not exist (false positives); and 3) the structure of a cluster in a given investigation (the alternative hypothesis). The methods chosen, which are two of the most commonly used, are specific to different hypotheses. They both show low power for the small number of cases which are typical of citizen reports to health departments.

Clinical Protocols

The internal validity of efficacy studies: design and statistical power in studies of language therapy for aphasics.

In this study the internal validity of efficacy studies of language therapy for aphasic patients is discussed. The lack of sufficient internal validity is demonstrated with respect to research designs used in these studies and the statistical power of their statistical significance tests. The internal validity problems are viewed as the major cause of the conflicting conclusions in the efficacy studies in the past three decades.

Aphasia

Statistical power in nursing research.

A power analysis was performed on 62 articles that were published in Nursing Research and Research in Nursing and Health during 1989. The analysis revealed that when effects were small, the mean power of the statistical tests being performed to test research hypotheses was .26, indicating a very high risk of committing a Type II error. When effects were moderate, the mean power increased to .71, which is still below the conventionally acceptable power of .80. Only when a study involved large effects was the power adequate (mean of .95). Of the 583 power estimates calculated, 53% were for small effects. These analyses indicate that a substantial number of published nursing studies, and presumably even more of unpublished studies, have insufficient power to detect real effects, primarily because the samples used are too small.

Nursing

Simulation program for estimating statistical power of Cox's proportional hazards model assuming no specific distribution for the survival time.

Small sample properties of the maximum partial likelihood estimates for Cox's proportional hazards model depend on the sample size, the true values of regression coefficients, covariate structure, censoring pattern and possibly baseline hazard functions. Therefore, it would be difficult to construct a formula or table to calculate the exact power of a statistical test for the treatment effect in any specific clinical trial. The simulation program, written in SAS/IML, described in this paper uses Monte-Carlo methods to provide estimates of the exact power for Cox's proportional hazards model. For illustrative purposes, the program was applied to real data obtained from a clinical trial performed in Japan. Since the program does not assume any specific function for the baseline hazard, it is, in principle, applicable to any censored survival data as long as they follow Cox's proportional hazards model.

Clinical Trials as Topic

PoweREST: Statistical Power Estimation for Spatial Transcriptomics Experiments to Detect Differentially Expressed Genes Between Two Conditions.

Recent advancements in Spatial Transcriptomics (ST) have significantly enhanced biological research in various domains. However, the high cost of current ST data generation techniques restricts its application in large-scale population studies. Consequently, there is a pressing need to maximize the use of available resources to achieve robust statistical power. One fundamental question in ST analysis is to detect differentially expressed genes (DEGs) among different conditions using ST data. Such DEG analysis is often performed but the associated power calculation is rarely discussed in the literature. To address this gap, we introduce, PoweREST (https://github.com/lanshui98/PoweREST), a power estimation tool designed to support power calculation of DEG detection with 10X Genomics Visium data. PoweREST enables power estimation both before any ST experiments or after preliminary data are collected, making it suitable for a wide variety of power analyses in ST studies. We also provide a user-friendly, program-free web application (https://lanshui.shinyapps.io/PoweREST/), allowing users to interactively calculate and visualize the study power along with relevant the parameters.

Differentially expressed genes