Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Sexual expression and dementia. Views of caregivers: A pilot study.

OBJECTIVE: To measure the attitudes of health professionals in nursing homes towards sexuality and sexual expression in cognitively impaired and cognitively intact residents. DESIGN: Postal survey. PARTICIPANTS: The staff (administrators, clinicians, social workers and others) of 300 randomly selected nursing homes located in three states. Of these, 114 representatives responded. MAIN OUTCOME MEASURE: A measure of attitudes towards resident sexuality developed during a prior study. RESULTS: Results suggest that respondents held a generally positive orientation towards residents' sexual expression which was expressed with respect to cognitively impaired residents as well as to those who were cognitively intact. Possibly because of the small sample size and resulting low statistical power, statistical analyses failed to demonstrate any significant differences among the groups of residents: administrators, clinicians, social workers, and undifferentiated 'others'. However, while non-significant, there was a consistent tendency for administrators to be relatively more conservative than were the other groups. Almost all respondents agreed that additional staff training should focus specifically on dealing with resident sexual expression. CONCLUSIONS: Overall, the sample reported generally positive attitudes towards resident sexuality and sexual expression.

Aged↗

The power of statistical tests in meta-analysis.

Calculations of the power of statistical tests are important in planning research studies (including meta-analyses) and in interpreting situations in which a result has not proven to be statistically significant. The authors describe procedures to compute statistical power of fixed- and random-effects tests of the mean effect size, tests for heterogeneity (or variation) of effect size parameters across studies, and tests for contrasts among effect sizes of different studies. Examples are given using 2 published meta-analyses. The examples illustrate that statistical power is not always high in meta-analysis.

Humans↗

The power of a statistical test. What does insignificance mean?

In statistical testing of data, the p value is a standard measure for reporting quantitative results. When a significant difference is reported, (e.g., P less than .05), most readers understand that there is less than a 5% chance that the authors have made a type I error (false positive or alpha) with their conclusion. In contrast, when nonsignificant differences between treatments, groups, or parameters of interest are reported (e.g., P greater than .05), many investigators and readers incorrectly interpret the 95% confidence interval for this conclusion as a 95% chance of making the correct decision. In fact, the alpha level of significance (in this example, .05) is only one of the parameters that determines the probability of committing a type II error (false negative or beta) when concluding statistical insignificance. Statistical power is the probability of having made a correct decision when the statistical tests reveal insignificance (P greater than .05) and the null hypothesis is true. The higher the power, the greater the chance that the decision is correct. Power depends on the alpha level of significance, the sample size, the standard deviation of the population or the sample, and the magnitude of the difference the investigators are trying to demonstrate.

Probability↗

The power of analysis: statistical perspectives. Part 1.

Failure to consider statistical power when achieving apparently "negative" results prevents accurate interpretation of the results. A nonsignificant result can be obtained when one includes an insufficient number of subjects to permit observation of a true effect (low power to detect an effect), or when one has an adequate number of subjects, but a meaningful effect does not exist (high power, no effect); one can also have a situation of lower power and no real effect. Without considering power, one is unable to distinguish a "negative" experiment from an inadequate one. This article examines 154 published nonsignificant t-test results. When power is calculated with an effect size equal to a standardized difference of unity, over 50% of the tests have inadequate power.

Clinical Trials as Topic↗

Statistical methods for analysing Barthel scores in trials of poststroke interventions: a review and computer simulations.

BACKGROUND AND PURPOSE: Arguments persist as to whether parametric or non-parametric methods should be used to analyse ordinal data in trials. This paper aims to assess methods used for presenting and analysing an ordinal scale, the Barthel Index, in trials of poststroke interventions. METHODS: All randomized controlled trials (RCTs) of poststroke interventions published from 1995 to 2004 in two journals (Stroke and Clinical Rehabilitation) were scrutinized for methods used to present and analyse Barthel scores. Computer simulations were used to compare the type I errors and the statistical power of different statistical methods under a range of assumed circumstances. RESULTS: One hundred and fifty-six RCTs were identified within the two journals. The central tendency of Barthel scores was measured by the median in 47 trials and by the mean in 35 trials. Non-parametric analyses of Barthel scores were conducted in 47 trials and parametric methods used in 18 trials. The results of computer simulations demonstrate that the t-test has a similar type I error rate and statistical power when compared with the rank sum test. However, when a zero final Barthel score is assigned to patients who have died, the statistical power of the t-test is much reduced. The possible maximal statistical power of dichotomization and ordinal regression is usually much lower than that of the rank sum test. CONCLUSIONS: To facilitate comparison and meta-analysis, we recommend that mean values (with standard deviations or standard errors) of Barthel scores should be routinely reported in trials of poststroke interventions. The rank sum test appears the most powerful inferential technique for detecting differences in Barthel scores.

Computer Simulation↗

Seven ways to increase power without increasing N.

Many readers of this monograph may wonder why a chapter on statistical power was included. After all, by now the issue of statistical power is in many respects mundane. Everyone knows that statistical power is a central research consideration, and certainly most National Institute on Drug Abuse grantees or prospective grantees understand the importance of including a power analysis in research proposals. However, there is ample evidence that, in practice, prevention researchers are not paying sufficient attention to statistical power. If they were, the findings observed by Hansen (1992) in a recent review of the prevention literature would not have emerged. Hansen (1992) examined statistical power based on 46 cohorts followed longitudinally, using nonparametric assumptions given the subjects' age at posttest and the numbers of subjects. Results of this analysis indicated that, in order for a study to attain 80-percent power for detecting differences between treatment and control groups, the difference between groups at posttest would need to be at least 8 percent (in the best studies) and as much as 16 percent (in the weakest studies). In order for a study to attain 80-percent power for detecting group differences in pre-post change, 22 of the 46 cohorts would have needed relative pre-post reductions of greater than 100 percent. Thirty-three of the 46 cohorts had less than 50-percent power to detect a 50-percent relative reduction in substance use. These results are consistent with other review findings (e.g., Lipsey 1990) that have shown a similar lack of power in a broad range of research topics. Thus, it seems that, although researchers are aware of the importance of statistical power (particularly of the necessity for calculating it when proposing research), they somehow are failing to end up with adequate power in their completed studies. This chapter argues that the failure of many prevention studies to maintain adequate statistical power is due to an overemphasis on sample size (N) as the only, or even the best, way to increase statistical power. It is easy to see how this overemphasis has come about. Sample size is easy to manipulate, has the advantage of being related to power in a straight-forward way, and usually is under the direct control of the researcher, except for limitations imposed by finances or subject availability. Another option for increasing power is to increase the alpha used for hypothesis-testing but, as very few researchers seriously consider significance levels much larger than the traditional .05, this strategy seldom is used. Of course, sample size is important, and the authors of this chapter are not recommending that researchers cease choosing sample sizes carefully. Rather, they argue that researchers should not confine themselves to increasing N to enhance power. It is important to take additional measures to maintain and improve power over and above making sure the initial sample size is sufficient. The authors recommend two general strategies. One strategy involves attempting to maintain the effective initial sample size so that power is not lost needlessly. The other strategy is to take measures to maximize the third factor that determines statistical power: effect size.

Data Interpretation, Statistical↗

Incorporating individual error rate into association test of unmatched case-control design.

OBJECTIVES: Genotyping error commonly occurs and could reduce the power and bias statistical inference in genetics studies. In addition to genotypes, some automated biotechnologies also provide quality measurement of each individual genotype. We studied the relationship between the quality measurement and genotyping error rate. Furthermore, we propose two association tests incorporating the genotyping quality information with the goal to improve statistical power and inference. METHODS: 50 pairs of DNA sample duplicates were typed for 232 SNPs by BeadArray technology. We used scatter plot, smoothing function and generalized additive models to investigate the relationship between genotype quality score (q) and inconsistency rate (ĩ) among duplicates. We constructed two association tests: (1) weighted contingency table test (WCT) and (2) likelihood ratio test (LRT) to incorporate individual genotype error rate (epsilon(i)), in unmatched case-control setting. RESULTS: In the 50 duplicates, we found q and ĩ were in strong negative association, suggesting the genotypes with low quality score were more likely to be mistyped. The WCT improved the statistical power and partially corrects the bias in point estimation. The LRT offered moderate power gain, but was able to correct the bias in odds ratio estimation. The two new methods also performed favorably in some scenarios when epsilon(i) was mis-specified. CONCLUSIONS: With increasing number of genetic studies and application of automated genotyping technology, there is a growing need to adequately account for individual genotype error rate in statistical analysis. Our study represents an initial step to address this need and points out a promising direction for further research.

Case-Control Studies↗

Efficacy of pressure topical anaesthesia in punctal occlusion by diathermy.

AIMS: To prospectively compare the efficacy and safety of pressure topical anaesthesia in punctal occlusion by using cautery in the treatment of dry eye syndrome (DES) with that of conventional treatment by using needle injection of anaesthetic agents. METHODS: In a randomised controlled trial, 18 consecutive adult patients with DES requiring punctal occlusion were recruited over a 10 month period. Consenting patients were randomised into two groups. Group A patients received pressure topical anaesthesia in the right eye followed by injection anaesthesia in the left eye. Group B was vice versa. Punctal occlusion using cautery was performed in each eye after a specified time following the application of anaesthesia. The main outcome measures were the pain experienced during application of anaesthesia and that during punctal occlusion. RESULTS: 36 eyes of 18 patients were randomised to receive injection anaesthesia in one eye and pressure topical anaesthesia in the other. Nine patients (nine females) were in group A and nine patients (seven females, two males) in group B. The mean age of group A patients was 45.3 (SD 13.5) years, and that of group B patients was 55.6 (12.6) years. The two groups were comparable in terms of mean age (p=0.117) and mean pain score for pressure topical anaesthesia application (p=0.612), injection anaesthesia application (p=0.454), diathermy in pressure anaesthetised eyes (p=0.113), and diathermy in injection anaesthetised eyes (p=0.289). Paired t test was used to compare the mean pain score for pressure topical anaesthesia application (16.8 (24.8)) with those for injection anaesthesia application (56.7 (30.0)). 18 eyes of 18 patients were compared with the fellow eye of the same 18 patients. The mean pain score for injection anaesthesia was greater than for pressure topical anaesthesia application (p<0.0001) (statistical power=0.87). No statistically significant difference was found in the mean pain score for diathermy for eyes that received pressure topical anaesthesia (20.5 (27.5)) compared with eyes that received injection anaesthesia (23.1 (26.3)) (p=0.760) (statistical power=0.96). All 18 patients preferred pressure topical anaesthesia to injection anaesthesia. CONCLUSION: Injection anaesthesia for punctal occlusion is more painful than pressure topical anaesthesia application. However, the pain experienced during diathermy application for punctal occlusion is similar between pressure anaesthetised eyes and injection anaesthetised eyes. Pressure topical anaesthesia is a less painful (in terms of anaesthesia application) but equally effective alternative to conventional injection anaesthesia when used for punctal occlusion.

Adult↗

Précis of statistical significance: rationale, validity, and utility.

The null-hypothesis significance-test procedure (NHSTP) is defended in the context of the theory-corroboration experiment, as well as the following contrasts: (a) substantive hypotheses versus statistical hypotheses, (b) theory corroboration versus statistical hypothesis testing, (c) theoretical inference versus statistical decision, (d) experiments versus nonexperimental studies, and (e) theory corroboration versus treatment assessment. The null hypothesis can be true because it is the hypothesis that errors are randomly distributed in data. Moreover, the null hypothesis is never used as a categorical proposition. Statistical significance means only that chance influences can be excluded as an explanation of data; it does not identify the nonchance factor responsible. The experimental conclusion is drawn with the inductive principle underlying the experimental design. A chain of deductive arguments gives rise to the theoretical conclusion via the experimental conclusion. The anomalous relationship between statistical significance and the effect size often used to criticize NHSTP is more apparent than real. The absolute size of the effect is not an index of evidential support for the substantive hypothesis. Nor is the effect size, by itself, informative as to the practical importance of the research result. Being a conditional probability, statistical power cannot be the a priori probability of statistical significance. The validity of statistical power is debatable because statistical significance is determined with a single sampling distribution of the test statistic based on H0, whereas it takes two distributions to represent statistical power or effect size. Sample size should not be determined in the mechanical manner envisaged in power analysis. It is inappropriate to criticize NHSTP for nonstatistical reasons. At the same time, neither effect size, nor confidence interval estimate, nor posterior probability can be used to exclude chance as an explanation of data. Neither can any of them fulfill the nonstatistical functions expected of them by critics.

Reproducibility of Results↗

The power of statistical studies in consultation-liaison psychiatry.

Several authors recently have proclaimed the need for empirically based research articles in consultation-liaison psychiatry. The authors report that although the proportion of empirically based studies published in Psychosomatics increased 148% from 1979 to 1989, the power of statistical analyses and the deleterious effect of multiple tests were often neglected. A power analysis of empirical studies published in the 1989 volume year of Psychosomatics is reported, showing statistical power to be low for all but the most robust of effect sizes.

Female↗

Incidence of childhood leukaemia and non-Hodgkin's lymphoma in the vicinity of nuclear sites in Scotland, 1968-93.

OBJECTIVES: The primary aims were to investigate the incidence of leukaemia and non-Hodgkin's lymphoma in children resident near seven nuclear sites in Scotland and to determine whether there was any evidence of a gradient in risk with distance of residence from a nuclear site. A secondary aim was to assess the power of statistical tests for increased risk of disease near a point source when applied in the context of census data for Scotland. METHODS: The study data set comprised 1287 cases of leukaemia and non-Hodgkin's lymphoma diagnosed in children aged under 15 years in the period 1968-93, validated for accuracy and completeness. A study zone around each nuclear site was constructed from enumeration districts within 25 km. Expected numbers were calculated, adjusting for sex, age, and indices of deprivation and urban-rural residence. Six statistical tests were evaluated. Stone's maximum likelihood ratio (unconditional application) was applied as the main test for general increased incidence across a study zone. The linear risk score based on enumeration districts (conditional application) was used as a secondary test for declining risk with distance from each site. RESULTS: More cases were observed (O) than expected (E) in the study zones around Rosyth naval base (O/E 1.02), Chapelcross electricity generating station (O/E 1.08), and Dounreay reprocessing plant (O/E 1.99). The maximum likelihood ratio test reached significance only for Dounreay (P = 0.030). The linear risk score test did not indicate a trend in risk with distance from any of the seven sites, including Dounreay. CONCLUSIONS: There was no evidence of a generally increased risk of childhood leukaemia and non-Hodgkin's lymphoma around nuclear sites in Scotland, nor any evidence of a trend of decreasing risk with distance from any of the sites. There was a significant excess risk in the zone around Dounreay, which was only partially accounted for by the sociodemographic characteristics of the area. The statistical power of tests for localised increased risk of disease around a point source should be assessed in each new setting in which they are applied.

Adolescent↗

Comparative trials on hybrid walking systems for people with paraplegia: an analysis of study methodology.

A new orthosis (SEPRIX) which combines user friendliness with low energy cost of walking has been developed and will be subject to a clinical comparison with conventional hip-knee-ankle-foot orthoses. In designing such comparative trials it was considered it may be worthwhile to use previous clinical studies as practical examples. A literature search was conducted in order to select all comparative trials which have studied two walking systems (hip-knee-ankle-foot orthoses) for patients with a complete thoracic lesion. Study population, intervention, study design, outcome measurement and statistical analyses were examined. Statistical power was calculated where possible. Of 12 selected studies, 7 were simple A-B comparisons, 2 A-B comparisons with a replication, 2 cross-over trials and 1 nonrandomised parallel group design, the last of which was considered internally invalid due to severe confounding by indication. All A-B comparisons were considered internally invalid as well, since they have not taken into account that a comparison of two orthoses requires a control for aspecific effects (like test effects) which may cause a difference. Statistical power could only be examined in 4 studies and the highest statistical power achieved in one study was 47%. It is concluded that statistical power was too low to be able to detect differences. Even analysis through interval estimation showed that the estimation of the difference was too imprecise to be useful. Since the majority of the surveyed papers have reported small studies (of only 4-6 patients), it is assumed that lack of statistical power is a more general problem. Three possibilities are discussed in order to enhance statistical power in comparative trials, i.e. multicentre studies, statistical pooling of results and improving the efficiency of study design by means of interrupted time series designs.

Equipment Design↗

Small-scale randomized controlled trials need more powerful methods of mediational analysis than the Baron-Kenny method.

OBJECTIVE: To devise more-effective physical activity interventions, the mediating mechanisms yielding behavioral change need to be identified. The Baron-Kenny method is most commonly used, but has low statistical power and may not identify mechanisms of behavioral change in small-to-medium size studies. More powerful statistical tests are available. STUDY DESIGN AND SETTING: Inactive adults (N=52) were randomized to either a print or a print-plus-telephone intervention. Walking and exercise-related social support were assessed at baseline, after the intervention, and 4 weeks later. The Baron-Kenny and three alternative methods of mediational analysis (Freedman-Schatzkin; MacKinnon et al.; bootstrap method) were used to examine the effects of social support on initial behavior change and maintenance. RESULTS: A significant mediational effect of social support on initial behavior change was indicated by the MacKinnon et al., bootstrap, and, marginally, Freedman-Schatzkin methods, but not by the Baron-Kenny method. No significant mediational effect of social support on maintenance of walking was found. CONCLUSIONS: Methodologically rigorous intervention studies to identify mediators of change in physical activity are costly and labor intensive, and may not be feasible with large samples. The use of statistically powerful tests of mediational effects in small-scale studies can inform the development of more effective interventions.

Aged↗

Sample size for detecting differentially expressed genes in microarray experiments.

BACKGROUND: Microarray experiments are often performed with a small number of biological replicates, resulting in low statistical power for detecting differentially expressed genes and concomitant high false positive rates. While increasing sample size can increase statistical power and decrease error rates, with too many samples, valuable resources are not used efficiently. The issue of how many replicates are required in a typical experimental system needs to be addressed. Of particular interest is the difference in required sample sizes for similar experiments in inbred vs. outbred populations (e.g. mouse and rat vs. human). RESULTS: We hypothesize that if all other factors (assay protocol, microarray platform, data pre-processing) were equal, fewer individuals would be needed for the same statistical power using inbred animals as opposed to unrelated human subjects, as genetic effects on gene expression will be removed in the inbred populations. We apply the same normalization algorithm and estimate the variance of gene expression for a variety of cDNA data sets (humans, inbred mice and rats) comparing two conditions. Using one sample, paired sample or two independent sample t-tests, we calculate the sample sizes required to detect a 1.5-, 2-, and 4-fold changes in expression level as a function of false positive rate, power and percentage of genes that have a standard deviation below a given percentile. CONCLUSIONS: Factors that affect power and sample size calculations include variability of the population, the desired detectable differences, the power to detect the differences, and an acceptable error rate. In addition, experimental design, technical variability and data pre-processing play a role in the power of the statistical tests in microarrays. We show that the number of samples required for detecting a 2-fold change with 90% probability and a p-value of 0.01 in humans is much larger than the number of samples commonly used in present day studies, and that far fewer individuals are needed for the same statistical power when using inbred animals rather than unrelated human subjects.

Animals↗

Numerical statistics of power dropouts based on the Lang-Kobayashi model.

The statistics of power dropouts in semiconductor lasers subjected to delayed optical feedback have been numerically investigated using the Lang-Kobayashi model. The data from the numerical simulations have then been used to calculate the probability distribution functions, mean values, and return maps of the time that elapses between dropouts. In addition, the transition from the "low frequency fluctuation" to the "coherence collapse" regime has also been investigated. The numerical simulations compare well with both experimental results, obtained from multilongitudinal mode lasers, and analytical results obtained from other theoretical models. Evidence of "excitability" within the Lang-Kobayashi model is also reported.

Journal Article↗

Dynamic wait-listed designs for randomized trials: new designs for prevention of youth suicide.

BACKGROUND: The traditional wait-listed design, where half are randomly assigned to receive the intervention early and half are randomly assigned to receive it later, is often acceptable to communities who would not be comfortable with a no-treatment group. As such this traditional wait-listed design provides an excellent opportunity to evaluate short-term impact of an intervention. We introduce a new class of wait-listed designs for conducting randomized experiments where all subjects receive the intervention, and the timing of the intervention is randomly assigned. We use the term "dynamic wait-listed designs" to describe this new class. PURPOSE: This paper examines a new class of statistical designs where random assignment to intervention condition occurs at multiple times in a trial. As an extension of a traditional wait-listed design, this dynamic design allows all subjects to receive the intervention at a random time. Motivated by our search for increased statistical power in an ongoing school-based trial that is testing a program of gatekeeper training to identify suicidal youth and refer them to treatment, this new design class is especially useful when the primary outcome is a count or rate of occurrence, such as suicidal behavior, whose rate can fluctuate over time due to uncontrolled factors. METHODS: Statistical power is computed for various dynamic wait-listed designs under conditions where the underlying rate of occurrence is allowed to vary nonsystematically. We also present as an example a large ongoing trial to evaluate a gatekeeper training suicide prevention program in 32 schools which we initially began as a classic randomized wait-listed design. The primary outcome of interest in this study is the count of the number of children who are identified by the school system as having suicidal thoughts or behaviors who are then validated as being suicidal by mental health professionals in the community. RESULTS: A general result shows that dynamic wait-listed designs always have higher statistical power over a traditional wait-listed design. This power increase can be substantial. Efficiency gains of 33% are easy to obtain for situations where the intervention has a small effect and the variation in rate across time is quite high. When the rate variation for an outcome is very low or the intervention effect is large, efficiency gains approach 100%. A small increase in the number of times where random assignment occurs from 2 - for the standard wait-listed design, to say 4 can provide a large reduction in variance. Efficiency gains can also be high when converting standard wait-listed design to a dynamic one half-way into the study. LIMITATIONS: As with all wait-listed designs, dynamic wait-listed designs can only be used to evaluate short-term impact. Since all subjects eventually receive the intervention, no comparison can be made after the end of the random assignment period. The statistical power benefits are primarily limited to outcomes that can be treated as count or time to event data. CONCLUSIONS: A dynamic design randomly assigns units - either individuals or groups - to start the intervention at varying times during the course of the study. This design is useful in testing interventions that screen for new or existing cases, as well as testing the scalability of interventions as they are disseminated or expanded system wide. They can improve on the traditional wait-listed design both in terms of statistical power and robustness in the presence of exogenous factors. This paper demonstrates that such designs yield smaller standard errors and can achieve higher statistical power than that of a standard wait-listed design. Just as important, dynamic designs can also help reduce the logistical challenges of implementing an intervention on a wide scale. When the intervention requires that significant training resources be allocated throughout the study, the dynamic wait-listed design is likely to increase the rate of training and lead to a higher level of program implementation.

Adolescent↗

More reliable outcome measures can reduce sample size requirements.

In the design of a clinical trial, considerations of statistical power primarily involve the evaluation of prospective sample sizes. Another strategy for increasing statistical power that is rarely used focuses on the selection of the outcome measure. When an outcome measure is selected, its reliability and validity must be carefully evaluated. Here the relationship between reliability and statistical power is explored empirically. We show that as the number of related items in an outcome scale increases, the internal consistency reliability of the scale also increases. As a consequence, the within-group variability decreases and, in turn, the between-group effect size increases and sample size requirements decrease. As a result, sample size requirements can be reduced and research costs decreased. We recommend careful consideration of the psychometric properties of outcome measures prior to sample size determination in any statistical power analyses.

Alprazolam↗

Charting patients' course: a comparison of statistics used to summarize patient course in longitudinal and repeated measures studies.

Investigators conducting longitudinal studies of psychiatric illnesses often analyze data based on psychiatric symptom scales that were administered at multiple time points. This study examines the statistical properties of seven indices that summarize patient long-term course. These indices can be used to compare differences between two or more groups or to test for changes in symptoms over time. They may also be treated as outcome measures and correlated with other clinical variables.The performance of each of the seven indices was assessed using data from two large ongoing studies of psychiatric patients: a longitudinal study of affective disorders and a longitudinal study of first-episode psychosis. These two datasets were subjected to bootstrapping techniques in order to calculate both type I error rates and statistical power for each summary statistic. Of the seven indices, Kendall's tau performed the best as a measure of patients' symptom course. Kendall's tau appears to offer more statistical power to detect change in course, yet its average type I error rate was comparable to the other indices.

Adult↗