Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

[Abecedary of the "annales". Part 16].

The terms included and detailed in the present part are: weighting, diagnostic accuracy, predictive value, prevalence, prevention, statistical power.

Statistics as Topic↗

Statistical development and evaluation of microarray gene expression data filters.

Filtering is a common practice used to simplify the analysis of microarray data by removing from subsequent consideration probe sets believed to be unexpressed. The m/n filter, which is widely used in the analysis of Affymetrix data, removes all probe sets having fewer than m present calls among a set of n chips. The m/n filter has been widely used without considering its statistical properties. The level and power of the m/n filter are derived. Two alternative filters, the pooled p-value filter and the error-minimizing pooled p-value filter are proposed. The pooled p-value filter combines information from the present-absent p-values into a single summary p-value which is subsequently compared to a selected significance threshold. We show that pooled p-value filter is the uniformly most powerful statistical test under a reasonable beta model and that it exhibits greater power than the m/n filter in all scenarios considered in a simulation study. The error-minimizing pooled p-value filter compares the summary p-value with a threshold determined to minimize a total-error criterion based on a partition of the distribution of all probes' summary p-values. The pooled p-value and error-minimizing pooled p-value filters clearly perform better than the m/n filter in a case-study analysis. The case-study analysis also demonstrates a proposed method for estimating the number of differentially expressed probe sets excluded by filtering and subsequent impact on the final analysis. The filter impact analysis shows that the use of even the best filter may hinder, rather than enhance, the ability to discover interesting probe sets or genes. S-plus and R routines to implement the pooled p-value and error-minimizing pooled p-value filters have been developed and are available from www.stjuderesearch.org/depts/biostats/index.html.

Computational Biology↗

Bioassays of shortened duration for drugs: statistical implications.

Declining survival rates in rodent carcinogenesis bioassays have raised a concern that continuing the practice of terminating such studies at 24 months could result in too few live animals at termination for adequate pathological evaluation. One option for ensuring sufficient numbers of animals at the terminal sacrifice is to shorten the duration of the bioassay, but this approach is accompanied by a reduction in statistical power for detecting carcinogenic potential. The present study was conducted to evaluate the loss of power associated with early termination. Data from drug studies in rats were used to formulate biologically based dose-response models of carcinogenesis using the 2-stage clonal expansion model as a context. These dose-response models, which were chosen to represent 6 variations of the initiation-promotion-completion cancer model, were employed to generate a large number of representative bioassay data sets using Monte Carlo simulation techniques. For a variety of tumor dose-response trends, tumor lethality, and competing risk-survival rates, the power of age-adjusted statistical tests to assess the significance of carcinogenic potential was evaluated at 18 and 21 months, and compared to the power at the normal 24-month stopping time. The results showed that stopping at 18 months would reduce power to an unacceptable level for all 6 submodels of the 2-stage clonal expansion model, with the pure-promoter and pure-completer models being most adversely affected. For the 21-month stopping time, the results showed that, unless pure promotion can be ruled out a priori as a potential carcinogenic mode of action, the loss of power is too great to warrant early stopping.

Animals↗

The PowerAtlas: a power and sample size atlas for microarray experimental design and research.

BACKGROUND: Microarrays permit biologists to simultaneously measure the mRNA abundance of thousands of genes. An important issue facing investigators planning microarray experiments is how to estimate the sample size required for good statistical power. What is the projected sample size or number of replicate chips needed to address the multiple hypotheses with acceptable accuracy? Statistical methods exist for calculating power based upon a single hypothesis, using estimates of the variability in data from pilot studies. There is, however, a need for methods to estimate power and/or required sample sizes in situations where multiple hypotheses are being tested, such as in microarray experiments. In addition, investigators frequently do not have pilot data to estimate the sample sizes required for microarray studies. RESULTS: To address this challenge, we have developed a Microrarray PowerAtlas. The atlas enables estimation of statistical power by allowing investigators to appropriately plan studies by building upon previous studies that have similar experimental characteristics. Currently, there are sample sizes and power estimates based on 632 experiments from Gene Expression Omnibus (GEO). The PowerAtlas also permits investigators to upload their own pilot data and derive power and sample size estimates from these data. This resource will be updated regularly with new datasets from GEO and other databases such as The Nottingham Arabidopsis Stock Center (NASC). CONCLUSION: This resource provides a valuable tool for investigators who are planning efficient microarray studies and estimating required sample sizes.

Algorithms↗

Impact of criticism of null-hypothesis significance testing on statistical reporting practices in conservation biology.

Over the last decade, criticisms of null-hypothesis significance testing have grown dramatically, and several alternative practices, such as confidence intervals, information theoretic, and Bayesian methods, have been advocated. Have these calls for change had an impact on the statistical reporting practices in conservation biology? In 2000 and 2001, 92% of sampled articles in Conservation Biology and Biological Conservation reported results of null-hypothesis tests. In 2005 this figure dropped to 78%. There were corresponding increases in the use of confidence intervals, information theoretic, and Bayesian techniques. Of those articles reporting null-hypothesis testing--which still easily constitute the majority--very few report statistical power (8%) and many misinterpret statistical nonsignificance as evidence for no effect (63%). Overall, results of our survey show some improvements in statistical practice, but further efforts are clearly required to move the discipline toward improved practices.

Conservation of Natural Resources↗

Randomized clinical trial: myths around elementary statistical principles.

In discussing design and results of randomized clinical trials, in particular with clinical oncologists, one often encounters the opinion that a phase III trial is a complicated, highly costly, and difficult task. Part of this opinion seems to originate in myths around underlying biostatistical principles such as randomization, sampling and sample size, statistical hypotheses, statistical error probabilities, and statistical power. This work clarifies basic statistical issues of randomized clinical trials and the interpretation of their results. Six issues ('myths') relevant for the design of clinical trials and the interpretation of their results are addressed. They concern choice of study design, choice of participating centers, and recruitment of patients as well as statistical questions of establishing study hypotheses and interpreting p values. These myths are shown to be caused primarily through a misunderstanding of statistical inference and statistical thinking that can be avoided when a rational understanding of statistical principles is translated into a clinical research approach. We also conclude that before clinical evidence is summarized from different studies each study should be examined thoroughly.

Bias↗

Selective exposure and dissonance after decisions.

Well-known literature reviews from the 1960s question whether cognitive dissonance underlies experimental participants' selective exposure of themselves to consonant messages and avoidance of dissonant ones. A meta-analytic review of 16 studies published from 1956 to 1996 and involving 1,922 total participants shows that experimental tests consistently support the supposition that dissonance is associated with selective exposure (r = .22, p < .001). Statistical power exceeded .99. Advances in statistical methodology and increased attention to selecting appropriate tests of dissonance theory were essential to finally resolving this question.

Cognition↗

Cost comparison of oral, nasogastric, and intramuscular cimetidine drug delivery systems.

Under diagnosis-related group prospective payment, the role of the hospital pharmacy department is to develop and promote cost effective and rational drug therapy. This study, conducted in a simulated fashion, evaluated the cost effectiveness of various cimetidine drug delivery systems: oral, nasogastric, and intramuscular. The evaluation was also extended to compare time efficacy between different dosage forms for each route of administration with the exception of the intramuscular route. Each of four nurses and four pharmacy technicians conducted 10 trials for each system to detect a statistical significance difference in pharmacy preparation and nursing administration time with at least 80% statistical power. The results showed no statistically significant difference (P > 0.1) between the total oral administration time of a unit-dose tablet or liquid. A significant difference was detected among the pharmacy-prepared liquids, and unit-dose liquid administered nasogastrically (P < 0.01). The most cost-effective system is the orally administered unit-dose tablet and the most expensive system is the unit-dose liquid administered nasogastrically.

Administration, Intranasal↗

Recommendations for appropriate statistical practice in toxicologic experiments.

Appropriate statistical practice in toxicologic research is reviewed. The problem centers on the antagonistic needs to discover toxic effects and avoid false indictments of harmless compounds. Specific problems which distort error rates include having many dependent variables, conducting many tests on a variable, data snooping, lack of statistical power and 5) violations of statistical assumptions. Solution strategies include top down planning, designing efficient experiments, balancing type I and type II error rates, selecting hypotheses and 5) using appropriate statistical analysis methods. Top down planning and the "leapfrog" design strategy are particularly emphasized.

Research Design↗

Power evaluation of disease clustering tests.

BACKGROUND: Many different test statistics have been proposed to test for spatial clustering. Some of these statistics have been widely used in various applications. In this paper, we use an existing collection of 1,220,000 simulated benchmark data, generated under 51 different clustering models, to compare the statistical power of several disease clustering tests. These tests are Besag-Newell's R, Cuzick-Edwards' k-Nearest Neighbors (k-NN), the spatial scan statistic, Tango's Maximized Excess Events Test (MEET), Swartz' entropy test, Whittemore's test, Moran's I and a modification of Moran's I. RESULTS: Except for Moran's I and Whittemore's test, all other tests have good power for detecting some kind of clustering. The spatial scan statistic is good at detecting localized clusters. Tango's MEET is good at detecting global clustering. With appropriate choice of parameter, Besag-Newell's R and Cuzick-Edwards' k-NN also perform well. CONCLUSION: The power varies greatly for different test statistics and alternative clustering models. Consideration of the power is important before we decide which test statistic to use.

Journal Article↗

Scale and shape issues in focused cluster power for count data.

BACKGROUND: Interest in the development of statistical methods for disease cluster detection has experienced rapid growth in recent years. Evaluations of statistical power provide important information for the selection of an appropriate statistical method in environmentally-related disease cluster investigations. Published power evaluations have not yet addressed the use of models for focused cluster detection and have not fully investigated the issues of disease cluster scale and shape. As meteorological and other factors can impact the dispersion of environmental toxicants, it follows that environmental exposures and associated diseases can be dispersed in a variety of spatial patterns. This study simulates disease clusters in a variety of shapes and scales around a centrally located single pollution source. We evaluate the power of a range of focused cluster tests and generalized linear models to detect these various cluster shapes and scales for count data. RESULTS: In general, the power of hypothesis tests and models to detect focused clusters improved when the test or model included parameters specific to the shape of cluster being examined (i.e. inclusion of a function for direction improved power of models to detect clustering with an angular effect). However, power to detect clusters where the risk peaked and then declined was limited. CONCLUSION: Findings from this investigation show sizeable changes in power according to the scale and shape of the cluster and the test or model applied. These findings demonstrate the importance of selecting a test or model with functions appropriate to detect the spatial pattern of the disease cluster.

Journal Article↗

Extreme selection strategies in gene mapping studies of oligogenic quantitative traits do not always increase power.

It is well known that obtaining adequate statistical power to detect linkage to or association with genes for complex quantitative traits can be very difficult. In response, investigators have developed a number of power-enhancing strategies that consider restraints such as genotyping (and/or phenotyping) costs. In the context of both association and sib pair linkage studies of quantitative traits, one of the most widely discussed techniques is the selective sampling of phenotypically extreme individuals. Several papers have demonstrated that such extreme sampling can markedly increase power (under certain circumstances). However, the parenthetical phrase in the previous sentence has generally not been made explicit and it appears to be implied that the more phenotypically extreme the individuals, the more power one has. In this paper, we show by simulation that this is not true under all circumstances. In particular, we show that under oligogenic models, where some biallelic quantitative trait loci (QTLs) have markedly asymmetric allele frequencies and large mean displacement among genotypes, and others have less asymmetric allele frequencies and smaller mean displacement among genotypes, power to detect linkage to or association with the latter QTL can actually decrease by sampling more extreme sib pairs. This suggests that more extreme sampling is not always better. The 'optimal' sampling scheme may depend on both what one suspects the underlying genetic architecture to be and which of the oligogenic QTL one has greatest interest in detecting.

Chromosome Mapping↗

Characteristics of simple sibship variance tests for the detection of major loci and application to height, weight and spatial performance.

Simple methods have been proposed as screening tests for major loci. These methods rely primarily upon the detection of differences in within sibship variances expected for segregating and non-segregating sibships when a major locus is present. In the present study, computer simulation was used to investigate power and robustness of the test statistics. Power of the analyses depends upon the specific major locus model, but, in general, their application is quite practical for small samples. The test statistics were shown to be sensitive to deviations from normality, but robust under the conditions of either a polygenic or environmental model. Application of the test procedures to sibship data from the Boulder Family Study led to significant results for a three-dimensional spatial rotation test, but results for height and weight were non-significant. The simplest interpretation of the results for the spatial performance test was in terms of a sex-linked major locus.

Adolescent↗

Benchmark data and power calculations for evaluating disease outbreak detection methods.

INTRODUCTION: Early detection of disease outbreaks enables public health officials to implement immediate disease control and prevention measures. Computer-based syndromic surveillance systems are being implemented to complement reporting by physicians and other health-care professionals to improve the timeliness of disease-outbreak detection. Space-time disease-surveillance methods have been proposed as a supplement to purely temporal statistical methods for outbreak detection to detect localized outbreaks before they spread to larger regions. OBJECTIVE: The aims of this study were twofold: 1) to design and make available benchmark data sets for evaluating the statistical power of space-time early detection methods and 2) to evaluate the power of the prospective purely temporal and space-time scan statistics by applying them to the benchmark data sets at different parameter settings. METHODS: Simulated data sets based on the geography and population of New York City were created, including effects of outbreaks of varying size and location. Data sets with no outbreak effects were also created. Scan statistics were then run on these data sets, and the resulting power performances were analyzed and compared. RESULTS: The prospective space-time scan statistic performs well for a spectrum of outbreak models. By comparison, the prospective purely temporal scan statistic has higher power for detecting citywide outbreaks but lower power for detecting geographically localized outbreaks. CONCLUSIONS: The benchmark data sets created for this study can be used successfully for formal statistical power evaluations and comparisons. If an anomaly caused by an outbreak is local, purely temporal surveillance methods might be unable to detect it, in which case space-time methods would be necessary for early detection.

Benchmarking↗

Power in randomized group comparisons: the value of adding a single intermediate time point to a traditional pretest-posttest design.

Adding a pretest as a covariate to a randomized posttest-only design increases statistical power, as does the addition of intermediate time points to a randomized pretest-posttest design. Although typically 5 waves of data are required in this instance to produce meaningful gains in power, a 3-wave intensive design allows the evaluation of the straight-line growth model and may reduce the effect of missing data. The authors identify the statistically most powerful method of data analysis in the 3-wave intensive design. If straight-line growth is assumed, the pretest-posttest slope must assume fairly extreme values for the intermediate time point to increase power beyond the standard analysis of covariance on the posttest with the pretest as covariate, ignoring the intermediate time point.

Humans↗

Theoretical and empirical power of regression and maximum-likelihood methods to map quantitative trait loci in general pedigrees.

Both theoretical calculations and simulation studies have been used to compare and contrast the statistical power of methods for mapping quantitative trait loci (QTLs) in simple and complex pedigrees. A widely used approach in such studies is to derive or simulate the expected mean test statistic under the alternative hypothesis of a segregating QTL and to equate a larger mean test statistic with larger power. In the present study, we show that, even when the test statistic under the null hypothesis of no linkage follows a known asymptotic distribution (the standard being chi(2)), it cannot be assumed that the distribution under the alternative hypothesis is noncentral chi(2). Hence, mean test statistics cannot be used to indicate power differences, and a comparison between methods that are based on simulated average test statistics may lead to the wrong conclusion. We illustrate this important finding, through simulations and analytical derivations, for a recently proposed new regression method for the analysis of general pedigrees to map quantitative trait loci. We show that this regression method is not necessarily more powerful nor computationally more efficient than a maximum-likelihood variance-component approach. We advocate the use of empirical power to compare trait-mapping methods.

Chromosome Mapping↗

Bayesian evaluation of group sequential clinical trial designs.

Clinical trial designs often incorporate a sequential stopping rule to serve as a guide in the early termination of a study. When choosing a particular stopping rule, it is most common to examine frequentist operating characteristics such as type I error, statistical power, and precision of confidence intervals (Statist. Med. 2005, in revision). Increasingly, however, clinical trials are designed and analysed in the Bayesian paradigm. In this paper, we describe how the Bayesian operating characteristics of a particular stopping rule might be evaluated and communicated to the scientific community. In particular, we consider a choice of probability models and a family of prior distributions that allows concise presentation of Bayesian properties for a specified sampling plan.

Antibodies↗