Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical power analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Publication and related bias in meta-analysis: power of statistical tests and prevalence in the literature.

Publication and selection biases in meta-analysis are more likely to affect small studies, which also tend to be of lower methodological quality. This may lead to "small-study effects," where the smaller studies in a meta-analysis show larger treatment effects. Small-study effects may also arise because of between-trial heterogeneity. Statistical tests for small-study effects have been proposed, but their validity has been questioned. A set of typical meta-analyses containing 5, 10, 20, and 30 trials was defined based on the characteristics of 78 published meta-analyses identified in a hand search of eight journals from 1993 to 1997. Simulations were performed to assess the power of a weighted regression method and a rank correlation test in the presence of no bias, moderate bias or severe bias. We based evidence of small-study effects on P < 0.1. The power to detect bias increased with increasing numbers of trials. The rank correlation test was less powerful than the regression method. For example, assuming a control group event rate of 20% and no treatment effect, moderate bias was detected with the regression test in 13.7%, 23.5%, 40.1% and 51.6% of meta-analyses with 5, 10, 20 and 30 trials. The corresponding figures for the correlation test were 8.5%, 14.7%, 20.4% and 26.0%, respectively. Severe bias was detected with the regression method in 23.5%, 56.1%, 88.3% and 95.9% of meta-analyses with 5, 10, 20 and 30 trials, as compared to 11.9%, 31.1%, 45.3% and 65.4% with the correlation test. Similar results were obtained in simulations incorporating moderate treatment effects. However the regression method gave false-positive rates which were too high in some situations (large treatment effects, or few events per trial, or all trials of similar sizes). Using the regression method, evidence of small-study effects was present in 21 (26.9%) of the 78 published meta-analyses. Tests for small-study effects should routinely be performed in meta-analysis. Their power is however limited, particularly for moderate amounts of bias or meta-analyses based on a small number of small studies. When evidence of small-study effects is found, careful consideration should be given to possible explanations for these in the reporting of the meta-analysis.

Bias↗

Power analysis of statistical methods for comparing treatment differences from limiting dilution assays.

Six different statistical methods for comparing limiting dilution assays were evaluated, using both real data and a power analysis of simulated data. Simulated data consisted of a series of 12 dilutions for two treatment groups with 24 cultures per dilution and 1,000 independent replications of each experiment. Data within each replication were generated by Monte Carlo simulation, based on a probability model of the experiment. Analyses of the simulated data revealed that the type I error rates for the six methods differed substantially, with only likelihood ratio and Taswell's weighted mean methods approximating the nominal 5% significance level. Of the six methods, likelihood ratio and Taswell's minimum Chi-square exhibited the best power (least probability of type II errors). Taswell's weighted mean test yielded acceptable type I and type II error rates, whereas the regression method was judged unacceptable for scientific work.

Cell Survival↗

Individual-specific risk factors for anorexia nervosa: a pilot study using a discordant sister-pair design.

BACKGROUND: The aim of this pilot study was to examine which unique factors (genetic and environmental) increase the risk for developing anorexia nervosa by using a case-control design of discordant sister pairs. METHODS: Forty-five sister-pairs, one of whom had anorexia nervosa and the other did not, were recruited. Both sisters completed the Oxford Risk Factor Interview for Eating Disorders and measures for eating disorder traits, and sib-pair differences. Blood or cheek cell samples were taken for genetic analysis. Statistical power of the genetic analysis of discordant same-sex siblings was calculated using a specially written program, DISCORD. RESULTS: The sisters with anorexia nervosa differed from their healthy sisters in terms of personal vulnerability traits and exposure to high parental expectations and sexual abuse. Factors within the dieting risk domain did not differ. However, there was evidence of poor feeding in childhood. No difference in the distribution of genotypes or alleles of the DRD4, COMT, the 5HT2A and 5HT2C receptor genes was detected. These results are preliminary because our calculations indicate that there is insufficient power to detect the expected effect on risk with this sample size. CONCLUSIONS: A combination of intrinsic and extrinsic factors increases the risk of developing anorexia nervosa. It would, therefore, be informative to undertake a larger study to examine in more detail the unique genetic and environmental factors that are associated with various forms of eating disorders.

Adolescent↗

Regression analysis in biological research: sample size and statistical power.

Regression analysis is often used to demonstrate associations among variables believed to be biologically related. Failure to demonstrate a "significant" relationship may be due to two factors: 1) the variables are truly unrelated, or 2) a relationship exists but goes undetected due to inadequate statistical power. Investigators must consider the second possibility since failure to detect a statistically significant relationship is often taken as evidence for no biological relationship. These issues are addressed in the context of the interrelationship between four features common to all statistical methods: the size of effect or relationship worth detecting, the Type I (alpha) error, the sample size, and the Type II (beta) error. An example derived from published data relating morphological characteristics of muscle fiber type and isokinetic strength performance illustrates the practical significance of this dilemma.

Regression Analysis↗

Sampling design for total and filterable reactive phosphorus monitoring in a lowland stream: considerations of spatial variability, measurement uncertainty and statistical power.

An analysis for spatial variation of phosphorus (P) concentrations in the dissolved and particulate compartments of Latrobe River water in Victoria, Australia is described. Water sampling was based on a nested hierarchical design and variation was measured at different spatial scales. Total variance of the dissolved and particulate P compartments was partitioned using analysis of variance (ANOVA) to determine the spatial scale that requires most sampling effort. Statistical power analysis was used to determine the optimum sample size for the spatial scale. An uncertainty budget was estimated from sampling and analytical uncertainty. Ten and twelve samples, at the smallest spatial scale, required the greatest sampling effort, and led to the greatest required statistical power for the determination of dissolved and particulate P, respectively, in the Latrobe River catchment. The results emphasize the need for aquatic chemists to be aware of the ramifications of different types of uncertainty and variance in environmental studies.

Environmental Monitoring↗

Increasing scientific power with statistical power.

A survey of basic ideas in statistical power analysis demonstrates the advantages and ease of using power analysis throughout the design, analysis, and interpretation of research. The power of a statistical test is the probability of rejecting the null hypothesis of the test. The traditional approach to power involves computation of only a single power value. The more general power curve allows examining the range of power determinants, which are sample size, population difference, and error variance, in traditional ANOVA. Power analysis can be useful not only in study planning, but also in the evaluation of existing research. An important application is in concluding that no scientifically important treatment difference exists. Choosing an appropriate power depends on: a) opportunity costs, b) ethical trade-offs, c) the size of effect considered important, d) the uncertainty of parameter estimates, and e) the analyst's preferences. Although precise rules seem inappropriate, several guidelines are defensible. First, the sensitivity of the power curve to particular characteristics of the study, such as the error variance, should be examined in any power analysis. Second, just as a small type I error rate should be demonstrated in order to declare a difference nonzero, a small type II error should be demonstrated in order to declare a difference zero. Third, when ethical and opportunity costs do not preclude it, power should be at least .84, and preferably greater than .90.

Data Interpretation, Statistical↗

Statistical power and sample size estimation for headache research: an overview and power calculation tools.

The present article reviews the concept of statistical power analysis for research designs in headache. First, we present a basic overview of the concepts of statistical hypothesis testing. Then we discuss the elements of power analysis and, where appropriate, we address conventions for power calculations. We offer, for public use, an applied power calculator for two applications that are often encountered in headache research. We intend to help headache researchers design trials with adequate statistical power by offering a conceptual overview and the power calculators. In closing, we briefly address the implications of the present trend toward reporting point estimates of effect sizes with confidence levels.

Headache Disorders↗

A refined in vitro model to study inflammatory responses in organotypic membrane culture of postnatal rat hippocampal slices.

BACKGROUND: Propagated tissue degeneration, especially during aging, has been shown to be enhanced through potentiation of innate immune responses. Neurodegenerative diseases and a wide variety of inflammatory conditions are linked together and several anti-inflammatory compounds considered as having therapeutic potential for example in Alzheimer's disease (AD). In vitro brain slice techniques have been widely used to unravel the complexity of neuroinflammation, but rarely, has the power of the model itself been reported. Our aim was to gain a more detailed insight and understanding of the behaviour of hippocampus tissue slices in serum-free, interface culture per se and after exposure to different pro- and anti-inflammatory compounds. METHODS: The responses of the slices to pro- and anti-inflammatory stimuli were monitored at various time points by measuring the leakage of lactate dehydrogenase (LDH) and the release of cytokines interleukin 6 (IL-6) and tumour necrosis factor alpha (TNF-alpha) and nitric oxide (NO) from the culture media. Histological methods were applied to reveal the morphological status after exposure to stimuli and during the time course of the culture period. Statistical power analysis were made with nQuery Advisor, version 5.0, (Statistical Solutions, Saugus, MA) computer program for Wilcoxon (Mann-Whitney) rank-sum test. RESULTS: By using the interface membrane culture technique, the hippocampal slices largely recover from the trauma caused by cutting after 4-5 days in vitro. Furthermore, the cultures remain stable and retain their responsiveness to inflammatory stimuli for at least 3 weeks. During this time period, cultures are susceptible to modification by inflammatory stimuli as assessed by quantitative biochemical assays and morphological characterizations. CONCLUSION: The present report outlines the techniques for studying immune responses using a serum-free slice culture model. Statistically powerful data under controlled culture conditions and with ethically justified use of animals can be obtained as soon as after 4-5 DIV. The model is most probably suitable also for studies of chronic inflammation.

Journal Article↗

[Usefulness of residuals in clinical research].

The simple linear regression analysis, multiple linear regression and logistic regression constitute powerful statistical analysis tools widely used in clinical research. These kinds of analyses are based upon mathematical models which at the same time are established on certain basic assumptions. The regression analysis assumptions are basically: a) that the model is really linear, b) that the distribution of data is normal (from a statistical point of view), c) that the variances of the employed data are homogeneous (homocedastics) and that the included data are independent. The regression diagnostic has become popular as a form to evaluate if the assumptions have been accomplished, one of its most important techniques is the residual analysis. A residual can be defined as the value which measures the distance between the regression line and the corresponding value of the variable "y". Among these kinds of residuals used to evaluate the assumptions of regression are: the crude residual, the standardized, of student and the jackknife. The most useful among them is the jackknife residual. The usefulness and limitations of the residuals in the evaluation of the regression analysis assumptions are described, basically referring to the identification and handling of extreme values (outliers).

Evaluation Studies as Topic↗

Weighing the results of differing 'low dose' studies of the mouse prostate by Nagel, Cagen, and Ashby: quantification of experimental power and statistical results.

Differing experimental findings with respect to "low dose" responses in the mouse prostate after in utero exposure have generated considerable controversy. An analysis of such controversies requires a broad strength and weight of the evidence approach. For example, a National Toxicology Program review panel acquired the raw data from nearly 50 studies and then statistically reanalyzed these data in a common and comparable approach. However, the statistical power of the various studies was not calculated and the quantitative p values were not reported in this reanalysis. Such calculations and values address vital strength- and weight-of-the-evidence questions: (1) how sensitive were the various studies to detect changes in prostate weight, particularly the negative replicate studies and (2) what were the p values; were negative studies robust or only marginal in their inability to find an effect? We first examined the statistical power of the studies to detect a positive effect on prostate weight. Preliminary calculations indicated that the two subsequent replicating studies were indeed more sensitive to changes in prostate weight in comparison to the original study, having reasonable power to detect an effect at only 50% of the response reported in the original study. Additional calculations were performed using the raw data available from one negative replicating study and the methods recommended by the statistics subpanel of the original review. This analysis used Dunnett's multiple comparison procedure for groups with p<0.05 to infer statistical significance, employed an analysis-of-covariance model with body weight as a covariate, and addressed litter as a nested random effect. The quantitated p values for this replicated study, comparing the two Bisphenol A treatment groups (2 and 20 microg/kg/day) to the control, were 0.821 and 0.972, respectively. This indicates this study was indeed robust in finding no treatment-related effect. Thus, the weight and strength of the evidence, based on sensitivity and quantitative p value, was that it is highly unlikely for this negative replicating study to have missed a true effect. In the future, we recommend a similar use of statistical power analysis for those designing experimental studies and for those conducting weight-of-the-evidence reviews, and we also recommend the clear quantitation and reporting of p values to support the review's interpretation and conclusions.

Algorithms↗

Children with disturbances in sensory processing: a pilot study examining the role of the parasympathetic nervous system.

This study was a preliminary investigation of parasympathetic nervous system (PNS) functioning in children with disturbances in sensory processing. The specific aims of this study were to (1) provide preliminary data about group differences in parasympathetic functions, as measured by the vagal tone index, between children with disturbances in sensory processing and those without; (2) determine effect size and power needed for future studies; and (3) to lay the foundation for further examination of the relations of parasympathetic functioning and functional behavior in children with disturbances in sensory processing. Participants were 15 children, nine with disturbances in sensory processing and six typically developing children. Heart period data were continuously collected for a 2-minute baseline and during administration of the 15-minute Sensory Challenge Protocol, a unique laboratory protocol designed to measure sensory reactivity (Miller, Reisman, McIntosh, & Simon, 2001). Groups were compared on vagal tone index, heart period, and heart rate using two-tailed, independent sample t tests. Children with disturbances in sensory processing had significantly lower vagal tone than the typically developing sample (t(13) = 2.4, p = .05). Statistical power analysis indicated that, for future studies, a sample size of 20 in each group would yield adequate statistical power. Although the number of subjects in this pilot study is small, the results from this study support further investigations of parasympathetic functions and functional behavior in children with disturbances in sensory processing.

Case-Control Studies↗

Conducting clinical trials to establish drug efficacy in chronic pain.

Design of clinical trials to establish drug efficacy in chronic pain is a complicated issue, and numerous factors must be considered in identifying an optimal study design. Investigators should begin by identifying a focused and testable research question, with the outcome variables operationalized in a way that allows appropriate quantitative analysis. The prospective, randomized, placebo-controlled, doubled-blind study using validated quantitative measures is considered the optimal clinical trial design. Within this general study type, the between-subjects design has less statistical power than does the crossover design, in which the patient serves as his or her own control. However, potential problems with the drug effects from the first condition carrying over into, and confounding, the second drug condition present a noteworthy limitation that must be addressed through adequate washout periods and statistical control if crossover designs are used. Retrospective designs may be useful primarily to take advantage of samples of convenience for development of pilot data that provide the basis for conducting better-controlled prospective studies. Sample selection issues must be considered during study design, including the sample size required (based on statistical power analysis), appropriate inclusion and exclusion criteria, and likely availability of qualifying patients. There are numerous statistical options for analyzing data that must be selected based on whether data are parametric or nonparametric, whether within-subject (crossover) or between-subject comparisons are used, and whether baseline values of outcome measures affect the degree of change in these measures over the course of the study. Involvement of a biostatistical consultant is recommended during all phases of clinical trials.

Analgesics↗

Kidney transplant rejection and tissue injury by gene profiling of biopsies and peripheral blood lymphocytes.

A major challenge for kidney transplantation is balancing the need for immunosuppression to prevent rejection, while minimizing drug-induced toxicities. We used DNA microarrays (HG-U95Av2 GeneChips, Affymetrix) to determine gene expression profiles for kidney biopsies and peripheral blood lymphocytes (PBLs) in transplant patients including normal donor kidneys, well-functioning transplants without rejection, kidneys undergoing acute rejection, and transplants with renal dysfunction without rejection. We developed a data analysis schema based on expression signal determination, class comparison and prediction, hierarchical clustering, statistical power analysis and real-time quantitative PCR validation. We identified distinct gene expression signatures for both biopsies and PBLs that correlated significantly with each of the different classes of transplant patients. This is the most complete report to date using commercial arrays to identify unique expression signatures in transplant biopsies distinguishing acute rejection, acute dysfunction without rejection and well-functioning transplants with no rejection history. We demonstrate for the first time the successful application of high density DNA chip analysis of PBL as a diagnostic tool for transplantation. The significance of these results, if validated in a multicenter prospective trial, would be the establishment of a metric based on gene expression signatures for monitoring the immune status and immunosuppression of transplanted patients.

Adult↗

Analysis of a multifactor microarray study using Partek genomics solution.

Partek Genomics Suite (Partek GS) is a powerful statistical analysis and interactive visualization software solution designed to analyze single channel oligonucleotide (Affymetrix) and two-color cDNA microarrays, as well as data from other emerging genomic and proteomic technologies. This chapter takes a simple study on obesity and susceptibility to type 2 diabetes and uses it as an example that demonstrates how Partek GS can be used to analyze data arising from a microarray experiment.

Animals↗

Critical evaluation of thioridazine bioequivalence.

The phenothiazines have exhibited a history of problems associated with the bioequivalence of solid oral dosage forms. The more recent availability of chemically equivalent forms of thioridazine has raised new and interesting questions about the appropriateness of generic product interchange, even among brands that have been designated "therapeutically equivalent" by the Food and Drug Administration. The scrutiny that has accompanied the consideration of thioridazine products for inclusion into various state generic substitution formularies has offered an opportunity to examine issues involving bioequivalency in considerable detail. Specific bioequivalency concerns relate to: correct analysis of drug in biological fluids; the importance of evaluating active metabolites: single-dose vs. multiple-dose crossover studies; appropriate statistical power analysis; the "70/70" rule, and comparison of product variabilities. Examples of problems are cited to illustrate that significant questions still remain about the appropriate factors that should be used to establish bioequivalency.

Biotransformation↗

Two-stage designs in case-control association analysis.

DNA pooling is a cost-effective approach for collecting information on marker allele frequency in genetic studies. It is often suggested as a screening tool to identify a subset of candidate markers from a very large number of markers to be followed up by more accurate and informative individual genotyping. In this article, we investigate several statistical properties and design issues related to this two-stage design, including the selection of the candidate markers for second-stage analysis, statistical power of this design, and the probability that truly disease-associated markers are ranked among the top after second-stage analysis. We have derived analytical results on the proportion of markers to be selected for second-stage analysis. For example, to detect disease-associated markers with an allele frequency difference of 0.05 between the cases and controls through an initial sample of 1000 cases and 1000 controls, our results suggest that when the measurement errors are small (0.005), approximately 3% of the markers should be selected. For the statistical power to identify disease-associated markers, we find that the measurement errors associated with DNA pooling have little effect on its power. This is in contrast to the one-stage pooling scheme where measurement errors may have large effect on statistical power. As for the probability that the disease-associated markers are ranked among the top in the second stage, we show that there is a high probability that at least one disease-associated marker is ranked among the top when the allele frequency differences between the cases and controls are not <0.05 for reasonably large sample sizes, even though the errors associated with DNA pooling in the first stage are not small. Therefore, the two-stage design with DNA pooling as a screening tool offers an efficient strategy in genomewide association studies, even when the measurement errors associated with DNA pooling are nonnegligible. For any disease model, we find that all the statistical results essentially depend on the population allele frequency and the allele frequency differences between the cases and controls at the disease-associated markers. The general conclusions hold whether the second stage uses an entirely independent sample or includes both the samples used in the first stage and an independent set of samples.

Algorithms↗

Non-invasive echocardiographic studies in mice: influence of anesthetic regimen.

Transgenic murine models of cardiovascular disease offer great potential insights regarding mechanisms of human disease, but efficient and reliable methods for phenotype evaluation are necessary. We employed non-invasive echocardiography to evaluate hemodynamic parameters in mice, and evaluated statistical reliability of these parameters with respect to anesthesia regimen. Male CF-1 mice received inhaled halothane (0.25-0.75% in 95% O2) or ketamine/xylazine (80/10 mg/kg i.p.) and 2-dimensional, M-mode, and Doppler ultrasound imaging were used to assess cardiac contractility and aortic flow velocities. Halothane was more convenient and reliable with respect to rate of induction, reversal, and control of anesthetic depth. At comparable levels of anesthesia, ketamine/xylazine produced significant reductions in heart rate (308 +/- 14 vs. 501 +/- 14 bpm, p<0.001), left ventricular fractional shortening (41.7 +/- 1.3 vs. 49.3 +/- 1.0%, p<0.001), and cardiac output (7.6 +/- 0.5 vs. 11.5 +/- 0.6 ml/min, p<0.001) when compared to halothane inhalation. No change in stroke volume or peak aortic velocity was observed. Correlation analyses revealed highly significant positive relationships between heart rate and fractional shortening (r=0.61, p<0.002) and cardiac output (r=0.88, p<0.001) but no relation to stroke volume or aortic velocity. Variability of intra-animal and intragroup parameter estimation were frequently 2-fold larger for ketamine/xylazine anesthesia vs. halothane. Statistical power analysis showed the increased measurement error for ketamine/xylazine leads to much larger numbers of mice/group to achieve identical statistical sensitivity. These data further illustrate the feasibility of echocardiography for rapid, non-invasive cardiovascular assessment in mice. However, several obtainable parameters are highly sensitive to both heart rate and anesthetic used and the choice and control of anesthetic are critical for physiologically relevant performance parameters and maximal ability to detect statistical differences among groups. Thus, for these non-invasive studies, inhalation anesthesia with agents such as halothane is superior to anesthesia induced by ketamine/xylazine administration.

Adrenergic alpha-Agonists↗