Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Sample size estimation: how many individuals should be studied?

The number of individuals to include in a research study, the sample size of the study, is an important consideration in the design of many clinical studies. This article reviews the basic factors that determine an appropriate sample size and provides methods for its calculation in some simple, yet common, cases. Sample size is closely tied to statistical power, which is the ability of a study to enable detection of a statistically significant difference when there truly is one. A trade-off exists between a feasible sample size and adequate statistical power. Strategies for reducing the necessary sample size while maintaining a reasonable power will also be discussed.

Humans↗

Incorporating biological knowledge into evaluation of causal regulatory hypotheses.

Biological data can be scarce and costly to obtain. The small number of samples available typically limits statistical power and makes reliable inference of causal relations extremely difficult. However, we argue that statistical power can be increased substantially by incorporating prior knowledge and data from diverse sources. We present a Bayesian framework that combines information from different sources and we show empirically that this lets one make correct causal inferences with small sample sizes that otherwise would be impossible.

Algorithms↗

Evaluation of old and new tests of heterogeneity in epidemiologic meta-analysis.

The identification of heterogeneity in effects between studies is a key issue in meta-analyses of observational studies, since it is critical for determining whether it is appropriate to pool the individual results into one summary measure. The result of a hypothesis test is often used as the decision criterion. In this paper, the authors use a large simulation study patterned from the key features of five published epidemiologic meta-analyses to investigate the type I error and statistical power of five previously proposed asymptotic homogeneity tests, a parametric bootstrap version of each of the tests, and tau2-bootstrap, a test proposed by the authors. The results show that the asymptotic DerSimonian and Laird Q statistic and the bootstrap versions of the other tests give the correct type I error under the null hypothesis but that all of the tests considered have low statistical power, especially when the number of studies included in the meta-analysis is small (<20). From the point of view of validity, power, and computational ease, the Q statistic is clearly the best choice. The authors found that the performance of all of the tests considered did not depend appreciably upon the value of the pooled odds ratio, both for size and for power. Because tests for heterogeneity will often be underpowered, random effects models can be used routinely, and heterogeneity can be quantified by means of R(I), the proportion of the total variance of the pooled effect measure due to between-study variance, and CV(B), the between-study coefficient of variation.

Epidemiologic Methods↗

Evaluation of rodent sperm, vaginal cytology, and reproductive organ weight data from National Toxicology Program 13-week studies.

Sperm morphology and vaginal cytology examinations (SMVCEs), which include evaluations of motility, concentration and head morphology of sperm from the cauda epididymis, and male reproductive organ weight data, were developed by the National Toxicology Program as a screening system for reproductive toxicants. An analysis was conducted of SMVCE studies carried out at the end of fifty 13-week studies (25 for rats, 25 for mice) over a 3-year period. Statistically significant changes in these studies were summarized, as were control data for each male endpoint (mean, SD, 95% confidence limits around the mean, median, and statistical power). Reproductive organ weights (testis, epididymis, cauda epididymis) and sperm motility were the most statistically powerful endpoints evaluated; sperm head morphology may also be a sensitive endpoint for detecting reproductive toxicants. For 24 chemicals tested in both rats and mice, the concordance of results [i.e., no adverse effect in either species, or at least one SMVCE endpoint (not necessarily the same one) adversely affected in both species] was 58%. These data suggest that detection of potential reproductive toxicants might be best when both species are used. Types of sperm head abnormalities and their relative proportion of the total did not differ among control and treatment groups. Estrous cycle data were obtained in the final week of forty-six 13-week studies (23 for mice, 23 for rats). Only 3 chemicals caused an increase in mean cycle length compared with the control group. More data from breeding studies in which female estrous cycle length is measured are needed to assess fully the association of cycle length with reproductive outcome; stages of the estrous cycle are so variable that they may not be useful in assessing potential toxicity. Interlaboratory variability in SMVCE values for many endpoints was documented. Very few of the chemicals that form the basis of this report have been evaluated in definitive reproductive toxicology protocols; a companion paper compares changes in SMVCE endpoints with the outcome of continuous breeding reproduction studies.

Animals↗

Power-law statistics for avalanches in a martensitic transformation.

We devise a two-dimensional model that mimics the recently observed power-law distributions for the amplitudes and durations of the acoustic emission signals observed during martensitic transformation [Vives et al., Phys. Rev. Lett. 72, 1694 (1994)]. We include a threshold mechanism, long-range interaction between the transformed domains, inertial effects, and dissipation arising due to the motion of the interface. The model exhibits thermal hysteresis and, more importantly, it shows that the energy is released in the form of avalanches with power-law distributions for their amplitudes and durations. Computer simulations also reveal morphological features similar to those observed in real systems.

Journal Article↗

Assessing statistical precision, power, and robustness of alternative experimental designs for two color microarray platforms based on mixed effects models.

Recommendations on experimental designs for two color microarray systems have been generally conflicting as they pertain to the general choice between reference and non-reference loop designs. This conflict may currently exist because many previously published assessments may not have effectively connected design layout with the level of biological relative to technical replication. We reassess various reference and non-reference designs for statistical efficiency in terms of standard errors of mean differences, power of test, and robustness using recently developed mixed model software tools. In minimally replicated cases (n = 2), it appears that the reference design outperforms the classical loop design whereby a sample from each animal is used for only one particular array hybridization. Alternatively, the reference design was consistently inferior to those connected loop designs in which a sample from each animal is used in two different hybridizations. Nevertheless, the gap in power between these two designs diminished as the biological to residual variance ratio increased. The statistical efficiency of a single large classical loop design for the comparison of many treatments was demonstrated to be highly sensitive to missing arrays relative to a common reference design (n = 2). However, the use of two loops within an interwoven loop design was shown to be substantially more robust to missing arrays and statistically more efficient relative to a common reference design. Furthermore, the use of more than one loop leads to less disparity in precision and power comparisons between any two treatments.

Animals↗

An efficient test for comparing sequence diversity between two populations.

We address the problem of comparing interindividual genomic sequence diversity between two populations. Although the methods are general, for concreteness we focus on comparing two human immunodeficiency virus (HIV) infected populations. From a viral isolate(s) taken from each individual in a sample of persons from each population, suppose one or multiple measurements are made on the genetic sequence of a coding region of HIV. Given a definition of genetic distance between sequences, the goal is to test if the distribution of interindividual distances differs between populations. If distances between all pairs of sequences within each group are used, then data-dependencies arising from the use of multiple sequences from individuals invalidates the use of a standard two-sample test such as the t-test. Where this problem has been recognized, a typical solution has been to apply a standard test to a reduced dataset comprised of one sequence or a consensus sequence from each patient. Disadvantages of this procedure are that the conclusion of the test depends on the choice of utilized sequences, often an arbitrary decision, and exclusion of replicate sequences from the analysis may needlessly sacrifice statistical power. We present a new test free of these drawbacks, which is based on a statistic that linearly combines all possible standard test statistics calculated from independent sequence subsamples. We describe statistical power advantages of the test and illustrate its use by application to nucleotide sequence distances measured from HIV-1 infected populations in southern Africa (GenBank accession numbers AF110959--AF110981) and North America/Europe. The test makes minimal assumptions, is maximally efficient and objective, and is broadly applicable.

Africa, Southern↗

Novel scaled average bioequivalence limits based on GMR and variability considerations.

PURPOSE: i) To develop novel approaches for the construction of bioequivalence (BE) limits incorporating both the intrasubject variability and the geometric mean ratio (GMR), and ii) to assess the performance of the novel approaches in comparison to several scaled BE procedures and the classic unscaled average BE. METHODS: Plots of the BE limits or the extreme GMR values accepted as a function of the coefficient of variation (CV) were constructed for published and the developed scaled procedures. Two-period crossover BE investigations with 12, 24, or 36 subjects were simulated with assumptions of a CV 10%, 20%, 30%, or 40%. The decline in the percentage of accepted studies was recorded as the true GMR for the two formulations was raised from 1.00 to 1.50. Acceptance of BE was evaluated by published and the developed scaled procedures, and, for comparison, by the unscaled average BE. RESULTS: Two GMR-dependent BE limits are proposed for the evaluation of average BE: i) BELscG1 with Ln(Upper, Lower BE limit) = +/-[(5 - 4GMR)0.496s + Ln(1.25)], and ii) BELscG2 with Ln(Upper, Lower BE limit) = +/-[(3 - 2GMR)(0.496s + Ln(1.25))], where s is the square root of the intrasubject variance. The range of BE limits becomes narrower as GMR values deviate from unity, and increases with variability. The two new approaches exhibit the highest statistical power at low CV values. At high levels of variability, BELscG1 and BELscG2 show high statistical power, as well as the lowest percentages of acceptance among the scaled methods when GMR = 1.25. The latter becomes more obvious when a large number of subjects is incorporated in the studies. CONCLUSIONS: The GMR and CV estimates of the BE study can be used in conjunction with the GMR vs. CV plot for the assessment of average BE. The new approaches, BELscG1 and BELscG2, appear to be highly effective at all levels of variation investigated.

Algorithms↗

A strategy to discover genes that carry multi-allelic or mono-allelic risk for common diseases: a cohort allelic sums test (CAST).

A method is described to discover if a gene carries one or more allelic mutations that confer risk for any specified common disease. The method does not depend upon genetic linkage of risk-conferring mutations to high frequency genetic markers such as single nucleotide polymorphisms. Instead, the sums of allelic mutation frequencies in case and control cohorts are determined and a statistical test is applied to discover if the difference in these sums is greater than would be expected by chance. A statistical model is presented that defines the ability of such tests to detect significant gene-disease relationships as a function of case and control cohort sizes and key confounding variables: zygosity and genicity, environmental risk factors, errors in diagnosis, limits to mutant detection, linkage of neutral and risk-conferring mutations, ethnic diversity in the general population and the expectation that among all exonic mutants in the human genome greater than 90% will be neutral with regard to any effect on disease risk. Means to test the null hypothesis for, and determine the statistical power of, each test are provided. For this "cohort allelic sums test" or "CAST", the statistical model and test are provided as an Excel program, CASTAT(c) at . Based on genetics, technology and statistics, a strategy of enumerating the mutant alleles carried in the exons and splice sites of the estimated approximately 25,000 human genes in case cohort samples of 10,000 persons for each of 100 common diseases is proposed and evaluated: A wide range of possible conditions of multi-allelic or mono-allelic and monogenic, multigenic or polygenic (including epistatic) risk are found to be detectable using the statistical criteria of 1 or 10 "false positive" gene associations approximately 25,000 gene-disease pair-wise trials and a statistical power of >0.8. Using estimates of the distribution of both neutral and gene-inactivating nondeleterious mutations in humans and the sensitivity of the test to multigenic or multicausal risk, it is estimated that about 80% of nullizygous, heterozygous and functionally dominant gene-common disease associations may be discovered. Limitations include relative insensitivity of CAST to about 60% of possible associations given homozygous (wild type) risk and, more rarely, other stochastic limits when the frequency of mutations in the case cohort approaches that of the control cohort and biases such as absence of genetic risk masked by risk derived from a shared cultural environment.

Alleles↗

Why rehabilitation research does not work (as well as we think it should)

Establishing treatment effectiveness is a high priority for rehabilitation research. The use of traditional quantitative null hypotheses to achieve this priority is reviewed. Three problems are identified in the analysis and interpretation of investigations based on statistical testing of hypotheses: (1) confusion of clinical and statistical significance, (2) low statistical power in detecting clinically important results, and (3) a failure to understand the importance of replication in developing a knowledge base for rehabilitation practice. Technical aspects associated with each problem are reviewed and examples presented illustrating the impact of low statistical power and the results of misinterpreting statistical significance tests. Several specific recommendations are made to improve the clinical usefulness of quantitative research conducted in rehabilitation.

Clinical Trials as Topic↗

Is hysterectomy a risk factor for vaginal cancer?

Several recent case series have called attention to a possible association between previous hysterectomy and the subsequent development of vaginal cancer. To study this relationship, we compared 49 patients with vaginal cancer with 49 controls matched for age, race, and prior cervical dysplasia or neoplasia. Patients and controls were alike in terms of exposure to estrogens. Twenty-four patients (49%) had had prior hysterectomies, of which 13 (27%) were for benign disease. Similarly, 24 controls had a history of a hysterectomy. The matched-pairs odds ratio relating prior hysterectomy to vaginal cancer was 1.00 based on these data, with a 95% confidence interval of 0.47 to 2.12. In the subsample of women without a history of cervical disease, a similar odds ratio appeared. Although the study sample size did not permit exclusion of a twofold increase in risk, the statistical power to detect an actual odds ratio of 2.5 is 76%. At this level of statistical power, our data suggest that hysterectomy has a low probability of being a risk factor for vaginal cancer when age and cervical disease are controlled for. In the absence of such a relationship, screening for vaginal cancer does not appear to be necessary for women who have had a hysterectomy for benign disease.

Adult↗

The usefulness of matched pair randomization for medical practice-based research.

To be feasible, study designs for most intervention research in primary care settings must limit the number of participating physicians, without sacrificing the statistical power required to test the research hypotheses. A model was developed to examine sample size and statistical power requirements when using the physicians' practice as the unit of analysis. Randomized designs using either matched or unmatched samples of practices were compared under varying conditions. When baseline variability is small or the number of practice pairs is large, matching at best marginally increases power. However, in the typical case when baseline variability is large or the number of practice pairs is small, matching substantially increases the power to find intervention effects with a smaller sample. Thus, matching prior to randomization could improve the design of many intervention studies in primary care settings.

Humans↗

Biomonitoring of organochlorines in women with benign and malignant breast disease.

Established risk factors for breast cancer explain breast cancer risk only partially. Organochlorines are considered to be a possible cause for hormone-dependent cancers. A hospital-based case-control study, the first from India, was conducted among 50 women undergoing surgery for breast disease to examine the association between organochlorine exposure and breast cancer risk. Blood, tumor, and surrounding adipose tissue of the breast were collected from the subjects with benign (control) and malignant breast (study) lesions and analyzed to determine organochlorine insecticides using a gas-liquid chromatograph equipped with an electron capture detector. The alpha, beta, gamma, and delta isomers of hexachlorocyclohexane (HCH), p,p'-dichlorodiphenyltrichloroethane (DDT), o,p'-DDT, p,p-dichlorodiphenyldichloroethylene, and p,p'-dichlorodiphenyldichloroethane were frequently detected in three specimens. Total HCH and total DDT levels were higher in the blood of the study group (25 cases) than in those of the controls (25 cases) with only gamma-HCH being significantly different (P<0.05). However, both total HCH and total DDT were higher in the tumor tissues of the controls than in those of the study group; gamma-HCH was significantly different (P<0.05). The level of total HCH (alpha-HCH was significantly different, P<0.05) was higher in the breast adipose tissue of the study group, whereas total DDT was higher in the breast adipose tissue of the control group. The distribution of known confounders of breast cancer including age, body mass index, age at menarche and menopause, duration of breast feeding, and family history related to breast disease did not differ significantly between benign and malignant groups. This pilot study with limited statistical power does not support a positive association between exposure to organochlorines and risk of breast cancer but paves the way for a larger Indian study with greater statistical power encompassing different regions of the country to enable statistically sound conclusions.

Adipose Tissue↗

Phosphor-stimulated computed cephalometry: reliability of landmark identification.

The aim of this randomized, controlled, prospective study was to determine the reliability of computed lateral cephalometry (Fuji Medical Systems, Tokyo, Japan) in terms of landmark identification compared to conventional lateral cephalometry (CAWO, Schrobenhausen, Germany). To assess the reliability of landmark identification on lateral cephalographs, 20 computed images, taken at 30 per cent reduced radiation (70 kV, 15 mA, 0.35 s) were compared to 20 conventional images (70 kV, 15 mA, 0.5 s). The 40 lateral cephalographs were taken from 20 orthodontic patients at immediate post-treatment and 1 year after retention. The order and type of imaging was randomized. Five orthodontists identified eight skeletal, four dental and five soft tissue landmarks on each of the 40 films. The error of identification was analysed in the XY Cartesian co-ordinate following digitization. Skeletal landmarks exhibited characteristic dispersion with respect to the Cartesian co-ordinates. Root apices were more variable than crown tips. Soft tissue landmarks were more consistent in the X co-ordinate. Two-way ANOVA shows that there is no significant difference between the two imaging systems in both co-ordinates (P > 0.05). Moreover, the differences are generally small (< 0.5 mm), and are unlikely to be of clinical significance. Most of the variables attained statistical power of at least 0.8 in the X-co-ordinate while only the dental landmarks achieved statistical power of at least 0.78 in the Y-co-ordinate. Based on the results of the study: (1) computed lateral cephalographs can be taken at 30 per cent radiation reduction, compared to conventional lateral cephalograph; (2) each anatomical landmark exhibits its characteristic dispersion of error in both the Cartesian co-ordinates; (3) there is no trend between the two imaging systems, with equivocal result, and none of the landmarks attained statistical significance when both raters and imaging systems are considered as factorial variables; (4) the random error of raters in landmark identification after replicate tracing was highlighted and needs to be taken into consideration in all studies involving landmark identification.

Adolescent↗

Enhancing power while controlling family-wise error: an illustration of the issues using electrocortical studies.

This study examined the relative family-wise error (FWE) rate and statistical power of multivariate permutation tests (MPTs), Bonferroni-adjusted alpha, and uncorrected-alpha tests of significance for bivariate associations. Although there are many previous applications of MPTs, this is the first to apply it to testing bivariate associations. Electrocortical studies were selected as an example class because the sample sizes that are typical of electrocortical studies published in 2001 and 2002 are small and their multiple significance tests are typically nonindependent. Because Bonferroni adjustments assume independent predictors, we expected that MPTs would be more powerful than the Bonferroni adjustment. Results support the following conclusions: (a) failure to control for multiple significance testing results in unacceptable FWE rates, (b) the FWE rate for the MPTs approximated the alpha set for the analyses, and (c) the statistical power advantage that MPTs provide over Bonferroni adjustments is important when using small sample sizes such as those that are typical of recent electrocortical studies.

Analysis of Variance↗

Effect of continuous versus dichotomous outcome variables on study power when sample sizes of orthopaedic randomized trials are small.

It is often not feasible to conduct large trials in orthopaedic surgery. Therefore, surgeons must identify strategies to optimize the statistical power of their smaller studies. The aim of this study was to compare study power in randomized trials with continuous versus dichotomous outcome variables. We performed a systematic review of the literature to identify randomized trials in orthopaedic trauma. Of these, we examined only those trials with small sample sizes (50 patients or less). The outcomes in each eligible study were categorized as continuous or dichotomous. Standard power calculations were performed for each study, and comparisons were made between continuous and dichotomous outcome variables. We identified 196 randomized trials in orthopaedic trauma. Of these, 76 trials had a sample size of 50 patients or fewer (29 trials with continuous outcomes, 47 trials with dichotomous outcomes). Studies that reported continuous outcomes had a significantly higher mean power than those that reported dichotomous variables (power 49% vs 38%, p=0.042). Twice as many trials with continuous outcome variables reached acceptable levels of study power (i.e. >80% power) when compared with trials with dichotomous variables (37% vs 18.6%, p=0.04). When orthopaedic surgeons anticipate small sample sizes for their study, they can optimize their study's statistical power by choosing a continuous outcome variable.

Cluster Analysis↗

Asymmetric integration of various cancer datasets for identifying risk-associated variants and genes.

MOTIVATION: Cancer genomic research provides an opportunity to identify cancer risk-associated genes, but often suffers from undesirable low statistical power due to a limited sample size. Integrated analysis with different cancers has the potential to enhance statistical power for identifying pan-cancer risk genes. However, substantial heterogeneity across various cancers makes this challenging. RESULTS: Recently, a novel asymmetric integration method was developed that can deal with data heterogeneity and exclude unhelpful datasets from the analysis. We adapted and applied this method to integrate genotype datasets with matched case and control individuals from the Michigan Genomics Initiative, using each cancer as the primary dataset of interest and the other cancers as auxiliary datasets, respectively. Conditional logistic regression models were coupled with the asymmetric integrated framework to handle the matched case-control study design and permutation tests were performed to control for false discovery rates (FDRs). At the same FDR level, the integrated analysis found more potential genetic variants and genes that are associated with the risks of various cancers, showcasing the promise of the proposed approach for integrated analysis of cancer datasets. AVAILABILITY AND IMPLEMENTATION: Our method is available as source code at https://github.com/rxxwang/integrate_cancer.

Journal Article↗

The genealogy of a sequence subject to purifying selection at multiple sites.

We investigate the effect of purifying selection at multiple sites on both the shape of the genealogy and the distribution of mutations on the tree. We find that the primary effect of purifying selection on a genealogy is to shift the distribution of mutations on the tree, whereas the shape of the tree remains largely unchanged. This result is relevant to the large number of coalescent estimation procedures, which generally assume neutrality for segregating polymorphisms--applying these estimators to evolutionarily constrained sequences could lead to a significant degree of bias. We also estimate the statistical power of several neutrality tests in detecting weak to moderate purifying selection and find that the power is quite good for some parameter combinations. This result contrasts with previous studies, which predicted low statistical power because of the minor effect that weak purifying selection has on the shape of a genealogy. Finally, we investigate the effect of Hill-Robertson interference among linked deleterious mutations on patterns of molecular variation. We find that dependence among selected loci can substantially reduce the efficacy of even fairly strong purifying selection.

Evolution, Molecular↗