Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

The RID2 biofidelic rear impact dummy: a pilot study using human subjects in low speed rear impact full scale crash tests.

STUDY DESIGN: Human subjects and the recently developed RID2 rear impact crash test dummy were exposed to a series of full scale, vehicle-to-vehicle crash tests. OBJECTIVE: To evaluate the biofidelity of the RID2 anthropometric test dummy on the basis of calculated neck injury criterion (NIC) values by comparing these values to those obtained from human subjects exposed in the very same crashes. SUMMARY OF BACKGROUND DATA: The widely used and familiar hybrid III dummy has been said to lack biofidelity in the special application of low speed rear impact crashes. Several attempts have been made to modify this dummy with only marginal success. Two completely new dummies have been developed; the BioRID and the RID2. Neither have been tested under real world crash boundary conditions in side-by-side comparisons with live human subjects. METHODS: Volunteer subjects, including a 50th percentile male, a 95th percentile male, and a 50th percentile female, were placed in the driver's seat of a vehicle and subjected to a series of three low speed rear impact crashes each. The RID2 dummy, which is modeled after a 50th percentile male, was placed in the passenger seat in each case. Both subjects and dummy were fully instrumented and acceleration-time histories were recorded. From this data, velocities of the heads and torsos were determined and both were used to calculate the NIC values for both crash test subjects and the RID2. RESULTS: The RID2 demonstrated generally higher head accelerations and NIC values than those of the human subjects. Most of the observed variations might be explained on the basis of differing head restraint geometry, posture, and body size. The RID2 NIC values compared most favorably with those of the 50th percentile male subject. For the whole group, the correlations between RID2 and human subjects did not reach statistical significance. CONCLUSIONS: The small number of test subjects and crash tests limited the statistical power of this pilot study, and the correlation between the RID2 and human subject NIC values were not statistically significant. The overall qualitative performance and biofidelity of the RID2 was reasonable when compared with the male human 50th percentile subject. Its overall higher ranges of head acceleration and calculated NIC values compared to all of the human subjects were generally consistent. This condition could likely be improved by increasing the stiffness of the RID2 neck. Biofidelic validation of the RID2 will require ongoing testing using a larger number of human subjects and varying boundary conditions. The results of this pilot study, while encouraging, should be considered preliminary.

Acceleration↗

Stroke assessment: morphometric infarct size versus neurologic deficit.

We presently examine the relation between histologic infarct size and neurologic deficit as endpoints and seek to clarify their sensitivity in defining stroke outcome. Neurologic deficits of 76 cats subjected to middle cerebral artery occlusion were assessed daily and correlated with the corresponding infarct sizes determined morphometrically after 2 weeks' survival. A five-item neurologic deficit score included the time elapsed until hemiparesis, and forced circling resolved (if ever), presence of impaired placing reactions and time elapsed until able to stand and being alert. We then evaluated the two endpoints' statistical powers to detect group differences using two sets of comparison groups. The neurologic deficit score correlated well with infarct size (r = 0.76, p < 0.001) and each of the individual deficit score components named above, in turn, correlated with decreasing power with infarct size. Even so, the number of study subjects required to achieve the same level of statistical significance in assessing group differences was two-fold greater when using the neurologic deficit than the infarct size data: Group sizes of eight and five animals were sufficient for significant infarct size differences while the groups needed be expanded to 15 and 10 animals to similarly achieve significant neurologic score differences. Thus, infarct size emerges as a more sensitive measure of stroke outcome than does the assessment of neurologic deficits.

Animals↗

Partner's adjustment to breast cancer: a critical analysis of intervention studies.

Partners of breast cancer patients do not have resources available for dealing with their concerns. An analysis of intervention studies with partners was conducted, spanning research published from 1966 to 2004. Although there is considerable descriptive research documenting the need for partner interventions in the context of breast cancer, only 4 studies met criteria for inclusion in this analysis. Two studies reported limited intervention efficacy, but none incorporated all characteristics of a rigorous clinical trial with adequate power to fully test the intervention. Future intervention research should incorporate randomized, controlled clinical trial designs; have adequate statistical power; clearly report eligibility criteria; delineate theoretically based, fully explicated, and consistently delivered interventions; and use outcome measures that are sensitive to empirically derived partner-adjustment issues.

Adaptation, Psychological↗

Variations in medical care for HIV-related Pneumocystis carinii pneumonia: a comparison of process and outcome at two hospitals.

BACKGROUND: Institutional variation in the quality of medical care may be evaluated by examining process measures, such as use of diagnostic procedures or treatment modalities, or outcome measures, such as mortality. We undertook this study to examine variations in both process and outcome of care for patients with HIV-related Pneumocystis carinii pneumonia (PCP) at two geographically diverse, HIV-experienced, public municipal hospitals. DESIGN: Retrospective review of hospitalized patients diagnosed as having PCP cared for at two municipal hospitals from 1988 to 1990. At hospital A, charts of all patients diagnosed as having PCP were abstracted (n=209); at hospital B, a random sample of 15% were abstracted (=136). RESULTS: Among all hospitalized patients diagnosed as having PCP, the frequency of making a definitive diagnosis of PCP (as opposed to treating empirically) differed markedly at the two hospitals (85% in hospital A vs 26% in hospital B; p<0.001), as did the use of intensive care (18% vs 3%; p<0.001) and "do-not-resuscitate" orders (39% vs 14%; p<0.001), although the timing of starting anti-Pneumocystis medications (89% vs 88% within the first 2 hospital days) and the use of corticosteroids (21% vs 23%) were similar. Despite differences in the process of care, survival rates were similar at the two institutions (75% vs 76%; p=0.8) and remained similar when logistic regression was used to control for demographic variables and severity of illness (odds ratio for survival, hospital B vs A, 1.2 [95% confidence interval, 0.7, 2.0]). The 95% confidence intervals (0.7, 2.0), however, were consistent with a considerable (and clinically significant) disparity in survival (from 30% lower to a twofold higher odds of survival). Sample size calculations showed that a sample of 10 cases in each hospital would be required to detect the observed difference in definitive diagnosis rates (85% vs 26%), but 722 cases in each hospital would be required to detect a relevant difference in mortality. CONCLUSIONS: The process of care for hospitalized patients with PCP in these two institutions differed considerably, but the survival rates were not significantly different, even after adjusting for confounding factors. While sample sizes available at the individual institutions were sufficient for evaluation of the process of care, they did not provide the power necessary to evaluate outcomes. Comparisons of outcomes such as mortality between individual hospitals may not have the statistical power to exclude important differences.

AIDS-Related Opportunistic Infections↗

Controversies in antimicrobial therapy: critical analysis of clinical trials.

Problems with design and statistical evaluation of clinical efficacy trials of antimicrobial agents are reviewed. Of the three major criteria used for evaluating antimicrobial agents (efficacy, toxicity, cost), the most important is efficacy. Clinical efficacy can be evaluated in uncontrolled or controlled clinical trials. Uncontrolled trials are often conducted to satisfy Food and Drug Administration requirements during premarketing testing; the response rate is typically high because only patients with susceptible infections may be treated and large doses are given. Controlled antibiotic trials should be randomized, blinded, parallel comparisons of an investigational agent versus the best available agent at an accepted dose. However, interpretation of these studies is frequently clouded by poor study design, small sample sizes, and heterogeneous patient populations. Controlled trials are usually centered around a null hypothesis (i.e., that no difference will be found between the agents being compared). All conclusions (to reject or not reject the null hypothesis) should be carefully evaluated by clinicians seeking to apply the available data to patient care. Researchers can incorrectly conclude that two therapies have equal efficacy because of insufficient statistical power (i.e., small sample size) or poor study design. Likewise, researchers may incorrectly conclude that there is a statistical difference between two therapies because of poor design or improper sample selection. For the clinician, clinical relevance takes precedence over statistical significance. Before the results of a study are allowed to affect drug use in an institution, strong similarities between subjects and methods in the study and patients and care in the institution should be demonstrated.

Anti-Bacterial Agents↗

Statistical significance versus clinical relevance. Part I. The essential role of the power of a statistical test.

When comparing two treatment groups, hypothesis testing is widely used. However, clinical trialists should be more interested in statistical methods which elicit the magnitude of the differences between treatment groups, rather than a simple indication of whether or not the differences are statistically significant. Statistical significance does not necessarily imply clinical relevance. If the true difference between two treatment groups is so small that it is clinically irrelevant, a sample size can be found for which this difference is statistically significant. On the other hand, if the difference between treatment groups is statistically non-significant, it may still be clinically important. The limitations of conventional hypothesis testing of equal true means as such are highlighted. The need to control the power of the test--which takes into account the difference in treatment means which is considered important (clinically relevant) by the researcher--is discussed.

Clinical Trials as Topic↗

Power estimation of multiple SNP association test of case-control study and application.

At the current stage, a large number of single nucleotide polymorphisms (SNPs) have been deployed in searching for genes underlying complex diseases. A powerful method is desirable for efficient analysis of SNP data. Recently, a novel method for multiple SNP association test using a combination of allelic association (AA) and Hardy-Weinberg disequilibrium (HWD) has been proposed. However, the power of this test has not been systematically examined. In this study, we conducted a simulation study to further evaluate the statistical power of the new procedure, as well as of the influence of the HWD on its performance. The simulation examined the scenarios of multiple disease SNPs among a candidate pool, assuming different parameters including allele frequencies and risk ratios, dominant, additive, and recessive genetic models, and the existence of gene-gene interactions and linkage disequilibrium (LD). We also evaluated the performance of this test in capturing real disease associated SNPs, when a significant global P value is detected. Our results suggest that this new procedure is more powerful than conventional single-point analyses with correction of multiple testing. However, inclusion of HWD reduces the power under most circumstances. We applied the novel association test procedure to a case-control study of preterm delivery (PTD), examining the effects of 96 candidate gene SNPs concurrently, and detected a global P value of 0.0250 by using Cochran-Armitage chi(2)s as "starting" statistics in the procedure. In the following single point analysis, SNPs on IL1RN, IL1R2, ESR1, Factor 5, and OPRM1 genes were identified as possible risk factors in PTD.

Algorithms↗

Willingness to pay and size of health benefit: an integrated model to test for 'sensitivity to scale'.

A key theoretical prediction concerning willingness to pay is that it is positively correlated with benefit size and is assessed by testing the 'sensitivity to scale (scope)'. 'External' (between-sample) sensitivity tests are usually regarded as less powerful than 'internal' (within-subject) tests. However, the latter may suffer from 'anchoring' effects. This paper studies the statistical power of these tests by questioning the distributional assumption of empirical data. We present an integrated model to capture both internal and external variations, while controlling for sample heterogeneity, applied to data from a survey estimating the value of reducing symptom-days. Results indicate that once data is properly transformed, WTP becomes 'scale sensitive' and consistent with diminishing marginal utility theory.

Adult↗

Non-invasive echocardiographic studies in mice: influence of anesthetic regimen.

Transgenic murine models of cardiovascular disease offer great potential insights regarding mechanisms of human disease, but efficient and reliable methods for phenotype evaluation are necessary. We employed non-invasive echocardiography to evaluate hemodynamic parameters in mice, and evaluated statistical reliability of these parameters with respect to anesthesia regimen. Male CF-1 mice received inhaled halothane (0.25-0.75% in 95% O2) or ketamine/xylazine (80/10 mg/kg i.p.) and 2-dimensional, M-mode, and Doppler ultrasound imaging were used to assess cardiac contractility and aortic flow velocities. Halothane was more convenient and reliable with respect to rate of induction, reversal, and control of anesthetic depth. At comparable levels of anesthesia, ketamine/xylazine produced significant reductions in heart rate (308 +/- 14 vs. 501 +/- 14 bpm, p<0.001), left ventricular fractional shortening (41.7 +/- 1.3 vs. 49.3 +/- 1.0%, p<0.001), and cardiac output (7.6 +/- 0.5 vs. 11.5 +/- 0.6 ml/min, p<0.001) when compared to halothane inhalation. No change in stroke volume or peak aortic velocity was observed. Correlation analyses revealed highly significant positive relationships between heart rate and fractional shortening (r=0.61, p<0.002) and cardiac output (r=0.88, p<0.001) but no relation to stroke volume or aortic velocity. Variability of intra-animal and intragroup parameter estimation were frequently 2-fold larger for ketamine/xylazine anesthesia vs. halothane. Statistical power analysis showed the increased measurement error for ketamine/xylazine leads to much larger numbers of mice/group to achieve identical statistical sensitivity. These data further illustrate the feasibility of echocardiography for rapid, non-invasive cardiovascular assessment in mice. However, several obtainable parameters are highly sensitive to both heart rate and anesthetic used and the choice and control of anesthetic are critical for physiologically relevant performance parameters and maximal ability to detect statistical differences among groups. Thus, for these non-invasive studies, inhalation anesthesia with agents such as halothane is superior to anesthesia induced by ketamine/xylazine administration.

Adrenergic alpha-Agonists↗

Statistical aspects of controlled single subject trials.

Randomized controlled trials in groups and single subjects differ in several statistical aspects. In group trials the experimental unit is a randomly selected subject from a predefined population and this subject is randomly assigned to a treatment. Outcome is confined to average effects which can be generalized to the specific population, but which do not necessarily apply to individual persons. In single subject trials the experimental unit is a treatment period and each treatment period is randomly allocated in a multiple cross-over sequence of periods. The single subject is only representative of itself, but similar responses in corresponding single subject trials may justify careful extrapolation of the results. Single subject trials have a high risk of Type II errors. However, the randomization procedure chosen and the type of statistical test applied may enhance the statistical power of such trials. Internal validity depends on modeling the trial design to the clinical features, drug properties and statistical requirements, while reliability is determined by the reproducibility of the trial response.

Humans↗

Childhood leukemia near nuclear plants in the United Kingdom: the evolution of a systematic approach to studying rare disease in small geographic areas.

A cluster of childhood leukemia in a village near a nuclear plant in northern England prompted further studies of cancer in the vicinity of other nuclear plants in the United Kingdom. These studies demonstrated that the risk of childhood leukemia was increased near certain other nuclear plants. Although the reasons for the increase are still unclear, the scientific debate stimulated by these findings has clarified some of the special methodological problems encountered when studying rare diseases in small areas. Firstly, unless a specific hypothesis is defined in advance, the relevance of a single geographic cluster of disease can rarely be interpreted. Even when a prior hypothesis exists, the small number of cases which generally occur in a small area make the findings highly sensitive to reporting, diagnostic, or classification errors. The statistical power of such investigations is also usually low and only marked increases in risk can be detected. Furthermore, conventional statistical tests may be inappropriate if the underlying spatial distribution of the disease is not random; and little is known about the background distribution of disease in small areas. Investigations of specific hypotheses about defined sources of environmental contamination, especially if they can be replicated, are more likely to result in conclusive findings that are in-depth studies of individual clusters.

Adolescent↗

Prospective, randomized, crossover comparison of sublingual apomorphine (3 mg) with oral sildenafil (50 mg) for male erectile dysfunction.

PURPOSE: We established the efficacy and safety of sublingual apomorphine compared with oral sildenafil in comparable groups of patients with erectile dysfunction (ED). MATERIALS AND METHODS: This prospective, randomized, crossover study included 77 heterosexual men with ED of various etiologies and severities. A total of 62 men were randomized but only 34 were evaluable for efficacy and tolerability. The study started with a run-in period of 2 to 4 weeks. The first 4 weeks of treatment were followed by a washout period of 4 weeks, after which patients changed to the alternate treatment for an additional 4-week period. The sequence of the 2 treatments was established by a randomization list in blocks in closed packets. The primary efficacy end point was the percent of attempts resulting in erection firm enough for intercourse. Additional variables were the percent of attempts resulting in intercourse and improvement in ED, as evaluated by the erectile function domain score of the International Index of Erectile Function questionnaire. RESULTS: Sildenafil was significantly more effective than apomorphine in regard to the percent of attempts resulting in erection firm enough for intercourse (85% vs 44%, p <0.0001) and actually resulting in intercourse (81% vs 43%, p <0.0001) as well as erectile function evaluated by the erectile function domain score of the International Index of Erectile Function (p <0.001). The incidence of adverse events was not significantly different for the 2 drugs. Although the number of patients was small, this study had strong statistical power due to the striking difference in results. CONCLUSIONS: Sildenafil was significantly more effective than apomorphine for ED. No statistical difference in adverse events was noted.

Administration, Oral↗

A statistical examination of historical controls for mouse bone marrow cytogenetic assays.

Data from 1,111 controls from assays run over 11 years are examined to determine a most powerful statistical procedure for detecting a mutagenic effect. It is concluded that the data do not show a constant probability of chromosomal abnormalities and that the data do not fit a simple Poisson or binomial distribution. Empirical Bayes techniques are used to derive a test that declares an effect if three or more cells in a group of 50 tested contain abnormalities.

Analysis of Variance↗

Genome scan meta-analysis for hypertension.

BACKGROUND: Genome scans for hypertension have yielded inconsistent results. The non-replication of significant or suggestive linkage might be due to lack of power of individual studies. Here, we conducted a genome scan meta-analysis for hypertension in an attempt to increase statistical power and to enhance evidence of linkage. METHODS: A newly developed Genome Search Meta-analysis (GSMA) method was applied to pool the results obtained from six scans reported in five papers. RESULTS: Our analysis did not find any regions with genome-wide significant linkage to hypertension. We did identify several regions with suggestive linkage, including 2p, 5q, 6q, 8p, 9p, 9q, and 11q. CONCLUSIONS: It seems that no region has a uniformly large impact on hypertension and that susceptibility genes for hypertension may be very difficult to detect.

Blood Pressure↗

Using cluster random assignment to measure program impacts. Statistical implications for the evaluation of education programs.

This article explores the possibility of randomly assigning groups (or clusters) of individuals to a program or a control group to estimate the impacts of programs designed to affect whole groups. This cluster assignment approach maintains the primary strength of random assignment--the provision of unbiased impact estimates--but has less statistical power than random assignment of individuals, which usually is not possible for programs focused on whole groups. To explore the statistical implications of cluster assignment, the authors (a) outline the issues involved, (b) present an analytic framework for studying these issues, and (c) apply this framework to assess the potential for using the approach to evaluate education programs targeted on whole schools. The findings suggest that cluster assignment of schools holds some promise for estimating the impacts of education programs when it is possible to control for the average performance of past student cohorts or the past performance of individual students.

Bias↗

Metaanalysis: methodology, biostatistic rules, application in onco-hematology.

Metaanalysis are proposed in order to increase statistical power and draw better conclusion about the efficacy of a treatment. After inclusion of trials corresponding to an objective, a statistical procedure is performed permitting to obtain an estimate of the common effect with the confidence interval after model validation. The main problem is posed by the selection of trials in order to avoid positive bias linked to the greater number of positive published studies. Application of metaanalysis in oncohematology should be proposed in case of weakness size of population in the individual trials.

Algorithms↗

A likelihood-based method for testing for nonstochastic variation of diversification rates in phylogenies.

Observed variations in rates of taxonomic diversification have been attributed to a range of factors including biological innovations, ecosystem restructuring, and environmental changes. Before inferring causality of any particular factor, however, it is critical to demonstrate that the observed variation in diversity is significantly greater than that expected from natural stochastic processes. Relative tests that assess whether observed asymmetry in species richness between sister taxa in monophyletic pairs is greater than would be expected under a symmetric model have been used widely in studies of rate heterogeneity and are particularly useful for groups in which paleontological data are problematic. Although one such test introduced by Slowinski and Guyer a decade ago has been applied to a wide range of clades and evolutionary questions, the statistical behavior of the test has not been examined extensively, particularly when used with Fisher's procedure for combining probabilities to analyze data from multiple independent taxon pairs. Here, certain pragmatic difficulties with the Slowinski-Guyer test are described, further details of the development of a recently introduced likelihood-based relative rates test are presented, and standard simulation procedures are used to assess the behavior of the two tests in a range of situations to determine: (1) the accuracy of the tests' nominal Type I error rate; (2) the statistical power of the tests; (3) the sensitivity of the tests to inclusion of taxon pairs with few species; (4) the behavior of the tests with datasets comprised of few taxon pairs; and (5) the sensitivity of the tests to certain violations of the null model assumptions. Our results indicate that in most biologically plausible scenarios, the likelihood-based test has superior statistical properties in terms of both Type I error rate and power, and we found no scenario in which the Slowinski-Guyer test was distinctly superior, although the degree of the discrepancy varies among the different scenarios. The Slowinski-Guyer test tends to be much more conservative (i.e., very disinclined to reject the null hypothesis) in datasets with many small pairs. In most situations, the performance of both the likelihood-based test and particularly the Slowinski-Guyer test improve when pairs with few species are excluded from the computation, although this is balanced against a decline in the tests' power and accuracy as fewer pairs are included in the dataset. The performance of both tests is quite poor when they are applied to datasets in which the taxon sizes do not conform to the distribution implied by the usual null model. Thus, results of analyses of taxonomic rate heterogeneity using the Slowinski-Guyer test can be misleading because the test's ability to reject the null hypothesis (equal rates) when true is often inaccurate and its ability to reject the null hypothesis when the alternative (unequal rates) is true is poor, particularly when small taxon pairs are included. Although not always perfect, the likelihood-based test provides a more accurate and powerful alternative as a relative rates test.

Biodiversity↗

Efficiency and power in genetic association studies.

We investigated selection and analysis of tag SNPs for genome-wide association studies by specifically examining the relationship between investment in genotyping and statistical power. Do pairwise or multimarker methods maximize efficiency and power? To what extent is power compromised when tags are selected from an incomplete resource such as HapMap? We addressed these questions using genotype data from the HapMap ENCODE project, association studies simulated under a realistic disease model, and empirical correction for multiple hypothesis testing. We demonstrate a haplotype-based tagging method that uniformly outperforms single-marker tests and methods for prioritization that markedly increase tagging efficiency. Examining all observed haplotypes for association, rather than just those that are proxies for known SNPs, increases power to detect rare causal alleles, at the cost of reduced power to detect common causal alleles. Power is robust to the completeness of the reference panel from which tags are selected. These findings have implications for prioritizing tag SNPs and interpreting association studies.

Algorithms↗