Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Pair chart test for an early survival difference.

The log-rank test is commonly used in comparing survival distributions between treatment and control groups in clinical trials. However, in many studies, the treatment is only effective at the early stage of the trial. Especially when the two survival curves cross, the log-rank test has a low statistical power to show the survival difference. We propose a test statistic for detecting such an early difference between the two treatment arms. The new test has an intuitive geometric interpretation based on a pair chart and is shown to have more power than the log-rank test when the treatment effect only appears in the early phase of the study. This advantage is evaluated for finite sample sizes in simulation studies. Finally, the proposed method is illustrated with a real data example of patients with gastric cancer.

Biometry↗

The use of predicted confidence intervals when planning experiments and the misuse of power when interpreting results.

Although there is a growing understanding of the importance of statistical power considerations when designing studies and of the value of confidence intervals when interpreting data, confusion exists about the reverse arrangement: the role of confidence intervals in study design and of power in interpretation. Confidence intervals should play an important role when setting sample size, and power should play no role once the data have been collected, but exactly the opposite procedure is widely practiced. In this commentary, we present the reasons why the calculation of power after a study is over is inappropriate and how confidence intervals can be used during both study design and study interpretation.

Bayes Theorem↗

Power comparisons between the TDT and two likelihood-based methods.

We compare the statistical power of the transmission disequilibrium test (TDT) with that of two likelihood-based linkage tests, the classical LOD score and a modified LOD score in which a linkage disequilibrium (LD) parameter is incorporated into the likelihood (LD-LOD). We hypothesize that, when LD is present, the LD-LOD will have the greatest power of the three tests because the TDT breaks a multiplex pedigree into triads, and the LOD score has previously been shown to have lower power when LD is present but not accounted for. We test this hypothesis using a simulation study in which we generate affected sib-pair (ASP) pedigrees under a range of genetic models, varying the genotypic relative risk (GRR) from 6 to 16. Because the likelihood-based tests require that a genetic model be specified, we compare the tests under two scenarios. First, we assume the true genetic model in the analysis, and second, we compare the tests when the LD-LOD (LOD) is maximized over two wrong genetic models. For the generating models we considered, we find that the LD-LOD has greater power than the TDT even when the genetic models is mis-specified and the results corrected for multiple tests. Extreme differences occur under the multiplicative and dominant models, for which the difference in power is as high as 40% at complete LD. The LOD score provides the lowest power in the presence of LD for the range of GRR considered here.

Adult↗

More power to you: simple power calculations for treatment effects with one degree of freedom.

Although numerous computer programs for statistical power analysis are available, power is an under-used aspect of experimental analysis, perhaps because of the perceived difficulty of performing the necessary calculations or because existing computer software can be expensive or complicated to learn. For single-degree-of-freedom tests, however, it is possible to calculate power in a straightforward manner, using the t distribution. Because these calculations are based on t, they use easily understood and readily available quantities. These calculations can be performed with a desk calculator; we also present a simple-to-use program called MorePower that will perform the necessary calculations. The straightforward nature of the calculations potentially will enable more researchers to consider issues of power when planning and reporting their experiments.

Humans↗

Randomized clinical comparison of Healon GV and Viscoat.

PURPOSE: To compare the ability of Healon GV (sodium hyaluronate 1.4%) and Viscoat (sodium chondroitin sulfate 4.0%-sodium hyaluronate 3.0%) to protect the corneal endothelium during endocapsular phacoemulsification and foldable intraocular lens (IOL) implantation. SETTING: A small ophthalmology group practice. METHODS: One hundred forty patients were randomized, 70 per group, in a prospective, partially masked study of cataract surgery using Healon GV or Viscoat. One ophthalmologist performed all surgery. Primary outcome variables were the 2 week postoperative changes in corneal thickness, endothelial cell density, mean endothelial cell size, and endothelial cell hexagonality. Several secondary variables were measured, and an analysis of the statistical power of the study was performed. RESULTS: There were no statistically significant differences between groups in terms of age (P = .856), cataract density (P = .117), preoperative best corrected visual acuity (BCVA) (P = .892), postoperative BCVA (P = .969), amount of viscoelastic material used during surgery (P = .444), amount of irrigating solution used (P = .125), or phacoemulsification time (P = .088). It took longer to remove the Viscoat than the Healon GV (P < .001), and total operating time for the Viscoat group was longer (P < .001). Two weeks after surgery, there were no significant differences between groups in corneal thickness (P = .362), endothelial cell density (P = .351), or mean endothelial cell size (P = .610). However, Viscoat preserved the hexagonal shape of endothelial cells slightly better than Healon GV (P = .043). The study had sufficient power to detect clinically significant differences in corneal thickness, endothelial cell density, and endothelial cell size. CONCLUSIONS: Healon GV and Viscoat were comparable in their ability to protect the corneal endothelium during endocapsular phacoemulsification and foldable IOL implantation. Results may vary, however, if phacoemulsification is performed anterior to the iris plane.

Aged↗

The Oxford Laser Prostate Trial: a double-blind randomized controlled trial of contact vaporization of the prostate against transurethral resection; preliminary results.

OBJECTIVE: To compare the results of contact laser vaporization and transurethral resection of the prostate (TURP) in a double-blind randomized controlled clinical trial. PATIENTS AND METHODS: The study comprised 148 patients with clinical benign prostatic hypertrophy (BPH) who were recruited and allocated randomly to undergo either TURP (72 patients) or laser ablation of the prostate (76 patients). The outcome was assessed using the American Urological Association (AUA -7) symptom score after 1 and 3 months as the primary measure and by urinary flow rates, haematological factors and the duration of hospital stay and length of catheterization. RESULTS: With 90% statistical power, the results at 3 months showed no clinical or statistical difference between the treatments in change in AUA symptom score. A lower blood loss, hospital stay and duration of catheterization significantly favoured the laser treatment, although the failure rate of trial without catheter and the rate of re-operation were higher after laser treatment. CONCLUSIONS: These early data are encouraging for this technique, although the outcome after one year requires evaluation before advocating the widespread uptake of this method.

Aged↗

Preliminary report: prescription of prism-glasses by the Measurement and Correction Method of H.-J. Haase or by conventional orthoptic examination: a multicenter, randomized, double-blind, cross-over study.

In a multicenter, randomized, double-blind, cross-over study in the Netherlands, the effectiveness of (prism-)glasses prescribed by the Measurement and Correction Method of H.-J. Haase (MKH) was compared to that of glasses prescribed by conventional orthoptic examination. Nine pairs of MKH-optometrists and orthoptists recruited patients who primarily presented with asthenopia, and each prescribed the patient (prism-)glasses. A questionnaire for asthenopia was developed that rated headache and tired eyes as 0-7 days per week and none-light-medium-severe, respectively. Light sensitivity, problems with focusing, near-work problems and burning eyes were each rated as: never-occasionally-often-always. A patient was eligible if he scored 'medium', 'often' or '5 days a week' twice; or 'medium' (etc.) once and 'light' (etc.) twice. Controls, in contrast to the patients, typically answered 'none' or 'never' to half of the complaints, but 37% of them would have passed the admission criteria. Among other criteria were: 18 to 40 years of age, horizontal angle < 4 degrees, vertical < 1.7 degrees, acuity > or = 0.8, stereopsis threshold disparity < 120". Seventy-two patients fulfilled all criteria and returned sufficient questionnaires. They wore the first glasses for six weeks, were without glasses for two weeks, and then wore the second glasses for six weeks. At the start, halfway and at the end of each 6-week period, questionnaires were filled out; 97% were returned. Only 19 of the orthoptists' glasses contained prisms (14 horizontal, 5 vertical; horizontal average of all glasses 0.49 PD, vertical 0.05 PD). Five of the orthoptists' glasses were plano. All MKH glasses contained prisms, 53 of 72 both horizontal and vertical, 18 only horizontal and one only vertical (horizontal average of all glasses 2.83 PD, vertical 0.79 PD). The starting levels of complaints were high and both glasses improved complaints dramatically. The starting levels were lower, but not significantly, in the second 6-week period and improvement was less outspoken. Because of these differences, the two periods had to be evaluated separately. The primary outcome of the study was defined as the difference between the effect of the MKH glasses and that of the orthoptists' glasses in the first and second 6-week periods. For problems with focusing, in the first 6-week period, and for tired eyes, in the second 6-week period, the difference exceeded the difference that had been defined as clinically significant (one day per week less headache or half the distance light-medium or half the distance occasionally-often), but it did not reach statistical significance. The statistical power was approximately 0.7 for demonstrating this clinically significant difference. Statistical significance was not reached in multivariate repeated measure ANOVA either. Forty-four patients preferred to keep the MKH glasses, 25 the orthoptists' glasses, including one plano. It is striking that 25% of the patients did not prefer the glasses that, according to the questionnaire, improved their complaints the most. A year after the study, the questionnaire was sent again to all patients: The levels of complaints after a year were similar to those at the end of the second 6-week period, whether they had preferred the MKH or the orthoptists' glasses, and were similar to the levels in controls. The most conspicuous finding was that both glasses improved the complaints dramatically. Apart from the prisms, other reasons could be: spherical and cylindrical correction, improved wearing comfort of the frame, placebo effect, Hawthorne effect and regression to the mean.

Adolescent↗

Uniform matrix stability study designs.

Uniform matrix designs for drug stability studies are introduced and compared to standard matrix designs. Comparisons are made on the basis of design moment, D-efficiency, uncertainty, G-efficiency, and statistical power. It is shown that uniform matrix designs provide superior statistical properties with the same or fewer design points than standard matrix designs.

Chemistry, Pharmaceutical↗

Regression-based association analysis with clustered haplotypes through use of genotypes.

Haplotype-based association analysis has been recognized as a tool with high resolution and potentially great power for identifying modest etiological effects of genes. However, in practice, its efficacy has not been as successfully reproduced as expected in theory. One primary cause is that such analysis tends to require a large number of parameters to capture the abundant haplotype varieties, and many of those are expended on rare haplotypes for which studies would have insufficient power to detect association even if it existed. To concentrate statistical power on more-relevant inferences, in this study, we developed a regression-based approach using clustered haplotypes to assess haplotype-phenotype association. Specifically, we generalized the probabilistic clustering methods of Tzeng to the generalized linear model (GLM) framework established by Schaid et al. The proposed method uses unphased genotypes and incorporates both phase uncertainty and clustering uncertainty. Its GLM framework allows adjustment of covariates and can model qualitative and quantitative traits. It can also evaluate the overall haplotype association or the individual haplotype effects. We applied the proposed approach to study the association between hypertriglyceridemia and the apolipoprotein A5 gene. Through simulation studies, we assessed the performance of the proposed approach and demonstrate its validity and power in testing for haplotype-trait association.

Apolipoprotein A-V↗

Assessing group change under conditions of anonymity and overlapping samples.

BACKGROUND: A commonly used research design in the social sciences involves the matching of observations over 2 time periods (i.e., Time 1 --> Time 2) to assess group change. Because coupled observations are usually correlated, a paired- or dependent-samples t test is generally recommended in such applications to determine if there has been a statistically significant change in mean scores across time. Consequently, it is typically believed that unless information for matching respondents' observations is available, researchers have no choice but to treat the observations as if they were independent. OBJECTIVES: To demonstrate alternative statistical approaches for employing the paired samples ttest when information for matching respondents' observations is unavailable and to illustrate the applicability of these alternatives to longitudinal designs in which respondents at Time 1 are partially replaced by new respondents at Time 2. METHOD: Theoretical arguments and examples are employed to achieve the specified objectives. RESULTS/DISCUSSION: Performing an independent-samples ttest when a paired-samples t test is more appropriate will lead to a loss of statistical power and, thus, increase the likelihood of a Type II statistical error. The statistical approaches that are demonstrated allow researchers to account for pair wise dependency across observations and, therefore, to obtain a fairer test of group change in means.

Humans↗

Should level of measurement considerations affect the choice of statistic?

The belief that the level of measurement (nominal, ordinal, interval, ratio) achieved in the data constrains the type of statistic that may be used legitimately for analysis is examined in general, and in an optometric-vision science context. Theoretical considerations indicate that statistical statements about the data may be made independently of the level of measurement, and that although researchers must be concerned about the quality of their measurement, the role of measurement theory is in the interpretation of the meaning of the investigation's results as a whole not in the governance of the choice of statistic. Empirical studies indicate that measurement considerations can be ignored for the purposes of testing the null hypothesis with little or no resultant error. Finally, adherence to the belief that level of measurement considerations limits statistical choice would result in the use of generally less powerful statistical tests--an undesirable and, all things considered, an unwarranted consequence.

Optics and Photonics↗

Prednisone withdrawal in kidney transplant recipients on cyclosporine and mycophenolate mofetil--a prospective randomized study. Steroid Withdrawal Study Group.

BACKGROUND: Prospective randomized trials have shown a reduced rate of acute rejection (AR) in mycophenolate mofetil-treated kidney transplant recipients. We hypothesized that this increased protection from AR could allow successful prednisone (P) withdrawal in cyclosporine/mycophenolate mofetil/P-treated recipients. METHODS: A multicenter, prospective, randomized, double-blind trial of P withdrawal at 3 months post-transplant was initiated. Entry criteria were: primary transplant, adult, no AR by 90 days, mycophenolate mofetil dose > or =2 g/day, cyclosporine dose = 5-15 mg/kg/ day, P dose = 10-15 mg/day. Study participants were randomized to have P tapered over 8 weeks (beginning at 3 months posttransplant) to 0 vs. 10 mg/day. Prestudy power analysis determined 500 recipients should be randomized for 80% statistical power to test equivalence of the primary endpoint, AR, or treatment failure at 1 year posttransplant. By design, the study was to be stopped if interim data precluded reaching equivalence. An established data safety monitoring board monitored the study. RESULTS: After 266 patients were enrolled, the patient enrollment was stopped (after safety monitoring board review) because of excess rejection in the P withdrawal group. The Kaplan-Meier estimate of the cumulative incidence of rejection or treatment failure within 1 year posttransplant (+/-95% confidence interval) for the maintenance group was 9.8% (4.4%; treatment failure, 14.9%); for the withdrawal group, 30.8% (21.0%; 39.3%). Treatment differences in the distribution of time to event were highly significant (P = 0.0007). Of note, risk was higher in blacks (39.6%) versus nonblacks (16.0%) (P<0.001). At 1 year post-transplant, there was no difference between groups in patient or graft survival. For the patients with functioning grafts at 6 months posttransplant, withdrawal patients had lower cholesterol (P = 0.0005), had higher creatinine (P = 0.03), and were less likely to use antihypertensives (P = 0.001). These differences persist to 1 yr posttransplant. CONCLUSIONS: We conclude that for recipients on cyclosporine/mycophenolate mofetil/P with no AR at 90 days, the chance of developing subsequent AR is small; if P is tapered and withdrawn, the risk increases (but the majority remain free of acute and chronic rejection). After withdrawal, the risk of AR is different for blacks versus nonblacks. Withdrawal patients had a lower cholesterol level and less need for antihypertensives.

Adolescent↗

The diaphragm with and without spermicide. A randomized, comparative efficacy trial.

OBJECTIVE: To determine the relative contraceptive efficacy of a diaphragm used with spermicide as compared to one used without. STUDY DESIGN: Two hundred sixteen women entered the study between September 1985 and December 1990. Of these, 84 were randomly assigned to the diaphragm-only group and 80 to the diaphragm-with-spermicide group as their primary method of contraception. In addition, a spermicide-only group was planned originally to serve as a control group to assess the contribution to efficacy made by a spermicide alone. Thirty-nine women were randomly assigned to this group, and 13 selected themselves for it. All were followed for a maximum of 12 months. The primary outcome variable was accidental pregnancy. The statistical difference between the two diaphragm groups was analyzed. RESULTS: The 12-month "typical use" failure rates for the diaphragm-only group were 28.6 per 100 women and for the diaphragm-with-spermicide group, 21.2. The 12-month cumulative consistent-use failure rates were 19.3 per 100 women for the diaphragm-only group as compared to 12.3 per 100 women for users of a diaphragm with spermicide. CONCLUSION: Although the consistent use rates were not significantly different, this study had low statistical power and hence gives no support to the hypothesis that adjunctive spermicide use fails to improve the effectiveness of the diaphragm method, especially in view of the magnitude and direction of the difference observed. Unless a study with sufficient power proves that the use of a diaphragm alone is statistically as effective as use of a diaphragm with spermicide, use of a spermicide in conjunction with the diaphragm continues to be the appropriate clinical recommendation.

Adolescent↗

Sample size calculations for studies with correlated observations.

Correlated data occur frequently in biomedical research. Examples include longitudinal studies, family studies, and ophthalmologic studies. In this paper, we present a method to compute sample sizes and statistical powers for studies involving correlated observations. This is a multivariate extension of the work by Self and Mauritsen (1988, Biometrics 44, 79-86), who derived a sample size and power formula for generalized linear models based on the score statistic. For correlated data, we appeal to a statistic based on the generalized estimating equation method (Liang and Zeger, 1986, Biometrika 73, 13-22). We highlight the additional assumptions needed to deal with correlated data. Some special cases that are commonly seen in practice are discussed, followed by simulation studies.

Biometry↗

Patching with carotid endarterectomy: when to do it and what to use.

This is an analysis of published data on the outcomes of carotid endarterectomy patch reconstruction with the commonly used materials--autologous greater saphenous vein, Dacron, and polytetrafluoroethylene. Taken individually, these prospective randomized, prospective nonrandomized, and retrospective studies have mixed findings with respect to both the advantage of patching over primary closure and the superiority, if any, of one patch material. A major problem with these reports is small sample size and the associated low statistical power of tests on outcomes. Comparative outcomes are often marginally statistically significant or not significant at the .05 level. However, when the data from similarly designed studies are pooled, as with a meta-analysis, the sample sizes are large enough to allow definitive statistical evaluation. This analysis indicates that obligatory patching or selectively patching more than 90% of carotid endarterectomies gives clearly superior outcomes (in terms of early postoperative thrombosis, perioperative stroke, and 50% or more residual or restenosis in the first year) when compared with primary closure. There is softer, but clearly important, evidence that carotid endarterectomy patch reconstruction with greater saphenous vein has better perioperative stroke and restenosis outcomes than that obtained with Dacron and polytetrafluoroethylene.

Blood Vessel Prosthesis↗

Confidence limit analyses should replace power calculations in the interpretation of epidemiologic studies.

Frequently, after an epidemiologic study is completed, statistical power to detect a relative risk of interest is recalculated using data obtained during the course of the study. A negative study may then be dismissed on the grounds that its power was too low. However, post hoc power calculations ignore the actual relative estimate and its variance, which are by then known. We present evidence that post-study power calculations have little value and should be replaced by a more informative method using the upper (1 - alpha)% confidence limit of the point estimate that touches the value of the relative risk of interest.

Confidence Intervals↗

Temporal variability of atrial tachyarrhythmia burden in bradycardia-tachycardia syndrome patients.

AIMS: Several studies have tested non-pharmacological therapies for atrial tachyarrhythmias (ATs) by measuring the cumulative time (burden) the patient spends in arrhythmia. Contradictory results questioned either therapy efficacy or statistical power of the trials. We studied AT burden variability in patients paced for sinus node disease (SND) in order to interpret currently published data appropriately and to evaluate reliable sample sizes. METHODS AND RESULTS: One hundred and five patients with AT and SND received a dual chamber pacemaker with antitachyarrhythmia-pacing capability, and were followed for 13 months. Seventy-eight patients (74%) suffered AT recurrences. Device-gathered diagnostic measures were used to simulate results of randomized studies both with crossover and parallel design. The sample size required for statistically significant results was calculated as a function of the expected therapy-induced burden reduction. AT burden intra-patient variability was high: 43% of patients showed intrinsic fluctuations hiding any therapy-induced burden reduction lower than 30%. Demonstrating therapeutic breakthrough through a 6 month study would require 290 patients with crossover design and 5800 patients with parallel design. Doubling the study period requires 400 and 3000 patients, respectively. CONCLUSION: Patients with AT and paced for SND showed high intra-patient burden variability, which could possibly hide an AT burden reduction induced by a therapy. Previous studies involving non-pharmacological therapies utilizing AT burden endpoints could lack the power to reach statistical significance.

Aged↗

The effects of selected sampling on the transmission disequilibrium test of a quantitative trait locus.

We investigate how sampling of parents or children based on their extreme phenotypic values selected from clinical databases would affect the power of identification of quantitative trait loci (QTL) by a transmission disequilibrium test (TDT). We consider three selective sampling schemes based on the selection of phenotypic values of parents or children in nuclear families: (1) two children, one of extreme value, the other random; (2) two children extremely discordant; (3) one parent of extreme value. Other family members not specified will be recruited randomly with regard to phenotypic values. Our study shows that the second sampling scheme can always enhance the power for QTL identification, sometimes dramatically so. The increase in the statistical power of the TDT is particularly dramatic when h2 at the QTL under test is small or intermediate (e.g. 0.05 or 0.10). For the other two sampling schemes, under dominant effects at the QTL, the power is always increased relative to random sampling; however, under recessive or additive genetic effects, the power gain is generally minor or even decreased a little sometimes. Allele frequencies at the QTL and the selection stringency are important for determining the effect of selective sampling on the power of QTL identification. Our study is useful as a practical guideline on how to perform the TDT efficiently in practice by taking advantage of the extensive databases accumulated that are enriched with people of extreme phenotypic values.

Disease Transmission, Infectious↗