Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Bench to bedside: the quest for quality in experimental stroke research.

Over the past decades, great progress has been made in clinical as well as experimental stroke research. Disappointingly, however, hundreds of clinical trials testing neuroprotective agents have failed despite efficacy in experimental models. Recently, several systematic reviews have exposed a number of important deficits in the quality of preclinical stroke research. Many of the issues raised in these reviews are not specific to experimental stroke research, but apply to studies of animal models of disease in general. It is the aim of this article to review some quality-related sources of bias with a particular focus on experimental stroke research. Weaknesses discussed include, among others, low statistical power and hence reproducibility, defects in statistical analysis, lack of blinding and randomization, lack of quality-control mechanisms, deficiencies in reporting, and negative publication bias. Although quantitative evidence for quality problems at present is restricted to preclinical stroke research, to spur discussion and in the hope that they will be exposed to meta-analysis in the near future, I have also included some quality-related sources of bias, which have not been systematically studied. Importantly, these may be also relevant to mechanism-driven basic stroke research. I propose that by a number of rather simple measures reproducibility of experimental results, as well as the step from bench to bedside in stroke research may be made more successful. However, the ultimate proof for this has to await successful phase III stroke trials, which were built on basic research conforming to the criteria as put forward in this article.

Animals↗

Methodological characteristics and quality of alcohol treatment outcome studies, 1970-98: an expanded evaluation.

AIMS: To examine the methodological quality of the research literature examining treatment for alcohol use disorders. DESIGN: In total, 701 studies first reported between 1970 and 1998 were evaluated with a scoring system covering 19 areas of research quality. FINDINGS: Methodological quality improved from a score of 8.2 out of a possible 28.5 in the 1970s, to 10.6 in the 1990s. Strengths included reporting the initial number of participants and conducting follow-ups of 12 months or longer. Although the percentage of studies with adequate power also increased overall, the average statistical power of comparative studies to detect a medium-sized treatment effect was low (0.54). CONCLUSIONS: Areas with room for improvement include: ensuring that follow-up data are collected when respondents are not under the influence of alcohol, testing for differential dropout among treatment groups with respect to participant background characteristics, reporting the number of individuals being treated in the programs from which samples are drawn, noting the reliability and validity of measures used and conducting process analyses to examine potential mechanisms underlying treatment effects.

Alcoholism↗

Raising the curtain on the gray region.

The gray region in EPA Document QA/G-4 is defined as the range of possible parameter values near the action level where the cost of determining that the alternative condition is true outweighs the expected consequences of a decision error. The gray region is also described as a range of true parameter values within the alternative condition near the action level where it is "too close to call." EPA Document QA/G-4HW states that during the planning stage the action level is based on an ideal decision rule, while during the assessment stage an operational decision rule is used. This paper analyzes the factors that define the gray region and the action level, including the errors of the first kind (a) and second kind (beta) and the number of samples taken to determine the mean result. The relationship between the Decision Performance Curve presented in EPA QA/G-4 and the statistical power curve is also discussed. The statistically derived critical level is identified as the concentration of importance for decision-making. The action level is defined in terms of the critical level so that its value is consistent for decisions made during both planning (a priori decisions) and assessment (a posteriori decisions).

Decision Support Techniques↗

Multi-level statistical models in studies of periodontal diseases.

Periodontal data typically have a hierarchical structure, with sites grouped within individuals, and individuals grouped within communities. Also, the occasion may be regarded as another level since the acquired knowledge indicates that periodontal disease activity may vary over time. Conventional statistical tests are based on unilevel analysis of data. However, this approach to statistical analysis is often inconvenient in periodontal research because of the variation in the outcome variables between the various levels in the hierarchy. Lately there have been important developments in the statistical theory which have made available powerful statistical techniques for analyzing multilevel or hierarchical data. This report describes a new approach for analyzing periodontal data and uses an illustrative example to build a model which explains part of the variability in the response variable. The results from this analysis are then compared to results from an earlier report which uses unilevel methods and the findings discussed. The present multilevel approach has several advantages over unilevel methods, mainly due to its statistical validity and efficiency. Further, it permits the incorporation of explanatory variables measured at the site and the subject levels, and those which vary across the time points. Multilevel analyses have a promising potential and are expected to have a significant impact on periodontal research.

Age Factors↗

European best practice guidelines for renal transplantation. Section IV: Long-term management of the transplant recipient. IV.13 Analysis of patient and graft survival.

GUIDELINES: A. It is important for a transplant unit to follow-up on the results of their transplant activities. In order to achieve correct reports on graft and patient outcome in all patients, it is necessary to have sufficient resources, such as a computerized database, and continuous updates of patient information. All data collected should be subjected to validation procedures to ensure completeness and accuracy. B. Improved outcomes following implementation of new protocols, based on evaluation of clinical multi-centre trials, should be verified at local transplant centres since centres often include a range of patients different from those selected for the trial. C. The most widely accepted descriptor of outcome is the Kaplan-Meier probability estimate of patient and graft survival. Survival estimates should be calculated at intervals of time after transplantation and should always be expressed with their 95% confidence intervals. D. Kaplan-Meier survival estimates may be calculated in three ways. (i) 'Patient survival' should be calculated from the date of transplantation to the date of death or the date of the last follow-up. (ii) 'Graft survival' (non-censored for death) should be calculated from the date of transplantation to the date of irreversible graft failure signified by return to long-term dialysis (or retransplantation) or the date of the last follow-up during the period when the transplant was still functioning or to the date of death. Here, death with graft function is treated as graft failure. (iii) 'Graft survival censored for death with a functioning graft' (death-censored graft survival) should be calculated from the date of transplantation to the date of irreversible graft failure signified by return to long-term dialysis (or retransplantation) or the date of last follow-up during the period when the transplant was still functioning. In the event of death with a functioning graft, the follow-up period is censored at the date of death. E. The outcome of transplants carried out at a centre should be compared with those achieved across a range of data from centres collated by national and international multi-centre registries. Interpretation of a centre's performance should take into account the number of transplants performed and the prevalence of major risk factors. F. Major risk factors that influence transplant outcome are identifiable by applying multivariate analytical methods to large multi-centre follow-up databases. Although these major risk factors may not be identifiable in individual centre data, they should nonetheless be taken into account in patient management. G. When designing a clinical trial or evaluating data from a recent trial, the expected improvement in graft survival resulting from a reduction in acute rejection may be estimated from a knowledge of the rejection and graft survival rates that existed prior to the introduction of the new therapeutic regimen. H. When designing or evaluating a clinical trial, it is important to analyse the power of the study to verify statistically the difference (in graft survival) that might be expected and its statistical significance. A study resulting in absence of statistically significant differences between two treatment groups with insufficient statistical power to verify a difference at the expected level should not be taken as evidence of absence of a true difference.

Clinical Trials as Topic↗

A meta-analysis of the reliability and validity of the Rorschach.

The results of a meta-analysis of Rorschach studies indicate that reliabilities in the order of .83 and higher and validity coefficients of .45 or .50 and higher can be expected for the Rorschach--when hypotheses supported by empirical or theoretical rationales are tested using reasonably powerful statistics. Three important determinants of variance accounted for in a variety of Rorschach scores were identified in 530 statistics from 39 papers published in the Journal of Personality Assessment from 1971 to 1980. The a priori theoretical or empirical evidence determining the likelihood of obtaining significant results, and the power of the statistic used to measure the results, as well as the interaction between the likelihood of results and the power of the statistic used, were all significant determinants of the proportion of variance accounted for in the Rorschach measures reported.

Journal Article↗

An integrated population-averaged approach to the design, analysis and sample size determination of cluster-unit trials.

While the mixed model approach to cluster randomization trials is relatively well developed, there has been less attention given to the design and analysis of population-averaged models for randomized and non-randomized cluster trials. We provide novel implementations of familiar methods to meet these needs. A design strategy that selects matching control communities based upon propensity scores, a statistical analysis plan for dichotomous outcomes based upon generalized estimating equations (GEE) with a design-based working correlation matrix, and new sample size formulae are applied to a large non-randomized study to reduce underage drinking. The statistical power calculations, based upon Wald tests for summary statistics, are special cases of a general power method for GEE.

Adolescent↗

Farm Scale Evaluations of spring-sown genetically modified herbicide-tolerant crops: a statistical assessment.

Primary results from the Farm Scale Evaluations (FSEs) of spring-sown genetically modified herbicide-tolerant crops were published in 2003. We provide a statistical assessment of the results for count data, addressing issues of sample size (n), efficiency, power, statistical significance, variability and model selection. Treatment effects were consistent between rare and abundant species. Coefficients of variation averaged 73% but varied widely. High variability in vegetation indicators was usually offset by large n and treatment effects, whilst invertebrate indicators often had smaller n and lower variability; overall, achieved power was broadly consistent across indicators. Inferences about treatment effects were robust to model misspecification, justifying the statistical model adopted. As expected, increases in n would improve detectability of effects whilst, for example, halving n would have resulted in a loss of significant results of about the same order. 40% of the 531 published analyses had greater than 80% power to detect a 1.5-fold effect; reducing n by one-third would most likely halve the number of analyses meeting this criterion. Overall, the data collected vindicated the initial statistical power analysis and the planned replication. The FSEs provide a valuable database of variability and estimates of power under various sample size scenarios to aid planning of more efficient future studies.

Agriculture↗

Statistical limitations to the Cornell model of latent tuberculosis infection for the study of relapse rates.

The Cornell model has been extensively used as a mouse model for studying the latent stage of Mycobacterium tuberculosis infection. In this model mice are infected and then given a course of chemotherapy prior to reactivation of infection. We discuss here the importance of using adequate mouse numbers in a Cornell model for the study of relapse rates in order to obtain sufficient statistical power to confirm a hypothesis. Experiments with small sample sizes are useful for 'screening' experiments, but will have very little value in 'confirming' the objective. When the objective of the experiment is confirmation of an effect through establishment of statistical significance, power calculations are critical in order to assure that the sample size will be sufficient to meet that objective.

Animals↗

The statistical performance of an MCF-7 cell culture assay evaluated using generalized linear mixed models and a score test.

Biological assays often utilize experimental designs where observations are replicated at multiple levels, and where each level represents a separate component of the assay's overall variance. Statistical analysis of such data usually ignores these design effects, whereas more sophisticated methods would improve the statistical power of assays. This report evaluates the statistical performance of an in vitro MCF-7 cell proliferation assay (E-SCREEN) by identifying the optimal generalized linear mixed model (GLMM) that accurately represents the assay's experimental design and variance components. Our statistical assessment found that 17beta-oestradiol cell culture assay data were best modelled with a GLMM configured with a reciprocal link function, a gamma error distribution, and three sources of design variation: plate-to-plate; well-to-well, and the interaction between plate-to-plate variation and dose. The gamma-distributed random error of the assay was estimated to have a coefficient of variation (COV) = 3.2 per cent, and a variance component score test described by X. Lin found that each of the three variance components were statistically significant. The optimal GLMM also confirmed the estrogenicity of five weakly oestrogenic polychlorinated biphenyls (PCBs 17, 49, 66, 74, and 128). Based on information criteria, the optimal gamma GLMM consistently out-performed equivalent naive normal and log-normal linear models, both with and without random effects terms. Because the gamma GLMM was by far the best model on conceptual and empirical grounds, and requires only trivially more effort to use, we encourage its use and suggest that naive models be avoided when possible.

Biological Assay↗

Prevention of missing data in clinical research studies.

Missing data is a problem that is ubiquitous to all clinical studies and a source of multiple problems from an analytic point of view (reduced statistical power, increased the type I error, bias) Statistical approaches have been developed to analyze data collected from trials with missing data. Understanding and implementing the appropriate statistical technique is essential but should be differentiated from preventive approaches that are designed to reduce rates of missing data In this article, we draw attention to these preventive efforts. Seven steps to minimizing the amount of missing data are defined as documentation, training, monitoring reports, patient contact, data entry and management, pilot studies, and communication. Although the implementation of these approaches is time consuming and costly, the overall quality of the study is increased. Despite efforts devoted to areas, no study is without missing data. Once the study is completed, it is essential to assess the pattern of missing data and apply the appropriate statistical analysis.

Bias↗

Subgroup analysis and other (mis)uses of baseline data in clinical trials.

BACKGROUND: Baseline data collected on each patient at randomisation in controlled clinical trials can be used to describe the population of patients, to assess comparability of treatment groups, to achieve balanced randomisation, to adjust treatment comparisons for prognostic factors, and to undertake subgroup analyses. We assessed the extent and quality of such practices in major clinical trial reports. METHODS: A sample of 50 consecutive clinical-trial reports was obtained from four major medical journals during July to September, 1997. We tabulated the detailed information on uses of baseline data by use of a standard form. FINDINGS: Most trials presented baseline comparability in a table. These tables were often unduly large, and about half the trials inappropriately used significance tests for baseline comparison. Methods of randomisation, including possible stratification, were often poorly described. There was little consistency over whether to use covariate adjustment and the criteria for selecting baseline factors for which to adjust were often unclear. Most trials emphasised the simple unadjusted results and covariate adjustment usually made negligible difference. Two-thirds of the reports presented subgroup findings, but mostly without appropriate statistical tests for interaction. Many reports put too much emphasis on subgroup analyses that commonly lacked statistical power. INTERPRETATION: Clinical trials need a predefined statistical analysis plan for uses of baseline data, especially covariate-adjusted analyses and subgroup analyses. Investigators and journals need to adopt improved standards of statistical reporting, and exercise caution when drawing conclusions from subgroup findings.

Bias↗

No contribution of angiotensin-converting enzyme (ACE) gene variants to severe obesity: a model for comprehensive case/control and quantitative cladistic analysis of ACE in human diseases.

Candidate gene analyses are often inconclusive owing to genetic or phenotypic heterogeneity, low statistical power, selection of nonfunctional SNPs, and inadequate statistical analysis of the genetic architecture. Angiotensin-converting enzyme (ACE) is involved in adipocyte growth and function and the ACE-processed angiotensin II inhibits adipocyte differentiation. Associations between body mass index (BMI) and ACE polymorphisms have been reported in general populations, but the contribution to severe obesity of this gene, which is located under an obesity genome-scan linkage peak on 17q23, is unknown. ACE is one of the most studied genes and markers responsible for variation in circulating ACE enzyme levels have been extensively characterised. Eight of these variants were genotyped in 1054 severely obese cases and 918 nonobese controls, as well as 116 nuclear families from the genome scan (n=447), enabling the known clades to be inferred. Qualitative analysis of individual single-nucleotide polymorphisms (SNPs), haplotypes, clades, and diploclades demonstrated no significant associations (P<0.05) after minimal correction for multiple testing. Quantitative analysis of clades and diploclades for BMI, waist-to-hip ratio, or ZBMI in children were also not significant. This rigorous, large-scale study of common, well-defined, severe polygenic obesity provides strong evidence that functionally relevant sequence variation in ACE, whether it is defined at the level of SNPs, haplotypes, or clades, is not associated with severe obesity in French Caucasians. Such a study design exemplifies the strategy needed to clearly define the contribution of the ACE gene to the plethora of complex genetic diseases where weak associations have been previously reported.

Case-Control Studies↗

Supporting medication adherence in renal transplantation (SMART): a pilot RCT to improve adherence to immunosuppressive regimens.

BACKGROUND: Although non-adherence to an immunosuppressive regimen (NAH) is a major risk factor for poor outcome after renal transplantation (RTx), very few studies have examined non-adherence intervention in this context. This pilot randomized controlled trial (RCT) tested the efficacy of an educational-behavioural intervention to increase adherence in non-adherent RTx patients. We also assessed how NAH evolves over time. METHODS: Eighteen RTx non-adherent patients (age: 45.6 +/- 1.2 yr; 78.6% male) were randomly assigned to either an intervention group (IG) (n = 6) or an enhanced usual care group (EUCG) (n = 12), the latter receiving the usual clinical care. The IG received one home visit and three telephone interviews. We assessed NAH through electronic monitoring (EM) of medication intake during a nine-month period (three months intervention, six months follow-up). RESULTS: Five of 18 patients withdrew. Inclusion in the study resulted in a remarkable decrease in NAH in both groups over the first three months (IG chi(2) = 3.97, df = 1, p = 0.04; EUCG chi(2) = 3.40, df = 1, p = 0.06). The IG showed the greatest decrease in NAH after three months, although this did not reach statistical significance (at 90 d, chi(2) = 1.05, df = 1, p = 0.31). Thereafter, NAH increased gradually in both groups, reaching comparable levels at the end of the six-month follow-up (i.e. at nine months). CONCLUSION: Our findings suggest an inclusion effect. Although the intervention in this pilot RCT appeared to add further benefit in medication compliance, a lack of statistical power prevented us from making a strong statistical statement.

Adolescent↗

Occupation and five cancers: a case-control study using death certificates.

A case-control approach has been used to examine mortality from five cancers--oesophagus, pancreas, cutaneous melanoma, kidney, and brain--among young and middle aged men resident in three English counties. The areas studied were chosen because they include major centres of chemical manufacture. By combining data from 20 years it was possible to look at local industries with greater statistical power than is possible using routine national statistics. Each case was matched with up to four controls of similar age who died in the same year from other causes. The occupations and industries recorded on death certificates were coded to standard classifications and risk estimates derived for each job category. Where positive associations were found the records of the cases concerned were examined in greater detail to see whether the risk was limited to specific combinations of occupation and industry. The most interesting findings to emerge were risks of brain cancer associated with the production of meat and fish products (relative risk (RR) = 9.7, 95% confidence interval (CI) 2.6-36.8) and with mineral oil refining (RR = 2.9, CI 1.2-7.0), and a cluster of four deaths from melanoma among refinery workers (RR = 16.0, CI CI 1.8-143.2). A job-exposure matrix was applied to the data but gave no strong indications of further disease associations. Local analyses of occupational mortality such as this can usefully supplement national statistics.

Adolescent↗

Multifactor dimensionality reduction: an analysis strategy for modelling and detecting gene-gene interactions in human genetics and pharmacogenomics studies.

The detection of gene-gene and gene-environment interactions associated with complex human disease or pharmacogenomic endpoints is a difficult challenge for human geneticists. Unlike rare, Mendelian diseases that are associated with a single gene, most common diseases are caused by the non-linear interaction of numerous genetic and environmental variables. The dimensionality involved in the evaluation of combinations of many such variables quickly diminishes the usefulness of traditional, parametric statistical methods. Multifactor dimensionality reduction (MDR) is a novel and powerful statistical tool for detecting and modelling epistasis. MDR is a non-parametric and model-free approach that has been shown to have reasonable power to detect epistasis in both theoretical and empirical studies. MDR has detected interactions in diseases such as sporadic breast cancer, multiple sclerosis and essential hypertension. As this method is more frequently applied, and was gained acceptance in the study of human disease and pharmacogenomics, it is becoming increasingly important that the implementation of the MDR approach is properly understood. As with all statistical methods, MDR is only powerful and useful when implemented correctly. Concerns regarding dataset structure, configuration parameters and the proper execution of permutation testing in reference to a particular dataset and configuration are essential to the method's effectiveness. The detection, characterisation and interpretation of gene-gene and gene-environment interactions are expected to improve the diagnosis, prevention and treatment of common human diseases. MDR can be a powerful tool in reaching these goals when used appropriately.

Environment↗

A model to select chemotherapy regimens for phase III trials for extensive-stage small-cell lung cancer.

BACKGROUND: Many more phase II studies have favorable outcomes than the subsequent phase III trials. We used historical data from phase II and phase III studies for patients with extensive-stage small-cell lung cancer (SCLC) to generate a statistical model to provide assistance in selecting chemotherapy regimens from phase II studies for subsequent use in phase III randomized studies. METHODS: Information from 21 phase III trials for patients with extensive-stage SCLC initiated during the period from 1972 through 1990 was reviewed to identify those that were preceded by phase II studies of the same regimen. We used data from all the trial pairs to develop a statistical model in which the number of patients, the median survival of patients, and the number of deaths observed in the phase II trial are used to estimate the statistical power of the subsequent phase III trial. All statistical tests were two-sided. RESULTS: Nine phase II studies were identified that preceded phase III trials of the same regimen. The regimens from two phase II studies with the greatest expected power in the phase III trial (0. 62 and 0.58) both demonstrated significantly prolonged survival when compared with standard treatment in subsequent phase III trials (P<. 001 and P =.002, respectively). The regimens from six of the other phase II studies, for which the median power expected in the phase III trial was 0.28 (range, 0.19-0.52), showed no difference when compared with standard treatment in a phase III trial. CONCLUSIONS: Phase II studies for particular regimens that have an expected power of greater than 0.55 provide a reasonable basis for proceeding with a phase III trial.

Antineoplastic Agents↗

The comparative efficacy of trazodone and imipramine in the treatment of depression.

OBJECTIVE: To review published clinical trials comparing the efficacy of trazodone with that of tricyclic antidepressant medication. DATA SOURCES: MEDLINE was searched for relevant articles published from 1983 to 1991. The bibliography of a review article was searched for further references. STUDY SELECTION: In all, 25 clinical trials were found. Six of these met the methodologic assessment criteria (adapted from the McMaster guidelines for the evaluation of clinical trials), which included the stipulation of a score of 18 or more on the Hamilton depression rating scale and a 50% reduction in that score as an outcome measure. DATA EXTRACTION: All six studies compared trazodone with imipramine. Data describing response to the treatments were extracted, and post-hoc power estimates were calculated. The analysis also involved statistical tests of a modified null hypothesis, the generation of confidence intervals (CIs) and a meta-analysis. DATA SYNTHESIS: All the studies found no significant difference in the efficacy of trazodone and imipramine. However, the statistical power of most of them was less than 50% and often less than 10%; thus there was a low probability that differences would be detected. The results of statistical tests of the modified null hypothesis, inspection of the CIs and the results of the meta-analysis all suggested that trazodone, and imipramine are equally efficacious. CONCLUSION: The application of various techniques for the analysis of equivalence data suggests that trazodone and imipramine are of approximately equivalent efficacy. The data are compatible with small differences in efficacy, but the differences are of a magnitude such that they are unlikely to be of clinical significance.

Clinical Trials as Topic↗