Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Statistical Power”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Prognostic factor studies in oncology: osteosarcoma as a clinical example.

OBJECTIVE: Prognostic factor studies are abundant in oncology. Nevertheless, most of them have very limited impact on clinical practice, in part because many of them have a low statistical power. The importance of statistical power is illustrated using bootstrap resampling of data from a series of osteosarcoma patients. METHODS AND MATERIALS: Osteosarcoma is a rare disease, the incidence being just a few cases per million person-years in the Western World. Very few Phase III studies have been conducted in the disease, and much knowledge on therapeutic progress has come from Phase II studies. This has caused controversy concerning the validity of historical controls, which again has stimulated interest in the identification of prognostic factors in this disease. A literature search in the National Library of Medicine MEDLINE database was performed to identify prognostic factor studies in osteosarcoma published between 1975 and 1998. Monte Carlo methods, so-called bootstrap resampling, are used to investigate the importance of sample size based on survival data for a previously published series of 158 osteosarcomas treated with surgery alone. RESULTS: Most published studies are too small to provide useful information on prognostic factors in osteosarcoma. Three-quarters of the papers reviewed included less than 100 patients with osteosarcoma. More than 20 different potential prognostic factors were included in these papers. Inherent differences between studies and poor reporting hamper a synthesis of information from various studies. The results from the bootstrap resampling illustrate how the majority of published studies would miss even quite significant prognostic factors. CONCLUSIONS: An effort is needed to improve the design, conduct, and reporting of prognostic studies in oncology.

Bone Neoplasms↗

Limitations of controlled augmentation trials in schizophrenia.

To assess the empirical basis for add-on augmentation treatments in schizophrenia, this study examined the experimental design components of pharmacologic augmentation trials in schizophrenia and compared them to conventional requirements. A search covering a 5-year period (1988-1992) identified 13 double-blind, placebo-controlled, parallel, add-on neuroleptic augmentation drug trials. The mean number of subjects per trial was 34.5, and the mean number of outcome measures examined was 25.0. The probability for a significant finding by chance was 63%. Mean effect size required to achieve conventional statistical power was 1.6. Mean statistical power (for effect sizes of .5-1.0) was .1-.4. The mean number of subjects actually required for power of .80 was 58-216. The majority of the 13 trials included too few patients and employed too many outcome measures to conclusively prove or disprove therapeutic efficacy. Conclusions drawn from such trials with less than 40-100 subjects or more than one hypothesis must remain tentative at best.

Antipsychotic Agents↗

Maximum urinary flow rate by uroflowmetry: automatic or visual interpretation.

We measured the maximum urinary flow rate monthly for 1 year by uroflowmetry in 1,645 patients in a double-blind, placebo-controlled study of finasteride therapy for benign prostatic hyperplasia. Patients were randomized to receive placebo (555) or finasteride (1,090). A total of 23,857 flow measurements was obtained. Because of the presence of artifacts on many uroflow curves, we read the maximum urinary flow rate values manually and compared them to the values provided electronically by the uroflowmeter. On average, the manually read values were 1.5 ml. per second lower than the machine read values. Artifacts causing a difference of 2 ml. per second or more between the 2 methods were found in 20% and of more than 3 ml. per second in 9% of the tracings. The difference between treatment groups in mean maximum urinary flow rate change at the end of the study was the same with both reading methods. However, confidence intervals were 15 to 25% larger for the machine read compared to the manually read values. This larger variability in machine read maximum urinary flow rate has a marked negative impact on the power of statistical tests to assess any given difference in maximum urinary flow rate between treatment groups. Furthermore, it increases sample size requirements by 50% to achieve any given statistical power. We conclude that maximum urinary flow rate artifacts contribute significantly to the variability of maximum urinary flow rate measurement by uroflowmetry. Manual reading of the maximum urinary flow rate eliminates an important fraction of such variability.

5-alpha Reductase Inhibitors↗

Quantitative trait locus mapping in chickens by selective DNA pooling with dinucleotide microsatellite markers by using purified DNA and fresh or frozen red blood cells as applied to marker-assisted selection.

Many large, half-sib sire families are an integral component of chicken genetic improvement programs. These family structures include a sufficient number of individuals for mapping quantitative trait loci (QTL) at high statistical power. However, realizing this statistical power through individual or selective genotyping is yet too costly to be feasible under current genotyping methodologies. Genotyping costs can be greatly reduced through selective DNA pooling, involving densitometric estimates of marker allele frequencies in pooled DNA samples. When using dinucleotide microsatellite markers, however, such estimates are often confounded by overlapping "shadow" bands and can be confounded further by differential amplification of alleles. In the present study a shadow correction procedure provided accurate densitometric estimates of allele frequency for dinucleotide microsatellite markers in pools made from chicken purified DNA samples, fresh blood samples, and frozen-thawed blood samples. In a retrospective study, selective DNA pooling with thawed blood samples successfully identified two QTL previously shown by selective genotyping to affect resistance in chickens to Marek's disease. It is proposed that use of selective DNA pooling can provide relatively low-cost mapping and use in marker-assisted selection of QTL that affect production traits in chickens.

Alleles↗

Repeated analysis of semen parameters in beagle dogs during a 2-year study with the HMG-CoA reductase inhibitor, atorvastatin.

Sperm analyses are often incorporated into reproductive toxicity studies in rats. Due to the relative ease of collecting multiple samples throughout a study, semen analysis in non-rodents such as dogs offers the opportunity to assess potential development of functional effects of compounds on male reproduction over time. In the present study, semen parameters were evaluated in beagle dogs during and at termination of a chronic toxicity study with the 3-hydroxy-3-methylglutaryl-coenzyme A (HMG-CoA) reductase inhibitor, atorvastatin. Male dogs received 0, 10, 40, or 120 mg/kg orally in gelatin capsules for up to 104 weeks (n = 10/group). After 52 weeks of dosing, 3 dogs/group were euthanized, and 2/group were withdrawn from treatment for a 12-week reversal period and euthanized at Week 64. The remaining 5/group continued treatment until Week 104. Semen was collected from all animals for 3 consecutive weeks prior to termination of the 52-week animals (Weeks 50, 51, 52) for analysis of sperm parameters, using manual methods of evaluation. Semen was collected from the remaining animals at Weeks 64, 78, 91, and 104, and was analyzed. At necropsy, testes, epididymides, and prostates were weighed and evaluated histologically, and epididymal sperm counts were determined. Serum cholesterol was decreased 25--60% at all doses during the study. There were no drug-related differences in semen volume and color, total sperm count, and sperm concentration, morphology, progressiveness, and percent motility during treatment with atorvastatin. There were also no effects on reproductive organ weights or histopathology, and no effects on epididymal sperm count. Thus, incorporation of semen analyses into this study allowed the evaluation of potential male reproductive effects in dogs at multiple time points during the study. Statistical power calculations demonstrated acceptable statistical power (> 80%) for semen sperm count, concentration, morphology, and motility with group sizes of 8--10 animals, and for semen sperm count and concentration or epididymal sperm count with group sizes of 3--5 animals, using the methodology described in this paper.

Acyl Coenzyme A↗

Testing statistical hypotheses about rat liver foci.

Tests of statistical hypotheses concerning treatment effect on the development of hepatocellular foci can be carried out directly on two-dimensional observations made on histologic sections or on estimates of the density and volume of foci in three dimensions. Inferences about differences in the density or size of foci from tests based on two-dimensional observations, however, can be misleading. This is because both the number of focus cross-sections observed in a tissue section and the percent area occupied by foci can be expressed in terms of the number of foci per unit volume of liver tissue and the mean focus size. As a consequence, a treatment difference may be caused by a difference in the density of foci, their average size, or both. Of more serious concern is the possibility that failure to detect a treatment effect may occur not only when there is no treatment effect but also when the density and size of foci differ between treatments in such a way that their product is unchanged. This can happen if the effect of treatment is to increase the number of foci and decrease their average size, or vice versa. A similar difficulty of interpretation is associated with hypothesis tests based on average focus cross-section area. Tests based on estimates of the number of foci per unit volume and mean focus volume allow direct inference about the quantities of interest, but these estimates are unstable because they have large variances. Empirical estimates of statistical power for the Wilcoxon rank sum test and the t-test from data on control rats suggest power may be limited in experiments with group sizes of ten and low observed numbers of focus cross-sections. If hypothesis tests based on estimates of the density and size of foci are to form the basis for a bioassay, then the power of statistical tests used to identify treatment effects should be investigated.

Animals↗

A structure-based approach to psychological assessment: matching measurement models to latent structure.

The present article sets forth the argument that psychological assessment should be based on a construct's latent structure. The authors differentiate dimensional (continuous) and taxonic (categorical) structures at the latentand manifest levels and describe the advantages of matching the assessment approach to the latent structure of a construct. A proper match will decrease measurement error, increase statistical power, clarify statistical relationships, and facilitate the location of an efficient cutting score when applicable. Thus, individuals will be placed along a continuum or assigned to classes more accurately. The authors briefly review the methods by which latent structure can be determined and outline a structure-based approach to assessment that builds on dimensional scaling models, such as item response theory, while incorporating classification methods as appropriate. Finally, the authors empirically demonstrate the utility of their approach and discuss its compatibility with traditional assessment methods and with computerized adaptive testing.

Bayes Theorem↗

Novel linkage mapping approach using DNA pooling in human and animal genetics. II. Detection of quantitative traits loci in dairy cattle.

Selective DNA pooling is an advanced methodology for linkage mapping of quantitative trait loci (QTL) in farm animals. The principle is based on densitometric estimates of marker allele frequency in pooled DNA samples of phenotypically extreme individuals from half-sib, backcross and F(2) experimental designs in farm animals. This methodology provides a rapid and efficient analysis of a large number of individuals with short tandem repeat markers that are essential to detect QTL through the genome - wide searching approach. Several strategies involving whole genome scanning with a high statistical power have been developed for systematic search to detect the quantitative traits loci and linked loci of complex traits. In recent studies, greater success has been achieved in mapping several QTLs in Israel-Holstein cattle using selective DNA pooling. This paper outlines the currently emerged novel strategies of linkage mapping to identify QTL based on selective DNA pooling with more emphasis on its theoretical pre-requisite to detect linked QTLs, applications, a general theory for experimental half-sib designs, the power of statistics and its feasibility to identify genetic markers linked QTL in dairy cattle. The study reveals that the application of selective DNA pooling in dairy cattle can be best exploited in the genome-wide detection of linked loci with small and large QTL effects and applied to a moderately sized half-sib family of about 500 animals.

Animals↗

A comprehensive power-analytic investigation of research in medical education.

A total of 230 major articles in volumes 55 through 57 of the Journal of Medical Education were reviewed for a comprehensive power-analytic investigation of research in medical education. Three statistical power determinations were made for each of 2,220 reported tests of significance, and the average power for detecting a range of possible treatment effects was calculated for each of the 100 studies subsequently included in the analysis. Among other findings, fully 91 percent of the 100 articles analyzed had less than a 50-50 chance of detecting a "small" treatment effect. Average power figures from similar surveys in other disciplines demonstrate that the problem of low statistical power is not unique to research in medical education. Additionally, the practical consequences of low statistical power are outlined, and workable guidelines for reporting the information necessary for the independent evaluation of published studies are provided.

Education, Medical↗

Targeting neuroprotection clinical trials to ischemic stroke patients with potential to benefit from therapy.

BACKGROUND AND PURPOSE: Clinical trials of neuroprotective drugs have had limited success. We investigated whether selecting patients according to prognostic features would improve the statistical power of a trial to identify an efficacious treatment. METHODS: Using placebo data from the Glycine Antagonist in Neuroprotection (GAIN) International and National Institute of Neurological Disorders and Stroke (NINDS) recombinant tissue plasminogen activator (rtPA) clinical trials, we developed and validated simple prognostic models for stroke trial end points: Barthel Index > or =95, modified Rankin Scale < or =1, National Institutes of Health Stroke Scale < or =1, and Glasgow Outcome Scale=1. Using these models, we simulated 1000 clinical trials and estimated, under several hypothetical treatment effect patterns of neuroprotection, the effect on statistical power of including only patients with moderate prognosis. We calculated the number of patients that would have to be enrolled to maintain the statistical power achieved in selecting the whole trial population. Reanalysis of actual data from the NINDS rtPA trials confirmed the results independently. RESULTS: Selecting patients with moderate prognosis (predicted probability of favorable outcome 0.2 to 0.8) enabled a sample size reduction, without loss of statistical power, of between 54.6% (51.3% to 57.6%) and 68.6% (66.0% to 71.1%), depending on the treatment effect pattern and outcome measure. These benefits were largely due to the exclusion of patients with poor prognosis. CONCLUSIONS: Targeting patients with potential to benefit enables a substantial sample size reduction without compromising statistical power or duration of recruitment. As part of a broader trial design strategy, informed use of prognostic data available acutely would help in identifying effective neuroprotective treatments.

Aged↗

The power of genome-wide association studies of complex disease genes: statistical limitations of indirect approaches using SNP markers.

Genome-wide association studies using a dense map of single nucleotide polymorphism (SNP) markers seem to enable us to detect a number of complex disease genes. In such indirect association studies, whether susceptibility genes can be detected is dependent not only on the degree of linkage disequilibrium between the disease variant and the SNP marker but also on the difference in their allele frequencies. These factors, as well as penetrance of the disease variant, influence the statistical power of such approaches. However, the power of indirect association studies is not well understood. We calculated the number of individuals necessary for the detection of the disease variant in both direct and indirect association studies with a case-control design. The result shows that a remarkable reduction in the statistical power of indirect studies, compared with that of direct ones, is unavoidable in the genome-wide screening of complex disease genes. If there is a large difference in allele frequency between the disease variant and the marker, the disease variant cannot be detected. Because the frequency of the disease variant is unknown, SNP markers with various allele frequencies, or a large number of SNP markers, must be used in indirect association studies. However, if the number of SNP markers is increased, the obtained P value may not reach the significance level due to the Bonferroni adjustment. Thus, to test a possible association between functional variants and a complex disease directly, we should identify such SNPs in as many genes as possible for use in genome-wide association studies.

Gene Frequency↗

Combined quality improvement ratio: a method for a more robust evaluation of changes in screening rates.

INTRODUCTION: It has been proposed that a ratio of the discordant cells from a McNemar's Chi-square table be used as a measure of quality improvement, and that this measure be called the Quality Improvement Ratio (QuIR). As proposed, patients enrolled in only one year of a two-year study are excluded from the McNemar's table of the QuIR. Since the original proposal of the McNemar's Chi-square in 1947 included application to matched pair data, a more comprehensive analysis would be possible if the single-year enrollees were matched into pairs. METHODS: Patients enrolled in only the first study year are matched and paired with patients enrolled in only the second study year. The pairs are matched on variables important to the disease or process being evaluated. The matched pairs are combined with the repeatedly measured subjects to increase the statistical power of the analysis. The Combined Quality Improvement Ratio (CQuIR) is demonstrated with parameters from the original articles, in a--Markov chain Monte-Carlo simulation, so a direct comparison can be made. RESULTS: CQuIR improved statistical power, especially in simulations of small populations. In some simulations the statistical power was double that of the QuIR alone. DISCUSSION: Although the QuIR provides important information, the CQuIR allows more of the data to be used to evaluate the effect of interventions in policy, delivery, and practice. The increase in statistical power of the CQuIR over the QuIR can facilitate successful evaluation of health care services.

Breast Neoplasms↗

Power of the classical twin design revisited.

Statistical power of the classical twin design was revisited. The approximate sampling variances of a least-squares estimate of the heritability in a univariate analysis and estimate of the genetic correlation coefficient in a bivariate analysis were derived analytically for the ACE model. Statistical power to detect additive genetic variation under the ACE model was derived analytically for least-squares, goodness-of-fit and maximum likelihood-based test statistics. The noncentrality parameter for the likelihood ratio test statistic is shown to be a simple function of the MZ and DZ intraclass correlation coefficients and the proportion of MZ and DZ twin pairs in the sample. All theoretical results were validated using simulation. The derived expressions can be used to calculate power of the classical twin design in a simple and rapid manner.

Analysis of Variance↗

Recent temporal trend monitoring of mercury in Arctic biota--how powerful are the existing data sets?

The goal of this paper is to describe and discuss statistical power with respect to mercury in Arctic biota, using data gathered during the past two or three decades, mostly under the auspices of AMAP Phases I and II. It will describe the current levels of power of existing data sets to detect temporal trends of Hg concentrations. If the desired power is fixed to an appropriate magnitude, the minimum size of a detectable trend within a specified time period or the number of years that is required to detect a certain trend could be estimated provided that the random between-year variation for the current time-series is known. These various measures of performance of the AMAP mercury time-series, derived from the power analysis, are discussed in some detail. The number of years required to detect a certain trend at a particular power at a specific Type I error rate (alpha) is compared with the actual number of years available when the AMAP Phase II assessment was carried out. In general the investigated time-series were too short to possess an acceptable statistical power. The effect of varying the Type-I error rate, the slope of a trend and the desired power is investigated to rank the importance of the various components regulating the statistical power. The consequence of sampling less frequently than once a year is considerable loss of power.

Animals↗

Is overall survival a realistic primary end point in advanced colorectal cancer studies? A critical assessment based on four clinical trials comparing fluorouracil plus leucovorin with the same treatment combined either with oxaliplatin or with CPT-11.

BACKGROUND: The adequacy of overall survival (OS) as study end point in phase III trials for advanced solid tumors is questionable. The present review highlights the limits of OS as study end point to evaluate the efficacy of new drugs. METHODS: Four phase III clinical trials comparing a fluorouracil-based regimen with the same regimen plus either CPT-11 or oxaliplatin in advanced colorectal cancer patients were reviewed. The primary aim of the critical assessment was to explain the lack of OS advantage observed in two of the four trials, despite the presence of increased response rate (RR) and time to progression (TTP). Four possible reasons for the lack of OS benefit (i.e. statistical power, cross-over, magnitude of the effect on RR and TTP, non-tumor-related deaths) were systematically reviewed in the trials, and the detectable 1-year OS difference, assuming a statistical power of 80%, was calculated for each. RESULTS: None of these reasons for the lack of OS advantage in presence of RR and TTP benefits convincingly explained the results of the evaluated trials. Three of the four trials had roughly the same statistical power to detect 1-year OS differences, while the fourth trial was underpowered to detect realistic OS differences. The lack of OS advantage observed in the two oxaliplatin trials is therefore likely fortuitous, and due to lack of statistical power. CONCLUSIONS: Although increase in OS remains the ultimate goal of many clinical trials, the choice of OS benefit as a mandatory requirement to register new compounds can lead to a serious underestimation of a drug's real efficacy.

Antimetabolites, Antineoplastic↗

Functional MRI using sensitivity-encoded echo planar imaging (SENSE-EPI).

Parallel imaging methods become increasingly available on clinical MR scanners. To investigate the potential of sensitivity-encoded single-shot EPI (SENSE-EPI) for functional MRI, five imaging protocols at different SENSE reduction factors (R) and matrix sizes were compared with respect to their noise characteristics and their sensitivity toward functional activation in a motor task examination. At constant echo times, SENSE-EPI was either used to shorten the single volume acquisition times (TR(min)) at matrix size 128 x 100 (22 slices) from 3.9 s (no SENSE) to 2.0 s at R = 3, or to increase the matrix size to 192 x 153 (22 slices), resulting in TR(min) = 5.3 s for R = 2 or TR(min) = 3.4 s for R = 3. At the lower resolution, the bisection of echo train length (R = 2) substantially reduced distortions and blurring, while signal-to-noise and statistical power (measured by cluster size and maximum t value per unit time) were hardly reduced. At R = 3 the additional gain in speed and distortion reduction was quite small, while signal-to-noise and statistical power dropped significantly. With enhanced spatial resolution the time course signal-to-noise was better than expected from theory for purely thermal noise because of a reduced contribution of physiological noise, and statistical power almost reached that of the regular, low-resolution single-shot EPI, with a slight drop off toward R = 3. Thus, SENSE-EPI allows to substantially increase speed and spatial resolution in fMRI. At SENSE reduction factors up to R = 2, the potential drawbacks regarding signal-to-noise and statistical power are almost negligible.

Adult↗

Central pathology review in clinical trials for patients with malignant glioma. A Report of Radiation Therapy Oncology Group 83-02.

BACKGROUND: Confounding biologic factors, including histologic grade, may influence the outcome of adult patients with malignant gliomas more than may modifications in therapeutic approach. Any clinical trial design for malignant gliomas in adults must account for such biologic factors, including the accurate identification of the two histologic subgroups astrocytoma with anaplastic foci (AAF) or glioblastoma multiforme (GBM), which are associated with distinctly different survival outcomes. This paper examines the need for a central pathology review before entry of patients in cooperative group clinical trials stratified by histologic grade. METHODS: Pathology slides from Radiation Therapy Oncology Group (RTOG) trial 83-02, a randomized Phase II study of hyperfractionated and accelerated hyperfractionated radiation therapy and carmustine for malignant gliomas, provided 747 analyzable cases, with 680 (91%) available for central pathology review. This review was performed by a single pathologist according to RTOG/Eastern Cooperative Oncology Group histopathologic criteria. The kappa statistic was used to measure agreement between the institutional and central classification of AAF and GBM. The influence of misclassification was examined using computer simulation of varying clinical trial sizes (n = 25, 50, or 200). The effect on the statistical power of trials (n = 200) with varying mixtures of AAF and GBM tumors was investigated using computer simulations. RESULTS: Of 159 tumors classified as AAF by institutional pathology review, only 66% (105) were classified as AAF (AAF/AAF) by central review, and 54 of these cases (34%) were classified as GBM (GBM/AAF), whereas 96% (501) of 521 institutionally classified as GBM (GBM/GBM) were similarly classified by central review. Computer simulations demonstrated a 59% underestimation in the median survival (1.82 vs. 4.49 years) for trials of patients with institutionally defined AAF compared to patients with centrally defined AAF in studies of 200 patients, resulting from the addition of poor prognosis of GBM in the trial. Misclassification can also substantially reduce the statistical power of a clinical trial. In one of the simulation studies, statistical power was reduced from 65% to 14% if 50% of the patients were to receive an inaccurate histologic classification. Even greater losses in power are possible in many plausible clinical settings. CONCLUSIONS: This examination of a central versus an institutional pathology review demonstrates a low level of agreement on AAF classification and a high level of concordance on GBM classification. The results indicate the need to adjust sample size for trials of both AAF and GBM tumors to have adequate statistical power. A central pathology review remains essential for trial entry for patients with AAF and could be omitted for trials enrolling patients with GBM only.

Adult↗