Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Understanding statistical power in the context of applied research.

Estimates of statistical power are widely used in applied research for purposes such as sample size calculations. This paper reviews the benefits of power and sample size estimation and considers several problems with the use of power calculations in applied research that result from misunderstandings or misapplications of statistical power. These problems include the use of retrospective power calculations and standardized measures of effect size. Methods of increasing the power of proposed research that do not involve merely increasing sample size (such as reduction in measurement error, increasing 'dose' of the independent variable and optimizing the design) are noted. It is concluded that applied researchers should consider a broader range of factors (other than sample size) that influence statistical power, and that the use of standardized measures of effect size should be avoided (except as intermediate stages in prospective power or sample size calculations).

Research↗

Prospective studies of diagnostic test accuracy when disease prevalence is low.

Prospective studies of diagnostic test accuracy have important advantages over retrospective designs. Yet, when the disease being detected by the diagnostic test(s) has a low prevalence rate, a prospective design can require an enormous sample of patients. We consider two strategies to reduce the costs of prospective studies of binary diagnostic tests: stratification and two-phase sampling. Utilizing neither, one, or both of these strategies provides us with four study design options: (1) the conventional design involving a simple random sample (SRS) of patients from the clinical population; (2) a stratified design where patients from higher-prevalence subpopulations are more heavily sampled; (3) a simple two-phase design using a SRS in the first phase and selection for the second phase based on the test results from the first; and (4) a two-phase design with stratification in the first phase. We describe estimators for sensitivity and specificity and their variances for each design, along with sample size estimation. We offer some recommendations for choosing among the various designs. We illustrate the study designs with two examples.

Journal Article↗

Measurement of motor recovery after stroke. Outcome assessment and sample size requirements.

BACKGROUND AND PURPOSE: The purpose of this study was to analyze recovery of motor function in a cohort of patients presenting with an acute occlusion in the carotid distribution. Analysis of recovery patterns is important for estimating patient care needs, establishing therapeutic plans, and estimating sample sizes for clinical intervention trials. METHODS: We prospectively measured the motor deficits of 104 stroke patients over a 6-month period to identify earliest measures that would predict subsequent motor recovery. Motor function was measured with the Fugl-Meyer Assessment. Fifty-four patients were randomly assigned to a training set for model development; 50 patients were assigned to a test set for model validation. In a second analysis, patients were stratified on basis of time and stroke severity. The sample size required to detect a 50% improvement in residual motor function was calculated for each level of impairment and at three points in time. RESULTS: At baseline the initial Fugl-Meyer motor scores accounted for only half the variance in 6-month motor function (r2 = 0.53, p less than 0.001). After 5 days, both the 5-day motor and sensory scores explained 74% of the variance (p less than 0.001). After 30 days, the 30-day motor score explained 86% of the variance (p less than 0.001). Application of these best models to the test set confirmed the results obtained with the training set. Sample-size calculations revealed that as severity and time since stroke increased, sample sizes required to detect a 50% improvement in residual motor deficits decreased. CONCLUSIONS: Most of the variability in motor recovery can be explained by 30 days after stroke. These findings have important implications for clinical practice and research.

Activities of Daily Living↗

Optimal power transformations for analysis of sperm concentration and other semen variables.

The nongaussian (or nonnormal) distribution of sperm concentration, and variables deriving from it, is a common practical problem in the statistical evaluation of semen data. Yet it has been little studied, and its importance to data analysis, as well as to practical remedies, is not widely appreciated. Inappropriate use of the raw scale of measurement produces inflated estimates of mean and variance, leading to false-negative (underpowered) statistical comparisons and excessive sample size estimates. This study employs the Box-Cox family of power transforms to illustrate by a simple graphical method how to identify optimal power transforms for semen data variables. Using robust statistical methods, it is shown that the nongaussian distribution is due to right skewing rather than multimodality or influential outliers. The optimal power transform, typically in the region of 0.15 to 0.35 (most easily implemented as a cube-root transformation), usually performs better than the logarithmic transformation in normalizing the data. In addition, the power transformation has an important practical advantage over the logarithmic transformation in the appropriate handling of zeros (azoospermia), a regular and important features of such data sets in practice.

Humans↗

Calculation of power and sample size with bounded outcome scores.

The two-sample Wilcoxon rank sum test is the most popular non-parametric test for the comparison of two samples when the underlying distributions are not normal. Although the underlying distributions need not be known in detail to calculate the null distribution of the test statistic, parametric assumptions are often made to determine the power of the test or the sample size. We encountered difficulties with this approach in the planning of a recent clinical trial in stroke patients. It is shown that, for power and sample size estimation, it can be dangerous to apply the classical formulae routinely, especially with outcome scores having a U-shaped or a J-shaped distribution. As an example we have taken the Barthel index, a quality-of-life outcome measure in stroke patients. Further, we have investigated alternative methods by means of Monte Carlo simulation. The distributional characteristics of the estimated powers were compared. Our findings suggest more appropriate computer software is necessary for the calculation of power and sample size when efficacy is measured by a non-parametric method.

Cerebrovascular Disorders↗

Negative results of randomized clinical trials published in the surgical literature: equivalency or error?

HYPOTHESIS: We hypothesized that review of randomized controlled clinical trials (RCTs) with nonstatistically significant or "negative" results published in the surgical literature do not have appropriate statistical power to demonstrate equivalency between treatment arms. DATA SOURCES AND STUDY SELECTION: The MEDLINE database was searched to obtain reports of all RCTs with negative results published in 3 surgical journals from 1988 to 1998. Manual review of one year (1997) of publications for each journal was performed to validate our search strategy. Equivalency was evaluated using the Two One-Sided Tests Procedure and post hoc power calculations. DATA SYNTHESIS: Ninety reports of RCTs with negative results were identified in the surgical literature between 1988 and 1998. The manual review of 1997 showed a 100% retrieval rate for our search strategy. After applying the Two One-Sided Tests Procedure, 35 reports (39%) met the criteria for demonstrating equivalency. The other 55 reports (61%) contained at least a 10% absolute difference in the 90% confidence interval of Delta. Using the power calculation method, only 22 (24%) articles had a power greater than.80 to detect a 50% difference in therapeutic effect. Only 29% of the reports included a formal sample size calculation and these studies were more likely to demonstrate equivalency than those without a sample size estimate (P<.01). CONCLUSIONS: Many reports from negative RCTs published in the surgical literature lack sufficient statistical power to establish that clinically important differences are not present. Surgeons should perform appropriate sample size calculations when designing RCTs and recognize the utility of confidence intervals when reporting negative results.

Confidence Intervals↗

Symmetry of the canine femur: implications for experimental sample size requirements.

The purpose of this study was to determine the degree of bilateral variability in cross-sectional geometric properties of the adult canine proximal femur and to use these data to determine minimum detectable treatment effects in paired and independent experimental designs. Thirteen pairs of canine femora were sectioned at nine locations and 16 cross-sectional geometric properties were determined for each section location. The canine femur was found to be bilaterally symmetrical. For a given sample size, the magnitude of the detectable treatment effect was (a) smaller for diameters than for areas and area moments of inertia and (b) smaller within the middiaphysis than proximally. The data from this study can be used to estimate sample size requirements for experiments in which the treatment effect is determined by using the contralateral femur as a control. It was found that an increase from 3 to 7 animals would have a much greater effect on improving the sensitivity of an experiment than would an increase from 7 to 11 animals.

Animals↗

Power approximation for the van Elteren test based on location-scale family of distributions.

The van Elteren test, as a type of stratified Wilcoxon-Mann-Whitney test for comparing two treatments accounting for stratum effects, has been used to replace the analysis of variance when the normality assumption was seriously violated. The sample size estimation methods for the van Elteren test have been proposed and evaluated previously. However, in designing an active-comparator trial where a sample of responses from the new treatment is available but the patient response data to the comparator are limited to summary statistics, the existing methods are either inapplicable or poorly behaved. In this paper we develop a new method for active-comparator trials assuming the responses from both treatments are from the same location-scale family. Theories and simulations have shown that the new method performs well when the location-scale assumption holds and works reasonably when the assumption does not hold. Thus, the new method is preferred when computing sample sizes for the van Elteren test in active-comparator trials.

Algorithms↗

Sample sizes for long-term medical trial with time-dependent dropout and event rates.

A general model is formulated that allows for time-dependent dropout and event rates in the determination of sample sizes for long-term medical trials when a therapy group and a control group are to be compared. The need for time-dependent event (dropout) rate is illustrated by using the Framingham Heart Study Mortality data to estimate sample size for the NHLBI (National Heart, Lung, and Blood Institute) Multiple Risk Factor Intervention Trial (MRFIT). A further generalization of the model allows for participants who drop out of the control group (e.g., because of therapeutic measures prescribed by their own physicians) to return (or not return) to the control group at a later date. The question of unequal sample sizes for the therapy and the control groups is also discussed.

Clinical Trials as Topic↗

A combination of distribution- and anchor-based approaches determined minimally important differences (MIDs) for four endpoints in a breast cancer scale.

OBJECTIVE: To determine distribution- and anchor-based minimal important difference (MID) estimates for four scores from the Functional Assessment of Cancer Therapy-Breast (FACT-B): the breast cancer subscale (BCS), Trial Outcome Index (TOI), FACT-G (the general version), and FACT-B. STUDY DESIGN AND SETTING: We used data from a Phase III clinical trial in metastatic breast cancer (ECOG study 1193; n=739) and a prospective observational study of pain in metastatic breast cancer (n=129). One third and one half of the standard deviation and 1 standard error of measurement were used as distribution-based criteria. Clinical indicators used to determine anchor-based differences included ECOG performance status, current pain, and response to treatment. RESULTS: FACT-B scores were responsive to performance status and pain anchors, but not to treatment response. By combining the results of distribution- and anchor-based methods, MID estimates were obtained: BCS=2-3 points, TOI=5-6 points, FACT-G=5-6 points, and FACT-B=7-8 points. CONCLUSION: Distribution- and anchor-based estimates of the MID do show convergence. These estimates can be used in combination with other measures of efficacy to determine meaningful benefit and provide a basis for sample size estimation in clinical trials.

Adult↗

Design and analysis of group-randomized trials: a review of recent practices.

We reviewed group-randomized trials (GRTs) published in the American Journal of Public Health and Preventive Medicine from 1998 through 2002 and estimated the proportion of GRTs that employ appropriate methods for design and analysis. Of 60 articles, 9 (15.0%) reported evidence of using appropriate methods for sample size estimation. Of 59 articles in the analytic review, 27 (45.8%) reported at least 1 inappropriate analysis and 12 (20.3%) reported only inappropriate analyses. Nineteen (32.2%) reported analyses at an individual or subgroup level, ignoring group, or included group as a fixed effect. Hence increased vigilance is needed to ensure that appropriate methods for GRTs are employed and that results based on inappropriate methods are not published.

Bibliometrics↗

Non-linearity of Parkinson's disease progression: implications for sample size calculations in clinical trials.

BACKGROUND: Estimation of sample size for long-term studies of neuroprotection in Parkinson's disease requires information on expected clinical decline. Values may be obtained by analyzing existing long-term data sets or by prediction models of clinical decline applied to available data from shorter-term trials. The most commonly used measure to track clinical decline is the Unified Parkinson's Disease Rating Scale (UPDRS) but this measure is also affected by symptomatic therapy. Models can help better understand behavior of the UPDRS after initiation of symptomatic therapy when scores will improve and eventually start deteriorating again. PURPOSE: To understand how UPDRS scores progress after initiation of symptomatic therapy and how this progression impacts sample size calculations. METHODS: We developed a non-linear model of UPDRS after introduction of symptomatic therapy. The model is specified as a non-linear mixed effects model and is applied to three different data sets from clinical trials. The model is then used to produce estimates for the change in UPDRS and its associated variance for a period of up to five years of follow-up. The estimates produced by the model serve as the basis for sample-size calculations for different lengths of follow-up (one through five years) and for different values of clinically meaningful change in UPDRS. RESULTS: Despite differences in the short-term benefit of the dopaminergic drugs, after a period of approximately six months UPDRS scores progress linearly at an estimated rate of approximately three points a year. The sample size that is required for a clinical trial where the baseline coincides with initiation of symptomatic therapy is very large. On the other hand, if baseline is set at six months after initiation of symptomatic therapy then the sample size required decreases with length of follow-up. LIMITATIONS: Model specification and estimation is based on a set of simplifying assumptions regarding the progression of individual level UPDRS scores. CONCLUSIONS: Sample size calculations based on these estimates indicate a substantial reduction in sample size if patients are required to be on symptomatic treatment for a period of time before being randomized to a neuroprotective trial.

Antiparkinson Agents↗

Clinical trials of behavioural interventions with heterogeneous teaching subgroup effects.

Behaviour modification is often delivered to teaching subgroups. For example, experimental and control smoking cessation programmes may be given to 15 classes (subgroups) with 10 (otherwise independent) individuals. We present general statistical tests and power estimates to compare continuous outcomes from two interventions in settings where the magnitude of teaching subgroup heterogeneity, number of subgroups and subgroup size can differ between intervention arms. An application is made to data from a trial to reduce disease-transmitting sexual behaviour. The statistical impact of teaching subgroup heterogeneity effect increases as the (a) number of participants in a subgroup increases, and (b) ratio of 'averaged experimental and control subgroup effect variance' to study subject variance increases. If plausible levels of subgroup teaching effect heterogeneity are ignored, the true sizes of tests with nominal 0.05 two-sided type I errors range from 0.055 to 0.47, while when planning studies, estimated sample sizes are only 11.1-95.2 per cent of the true requirements.

Behavior Therapy↗

Design issues for conducting cost-effectiveness analyses alongside clinical trials.

In response to rising demands for timely economic data on new medical technologies, cost-effectiveness studies are increasingly being conducted alongside clinical trials. Because of the historical differences in perspective and methods between cost-effectiveness studies and clinical trials, the design phase of these hybrid trials requires special consideration. Cost-effectiveness studies require more comprehensive evaluations of outcomes than the endpoints typically measured in clinical trials. Often, these comprehensive outcome measures (such as quality of life) prove useful for interpreting the other endpoints measured in the trial, as well as for estimating the cost-effectiveness of the intervention. In this manuscript, we discuss several aspects related to the design of joint clinical/economic trials, including study perspective, hypothesis testing, sample size estimation, and methods for collecting cost and outcome data. We also discuss issues that may limit the external validity of the cost-effectiveness results of these trials. Many potential threats to external validity can be successfully addressed if they are identified and accounted for in the design phase of the study.

Clinical Trials as Topic↗

Advanced statistics: statistical methods for analyzing cluster and cluster-randomized data.

Sometimes interventions in randomized clinical trials are not allocated to individual patients, but rather to patients in groups. This is called cluster allocation, or cluster randomization, and is particularly common in health services research. Similarly, in some types of observational studies, patients (or observations) are found in naturally occurring groups, such as neighborhoods. In either situation, observations within a cluster tend to be more alike than observations selected entirely at random. This violates the assumption of independence that is at the heart of common methods of statistical estimation and hypothesis testing. Failure to account for the dependence between individual observations and the cluster to which they belong can have profound implications on the design and analysis of such studies. Their p-values will be too small, confidence intervals too narrow, and sample size estimates too small, sometimes to a dramatic degree. This problem is similar to that caused by the more familiar "unit of analysis error" seen when observations are repeated on the same subjects, but are treated as independent. The purpose of this paper is to provide an introduction to the problem of clustered data in clinical research. It provides guidance and examples of methods for analyzing clustered data and calculating sample sizes when planning studies. The article concludes with some general comments on statistical software for cluster data and principles for planning, analyzing, and presenting such studies.

Cluster Analysis↗

Sample size calculation for complex clinical trials with survival endpoints.

Sample size estimation is important in planning clinical trials. The purpose of this paper is to describe features and use of SIZE, a comprehensive computer program for calculating sample size, power, and duration of study in clinical trials with time-dependent rates of event, crossover, and loss to follow-up. SIZE covers a wide range of complexities commonly occurring in clinical trials, such as nonproportional hazards, lag in treatment effect, and uncertainties in treatment benefit. The use of SIZE is illustrated by several hypothetical examples as well as applications to real study designs, each featuring a statistical issue.

Algorithms↗

Measuring atrophy in Alzheimer disease: a serial MRI study over 6 and 12 months.

BACKGROUND: Global brain atrophy rate calculated from serial MRI scans may be a surrogate marker of Alzheimer disease (AD) progression. Few studies have assessed atrophy in AD over short intervals. METHODS: Thirty-eight patients with AD and 19 control subjects had MRI scans at baseline, 6 months, and 1 year. Ventricular change and whole-brain volume loss were calculated directly from the regions manually outlined on registered scans and using the automated (boundary shift integral [BSI]) technique. Sample sizes required to power placebo-controlled treatment trials over 6 months and 1 year were calculated using these techniques. RESULTS: Increased rates of ventricular expansion and whole-brain atrophy were seen in AD compared with control subjects at both 6 and 12 months using manual and automated techniques (p < 0.001). Using the BSI consistently reduced measurement variability especially for whole-brain change. In clinical trials, at 6 months, significantly fewer patients would be required using the ventricular BSI (VBSI) compared with the brain BSI (BBSI) (e.g., 165 vs 410 per arm to provide 90% power to detect a 20% reduction in rate of change). At 1 year, sample size estimates were smaller than at 6 months, and the advantage of using VBSI instead of BBSI was less marked. CONCLUSIONS: In short-interval studies, using the ventricular boundary shift integral instead of the brain boundary shift integral may allow for disease-modifying effects to be demonstrated using significantly smaller sample sizes. This potential benefit must be balanced against the possibility that ventricular volumes may be more likely to be affected by factors other than neurodegeneration.

Aged↗

Increasing the degrees of freedom in future group randomized trials: the df* approach.

This article builds on the previous article by Blitstein et al. (2005), which showed how external estimates of intraclass correlation can be used to improve the precision for the analysis of an existing group randomized trial. The authors extend that work to sample size estimation and power analysis for future group-randomized trials. Often this approach will allow a smaller study than would otherwise be possible without sacrificing statistical power. Such studies are needed, for example, as pilot studies to help plan for a full-scale efficacy trial, as replication studies, or in situations in which resource constraints prohibit a larger trial. The authors discuss the circumstances under which this strategy will be most helpful and the risks associated with conducting smaller studies.

Humans↗