Search PubMed⌕ Search

PubMed · 15008067

Estimating sample size in clinical studies: basic methodological principles.

Abstract

In order to be valid, clinical studies must be methodologically rigorous. The internal validity of a study is of crucial importance: a study is valid if its results are an unbiased estimation of the true result. In this case, the validity is internal because it refers to the group of patients under study and not necessarily different ones (external validity or applicability). Internal validity in clinical research is achieved through rigorous design, data collection and appropriate analysis, and is threatened by bias (systematic errors) or chance (random variation of the phenomena under study). Regardless of the type of study (analytic, descriptive, etc.), the characteristics of its sample are fundamental for the validity of the results. The sampling methods are crucial if the study patients are to be representative of the population to which one desires to extrapolate the results. One of the most fundamental characteristics of a sample is its size. Even the best executed study may fail to answer the research question if the sample size is too small. On the other hand, a study with too large a sample is harder to conduct and more costly. The goal of planning the sample size is to estimate the appropriate number of research subjects for the study. In this paper we will present and discuss the methodological principles underlying calculation of sample size: outcomes, type I and II error, alpha and beta, study power and variability.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

António Vaz Carneiro. 2003. Estimating sample size in clinical studies: basic methodological principles.. https://pubmed.ncbi.nlm.nih.gov/15008067/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Identification and impact of outcome selection bias in meta-analysis.

The systematic review community has become increasingly aware of the importance of addressing the issues of heterogeneity and publication bias in meta-analyses. A potentially bigger threat to the validity of a meta-analysis appears relatively unnoticed. The within-study selective reporting of outcomes, defined as the selection of a subset of the original variables recorded for inclusion in publication of trials, can theoretically have a substantial impact on the results. A cohort of meta-analyses on the Cochrane Library was reviewed to examine how often this form of within-study publication bias was suspected and explained some of the evident funnel plot asymmetry. In cases where the level of suspicion was high, sensitivity analysis was undertaken to assess the robustness of the conclusion to this bias. Although within-study selection was evident or suspected in several trials, the impact on the conclusions of the meta-analyses was minimal. This paper deals with the identification of, sensitivity analysis for, and impact of within-study selective reporting in meta-analysis.

Clinical Trials as Topic↗

Sample size calculations for comparative clinical trials with over-dispersed Poisson process data.

This paper develops a new formula for sample size calculations for comparative clinical trials with Poisson or over-dispersed Poisson process data. The criteria for sample size calculations is developed on the basis of asymptotic approximations for a two-sample non-parametric test to compare the empirical event rate function between treatment groups. This formula can accommodate time heterogeneity, inter-patient heterogeneity in event rate, and also, time-varying treatment effects. An application of the formula to a trial for chronic granulomatous disease is provided.

Clinical Trials as Topic↗

A permutation test for inference in logistic regression with small- and moderate-sized data sets.

Inference based on large sample results can be highly inaccurate if applied to logistic regression with small data sets. Furthermore, maximum likelihood estimates for the regression parameters will on occasion not exist, and large sample results will be invalid. Exact conditional logistic regression is an alternative that can be used whether or not maximum likelihood estimates exist, but can be overly conservative. This approach also requires grouping the values of continuous variables corresponding to nuisance parameters, and inference can depend on how this is done. A simple permutation test of the hypothesis that a regression parameter is zero can overcome these limitations. The variable of interest is replaced by the residuals from a linear regression of it on all other independent variables. Logistic regressions are then done for permutations of these residuals, and a p-value is computed by comparing the resulting likelihood ratio statistics to the original observed value. Simulations of binary outcome data with two independent variables that have binary or lognormal distributions yield the following results: (a) in small data sets consisting of 20 observations, type I error is well-controlled by the permutation test, but poorly controlled by the asymptotic likelihood ratio test; (b) in large data sets consisting of 1000 observations, performance of the permutation test appears equivalent to that of the asymptotic test; and (c) in small data sets, the p-value for the permutation test is usually similar to the mid-p-value for exact conditional logistic regression.

Clinical Trials as Topic↗