Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Adjustment for strong predictors of outcome in traumatic brain injury trials: 25% reduction in sample size requirements in the IMPACT study.

The aim of this study was to quantify the potential reduction in sample size that can be achieved by adjustment for predictors of outcome in traumatic brain injury (TBI) trials. We used individual patient data from seven therapeutic phase III randomized clinical trials (RCTs; n = 6166) in moderate or severe TBI, and three TBI surveys (n = 2238). The primary outcome was the dichotomized Glasgow Outcome Scale at 6 months (favorable/unfavorable). Baseline predictors of outcome considered were age, motor score, pupillary reactivity, computed tomography (CT) classification, traumatic subarachnoid hemorrhage, hypoxia, hypotension, glycemia, and hemoglobin. We calculated the potential sample size reduction obtained by adjustment of a hypothetical treatment effect for one to seven predictors with logistic regression models. The distribution of predictors was more heterogeneous in surveys than in trials. Adjustment of the treatment effect for the strongest predictors (age, motor score, and pupillary reactivity) yielded a reduction in sample size of 16-23% in RCTs and 28-35% in surveys. Adjustment for seven predictors yielded a reduction of about 25% in most studies: 20-28% in RCTs and 32-39% in surveys. A major reduction in sample size can be obtained with covariate adjustment in TBI trials. Covariate adjustment for strong predictors should be incorporated in the analysis of future TBI trials.

Age Factors↗

Sample sizes for randomized trials measuring quality of life in cancer patients.

This paper describes the methods appropriate for calculating sample sizes for clinical trials assessing quality of life (QOL). An example from a randomized trial of patients with small cell lung cancer completing the Hospital Anxiety and Depression Scale (HADS) is used for illustration. Sample size estimates calculated assuming that the data are either of the Normal form or binary are compared to estimates derived using an ordered categorical approach. In our example, since the data are very skewed, the Normal and binary approaches are shown to be unsatisfactory: binary methods may lead to substantial over estimates of sample size and Normal methods take no account of the asymmetric nature of the distribution. When summarizing normative data for QOL scores the frequency distributions should always be given so that one can assess if non-parametric methods should be used for sample size calculations and analysis. Further work is needed to discover what changes in QOL scores represent clinical importance for health technology interventions.

Adult↗

On the sample size for studies based upon McNemar's test.

When computing the sample size for studies using McNemar's test, one needs to know the probability of discordance and the odds ratio to be detected. In many studies, the investigator is unable to specify the probability of discordance, but can state, at least approximately, the marginal probabilities of each variable. This information leads to restrictions on the possible values of the cell probabilities and provides a range of admissible values for the off-diagonal cells. We compute the sample size needed in these circumstances and compare them to the results cited by Schlesselman and Connett et al. These sample sizes for the method are quite close to those found in the Monte Carlo study of Connett et al.

Antibodies, Monoclonal↗

Two-stage sampling for etiologic studies. Sample size and power.

Preexisting computerized databases are potentially valuable sources of epidemiologic data. Since such databases are infrequently created specifically for etiologic research, data may be available for the exposure of interest and, through record linkage, for the endpoint of interest, but lacking for potential confounders. Because of the size of these databases, two-stage sampling is an efficient alternative to surveying the entire study population for confounder data. At stage 1, information on exposure and disease status is obtained for the entire study population. Confounder data are collected for probability-selected subsamples at stage 2. Logistic regression is performed on the stage 2 samples, with the parameter estimates and variances appropriately corrected to account for the stage 1 data. In this paper, the authors present methods for determining the required stage 2 sample size in the case of categorical exposure and confounding variables. Sample size tables, power curves, and a computer program have been produced to accommodate a binary exposure and a single binary confounder. With the increasing availability of preexisting yet incomplete databases, the potential for use of two-stage sampling will greatly increase in the future. This investigation provides a basis for estimating the number of participants to sample for the collection of confounder data at the second stage.

Case-Control Studies↗

Sample size and mass range effects on the allometric exponent of basal metabolic rate.

The controversial relationship between body mass and basal metabolic rate in animals revolves around two questions: what is the allometric scaling exponent and what is the functional basis for it? For mammals, the first question could be resolved if measurements from all 4600 extant species were available, but this study shows that data for only 150 species, spanning three to four orders of magnitude variation in body mass, are sufficient to accurately determine the exponent. Because the currently available data set includes about 600 species that vary over five orders of magnitude in body size, further increases in sample size are unlikely to change the estimate of the scaling exponent.

Animals↗

Sample size correction for treatment crossovers in randomized clinical trials with a survival endpoint.

Sample size determination in randomized clinical trials usually relies on the determination of survival rates at the time of analysis in both groups, under the null and the alternative hypotheses, the type I and II error rates and on other assumptions, such as proportional hazards in most cases. However, in numerous clinical trials for malignant chronic diseases, it is currently common that a patient allocated to the conventional treatment group would receive the experimental treatment in case of disease progression or relapse. Such crossovers are usually not taken into account when computing the sample size of the trial, but generally result in a decreased power of the trial. In this work, we aimed to correct the sample size of such trials to control the power, under an exponential survival assumption.

Cross-Over Studies↗

Adjusting sample size for anticipated dropouts in clinical trials.

Statistical models for calculating sample sizes for controlled clinical trials often fail to take into account the negative impact that dropouts have on the power of intent-to-treat analyses. Empirically defined dropout correction coefficients are proposed to adjust sample sizes for endpoint analysis of variance (ANOVA) and analysis of covariance (ANCOVA) that have been initially calculated assuming complete data. The implications of type of analysis (change-score ANOVA or ANCOVA), correlational structure of the repeated measurements (compound symmetry or autoregressive), and percentage of dropouts (20% or 30%) are considered, together with other less influential design and data parameters. We recommend the use of ANCOVA to correct for baseline differences and for time-in-study if there is a nonspecific change across time. Given a realistic autoregressive (order 1) correlational structure for the repeated measurements and a proposed endpoint ANCOVA, the empirical results support the common practice of increasing calculated sample size by the anticipated number of dropouts. The previous rationale has been to retain a requisite number of "completers" on which to base statistical inferences. We believe the present results provide the first documentation of the relevance of that strategy for intent-to-treat analyses in which the incomplete data for dropouts must be included. Based on comparative power analyses, the strategy also seems appropriate for maintaining the power of mixed-model regression analyses, simple regression on a normalized time scale, and analyses of trends fitted to imputed scores for dropouts.

Clinical Trials as Topic↗

Practical FDR-based sample size calculations in microarray experiments.

MOTIVATION: Owing to the experimental cost and difficulty in obtaining biological materials, it is essential to consider appropriate sample sizes in microarray studies. With the growing use of the False Discovery Rate (FDR) in microarray analysis, an FDR-based sample size calculation is essential. METHOD: We describe an approach to explicitly connect the sample size to the FDR and the number of differentially expressed genes to be detected. The method fits parametric models for degree of differential expression using the Expectation-Maximization algorithm. RESULTS: The applicability of the method is illustrated with simulations and studies of a lung microarray dataset. We propose to use a small training set or published data from relevant biological settings to calculate the sample size of an experiment. AVAILABILITY: Code to implement the method in the statistical package R is available from the authors.

Algorithms↗

Sample size requirements for evaluating a conservative therapy.

Determination of an adequate sample size for a clinical trial has traditionally involved the specification of type I (false positive) and type II (false negative) error rates, and a difference that one wishes to detect. Because newer therapy has generally been more invasive or more toxic, it is conventional for the type I error to be 0.05 in order that new therapy not be accepted as superior unless its advantages are definitively established. Recently, many new trials have been directed toward showing that a more conservative treatment is equivalent in efficacy to a standard intensive therapy. In this paper, we provide formulas which prescribe the sample size necessary to meet certain criteria specified by the investigator for this alternative type of clinical trial. In addition, the percent increase in total sample size is described when more patients are allocated to one treatment than the other.

Biometry↗

Sample size calculation for clinical trials: the impact of clinician beliefs.

The UK Medical Research Council (MRC) randomized trial of gastric surgery, ST01, compared conventional (D1) with radical (D2) surgery. Sample size estimation was based upon the consensus opinion of the surgical members of the design team, which suggested that a change in 5-year survival from 20% (D1) to 34% (D2) could be realistic and medically important. On the basis of these survival rates, the sample size for the trial was 400 patients. However, this trial was exceptional in the way that a survey of surgeons' opinions was made at the start of the trial, in 1986, and again before results were analysed but after termination of the trial in 1994. At the initial survey, the three surgeons from the trial steering committee and 23 other surgeons experienced in treating gastric carcinoma were given detailed questionnaires. They were asked about the expected survival rate in the D1 group, anticipated difference in survival from D2 surgery, and what difference would be medically important and influence future treatment of patients. The consensus opinion of those surveyed was that there might be a survival improvement of 9.4%. In 1994, prior to closure of the trial, and before any survival information was disclosed, the survey was repeated with 21 of the original 26 surgeons. At this second survey, the opinion of the trial steering committee was that 9.5% difference was more realistic. This was in accord with the opinion of the larger group, which remained little changed since 1986. The baseline 5-year D1 survival was thought likely to be about 32%, which corresponded closely to the actual survival of recruited patients. Revised sample size calculations suggested that, on the basis of these more recent opinions, between 800 and 1200 patients would have been required. Both surveys assessed the level of treatment benefit that was deemed to be sufficient for causing surgeons to change their practice. This showed that the 13% difference in survival used as the study target was clinically relevant, but also indicated that many clinicians would remain unwilling to change their practice if the difference is only 9.5%. The experience of this carefully designed trial illustrates the problems of designing long-term, randomized trials. It raises interesting questions about the common practice of basing sample size estimates upon the beliefs of a trial design committee that may include a number of enthusiasts for the trial treatment. If their opinion of anticipated effect sizes drives the design of the trial, rather than the opinion of a larger community of experts that includes sceptics as well as enthusiasts, there is likely to be a serious miscalculation of sample size requirements.

Bayes Theorem↗

Statistical methods in epidemiology. II: A commonsense approach to sample size estimation.

PURPOSE: It has been argued, by many, that mathematical formulae for estimating sample size are unnecessarily complex, so much so that researchers may be reluctant to seek statistical advice. METHOD: This paper reviews methods of sample size estimation arguing that two formulae (one based on comparison of proportion of 'successes', the other based on comparison of means of normally distributed data) suffice for many situations. This paper argues the case by taking examples drawn mainly from clinical trials research. However, the methods outlined can also be used in epidemiology specifically in both case-control and cohort studies with no loss of information. RESULTS: For the situations outlined, worked examples are provided. CONCLUSIONS: Sample size estimation need not necessarily be a complex process. Simple techniques exist which enable the clinician and the statistician to work together. Continued dialogue between both parties is required so that good ideas do not go to waste.

Data Interpretation, Statistical↗

Adapting the sample size planning of a phase III trial based on phase II data.

Traditionally, in clinical development plan, phase II trials are relatively small and can be expected to result in a large degree of uncertainty in the estimates based on which Phase III trials are planned. Phase II trials are also to explore appropriate primary efficacy endpoint(s) or patient populations. When the biology of the disease and pathogenesis of disease progression are well understood, the phase II and phase III studies may be performed in the same patient population with the same primary endpoint, e.g. efficacy measured by HbA1c in non-insulin dependent diabetes mellitus trials with treatment duration of at least three months. In the disease areas that molecular pathways are not well established or the clinical outcome endpoint may not be observed in a short-term study, e.g. mortality in cancer or AIDS trials, the treatment effect may be postulated through use of intermediate surrogate endpoint in phase II trials. However, in many cases, we generally explore the appropriate clinical endpoint in the phase II trials. An important question is how much of the effect observed in the surrogate endpoint in the phase II study can be translated into the clinical effect in the phase III trial. Another question is how much of the uncertainty remains in phase III trials. In this work, we study the utility of adaptation by design (not by statistical test) in the sense of adapting the phase II information for planning the phase III trials. That is, we investigate the impact of using various phase II effect size estimates on the sample size planning for phase III trials. In general, if the point estimate of the phase II trial is used for planning, it is advisable to size the phase III trial by choosing a smaller alpha level or a higher power level. The adaptation via using the lower limit of the one standard deviation confidence interval from the phase II trial appears to be a reasonable choice since it balances well between the empirical power of the launched trials and the proportion of trials not launched if a threshold lower than the true effect size of phase III trial can be chosen for determining whether the phase III trial is to be launched.

Clinical Trials, Phase II as Topic↗

A simple approximation for calculating sample sizes for detecting linear trend in proportions.

A simple approximate formula for sample sizes for detecting a linear trend in proportions is derived. The formulas for both the uncorrected and corrected Cochran-Armitage test are given. For two binomial proportions these reduce to those given by Casagrande, Pike, and Smith (1978, Biometrics 34, 483-486). Some numerical results of a power study for small sample sizes show that the nominal power corresponding to the approximate sample size is a reasonably good approximation to the actual power.

Analysis of Variance↗

A simple approach to power and sample size calculations in logistic regression and Cox regression models.

For a given regression problem it is possible to identify a suitably defined equivalent two-sample problem such that the power or sample size obtained for the two-sample problem also applies to the regression problem. For a standard linear regression model the equivalent two-sample problem is easily identified, but for generalized linear models and for Cox regression models the situation is more complicated. An approximately equivalent two-sample problem may, however, also be identified here. In particular, we show that for logistic regression and Cox regression models the equivalent two-sample problem is obtained by selecting two equally sized samples for which the parameters differ by a value equal to the slope times twice the standard deviation of the independent variable and further requiring that the overall expected number of events is unchanged. In a simulation study we examine the validity of this approach to power calculations in logistic regression and Cox regression models. Several different covariate distributions are considered for selected values of the overall response probability and a range of alternatives. For the Cox regression model we consider both constant and non-constant hazard rates. The results show that in general the approach is remarkably accurate even in relatively small samples. Some discrepancies are, however, found in small samples with few events and a highly skewed covariate distribution. Comparison with results based on alternative methods for logistic regression models with a single continuous covariate indicates that the proposed method is at least as good as its competitors. The method is easy to implement and therefore provides a simple way to extend the range of problems that can be covered by the usual formulas for power and sample size determination.

Breast Neoplasms↗

[Optimum sample size in the crossing experiments].

Crossing experiments are time-consuming and costly, hence, it is very essential to make a through plan and provide the necessary and sufficient sample size in advance. The common formula of the sample size in statistics is not suitable for the crossing experiments. This paper discussed two cases and put forward a corresponding estimate formula of sample size for crossing experiments, by utilizing the sample size derived from the estimate formula to arrange the crossing experiments; thus the total experimentation cost could be reduced to the lowest, or the total number of tested livestocks may be the fewest under the precondition of satisfying the requirements of the experiment designer.

Costs and Cost Analysis↗

Simulation model for enumeration of Salmonella on chicken as a function of PCR detection time score and sample size: implications for risk assessment.

A data gap commonly identified in risk assessments is the lack of quantitative information on the contamination of food with pathogens. A simulation model that predicts the incidence and distribution of Salmonella contamination on chicken as a function of PCR detection time score and sample size was developed with data from challenge studies with preenrichment samples that were composed of 25 g of chicken and 225 ml of buffered peptone water inoculated with 10(0.7) to 10(6) Salmonella and incubated at 37 degrees C. At 0, 2, 4, 6, 8, 10, 12, and 24 h of incubation, subsamples were collected and tested for Salmonella by PCR, and a PCR detection time score based on the widths of the bands in the electrophoresis gel was obtained for each preenrichment sample. Standard curves relating PCR detection time score to initial density of Salmonella inoculated were developed for sterile and nonsterile preenrichment samples. Presence of other microorganisms in the preenrichment sample decreased the PCR detection time score at low (<10(2) per 25 g) but not at high (>10(2) per 25 g) initial densities of Salmonella and resulted in a nonlinear standard curve rather than the linear standard curve obtained for sterile samples. The predicted incidence and distribution of Salmonella contamination on chicken increased in a nonlinear manner as sample size increased from 25 to 500 g. The new method reduced the time and cost of Salmonella enumeration by eliminating the selective enrichment, selective plating, and confirmation steps of the traditional most-probable-number method. Results are useful for risk assessment because they consider the uncertainty of the standard curve predictions and because they provide distributions of Salmonella contamination for different size samples of chicken that can be directly used in risk assessment.

Animals↗

Power and sample size evaluation for the McNemar test with application to matched case-control studies.

Various expressions have appeared for sample size calculation based on the power function of McNemar's test for paired or matched proportions, especially with reference to a matched case-control study. These differ principally with respect to the expression for the variance of the statistic under the alternative hypothesis. In addition to the conditional power function, I identify and compare four distinct unconditional expressions. I show that the unconditional calculation of Schlesselman for the matched case-control study can be expressed as a first-order unconditional calculation as described by Miettinen. Corrections to Schlesselman's unconditional expression presented by Fleiss and Levin and by Dupont, which use different models to describe exposure association among matched cases and controls, are also equivalent to a first-order unconditional calculation. I present a simplification of these corrections that directly provides the underlying table of cell probabilities, from which one can perform any of the alternative sample size calculations. Also, I compare the four unconditional sample size expressions relative to the exact power function. The conclusion is that Miettinen's first-order expression tends to underestimate sample size, while his second-order expression is usually fairly accurate, though possibly slightly anti-conservative. A multinomial-based expression presented by Connor, among others, is also fairly accurate and is usually slightly conservative. Finally, a local unconditional expression of Mitra, among others, tends to be excessively conservative.

Case-Control Studies↗

Sample size for FDR-control in microarray data analysis.

We consider identifying differentially expressing genes between two patient groups using microarray experiment. We propose a sample size calculation method for a specified number of true rejections while controlling the false discovery rate at a desired level. Input parameters for the sample size calculation include the allocation proportion in each group, the number of genes in each array, the number of differentially expressing genes and the effect sizes among the differentially expressing genes. We have a closed-form sample size formula if the projected effect sizes are equal among differentially expressing genes. Otherwise, our method requires a numerical method to solve an equation. Simulation studies are conducted to show that the calculated sample sizes are accurate in practical settings. The proposed method is demonstrated with a real study.

Algorithms↗