Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Sample sizes for cancer trials where Health Related Quality of Life is the primary outcome.

Health Related Quality of Life (HRQoL) instruments are increasingly important in evaluating health care, especially in cancer trials. When planning a trial, one essential step is the calculation of a sample size, which will allow a reasonable chance (power) of detecting a pre-specified difference (effect size) at a given level of statistical significance. It is almost mandatory to include this calculation in research protocols. Many researchers quote means and standard deviations to determine effect sizes, and assume the data will have a Normal distribution to calculate their required sample size. We have investigated the distribution of scores for two commonly used HRQoL instruments completed by lung cancer patients, and have established that scores do not have the Normal distribution form. We demonstrate that an assumption of Normality can lead to unrealistically sized studies. Our recommendation is to use a technique that is based on the fact that the HRQoL data are ordinal and makes minimal but realistic assumptions.

Antineoplastic Combined Chemotherapy Protocols↗

Sample size determination for estimation of the accuracy of two conditionally independent tests in the absence of a gold standard.

We developed an Excel spreadsheet template (available at http://www.epi.ucdavis.edu/diagnostictests/) to calculate sample sizes to estimate sensitivity and specificity with desired precision in the absence of a gold standard. Calculations are predicated on the use of two conditionally independent tests for screening animals from two populations and are based on the methods of Hui and Walter(1980). Sample size calculations rely on asymptotic normality of maximum likelihood (ML) estimates of parameters. Spreadsheets for calculating standard errors for the parameter estimates and for providing ML estimates using cross-tabulated data also are included. An example of application of the methods to bovine paratuberculosis is presented.

Animals↗

Sample-size calculation for a log-transformed outcome measure.

The outcome measure of interest in clinical trials sometimes requires transformation to the logarithmic scale for analysis. This paper examines sample-size calculation for both independent groups and matched-pairs trials for log-transformed outcomes. For both types of trial, we demonstrate how the calculation can be formulated in terms of a relative treatment effect and a statement of relative variability, both specified on the original scale of measurement. For a comparison of two independent groups, the relative treatment effect is the ratio of group geometric means (or alternatively, group arithmetic means) and the coefficient of variation is used as a summary of relative variability. For a matched-pairs comparison, the appropriate relative treatment effect is the geometric mean of the within-pair ratios, and relative variability can be specified as an upper bound on within-pair ratios under a null hypothesis of the relative effect being equal to 1 (i.e., no difference). We discuss the clinical study that motivated this work and demonstrate the application of the sample-size calculation to this study.

Algorithms↗

[Estimation of sample size in randomized controlled clinical trials--elimination of various systematic biases by the utilization of computer network system].

Although sample size calculation is mandatory in clinical trials, the power of most of the published small size clinical studies, was only about 0.25 instead of 0.80 which is normally required in a standard randomized controlled trial. This phenomenon is probably due to the publication bias. In a cancer clinical study, the ideal sample size needed in a trial exceeds over a thousand, when significance level was fixed under 0.05 and treatment difference is estimated from 10 to 15%. In such cases, evaluation of a new treatment in less common cancers or in a specified strata become unfeasible. Although statistically plausible clinical trial is difficult, small trials of newly advocated treatment is likely to be performed elsewhere, and presentation of these haphazard results might propagate a wrong information about the new treatment. To prevent the dissemination of these biased results, 1) preregistration of all planned clinical trials to an authorized organization before the initiation of the trial, 2) randomization of the patients entered in the trial from the first case, and 3) registration of all available data of the patients and meta-analysis of preregistered multiple trials, should be the most effective counterplans. In order to achieve those functions, installation of a computer assisted coordinating center is considered to be the best solution for the proper evaluation of the clinical trial as well as for the evaluation of a new treatment. With the collaboration of regional affiliating hospitals, a pilot study have started to establish a computer network system.

Database Management Systems↗

Covariate adjustment in randomized controlled trials with dichotomous outcomes increases statistical power and reduces sample size requirements.

OBJECTIVE: Randomized controlled trials (RCTs) with dichotomous outcomes may be analyzed with or without adjustment for baseline characteristics (covariates). We studied type I error, power, and potential reduction in sample size with several covariate adjustment strategies. STUDY DESIGN AND SETTING: Logistic regression analysis was applied to simulated data sets (n=360) with different treatment effects, covariate effects, outcome incidences, and covariate prevalences. Treatment effects were estimated with or without adjustment for a single dichotomous covariate. Strategies included always adjusting for the covariate ("prespecified"), or only when the covariate was predictive or imbalanced. RESULTS: We found that the type I error was generally at the nominal level. The power was highest with prespecified adjustment. The potential reduction in sample size was higher with stronger covariate effects (from 3 to 46%, at 50% outcome incidence and covariate prevalence) and independent of the treatment effect. At lower outcome incidences and/or covariate prevalences, the reduction was lower. CONCLUSION: We conclude that adjustment for a predictive baseline characteristic may lead to a potentially important increase in power of analyses of treatment effect. Adjusted analysis should, hence, be considered more often for RCTs with dichotomous outcomes.

Data Interpretation, Statistical↗

An evaluation of HapMap sample size and tagging SNP performance in large-scale empirical and simulated data sets.

A substantial investment has been made in the generation of large public resources designed to enable the identification of tag SNP sets, but data establishing the adequacy of the sample sizes used are limited. Using large-scale empirical and simulated data sets, we found that the sample sizes used in the HapMap project are sufficient to capture common variation, but that performance declines substantially for variants with minor allele frequencies of <5%.

Chromosome Mapping↗

Sample size calculation for multiple testing in microarray data analysis.

Microarray technology is rapidly emerging for genome-wide screening of differentially expressed genes between clinical subtypes or different conditions of human diseases. Traditional statistical testing approaches, such as the two-sample t-test or Wilcoxon test, are frequently used for evaluating statistical significance of informative expressions but require adjustment for large-scale multiplicity. Due to its simplicity, Bonferroni adjustment has been widely used to circumvent this problem. It is well known, however, that the standard Bonferroni test is often very conservative. In the present paper, we compare three multiple testing procedures in the microarray context: the original Bonferroni method, a Bonferroni-type improved single-step method and a step-down method. The latter two methods are based on nonparametric resampling, by which the null distribution can be derived with the dependency structure among gene expressions preserved and the family-wise error rate accurately controlled at the desired level. We also present a sample size calculation method for designing microarray studies. Through simulations and data analyses, we find that the proposed methods for testing and sample size calculation are computationally fast and control error and power precisely.

Computer Simulation↗

Sample size and power determination for clustered repeated measurements.

It is common in epidemiological and clinical studies that each subject has repeated measurements on a single common variable, while the subjects are also 'clustered'. To compute sample size or power of a test, we have to consider two types of correlation: correlation among repeated measurements within the same subject, and correlation among subjects in the same cluster. We develop, based on generalized estimating equations, procedures for computing sample size and power with clustered repeated measurements. Explicit formulae are derived for comparing two means, two slopes and two proportions, under several simple correlation structures.

Alveolar Bone Loss↗

Design and analysis of genetic association studies to finely map a locus identified by linkage analysis: sample size and power calculations.

Association (e.g. case-control) studies are often used to finely map loci identified by linkage analysis. We investigated the influence of various parameters on power and sample size requirements for such a study. Calculations were performed for various values of a high-risk functional allele (fA), frequency of a marker allele associated with the high risk allele (f1), degree of linkage disquilibrium between functional and marker alleles (D') and trait heritability attributable to the functional locus (h2). The calculations show that if cases and controls are selected from equal but opposite extreme quantiles of a quantitative trait, the primary determinants of power are h2 and the specific quantiles selected. For a dichotomous trait, power also depends on population prevalence. Power is optimal if functional alleles are studied (fA= f1 and D'= 1.0) and can decrease substantially as D' diverges from 1.0 or as f(1) diverges from fA. These analyses suggest that association studies to finely map loci are most powerful if potential functional polymorphisms are identified a priori or if markers are typed to maximize haplotypic diversity. In the absence of such information, expected minimum power at a given location for a given sample size can be calculated by specifying a range of potential frequencies for fA (e.g. 0.1-0.9) and determining power for all markers within the region with specification of the expected D' between the markers and the functional locus. This method is illustrated for a fine-mapping project with 662 single nucleotide polymorphisms in 24 Mb. Regions differed by marker density and allele frequencies. Thus, in some, power was near its theoretical maximum and little additional information is expected from additional markers, while in others, additional markers appear to be necessary. These methods may be useful in the analysis and interpretation of fine-mapping studies.

Alleles↗

Cohort versus cross-sectional design in large field trials: precision, sample size, and a unifying model.

In planning large longitudinal field trials, one is often faced with a choice between a cohort design and a cross-sectional design, with attendant issues of precision, sample size, and bias. To provide a practical method for assessing these trade-offs quantitatively, we present a unifying statistical model that embraces both designs as special cases. The model takes account of continuous and discrete endpoints, site differences, and random cluster and subject effects of both a time-invariant and a time-varying nature. We provide a comprehensive design equation, relating sample size to precision for cohort and cross-sectional designs, and show that the follow-up cost and selection bias attending a cohort design may outweigh any theoretical advantage in precision. We provide formulae for the minimum number of clusters and subjects. We relate this model to the recently published prevalence model for COMMIT, a multi-site trial of smoking cessation programmes. Finally, we tabulate parameter estimates for some physiological endpoints from recent community-based heart-disease prevention trials, work an example, and discuss the need for compiling such estimates as a basis for informed design of future field trials.

Analysis of Variance↗

Economics in sample size determination for clinical trials.

In the design of clinical trials, sample size determination is usually undertaken by statisticians and clinicians. It is rare for health economists to be involved in this aspect of trial design. However, there are a number of outcome changes that are of 'economic significance', and it is important for trial designers and funders to be aware of these before planning, funding and mounting a trial. In this paper we demonstrate through the use of three examples (prevention of osteoporosis, management of infertility, and endometriosis) how economics can be used to influence the size of a clinical trial. Trials that are too small or too large waste research resources; health economics can lead to more efficient trial designs.

Adult↗

Evaluation of sample size and power for multi-arm survival trials allowing for non-uniform accrual, non-proportional hazards, loss to follow-up and cross-over.

We present a general framework for sample size calculation in survival studies based on comparing two or more survival distributions using any one of a class of tests including the logrank test. Incorporated within this framework are the possible presence of non-uniform staggered patient entry, non-proportional hazards, loss to follow-up and treatment changes including cross-over between treatment arms. The framework is very general in nature and is based on using piecewise exponential distributions to model the survival distributions. We illustrate the use of the approach and explore its validity using simulation studies. These studies have shown that not adjusting for loss to follow-up, non-proportional hazards or cross-over can lead to significant alterations in power or equivalently, a marked effect on sample size. The approach has been implemented in the freely available program ART (for Stata). Our investigations suggest that ART is the first software to allow incorporation of all these elements. Further extensions to the methodology such as non-local alternatives for the logrank test are also considered.

Antiretroviral Therapy, Highly Active↗

Power and sample size calculations for discrete bounded outcome scores.

We consider power and sample size calculations for randomized trials with a bounded outcome score (BOS) as primary response adjusted for a priori chosen covariates. We define BOS to be a random variable restricted to a finite interval. Typically, a BOS has a J- or U-shaped distribution hindering traditional parametric methods of analysis. When no adjustment for covariates is needed, a non-parametric test could be chosen. However, there is still a problem with calculating the power since the common location-shift alternative does not hold in general for a BOS. In this paper, we consider a parametric approach and assume that the observed BOS is a coarsened version of a true BOS, which has a logit-normal distribution in each treatment group allowing correction for covariates. A two-step procedure is used to calculate the power. Firstly, the power function is defined conditionally on the covariate values. Secondly, the marginal power is obtained by averaging the conditional power with respect to an assumed distribution for the covariates using Monte Carlo integration. A simulation study evaluates the performance of our method which is also applied to the ECASS-1 stroke study.

Computer Simulation↗

Planning genetic studies in human stroke: sample size estimates based on family history data.

BACKGROUND: Identification of stroke risk genes in humans has relied on case-control methods to determine the association between candidate genes and disease. Alternative approaches include linkage analysis using affected sibling pairs, transmission disequilibrium testing (TDT), and sibling TDT (S-TDT). Despite theoretical benefits, the feasibility of these methods in stroke remains unknown. METHODS: Family history was determined in 727 patients with ischemic stroke and 623 control subjects. These data were used to estimate the number of stroke patients required for the different study designs. RESULTS: A family history of any stroke occurring at < or =65 years was an independent risk factor for ischemic stroke at all ages (OR 1.47, 95% CI 1.02 to 2.12, p = 0.04) and a stronger risk factor for young (< or =65 years) ischemic stroke (OR 2.25, 95% CI 1.43 to 3.55, p < 0.0001). For early-onset ischemic stroke, the sibling risk ratio was estimated to be 3.08. Assuming three major stroke loci, collection of 953 affected sibling pairs (both < or =65 years) would be needed for a linkage study, and 115,472 ischemic stroke patients would have to be screened to achieve this sample size from the authors' population. The predicted sample sizes for association studies to detect a gene conferring an OR of 2.0 were case-control methodology (414), TDT (414), and S-TDT (617), which would require screening of 820, 31,680, and 3,062 cases. CONCLUSION: Alternative genetic approaches are feasible, but TDT and linkage studies using the affected sib-pair methodology may require large multicenter collaborations. S-TDT approaches appear more practical. These estimates will aid in planning of such studies.

Aged↗

Estimated coefficient of variation values for sample size planning in bioequivalence studies.

OBJECTIVE: The aim of the present communication is to provide information regarding the intrasubject coefficent of variation obtained from 30 bioequivalence studies covering 16 drugs which can be used for estimation of sample size. Additionally, an attempt was also made to estimate the test power of each of the studies conducted. METHODS: The intrasubject coefficient of variation was estimated from the residual mean square error obtained from analysis of variance of the parameters AUC0-infinity, Cmax and Cmax/AUC0-infinity after logarithmic transformation. The test power in the analyses of the above parameters was subsequently estimated using nomograms provided by Diletti et al. [1991]. RESULTS AND CONCLUSION: Thirty products covering 16 drugs were studied in which 22 were immediate-release (including one dispersible tablet) and 8 were sustained-release formulations. The intrasubject coefficient of variation for the parameter AUC0-infinity was smaller than Cmax, and hence considerably more studies were able to attain a power of greater than 80% using 12 volunteers for the AUC0-infinity, compared to the Cmax. However, the variability in the Cmax could be reduced by using the parameter Cmax/ AUC0-infinity, and thus, provide a more realistic estimation of sample size, since the latter reflects only the rate of absorption and not both the rate and extent as in the case of Cmax [Endrenyi et al. 1991].

Absorption↗

A simple formula for sample size calculation in equivalence studies.

Bioequivalence and clinical equivalence can be claimed based on the two one-sided test approach or the confidence interval approach. Consequently the power function of the equivalence test can be derived from either noncentral t-distribution or central t-distribution. The sample size is then determined from the power function either by numerical method or closed formulas. In this paper, we propose a simple formula for sample size calculation based on central t-distribution. The proposed formula has better properties than those currently available and it can be easily applied in all equivalence studies.

Confidence Intervals↗

Adaptive sample size calculations in group sequential trials.

A method for group sequential trials that is based on the inverse normal method for combining the results of the separate stages is proposed. Without exaggerating the Type I error rate, this method enables data-driven sample size reassessments during the course of the study. It uses the stopping boundaries of the classical group sequential tests. Furthermore, exact test procedures may be derived for a wide range of applications. The procedure is compared with the classical designs in terms of power and expected sample size.

Acne Vulgaris↗

Substantial effective sample sizes were required for external validation studies of predictive logistic regression models.

BACKGROUND AND OBJECTIVES: The performance of a prediction model is usually worse in external validation data compared to the development data. We aimed to determine at which effective sample sizes (i.e., number of events) relevant differences in model performance can be detected with adequate power. METHODS: We used a logistic regression model to predict the probability that residual masses of patients treated for metastatic testicular cancer contained only benign tissue. We performed standard power calculations and Monte Carlo simulations to estimate the numbers of events that are required to detect several types of model invalidity with 80% power at the 5% significance level. RESULTS: A validation sample with 111 events was required to detect that a model predicted too high probabilities, when predictions were on average 1.5 times too high on the odds scale. A decrease in discriminative ability of the model, indicated by a decrease in the c-statistic from 0.83 to 0.73, required 81 to 106 events, depending on the specific scenario. CONCLUSION: We suggest a minimum of 100 events and 100 nonevents for external validation samples. Specific hypotheses may, however, require substantially higher effective sample sizes to obtain adequate power.

Humans↗