Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Sample size planning for developing classifiers using high-dimensional DNA microarray data.

Many gene expression studies attempt to develop a predictor of pre-defined diagnostic or prognostic classes. If the classes are similar biologically, then the number of genes that are differentially expressed between the classes is likely to be small compared to the total number of genes measured. This motivates a two-step process for predictor development, a subset of differentially expressed genes is selected for use in the predictor and then the predictor constructed from these. Both these steps will introduce variability into the resulting classifier, so both must be incorporated in sample size estimation. We introduce a methodology for sample size determination for prediction in the context of high-dimensional data that captures variability in both steps of predictor development. The methodology is based on a parametric probability model, but permits sample size computations to be carried out in a practical manner without extensive requirements for preliminary data. We find that many prediction problems do not require a large training set of arrays for classifier development.

Computer Simulation↗

The study of candidate genes in drug trials: sample size considerations.

With discovery of an increasing number of candidate genes that may affect inter-individual variability in response to drugs, the design of drug trials that incorporate their study has become relevant. We discuss the determination of sample size for such studies when the number of tests to perform is given, or, alternatively, the number of tests to perform when the sample size is given. In many cases, a uniformly most powerful test does not exist and normal approximations are not sufficiently accurate to determine sample size. We discuss briefly various tests of interest and we give simple examples to illustrate some of the problems that arise.

Clinical Trials, Phase I as Topic↗

Sample sizes for constructing confidence intervals and testing hypotheses.

Although estimation and confidence intervals have become popular alternatives to hypothesis testing and p-values, statisticians usually determine sample sizes for randomized clinical trials by controlling the power of a statistical test at an appropriate alternative, even those statisticians who recommend the use of confidence intervals for inference. There is merit in achieving consistency in the techniques for data analysis and sample size determination. To that end, this paper compares sample size determination with use of the length of the confidence interval with that obtained by control of power.

Angina Pectoris↗

Sample size requirements for addressing the population genetic issues of forensic use of DNA typing.

DNA typing offers a unique opportunity to identify individuals for medical and forensic purposes. Probabilistic inference regarding the chance occurrence of a match between the DNA type of an evidentiary sample and that of an accused suspect, however, requires reliable estimation of genotype and allele frequencies in the population. Although population-based data on DNA typing at several hypervariable loci are being accumulated at various laboratories, a rigorous treatment of the sample size needed for such purposes has not been made from population genetic considerations. It is shown here that the loci that are potentially most useful for forensic identification of individuals have the intrinsic property that they involve a large number of segregating alleles, and a great majority of these alleles are rare. As a consequence, because of the large number of possible genotypes at the hypervariable loci that offer the maximum potential for individualization, the sample size needed to observe all possible genotypes in a sample is large. In fact, the size is so large that even if such a huge number of individuals could be sampled, it could not be guaranteed that such a sample was drawn from a single homogeneous population. Therefore adequate estimation of genotypic probabilities must be based on allele frequencies, and the sample size needed to represent all possible alleles is far more reasonable. Further economization of sample size is possible if one wants to have representation of only the frequent alleles in the sample, so that the rare allele frequencies can be approximated by an upper bound for forensic applications.

Alleles↗

Effects of reducing sample size on density estimates of citrus rust mite (Acari: Eriophyidae) on citrus fruit: simulated sampling.

The consequence of reducing sample size on the accuracy and precision of estimates of citrus rust mite, Phyllocoptruta oleivora (Ashmead), densities on oranges was investigated. The sample unit was a 1-cm2 surface area on fruit. Sampling plans consisting of 360, 300, 200, 160, 80, 48, 36, or 20 samples per 4 ha were evaluated through computer simulations by using real count data from 32 data sets of 600 sample units per 4 ha. The original and reduced sampling plans were hierarchical with different numbers of sample areas per 4 ha, trees per area, fruit per tree, and samples per fruit. Individual estimates (n=100 simulations per data set) using each plan were sometimes considerably below or above target densities. In an original set of count data with a mean of six mites per cm2, simulations of 36 samples per 4 ha produced individual estimates ranging from one to 16 mites per cm2, whereas 80 samples per 4 ha produced estimates ranging from two to 10 mites per cm2. The plans consisting of 36 or more samples were projected to provide precision levels of 0.25 (SEM/mean) or better at densities of five or more mites per cm2 based on log-data, a projection that needs to be verified under real-grove situations. Each plan consistently provided mite detection in these sampling simulations except those consisting of 20 or 36 samples, which sometimes failed to detect mites when the target density was less than five mites per cm2. The study provided insight into the probable precision, accuracy and detection thresholds for eight candidate sampling plans varying from relatively low to high resource input.

Acari↗

Sample size and power estimation for studies with health related quality of life outcomes: a comparison of four methods using the SF-36.

We describe and compare four different methods for estimating sample size and power, when the primary outcome of the study is a Health Related Quality of Life (HRQoL) measure. These methods are: 1. assuming a Normal distribution and comparing two means; 2. using a non-parametric method; 3. Whitehead's method based on the proportional odds model; 4. the bootstrap. We illustrate the various methods, using data from the SF-36. For simplicity this paper deals with studies designed to compare the effectiveness (or superiority) of a new treatment compared to a standard treatment at a single point in time. The results show that if the HRQoL outcome has a limited number of discrete values (< 7) and/or the expected proportion of cases at the boundaries is high (scoring 0 or 100), then we would recommend using Whitehead's method (Method 3). Alternatively, if the HRQoL outcome has a large number of distinct values and the proportion at the boundaries is low, then we would recommend using Method 1. If a pilot or historical dataset is readily available (to estimate the shape of the distribution) then bootstrap simulation (Method 4) based on this data will provide a more accurate and reliable sample size estimate than conventional methods (Methods 1, 2, or 3). In the absence of a reliable pilot set, bootstrapping is not appropriate and conventional methods of sample size estimation or simulation will need to be used. Fortunately, with the increasing use of HRQoL outcomes in research, historical datasets are becoming more readily available. Strictly speaking, our results and conclusions only apply to the SF-36 outcome measure. Further empirical work is required to see whether these results hold true for other HRQoL outcomes. However, the SF-36 has many features in common with other HRQoL outcomes: multi-dimensional, ordinal or discrete response categories with upper and lower bounds, and skewed distributions, so therefore, we believe these results and conclusions using the SF-36 will be appropriate for other HRQoL measures.

Algorithms↗

REPLI: a program in BASIC for determination of approximate sample size.

REPLI, a program written in elementary BASIC, calculates the approximate sample size, which is required to detect a desired difference between any two group means in an experiment with n groups for a given probability and at three significance levels of the means difference. A prior knowledge of the variability of data in the groups is expected in order to base the estimate on a rational footing. If this knowledge does not exist, an educated guess and/or several trials with different assumptions on the most likely variability can be used. The a priori estimate prevents that sample sizes are completely out of a reasonable range. Since the program is applicable for experimental settings where several groups need be investigated, it is particularly interesting to users of ANOVA and/or comparable non-parametric tests.

Analysis of Variance↗

Sample size determination under an exponential model in the presence of a confounder and type I censoring.

In controlled clinical trials, random assignment of treatments to individuals is usually used to eliminate the effects of confounding variables. When there is censorship in data, however, confounding effects may not be automatically removed solely by random assignment of treatments to individuals under the exponential model. Therefore, it is important to incorporate the confounding effect into the sample size calculation even after randomization of treatments to individuals. In this paper, the discussion is restricted only to the situation where there are two comparison groups and one single Bernoulli confounding variable. Based on an exponential covariate model, an explicit sample size formula considering the confounding effect has been derived for the design of trials with type I censoring, in which an end time is fixed in advance and all responses occurring after that time are censored. The resulting sample size formula can also be applied to nonrandomized clinical trials. Finally, to provide insight into the influence of different factors on sample size calculation, a discussion on the effects of treatments, the confounder, the length of follow-up times for studied individuals, and the joint distribution of the treatment and the confounder has been included.

Clinical Trials as Topic↗

Sample size calculations for controlled clinical trials using generalized estimating equations (GEE).

OBJECTIVES: Clinical trials with correlated response data based on generalized estimating equations (GEE) have become increasingly popular as they require smaller samples than classical methods that ignore the clustered nature of the data. We have recently derived the recommendation to use the independence estimating equations (IEE) as primary analysis in most controlled clinical trials instead of GEE with estimated correlations. Although several approaches for sample size and power calculation have been proposed, we have shown that most of these procedures are very specific and not as general as required for designing clinical trials. METHODS: We extended the previously developed SAS macro GEESIZE to overcome this restriction. Specifically, we have added the option of an independence working correlation matrix required for the IEE. Additionally, we have reformulated the hypotheses to allow for coding that includes an intercept term instead of the previously used analysis of variance coding. RESULTS: To demonstrate the validity of GEESIZE we investigate the calculated sample sizes for specific models where closed formulae are available. For illustration, we utilize GEESIZE for planning a new trial on the treatment of hypertension and thereby exemplify its flexibility. CONCLUSIONS: We show that our freely available macro is a very general and useful tool for sample size calculation purposes in clinical trials with correlated data.

Cluster Analysis↗

Sample size determination for phase II studies of new vaccines.

Prior to the evaluation of protective efficacy, experimental vaccines conventionally undergo phase II randomized controlled clinical trials to evaluate safety and immunogenicity. Typically, an experimental vaccine is compared to another vaccine or to a placebo with respect to adverse events or immune responses, or both. Various strategies and methods are available for design and analysis of such studies. A key aspect of design is the determination of sample size. Often a sample size is chosen that gives a high probability ("power") of finding a statistically significant difference in an outcome of interest, if a difference of a specified size exists. This approach is appropriate when the primary goal of the study is to demonstrate that a difference exists between two groups or treatments. It may not, however, give adequate assurance that a confidence interval around the observed difference will be narrow enough to exclude the possibility of an unacceptably low immune response or unacceptably high adverse event frequency in recipients of the experimental vaccine. In this paper, we apply the "non-inferiority" trial design to phase II vaccine studies; that is, we design the trial to rule out a difference between the vaccine and control in immunogenicity or reactogenicity that is considered unacceptable. We also consider a setting in which the desire is to show that the difference between immune response rates for vaccine and control is greater than a specified value.

Cholera Vaccines↗

Bayesian sample size calculations in phase II clinical trials using a mixture of informative priors.

A number of researchers have discussed phase II clinical trials from a Bayesian perspective. A recent article by Mayo and Gajewski focuses on sample size calculations, which they determine by specifying an informative prior distribution and then calculating a posterior probability that the true response will exceed a prespecified target. In this article, we extend these sample size calculations to include a mixture of informative prior distributions. The mixture comes from several sources of information. For example consider information from two (or more) clinicians. The first clinician is pessimistic about the drug and the second clinician is optimistic. We tabulate the results for sample size design using the fact that the simple mixture of Betas is a conjugate family for the Beta- Binomial model. We discuss the theoretical framework for these types of Bayesian designs and show that the Bayesian designs in this paper approximate this theoretical framework.

Algorithms↗

Extremely small sample size in some toxicity studies: an example from the rabbit eye irritation test.

The conventional sample-size equations based on either the precision of estimation or the power of testing a hypothesis may not be appropriate to determine sample size for a "diagnostic" testing problem, such as the eye irritant Draize test. When the animals' responses to chemical compounds are relatively uniform and extreme and the objective is to classify a compound as either irritant or nonirritant, the test using just two or three animals may be adequate.

Animals↗

A sample size adjustment procedure for clinical trials based on conditional power.

When designing clinical trials, researchers often encounter the uncertainty in the treatment effect or variability assumptions. Hence the sample size calculation at the planning stage of a clinical trial may also be questionable. Adjustment of the sample size during the mid-course of a clinical trial has become a popular strategy lately. In this paper we propose a procedure for calculating additional sample size needed based on conditional power, and adjusting the final-stage critical value to protect the overall type-I error rate. Compared to other previous procedures, the proposed procedure uses the definition of the conditional type-I error directly without appealing to an extra special function for it. It has better flexibility in setting up interim decision rules and the final-stage test is a likelihood ratio test.

Journal Article↗

Approximate estimation of minimal sample size required for marker-assisted backcross breeding.

Backcross breeding is a useful method to transfer favorable alleles from a donor parent into a recipient parent. Marker-assisted selection (MAS) can speed up the process. To make an appropriate plan before using MAS in a breeding program, breeders need to know the minimal sample size of the progeny generation required. A method to estimate the minimal sample size required for marker-assisted backcross breeding when both foreground selection and background selection are conducted is proposed. On the basis of a simplified assumption that the target loci are introgressed and the genetic background are independent, the probability of selecting individuals with desired genotypes in each generation is approximately estimated combining analytical approach (for foreground selection) and simulation according to the graphic genotypes of backcrossing parents (for background selection). The minimal sample size required to obtain at least one desired individual with a given probability is estimated. Application of the method is demonstrated with hypothesized examples. The method can be conveniently applied to practical backcross breeding programs.

Alleles↗

Power and sample size calculations for case-control studies of gene-environment interactions with a polytomous exposure variable.

Genetic polymorphisms may appear to the epidemiologist most commonly as different levels of susceptibility to exposure. Epidemiologic studies of heterogeneity in exposure susceptibility aim at estimating the parameter quantifying the gene-environment interaction. In this paper, the authors use a general approach to power and sample size calculations for case-control studies, which is applicable to settings where the exposure variable is polytomous and where the assumption of independence between the distribution of the genotype and the environmental factor may not be met. It was found through exploration of different scenarios that in the cases explored, power calculations were relatively insensitive to assumptions about the odds ratio for the exposure in the referent genotype category and to assumptions about the odds ratio for the genetic factor in the lowest exposure category, yet they were relatively sensitive to assumptions about gene frequency, particularly when gene frequency was low. In general, to detect a small to moderate gene-environment interaction effect, large sample sizes are needed. Because the examples studied represent only a small subset of possible scenarios that could occur in practice, the authors encourage the use of their user-friendly Fortran program for calculating power and sample size for gene-environment interactions with exposures grouped by quantiles that are explicitly tailored to the study at hand.

Case-Control Studies↗

Sample size calculations when outcomes will be compared with an historical control.

The usual formulae for calculating sample size are not valid when a treated cohort will be compared with an historical control population. Methods appropriate for a single group study may greatly underestimate the required sample size because they, in effect, assume a control population of infinite size. Methods appropriate for a two-group study require specification of the proportions of patients in the two groups--a requirement that can not be met when the size of the historical control group is fixed. An easily programmed iterative solution is presented and simulations indicating the validity of the method are performed.

Algorithms↗

Sample size nomograms for interpreting negative clinical studies.

In recent years there has been increasing attention to the appropriate interpretation of a clinical study. One special concern has been the difficulty inherent in interpreting studies that were not statistically significant: Was the sample size sufficient to detect a clinically important effect if, in fact, it existed? This concern is further complicated because readers may have differing opinions of what size effect is clinically important. A pair of sample size nomograms has been developed, using common levels of statistical significance, to assist in this interpretation. The nomograms are intended to provide the clinician with a handy and easy-to-use reference for ascertaining whether an apparently negative study has a sample size adequate to detect reliably any difference between treatment groups that the clinician believes is clinically important. Examples are provided to show these principles and the use of the nomograms in interpreting negative studies.

Clinical Trials as Topic↗

Rule-of-thumb adjustment of sample sizes to accommodate dropouts in a two-stage analysis of repeated measurements.

Recent contributions to the statistical literature have provided elegant model-based solutions to the problem of estimating sample sizes for testing the significance of differences in mean rates of change across repeated measures in controlled longitudinal studies with differentially correlated error and missing data due to dropouts. However, the mathematical complexity and model specificity of these solutions make them generally inaccessible to most applied researchers who actually design and undertake treatment evaluation research in psychiatry. In contrast, this article relies on a simple two-stage analysis in which dropout-weighted slope coefficients fitted to the available repeated measurements for each subject separately serve as the dependent variable for a familiar ANCOVA test of significance for differences in mean rates of change. This article is about how a sample of size that is estimated or calculated to provide desired power for testing that hypothesis without considering dropouts can be adjusted appropriately to take dropouts into account. Empirical results support the conclusion that, whatever reasonable level of power would be provided by a given sample size in the absence of dropouts, essentially the same power can be realized in the presence of dropouts simply by adding to the original dropout-free sample size the number of subjects who would be expected to drop from a sample of that original size under conditions of the proposed study.

Analysis of Variance↗