Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Determining the sample size for co-dominant molecular marker-assisted linkage detection for a monogenic qualitative trait by controlling the type-I and type-II errors in a segregating F2 population.

Tests for linkage are usually performed using the lod score method. A critical question in linkage analyses is the choice of sample size. The appropriate sample size depends on the desired type-I error and power of the test. This paper investigates the exact type-I error and power of the lod score method in a segregating F(2) population with co-dominant markers and a qualitative monogenic dominant-recessive trait. For illustration, a disease-resistance trait is considered, where the susceptible allele is recessive. A procedure is suggested for finding the appropriate sample size. It is shown that recessive plants have about twice the information content of dominant plants, so the former should be preferred for linkage detection. In some cases the exact alpha-values for a given nominal alpha may be rather small due to the discrete nature of the sampling distribution in small samples. We show that a gain in power is possible by using exact methods.

Crosses, Genetic↗

Sample size tables for receiver operating characteristic studies.

OBJECTIVE: I provide researchers with tables of sample size for multiobserver receiver operating characteristic (ROC) studies that compare the diagnostic accuracies of two imaging techniques. MATERIALS AND METHODS: I computed the number of patients and observers needed as a function of five parameters: the measure of diagnostic accuracy (area under the ROC curve, sensitivity at a false-positive rate </= 0.10, or specificity at a false-negative rate </= 0.10), conjectured level of accuracy, suspected difference in accuracy between the two imaging techniques, observer variability, and ratio of patients without to patients with the condition. RESULTS: The numbers of patients and observers required vary dramatically with these five parameters, increasing with more refined measures of accuracy, with lower accuracy levels, with smaller suspected differences, with greater observer variability, and with less balanced designs. The number of patients required for a study can be reduced by increasing the number of observers, and vice versa. When the intra- and interobserver variability is large, a study design with just four observers is usually inadequate. CONCLUSION: Many factors must be considered when determining the appropriate sample sizes for multiobserver ROC studies. My tables serve only as initial ballpark estimates. Investigators should compute sample size using parameters that reflect their clinical application.

ROC Curve↗

Type I error in sample size re-estimations based on observed treatment difference.

Sample size re-estimation based on an observed difference can ensure an adequate power and potentially save a large amount of time and resources in clinical trials. One of the concerns for such an approach is that it may inflate the type I error. However, such a possible inflation has not been mathematically quantified. In this paper the mathematical mechanism of this inflation is explored for two-sample normal tests. A (conditional) type I error function based on normal data is derived. This function not only provides the quantification but also gives mathematical mechanisms of possible inflation in the type I error due to the sample size re-estimation. Theoretically, based on their decision rules (certain upper and lower bounds), people can calculate this function and exactly visualize the changes in type I error. Computer simulations are performed to ensure the results. If there are no bounds for the adjustment, the inflation is evident. If proper adjusting rules are used, the inflation can be well controlled. In some cases the type I error can even be reduced. The trade-off is to give up some 'unrealistic power'. We investigated several scenarios in which the mechanisms to change the type I error are different. Our simulations show that similar results may apply to other distributions.

Clinical Trials as Topic↗

Multiplicity-adjusted sample size requirements: a strategy to maintain statistical power with Bonferroni adjustments.

BACKGROUND: A researcher must carefully balance the risk of 2 undesirable outcomes when designing a clinical trial: false-positive results (type I error) and false-negative results (type II error). In planning the study, careful attention is routinely paid to statistical power (i.e., the complement of type II error) and corresponding sample size requirements. However, Bonferroni-type alpha adjustments to protect against type I error for multiple tests are often resisted. Here, a simple strategy is described that adjusts alpha for multiple primary efficacy measures, yet maintains statistical power for each test. METHOD: To illustrate the approach, multiplicity-adjusted sample size requirements were estimated for effects of various magnitude with statistical power analyses for 2-tailed comparisons of 2 groups using chi2 tests and t tests. These analyses estimated the required sample size for hypothetical clinical trial protocols in which the prespecified number of primary efficacy measures ranged from 1 to 5. Corresponding Bonferroni-adjusted alpha levels were used for these calculations. RESULTS: Relative to that required for 1 test, the sample size increased by about 20% for 2 dependent variables and 30% for 3 dependent variables. CONCLUSION: The strategy described adjusts alpha for multiple primary efficacy measures and, in turn, modifies the sample size to maintain statistical power. Although the strategy is not novel, it is typically overlooked in psychopharmacology trials. The number of primary efficacy measures must be prespecified and carefully limited when a clinical trial protocol is prepared. If multiple tests are designated in the protocol, the alpha-level adjustment should be anticipated and incorporated in sample size calculations.

Clinical Protocols↗

Sample size requirements for association studies of gene-gene interaction.

In the study of complex diseases, it may be important to test hypotheses related to gene-gene (G x G) interaction. The success of such studies depends critically on obtaining adequate sample sizes. In this paper, the author investigates sample size requirements for studies of G x G interaction, focusing on four study designs: the matched-case-control design, the case-sibling design, the case-parent design, and the case-only design. All four designs provide an estimate of interaction on a multiplicative scale, which is used as a unifying theme in the comparison of sample size requirements. Across a variety of genetic models, the case-only and case-parent designs require fewer sampling units (cases and case-parent trios, respectively) than the case-control (pairs) or case-sibling (pairs) design. For example, the author describes an asthma study of two common recessive genes for which 270 matched case-control pairs would be required to detect a G x G interaction of moderate magnitude with 80% power. By comparison, the same study would require 319 case-sibling pairs but only 146 trios in the case-parent design or 116 cases in the case-only design. A software program that computes sample size for studies of G x G interaction and for studies of gene-environment (G x E) interaction is freely available (http://hydra.usc.edu/gxe).

Case-Control Studies↗

Sample size determination for pair-matched case-control studies where the goal is interval estimation of the odds ratio.

Samples sizes are calculated for case-control studies where 1:1 matching has been employed, and where the goal is the interval estimation of the odds ratio. The optimal sample size is defined to be the smallest value for which a 100(1 - alpha)% confidence interval for the log odds ratio will not exceed a specified width 2 delta with specified probability (1 - gamma). This approach is similar in spirit to the power-based approach for sample size determination when significance testing is the goal. Tables of sample sizes are presented for various choices of parameters. We also find considerable disagreement with a published method based on expected numbers of discordant pairs.

Case-Control Studies↗

The CellSoft computerized semen analysis system. I. Consistency of measurements and stability of results in relation to sample size analyzed.

The CellSoft TM computer-assisted digital image analysis system was evaluated for reproducibility of measurements and for the sample size needed to obtain stable results for seven parameters characterizing human spermatozoa. The parameters were sperm density, velocity, linearity, percent motile cells, maximum and mean amplitude of sperm head's lateral displacement (ALH max and ALH mean) and beat cross frequency. Thirty men were assigned to three equally large groups according to sperm density, and the groups were studied separately. Consistency of measurements was highly acceptable for all parameters (coefficient of variation ranging from 0.9% to 3.7%). The average minimum number of sample size needed for stable results varied between the groups and within each group dependent upon the parameter studied. Velocity and linearity seemed to require the lowest average cell number, while percentage motile cells and ALH max needed the highest sample size. Our current recommendation of sample sizes is outlined in detail and limitations and future utility of CellSoft is discussed.

Evaluation Studies as Topic↗

Sample size determination for multiple comparison studies treating confidence interval width as random.

Methods for optimal sample size determination are developed using four popular multiple comparison procedures (Scheffe's, Bonferroni's, Tukey's and Dunnett's procedures), where random samples of the same size n are to be selected from k (>/=2) normal populations with common variance sigma2, and where primary interest concerns inferences about a family of L linear contrasts among the k population means. For a simultaneous coverage probability of (1-alpha), the optimal sample size is defined to be the smallest integer value n*m such that, simultaneously for all L confidence intervals, the width of the lth confidence interval will be no greater than tolerance 2deltal (l=1,2,...,L) with tolerance probability at least (1-gamma), treating the pooled sample variance S2p as a random variable. Using Scheffe's procedure as an illustration, comparisons are made to usual sample size methods that incorrectly ignore the stochastic nature of S2p. The latter approach can lead to serious underestimation of required sample sizes and hence to unacceptably low values of the actually tolerance probability (1-gamma'). Our approach guarantees a lower bound of [1-(alpha+gamma)] for the probability that the L confidence intervals will both cover the parametric functions of interest and also be sufficiently narrow. Recommendations are provided regarding the choices among the four multiple comparison procedures for sample size determination and inference-making.

Confidence Intervals↗

On power and sample size for studying features of the relative odds of disease.

Estimates of sample size and statistical power are essential ingredients in the design of epidemiologic studies. Once an association between disease and exposure has been demonstrated, additional studies are often needed to investigate special features of the relation between exposure, other covariates, and risk of disease. The authors present a general formulation to compute sample size and power for case-control and cohort studies to investigate more complex patterns in the odds ratios, such as to distinguish between two different slopes of linear trend, to distinguish between two possible dose-response relations, or to distinguish different models for the joint effects of two important exposures or of one exposure factor adjusting for another. Such special studies of exposure-response relations may help investigators to distinguish between plausible biologic models and may lead to more realistic models for calculating attributable risk and lifetime disease risk. The sample size formulae are applied to studies of indoor radon exposure and lung cancer and suggest that epidemiologic studies may not be feasible for addressing some issues. For example, if the risk estimates from underground miners' studies are, in truth, not applicable to home exposures and overestimate the gradient of risk from home exposure to radon by, for example, a factor of 2, then enormously large numbers of subjects would be required to detect the difference. Furthermore, if the true interaction between smoking and radon exposure is less than multiplicative, only the largest investigations will have sufficient power to reject additivity. For the simple case of testing for no exposure effect, when exposure is either dichotomous or continuous, these methods yield well-known formulae.

Aged↗

Sample sizes for individually matched case-control studies: a group sequential approach.

This paper proposes the use of group sequential methods to calculate sample sizes for individually matched case-control study designs. A table is presented in which the average sample size required for a group sequential (i.e., multistage) matched pair design is compared to that of the conventional matched pair fixed sample size plan for the usual constant relative risk situation. The table shows that group sequential designs are in general more efficient than fixed sample size plans. Computer simulations showed that group sequential methods yield the appropriate type I and type II error rates not only for matching on a one-to-one basis, but also more generally with multiple matched controls per case. Further simulation studies indicated that there may be only a small loss of power when the matching variable(s) is associated with the probability of exposure but not with the disease. This is shown for both the multistage and fixed sample tests.

Epidemiologic Methods↗

Variance estimation in clinical studies with interim sample size re-estimation.

We consider clinical studies with a sample size re-estimation based on the unblinded variance estimation at some interim point of the study. Because the sample size is determined in such a flexible way, the usual variance estimator at the end of the trial is biased. We derive sharp bounds for this bias. These bounds have a quite simple form and can help for the decision if this bias is negligible for the actual study or if a correction should be done. An exact formula for the bias is also provided. We discuss possibilities to get rid of this bias or at least to reduce the bias substantially. For this purpose, we propose a certain additive correction of the bias. We see in an example that the significance level of the test can be controlled when this additive correction is used.

Analysis of Variance↗

Sample sizes for estimation of exposure-specific disease rates in population-based case-control studies.

This paper discusses sample sizes for estimation of exposure-specific disease rates for population-based case-control studies. Neutra and Drolette's confidence limits, which are based on the approximate normality of the logarithm of the ratio of independent binomial exposure rates, are used to determine the sample sizes required for precise estimation of exposure-specific disease rates. It is shown that, for large sample sizes, the disease rate in the exposed population is more precisely estimated than the disease rate in the unexposed population when more than 50% of the cases are exposed, and that the converse is true when fewer than 50% of the cases are exposed. Expressions are derived for the optimal case and control sample sizes that ensure the required level of precision and minimize the total study size. The optimum control-to-case ratio is found to be equal to the square root of the exposure odds ratio. The optimum number of cases and the total study size are found to be smaller for precise estimation of the disease rate in the exposed population than for precise estimation of the exposure odds ratio when the disease is rare.

Humans↗

Sample sizes for bioequivalence studies.

In recent years a number of decision rules, based on sound statistical principles, have been proposed for deciding if a test formulation is bioequivalent to a reference formulation. The decision rule based on confidence intervals has been accepted by regulatory agencies, at least by the Food and Drug Administration of the United States. A useful property of this decision rule is that the regulatory agency need not require a certain sample size, since the level of protection against wrongly deciding bioequivalence is set by the choice of the alpha level used to compute the confidence intervals. The manufacturer claiming bioequivalence is concerned about sample size, for sample size determines the probability of falsely deciding non-bioequivalence when the test formulation does indeed have an acceptable relative bioavailability. Curves of probability of rejecting bioequivalence have been computed for error coefficient of variation of 10, 20 and 30 per cent, for relative bioavailability from 70 to 130 per cent, and for protection levels of 90 and 95 per cent. These curves can be used for choosing the sample size for a bioequivalence study.

Bayes Theorem↗

MSurvPow: a FORTRAN program to calculate the sample size and power for cluster-randomized clinical trials with survival outcomes.

Manatunga and Chen [A.K. Manatunga, S. Chen, Sample size estimation for survival outcomes in cluster-randomized studies with small cluster sizes, Biometrics 56 (2000) 616-621] proposed a method to estimate sample size and power for cluster-randomized studies where the primary outcome variable was survival time. The sample size formula was constructed by considering a bivariate marginal distribution (Clayton-Oakes model) with univariate exponential marginal distributions. In this paper, a user-friendly FORTRAN 90 program was provided to implement this method and a simple example was used to illustrate the features of the program.

Cluster Analysis↗

Sample size determinations using logistic regression with pilot data.

Suppose the goal of a projected study is to estimate accurately the value of a 'prediction' proportion p that is specific to a given set of covariates. Available pilot data show that (1) the covariates are influential in determining the value of p and (2) their relationship to p can be modelled as a logistic regression. A sample size justification for the projected study can be based on the logistic model; the resulting sample sizes not only are more reasonable than the usual binomial sample size values from a scientific standpoint (since they are based on a model that is more realistic), but also give smaller prediction standard errors than the binomial approach with the same sample size. In appropriate situations, the logistic-based sample sizes could make the difference between a feasible proposal and an unfeasible, binomial-based proposal. An example using pilot study data of dental radiographs demonstrates the methods.

Confidence Intervals↗

Increasing the sample size during clinical trials with t-distributed test statistics without inflating the type I error rate.

In clinical trials with t-distributed test statistics the required sample size depends on the unknown variance. Taking estimates from previous studies often leads to a misspecification of the true value of the variance. Hence, re-estimation of the variance based on the collected data and re-calculation of the required sample size is attractive. We present a flexible method for extensions of fixed sample or group-sequential trials with t-distributed test statistics. The method can be applied at any time during the course of the trial and does not require the necessity to pre-specify a sample size re-calculation rule. All available information can be used to determine the new sample size. The advantage of our method when compared with other adaptive methods is maintenance of the efficient t-test design when no extensions are actually made. We show that the type I error rate is preserved.

Computer Simulation↗

Necessary sample size for method comparison studies based on regression analysis.

BACKGROUND: In method comparison studies, it is of importance to assure that the presence of a difference of medical importance is detected. For a given difference, the necessary number of samples depends on the range of values and the analytical standard deviations of the methods involved. For typical examples, the present study evaluates the statistical power of least-squares and Deming regression analyses applied to the method comparison data. METHODS: Theoretical calculations and simulations were used to consider the statistical power for detection of slope deviations from unity and intercept deviations from zero. For situations with proportional analytical standard deviations, weighted forms of regression analysis were evaluated. RESULTS: In general, sample sizes of 40-100 samples conventionally used in method comparison studies often must be reconsidered. A main factor is the range of values, which should be as wide as possible for the given analyte. For a range ratio (maximum value divided by minimum value) of 2, 544 samples are required to detect one standardized slope deviation; the number of required samples decreases to 64 at a range ratio of 10 (proportional analytical error). For electrolytes having very narrow ranges of values, very large sample sizes usually are necessary. In case of proportional analytical error, application of a weighted approach is important to assure an efficient analysis; e.g., for a range ratio of 10, the weighted approach reduces the requirement of samples by >50%. CONCLUSIONS: Estimation of the necessary sample size for a method comparison study assures a valid result; either no difference is found or the existence of a relevant difference is confirmed.

Clinical Laboratory Techniques↗

Sample size calculations for surveys to substantiate freedom of populations from infectious agents.

We develop a Bayesian approach to sample size computations for surveys designed to provide evidence of freedom from a disease or from an infectious agent. A population is considered "disease-free" when the prevalence or probability of disease is less than some threshold value. Prior distributions are specified for diagnostic test sensitivity and specificity and we test the null hypothesis that the prevalence is below the threshold. Sample size computations are developed using hypergeometric sampling for finite populations and binomial sampling for infinite populations. A normal approximation is also developed. Our procedures are compared with the frequentist methods of Cameron and Baldock (1998a, Preventive Veterinary Medicine34, 1-17.) using an example of foot-and-mouth disease. User-friendly programs for sample size calculation and analysis of survey data are available at http://www.epi.ucdavis.edu/diagnostictests/.

Animals↗