Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Errors and linkage disequilibrium interact multiplicatively when computing sample sizes for genetic case-control association studies.

Single nucleotide polymorphisms (SNP) may be used in case-control designs to test for association between a SNP marker and a disease. Such designs may assume that the genotype data are reported without error. Our goal is quantifying the effects that errors have on sample size for case-control studies with haplotypes formed by a disease locus and a SNP marker locus in the presence of linkage disequilibrium (LD). We consider the effects of a recently published error model on 2x3 chi-square analysis. We study the joint relation of LD and errors with sample size for three specific genetic disease models and two settings each of marker allele frequencies (total of 6 studies). Minimal sample size necessary for fixed asymptotic power is estimated as a 4th degree polynomial in the variables S (error) and D' (LD measure) via a backward step-wise regression. We find that increased error rates lower power. In all studies, we observe that LD and errors interact in a non-linear fashion. In particular, regression analyses shows that several higher order interaction terms have coefficients significantly different from 0 in each study, with fraction of variance explained greater than 0.9999. Finally, the increase in sample size necessary to maintain constant asymptotic power and level of significance as a function of S is smallest when D' = 1 (perfect LD). The increase grows monotonically as D' decreases to 0.5 for all studies.

Case-Control Studies↗

Gene genealogies when the sample size exceeds the effective size of the population.

We study the properties of gene genealogies for large samples using a continuous approximation introduced by R. A. Fisher. We show that the major effect of large sample size, relative to the effective size of the population, is to increase the proportion of polymorphisms at which the mutant type is found in a single copy in the sample. We derive analytical expressions for the expected number of these singleton polymorphisms and for the total number of polymorphic, or segregating, sites that are valid even when the sample size is much greater than the effective size of the population. We use simulations to assess the accuracy of these predictions and to investigate other aspects of large-sample genealogies. Lastly, we apply our results to some data from Pacific oysters sampled from British Columbia. This illustrates that, when large samples are available, it is possible to estimate the mutation rate and the effective population size separately, in contrast to the case of small samples in which only the product of the mutation rate and the effective population size can be estimated.

DNA, Mitochondrial↗

The use of bootstrap methods for estimating sample size and analysing health-related quality of life outcomes.

Health-related quality of life (HRQoL) measures are increasingly used in trials as primary outcome measures. Investigators are now asking statisticians for advice on how to plan and analyse studies using such outcomes. HRQoL outcomes, like the SF-36, are usual measured on an ordinal scale, although most investigators assume that there exists an underlying continuous latent variable and that the actual measured outcomes (the ordered categories) reflect contiguous intervals along this continuum. The ordinal scaling of HRQoL measures means they tend to generate data that have discrete, bounded and skewed distributions. Thus, standard methods of analysis that assume Normality and constant variance may not be appropriate. For this reason, conventional statistical advice would suggest non-parametric methods be used to analyse HRQoL data. The bootstrap is one such computer intensive non-parametric method for estimating sample sizes and analysing data. We describe three methods of estimating sample sizes for two-group cross-sectional comparisons of HRQoL outcomes. We then compared the power of the three methods for a two-group cross-sectional study design using bootstrap simulation. The results showed that under the location shift alternative hypothesis, conventional methods of sample size estimation performed well, particularly Whitehead's method. Whitehead's method is recommended if the HRQoL outcome has a limited number of discrete values (<7) and/or the expected proportion of cases at either of the bounds is high. If a pilot data set is readily available then bootstrap simulation will provide a more accurate and reliable estimate, than conventional methods.Finally, we used the bootstrap for hypothesis testing and the estimation of standard errors and confidence intervals for parameters, in an example data set. We then compared and contrasted the bootstrap with standard methods of analysing HRQoL outcomes. In the data set studied, with the SF-36 outcome, the use of the bootstrap for estimating sample sizes and analysing HRQoL data produces results similar to conventional statistical methods. These results suggest that bootstrap methods are not more appropriate for analysing HRQoL outcome data than standard methods.

Adolescent↗

Comparative evaluation of two models for estimating sample sizes for tests on trends across repeated measurements.

Two equations for calculating sample sizes that are required for power in testing differences in rates of change in repeated measurement designs have been presented by different authors. One equation provides support for the conclusion that increased frequency of measurements across a treatment period of fixed duration enhances power of the tests. The other equation supports the counterintuitive conclusion that increased frequency of measurements actually tends to decrease power in the presence of realistic serial dependencies in the data. Monte Carlo methods confirm that the equation providing support for the latter conclusion is accurate, whereas the alternative equation tends to underestimate sample sizes required for power in testing differences in slopes of regression lines fitted to changes in the repeated measurements across time when symmetry is absent from the covariance structure.

Clinical Trials as Topic↗

Sample size for cohort studies in pharmacoepidemiology.

OBJECT: Cohort studies in pharmacoepidemiology can result in a unique type of study, where subjects have complex types of exposure to drugs (with periods of non-exposure as well). The object of this paper is to explain how to calculate the sample size of such a study. METHOD: It is assumed that adverse events follow Poisson distributions in the two study groups. The null hypothesis is that the two groups have equal rates of disease. Formulae are provided to calculate the sample size required to significantly reject the null hypothesis. Sample size is given as the number of events, rather than the number of subjects entered. In a Poisson study, it is the ratio of the amount of person-years exposure in the two groups that is important to calculate sample size, rather than the actual amounts of exposure (or number of subjects in the study). Some examples are included.

Journal Article↗

A note on sample size calculation for mean comparisons based on noncentral t-statistics.

One-sample and two-sample t-tests are commonly used in analyzing data from clinical trials in comparing mean responses from two drug products. During the planning stage of a clinical study, a crucial step is the sample size calculation, i.e., the determination of the number of subjects (patients) needed to achieve a desired power (e.g., 80%) for detecting a clinically meaningful difference in the mean drug responses. Based on noncentral t-distributions, we derive some sample size calculation formulas for testing equality, testing therapeutic noninferiority/superiority, and testing therapeutic equivalence, under the popular one-sample design, two-sample parallel design, and two-sample crossover design. Useful tables are constructed and some examples are given for illustration.

Algorithms↗

Sample size and power in psychiatric research.

The conclusions drawn by study are susceptible to two types of errors. The more familiar one occurs when it is believed that there was a true difference between the groups or an association between two variables, when in fact this observation was due to chance (a Type I error). The second potential error consists of falsely concluding that there was no difference or association when indeed there was one (a Type II error). Most researchers know that the probability of a Type I error can be controlled by the setting of alpha level of statistical significance; however, many are unaware of methods to control or estimate Type II errors, based on estimates of an appropriate sample size. This paper discusses techniques researchers can use to calculate the sample sizes required for studies, and the effects of sample sizes which are too small or too large. If it is too small, there is an increased risk of a Type II error, whereas if it is too large, there may be a needless waste of time, money, and effort. The paper also discusses how readers of research articles can determine whether or not negative findings reported by a study are a true reflection of the lack of any difference between groups, or a result of insufficient sample size.

Canada↗

The methods for handling missing data in clinical trials influence sample size requirements.

OBJECTIVE: Results of studies estimating osteoarthritis progression may be affected by missing values. In clinical trials assessing disease-modifying osteoarthritis drugs, sample sizes should be calculated using close estimates of outcome variables. STUDY DESIGN AND SETTING: Supposing a two-parallel group design in hip osteoarthritis clinical trials, we estimated sample sizes using the joint space width (JSW), number of patients with JSW progression >0.5 mm (JSN), time to total hip arthroplasty (THA), and time to JSN or THA using several approaches to deal with missing data. RESULTS: Three-year clinical trials testing a treatment effect of 50%, with a power of 80%, could require sample sizes of 121 patients for JSW, 57 for JS progression using multiple imputation for handling missing values; 200 for THA; and 47 for JSN or THA. These numbers vary greatly depending on the approach chosen for handling missing data. CONCLUSIONS: These results can help investigators plan clinical trials to select the primary outcome and a priori specify the way missing data will be handled.

Aged↗

Meta-analysis of data on costs from trials of counselling in primary care: using individual patient data to overcome sample size limitations in economic analyses.

OBJECTIVE: To assess the feasibility of overcoming sample size limitations in economic analyses of clinical trials through meta-analysis of data on individual patients from multiple trials. DESIGN: Meta-analysis of individual patient data from trials of counselling in primary care compared with usual care by a general practitioner. SETTING: Primary care. PATIENTS: People with mental health problems. MAIN OUTCOME MEASURES: Direct treatment costs, depressive symptoms, and cost effectiveness. RESULTS: Meta-analysis of individual patient data proved feasible. The results showed that the previous analyses of individual trials were underpowered to provide useful conclusions about the cost comparisons. The results are sensitive to assumptions made about the costs of sessions with a counsellor and the management of patients by a general practitioner. CONCLUSIONS: Meta-analysis of individual patient data may assist in overcoming sample size limitations in economic analyses. Although feasible, such analysis has shortcomings that may limit the validity of the results. The relative costs and benefits of this method, as opposed to further collection of primary data, are as yet unclear.

Clinical Trials as Topic↗

Modeling motor vehicle crashes using Poisson-gamma models: examining the effects of low sample mean values and small sample size on the estimation of the fixed dispersion parameter.

There has been considerable research conducted on the development of statistical models for predicting crashes on highway facilities. Despite numerous advancements made for improving the estimation tools of statistical models, the most common probabilistic structure used for modeling motor vehicle crashes remains the traditional Poisson and Poisson-gamma (or Negative Binomial) distribution; when crash data exhibit over-dispersion, the Poisson-gamma model is usually the model of choice most favored by transportation safety modelers. Crash data collected for safety studies often have the unusual attributes of being characterized by low sample mean values. Studies have shown that the goodness-of-fit of statistical models produced from such datasets can be significantly affected. This issue has been defined as the "low mean problem" (LMP). Despite recent developments on methods to circumvent the LMP and test the goodness-of-fit of models developed using such datasets, no work has so far examined how the LMP affects the fixed dispersion parameter of Poisson-gamma models used for modeling motor vehicle crashes. The dispersion parameter plays an important role in many types of safety studies and should, therefore, be reliably estimated. The primary objective of this research project was to verify whether the LMP affects the estimation of the dispersion parameter and, if it is, to determine the magnitude of the problem. The secondary objective consisted of determining the effects of an unreliably estimated dispersion parameter on common analyses performed in highway safety studies. To accomplish the objectives of the study, a series of Poisson-gamma distributions were simulated using different values describing the mean, the dispersion parameter, and the sample size. Three estimators commonly used by transportation safety modelers for estimating the dispersion parameter of Poisson-gamma models were evaluated: the method of moments, the weighted regression, and the maximum likelihood method. In an attempt to complement the outcome of the simulation study, Poisson-gamma models were fitted to crash data collected in Toronto, Ont. characterized by a low sample mean and small sample size. The study shows that a low sample mean combined with a small sample size can seriously affect the estimation of the dispersion parameter, no matter which estimator is used within the estimation process. The probability the dispersion parameter becomes unreliably estimated increases significantly as the sample mean and sample size decrease. Consequently, the results show that an unreliably estimated dispersion parameter can significantly undermine empirical Bayes (EB) estimates as well as the estimation of confidence intervals for the gamma mean and predicted response. The paper ends with recommendations about minimizing the likelihood of producing Poisson-gamma models with an unreliable dispersion parameter for modeling motor vehicle crashes.

Accidents, Traffic↗

A Markov model for sample size calculation and inference in vaccine cost-effectiveness studies.

Sample size and power formulae are derived for designing trials of vaccine cost-effectiveness in healthy working adults. When existing trials of influenza vaccine have presented sample size calculations, these have been based on vaccine effectiveness. A Markov model is introduced to account for non-independence of working days lost. Two effect sizes need to be specified, one for a reduction in the number of episodes of absence and the second for a reduction in the mean duration of episodes. The relative cost of vaccine provision is also incorporated. Applying the resulting formulae to published studies indicates that most were underpowered to detect any financial benefit. Moreover, biased estimates of variance have led to led to incorrect inference.

Adult↗

Sample size and power issues in estimating incremental cost-effectiveness ratios from clinical trials data.

It is becoming increasingly more common for a randomized controlled trial of a new therapy to include a prospective economic evaluation. The advantage of such trial-based cost-effectiveness is that conventional principles of statistical inference can be used to quantify uncertainty in the estimate of the incremental cost-effectiveness ratio (ICER). Numerous articles in the recent literature have outlined and compared various approaches for determining confidence intervals for the ICER. In this paper we address the issue of power and sample size in trial-based cost-effectiveness analysis. Our approach is to determine the required sample size to ensure that the resulting confidence interval is narrow enough to distinguish between two regions in the cost-effectiveness plane: one in which the new therapy is considered to be cost-effective and one in which it is not. As a result, for a given sample size, the cost-effectiveness plane is divided into two regions, separated by an ellipse centred at the origin, such that the sample size is adequate only if the truth lies on or outside the ellipse.

Antineoplastic Agents↗

Effect of sample size on reproducibility of behavioral teratological study results: a computer simulation experiment using data from the Collaborative Behavioral Teratology Study of the National Center for Toxicological Research.

A computer simulation experiment which attempted to examine the effect of sample size on reproducibility of the effect of treatment was performed on the basis of actual data obtained from the Collaborative Behavioral Teratology Study of the National Center for Toxicological Research. The degree of the treatment effect was assessed in terms of the strength of the association (eta square). The results indicate that sample size has a large effect on the reproducibility of results which are assessed with the magnitude of SD for eta squares obtained from replication experiments. Suitable sample sizes to obtain relatively consistent results across studies were discussed, pointing out that not enough attention has been paid to the effect of sample size in the issue of reproducibility of results in some behavioral teratology studies.

Computer Simulation↗

Exact sample sizes needed to detect dependence in 2 x 3 tables.

Many medical and biological studies entail classifying a number of observations according to two factors, where one has two and the other three possible categories. This is the case of, for example, genetic association studies of complex traits with single-nucleotide polymorphisms (SNPs), where the a priori statistical planning, analysis, and interpretation of results are of critical importance. Here, we present methodology to determine the minimum sample size required to detect dependence in 2 x 3 tables based on Fisher's exact test, assuming that neither of the two margins is fixed and only the grand total N is known in advance. We provide the numerical tools necessary to determine these sample sizes for desired power, significance level, and effect size, where only the computational time can be a limitation for extreme parameter values. These programs can be accessed at . This solution of the sample size problem for an exact test will permit experimentalists to plan efficient sampling designs, determine the extent of statistical support for their hypotheses, and gain insight into the repeatability of their results. We apply this solution to the sample size problem to three empirical studies, and discuss the results with specified power and nominal significance levels.

DNA, Mitochondrial↗

Use of randomised controlled trials for producing cost-effectiveness evidence: potential impact of design choices on sample size and study duration.

BACKGROUND: A number of approaches to conducting economic evaluations could be adopted. However, some decision makers have a preference for wholly stochastic cost-effectiveness analyses, particularly if the sampled data are derived from randomised controlled trials (RCTs). Formal requirements for cost-effectiveness evidence have heightened concerns in the pharmaceutical industry that development costs and times might be increased if formal requirements increase the number, duration or costs of RCTs. Whether this proves to be the case or not will depend upon the timing, nature and extent of the cost-effectiveness evidence required. OBJECTIVE: To illustrate how different requirements for wholly stochastic cost-effectiveness evidence could have a significant impact on two of the major determinants of new drug development costs and times, namely RCT sample size and study duration. DESIGN: Using data collected prospectively in a clinical evaluation, sample sizes were calculated for a number of hypothetical cost-effectiveness study design scenarios. The results were compared with a baseline clinical trial design. RESULTS: The sample sizes required for the cost-effectiveness study scenarios were mostly larger than those for the baseline clinical trial design. Circumstances can be such that a wholly stochastic cost-effectiveness analysis might not be a practical proposition even though its clinical counterpart is. In such situations, alternative research methodologies would be required. For wholly stochastic cost-effectiveness analyses, the importance of prior specification of the different components of study design is emphasised. However, it is doubtful whether all the information necessary for doing this will typically be available when product registration trials are being designed. CONCLUSIONS: Formal requirements for wholly stochastic cost-effectiveness evidence based on the standard frequentist paradigm have the potential to increase the size, duration and number of RCTs significantly and hence the costs and timelines associated with new product development. Moreover, it is possible to envisage situations where such an approach would be impossible to adopt. Clearly, further research is required into the issue of how to appraise the economic consequences of alternative economic evaluation research strategies.

Cost-Benefit Analysis↗

Power and sample size computations in simultaneous tests for non-inferiority based on relative margins.

In this paper, we address the problem of calculating power and sample sizes associated with simultaneous tests for non-inferiority. We consider the case of comparing several experimental treatments with an active control. The approach is based on the ratio view, where the common non-inferiority margin is chosen to be some percentage of the mean of the control treatment. Two power definitions in multiple hypothesis testing, namely, complete power and minimal power, are used in the computations. The sample sizes associated with the ratio-based inference are also compared with that of a comparable inference based on the difference of means for various scenarios. It is found that the sample size required for ratio-based inferences is smaller than that of difference-based inferences when the relative non-inferiority margin is less than one and when large response values indicate better treatment effects. The results are illustrated with examples.

Data Interpretation, Statistical↗

Transcervical CVS sample size: correlation with placental location, cytogenetic findings, and pregnancy outcome.

Data from 2907 transcervical CVS cases performed on singleton pregnancies were reviewed retrospectively and villus sample size was correlated with cytogenetic results, placental location, maternal age at the expected date of confinement (EDC), gestational age at the time of sampling, birth weight, gestational age at the time of delivery, and pregnancy outcome. No correlation was noted between villus sample size and maternal age, gestational age at sampling, gestational age at delivery, birth weight, or pregnancy outcome. An inverse correlation between villus sample size and percentage of abnormal cytogenetic findings was statistically significant (chi 2 = 8.53, p < 0.01). The percentage of small samples was greater when the placenta was anterior, lateral, or fundal than when the placenta was posterior.

Adult↗