Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Choosing the number of controls in a matched case-control study, some sample size, power and efficiency considerations.

This paper investigates the efficiency of using multiple controls in a case-control study, when there is a single binary exposure variable. Specifically, we consider the asymptotic power of the Cochran test statistic against non-local alternatives of interest. When it is desirable to take multiple controls per case, we show that the marginal return rapidly diminishes as the number of controls per case increases. The effect is as strong, if not stronger, for non-local alternatives as it is for local alternatives. Hence, it is rarely worth choosing more than three controls per case. We also provide a table of sample sizes necessary to achieve 80 per cent power for some odds ratios not equal to one. We extend the results to a special case when there are two binary exposure variables.

Biometry↗

Sample size calculations for classical association and TDT-type methods using family data.

Transmission Disequilibrium Test (TDT)-based methods have been advocated by several authors for testing that a marker-phenotype association is actually due to linkage and not to uncontrolled stratification. As a pre-requisite of TDT-type methods is the presence of an association between marker and phenotype, one may wish to first investigate the association using a classical association study, and then to check by a TDT approach whether this association is actually due to linkage. We propose an estimating equation (EE) procedure, to compute analytically the minimum sample size of sibship data required to detect the association between a marker and a quantitative phenotype, and that required to confirm it by two TDT methods. We show that, when the marker allele frequency is low or high, the number of informative sibs needed in TDT-type methods can be lower than the number required in an association analysis, and even more so when the familial clustering is strong. However, in all cases, the number of sibs that need to be sampled to get the appropriate number of informative sibs for analysis is always larger for TDT methods than for an association study. In a phenotype-first strategy, this number may be critical when investigating costly phenotypes.

Alleles↗

A Monte Carlo investigation of homogeneity tests of the odds ratio under various sample size configurations.

Epidemiologic data for case-control studies are often summarized into K 2 x 2 tables. Given a fixed number of cases and controls, the degree of sparseness in the data depends on the number of strata, K. The effect of increasing stratification on size and power of seven tests of homogeneity of the odds ratio is studied using Monte Carlo methods. In all the designs considered here, the numbers of cases and controls per stratum are the same. Considering both size and power in non-sparse-data settings, we recommend the Breslow-Day statistic (1980, Statistical Methods in Cancer Research, 1. The Analysis of Case-Control Studies, p. 142; Lyon: International Agency for Research on Cancer) for general use. In sparse-data settings the T4 statistic of Liang and Self (1985, Biometrika 72, 353-358) performs the best when all tables, regardless of sample size, have odds ratios generated from the same distribution. In sparse-data settings characterized by a large table with an odds ratio of 1 and many small tables with odds ratios greater than 1, the T5 statistic of Liang and Self (1985) performs the best. One of the most important results of this study is the generally low power for all homogeneity tests especially when the data are sparse.

Computer Simulation↗

Data analysis and sample size issues in evaluations of community-based health promotion and disease prevention programs: a mixed-model analysis of variance approach.

The growing interest in community-based approaches to health promotion and disease prevention (HP/DP) has been accompanied by a growing need to evaluate the effectiveness of such programs. Special issues that arise in these evaluation studies include (1) entire communities are assigned to intervention and control groups, (2) only a small number of communities can usually be studied, (3) the time course of changes in behavior and other outcomes is often of interest, and (4) surveys to measure such changes over time can be conducted with either repeated cross-sectional samples or with longitudinal samples. This paper shows how these issues can be addressed under a mixed-model analysis of variance approach. This approach serves to unify several ideas in the literature on evaluation of community studies, including use of time-series regression and the question of whether the individual or the community should be the unit of analysis. We also describe how the method can be used to estimate sample size requirements, statistical power, or minimum detectable program effect.

Analysis of Variance↗

Power and sample size requirements for two-part models.

Two-part models assume the data has a probability mass at zero and a continuous response for values greater than 0. We construct tests based on the proportion of zeros and the difference among the positive values of a response giving rise to a two degree of freedom chi(2) test. This note gives the non-centrality parameter and finds the power for the test. The power is compared to the simulation results in a recent paper. We derive sample size calculations. We note that some modifications of this procedure may provide better tests if the non-zero responses are discrete. Published in 2001 by John Wiley & Sons, Ltd.

Data Interpretation, Statistical↗

Comparison of the proportions of affected relatives of cases and controls: analysis and minimum sample size formula.

The problem of familial aggregation has frequently been approached by comparing the proportion of affected relatives of index cases with the proportion of affected relatives of index controls. This type of study has an analogy with the case-control design and is frequently analysed as if it was one. It is however an essentially different design which may involve dependent observations. We show that conventional tests for comparing proportions of affected subjects are still valid with this design; we give a formula for computing the minimum sample size required to detect a given degree of aggregation measured by the intracluster correlation coefficient, which can be estimated very simply by the difference in proportions of affected relatives in the two groups.

Case-Control Studies↗

An integrated population-averaged approach to the design, analysis and sample size determination of cluster-unit trials.

While the mixed model approach to cluster randomization trials is relatively well developed, there has been less attention given to the design and analysis of population-averaged models for randomized and non-randomized cluster trials. We provide novel implementations of familiar methods to meet these needs. A design strategy that selects matching control communities based upon propensity scores, a statistical analysis plan for dichotomous outcomes based upon generalized estimating equations (GEE) with a design-based working correlation matrix, and new sample size formulae are applied to a large non-randomized study to reduce underage drinking. The statistical power calculations, based upon Wald tests for summary statistics, are special cases of a general power method for GEE.

Adolescent↗

Sample sizes for the exact test of 'no interaction' in 2 X 2 X 2 tables.

The exact two-sided test of the hypothesis of 'no interaction' in a 2 X 2 X 2 table with fixed totals of rows X columns is considered. A method of approximating the power of the test, based on the limiting conditional distribution of a cell entry (Godambe and Harkness, 1975, Communications in Statistics 4, 699-709), is introduced and found to provide a close fit. This method can be utilized to calculate the sample sizes required to detect interaction effects with a pregiven power.

Breast Neoplasms↗

Confidence intervals and sample size calculations to compare variant frequencies.

Direct mutagenicity tests offer the opportunity of monitoring human populations to detect evidence of genetic damage that occurs in vivo. As such these tests offer the potential of linking earlier exposures to mutagenic agents to subsequent health effects. One such test detects mutant T-lymphocytes that arise in vivo in human peripheral blood. Statistical analysis of the value observed, the variant frequency (Vf), is the subject of this paper. We present and illustrate procedures for finding confidence intervals for a single variant frequency and for the ratio of two independent variant frequencies. We also derive formulae for required sample size to ensure adequate power to detect a specified change in a single variant frequency or a specified difference between two variant frequencies. Our approach is to exploit the approximate normality of the natural logarithm of a Poisson distributed variable. The procedures developed appear to be quite accurate even for the small (5-15) values often observed in the variant frequency assay. Moreover, the procedures are very easy to use and should prove valuable to investigators involved in direct mutagenicity testing.

Drug Resistance↗

Choice of sample size in parasitological experiments.

One of the most awkward practical problems confronting a research worker setting up an experiment is to decide how big the experiment should be. The larger the number of subjects then the smaller is the degree of uncertainty in the conclusions; but large trials are costly and may have to be of long duration, so there is every incentive to keep trials as small as possible. Statistics is the science of coping; with uncertainty, and in this article Michael Healy explains the statistical basis that helps to decide how big a sample size should be - how to balance minimum error rates with optimal cost efficiency.

Journal Article↗

Performance estimation of diagnostic tests for cervical precancer based on fluorescence spectroscopy: effects of tissue type, sample size, population, and signal-to-noise ratio.

Fluorescence spectroscopy may provide a cost-effective tool to improve precancer detection. We describe a method to estimate the diagnostic performance of classifiers based on optical spectra, and to explore the sensitivity of these estimations to factors affecting spectrometer cost. Fluorescence spectra were obtained at three excitation wavelengths in 92 patients with an abnormal Papanicolaou smear and 51 patients with no history of an abnormal smear. Bayesian classification rules were developed and evaluated at multiple misclassification costs. We explored the sensitivity of classifier performance to variations in tissue type, sample size, tested population, signal to noise ratio (SNR), and number of excitation and emission wavelengths. Sensitivity and specificity could be evaluated within +/- 7%. Minimal decrease in diagnostic performance is observed as SNR is reduced to 15, the number of excitation-emission wavelength combinations is reduced to 15 or the number of excitation wavelengths is reduced to one. Diagnostic performance is compromised when ultraviolet excitation is not included. Significant spectrometer cost reduction is possible without compromising diagnostic ability. Decision-analytic methods can be used to rate designs based on incremental cost-effectiveness.

Algorithms↗

Suction applied to a muscle biopsy maximizes sample size.

A method for increasing the size of a percutaneous needle biopsy specimen of skeletal muscle is described. Suction (700 TORR) is applied to the inner bore of the biopsy needle after the needle has been inserted into the subject's muscle. The suction pulls the surrounding muscle tissue into the needle, thus insuring the taking of a larger piece (X = 78.5 mg). In most cases, this technique will eliminate the need for repeated biopsies because of inadequate muscle sample size and enhance the validity of subsequent analysis procedures.

Biopsy, Needle↗

Sample size requirements for calibration studies of dietary intake measurements in prospective cohort investigations.

Pooling data from multiple cohort studies on diet and cancer has the advantage that it allows detailed cross-validation of relative risk estimates between different study populations. If there is reasonable agreement between cohort-specific relative risk estimates, a more powerful pooled summary estimate can be obtained. A complication, however, is that in different cohorts, relative risk estimates may be biased to a different degree as a result of errors in the baseline assessments of habitual dietary intake levels. Such divergent biases can be adjusted for by means of "calibration" studies, using standardized reference measurements obtained in a subgroup of each cohort. These adjustments entail a cost, however, in terms of an increase in the confidence interval of relative risk estimates within each cohort separately. In this paper, the authors evaluate the possible magnitude of such intracohort losses in precision and discuss the approximate sample size required to have a sufficient level of accuracy in dietary calibration studies to adjust for bias.

Calibration↗

Sample size calculation for clinical trials in which entry criteria and outcomes are counts of events. ACIP Investigators. Asymptomatic Cardiac Ischemia Pilot.

In many chronic diseases, therapy aims to prevent or reduce the frequency of episodes of a disease manifestation, for example cardiac ischaemic episodes or epileptic seizures. Entry criteria for clinical trials typically include a minimum number of episodes within a baseline period, and regression to the mean should be anticipated. The distribution of the number of episodes at follow-up, the statistical power for treatment comparisons, and the difficulty of recruitment will depend on the entry criterion chosen. A gamma-Poisson mixture model is employed to describe the regression to the mean when the entry criterion and outcome measure are counts of discrete events. Sample size formulae which take account of the entry criterion are derived for comparison of mean number of events at follow-up and the proportion of patients with zero events at follow-up. Application of these formulae to screening data from the Asymptomatic Cardiac Ischemia Pilot (ACIP) Study is presented as an example.

Clinical Trials as Topic↗

Sample size calculations for Bayesian prediction of bovine viral-diarrhoea-virus infection in beef herds.

We used a Bayesian classification approach to predict the bovine viral-diarrhoea-virus infection status of a herd when the prevalence of persistently infected animals in such herds is very small (e.g. <1%). An example of the approach is presented using data on beef herds in Wyoming, USA. The approach uses past covariate information (serum-neutralization titres collected on animals in 16 herds) within a predictive model for classification of a future observable herd. Simulations to estimate misclassification probabilities for different misclassification costs and prevalences of infected herds can be used as a guide to the sample size needed for classification of a future herd.

Animals↗

Sample size determination for bioequivalence assessment by means of confidence intervals.

The statistical analysis of bioequivalence assessment has been consolidated in recent years through the work of Schuirmann [1987], Westlake [1988] and Hauschke et al. [1990], and this has been reflected in the CPMP Note for Guidance on Bioavailability and Bioequivalence and in the joint recommendations of the APV (International Association for Pharmaceutical Technology) and ZL (Central Laboratories of German Pharmacists) during a recent workshop in support of EC-Guidelines [Blume et al. 1990]. Since the decision procedure based on the inclusion of the shortest 90%-confidence interval in the bioequivalence range is the procedure of choice, and as this is equivalent to the two one-sided tests procedure, the sample size determination is based on the power of the latter. Following the approach of Phillips [1990] for the additive model, corresponding nomograms for the more relevant multiplicative model are given in this paper for various ratios of the expected means for test and reference and various coefficients of variation.

Confidence Intervals↗

Sample size determination for bioequivalence assessment by means of confidence intervals.

The statistical analysis of bioequivalence assessment has been consolidated in recent years through the work of Schuirmann [1987], Westlake [1988] and Hauschke et al. [1990], and this has been reflected in the CPMP Note for Guidance on Bioavailability and Bioequivalence and in the joint recommendations of the APV (International Association for Pharmaceutical Technology) and ZL (Central Laboratories of German Pharmacists) during a recent workshop in support of EC-Guidelines [Blume et al. 1990]. Since the decision procedure based on the inclusion of the shortest 90%-confidence interval in the bioequivalence range is the procedure of choice, and as this is equivalent to the two one-sided tests procedure, the sample size determination is based on the power of the latter. Following the approach of Phillips [1990] for the additive model, corresponding nomograms for the more relevant multiplicative model are given in this paper for various ratios of the expected means for test and reference and various coefficients of variation.

Humans↗