Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Sample size determination for a t test given a t value from a previous study: A FORTRAN 77 program.

When uncertain about the magnitude of an effect, researchers commonly substitute in the standard sample-size-determination formula an estimate of effect size derived from a previous experiment. A problem with this approach is that the traditional sample-size-determination formula was not designed to deal with the uncertainty inherent in an effect-size estimate. Consequently, estimate-substitution in the traditional sample-size-determination formula can lead to a substantial loss of power. A method of sample-size determination designed to handle uncertainty in effect-size estimates is described. The procedure uses the t value and sample size from a previous study, which might be a pilot study or a related study in the same area, to establish a distribution of probable effect sizes. The sample size to be employed in the new study is that which supplies an expected power of the desired amount over the distribution of probable effect sizes. A FORTRAN 77 program is presented that permits swift calculation of sample size for a variety of t tests, including independent t tests, related t tests, t tests of correlation coefficients, and t tests of multiple regression b coefficients.

Data Interpretation, Statistical↗

Sample size for beginners.

The common failure to include an estimation of sample size in grant proposals imposes a major handicap on applicants, particularly for those proposing work in any aspect of research in the health services. Members of research committees need evidence that a study is of adequate size for there to be a reasonable chance of a clear answer at the end. A simple illustrated explanation of the concepts in determining sample size should encourage the faint hearted to pay more attention to this increasingly important aspect of grantsmanship.

Analysis of Variance↗

Estimating effective population size from samples of sequences: a bootstrap Monte Carlo integration method.

We would like to use maximum likelihood to estimate parameters such as the effective population size N(e) or, if we do not know mutation rates, the product 4N(e) mu of mutation rate per site and effective population size. To compute the likelihood for a sample of unrecombined nucleotide sequences taken from a random-mating population it is necessary to sum over all genealogies that could have led to the sequences, computing for each one the probability that it would have yielded the sequences, and weighting each one by its prior probability. The genealogies vary in tree topology and in branch lengths. Although the likelihood and the prior are straightforward to compute, the summation over all genealogies seems at first sight hopelessly difficult. This paper reports that it is possible to carry out a Monte Carlo integration to evaluate the likelihoods approximately. The method uses bootstrap sampling of sites to create data sets for each of which a maximum likelihood tree is estimated. The resulting trees are assumed to be sampled from a distribution whose height is proportional to the likelihood surface for the full data. That it will be so is dependent on a theorem which is not proven, but seems likely to be true if the sequences are not short. One can use the resulting estimated likelihood curve to make a maximum likelihood estimate of the parameter of interest, N(e) or of 4N(e) mu. The method requires at least 100 times the computational effort required for estimation of a phylogeny by maximum likelihood, but is practical on today's work stations. The method does not at present have any way of dealing with recombination.

Base Sequence↗

Estimating population size and drag sampling efficiency for the blacklegged tick (Acari: Ixodidae).

Estimates of absolute density were determined over a 5-yr period (1990-1994) for a population of Ixodes scapularis Say located in Westchester County, NY, by mark-release-recapture (nymphs and adults) and removal (larvae) methods. Density estimates for larvae ranged from 5.2 to 16.5/m2 and averaged 11.5/m2. Values for nymphs varied as much as fourfold among successive years, ranging from 0.5 to 2.3/m2 and averaging 1.2/m2, whereas adult density ranged from 0.3 to 0.4/m2, averaging 0.33/m2. Natural mortality of nymphs and adults was measured in experimental cages during population estimation periods, and indicated that survival declined linearly over the short-term and did not significantly influence estimates. Drag sampling efficiency, the proportion of the estimated population obtained in a single sample, averaged 6.3% among all stages. Efficiency was not significantly different among stages and was independent of tick density within a given life stage. The population estimation techniques employed in this study are well suited for use with I. scapularis and can provide data that offer insights into mortality patterns in individual populations.

Animals↗

On power and sample size for studying features of the relative odds of disease.

Estimates of sample size and statistical power are essential ingredients in the design of epidemiologic studies. Once an association between disease and exposure has been demonstrated, additional studies are often needed to investigate special features of the relation between exposure, other covariates, and risk of disease. The authors present a general formulation to compute sample size and power for case-control and cohort studies to investigate more complex patterns in the odds ratios, such as to distinguish between two different slopes of linear trend, to distinguish between two possible dose-response relations, or to distinguish different models for the joint effects of two important exposures or of one exposure factor adjusting for another. Such special studies of exposure-response relations may help investigators to distinguish between plausible biologic models and may lead to more realistic models for calculating attributable risk and lifetime disease risk. The sample size formulae are applied to studies of indoor radon exposure and lung cancer and suggest that epidemiologic studies may not be feasible for addressing some issues. For example, if the risk estimates from underground miners' studies are, in truth, not applicable to home exposures and overestimate the gradient of risk from home exposure to radon by, for example, a factor of 2, then enormously large numbers of subjects would be required to detect the difference. Furthermore, if the true interaction between smoking and radon exposure is less than multiplicative, only the largest investigations will have sufficient power to reject additivity. For the simple case of testing for no exposure effect, when exposure is either dichotomous or continuous, these methods yield well-known formulae.

Aged↗

Relationship between case-control studies and the transmission/disequilibrium test.

Case-control studies provide a powerful approach for detecting disease susceptibility loci that have only a weak to moderate impact on the risk of disease, or markers that are in linkage disequilibrium with such loci. However, since any association detected in a case-control study may result from uncontrolled confounding, evidence for disease-marker associations obtained from such studies must be confirmed by alternative methods. Since studies that use the transmission/disequilibrium test or TDT are frequently employed to confirm disease-marker associations detected in case-control studies, data are increasingly available from both case-control studies and "TDT studies" of the same disease-marker association. It would, therefore, be useful to have a single measure of the magnitude of the disease-marker association that would allow for comparison of results from these two study designs. Such a measure could also be used to estimate minimum sample size requirements for TDT studies of previously reported disease-marker associations. An obvious measure of the disease-marker association in TDT studies is the frequency (T) with which heterozygous parents transmit the putative, high-risk marker allele to affected offspring. In this paper, it is shown that T can also be estimated from case-control data with a minimum of assumptions, and that T is the critical parameter for determining power and estimating sample sizes for the TDT.

Case-Control Studies↗

Variance and confidence limits in validation studies based on comparison between three different types of measurements.

BACKGROUND: The methods used in epidemiological studies to assess exposure are often affected by a conspicuous amount of measurement error. Exposure-measurement error is recognised to cause attenuation in the association between exposure and disease. Among different possible approaches, the validity coefficient of a measurement can be estimated by a comparison of three types of measurements, using either structural equation models or factor analysis (the triads method). These approaches assume that the measurements are linearly related to true intake and have independent random errors. METHODS: In this paper we present an estimator of the variance of the estimated validity coefficient to compute the associated confidence intervals. Standard error for the validity coefficient allows the efficiency of validation studies to be evaluated. Our work was motivated by the fact that existing software does not provide correct standard errors for the estimated validity coefficient. The approach is illustrated using selected examples from dietary validation studies. RESULTS: The accuracy of our formula is evaluated by comparison with the results of a simulation study, which shows that our variance estimator provides good results for sample sizes of at least n = 100 and when the expected value of the validity coefficient is not too close to 1.0, independent of the sample size. Our estimator formula performs better than either a naïve approach, that computes the standard error for a validity coefficient as if it is a straightforward correlation coefficient, or the SAS-CALIS procedure, which uses a maximum likelihood method. CONCLUSIONS: In evaluating the validity of the type of measurement chosen to assess exposure in an epidemiological study, it is important to provide an estimate of the precision of the validity coefficient of the measurement. Our variance estimator may help calculate sample size requirements for validation studies.

Analysis of Variance↗

Confidence interval estimation of a rate and the choice of sample size.

The problem of estimating a rate or proportion is considered. Four methods for constructing an approximate confidence interval are discussed and compared via a simulation study. The most accurate method is found. Also, for each method a sharp upper bound (dependent only on the sample size) is given for the length of the confidence interval. By choosing an appropriate sample size this bound enables the practitioner to achieve a prespecified maximum length for the confidence interval without knowing the population rate. The striking result is that the most accurate method has the smallest bound, thus requiring the least sample units.

Computer Simulation↗

Measures of reliability in sports medicine and science.

Reliability refers to the reproducibility of values of a test, assay or other measurement in repeated trials on the same individuals. Better reliability implies better precision of single measurements and better tracking of changes in measurements in research or practical settings. The main measures of reliability are within-subject random variation, systematic change in the mean, and retest correlation. A simple, adaptable form of within-subject variation is the typical (standard) error of measurement: the standard deviation of an individual's repeated measurements. For many measurements in sports medicine and science, the typical error is best expressed as a coefficient of variation (percentage of the mean). A biased, more limited form of within-subject variation is the limits of agreement: the 95% likely range of change of an individual's measurements between 2 trials. Systematic changes in the mean of a measure between consecutive trials represent such effects as learning, motivation or fatigue; these changes need to be eliminated from estimates of within-subject variation. Retest correlation is difficult to interpret, mainly because its value is sensitive to the heterogeneity of the sample of participants. Uses of reliability include decision-making when monitoring individuals, comparison of tests or equipment, estimation of sample size in experiments and estimation of the magnitude of individual differences in the response to a treatment. Reasonable precision for estimates of reliability requires approximately 50 study participants and at least 3 trials. Studies aimed at assessing variation in reliability between tests or equipment require complex designs and analyses that researchers seldom perform correctly. A wider understanding of reliability and adoption of the typical error as the standard measure of reliability would improve the assessment of tests and equipment in our disciplines.

Algorithms↗

Adaptive statistical analysis following sample size modification based on interim review of effect size.

In designing a comparative clinical trial, the required sample size is a function of the effect size, the value of which is unknown and at best may be estimated from historical data. Insufficiency in sample size as a result of overestimating the effect size can be destructive to the success of the clinical trial. Sample size re-estimation may need to be properly considered as a part of clinical trial planning. This paper is intended to give the motivations for the sample size re-estimation based partly on the effect size observed at an interim analysis and for a resulting simple adaptive test strategy. The performance of this adaptive design strategy is assessed by comparing it with a fixed maximum sample size design that is properly adjusted in anticipation of the possible sample size adjustment.

Algorithms↗

Adapting the sample size planning of a phase III trial based on phase II data.

Traditionally, in clinical development plan, phase II trials are relatively small and can be expected to result in a large degree of uncertainty in the estimates based on which Phase III trials are planned. Phase II trials are also to explore appropriate primary efficacy endpoint(s) or patient populations. When the biology of the disease and pathogenesis of disease progression are well understood, the phase II and phase III studies may be performed in the same patient population with the same primary endpoint, e.g. efficacy measured by HbA1c in non-insulin dependent diabetes mellitus trials with treatment duration of at least three months. In the disease areas that molecular pathways are not well established or the clinical outcome endpoint may not be observed in a short-term study, e.g. mortality in cancer or AIDS trials, the treatment effect may be postulated through use of intermediate surrogate endpoint in phase II trials. However, in many cases, we generally explore the appropriate clinical endpoint in the phase II trials. An important question is how much of the effect observed in the surrogate endpoint in the phase II study can be translated into the clinical effect in the phase III trial. Another question is how much of the uncertainty remains in phase III trials. In this work, we study the utility of adaptation by design (not by statistical test) in the sense of adapting the phase II information for planning the phase III trials. That is, we investigate the impact of using various phase II effect size estimates on the sample size planning for phase III trials. In general, if the point estimate of the phase II trial is used for planning, it is advisable to size the phase III trial by choosing a smaller alpha level or a higher power level. The adaptation via using the lower limit of the one standard deviation confidence interval from the phase II trial appears to be a reasonable choice since it balances well between the empirical power of the launched trials and the proportion of trials not launched if a threshold lower than the true effect size of phase III trial can be chosen for determining whether the phase III trial is to be launched.

Clinical Trials, Phase II as Topic↗

Unbiased measures of transmitted information and channel capacity from multivariate neuronal data.

Two measures from information theory, transmitted information and channel capacity, can quantify the ability of neurons to convey stimulus-dependent information. These measures are calculated using probability functions estimated from stimulus-response data. However, these estimates are biased by response quantization, noise, and small sample sizes. Improved estimators are developed in this paper that depend on both an estimate of the sample-size bias and the noise in the data.

Animals↗

Sample size calculations based on generalized estimating equations for population pharmacokinetic experiments.

We present a method for calculating the sample size of a pharmacokinetic study analyzed using a mixed effects model within a hypothesis testing framework. A sample size calculation method for repeated measurement data analyzed using generalized estimating equations has been modified for nonlinear models. The Wald test is used for hypothesis testing of pharmacokinetic parameters. A marginal model for the population pharmacokinetic is obtained by linearizing the structural model around the subject specific random effects. The proposed method is general in that it allows unequal allocation of subjects to the groups and accounts for situations where different blood sampling schedules are required in different groups of patients. The proposed method has been assessed using Monte Carlo simulations under a range of scenarios. NONMEM was used for simulations and data analysis and the results showed good agreement.

Computer Simulation↗

Sample size requirements for precise estimates of reliability, generalizability, and validity coefficients.

Precision of the reliability coefficient (r) is investigated. The width of the confidence interval for r as a function of sample size (N) is shown for retest, alternate-form, split-half, alpha, intraclass, interrater, and validity coefficients. Although the determination of the N needed for reliability studies is somewhat subjective, a minimum of 400 subjects is recommended. Much larger Ns may be needed for validity studies. A survey of published reliability studies shows that 59% of the sample sizes were less than 100. Confidence intervals for obtained test scores are used as a practical application measure that also leads to the conclusion of a minimum of 400 subjects.

Bias↗

The methods for handling missing data in clinical trials influence sample size requirements.

OBJECTIVE: Results of studies estimating osteoarthritis progression may be affected by missing values. In clinical trials assessing disease-modifying osteoarthritis drugs, sample sizes should be calculated using close estimates of outcome variables. STUDY DESIGN AND SETTING: Supposing a two-parallel group design in hip osteoarthritis clinical trials, we estimated sample sizes using the joint space width (JSW), number of patients with JSW progression >0.5 mm (JSN), time to total hip arthroplasty (THA), and time to JSN or THA using several approaches to deal with missing data. RESULTS: Three-year clinical trials testing a treatment effect of 50%, with a power of 80%, could require sample sizes of 121 patients for JSW, 57 for JS progression using multiple imputation for handling missing values; 200 for THA; and 47 for JSN or THA. These numbers vary greatly depending on the approach chosen for handling missing data. CONCLUSIONS: These results can help investigators plan clinical trials to select the primary outcome and a priori specify the way missing data will be handled.

Aged↗

Sample size and power for case-control studies when exposures are continuous.

In estimating the sample size for a case-control study, epidemiologic texts present formulae that require a binary exposure of interest. Frequently, however, important exposures are continuous and dichotomization may result in a 'not exposed' category that has little practical meaning. In addition, if risks vary monotonically with exposure, then dichotomization will obscure risk effects and require a greater number of subjects to detect differences in the exposure distributions among cases and controls. Starting from the usual score statistic to detect differences in exposure, this paper develops sample size formulae for case-control studies with arbitrary exposure distributions; this includes both continuous and dichotomous exposure measurements as special cases. The score statistic is appropriate for general differentiable models for the relative odds, and, in particular, for the two forms commonly used in prospective disease occurrence models: (1) the odds of disease increase linearly with exposure; or (2) the odds increase exponentially with exposure. Under these two models we illustrate calculation of sample sizes for a hypothetical case-control study of lung cancer among non-smokers who are exposed to radon decay products at home.

Environmental Exposure↗

Estimated coefficient of variation values for sample size planning in bioequivalence studies.

OBJECTIVE: The aim of the present communication is to provide information regarding the intrasubject coefficent of variation obtained from 30 bioequivalence studies covering 16 drugs which can be used for estimation of sample size. Additionally, an attempt was also made to estimate the test power of each of the studies conducted. METHODS: The intrasubject coefficient of variation was estimated from the residual mean square error obtained from analysis of variance of the parameters AUC0-infinity, Cmax and Cmax/AUC0-infinity after logarithmic transformation. The test power in the analyses of the above parameters was subsequently estimated using nomograms provided by Diletti et al. [1991]. RESULTS AND CONCLUSION: Thirty products covering 16 drugs were studied in which 22 were immediate-release (including one dispersible tablet) and 8 were sustained-release formulations. The intrasubject coefficient of variation for the parameter AUC0-infinity was smaller than Cmax, and hence considerably more studies were able to attain a power of greater than 80% using 12 volunteers for the AUC0-infinity, compared to the Cmax. However, the variability in the Cmax could be reduced by using the parameter Cmax/ AUC0-infinity, and thus, provide a more realistic estimation of sample size, since the latter reflects only the rate of absorption and not both the rate and extent as in the case of Cmax [Endrenyi et al. 1991].

Absorption↗

[Practical aspects regarding sample size in clinical research].

BACKGROUND: The knowledge of the right sample size let us to be sure if the published results in medical papers had a suitable design and a proper conclusion according to the statistics analysis. To estimate the sample size we must consider the type I error, type II error, variance, the size of the effect, significance and power of the test. To decide what kind of mathematics formula will be used, we must define what kind of study we have, it means if its a prevalence study, a means values one or a comparative one. In this paper we explain some basic topics of statistics and we describe four simple samples of estimation of sample size.

Sample Size↗