Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Power and sample size for survival analysis under the Weibull distribution when the whole lifespan is of interest.

Accessible and readily utilized software, tables and approximation formulae have been developed to estimate power and sample size for studies of time to event (survival times) when the survival times are assumed to be exponential. These methods can markedly misestimate power when the distribution is Weibull and not exponential. The Weibull distribution with increasing hazard is common in aging research, especially when the whole life span of the subjects is of interest. This note considers an extension of power and sample size calculations, previously developed under the exponential distributional assumption, to the more general case of the Weibull distribution for a prospective comparative follow-up study. The hypotheses are defined in terms of the ratio of the median survival times between two groups. It is shown that the power and sample sizes are heavily dependent on the shape parameter of the Weibull distribution. Using the extensions developed, investigators can use existing software and tables to calculate power and sample size under the assumption of a Weibull distribution.

Algorithms↗

Spreadsheet method for determining sample sizes for heart valve studies.

Grunkemeier, Johnson, and Naftel have given a method for computing sample size requirements for a clinical study of a new heart valve. This paper gives an implementation of the method on a computer spreadsheet. Moreover, it computes the sample size for the most powerful test with exactly the prescribed level of significance and power; all other tests will necessarily have smaller power and will need larger sample sizes. All graphs and tables are produced on the spreadsheet, and no use of special statistical functions is necessary.

Computers↗

Sample size estimation using repeated measurements on biomarkers as outcomes.

The objectives of this paper are to (1) examine methods of using longitudinal data in designing comparative trials and calculating sample sizes or power and (2) show the effect of autocorrelation of repeated measures on the assessment of sample sizes. A statistical model with a simple regression structure for the mean trajectory of the longitudinal data and a two-parameter model for the correlations of within-individual observations given by corr(yt,yt+s) = gamma s theta is used. The methods are illustrated by considering a two-group trial and investigating the effect of different values of the correlation parameters, gamma and theta on the sample size. The results show that taking account of the autocorrelation structure of longitudinal data may lead to more efficient designs. Specifically, the stronger the autocorrelation is, the smaller the sample size that is required.

AIDS Vaccines↗

Optimal sample size for a series of pilot trials of new agents.

A new approach is presented for determining the appropriate sample sizes for a series of screening trials to identify promising new therapeutic agents. The formulation of the problem is motivated by recognition of the fact that screening of new agents is a continuing process. Consequently, it does not seem ideal to fix the overall total sample size, as previous authors have done. Instead we fix the error rates and optimize the individual sample sizes to minimize the time to identify a promising agent, using an empirical Bayes formulation. When applied to data from the large historical experience of exploratory vaccination trials at Memorial Sloan-Kettering Cancer Center, the method demonstrates that relatively small individual screening trials are optimal in this setting. The reliability of the results is evaluated using bootstrapping techniques.

Bayes Theorem↗

Smallest detectable and minimal clinically important differences of rehabilitation intervention with their implications for required sample sizes using WOMAC and SF-36 quality of life measurement instruments in patients with osteoarthritis of the lower extremities.

OBJECTIVE: To discuss the concepts of the minimal clinically important difference (MCID) and the smallest detectable difference (SDD) and to examine their relation to required sample sizes for future studies using concrete data of the condition-specific Western Ontario and McMaster Universities Osteoarthritis Index (WOMAC) and the generic Medical Outcomes Study 36-Item Short Form (SF-36) in patients with osteoarthritis of the lower extremities undergoing a comprehensive inpatient rehabilitation intervention. METHODS: SDD and MCID were determined in a prospective study of 122 patients before a comprehensive inpatient rehabilitation intervention and at the 3-month followup. MCID was assessed by the transition method. Required SDD and sample sizes were determined by applying normal approximation and taking into account the calculation of power. RESULTS: In the WOMAC sections the SDD and MCID ranged from 0.51 to 1.33 points (scale 0 to 10), and in the SF-36 sections the SDD and MCID ranged from 2.0 to 7.8 points (scale 0 to 100). Both questionnaires showed 2 moderately responsive sections that led to required sample sizes of 40 to 325 per treatment arm for a clinical study with unpaired data or total for paired followup data. CONCLUSION: In rehabilitation intervention, effects larger than 12% of baseline score (6% of maximal score) can be attained and detected as MCID by the transition method in both the WOMAC and the SF-36. Effects of this size lead to reasonable sample sizes for future studies lying below n = 300. The same holds true for moderately responsive questionnaire sections with effect sizes higher than 0.25. When designing studies, assumed effects below the MCID may be detectable but are clinically meaningless.

Aged↗

Effects of interrater reliability of psychopathologic assessment on power and sample size calculations in clinical trials.

Although rater training is increasingly used to improve the quality of the investigated outcome parameters, the reliability of assessments is not perfect. Thus, empirical reliability estimates should be used instead of theoretically assumed perfect reliability. Implications of the reliability of psychiatric assessments for sample size and power calculations in clinical trials are presented. The theoretical basis of sample size and power calculations using empirical reliability scores is delineated. Examples from contemporary research on schizophrenia and depression are used to illustrate several implications for study design and interpretation of results. The tremendous impact of the lack of reliability of psychopathologic assessments on sample size, power, and detectable true score differences in clinical trials is shown. The problem of multiple outcome variables with different reliabilities is addressed. Studies lacking power because of unreliable assessments carry the risk of false-negative findings and raise ethical questions. Rater training is strongly recommended to assess and improve interrater reliability whenever necessary and possible before trials are started. Sample size calculations and power analysis should be based on empirical reliability values of outcome parameters as part of quality assurance and cost savings.

Clinical Trials as Topic↗

Mid-course sample size modification in clinical trials based on the observed treatment effect.

It is not uncommon to set the sample size in a clinical trial to attain specified power at a value for the treatment effect deemed likely by the experimenters, even though a smaller treatment effect would still be clinically important. Recent papers have addressed the situation where such a study produces only weak evidence of a positive treatment effect at an interim stage and the organizers wish to modify the design in order to increase the power to detect a smaller treatment effect than originally expected. Raising the power at a small treatment effect usually leads to considerably higher power than was first specified at the original alternative. Several authors have proposed methods which are not based on sufficient statistics of the data after the adaptive redesign of the trial. We discuss these proposals and show in an example how the same objectives can be met while maintaining the sufficiency principle, as long as the eventuality that the treatment effect may be small is considered at the design stage. The group sequential designs we suggest are quite standard in many ways but unusual in that they place emphasis on reducing the expected sample size at a parameter value under which extremely high power is to be achieved. Comparisons of power and expected sample size show that our proposed methods can out-perform L. Fisher's 'variance spending' procedure. Although the flexibility to redesign an experiment in mid-course may be appealing, the cost in terms of the number of observations needed to correct an initial design may be substantial.

Clinical Trials, Phase III as Topic↗

Design for sample size re-estimation with interim data for double-blind clinical trials with binary outcomes.

Estimation of sample size in clinical trials requires knowledge of parameters that involve the treatment effect and variability, which are usually uncertain to medical researchers. The recent release within the European Union of a Note for Guidance from the Commission for Proprietary Medical Products (CPMP) highlights the importance of this issue. Most previous papers considered the case of continuous response variables that assume a normal distribution; some regarded the portion up to the interim stage as an 'internal pilot study' and required unblinding. In this paper, our concern is with the case of binary response variables, which is more difficult than the normal case since the mean and variance are not distinct parameters. We offer a design with a simple stratification strategy that enables us to verify and update the assumption of the response rates given initially in the protocol. The design provides a method to re-estimate the sample size based on interim data while preserving the trial's blinding. An illustrative numerical example and simulation results show slight effect on the type I error rate and the decision making characteristics on sample size adjustment.

Clinical Trials as Topic↗

Minimization of sample size when comparing two small probabilities in a non-inferiority safety trial.

In clinical trials success rates of two treatments to be compared often range from 10 to 90 per cent. When the comparison probabilities are (much) smaller than 10 per cent, standard methods for sample size and power calculations may provide invalid results. This situation may occur when there is interest in safety rather than in efficacy. In such trials, no more patients should be included than strictly necessary. We compared the results of maximum likelihood methods for the computation of sample sizes in a non-inferiority trial, including exact procedures and considered unequal sample sizes for experimental and reference treatment. An exact, unequal sample size maximum likelihood procedure is advocated when the specified non-zero risk difference under the null hypothesis is not too large. Such a procedure is also indicated when the parameter of interest is the relative risk, rather than the risk difference.

Clinical Trials, Phase II as Topic↗

Reference range determination: the problem of small sample sizes.

The process of developing and validating a quantitative test includes determination of a reference range. Traditionally this has been taken as the mean +/- 2 standard deviations for a random sampling from a reference population. However, this method fails to recognize the substantial variability in the sample mean and standard deviation for the small sample sizes frequently encountered in nuclear medicine. A new approach, which involves calculating confidence intervals for the upper and lower bounds of the traditionally defined range, recognizes three ranges of values: normal, indeterminate, and abnormal. The principles of this approach are illustrated using differential renal function in twelve renal transplant donors. The 99mTc-DTPA differential uptake between 1 and 2 min gave a traditionally-defined single-kidney range of 50% +/- 8%, whereas with our method the normal range would be 50% +/- 6% with indeterminate ranges of 37%-44% and 56%-63%. These values are consistent with the wide variation in reference ranges reported in the literature, and suggest that much of this variability may be a statistical artifact resulting from inadequate sample sizes. A nomogram has been derived that permits the power of the reference range determination to be easily calculated from the sample size. Analysis of the effect of sample size on the accuracy of the upper and lower bounds of the reference range is advocated whenever small reference populations are used.

Humans↗

Effects of sample size on the reliability of noise floor and DPOAE.

This study investigated the effects of sample size on the test-retest reliability of the amplitude of distortion product otoacoustic emissions (DPOAE) (2f1-f2) and on the noise floor. Four pairs of primary frequencies (fl and f2) with geometric means of 531, 1000, 2000 and 4000 Hz were presented to 55 normal-hearing women at intensity levels of 35, 45 and 55 dB SPL (L1 = L2). Sample sizes of 12, 25, 50, 100, 200 and 400 sweeps were averaged. The results revealed that sample size, frequency, and intensity had little effect on the standard error of measurement. Thus, the DPOAE data were combined across all conditions and yielded a standard error of measurement of 2.2 dB. To assess whether two DPOAE measurements are statistically significant (e.g. before and after drug administration), the standard error of measurement of the difference between two values was calculated (3.1 dB). Thus, by use of the 95% confidence interval, the difference between two DPOAE is statistically significant if it exceeds approximately 6 dB.

Acoustic Stimulation↗

Computing sample size for receiver operating characteristic studies.

RATIONALE AND OBJECTIVES: Hanley and McNeil (1982) proposed a nonparametric method for computing the standard error of the area under the receiver operating characteristic (ROC) curve. The method has been important in planning the minimum sample size for ROC studies. However, the validity of this method for rating data with various standard deviation ratios has not been investigated. METHODS: A simulation study was conducted to compare the empirical standard error of the area under the curve with Hanley and McNeil's estimate over a range of parameters. An alternative method of computing the standard error based on a binormal distribution is proposed. RESULTS: The method of Hanley and McNeil can lead to underestimation of the minimum sample size. The proposed method provides more appropriate estimates of sample size. CONCLUSIONS: When determining sample size for a study of the area under the ROC curve where rating data are used, the standard error estimator based on the binormal distribution should be used.

Mathematics↗

Two-stage sample size re-estimation based on a nuisance parameter: a review.

Sample size calculations are important and difficult in clinical trails because they depend on the nuisance parameter and treatment effect. Recently, much attention has been focused on two-stage methods whereby the first stage constitutes an internal pilot study used to estimate parameters and revise the final sample size. This paper reviews two-stage methods based on estimation of nuisance parameters in either a continuous or dichotomous outcome setting.

Algorithms↗

The sample size required for intervention studies on fracture prevention can be decreased by using a bone resorption marker in the inclusion criteria: prospective study of a subset of the Nagano Cohort, on behalf of the Adequate Treatment of Osteoporosis (A-TOP) Research Group.

In drug developments for osteoporosis, large-scale and longterm fracture prevention studies have been required. We investigated whether or not it was possible to reduce the sample size and observation period under new selection criteria for an osteoporotic fracture-prevention study. A Poisson regression model was used to identify independent risks for incident vertebral fracture in 515 postmenopausal women who had had no intervention for osteoporosis; this group was a subset of Nagano Cohort participants. The total observation period for this group was 2,577 person-years, and a total of 146 new vertebral fractures were observed. Risk assessment for incident vertebral fracture among numerical covariates revealed that the following items showed significant independent risks for incident fractures; namely, baseline age (hazard ratio [HR]; 1.84; 95% confidence interval (CI), 1.44-2.35; P < 0.001), number of preexisting vertebral fractures (HR, 1.28; 95% CI, 1.17-1.40; P < 0.001), baseline lumbar bone mineral density (LBMD) (HR, 0.79; 95% CI, 0.71-0.88; P < 0.001), and urinary excretion of deoxypyridinoline (DPD) (HR, 1.18; 95% CI, 1.03-1.35; P = 0.016). Because the initial urinary excretion of DPD was found to be a risk for incident vertebral fracture, in addition to the conventional risks, we assessed whether or not the sample size or observation period could be reduced by the incorporation of the urinary excretion of DPD into the selection criteria of a fracture-prevention study. The assessment of sample size was calculated, using the log rank test, at a two-tailed significance level of 5% and with a power of 80%. When osteoporotic patients with preexisting fracture were selected (conventional criteria), the 3-year probability of vertebral fracture was estimated as 14.3% in the present population. On the other hand, the new vertebral fracture rate during 3 years in the osteoporotic patients with preexisting fracture plus high urinary DPD (HR, above 1.0); (new selection criteria) was estimated as 23.2%. When the HR between test drug and placebo was changed from 0.4 to 0.8, the required sample size for any level of HR showed a 40% reduction for the new selection criteria compared to the conventional criteria. Therefore, the addition of urinary DPD level to the selection criteria is useful to reduce sample size in an osteoporosis fracture-prevention study.

Aged↗

The design of case-control studies: the effect of confounding on sample size requirements.

This paper considers the extent to which confounding effects of covariates, which are not controlled for by matching in the design, may influence the sample size necessary for case-control studies. The quantitative calculations are performed for an age-matched case-control study on lung cancer and air pollution, and are based on different evaluation methods. For illustrative purposes attention is confined to a dichotomous risk factor and a single dichotomous covariate. By using the numerical values of a pilot study investigating lung cancer and air pollution, it turns out that the sample size required for detecting a relative risk as close as 1.15 to 1 is only slightly influenced by the strength of the association between confounder and risk factor for reasonable variations around our empirical values. On the other hand, sample size considerably increases with increasing relative risk of a confounder even when the association remains small. The sample size required for an individually matched analysis practically equals that for an age-stratified analysis when the relative risk of the covariate is one. With a relative risk greater than one, however, the size for a matched analysis exceeds that for a stratified analysis and the ratio between them increases with increasing relative risk.

Air Pollutants↗

Reduction in sample size for studies of remodeling in heart failure by the use of cardiovascular magnetic resonance.

Fast breathhold cardiovascular magnetic resonance (CMR) has become a reference standard for the measurement of cardiac volumes, function, and mass. The implications of this for sample sizes for remodeling studies in heart failure (HF) have not been elucidated. We determined the reproducibility of CMR in HF and calculated the sample size requirements and compared them with published values for echocardiography. Breathhold gradient echo cines of the left ventricle were acquired in 20 patients with HF and 20 normal subjects. Sample size values were calculated from the interstudy standard deviation of the difference. The percentage variability of the measured parameters in our HF group of intraobserver (2.0-7.4%), interobserver (3.3-7.7%), and interstudy (2.5-4.8%) measurements was slightly larger than for our normal group (1.6-6.6%, 1.6-7.3%, and 2.0-7.3%, respectively) but remained comparable with previous studies in normal subjects. The calculated sample sizes in patients with HF for CMR to detect a 10-ml change in end-diastolic volume (n = 12) and end-systolic volume (n = 10), a 3% change in ejection fraction (n = 15), and a 10-g change in mass was (n = 9) were substantially smaller than recently published values for two-dimensional echocardiography (reduction of 81-97%). Breathhold CMR is a fast comprehensive technique for the assessment of cardiac volumes, function, and mass in HF that is accurate but also highly reproducible. This allows a considerable reduction in the patient numbers required to prove a hypothesis in research studies, which suggests a potential for important research cost savings.

Cardiac Volume↗

Sample size needed for student ratings of instruction.

The number of evaluation forms students are asked to complete is multiplying. To reduce that number, the present study determines the minimum sample size needed for accurate student ratings of instruction. Typical questionnaire items using four- and seven-category ratings scales were studied. Data for four class sizes (40, 60, 100, 140) were sampled in graduated sizes, and a standard error of the mean was computed for each sample size. A permissible error index was computed to estimate the accuracy of ratings obtained from any sample size needed for the four different class sizes. Figures are presented from which minimum sample sizes necessary for accurate student evaluation of instruction can be computed. The figures show that sampling only one-third of classes of 100-140 students is sufficient to obtain accurate evaluations.

Evaluation Studies as Topic↗

[Sample size and number needed to treat--a statistical case study].

The sample size necessary in each group to detect a significant effect of a vaccine on a given disease in a randomised clinical trial (RCT) can be calculated from the observed disease incidence without vaccine, the expected incidence with vaccine, and the desired level of significance and statistical power of the study. In this example, sample size is calculated for an RCT of the effect of pneumococcal vaccine on the incidence of pneumococcal sepsis. The number of persons needed to be vaccinated to prevent one episode of pneumococcal sepsis can be similarly calculated from the observed incidences with and without vaccine.

Humans↗