Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Mammographic Localization Biopsy for Ductal Carcinoma In Situ: A Simple Mapping Technique for Sampling and Size Estimation.

Size, grade, margin status, and microscopic invasion are currently significant parameters for management of ductal carcinoma in situ (DCIS). Size estimation of DCIS is difficult or impossible if tissue is sampled haphazardly. Histologic examination of the entire biopsy to exclude microscopic invasion is cumbersome and expensive, and may not yield more information than less extensive but planned sampling. One hundred twenty-four mammographic localization biopsies with DCIS including 4 cases with 3 mm or less invasion are presented to address these issues. All were examined by a mapping technique utilizing specimen radiograph as a guide for sampling and a schematic drawing to record the sections. This involved sequential sampling of the entire tissue for smaller biopsies, and en bloc sampling of the mammographic abnormality and surrounding tissue with end-to-end sampling of the remaining tissue at regular intervals for larger biopsies. All tissue was examined histologically initially in 55 cases, while tissue away from the lesion was selectively omitted in 71 cases. To address the issue of occult invasion, all initially unsampled tissue of 44 biopsies was submitted for histologic examination after completion of the case and the findings were recorded separately. Size estimates ranged from 4 to 70 mm. The solitary focus of microscopic invasion in four cases was present in the initial sections at the site corresponding to the mammographic abnormality. No invasion was identified in the additional tissue in any of the 44 cases. Margin status did not change for any. Using specimen radiographs as a guide, all the necessary information for DCIS, including the size and microscopic invasion, can be obtained by a planned sampling of tissue with diagrammatic documentation. Sequential sections of the entire tissue for small biopsies and sampling at regular intervals to include tissue at and around the mammographic abnormality for larger biopsies is appropriate. Microscopic invasion, when focal, is likely to be identified at the site of mammographic abnormality in the initial en bloc sections.

Journal Article↗

Sample size calculations for trials in health services research.

The current orthodox way of estimating sample size for a trial is through a power calculation based on a significance test. It therefore carries the assumption that this test should be the centerpiece of the statistical analysis. However, it is increasingly the case that confidence intervals are preferred to significance tests in summarising the results of trials, particularly in health services research. We believe that the way sample size is estimated should reflect this change and focus on the width of the confidence interval rather than on the outcome of a significance test. Such a method of estimation is described here and shown to have additional advantages of simplicity and transparency, enabling a more informed debate about the proposed size of trials.

Confidence Intervals↗

Musculoskeletal parameters of muscles crossing the shoulder and elbow and the effect of sarcomere length sample size on estimation of optimal muscle length.

BACKGROUND: Knowledge of musculoskeletal parameters is essential to understanding and modeling a muscle's force generating capability. A study of musculoskeletal parameters was conducted in two parts: (I) Empirical measurement of upper extremity musculoskeletal parameters. (II) Computational bootstrap simulation to examine statistical power of detecting optimal muscle length as a function of sarcomere length sample size and effect size. METHODS: Parameters were determined with a cadaver model. Sarcomere lengths were measured for 120 samples per muscle using laser diffraction and the mean sarcomere length used to estimate optimal muscle length. A bootstrap computational simulation was conducted to estimate variance in mean sarcomere length as a function of sample size. Statistical power for detecting optimal muscle length as a function of sample size and effect size was then determined. FINDINGS: Parameters are reported in tabular format. Power is 80% at approximately 85, 50, 40 and 25 samples for effect sizes of 0.5, 0.75, 1.0 and 1.5 mm respectively. INTERPRETATION: Musculoskeletal parameters for predicting muscle forces can be adequately measured in a cadaver model. Measurement of 40-60 sarcomere lengths per muscle is sufficient to calculate mean sarcomere length for estimating optimal muscle length with power of 80% for an effect size of 0.75-1.0 mm.

Adult↗

How responsive is the Multiple Sclerosis Impact Scale (MSIS-29)? A comparison with some other self report scales.

OBJECTIVES: To compare the responsiveness of the Multiple Sclerosis Impact Scale (MSIS-29) with other self report scales in three multiple sclerosis (MS) samples using a range of methods. To estimate the impact on clinical trials of differing scale responsiveness. METHODS: We studied three discrete MS samples: consecutive admissions for rehabilitation; consecutive admissions for steroid treatment of relapses; and a cohort with primary progressive MS (PPMS). All patients completed four scales at two time points: MSIS-29; Short Form 36 (SF-36); Functional Assessment of MS (FAMS); and General Health Questionnaire (GHQ-12). We determined: (1) the responsiveness of each scale in each sample (effect sizes): (2) the relative responsiveness of competing scales within each sample (relative efficiency): (3) the differential responsiveness of competing scales across the three samples (relative precision); and (4) the implications for clinical trials (samples size estimates scales to produce the same effect size). RESULTS: We studied 245 people (64 rehabilitation; 77 steroids; 104 PPMS). The most responsive physical and psychological scales in both rehabilitation and steroids samples were the MSIS-29 physical scale and the GHQ-12. However, the relative ability of different scales to detect change in the two samples was variable. Differing responsiveness implied more than a twofold impact on sample size estimates. CONCLUSIONS: The MSIS-29 was the most responsive physical and second most responsive psychological scale. Scale responsiveness differs notably within and across samples, which affects sample size calculations. Results of clinical trials are scale dependent.

Adolescent↗

Statistical significance and statistical power in hypothesis testing.

Experimental design requires estimation of the sample size required to produce a meaningful conclusion. Often, experimental results are performed with sample sizes which are inappropriate to adequately support the conclusions made. In this paper, two factors which are involved in sample size estimation are detailed--namely type I (alpha) and type II (beta) error. Type I error can be considered a "false positive" result while type II error can be considered a "false negative" result. Obviously, both types of error should be avoided. The choice of values for alpha and beta is based on an investigator's understanding of the experimental system, not on arbitrary statistical rules. Examples relating to the choice of alpha and beta are presented, along with a series of suggestions for use in experimental design.

Research Design↗

Cluster randomised trials in maternal and child health: implications for power and sample size.

BACKGROUND: Interventions based in the community can be evaluated by randomising clusters, such as general practices, rather than individuals, as in conventional randomised trials. This increases the sample size needed because of intracluster correlation. AIMS: To estimate sample size requirements for cluster randomised trials of interventions based in general practice directed at common health problems affecting mothers and infants. METHODS: Data were collected from a pilot trial of the effect of Citizen's Advice Bureau services involving six general practices. Outcome measures included the Edinburgh postnatal depression score, the Warwick child health and morbidity profile, number of visits to the general practitioner, and two questionnaires delivered at the beginning and end of the study. Intracluster correlation coefficients and inflation factors (the ratio of the sample size required for a cluster randomised trial to that required for an individually randomised trial) were calculated. RESULTS: Intracluster correlation coefficients ranged from 0 (sleeping problems, accidental injury, hospitalisation) to 0.09 (maternal smoking), with most being < 0.04 (for example, maternal depression, breast feeding, general health, minor illness, behavioural problems, and visits to the general practitioner). Assuming 50 cases/practice, cluster randomised trials require sample sizes up to 3 times greater than individually randomised trials for most health outcomes measured. CONCLUSIONS: These data enable sample sizes to be estimated for cluster randomised trials into a range of maternal and child health outcomes. Using such a design, approximately 40 practices would be sufficient to evaluate the effect of an intervention on maternal depression, sleeping, and behavioural problems, and non-routine visits to the general practitioner.

Analysis of Variance↗

Sample size re-estimation: recent developments and practical considerations.

Interim findings of a clinical trial often will be useful for increasing the sample size if necessary to provide the required power against the null hypothesis when the alternative hypothesis is true. Strategies for carrying out the interim examination that have been described over the past several years include "internal pilot studies", blinded interim sample size adjustment and conditional power. Simulation studies show that the alternative methods generally control the type I error rate satisfactorily, although the power properties are more variable. The important issues associated with sample size re-estimation are strategic, not numeric. Clearly expressed regulatory preferences suggest that methods not requiring unblinding the data before completion of the trial would be most appropriate. Extending a trial has its risks. The investigators/patients enrolled later in the course of a trial are not necessarily the same as those recruited/entered early. Re-activating the enrollment process may be sufficiently complicated and expensive to justify enrolling more investigators/patients at the outset. Since sample size re-estimation adjusts the sample size on the basis of variability while efficacy interim analysis adjusts the sample size based on the basis of estimated effect size, both principles can be used in the same trial. Sample size re-estimation may not be advisable for trials involving extended follow-up of individual patients or, more generally, when the follow-up time is long relative to the recruitment time. In such cases, it may be better to estimate the sample size conservatively and introduce an interim efficacy evaluation.

Clinical Trials as Topic↗

Sample size planning for developing classifiers using high-dimensional DNA microarray data.

Many gene expression studies attempt to develop a predictor of pre-defined diagnostic or prognostic classes. If the classes are similar biologically, then the number of genes that are differentially expressed between the classes is likely to be small compared to the total number of genes measured. This motivates a two-step process for predictor development, a subset of differentially expressed genes is selected for use in the predictor and then the predictor constructed from these. Both these steps will introduce variability into the resulting classifier, so both must be incorporated in sample size estimation. We introduce a methodology for sample size determination for prediction in the context of high-dimensional data that captures variability in both steps of predictor development. The methodology is based on a parametric probability model, but permits sample size computations to be carried out in a practical manner without extensive requirements for preliminary data. We find that many prediction problems do not require a large training set of arrays for classifier development.

Computer Simulation↗

Power and sample size of therapeutic trials in procedural dermatology: how many patients are enough?

BACKGROUND: Many new devices and therapeutic interventions are continually introduced in cutaneous surgery. The efficacy of these new techniques must be compared with that of preexisting standards so that patients can be appropriately counseled. OBJECTIVES: The purpose of this article is to (1) review methods for estimating sample size and power, (2) estimate the range of sample sizes sufficient to ensure that true differences are not missed in clinical trials of new procedural dermatologic therapies, and (3) consider the reasons why the sample size may be too small in procedural dermatology trials and how this problem can be addressed. METHODS: (1) Selective review of textbooks and other relevant literature, presentation of a brief tutorial describing sample size and power determination for therapeutic clinical trials comparing two groups with continuous outcomes variables; (2) implementation of standard formulae and assumptions to estimate sample size in cutaneous surgery therapeutic trials. RESULTS: Assuming that one group receives a standard surgical intervention and another group undergoes a new technique, to identify a moderate difference in efficacy between groups, at least 50 to 200 subjects will need to be enrolled if conventional strategies are used to reduce the likelihood of finding a difference that does not really exist (Type I error), as well as the likelihood of missing a true difference (Type II error). CONCLUSION: By face validity, it is apparent that most efficacy comparisons in procedural dermatology have low sample size and a concomitant risk of failing to detect actual differences between therapeutic arms. Owing to the limitations that restrict surgeons from frequently performing large randomized controlled trials in procedural dermatology, meta-analyses may be needed to pool the results of smaller studies. When it is critically important that differences between groups be accurately identified, dermatologic surgeons may consider eschewing smaller trials in favor of collaborating on larger trials with an adequate sample size.

Dermatology↗

Use of the standard error as a reliability index of interest: an applied example using elbow flexor strength data.

The intraclass correlation coefficient (ICC) and the standard error of measurement (SEM) are two reliability coefficients that are reported frequently. Both measures are related; however, they define distinctly different properties. The magnitude of the ICC defines a measure's ability to discriminate among subjects, and the SEM quantifies error in the same units as the original measurement. Most of the statistical methodology addressing reliability presented in the physical therapy literature (eg, point and interval estimations, sample size calculations) focuses on the ICC. Using actual elbow flexor make and break strength measurements, this article illustrates a method for estimating a confidence interval for the SEM, shows how an a priori specification of confidence interval width can be used to estimate sample size, and provides several approaches for comparing error variances (and square root of the error variance, or the SEM).

Analysis of Variance↗

Are KISS data representative of German intensive care units? Statistical issues.

OBJECTIVES: Data collected within the German nosocomial infection surveillance system KISS are recommended as reference data for judging nosocomial infection rates in German intensive care units (ICUs). It is unknown whether the KISS data tend to under- or overestimate the true infection incidence rates. In this article, methodological aspects of the SIR1 study on the incidence of nosocomial infections are discussed, with the aim of estimating unbiased incidence rates of nosocomial infections in interdisciplinary German ICUs and examining whether the KISS data are representative. METHODS: We discuss the following methodological issues: 1) Sample size estimation. 2) Stratified random sampling of German ICUs. 3) Investigation of seasonal effects. 4) Statistical modeling of incidence rates using a negative binomial regression model. 5) Comparison of weighted incidence rates with the standardized rate ratio (SRR). RESULTS: Random sampling proved difficult to realize in practice since many ICUs refused to participate, particularly those in small hospitals. Analysis was adjusted for hospital size. No seasonal trends were found in the KISS data. Due to marked differences between ICUs, the number of infections is over-dispersed compared to a Poisson model, so negative binomial regression was used. Fifty ICUs were observed for two consecutive months each, corresponding to 21,832 patient days, during which 262 infections occurred. Infections were more frequent in large hospitals. The incidence rates provided by the SIR study are on average (SRR) 1.89 (1.63-2.20) times as large as those estimated by the KISS system. CONCLUSION: For estimating nosocomial infection incidence rates, random sampling and statistical modeling of over-dispersion were successfully performed. The study provides evidence that the KISS surveillance system tends to underestimate the true incidence rates of nosocomial infections in German ICUs.

Binomial Distribution↗

Sample size calculation for clinical trials: the impact of clinician beliefs.

The UK Medical Research Council (MRC) randomized trial of gastric surgery, ST01, compared conventional (D1) with radical (D2) surgery. Sample size estimation was based upon the consensus opinion of the surgical members of the design team, which suggested that a change in 5-year survival from 20% (D1) to 34% (D2) could be realistic and medically important. On the basis of these survival rates, the sample size for the trial was 400 patients. However, this trial was exceptional in the way that a survey of surgeons' opinions was made at the start of the trial, in 1986, and again before results were analysed but after termination of the trial in 1994. At the initial survey, the three surgeons from the trial steering committee and 23 other surgeons experienced in treating gastric carcinoma were given detailed questionnaires. They were asked about the expected survival rate in the D1 group, anticipated difference in survival from D2 surgery, and what difference would be medically important and influence future treatment of patients. The consensus opinion of those surveyed was that there might be a survival improvement of 9.4%. In 1994, prior to closure of the trial, and before any survival information was disclosed, the survey was repeated with 21 of the original 26 surgeons. At this second survey, the opinion of the trial steering committee was that 9.5% difference was more realistic. This was in accord with the opinion of the larger group, which remained little changed since 1986. The baseline 5-year D1 survival was thought likely to be about 32%, which corresponded closely to the actual survival of recruited patients. Revised sample size calculations suggested that, on the basis of these more recent opinions, between 800 and 1200 patients would have been required. Both surveys assessed the level of treatment benefit that was deemed to be sufficient for causing surgeons to change their practice. This showed that the 13% difference in survival used as the study target was clinically relevant, but also indicated that many clinicians would remain unwilling to change their practice if the difference is only 9.5%. The experience of this carefully designed trial illustrates the problems of designing long-term, randomized trials. It raises interesting questions about the common practice of basing sample size estimates upon the beliefs of a trial design committee that may include a number of enthusiasts for the trial treatment. If their opinion of anticipated effect sizes drives the design of the trial, rather than the opinion of a larger community of experts that includes sceptics as well as enthusiasts, there is likely to be a serious miscalculation of sample size requirements.

Bayes Theorem↗

Power analyses for correlations from clustered study designs.

Power analysis constitutes an important component of modern clinical trials and research studies. Although a variety of methods and software packages are available, almost all of them are focused on regression models, with little attention paid to correlation analysis. However, the latter is arguably a simpler and more appropriate approach for modelling concurrent events, especially in psychosocial research. In this paper, we discuss power and sample size estimation for correlation analysis arising from clustered study designs. Our approach is based on the asymptotic distribution of correlated Pearson-type estimates. Although this asymptotic distribution is easy to use in data analysis, the presence of a large number of parameters creates a major problem for power analysis due to the lack of real data to estimate them. By introducing a surrogacy-type assumption, we show that all nuisance parameters can be eliminated, making it possible to perform power analysis based only on the parameters of interest. Simulation results suggest that power and sample size estimates obtained under the proposed approach are robust to this assumption.

Clinical Trials as Topic↗

On the inappropriateness of an EM algorithm based procedure for blinded sample size re-estimation.

When planning a clinical trial the sample size calculation is commonly based on an a priori estimate of the variance of the outcome variable. Misspecification of the variance can have substantial impact on the power of the trial. It is therefore attractive to update the planning assumptions during the ongoing trial using an internal estimate of the variance. For this purpose, an EM algorithm based procedure for blinded variance estimation was proposed for normally distributed data. Various simulation studies suggest a number of appealing properties of this procedure. In contrast, we show that (i) the estimates provided by this procedure depend on the initialization, (ii) the stopping rule used is inadequate to guarantee that the algorithm converges against the maximum likelihood estimator, and (iii) the procedure corresponds to the special case of simple randomization which, however, in clinical trials is rarely applied. Further, we show that maximum likelihood estimation leads to no reasonable results for blinded sample size re-estimation due to bias and high variability. The problem is illustrated by a clinical trial in asthma.

Administration, Inhalation↗

Research cost analyses to aid in decision making in the conduct of a large prevention trial, CARET. Carotene and Retinol Efficacy Trial.

Because of their larger study populations and longer durations, prevention trials typically are more costly than treatment trials. Thus it is important to analyze costs systematically to aid in making cost-effective decisions during the conduct of prevention trials as well as in the original design. Cost analysis must be tied to sample size estimation because costs depend on such factors as the total number of person-years of follow-up and the number of trial outcomes, which are not basic design parameters but are derived quantities resulting from sample size estimation. We illustrate the use of cost analysis to decide among options for future conduct of an ongoing prevention trial with three issues that have arisen during the Carotene and Retinol Efficacy Trial (CARET): the trade-off between extending the duration of the trial or increasing the number of participants, the effect on costs of delay in accrual, and the cost effectiveness of particular retention activities.

Anticarcinogenic Agents↗

[Estimation of sample size in randomized controlled clinical trials--elimination of various systematic biases by the utilization of computer network system].

Although sample size calculation is mandatory in clinical trials, the power of most of the published small size clinical studies, was only about 0.25 instead of 0.80 which is normally required in a standard randomized controlled trial. This phenomenon is probably due to the publication bias. In a cancer clinical study, the ideal sample size needed in a trial exceeds over a thousand, when significance level was fixed under 0.05 and treatment difference is estimated from 10 to 15%. In such cases, evaluation of a new treatment in less common cancers or in a specified strata become unfeasible. Although statistically plausible clinical trial is difficult, small trials of newly advocated treatment is likely to be performed elsewhere, and presentation of these haphazard results might propagate a wrong information about the new treatment. To prevent the dissemination of these biased results, 1) preregistration of all planned clinical trials to an authorized organization before the initiation of the trial, 2) randomization of the patients entered in the trial from the first case, and 3) registration of all available data of the patients and meta-analysis of preregistered multiple trials, should be the most effective counterplans. In order to achieve those functions, installation of a computer assisted coordinating center is considered to be the best solution for the proper evaluation of the clinical trial as well as for the evaluation of a new treatment. With the collaboration of regional affiliating hospitals, a pilot study have started to establish a computer network system.

Database Management Systems↗

An iterative approach to the analysis of EM autoradiographs. II. Estimates of sample sizes and confidence limits.

The errors inherent in EM autoradiography are discussed and certain of them deemed to be of particular practical significance in the quantitative assessment of preparations. A method is described for estimating the standard errors attributable to each of several sources of variation and thence for obtaining the overall standard error value to be attached to relative activity estimates obtained in the method of Downs & Williams (1978). In an appendix, a fully worked example is given illustrating clearly the strategy of the method and the magnitudes of error estimates that are to be attached to the final specific activity values.

Autoradiography↗

On sample sizes to estimate the protective efficacy of a vaccine.

To estimate vaccine protective efficacy, defined as VE = 1 - ARV/ARU where ARV is the disease attack rate in the vaccinated group and ARU is the disease attack rate in the controls, investigators have used both cohort and case-control designs. For each design, we present a method for calculation of the sample size required to provide an approximate confidence interval for VE of predetermined width and probability of coverage. The required sample size is a function of the desired width of the confidence interval, the probability of coverage, the assumed VE, and, for cohort designs, the assumed disease attack rate in the controls, and for case-control designs, the assumed vaccine exposure prevalence for the controls.

Child, Preschool↗