Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Sample size calculations in scleroderma: a rational approach to choosing outcome measurements in scleroderma trials.

Subjects with both diffuse and limited scleroderma were studied to calculate the baseline characteristics of several commonly used outcome measurements in order to provide parameters for sample size calculations for scleroderma clinical trials. From these estimates, outcome measurements were chosen as potentially responsive to change in clinical trials if their sample sizes were not prohibitively large. Forty-five patients with scleroderma were systematically assessed to determine the means and standard deviations whereby sample size calculations can be performed using this information. Examples of sample sizes were determined for the entire group, and for 2 subsets: those with diffuse scleroderma and those with diffuse disease of recent onset. Many baseline characteristics were significantly different in patients with diffuse compared to limited systemic sclerosis. The baseline values were different for the Health Assessment Questionnaire (HAQ) disability score, Functional Index, grip strength, oral aperture, finger-to-palm distance, skin score, and physician global assessment. Sample sizes can vary widely depending upon the outcome measurement chosen and the range of deltas used within the different scleroderma subsets. Sample size requirements for many outcome measures are extremely large due to marked variability in the baseline measures. Primary outcome measures in scleroderma trials should be chosen which have adequate power to detect a minimal clinically relevant change in the primary outcome measurements at the sample sizes employed. All other outcome measures should be ranked as secondary. Skin scores, global assessments, and grip strength measurements require smaller sample sizes than the other outcome measurements which were studied. Sample sizes in future trials will vary depending upon the proportion of patients with diffuse and limited scleroderma who are included.

Female↗

A comparative study of conditional maximum likelihood estimation of a common odds ratio.

The finite-sample properties of various point estimators of a common odds ratio from multiple 2 X 2 tables have been considered in a number of simulation studies. However, the conditional maximum likelihood estimator has received only limited attention. That omission is partially rectified here for cases of relatively small numbers of tables and moderate to large within-table sample sizes. The conditional maximum likelihood estimator is found to be superior to the unconditional maximum likelihood estimator, and equal or superior to the Mantel-Haenszel estimator in both bias and precision.

Biometry↗

How should we measure maternal mortality in the developing world? A comparison of household deaths and sibling history approaches.

OBJECTIVE: A reduction in the maternal mortality ratio (MMR) is one of six health-related Millennium Development Goals (MDGs). However, there is no consensus about how to measure MMR in the many countries that do not have complete registration of deaths and accurate ascertainment of cause of death. In this study, we compared estimates of pregnancy-related deaths and maternal mortality in a developing country from three different household survey measurement approaches: a module collecting information on deaths of respondents' sisters; collection of information about recent household deaths with a time-of-death definition of maternal deaths; and a verbal autopsy instrument to identify maternal deaths. METHODS: We used data from a very large nationally-representative household sample survey conducted in Bangladesh in 2001. A total of 104 323 households were selected for participation, and 99 202 households (95.1% of selected households, 98.8% of contacted households) were successfully interviewed. FINDINGS: The sisterhood and household death approaches gave very similar estimates of all-cause and pregnancy-related mortality; verbal autopsy gave an estimate of maternal deaths that was about 15% lower than the pregnancy-related deaths. Even with a very large sample size, however, confidence intervals around mortality estimates were similar for all approaches and exceeded +/- 15%. CONCLUSION: Our findings suggest that with improved training for survey data collectors, both the sisterhood and household deaths methods are viable approaches for measuring pregnancy-related mortality. However, wide confidence intervals around the estimates indicate that routine sample surveys cannot provide the information needed to monitor progress towards the MDG target. Other approaches, such as inclusion of questions about household deaths in population censuses, should be considered.

Adolescent↗

Therapy-related acute myeloid leukemia following treatment with epipodophyllotoxins: estimating the risks.

In the past decade, therapy-related acute myeloid leukemia (t-AML) following treatment with regimens that include inhibitors of topoisomerase-II (TOPO-II) has been reported with increasing frequency. These cases of t-AML generally have a shorter latency period than t-AML following alkylator therapy, are associated with chromosomal translocations (especially involving chromosome band 11q23), and usually present as M4 or M5 FAB subtype. Although the epipodophyllotoxins (etoposide and teniposide) have been most often implicated, similar cases of t-AML occur following therapy with other classes of Topo-II inhibitors (e.g., anthracyclines). There is wide variation in published studies in the estimates of risk of t-AML following epipodophyllotoxin therapy. These varying estimates may reflect a number of factors, including: small sample size leading to large confidence intervals around risk estimates; varying susceptibility of different patient populations; varying schedules of epipodophyllotoxin administration; different cumulative doses of epipodophyllotoxins; and administration of epopodophyllotoxins with additional agents that may alter the leukemogenic effect of the epipodophyllotoxins. Available data suggest that children with acute lymphocytic leukemia (ALL) treated with high cumulative doses of epipodophyllotoxins using either weekly or twice-weekly schedules of administration have a relatively high risk of developing t-AML (5-12% cumulative risk). On the other hand, germ cell patients treated with relatively low cumulative doses of etoposide (usually 1,500-2,500 mg/m2) appear to have a low risk for developing t-AML. There is inadequate experience at this time with higher cumulative doses of etoposide (e.g., 4,000-5,000 mg/m2 as used for pediatric solid tumors) given on a daily x 5 schedule to allow estimates of risk to be developed for this schedule and cumulative dose. The Cancer Therapy Evaluation Program (CTEP) of the National Cancer Institute (NCI) has developed a monitoring plan designed to obtain reliable estimates of the risk of t-AML following epipodophyllotoxin treatment. Twelve Cooperative Group clinical trials that use epipodophyllotoxins at either low (< 1,500 mg/m2), moderate (1,500-3,999 mg/m2), or higher cumulative doses (> 4,000 mg/m2) are being prospectively monitored for cases of t-AML occurring among patients entered onto the trials.

Alkylating Agents↗

Estimating the size of subpopulations of heroin users: applications of log-linear models to capture/recapture sampling.

This article reviews two of the major methodologies applied to estimation of the number of heroin abusers: survey research methods and the capture/recapture technique. The main focus of the paper is to show the flexibility of the capture/recapture approach in handling not only the dependence of samples of heroin users but also the nonhomogeneity of sampling probabilities, allowing estimation in populations which are mixtures of qualitatively different heroin user types. Models with these features are illustrated using both simulated and real heroin abuse data.

Data Collection↗

Sample size and mass range effects on the allometric exponent of basal metabolic rate.

The controversial relationship between body mass and basal metabolic rate in animals revolves around two questions: what is the allometric scaling exponent and what is the functional basis for it? For mammals, the first question could be resolved if measurements from all 4600 extant species were available, but this study shows that data for only 150 species, spanning three to four orders of magnitude variation in body mass, are sufficient to accurately determine the exponent. Because the currently available data set includes about 600 species that vary over five orders of magnitude in body size, further increases in sample size are unlikely to change the estimate of the scaling exponent.

Animals↗

Interaction, subgroup analysis and sample size.

The term "interaction" has both statistical and scientific connotations that do not always coincide. Multistage models are used as a bridge between these two viewpoints and as a way of illustrating the different types of qualitative interaction. The most important type of interaction for metabolic polymorphisms is between a genetic trait and an environmental or lifestyle exposure. In some cases these factors act multiplicatively on the risk ratios and so do not interact in the statistical sense. Nevertheless, the risk can be much higher when both are present. Interactions present severe problems of interpretation, as naive comparisons in subgroups can be very misleading and can produce false positive results because of the large number of comparisons. Methods for combating this are discussed. The main requirement, however, is for a substantially larger sample size than would be required for estimating the main effects.

Bayes Theorem↗

Bootstrap standard error and confidence intervals for the correlation corrected for range restriction: a simulation study.

The standard Pearson correlation coefficient is a biased estimator of the true population correlation, rho, when the predictor and the criterion are range restricted. To correct the bias, the correlation corrected for range restriction, rc, has been recommended, and a standard formula based on asymptotic results for estimating its standard error is also available. In the present study, the bootstrap standard-error estimate is proposed as an alternative. Monte Carlo simulation studies involving both normal and nonnormal data were conducted to examine the empirical performance of the proposed procedure under different levels of rho, selection ratio, sample size, and truncation types. Results indicated that, with normal data, the bootstrap standard-error estimate is more accurate than the traditional estimate, particularly with small sample size. With nonnormal data, performance of both estimates depends critically on the distribution type. Furthermore, the bootstrap bias-corrected and accelerated interval consistently provided the most accurate coverage probability for rho.

Adolescent↗

Sample size recalculation in internal pilot study designs: a review.

The adequacy of sample size is important to clinical trials. In the planning phase of a trial, however, the investigators are often quite uncertain about the sizes of parameters which are needed for sample size calculations. A solution to this problem is mid-course recalculation of the sample size during the ongoing trial. In internal pilot study designs, nuisance parameters are estimated on the basis of interim data and the sample size is adjusted accordingly. This review attempts to give an overview on the available methods. It is written not only for biometricians who are already familar with the the topic and wish to update their knowledge but also for users new to the subject.

Clinical Trials as Topic↗

Stereologic estimation of breast tumor size.

OBJECTIVE: The largest tumor diameter, D(T), is a variable of great clinical value in breast cancer but a biased and imprecise estimator of real tumor size. Three-dimensional, shape-independent estimates would more realistically reflect the tumor bulk and provide more accurate clinical staging. For experimental oncology, the measurements may be useful for precise assessment of tumor burden. STUDY DESIGN: In 64 prospectively collected breast cancers, unbiased stereology was used for estimating the gross tumor volume, V(T), cutting specimens into parallel, equally thick sections with subsequent determination of total tumor sectional area. For comparison, the volume of invasive tumor epithelium, "V"(epi), was obtained by microscopic examination of systematically sampled tissue fractions of the same tumors, embedded in both methacrylate and paraffin. RESULTS: The median D(T) was 2.2 cm, and the median V(T) was 6.72 cm3. The correlation between these variables was not very close (r = .77), and the slope of the regression line was steeper than expected, presumably reflecting a change in tumor shape with growing size. The sampling scheme used for estimation of V(T) proved highly efficient, yielding a mean error coefficient of 9%. The median "V"(epi) in methacrylate was 1.19 cm3, 21% larger than in paraffin. Estimates of "V" (epi) were highly reproducible (r = .97) and correlated highly with point counting-based estimates of the feature (r = .96). "V"(epi) correlated with V(T) (r > or = .75), but the slopes of the regression lines were steeper than expected, corresponding with the correlation found between epithelial volume fraction and tumor size (r = .26). On average, about 25% of the gross tumor was composed of invasive epithelium, but with a wide range. CONCLUSION: In breast cancer, realistic estimates of tumor volume, volume of invasive epithelium and epithelial volume fraction can be obtained by efficient stereologic techniques, which seem useful for clinical and experimental oncology. In the present methodologic study, baseline data were generated. Further studies are needed to assess the clinical value of stereologic tumor size estimates as compared with traditional staging parameters.

Breast Neoplasms↗

Four-fold table cell frequencies imputation in meta analysis.

Meta analysis is a collection of quantitative methods devoted to combine summary information from related but independent studies. Because research reports usually present only data reductions and summary statistics rather than detailed data, the reviewer must often resort to rather crude methods for constructing summary effect estimate suitable for meta analysis pooling methods. When the studies involve a binary variable, both number of events and sample sizes are required to compute pooled estimate and its confidence interval. Sometimes, only summary statistics and related confidence intervals are provided in the publication. Although it is possible to estimate the standard error of each study's effect measure using the confidence interval from each study, this lack of detailed data compels the reviewers to use the inverse variance method to perform meta analysis, or to exclude the works with incomplete data. This paper shows three methods to reconstruct four-fold tables when summary measures for binary data and related confidence intervals and sample sizes are provided. The methods are discussed through a wider application example to assess the reconstruction precision, and the impact of using reconstructed data on meta analysis results. These methods seem to yield a correct reconstruction if original measures are reported at least with two decimal places. Meta analysis results do not seem seriously affected by the use of reconstructed data. These methods allow the reviewer to use full meta analysis statistical tools, instead of the simple inverse variance method, and can greatly contribute to the completeness of systematic reviews.

Confidence Intervals↗

How to fit a response time distribution.

Among the most valuable tools in behavioral science is statistically fitting mathematical models of cognition to data--response time distributions, in particular. However, techniques for fitting distributions very widely, and little is known about the efficacy of different techniques. In this article, we assess several fitting techniques by simulating six widely cited models of response time and using the fitting procedures to recover model parameters. The techniques include the maximization of likelihood and least squares fits of the theoretical distributions to different empirical estimates of the simulated distributions. A running example is used to illustrate the different estimation and fitting procedures. The simulation studies reveal that empirical density estimates are biased even for very large sample sizes. Some fitting techniques yield more accurate and less variable parameter estimates than do others. Methods that involve least squares fits to density estimates generally yield very poor parameter estimates.

Chi-Square Distribution↗

Modelling macroparasite aggregation using a nematode-sheep system: the Weibull distribution as an alternative to the negative binomial distribution?

Macroparasites are almost always aggregated across their host populations, hence the Negative Binomial Distribution (NBD) with its exponent parameter k is widely used for modelling, quantifying or analysing parasite distributions. However, many studies have pointed out some drawbacks in the use of the NBD, with respect to the sensitivity of k to the mean number of parasites per host or the under-representation of the heavily infected hosts in the estimate of k. In this study, we compare the fit of the NBD with 4 other widely used distributions on observed parasitic gastrointestinal nematode distributions in their sheep host populations (11 datasets). Distributions were fitted to observed data using maximum likelihood estimator and the best fits were selected using the Akaike's Information Criterion (AIC). A simulation study was also conducted in order to assess the possible bias in parameter estimations especially in the case of small sample sizes. We found that the NBD is seldom the best fit for gastrointestinal nematode distributions. The Weibull distribution was clearly more appropriate over a very wide range of degrees of aggregation, mainly because it was more flexible in fitting the heavily infected hosts. Moreover, the Weibull distribution estimates are less sensitive to sample size. Thus, when possible, we suggest to carefully check on observed data if the NBD is appropriate before conducting any further analysis on parasite distributions.

Animals↗

Evaluation of decision rules for frequency-doubling technology screening tests.

PURPOSE: Frequency-doubling technology (FDT) perimetry has shown promise as a screening test for glaucoma. This study investigates different possible decision rules for FDT screening by applying them to groups of normal and glaucoma subjects. METHODS: Within three centers, 218 subjects (aged 15-88 years; 78 with glaucoma, 140 without ocular disease) were each tested twice with the screening program of the FDT perimeter. The subjects consisted of 140 normal subjects with no evidence of glaucoma or other ocular disease likely to affect the visual field and 78 subjects with a diagnosis of glaucoma and no other ocular disease. Fifteen decision rules were applied to the data to compare their sensitivity and specificity. RESULTS: Estimated specificities of the different decision rules ranged from 78% to 99%, although with this sample size, the confidence intervals for these estimates are quite large. Estimated sensitivities ranged from 40% to 72%. Suggested criteria for distinguishing normal subjects from those with glaucoma seem to be either a cluster of two or more adjacent locations abnormal at the p < 2% level with at least one location confirmed or a single location very abnormal (p < 1%) and confirmed. CONCLUSIONS: Specificity was clearly improved by confirming an apparently abnormal test result by repeating the screening test outweighing the resultant small loss in sensitivity. These findings provide useful information for making an informed choice of decision rules for FDT screening results.

Adolescent↗

Strategies for genome-wide association studies: optimization of study designs by the stepwise focusing method.

Recently, the use of genome-wide linkage disequilibrium (LD) analysis to localize traits has attracted much attention because of the introduction of high-throughput genotyping systems. However, a limitation of such studies is often the total cost of genotyping in addition to sample size. Therefore, it is important to estimate optimal conditions for such a study given the total cost of genotyping. In the present study, we have introduced the "stepwise focusing method," in which candidate markers are selected in a stepwise fashion. In the first focusing step, samples from both case and control groups are genotyped at a certain number of single-nucleotide polymorphisms (SNPs) (for example, 50000), and the markers that exhibit significant intergroup differences by a chi(2) test are selected. In the first step, the risk of type I error is set rather high (for example, 0.1), and, therefore, most of the selected markers are false positives. In the second step, the markers selected in the first step are tested by using samples obtained from a different set of case-control samples. We performed extensive simulation studies to estimate both the type I error and the power of the test by changing parameters such as genotype relative risk, disease allele frequency, and sample size. If the total number of genotypings was limited, the stepwise focusing method yielded optimal conditions and was more powerful than conventional methods.

Genetic Markers↗

Subgroup analyses in randomised controlled trials: quantifying the risks of false-positives and false-negatives.

BACKGROUND: Subgroup analyses are common in randomised controlled trials (RCTs). There are many easily accessible guidelines on the selection and analysis of subgroups but the key messages do not seem to be universally accepted and inappropriate analyses continue to appear in the literature. This has potentially serious implications because erroneous identification of differential subgroup effects may lead to inappropriate provision or withholding of treatment. OBJECTIVES: (1) To quantify the extent to which subgroup analyses may be misleading. (2) To compare the relative merits and weaknesses of the two most common approaches to subgroup analysis: separate (subgroup-specific) analyses of treatment effect and formal statistical tests of interaction. (3) To establish what factors affect the performance of the two approaches. (4) To provide estimates of the increase in sample size required to detect differential subgroup effects. (5) To provide recommendations on the analysis and interpretation of subgroup analyses. METHODS: The performances of subgroup-specific and formal interaction tests were assessed by simulating data with no differential subgroup effects and determining the extent to which the two approaches (incorrectly) identified such an effect, and simulating data with a differential subgroup effect and determining the extent to which the two approaches were able to (correctly) identify it. Initially, data were simulated to represent the 'simplest case' of two equal-sized treatment groups and two equal-sized subgroups. Data were first simulated with no differential subgroup effect and then with a range of types and magnitudes of subgroup effect with the sample size determined by the nominal power (50-95%) for the overall treatment effect. Additional simulations were conducted to explore the individual impact of the sample size, the magnitude of the overall treatment effect, the size and number of treatment groups and subgroups and, in the case of continuous data, the variability of the data. The simulated data covered the types of outcomes most commonly used in RCTs, namely continuous (Gaussian) variables, binary outcomes and survival times. All analyses were carried out using appropriate regression models, and subgroup effects were identified on the basis of statistical significance at the 5% level. RESULTS: While there was some variation for smaller sample sizes, the results for the three types of outcome were very similar for simulations with a total sample size of greater than or equal to 200. With simulated simplest case data with no differential subgroup effects, the formal tests of interaction were significant in 5% of cases as expected, while subgroup-specific tests were less reliable and identified effects in 7-66% of cases depending on whether there was an overall treatment effect. The most common type of subgroup effect identified in this way was where the treatment effect was seen to be significant in one subgroup only. When a simulated differential subgroup effect was included, the results were dependent on the nominal power of the simulated data and the type and magnitude of the subgroup effect. However, the performance of the formal interaction test was generally superior to that of the subgroup-specific analyses, with more differential effects correctly identified. In addition, the subgroup-specific analyses often suggested the wrong type of differential effect. The ability of formal interaction tests to (correctly) identify subgroup effects improved as the size of the interaction increased relative to the overall treatment effect. When the size of the interaction was twice the overall effect or greater, the interaction tests had at least the same power as the overall treatment effect. However, power was considerably reduced for smaller interactions, which are much more likely in practice. The inflation factor required to increase the sample size to enable detection of the interaction with the same power as the overall effect varied with the size of the interaction. For an interaction of the same magnitude as the overall effect, the inflation factor was 4, and this increased dramatically to of greater than or equal to 100 for more subtle interactions of < 20% of the overall effect. Formal interaction tests were generally robust to alterations in the number and size of the treatment and subgroups and, for continuous data, the variance in the treatment groups, with the only exception being a change in the variance in one of the subgroups. In contrast, the performance of the subgroup-specific tests was affected by almost all of these factors with only a change in the number of treatment groups having no impact at all. CONCLUSIONS: While it is generally recognised that subgroup analyses can produce spurious results, the extent of the problem is almost certainly under-estimated. This is particularly true when subgroup-specific analyses are used. In addition, the increase in sample size required to identify differential subgroup effects may be substantial and the commonly used 'rule of four' may not always be sufficient, especially when interactions are relatively subtle, as is often the case. CONCLUSIONS--RECOMMENDATIONS FOR SUBGROUP ANALYSES AND THEIR INTERPRETATION: (1) Subgroup analyses should, as far as possible, be restricted to those proposed before data collection. Any subgroups chosen after this time should be clearly identified. (2) Trials should ideally be powered with subgroup analyses in mind. However, for modest interactions, this may not be feasible. (3) Subgroup-specific analyses are particularly unreliable and are affected by many factors. Subgroup analyses should always be based on formal tests of interaction although even these should be interpreted with caution. (4) The results from any subgroup analyses should not be over-interpreted. Unless there is strong supporting evidence, they are best viewed as a hypothesis-generation exercise. In particular, one should be wary of evidence suggesting that treatment is effective in one subgroup only. (5) Any apparent lack of differential effect should be regarded with caution unless the study was specifically powered with interactions in mind. CONCLUSIONS--RECOMMENDATIONS FOR RESEARCH: (1) The implications of considering confidence intervals rather than p-values could be considered. (2) The same approach as in this study could be applied to contexts other than RCTs, such as observational studies and meta-analyses. (3) The scenarios used in this study could be examined more comprehensively using other statistical methods, incorporating clustering effects, considering other types of outcome variable and using other approaches, such as Bootstrapping or Bayesian methods.

Data Interpretation, Statistical↗

Increasing the sample size during clinical trials with t-distributed test statistics without inflating the type I error rate.

In clinical trials with t-distributed test statistics the required sample size depends on the unknown variance. Taking estimates from previous studies often leads to a misspecification of the true value of the variance. Hence, re-estimation of the variance based on the collected data and re-calculation of the required sample size is attractive. We present a flexible method for extensions of fixed sample or group-sequential trials with t-distributed test statistics. The method can be applied at any time during the course of the trial and does not require the necessity to pre-specify a sample size re-calculation rule. All available information can be used to determine the new sample size. The advantage of our method when compared with other adaptive methods is maintenance of the efficient t-test design when no extensions are actually made. We show that the type I error rate is preserved.

Computer Simulation↗

Reassessing Instrument Strength in Two-Sample Mendelian Randomization Analysis.

Mendelian randomization (MR) analysis is widely used to estimate causal relationships between risk factors and outcomes of interest. Two-sample MR approaches have gained increasing attention in genetic epidemiology due to the growing availability of Genome-Wide Association Study (GWAS) summary statistics from public databases. A critical step in two-sample MR is the selection of genetic variants as instrumental variables (IVs). Although genome-wide significant variants are typically preferred, the inclusion of variants with weaker association p-values is considered, as they may potentially improve power through an increased instrument number of instruments, while they may introduce weak instrument bias and attenuate effect estimates towards the null. Our simulation results show that even modest levels of pleiotropy substantially increase the variability of causal effect estimates, while the inclusion of weak IVs does not substantially affect the direction and variability of causal effect estimates in most cases. In real data analyses, we used two released versions of FinnGen GWAS summary statistics with different sample sizes as exposure GWASs to assess the influence of weak IVs. Here, the inclusion of IVs with higher exposure-association p-values resulted in weakened estimated effect sizes, particularly when the exposure GWAS sample size was small. These findings suggest that incorporating weak IVs is reasonable when the exposure GWAS sample size is large, but it poses a risk of falsely concluding null associations when the exposure GWAS sample size is small.

Journal Article↗