Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

A stochastic model to estimate the prevalence of scrapie in Great Britain using the results of an abattoir-based survey.

In 1997/1998, an abattoir survey was conducted to determine the likely exposure of the human population to transmissible spongiform encephalopathy (TSE) infection in sheep submitted for slaughter in Great Britain. The survey examined brain material from 2809 sheep processed through British abattoirs. Sampling was targeted by age: 45% of animals tested were > or =15 months old. All samples of adequate quality (98%) were tested for signs of scrapie infection using histopathology and scrapie-associated fibril (SAF) detection and 500 were tested using immunohistochemistry (IHC). No conclusive positive animals were found using either histology or IHC. Ten animals were positive by SAF. Standard statistical analyses suggest (with 95% confidence) that the prevalence of detectable (by histopathology) infection in the slaughter population was < or =0.11%. However, the incubation period of scrapie is long (usually around 2-3 years) and none of the tests used in the survey is capable of detecting scrapie infection in the early stages of infection. We present an age-structured stochastic model incorporating parameters for the incubation period of scrapie, prevalence of infection by age and test sensitivity. Using the model, we demonstrate that the negative results obtained for all samples using IHC and histopathology are consistent with a true prevalence of infection in the slaughter population of up to 11%. This suggests that up to 300 of the animals tested might have been infected but the infection was not sufficiently advanced in these animals to be detectable by IHC or histopathology. The survey was designed to detect a prevalence of 1% with a precision of +/-0.5% and a confidence level of 95% in each age group assuming that diagnostic tests were 100% specific and sensitive from a known stage in the incubation period. The results of the model demonstrate that to estimate a true prevalence of scrapie infection of 1% with an accuracy of +/-0.5% would have required a far larger sample size. An accurate estimate of the required sample size is complicated by uncertainty about test sensitivity and the underlying infection dynamics of scrapie. A pre-requisite for any future abattoir survey is validation of the diagnostic tests used in relation to both stage of incubation and genotype. Sampling in the <15-month age group was of no value in this survey because the diagnostic tests used were thought to be ineffective in most of the animals in this age group.

Abattoirs↗

[The Madrid autonomous community epidemiological bulletin. A survey on its dissemination and opinion thereof on among primary care physicians for the year 2000].

BACKGROUND: The Autonomous Community of Madrid Epidemiological Bulletin is the main communications link between epidemiological monitoring system and health care professionals. The purpose of this study is that of ascertaining the dissemination and opinion of this Autonomous Community of Madrid Epidemiological Bulletin among primary care physicians for the purpose of adapting this publication to its readers' interests. METHOD: A telephone survey among primary care physicians in the Autonomous Community of Madrid, asking how often they read the Bulletin, the interest and usefulness of the information included in it. The sample size was estimated at 346 physicians. A two-stage sampling process was carried out-by cluster sampling in the first stage, randomly selecting 125 health care centers and 2.7 physicians per center, 17% being primary care team coordinators. A comparison is made of the results among physicians and coordinators by means of the Chi-square and Fisher's Exact Test method, with Epi-Info v.6. RESULTS: A total of 305 surveys were conducted (245 physicians and 60 coordinators). There was an awareness of the existence of the Autonomous Community of Madrid Epidemiological Bulletin on the part of 91.5% (CI 95%: 88.1-94.8), and 27.2% (CI 95%: 21.9-32.5) were familiar with more than 50% of the last issues published. A total of 92.4% (CI 95%: 89.4-95.8) considered the Bulletin to be interesting or highly interesting, grading its usefulness an average of 3.5 on a maximum scale of 5. Of the permanent sections, the most highly-valued was Epidemic Outbreaks, those reports related to meningococcal infection, tuberculosis and HIV/AIDS being the most highly-valued. CONCLUSIONS: The Autonomous Community of Madrid Epidemiological Bulletin is a publication which, although not widely-known by the primary care physicians in the Community, is well-valued when it is read, thus being a useful feedback tool within the Epidemiological Monitoring System.

Adult↗

Empirical Bayes estimation of gene-specific effects in micro-array research.

Micro-array technology allows investigators the opportunity to measure expression levels of thousands of genes simultaneously. However, investigators are also faced with the challenge of simultaneous estimation of gene expression differences for thousands of genes with very small sample sizes. Traditional estimators of differences between treatment means (ordinary least squares estimators or OLS) are not the best estimators if interest is in estimation of gene expression differences for an ensemble of genes. In the case that gene expression differences are regarded as exchangeable samples from a common population, estimators are available that result in much smaller average mean-square error across the population of gene expression difference estimates. We have simulated the application of such an estimator, namely an empirical Bayes (EB) estimator of random effects in a hierarchical linear model (normal-normal). Simulation results revealed mean-square error as low as 0.05 times the mean-square error of OLS estimators (i.e., the difference between treatment means). We applied the analysis to an example dataset as a demonstration of the shrinkage of EB estimators and of the reduction in mean-square error, i.e., increase in precision, associated with EB estimators in this analysis. The method described here is available in software that is available at http://www.soph.uab.edu/ssg.asp?id=1087.

Bayes Theorem↗

Spatial pattern and sequential sampling of squash bug (Heteroptera: Coreidae) adults in watermelon.

Spatial distribution patterns of adult squash bugs were determined in watermelon, Citrullus lanatus (Thunberg) Matsumura and Nakai, during 2001 and 2002. Results of analysis using Taylor's power law regression model indicated that squash bugs were aggregated in watermelon. Taylor's power law provided a good fit with r2 = 0.94. A fixed precision sequential sampling plan was developed for estimating adult squash bug density at fixed precision levels in watermelon. The plan was tested using a resampling simulation method on nine and 13 independent data sets ranging in density from 0.15 to 2.52 adult squash bugs per plant. Average estimated means obtained in 100 repeated simulation runs were within the 95% CI of the true means for all the data. Average estimated levels of precision were similar to the desired level of precision, particularly when the sampling plan was tested on data having an average mean density of 1.19 adult squash bugs per plant. Also, a sequential sampling for classifying adult squash bug density as below or above economic threshold was developed to assist in the decision-making process. The classification sampling plan is advantageous in that it requires smaller sample sizes to estimate the population status when the population density differs greatly from the action threshold. However, the plan may require excessively large sample sizes when the density is close to the threshold. Therefore, an integrated sequential sampling plan was developed using a combination of a fixed precision and classification sequential sampling plans. The integration of sampling plans can help reduce sampling requirements.

Animals↗

Sample size calculations for clinical studies allowing for uncertainty about the variance.

One of the most important steps in the design of a pharmaceutical clinical trial is the estimation of the sample size. For a superiority trial the sample size formula (to achieve a stated power) would be based on a given clinically meaningful difference and a value for the population variance. The formula is typically used as though this population variance is known whereas in reality it is unknown and is replaced by an estimate with its associated uncertainty. The variance estimate would be derived from an earlier similarly designed study (or an overall estimate from several previous studies) and its precision would depend on its degrees of freedom. This paper provides a solution for the calculation of sample sizes that allows for the imprecision in the estimate of the sample variance and shows how traditional formulae give sample sizes that are too small since they do not allow for this uncertainty with the deficiency being more acute with fewer degrees of freedom. It is recommended that the methodology described in this paper should be used when the sample variance has less than 200 degrees of freedom.

Clinical Trials as Topic↗

Sample size determination. Influencing factors and calculation strategies for survey research.

The paper reviews both the influencing factors and calculation strategies of sample size determination for survey research. It indicate the factors that affect the sample size determination procedure and explains how. It also provides calculation methods (including formulas) that can be applied directly and easily to estimate the sample size needed in most popular situations.

Sample Size↗

Variability of food intakes. An analysis of a 12-day data series using persistence measures.

Data from the US Department of Agriculture's Exploratory Study of Longitudinal Measures of Individual Food Intake, conducted in 1982, were used to evaluate individual intakes for day-to-day patterns and to relate these patterns to the reliability of estimated daily energy and nutrient intake means. Generalized least squares estimators incorporating a simple persistence hypothesis showed the importance of including day-to-day patterns in estimating mean daily intake levels. The simple persistence hypothesis was that day-to-day intakes follow a first-order autoregressive process. Results on the accuracy of mean intake estimates and sample size suggested that for the dietary components examined, greatest gains in accuracy of estimated mean daily intake were generally obtained with the first six days of intake data. This conclusion was, however, highly conditioned by the inclusion of the persistence hypothesis in the calculations. It was concluded that additional investigations exploring alternative and perhaps more physiologically justified persistence hypotheses for food intakes would improve the use of sample data in estimating mean daily intakes of diet components.

Adult↗

Sample sizes needed for specified margins of relative error in the estimates of the repeatability and reproducibility standard deviations.

Sample size formulas are developed to estimate the repeatability and reproducibility standard deviations (Sr and S(R)) such that the actual error in (Sr and S(R)) relative to their respective true values, sigmar and sigmaR, are at predefined levels. The statistical consequences associated with AOAC INTERNATIONAL required sample size to validate an analytical method are discussed. In addition, formulas to estimate the uncertainties of (Sr and S(R)) were derived and are provided as supporting documentation. Formula for the Number of Replicates Required for a Specified Margin of Relative Error in the Estimate of the Repeatability Standard Deviation.

Chemistry Techniques, Analytical↗

Calculating sample size bounds for logistic regression.

The calculation of a study's required sample size is one of the most important aspects of the validity of an epidemiological study. Logistic regression often is used in modelling in epidemiology. A simplified method to calculate the sample size for the multiple logistic-regression model was proposed by Hsieh et al. [Stat. Med. 17 (1998) 1623]. The approach for estimating the sample size is described and then applied in the planning of an epidemiological cross-sectional study of the associations of different risk factors with Toxoplasma infection among pregnant women. Although the method demands some additional information which is often difficult to obtain, it is a very useful tool in veterinary epidemiology.

Animals↗

r equivalent: A simple effect size indicator.

The purpose of this article is to propose a simple effect size estimate (obtained from the sample size, N, and a p value) that can be used (a) in meta-analytic research where only sample sizes and p values have been reported by the original investigator, (b) where no generally accepted effect size estimate exists, or (c) where directly computed effect size estimates are likely to be misleading. This effect size estimate is called r(equivalent) because it equals the sample point-biserial correlation between the treatment indicator and an exactly normally distributed outcome in a two-treatment experiment with N/2 units in each group and the obtained p value. As part of placing r(equivalent) into a broader context, the authors also address limitations of r(equivalent).

Analysis of Variance↗

Problems in design of stroke treatment trials.

Critical evaluation of the literature was use to identify remediable flaws in the design of clinical trials of stroke treatment. Trials of dexamethasone, dextran, and glycerol were reviewed. Available studies have in common major weaknesses in case selection (failure to exclude arteriolar strokes due to hemorrhage or lacunar infarction), and failure to estimate required sample size. Problems of case selection can be avoided with computerized tomography; the sample size required to show superiority of active treatment over placebo can be estimated using standard formulas. Prognostic stratification is suggested as a method of overcoming problems of unbalanced allocation. Further studies with improved design are required to evaluate the prospects for medical limitation of cerebral infarct size.

Adult↗

Understanding statistical power.

This article provides an introduction to power analysis so that readers have a basis for understanding the importance of statistical power when planning research and interpreting the results. A simple hypothetical study is used as the context for discussion. The concepts of false findings and missed findings are introduced as a way of thinking about type I and type II errors. The primary factors that affect power are described and examples are provided. Finally, examples are presented to demonstrate 2 uses of power analysis, 1 for prospectively estimating the sample size needed to insure finding effects of a known magnitude in a study and 1 for retrospectively estimating power to gauge the likelihood that an effect was missed.

Data Interpretation, Statistical↗

Cluster trials in implementation research: estimation of intracluster correlation coefficients and sample size.

The cluster randomized trial with a concurrent economic evaluation is considered the gold standard evaluative design for the conduct of implementation research evaluating different strategies to promote the transfer of research findings into clinical practice. This has implications for the planning of such studies, as information is needed on the effects of clustering on both effectiveness and efficiency outcomes. This paper describes the design considerations specific to implementation research studies, focusing particularly on the estimation of sample size requirements and on the need for reliable information on intracluster correlation coefficients for both effectiveness and efficiency outcomes.

Cluster Analysis↗

Sample size considerations for establishing clinical bioequivalence of allergen formulations.

Bioequivalence of formulations must be established by proving that the differences between the formulations are within a specified interval according to Equation 1, the Interval Hypothesis. Explicit estimates of sample size determined from Equation 8 and listed in Table 1 are qualitatively larger than those that would be determined from Equation 2, the Hypothesis of No Difference. Equation 8 was derived from the TOST procedure; other valid methods should yield comparable results. In any context, this discussion has illustrated that the failure to demonstrate a difference is not sufficient to demonstrate equivalence, and that a properly powered equivalence study of allergen formulations will generally demand many more than four study subjects.

Allergens↗

Statistical inferences for a twin correlation with multinomial outcomes.

Current methods for statistical analysis of twin studies focus on continuous and dichotomous data, while only limited methodology exists for analysing multinomial data. As a consequence, investigators are often tempted to collapse multinomial data into two categories simply to facilitate the analysis. We address this problem by developing and evaluating two approaches to the assessment of twin correlation for an outcome variable having more than two nominal categories. One method developed is an extension of the goodness-of-fit approach, while the other method is based on large sample normal theory. Procedures for confidence interval construction are developed and compared using Monte Carlo simulation. The results show that either method may be safely used for confidence interval construction provided the number of twin pairs is large (> or =100) but that in smaller sample sizes the goodness-of-fit procedure is to be preferred on the grounds of validity. Other inference problems are also discussed, including point estimation, hypothesis testing and sample size estimation. An example is included.

Computer Simulation↗

Sample size for cluster randomized trials: effect of coefficient of variation of cluster size and analysis method.

BACKGROUND: Cluster randomized trials are increasingly popular. In many of these trials, cluster sizes are unequal. This can affect trial power, but standard sample size formulae for these trials ignore this. Previous studies addressing this issue have mostly focused on continuous outcomes or methods that are sometimes difficult to use in practice. METHODS: We show how a simple formula can be used to judge the possible effect of unequal cluster sizes for various types of analyses and both continuous and binary outcomes. We explore the practical estimation of the coefficient of variation of cluster size required in this formula and demonstrate the formula's performance for a hypothetical but typical trial randomizing UK general practices. RESULTS: The simple formula provides a good estimate of sample size requirements for trials analysed using cluster-level analyses weighting by cluster size and a conservative estimate for other types of analyses. For trials randomizing UK general practices the coefficient of variation of cluster size depends on variation in practice list size, variation in incidence or prevalence of the medical condition under examination, and practice and patient recruitment strategies, and for many trials is expected to be approximately 0.65. Individual-level analyses can be noticeably more efficient than some cluster-level analyses in this context. CONCLUSIONS: When the coefficient of variation is <0.23, the effect of adjustment for variable cluster size on sample size is negligible. Most trials randomizing UK general practices and many other cluster randomized trials should account for variable cluster size in their sample size calculations.

Cluster Analysis↗

Quantitative genetics of plastron shape in slider turtles (Trachemys scripta).

Shape variation is widespread in nature and embodies both a response to and a source for evolution and natural selection. To detect patterns of shape evolution, one must assess the quantitative genetic underpinnings of shape variation as well as the selective environment that the organisms have experienced. Here we used geometric morphometrics to assess variation in plastron shell shape in 1314 neonatal slider turtles (Trachemys scripta) from 162 clutches of laboratory-incubated eggs from two nesting areas. Multivariate analysis of variance indicated that nesting area has a limited role in describing plastron shape variation among clutches, whereas differences between individual clutches were highly significant, suggesting a prominent clutch effect. The covariation between plastron shape and several possible maternal effect variables (yolk hormone levels and egg dimensions) was assessed for a subset of clutches and found to be negligible. We subsequently employed several recently proposed methods for estimating heritability from shape variables, and generalized a univariate approach to accommodate unequal sample sizes. Univariate estimates of shape heritability based on Procrustes distances yielded large values for both nesting populations (h2 approximately 0.86), and multivariate estimates of maximal additive heritability were also large for both nesting populations (h2max approximately 0.57). We also estimated the dominant trend in heritable shape change for each nesting population and found that the direction of shape evolution was not the same for the two sites. Therefore, although the magnitude of shape evolution was similar between nesting populations, the manner in which plastron shape is evolving is not. We conclude that the univariate approach for assessing quantitative genetic parameters from geometric morphometric data has limited utility, because it is unable to accurately describe how shape is evolving.

Animals↗

Approaches for sampling the twospotted spider mite (Acari: Tetranychidae) on clementines in Spain.

Tetranychus urticae Koch (Acari: Tetranychidae) is an important pest of clementine mandarins, Citrus reticulata Blanco, in Spain. As a first step toward the development of an integrated crop management program for clementines, dispersion patterns of T. urticae females were determined for different types of leaves and fruit. The study was carried out between 2001 and 2003 in different commercial clementine orchards in the provinces of Castelló and Tarragona (northeastern Spain). We found that symptomatic leaves (those exhibiting typical chlorotic spots) harbored 57.1% of the total mite counts. Furthermore, these leaves were representative of mite dynamics on other leaf types. Therefore, symptomatic leaves were selected as a sampling unit. Dispersion patterns generated by Taylor's power law demonstrated the occurrence of aggregated patterns of spatial distribution (b > 1.21) on both leaves and fruit. Based on these results, the incidence (proportion of infested samples) and mean density relationship were developed. We found that optimal binomial sample sizes for estimating low populations of T. urticae on leaves (up to 0.2 female per leaf) were very large. Therefore, enumerative sampling would be more reliable within this range of T. urticae densities. However, binomial sampling was the only valid method for estimating mite density on fruit.

Animals↗