Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Sample size considerations for superiority trials in systemic lupus erythematosus (SLE).

For reasons of efficiency and ethics, sample size calculations are an important part of the design of all clinical trials. This paper highlights the statistical issues inherent to the estimation of sample size requirements in superiority trials particular to SLE. Calculations based on statistical power for testing hypotheses have historically been the method of choice for sample size determination in clinical trials. The advantages of using confidence intervals (CI's) rather than P-values in reporting results of clinical trials is now well established. Since the design of a trial should match the analysis that will eventually be performed, sample size methods based on ensuring accurate estimation of important parameters via sufficiently narrow CI widths should be preferred to methods based on hypothesis testing. Methods and examples are given for sample size calculations for continuous and dichotomous outcomes from both a power and confidence interval width viewpoint. An understanding of sample size calculations in association with expert statistical consultation will result in better designed clinical trials that accurately estimate clinically relevant differences between treatment outcomes, thereby furthering the treatment of patients with SLE.

Confidence Intervals↗

Sample size-based indication of normality in lognormally distributed populations.

Occupational and environmental hygiene sampling strategies are usually dictated by factors that limit sample sizes to relatively small numbers. Often, parameters estimated from small sample sizes are then used to make further estimates of the occurrence of extreme events, which are governed by the underlying exposure distribution. We investigated the limitations superimposed by the number of samples in distinguishing an asymmetric (Lognormal) distribution through the rejection of a hypothesized symmetric (Normal) distribution. Sets of 5 to 250 synthetic samples from underlying Lognormal distributions with unit median were generated for 24 separate geometric standard deviations (GSDs), ranging from 1.25 to 7.00. Each simulated combination was repeated in blocks of 200 and each block was repeated tenfold. The synthetic samples were then tested for goodness of fit for Normality by using the Shapiro and Wilk's W Test. Results indicated that the number of samples required to distinguish between Normal and Lognormal distributions was inversely related to GSD. When GSD = 1.25, 169 samples were required for 90 percent distinction at alpha = 0.05. The criteria for success for GSD of 2.00 and 4.00 were 25 and 15 samples, respectively. These results led to the conclusion that the general inability to distinguish an underlying distribution may impose serious difficulties in the estimation of extreme events associated with occupational and environmental hygiene-related sampling.

Environmental Monitoring↗

Estimating morbidity risks with variable age of onset: review of methods and a maximum likelihood approach.

Various methods for the estimation of morbidity risk in a disease with late variable onset are described, along with a maximum likelihood approach. It is shown that the Strömgren estimator is nearly as efficient as the maximum likelihood estimator when the true risk is low, but may be significantly less efficient for high morbidity risks. The maximum likelihood estimator offers greater protection against risk estimates greater than or equal to 1, and for small samples, may also be less biased than the Strömgren estimator, especially when risk is high. For reasonable sample sizes both estimates are nearly unbiased. The modified Strömgren estimator is too biased, in general, to be practical. Methods for comparing morbidity risks are also described. If an age-of-onset distribution is estimated from the same sample as morbidity risk, a single maximum likelihood procedure is advocated. The methods are applied to a data set on major affective disorder. The sensitivity of morbidity-risk estimates and tests of hypotheses to the form of onset function assumed is examined.

Age Factors↗

Cluster versus individual randomization in adolescent tobacco and alcohol studies: illustrations for design decisions.

BACKGROUND: The decision to randomize by clusters of subjects such as a classroom or clinic versus individual randomization where some contamination may occur is examined within the framework of sample size issues. Estimates for background rates and intraclass correlations are also provided for adolescent tobacco and alcohol outcomes derived from a recent study using cluster randomization. METHODS: A ratio of adjusted sample sizes is derived which is a function of the intraclass correlation and cluster size for cluster randomization and total amount of contamination for individual randomization. Using estimated incidence rates and intraclass correlations, we provide a comparison of sample sizes for two plausible study outcomes. RESULTS: Small clusters such as a family or small classroom tend to have stronger within cluster dependence and cluster randomization would be clearly favoured over individual randomization. For moderately sized clusters, if contamination levels are likely to be high then cluster randomization would be a better choice. However in some situations where lower levels of contamination are expected, individual randomization may be preferred. With larger clusters, individual randomization should be considered when contamination rates are expected to be low. CONCLUSIONS: Investigators must carefully consider the choice of cluster randomization versus individual randomization in the context of likely contamination. In this paper we provided a basis for making this decision as well as examples to illustrate these decisions, and parameter estimates that will be especially useful for investigators in adolescent tobacco and alcohol studies.

Adolescent↗

Implications of genetic traits on vaccine efficacy.

Literature on genetic screening in the community suggests that people having specific genotypes may either get or protect from infection, for example, malaria, human papilloma virus, and haemophilic influenza, for which vaccines are either already developed or being targeted. In such a situation, the evaluation of the efficacy of vaccine in the community needs to be examined with caution. In this paper, I present a method for the estimation of vaccine efficacy (VE) in the presence of genetic traits/component (theta) and the sample size required to estimate the 95 per cent CI with a given relative width for the estimated vaccine efficacy. Considering true efficacy ranging from 40 to 80 per cent and the possible values of the genetic component (theta) ranging from 0 to 60 per cent, the VE was estimated. The 95 per cent confidence intervals (CI) for the estimated VE for relative widths (R) 1.0 and 0.1 were computed. The sample sizes required for each of the unvaccinated and vaccinated cohorts were computed for estimating the 95 per cent CI for given incidence rates in the unvaccinated (Iu) cohort. In the presence of genetic traits I found that the VE was consistently overestimated. There existed change in the location as well as the asymmetry of the 95 per cent CIs over the point estimate of VE. The sample size required for estimating 95 per cent CI of VE was substantially reduced, resulting in savings. The more the genetic component (theta) affecting disease in the community, the more the savings in sample size. I examined the above estimators for (i) VE, (ii) 95 per cent CI for VE and (iii) sample size required for estimating 95 per cent CI of VE using the real-life data from the Haemophilus influenzae type b vaccine trial conducted in Finland and the global genetic structure of encapsulated H. influenza. Because of escalated VE and large savings in sample size for estimating the 95 per cent CI for VE, I recommend that the design should consider the genetic component that causes/protects from infection/disease for the evaluation of efficacy of vaccine in the field.

Bacterial Capsules↗

Power and sample size calculations for clinical trials of myofascial pain of jaw muscles.

When a clinical trial is planned, the approximate number of subjects needed for significant differences between/among groups to be detected must be estimated. Sample-size calculations provide the investigator with this information. This paper discusses the choice of outcome measures and describes the steps used to estimate the numbers of subjects necessary for a study that compares treatments for patients with chronic myofascial pain of jaw muscles. Within- and between-subject variances were estimated for the chosen variables, the subjects' pain ratings on visual analogue scales. Sample sizes were then calculated for theoretical differences between groups by pre-treatment means and overall standard deviations (Cohen, 1977). The results of this analysis can be used by other researchers when planning studies involving these types of patients.

Adolescent↗

A two one-sided tests procedure for assessment of individual bioequivalence.

In this paper we propose a two one-sided tests procedure for assessment of individual bioequivalence based on the concept of individual equivalence ratios proposed by Anderson and Hauck. The proposed procedure is derived under the normality assumption for the logarithmic transformation of pharmacokinetic responses obtained from a standard two-sequence, two-period crossover design. We show that the hypotheses for individual bioequivalence are equivalent to the hypotheses for testing whether the upper (or lower) pth quantile of the distribution of the differences between the test and reference formulations from the same subject is not greater (or not smaller) than some prespecified equivalence limits. Under this setting, we examine the relationship between average and individual bioequivalence. There exists the uniformly most powerful invariant test for each of the two one-sided hypotheses. In addition, the proposed two one-sided tests procedure is a test of size alpha (i.e., < or = alpha). We demonstrate that the determination of critical values, the enumeration of power, and the estimation of sample sizes requires noncentral t-distributions but does not necessarily require the estimation of unknown population mean and variance for noncentrality parameters. We discuss possible extensions to other crossover and replicated crossover designs. A numerical example illustrates the proposed procedure.

Biotransformation↗

Sampling methods, dispersion patterns, and fixed precision sequential sampling plans for western flower thrips (Thysanoptera: Thripidae) and cotton fleahoppers (Hemiptera: Miridae) in cotton.

A 2-yr field study was conducted to examine the effectiveness of two sampling methods (visual and plant washing techniques) for western flower thrips, Frankliniella occidentalis (Pergande), and five sampling methods (visual, beat bucket, drop cloth, sweep net, and vacuum) for cotton fleahopper, Pseudatomoscelis seriatus (Reuter), in Texas cotton, Gossypium hirsutum (L.), and to develop sequential sampling plans for each pest. The plant washing technique gave similar results to the visual method in detecting adult thrips, but the washing technique detected significantly higher number of thrips larvae compared with the visual sampling. Visual sampling detected the highest number of fleahoppers followed by beat bucket, drop cloth, vacuum, and sweep net sampling, with no significant difference in catch efficiency between vacuum and sweep net methods. However, based on fixed precision cost reliability, the sweep net sampling was the most cost-effective method followed by vacuum, beat bucket, drop cloth, and visual sampling. Taylor's Power Law analysis revealed that the field dispersion patterns of both thrips and fleahoppers were aggregated throughout the crop growing season. For thrips management decision based on visual sampling (0.25 precision), 15 plants were estimated to be the minimum sample size when the estimated population density was one thrips per plant, whereas the minimum sample size was nine plants when thrips density approached 10 thrips per plant. The minimum visual sample size for cotton fleahoppers was 16 plants when the density was one fleahopper per plant, but the sample size decreased rapidly with an increase in fleahopper density, requiring only four plants to be sampled when the density was 10 fleahoppers per plant. Sequential sampling plans were developed and validated with independent data for both thrips and cotton fleahoppers.

Animals↗

Intra-cluster correlation coefficients in adults with diabetes in primary care practices: the Vermont Diabetes Information System field survey.

BACKGROUND: Proper estimation of sample size requirements for cluster-based studies requires estimates of the intra-cluster correlation coefficient (ICC) for the variables of interest. METHODS: We calculated the ICC for 112 variables measured as part of the Vermont Diabetes Information System, a cluster-randomized study of adults with diabetes from 73 primary care practices (the clusters) in Vermont and surrounding areas. RESULTS: ICCs varied widely around a median value of 0.0185 (Inter-quartile range: 0.006, 0.037). Some characteristics (such as the proportion having a recent creatinine measurement) were highly associated with the practice (ICC = 0.288), while others (prevalence of some comorbidities and complications and certain aspects of quality of life) varied much more across patients with only small correlation within practices (ICC<0.001). CONCLUSION: The ICC values reported here may be useful in designing future studies that use clustered sampling from primary care practices.

Adult↗

Determining sample size and power in clinical trials: the forgotten essential.

Estimation of the sample size is a fundamental but usually ignored requirement of the randomized controlled trial (RCT). Indeed, the publication of small trials without consideration of sample size is worrisome from both medical and ethical viewpoints. Type II errors are common, and readers and investigators may reject worthwhile treatments and interventions. Before embarking on an RCT, the investigator must choose an alpha, a beta, and the rates of outcomes anticipated in both treatment groups. This should reflect the characteristics of the condition and its treatment. If limited sample size or available resources pose a problem, the use of continuous outcome measures, paired before-after measurements, and more common outcome measures can minimize sample size requirements. If these approaches are not satisfactory, then a multicenter trial may be in order.

Clinical Trials as Topic↗

Interval estimation and optimal design for the within-subject coefficient of variation for continuous and binary variables.

BACKGROUND: In this paper we propose the use of the within-subject coefficient of variation as an index of a measurement's reliability. For continuous variables and based on its maximum likelihood estimation we derive a variance-stabilizing transformation and discuss confidence interval construction within the framework of a one-way random effects model. We investigate sample size requirements for the within-subject coefficient of variation for continuous and binary variables. METHODS: We investigate the validity of the approximate normal confidence interval by Monte Carlo simulations. In designing a reliability study, a crucial issue is the balance between the number of subjects to be recruited and the number of repeated measurements per subject. We discuss efficiency of estimation and cost considerations for the optimal allocation of the sample resources. The approach is illustrated by an example on Magnetic Resonance Imaging (MRI). We also discuss the issue of sample size estimation for dichotomous responses with two examples. RESULTS: For the continuous variable we found that the variance stabilizing transformation improves the asymptotic coverage probabilities on the within-subject coefficient of variation for the continuous variable. The maximum like estimation and sample size estimation based on pre-specified width of confidence interval are novel contribution to the literature for the binary variable. CONCLUSION: Using the sample size formulas, we hope to help clinical epidemiologists and practicing statisticians to efficiently design reliability studies using the within-subject coefficient of variation, whether the variable of interest is continuous or binary.

Algorithms↗

Strategies for selecting subsets of single-nucleotide polymorphisms to genotype in association studies.

In genetic association studies, linkage disequilibrium (LD) within a region can be exploited to select a subset of single-nucleotide polymorphisms (SNPs) to genotype with minimal loss of information. A novel entropy-based method for selecting SNPs is proposed and compared to an existing method based on the coefficient of determination (R2) using simulated data from Genetic Analysis Workshop 14. The effect of the size of the sample used to investigate LD (by estimating haplotype frequencies) and hence select the SNPs is also investigated for both measures. It is found that the novel method and the established method select SNP subsets that do not differ greatly. The entropy-based measure may thus have value because it is easier to compute than R2. Increasing the sample size used to estimate haplotype frequencies improves the predictive power of the subset of SNPs selected. A smaller subset of SNPs chosen using a large initial sample to estimate LD can in some instances be more informative than a larger subset chosen based on poor estimates of LD (using a small initial sample). An initial sample size of 50 individuals is sufficient in most situations investigated, which involved selection from a set of 7 SNPs, although to select a larger number of SNPs, a larger initial sample size may be required.

Genetic Loci↗

Linkage analysis of complex traits using affected sibpairs: effects of single-locus approximations on estimates of the required sample size.

We investigated the power of the affected sibpair method for detecting a disease locus when the disease is inherited through two bi-allelic loci. The power was computed for all possible values of the gene frequencies and penetrances that lead to a given population prevalence and a given sibling relative risk. A method to generate rapidly all possible models that give a specific population prevalence and relative risk is provided. We applied it to the case of a two-locus disease with a prevalence of 10% and a low sibling relative risk of 1.5. For this particular example, regardless of the true underlying model, a sample size (N = 450 for alpha = 0.05, N = 1,500 for alpha = 0.0001) may be determined such that one would expect enough power (0.80) to detect at least one of the two disease genes. In addition to the general case, we examined a special class of models in which the marginal penetrances at each locus are either recessive or dominant. In this instance, the gene frequencies were excellent predictors of the power afforded by a particular sample size. These methods have been implemented in a C program called SIBPOWER which is freely available from the first author. With this program, investigators can perform their own power calculations for any two-locus model of their choice thus avoiding the need to use single-locus approximations that may grossly underestimate the necessary sample size.

Gene Frequency↗

Sample size requirements for studies estimating odds ratios or relative risks.

This paper presents formulae for determining the number of subjects necessary, in either a case-control or a cohort study, to estimate the odds ratio or relative risk, respectively, to within a selected percentage (epsilon) of the true population value with some specified probability. This approach differs somewhat from previous comparable work that estimated the log odds ratio within a stated fixed distance rather than as a percentage of the actual odds ratio. Comparable development for relative risk has not previously appeared in the literature. These formulae provide guidelines for determination of study size that does not depend on hypothesis testing considerations.

Epidemiologic Methods↗

[Criterion quality and estimation of the expected sample size in sequential analysis of linkage].

It is shown that, when conducting sequential testing of linkage in pedigree samples, (1) type I and type II errors observed are less than expected and (2) the generally accepted method for determining the average sample size, E(N), required for sequential analysis of linkage, underestimates it. A less biased approximation of E(N) is proposed. A wide scattering of actual sample sizes required for completion of sequential analysis is demonstrated, which puts practical use of E(N) into question.

Evaluation Studies as Topic↗

A method for determining the size of internal pilot studies.

The assessment of sample size in clinical trials comparing population means requires a variance estimate of the main efficacy variable. When this variance estimate has a low precision, it may be appropriate to use the data from the first patients entered in the trial ('internal pilot study') to estimate the sample size. We suggest a method for determining the size of internal pilot studies, which aims at ensuring that this size is as large as possible, but not larger than 'the optimal size' of the planned study. Advantages and limitations of the method are discussed.

Bias↗

Calibrating a molecular clock from phylogeographic data: moments and likelihood estimators.

We present moments and likelihood methods that estimate a DNA substitution rate from a group of closely related sister species pairs separated at an assumed time, and we test these methods with simulations. The methods also estimate ancestral population size and can test whether there is a significant difference among the ancestral population sizes of the sister species pairs. Estimates presented in the literature often ignore the ancestral coalescent prior to speciation and therefore should be biased upward. The simulations show that both methods yield accurate estimates given sample sizes of five or more species pairs and that better likelihood estimates are obtained if there is no significant difference among ancestral population sizes. The model presented here indicates that the larger than expected variation found in multitaxa datasets can be explained by variation in the ancestral coalescence and the Poisson mutation process. In this context, observed variation can often be accounted for by variation in ancestral population sizes rather than invoking variation in other parameters, such as divergence time or mutation rate. The methods are applied to data from two groups of species pairs (sea urchins and Alpheus snapping shrimp) that are thought to have separated by the rise of Panama three million years ago.

Animals↗

Sample sizes for identifying the key types of container occupied by dengue-vector pupae: the use of entropy in analyses of compositional data.

A method has been developed for estimating the sample sizes needed to identify categories that comprise a large proportion of a compositional data-set. The method is to be used in the design of surveys of mosquito pupae, for identifying the key container types from which the majority of adult dengue vectors emerge. Although a finite-population correction was devised for estimating the mean of a negative binomial distribution, other complications of parametric approaches make them unlikely to yield methods simple enough to be practically applicable. The Shannon-Wiener index was therefore investigated as a more useful alternative, at the cost of theoretical generalizability, in an approach based on re-sampling methods in conjunction with the use of entropy. This index can be used to summarize the degree to which pupae are either concentrated in a few container types, or dispersed among many. An empirical relationship between the index and the repeatability of surveys of differing sample sizes was observed. A step-wise rule, based on the entropy of the cumulative data, was devised for determining the sample size, in terms of the number of houses positive for pupae, at which a pupal survey might reasonably be stopped.

Aedes↗