Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Determining dislocation cell sizes for high-strain deformation microstructures using the EBSP technique.

The effect of several data collection and processing choices has been examined for high-resolution electron back-scatter pattern (EBSP) investigation of a highly deformed sample. The results were compared with a transmission electron microscope (TEM) investigation of the same sample. The estimated dislocation cell size was examined as a function of data cleaning strategy, line intercept vs. reconstruction method, critical misorientation angle definition and step-size. The best agreement with the TEM results was obtained using a modified relative reconstruction algorithm on fine step-size maps allowing some of the noise in the data to be overcome. Step sizes of up to one-quarter the average cell size yielded similar values for the estimated average cell size. As a result of the mixture of both high- and low-angle boundaries, single diffraction condition TEM images may give larger cell size estimates than the EBSP data. Orientation noise in the EBSP data, however, still limits the extent to which quantitative information can be extracted.

Journal Article↗

The effect of different sensitivity, specificity and cause-specific mortality fractions on the estimation of differences in cause-specific mortality rates in children from studies using verbal autopsies.

BACKGROUND: Verbal autopsies (VA) are increasingly being used in developing countries to determine causes of death, but little attention is generally given to the misclassification effects of the VA. This paper considers the effect of misclassification on the estimation of differences in cause-specific mortality rates between two populations. METHODS: The bias in the percentage difference in cause-specific mortality between two populations has been explored under two different models: i) assuming that mortality from all other causes does not differ between the two populations; ii) allowing for a difference in mortality from all other causes. The bias is described in terms of the sensitivity and specificity of the VA diagnosis and the proportion of mortality due to the cause of interest. Methods for adjustment of sample size and adjusting the estimate of effect are also explored. RESULTS: The results are illustrated for a range of plausible values for these parameters. The bias is more extreme as both sensitivity and specificity fall, and is particularly affected even by a small loss of specificity. The bias also increases as the proportion of all deaths due to the cause of interest decreases, and is affected by the size of the true change in mortality due to the cause of interest relative to the change in mortality from other causes. Calculations from existing data suggest prohibitively large sample sizes may often be required to detect important differences in cause-specific mortality rates in studies using existing VA. CONCLUSIONS: Highly specific VA tools are needed before observed differences in cause-specific mortality can be interpreted. Loss of power due to misclassification may obscure real differences in cause-specific mortality.

Autopsy↗

Trend tests for case-control studies of genetic markers: power, sample size and robustness.

The Cochran-Armitage trend test is commonly used as a genotype-based test for candidate gene association. Corresponding to each underlying genetic model there is a particular set of scores assigned to the genotypes that maximizes its power. When the variance of the test statistic is known, the formulas for approximate power and associated sample size are readily obtained. In practice, however, the variance of the test statistic needs to be estimated. We present formulas for the required sample size to achieve a prespecified power that account for the need to estimate the variance of the test statistic. When the underlying genetic model is unknown one can incur a substantial loss of power when a test suitable for one mode of inheritance is used where another mode is the true one. Thus, tests having good power properties relative to the optimal tests for each model are useful. These tests are called efficiency robust and we study two of them: the maximin efficiency robust test is a linear combination of the standardized optimal tests that has high efficiency and the MAX test, the maximum of the standardized optimal tests. Simulation results of the robustness of these two tests indicate that the more computationally involved MAX test is preferable.

Case-Control Studies↗

Exact test size and power of a Gaussian error linear model for an internal pilot study.

Wittes and Brittain recommended using an 'internal pilot study' to adjust sample size. The approach involves five steps in testing a general linear hypothesis for a general linear univariate model, with Gaussian errors. First, specify the design, hypothesis, desired test size, power, a smallest 'clinically meaningful' effect, and a speculated error variance. Second, conduct a power analysis to choose provisionally a planned sample size. Third, collect a specified proportion of the planned sample as the internal pilot sample, and estimate the variance (but do not test the hypothesis). Fourth, update the power analysis with the variance estimate to adjust the total sample size. Fifth, finish the study and test the hypothesis with all data. We describe methods for computing exact test size and power under this scenario. Our analytic results agree with simulations of Wittes and Brittain. Furthermore, our exact results apply to any general linear univariate model with fixed predictors, which is much more general than the two-sample t-test considered by Wittes and Brittain. In addition, our results allow for examination of the impact on test size of internal pilot studies for more complicated designs in the framework of the general linear model. We examine the impact of (i) small samples, (ii) allowing the planned sample size to decrease, (iii) the choice of internal pilot sample size, and (iv) the maximum allowable size of the second sample. All affect test size, power and expected total sample size. We present a number of examples including one that uses an internal pilot study in a three-group analysis of variance.

Analysis of Variance↗

Comparison of adjusted attributable risk estimators.

The estimation of attributable risk in the presence of confounding and effect modification is studied in this paper. Different adjustment methods for the attributable risk are reviewed. The results of a stimulation study comparing these methods under the unrestricted multinomial sampling model are reported. From the study it is concluded that the maximum likelihood estimator resulting from the 'case load weighting' of the stratum-specific attributable risk estimates will be the best overall choice in all practical situations with relatively large sample sizes. Other adjusted attributable risk estimators depend heavily on the underlying structure of the multinomial model.

Humans↗

The effect of misclassification of disease status in follow-up studies: implications for selecting disease classification criteria.

For many diseases, a set of diagnostic criteria with perfect sensitivity and specificity does not exist. In the design of a follow-up study of such a disease, one often has a choice between using a set of narrow classification criteria for the disease outcome (i.e., a test with relatively high specificity and relatively low sensitivity) or a broader set of criteria (i.e., a more sensitive, less specific test). A model was investigated which simulated choices one may have between disease classification tests, to determine how the required sample size and bias in the estimates of the risk ratio and risk difference varied between tests. A two-sample study with nondifferential misclassification of disease outcome was assumed. Based on the model, the bias in the risk ratio increases as one increases the sensitivity of the diagnostic test at the expense of specificity. Conversely, the bias in the risk difference decreases with increasing sensitivity and declining specificity. The required sample size is minimized at relatively high sensitivity and relatively low specificity. Selection of the disease classification test as that at which the required sample size is minimized could reduce some of the large data collection costs of follow-up studies. The advantages and limitations of applying this technique to actual studies are discussed.

Clinical Trials as Topic↗

Different disease rates in two populations: how much is due to differences in risk factors?

Two populations with different disease rates may differ in their risk factors for the disease. If so, it is desirable to know what proportion of the disease excess in the high-risk population is attributable to its greater exposure to the risk factors. This proportion has been called the relative attributable risk (RAR). A related measure is the adjusted relative risk (ARR), defined as the ratio of rates in high-risk to low-risk populations that would be observed if the distribution of risk factors in the high-risk population equalled that of the low-risk population. We present methods for obtaining consistent estimates and asymptotic confidence intervals for both the RAR and the ARR using data from case-control studies in the two populations. The methods are applied to the problem of estimating the differences in ovarian cancer incidence between U.S. white women (high-risk) and U.S. black women (low-risk) attributable to differences in reproductive risk factors. Simulations show that the methods perform well; however, when the true RAR is close to 0 or 1 or when sample sizes are small, RAR estimates may fall outside the unit interval. We discuss circumstances when the true RAR lies outside the unit interval; in such circumstances the ARR is easier to interpret.

Black or African American↗

Degrees of freedom in interspecific allometry: an adjustment for the effects of phylogenetic constraint.

The data used in studies of bivariate interspecific allometry usually violate the assumption of statistical independence. Although the traits of each species are commonly treated as independent, the expression of a trait among species within a genus may covary because of shared common ancestry. The same effect exists for genera within a family and so on up the phylogenetic hierarchy. Determining sample size by counting data points overestimates the effective sample size, which then leads to overestimating the degrees of freedom that should be used in calculating probabilities and confidence intervals. This results in an inflated Type 1 error rate. Although some workers (e.g., Felsenstein [1985] Am. Nat. 125:1-15) have suggested that this issue may invalidate interspecific allometry as a comparative method, a correction for the problem can be approximated with variance components from a nested analysis of variance. Variance components partition the total variation in the data set among the levels of the nested hierarchy. If the variance component for each nested level is weighted by the number of groups at that level, the sum of these values is an estimate of an effective sample size for the data set which reflects the effects of phylogenetic constraint. Analysis of two data sets, using taxonomy to define levels of the nested hierarchy, suggests that it has been common for published studies of interspecific allometry to severely overestimate the number of degrees of freedom. Interspecific allometry remains an important comparative method for evaluating questions concerning individual species that are not similarly addressed by the format of most of the newer comparative methods. With the correction proposed here for estimating degrees of freedom, the major statistical weakness of the procedure is substantially reduced.

Analysis of Variance↗

Joint versus separate estimation of state and change in category frequencies from repeat stratified two-phase sampling with a fallible classifier.

Joint maximum likelihood estimates (JML) of category frequencies and change from repeat stratified two-phase sampling surveys with a fallible classifier are often seriously biased and have large root mean square errors when they are obtained for small populations (<5,000) with three or more categories and a moderate to small phase II sample size (<1,000). JML estimates of state also depend on antecedent or posterior data, a recipe for inconsistency. In these situations, a separate maximum likelihood estimation (SML) of category frequencies at each survey date appears preferable. SML estimates of net change are obtained as the difference in states. SML standard errors of change are obtained via an estimate of the temporal correlation and variances of state. A bivariate binary logistic model of change provided the estimate of temporal correlation. SML generally outperformed JML significantly in terms of bias and root mean square errors in eight case studies.

Data Collection↗

Sample size determination for studies of gene-environment interaction.

BACKGROUND: The search for interaction effects is common in epidemiological studies, but the power of such studies is a major concern. This is a practical issue as many future studies will wish to investigate potential gene-gene and gene-environment interactions and therefore need to be planned on the basis of appropriate sample size calculations. METHODS: The underlying model considered in this paper is a simple linear regression and relating a continuous outcome to a continuously distributed exposure variable. RESULTS: The slope of the regression line is taken to be dependent on genotype, and the ratio of the slopes for each genotype is considered as the interaction parameter. Sample size is affected by the allele frequency and whether the genetic model is dominant or recessive. It is also critically dependent upon the size of the association between exposure and outcome, and the strength of the interaction term. The link between these determinants is graphically displayed to allow sample size and power to be estimated. An example of the analysis of the association between physical activity and glucose intolerance demonstrates how information from previous studies can be used to determine the sample size required to examine gene-environment interactions. CONCLUSIONS: The formulae allowing the computation of the sample size required to study the interaction between a continuous environmental exposure and a genetic factor on a continuous outcome variable should have a practical utility in assisting the design of studies of appropriate power.

Effect Modifier, Epidemiologic↗

Sample size for multiple regression: obtaining regression coefficients that are accurate, not simply significant.

An approach to sample size planning for multiple regression is presented that emphasizes accuracy in parameter estimation (AIPE). The AIPE approach yields precise estimates of population parameters by providing necessary sample sizes in order for the likely widths of confidence intervals to be sufficiently narrow. One AIPE method yields a sample size such that the expected width of the confidence interval around the standardized population regression coefficient is equal to the width specified. An enhanced formulation ensures, with some stipulated probability, that the width of the confidence interval will be no larger than the width specified. Issues involving standardized regression coefficients and random predictors are discussed, as are the philosophical differences between AIPE and the power analytic approaches to sample size planning.

Humans↗

Economic evaluation of aquatic exercise for persons with osteoarthritis.

OBJECTIVES: To estimate cost and outcomes of the Arthritis Foundation aquatic exercise classes from the societal perspective. DESIGN: Randomized trial of 20-week aquatic classes. Cost per quality-adjusted life year (QALY) gained was estimated using trial data. Sample size was based on 80% power to reject the null hypothesis that the cost/QALY gained would not exceed $50,000. SUBJECTS AND METHODS: Recruited 249 adults from Washington State aged 55 to 75 with a doctor-confirmed diagnosis of osteoarthritis to participate in aquatic classes. The Quality of Well-Being Scale (QWB) and Current Health Desirability Rating (CHDR) were used for economic evaluation, supplemented by the arthritis-specific Health Assessment Questionnaire (HAQ), Center for Epidemiologic Studies-Depression Scale (CES-D), and Perceived Quality of Life Scale (PQOL) collected at baseline and postclass. Outcome results applied to life expectancy tables were used to estimate QALYs. Use of health care facilities was assessed from diaries/questionnaires and Medicare reimbursement rates used to estimate costs. Nonparametric bootstrap sampling of costs/QALY ratios established the 95% CI around the estimates. RESULTS: Aquatic exercisers reported equal (QWB) or better (CHDR, HAQ, PQOL) health-related quality of life compared with controls. Outcomes improved with regular class attendance. Costs/QALY gained discounted at 3% were $205,186 using the QWB and $32,643 using the CHRD. CONCLUSION: Aquatic exercise exceeded $50,000 per QALY gained using the community-weighted outcome but fell below this arbitrary budget constraint when using the participant-weighted measure. Confidence intervals around these ratios suggested wide variability of cost effectiveness of aquatic exercise.

Aged↗

Quantification of variability and uncertainty for censored data sets and application to air toxic emission factors.

Many environmental data sets, such as for air toxic emission factors, contain several values reported only as below detection limit. Such data sets are referred to as "censored." Typical approaches to dealing with the censored data sets include replacing censored values with arbitrary values of zero, one-half of the detection limit, or the detection limit. Here, an approach to quantification of the variability and uncertainty of censored data sets is demonstrated. Empirical bootstrap simulation is used to simulate censored bootstrap samples from the original data. Maximum likelihood estimation (MLE) is used to fit parametric probability distributions to each bootstrap sample, thereby specifying alternative estimates of the unknown population distribution of the censored data sets. Sampling distributions for uncertainty in statistics such as the mean, median, and percentile are calculated. The robustness of the method was tested by application to different degrees of censoring, sample sizes, coefficients of variation, and numbers of detection limits. Lognormal, gamma, and Weibull distributions were evaluated. The reliability of using this method to estimate the mean is evaluated by averaging the best estimated means of 20 cases for small sample size of 20. The confidence intervals for distribution percentiles estimated with bootstrap/MLE method compared favorably to results obtained with the nonparametric Kaplan-Meier method. The bootstrap/MLE method is illustrated via an application to an empirical air toxic emission factor data set.

Air↗

Comparison of different maximum likelihood estimators in a small sample logistic regression with two independent binary variables.

In order to examine the bias of the estimate of the log odds ratio in a 2 x 2 contingency table, Walter computed the entire distribution of the estimated log odds ratio using various small sample sizes. This is equivalent to computing the distribution of the estimated parameter b1 in a logistic regression with one independent binary variable. In this paper, the distributions of the estimated parameters b1 and b2 for two independent binary variables are computed for some small sample logistic regressions using six different estimation methods based on maximum likelihood. These estimates are then compared to the true parameter values. The best estimation method depends on the frequency of the outcome of interest and on whether the bias or mean square error is considered more important.

Bias↗

Homotypic variation of canine flexor tendons: implications for the design of experimental studies in animal models.

Water, collagen and glycosamimoglycan contents, cross-sectional area, stiffness and elastic modulus were carefully quantitated in flexor digitorum superficialis tendons from mature canines. From these data the within- and between-animal variability was estimated and used to demonstrate sample size calculations for both two-group and paired (within-animal) study designs. The estimated between-dog variance was typically 50% or less of the total variance for the parameters investigated. In other words, the correlation among the tendons within an animal for most measures was not strong. Therefore, for some variables (e.g., elastic modulus) in this animal and tendon model, there is no appreciable gain in statistical power by using a paired study design. A two-group design could be used, but any within-animal correlation must be accounted for in the analysis. For other variables such as collagen content, a paired design would gain substantial power.

Animals↗

dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference charts.

MOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS.

Algorithms↗

Sample size calculations for cluster randomised trials. Changing Professional Practice in Europe Group (EU BIOMED II Concerted Action).

OBJECTIVES: Cluster randomised trials, in which groups of individuals are randomised, are increasingly being used in the health field. Adopting a clustered approach has implications for the design of such trials, and sample size calculations need to be inflated to accommodate for the clustering effect. Reliable estimates of intracluster correlation coefficients (ICCs) are required for robust sample size calculations to be made; however, little empirical evidence is available on their likely size, and on factors which influence their magnitude. The aim of this study was to generate empirical estimates of ICCs and to explore factors which may affect their magnitude. METHODS: Empirical estimates of ICCs were calculated for both process variables and patient outcomes from a number of datasets of primary and secondary care implementation studies. RESULTS: Estimates of ICCs varied according to setting and type of outcome. Estimates of ICCs for process variables were higher than those for patient outcomes, and estimates derived from secondary care were higher than those from primary care. ICCs for process variables in primary care were of the order of 0.05-0.15, whilst those in secondary care were of the order of 0.3. Estimates for patient outcomes in primary care were generally lower than 0.05. CONCLUSIONS: Adopting cluster randomisation has implications for the design, size and analysis of clinical trials. This study gives an insight into the potential size of ICCs in primary and secondary care, and provides a practical guide to researchers to aid the planning of future studies in this area.

Cluster Analysis↗

Determination of mean particle volume, a Monte Carlo simulation.

Either length measurements or area measurements may be made on a sample of profiles for the purpose of estimating the mean volume of a population of convex particles. Diameters of spheres, caliper diameters of ellipsoids and intercept lengths are available length measurements. Profile areas can be evaluated by planimetry or point counting. Either all the available profiles in random sections or point sampled profiles can be utilized. We have applied a Monte Carlo simulation to compare several of the stereologic methods for the estimation of the mean volumes of spheres and ellipsoids. Populations of spherical, prolate ellipsoidal and oblate ellipsoidal particles were subjected to random sectioning and measurement. Diameter, point sampled intercept length, area and point sampled area were measured in the case of the spherical particles. With the ellipsoids, the same measurement excepting diameters were performed. The measurements were converted to volumes by the appropriate equations, and the means, the standard deviations of the means and the 95% confidence intervals were determined for increasing sample sizes. All the methods provide estimates that converge on their theoretical mean volumes. The area measurements and particularly the point sampled area measurement show some advantage over the length measurements, but differences among the methods are small, not entirely consistent over the different cases and unlikely to be significant in most real applications.

Algorithms↗