Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Estimating effective population size from samples of sequences: inefficiency of pairwise and segregating sites as compared to phylogenetic estimates.

It is known that under neutral mutation at a known mutation rate a sample of nucleotide sequences, within which there is assumed to be no recombination, allows estimation of the effective size of an isolated population. This paper investigates the case of very long sequences, where each pair of sequences allows a precise estimate of the divergence time of those two gene copies. The average divergence time of all pairs of copies estimates twice the effective population number and an estimate can also be derived from the number of segregating sites. One can alternatively estimate the genealogy of the copies. This paper shows how a maximum likelihood estimate of the effective population number can be derived from such a genealogical tree. The pairwise and the segregating sites estimates are shown to be much less efficient than this maximum likelihood estimate, and this is verified by computer simulation. The result implies that there is much to gain by explicitly taking the tree structure of these genealogies into account.

Likelihood Functions↗

Tobacco, nicotine, and cannabis use and exposure in an Australian Indigenous population during pregnancy: A protocol to measure parental and foetal exposure and outcomes.

BACKGROUND: The Australian National Perinatal Data Collection collates all live and stillbirths from States and Territories in Australia. In that database, maternal cigarette smoking is noted twice (smoking <20 weeks gestation; smoking >20 weeks gestation). Cannabis use and other forms of nicotine use, for example vaping and nicotine replacement therapy, are nor reported. The 2021 report shows the rate of smoking for Australian Indigenous mothers was 42% compared with 11% for Australian non-Indigenous mothers. Evidence shows that Indigenous babies exposed to maternal smoking have a higher rate of adverse outcomes compared to non-Indigenous babies exposed to maternal smoking (S1 File). OBJECTIVES: The reasons for the differences in health outcome between Indigenous and non-Indigenous pregnancies exposed to tobacco and nicotine is unknown but will be explored in this project through a number of activities. Firstly, the patterns of parental and household tobacco, nicotine and cannabis use and exposure will be mapped during pregnancy. Secondly, a range of biological samples will be collected to enable the first determination of Australian Indigenous people's nicotine and cannabis metabolism during pregnancy; this assessment will be informed by pharmacogenomic analysis. Thirdly, the pharmacokinetic and pharmacogenomic findings will be considered against maternal, placental, foetal and neonatal outcomes. Lastly, an assessment of population health literacy and risk perception related to tobacco, nicotine and cannabis products peri-pregnancy will be undertaken. METHODS: This is a community-driven, co-designed, prospective, mixed-method observational study with regional Queensland parents expecting an Australian Indigenous baby and their close house-hold contacts during the peri-gestational period. The research utilises a multi-pronged and multi-disciplinary approach to explore interlinked objectives. RESULTS: A sample of 80 mothers expecting an Australian Indigenous baby will be recruited. This sample size will allow estimation of at least 90% sensitivity and specificity for the screening tool which maps the patterns of tobacco and nicotine use and exposure versus urinary cotinine with 95% CI within &#xb1;7% of the point estimate. The sample size required for other aspects of the research is less (pharmacokinetic and genomic n = 50, and the placental aspects n = 40), however from all 80 mothers, all samples will be collected. CONCLUSIONS: Results will be reported using the STROBE guidelines for observational studies. FORWARD: We acknowledge the Traditional Custodians, the Butchulla people, of the lands and waters upon which this research is conducted. We acknowledge their continuing connections to country and pay our respects to Elders past, present and emerging. Notation: In this document, the terms Aboriginal and Torres Strait Islander and Indigenous are used interchangeably for Australia's First Nations People. No disrespect is intended, and we acknowledge the rich cultural diversity of the groups of peoples that are the Traditional Custodians of the land with which they identify and with whom they share a connection and ancestry.

Adult↗

Comparison of ratio-synthetic, sample-size dependent and EBLUP estimators as estimators of food-animal productivity parameters.

A comparison was made of three small-area sampling methods [two traditional design-based methods (ratio-synthetic, sample-size dependent) and one model-based method (EBLUP)] in estimation of some cow and sow population productivity parameters. Performance was evaluated in estimating both farm-specific mean responses and mean animal response over all farms using sample sizes of 100 and 25. Differences in results obtained with the cow and sow data are discussed in terms of the impact of sample size and population size on sampling method. There was a tendency for the model-based method to be the best performer in situations most likely to be operational when the sampling is done as part of a food-animal monitoring scheme. The situations are identified where the sample-size-dependent method performed best.

Animals↗

Sample size for testing and estimating the difference between two paired and unpaired proportions: a 'two-step' procedure combining power and the probability of obtaining a precise estimate.

Clinical trials and scientific research studies are currently planned calculating sample sizes to fulfill power requirements, but the simultaneous need to obtain a satisfactorily precise effect estimate is not widely recognized. I have devised a 'two-step' iterative procedure for comparing two binomial parameters for two paired and unpaired proportions (the most frequent situations in scientific research), which takes into account power and the probability of obtaining a predetermined precision of the effect estimate. The first step provides the sample size for the power of the statistical test, the expected width of its corresponding confidence interval, and the probability of obtaining, under the alternative hypothesis, confidence intervals whose width is less than that expected. The second step iteratively increases this sample size until the probability of obtaining such confidence intervals exceeds a required threshold.

Clinical Trials as Topic↗

The use of item parcels in structural equation modelling: non-normal data and small sample sizes.

Maximum likelihood estimation in confirmatory factor analysis requires large sample sizes, normally distributed item responses, and reliable indicators of each latent construct, but these ideals are rarely met. We examine alternative strategies for dealing with non-normal data, particularly when the sample size is small. In two simulation studies, we systematically varied: the degree of non-normality; the sample size from 50 to 1000; the way of indicator formation, comparing items versus parcels; the parcelling strategy, evaluating uniformly positively skews and kurtosis parcels versus those with counterbalancing skews and kurtosis; and the estimation procedure, contrasting maximum likelihood and asymptotically distribution-free methods. We evaluated the convergence behaviour of solutions, as well as the systematic bias and variability of parameter estimates, and goodness of fit.

Factor Analysis, Statistical↗

Estimates, power and sample size calculations for two-sample ordinal outcomes under before-after study designs.

Sample size calculations are given for comparing two groups of subjects, typically referring to active and non-active intervention groups, on an ordinal outcome in experiments where the subjects are measured before and after intervention. These calculations apply to log-odds models with random intercepts, treatment, time and treatment-by-time interaction terms, the latter being the term of interest. The assumed forms of the odds ratios are flexible, allowing for proportional odds, adjacent categories, or other conditional models for ordinal responses. Simulations studies show that, for given sample sizes, the nominal and actual powers of the proposed test are similar.

Clinical Trials as Topic↗

Monitoring the impact of Bt maize on butterflies in the field: estimation of required sample sizes.

The monitoring of genetically modified organisms (GMOs) after deliberate release is important in order to assess and evaluate possible environmental effects. Concerns have been raised that the transgenic crop, Bt maize, may affect butterflies occurring in field margins. Therefore, a monitoring of butterflies was suggested accompanying the commercial cultivation of Bt maize. In this study, baseline data on the butterfly species and their abundance in maize field margins is presented together with implications for butterfly monitoring. The study was conducted in Bavaria, South Germany, between 2000-2002. A total of 33 butterfly species was recorded in field margins. A small number of species dominated the community, and butterflies observed were mostly common species. Observation duration was the most important factor influencing the monitoring results. Field margin size affected the butterfly abundance, and habitat diversity had a tendency to influence species richness. Sample size and statistical power analyses indicated that a sample size in the range of 75 to 150 field margins for treatment (transgenic maize) and control (conventional maize) would detect (power of 80%) effects larger than 15% in species richness and the butterfly abundance pooled across species. However, a much higher number of field margins must be sampled in order to achieve a higher statistical power, to detect smaller effects, and to monitor single butterfly species.

Agriculture↗

Sample size required for predefined linkage decision quality.

A method for estimating the sample size required to attain a predefined linkage decision quality (type I and type II errors) is proposed using the linkage test power estimate developed by Ginsburg et al. [(1996) Genet Epidemiol 13:355-366]. The method is applicable for samples of arbitrarily structured pedigrees collected via proband. Comparison of different ascertainment schemes and pedigree structures by their consequent minimal sample size was performed. For recessive and dominant inheritance with complete penetrance, the relative ranks of the ascertainment schemes are invariant regardless of the true recombination fraction value and the trait and marker gene frequencies, which enables one to point out the better scheme. The feasibility of evaluating a sampling strategy by the cost of pedigree collection is also considered, and comparison between these two methods of sample planning is performed.

Gene Frequency↗

The PowerAtlas: a power and sample size atlas for microarray experimental design and research.

BACKGROUND: Microarrays permit biologists to simultaneously measure the mRNA abundance of thousands of genes. An important issue facing investigators planning microarray experiments is how to estimate the sample size required for good statistical power. What is the projected sample size or number of replicate chips needed to address the multiple hypotheses with acceptable accuracy? Statistical methods exist for calculating power based upon a single hypothesis, using estimates of the variability in data from pilot studies. There is, however, a need for methods to estimate power and/or required sample sizes in situations where multiple hypotheses are being tested, such as in microarray experiments. In addition, investigators frequently do not have pilot data to estimate the sample sizes required for microarray studies. RESULTS: To address this challenge, we have developed a Microrarray PowerAtlas. The atlas enables estimation of statistical power by allowing investigators to appropriately plan studies by building upon previous studies that have similar experimental characteristics. Currently, there are sample sizes and power estimates based on 632 experiments from Gene Expression Omnibus (GEO). The PowerAtlas also permits investigators to upload their own pilot data and derive power and sample size estimates from these data. This resource will be updated regularly with new datasets from GEO and other databases such as The Nottingham Arabidopsis Stock Center (NASC). CONCLUSION: This resource provides a valuable tool for investigators who are planning efficient microarray studies and estimating required sample sizes.

Algorithms↗

Sample size required for various methods of assessing bone status in commercial leghorn hens.

A study was conducted to determine the appropriate sample size required for various methods used to assess tibial bone status in commercial Leghorn hens. The methods used were in vivo bone mineral content (BMC), in vivo bone density (BD), in vitro BMC, in vitro BD, tibia bone breaking strength (TBS), and percentage bone ash (BA). Dietary total P levels of .4, .45, .5, .55, and .7% were used as treatment source of variation. Twenty hens were sampled randomly to represent each dietary treatment. The CV for each bone status comparison method was estimated and was used in a procedure to estimate the sample size requirement for detecting a difference of delta between treatments. The sample size required to detect the difference between treatment means varied depending on 1) the method used to compare bone status 2) the difference between the treatment means to be detected as significant (delta); and 3) the level of significance (alpha) assumed. The sample size required for various methods are tabulated at .01, .05, and .1 level of significance and for 2.5, 5,7.5, 10, 15, and 20% delta. To detect an actual difference of 5% from the mean to be significant, at the .05 level of significance, a sample size of 44, 22, 31, 23, 47, and 85 hens per treatment would be necessary for in vivo BMC, in vivo BD, in vitro BMC, in vitro BD, TBS, and BA methods, respectively. The estimated sample size values would help researchers in designing experiments that involve bone status comparison of commercial Leghorn hens.

Animals↗

Sample size calculations for clustered binary data.

In this paper we propose a sample size calculation method for testing on a binomial proportion when binary observations are dependent within clusters. In estimating the binomial proportion in clustered binary data, two weighting systems have been popular: equal weights to clusters and equal weights to units within clusters. When the number of units varies cluster by cluster, performance of these two weighting systems depends on the extent of correlation among units within each cluster. In addition to them, we will also use an optimal weighting method that minimizes the variance of the estimator. A sample size formula is derived for each of the estimators with different weighting schemes. We apply these methods to the sample size calculation for the sensitivity of a periodontal diagnostic test. Simulation studies are conducted to evaluate a finite sample performance of the three estimators. We also assess the influence of misspecified input parameter values on the calculated sample size. The optimal estimator requires equal or smaller sample sizes and is more robust to the misspecification of an input parameter than those assigning equal weights to units or clusters.

Bacteroides Infections↗

Sample size determination for a t test given a t value from a previous study: A FORTRAN 77 program.

When uncertain about the magnitude of an effect, researchers commonly substitute in the standard sample-size-determination formula an estimate of effect size derived from a previous experiment. A problem with this approach is that the traditional sample-size-determination formula was not designed to deal with the uncertainty inherent in an effect-size estimate. Consequently, estimate-substitution in the traditional sample-size-determination formula can lead to a substantial loss of power. A method of sample-size determination designed to handle uncertainty in effect-size estimates is described. The procedure uses the t value and sample size from a previous study, which might be a pilot study or a related study in the same area, to establish a distribution of probable effect sizes. The sample size to be employed in the new study is that which supplies an expected power of the desired amount over the distribution of probable effect sizes. A FORTRAN 77 program is presented that permits swift calculation of sample size for a variety of t tests, including independent t tests, related t tests, t tests of correlation coefficients, and t tests of multiple regression b coefficients.

Data Interpretation, Statistical↗