Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Correlated response, competition, and female canine size in primates.

Recently, comparative analyses of female canine tooth size in primates have yielded two hypotheses to explain interspecific variation in female relative canine size. Greenfield ([1992] Int. J. Primatol. 13:631-657; [1992] Yrbk. Phys. Anthropol. 35:153-184; [1996] J. Hum. Evol. 31:1-19) suggested that covariation in male and female canine size across species indicates that female canine size reflects correlated response (in which the expression of a trait in one sex causes the expression of the same trait in the other sex). Plavcan et al. ([1995] J. Hum. Evol. 28:245-276) noted that female canine size in primates is associated with variation in categorical estimates of the intensity of female-female agonistic competition, suggesting that selection favors large female canine size in many species. While it may seem that the two models are in conflict, they are not. To simultaneously evaluate these two models, this analysis examines the joint relations between male canine size, female canine size, and estimates of female-female competition in a sample of 108 primate species. Overall, female canine size is correlated with variation in male canine size. Controlling for variation in male canine size, female canine size is also correlated with estimates of the intensity of female-female agonistic competition. The relation between these variables differs strongly between anthropoid and strepsirhine primates. In anthropoids, the data suggest that selection for the development of large canines in females is not constrained by any affect of correlated response. In strepsirhines, the evidence suggests that sexual selection may affect male canine size but that correlated response affects female canine size, resulting in monomorphism for most species. These observations help reconcile the observations of Greenfield ([1992] Int. J. Primatol. 13:631-657; [1996] J. Hum. Evol. 31:1-19) and Plavcan et al. ([1995] J. Hum. Evol. 28:245-276) and provide a more precise model for understanding interspecific variation in female canine size and hence canine dimorphism.

Animals↗

Gene-dropping vs. empirical variance estimation for allele-sharing linkage statistics.

In this study, we compare the statistical properties of a number of methods for estimating P-values for allele-sharing statistics in non-parametric linkage analysis. Some of the methods are based on the normality assumption, using different variance estimation methods, and others use simulation (gene-dropping) to find empirical distributions of the test statistics. For variance estimation methods, we consider the perfect variance approximation and two empirical variance estimates. The simulation-based methods are gene-dropping with and without conditioning on the observed founder alleles. We also consider the Kong and Cox linear and exponential models and a Monte Carlo method modified from a method for finding genome-wide significance levels. We discuss the analytical properties of these various P-value estimation methods and then present simulation results comparing them. Assuming that the sample sizes are large enough to justify a normality assumption for the linkage statistic, the best P-value estimation method depends to some extent on the (unknown) genetic model and on the types of pedigrees in the sample. If the sample sizes are not large enough to justify a normality assumption, then gene-dropping is the best choice. We discuss the differences between conditional and unconditional gene-dropping.

Alleles↗

Long-term effective population sizes, temporal stability of genetic composition and potential for local adaptation in anadromous brown trout (Salmo trutta) populations.

We examined the long-term temporal (1910s to 1990s) genetic variation at eight microsatellite DNA loci in brown trout (Salmo trutta L) collected from five anadromous populations in Denmark to assess the long-term stability of genetic composition and to estimate effective population sizes (Ne). Contemporary and historical samples consisted of tissue and archived scales, respectively. Pairwise thetaST estimates, a hierarchical analysis of molecular variance (amova) and multidimensional scaling analysis of pairwise genetic distances between samples revealed much closer genetic relationships among temporal samples from the same populations than among samples from different populations. Estimates of Ne, using a likelihood-based implementation of the temporal method, revealed Ne >or= 500 in two of three populations for which we have historical data. A third population in a small (3 km) river showed Ne >or= 300. Assuming a stepping-stone model of gene flow we considered the relative roles of gene flow, random genetic drift and selection to assess the possibilities for local adaptation. The requirements for local adaptation were fulfilled, but only adaptations resulting from strong selection were expected to occur at the level of individual populations. Adaptations resulting from weak selection were more likely to occur on a regional basis, i.e. encompassing several populations. Ne appears to have declined recently in at least one of the studied populations, and the documented recent declines of many other anadromous brown trout populations may affect the persistence of local adaptation.

Animals↗

Farm Scale Evaluations of spring-sown genetically modified herbicide-tolerant crops: a statistical assessment.

Primary results from the Farm Scale Evaluations (FSEs) of spring-sown genetically modified herbicide-tolerant crops were published in 2003. We provide a statistical assessment of the results for count data, addressing issues of sample size (n), efficiency, power, statistical significance, variability and model selection. Treatment effects were consistent between rare and abundant species. Coefficients of variation averaged 73% but varied widely. High variability in vegetation indicators was usually offset by large n and treatment effects, whilst invertebrate indicators often had smaller n and lower variability; overall, achieved power was broadly consistent across indicators. Inferences about treatment effects were robust to model misspecification, justifying the statistical model adopted. As expected, increases in n would improve detectability of effects whilst, for example, halving n would have resulted in a loss of significant results of about the same order. 40% of the 531 published analyses had greater than 80% power to detect a 1.5-fold effect; reducing n by one-third would most likely halve the number of analyses meeting this criterion. Overall, the data collected vindicated the initial statistical power analysis and the planned replication. The FSEs provide a valuable database of variability and estimates of power under various sample size scenarios to aid planning of more efficient future studies.

Agriculture↗

[Comparison of family physicians' and gynecologists' use of the intrauterine device (IUD)].

OBJECTIVE: To quantify differences between general practitioners (GPs) and gynaecologists in the technique of insertion and follow-up of the intra-uterine device (IUD). DESIGN: Multicentred, descriptive, longitudinal study. SETTING: Two urban health centres and a family guidance clinic. PARTICIPANTS: Target population (n = 1700) between January 1993 and January 1996. Estimated mean of complications was 25%. Sample size was 247 for alpha = 0.05 and 1-alpha = 0.95. The sample was extended to 300 to allow for possible losses of files, estimated at 20%. MEASUREMENTS AND MAIN RESULTS: The variables age, sex, marital status, educational level, parity, abortions, previous contraception, type of job, type of IUD, post-insertion and follow-up complications, subjective evaluation, removal and average follow-up time, were analysed. 158 (54.9%) of the 288 IUDs finally studied were inserted by GPs, and 130 (45.1%) by gynaecologists. 69.5% were anchor-shaped, and 30.5% T-shaped. In 85.5% no immediate complications were found. Mean follow-up time was 22.67 months (CI 95%, 21.3-24.0), during which time 36.6% had complications detected, which led to removal of the device in 22.3% of complications. We found no statistically significant differences between the two populations for age, marital status, subjective evaluation, number of abortions, parity or previous contraception. Likewise, no differences between G.P.s and gynaecologists were detected for post-insertion or follow-up complications, percentage of IUDs removed, or period of time evaluated. There were differences found for the type of IUD used, with more anchor-shaped IUDs in primary care. There were no differences for the type of IUD or complications requiring its removal. CONCLUSIONS: In the population studied we found no differences in immediate or later complications between IUDs inserted by GPs and by gynaecologists.

Adult↗

How reporting delay, duration of follow-up and number of cases affect the estimates of the incubation time of transfusion-associated AIDS cases.

The authors discuss the impact of reporting delay, duration of follow-up, and number of cases in a sample on estimates of the incubation time of transfusion-associated AIDS cases. "This article comes to the conclusion that the accuracy of the incubation time estimate would depend on the sample size rather than on the duration of follow-up." (SUMMARY IN FRE)

Acquired Immunodeficiency Syndrome↗

Is there an epidemic of child or adolescent depression?

BACKGROUND: Both the professional and the general media have recently published concerns about an 'epidemic' of child and adolescent depression. Reasons for this concern include (1) increases in antidepressant prescriptions, (2) retrospective recall by successive birth cohorts of adults, (3) rising adolescent suicide rates until 1990, and (4) evidence of an increase in emotional problems across three cohorts of British adolescents. METHODS: Epidemiologic studies of children born between 1965 and 1996 were reviewed and a meta-analysis conducted of all studies that used structured diagnostic interviews to make formal diagnoses of depression on representative population samples of participants up to age 18. The effect of year of birth on prevalence was estimated, controlling for age, sex, sample size, taxonomy (e.g., DSM vs. ICD), measurement instrument, and time-frame of the interview (current, 3 months, 6 months, 12 months). RESULTS: Twenty-six studies were identified, generating close to 60,000 observations on children born between 1965 and 1996 who had received at least one structured psychiatric interview capable of making a formal diagnosis of depression. Rates of depression showed no effect of year of birth. There was little effect of taxonomy, measurement instrument, or time-frame of interview. The overall prevalence estimates were: under 13, 2.8% (standard error (SE) .5%); 13-18 5.6% (SE .3%); 13-18 girls: 5.9% (SE .3%); 13-18 boys: 4.6% (SE .3%). CONCLUSIONS: When concurrent assessment rather than retrospective recall is used, there is no evidence for an increased prevalence of child or adolescent depression over the past 30 years. Public perception of an 'epidemic' may arise from heightened awareness of a disorder that was long under-diagnosed by clinicians.

Adolescent↗

Maternal asthma and risk of preeclampsia: a case-control study.

OBJECTIVE: To quantify the associations between asthma characteristics and the risk of preeclampsia. STUDY DESIGN: In this case-control study, asthma history among 286 preeclampsia cases and 470 normotensive controls in Seattle was assessed by postpartum interview and medical record abstraction. OR and 95% CI were estimated using logistic regression. The sample size was adequate to detect unadjusted asthma history with ORs of > or =1.6 at a power of 80%. RESULTS: After adjustment, women with a history of prepregnancy asthma diagnosis were not at increased preeclampsia risk (OR 0.94, 95% CI 0.58-1.52). Women experiencing asthma symptoms during pregnancy were more likely than pregnant nonasthmatics to have preeclampsia (OR 2.20, 95% CI 0.79-6.10). Those with long-term pre-pregnancy asthma and symptoms during pregnancy were at particularly increased risk (OR 9.09, 95% CI 1.02-81.6). Point estimates were generally higher after restriction to women withfull-term deliveries. CONCLUSION: This analysis suggests that asthmatics, particularly those who are symptomatic during pregnancy, may be at higher risk of developing preeclampsia.

Adult↗

Implications of measurement error in exposure for the sample sizes of case-control studies.

In this paper, recent results describing the effects of measurement error on estimation of the association between an exposure and a disease are applied to sample size calculation in case-control studies. Models of the relation between true exposure and a surrogate exposure measure assessed with error are used to derive equations for sample size determination. The results show that the sample size of a study based on an exposure variable which is measured with error must be larger by a factor of 1/rho 2 than if exposure were measured without error, where rho is the correlation between the true exposure and the surrogate exposure measure. Review of the magnitude of measurement error in dietary assessments suggests that failure to account for measurement error in sample size determination for case-control studies of diet and disease could lead to marked underestimation of the required sample size.

Bias↗

Sampling design for total and filterable reactive phosphorus monitoring in a lowland stream: considerations of spatial variability, measurement uncertainty and statistical power.

An analysis for spatial variation of phosphorus (P) concentrations in the dissolved and particulate compartments of Latrobe River water in Victoria, Australia is described. Water sampling was based on a nested hierarchical design and variation was measured at different spatial scales. Total variance of the dissolved and particulate P compartments was partitioned using analysis of variance (ANOVA) to determine the spatial scale that requires most sampling effort. Statistical power analysis was used to determine the optimum sample size for the spatial scale. An uncertainty budget was estimated from sampling and analytical uncertainty. Ten and twelve samples, at the smallest spatial scale, required the greatest sampling effort, and led to the greatest required statistical power for the determination of dissolved and particulate P, respectively, in the Latrobe River catchment. The results emphasize the need for aquatic chemists to be aware of the ramifications of different types of uncertainty and variance in environmental studies.

Environmental Monitoring↗

[Mapping the trait controlled by two duplicate genes in the DH or RIL population].

While there is linkage between molecular marker and trait controlled by two duplicate genes in the DH or RIL population, the recombination rate (RR) between molecular marker and one gene controlling the above trait may be estimated by the maximum likelihood method. Moreover, the standard deviation of RR was also obtained in this paper. Finally, the results from Monte Carlo simulation with 3000 replications showed that the unbiasedness of RR for various sample size and RR was good, and the variation of the estimated value of RR decreased with the increase of sample size or RR.

English Abstract↗

Confidence intervals for a ratio of binomial proportions based on paired data.

Four interval estimation methods for the ratio of marginal binomial proportions are compared in terms of expected interval width and exact coverage probability. Two new methods are proposed that are based on combining two Wilson score intervals. The new methods are easy to compute and perform as well or better than the method recently proposed by Nam and Blackwelder. Two sample size formulas are proposed to approximate the sample size required to achieve an interval estimate with desired confidence level and width.

Adolescent↗

Sample size and power calculations with correlated binary data.

Correlated binary data are common in biomedical studies. Such data can be analyzed using Liang and Zeger's generalized estimating equations (GEE) approach. An attractive point of the GEE approach is that one can use a misspecified working correlation matrix, such as the working independence model (i.e., the identity matrix), and draw (asymptotically) valid statistical inference by using the so-called robust or sandwich variance estimator. In this article we derive some explicit formulas for sample size and power calculations under various common situations. The given formulas are based on using the robust variance estimator in GEE. We believe that these formulas will facilitate the practice in planning two-arm clinical trials with correlated binary outcome data.

Clinical Trials as Topic↗

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results↗

Inferring population history from genealogical trees.

Inference about population history from DNA sequence data has become increasingly popular. For human populations, questions about whether a population has been expanding and when expansion began are often the focus of attention. For viral populations, questions about the epidemiological history of a virus, e.g., HIV-1 and Hepatitis C, are often of interest. In this paper I address the following question: Can population history be accurately inferred from single locus DNA data? An idealised world is considered in which the tree relating a sample of n non-recombining and selectively neutral DNA sequences is observed, rather than just the sequences themselves. This approach provides an upper limit to the information that possibly can be extracted from a sample. It is shown, based on Kingman's (1982a) coalescent process, that consistent estimation of parameters describing population history (e.g., a growth rate) cannot be achieved for increasing sample size, n. This is worse than often found for estimators of genetic parameters, e.g., the mutation rate typically converges at rate under the assumption that all historical mutations can be observed in the sample. In addition, various results for the distribution of maximum likelihood estimators are presented.

DNA↗

Accuracy of coalescent likelihood estimates: do we need more sites, more sequences, or more loci?

A computer simulation study has been made of the accuracy of estimates of Theta = 4Nemu from a sample from a single isolated population of finite size. The accuracies turn out to be well predicted by a formula developed by Fu and Li, who used optimistic assumptions. Their formulas are restated in terms of accuracy, defined here as the reciprocal of the squared coefficient of variation. This should be proportional to sample size when the entities sampled provide independent information. Using these formulas for accuracy, the sampling strategy for estimation of Theta can be investigated. Two models for cost have been used, a cost-per-base model and a cost-per-read model. The former would lead us to prefer to have a very large number of loci, each one base long. The latter, which is more realistic, causes us to prefer to have one read per locus and an optimum sample size which declines as costs of sampling organisms increase. For realistic values, the optimum sample size is 8 or fewer individuals. This is quite close to the results obtained by Pluzhnikov and Donnelly for a cost-per-base model, evaluating other estimators of Theta. It can be understood by considering that the resources spent collecting larger samples prevent us from considering more loci. An examination of the efficiency of Watterson's estimator of Theta was also made, and it was found to be reasonably efficient when the number of mutants per generation in the sequence in the whole population is less than 2.5.

Genetic Variation↗

Design issues and sample size when exposure measurement is inaccurate.

Measurement error often leads to biased estimates and incorrect tests in epidemiological studies. These problems can be corrected by design modifications which allow for refined statistical models, or in some situations by adjusted sample sizes to compensate a power reduction. The design options are mainly an additional replication or internal validation study. Sample size calculations for these designs are more complex, since usually there is no unique design solution to obtain a prespecified power. Thus, additionally to a power requirement, an optimal design should also fulfill the criteria of minimizing overall costs. In this review corresponding strategies and formulae are described and appraised.

Bias↗

Gene and haplotype frequencies for the loci hLA-A, hLA-B, and hLA-DR based on over 13,000 german blood donors.

Numerous applications in clinical medicine and forensic sciences depend on reliable data concerning the frequencies of human leukocyte antigen (HLA) genes and haplotypes. Assuming a Hardy-Weinberg equilibrium of the underlying population, these frequencies can be estimated from phenotype data using an expectation-maximization-algorithm also known under the name "gene counting." We have refined this algorithm in order to cope with the heterogeneous resolution of HLA phenotypes frequently occurring in large datasets due to the structure of the HLA nomenclature. This was a prerequisite to analyze a set of 13,386 blood donors contributed by over 40 blood banks who were tested for HLA-DR when they volunteered to become marrow donors. This data set is still unique in the German national donor registry because their HLA-DR-typing was not biased by patient oriented searches or other strategies for selective typing. As a consequence of the size of the sample, the frequency estimates for the genes and the two- and three-locus haplotypes of HLA-A, HLA-B, and HLA-DR are of unprecedented precision and allow interesting projections concerning the efficiency and economic aspects of the development of a large donor registry in Germany.

Algorithms↗