Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Correcting for ascertainment bias of relative-risk estimates obtained using affected-sib-pair linkage data.

Locus-specific sibling relative risk is often estimated using affected-sib-pair lod score analysis of affected sibships and may be used to decide whether to continue or discontinue the search for additional susceptibility genes. We showed that relative-risk estimates obtained using affected-sib-pair data are asymptotically unbiased when each pair is given a weight inversely proportional to the sibship ascertainment probability. Here we show by simulation that the extent of the bias of relative risks estimated using the incorrect ascertainment weights is small for dominant models, but large for single-locus recessive models and some two-locus heterogeneity models. Since in practice the ascertainment scheme is often unknown, we investigate methods for jointly estimating ascertainment and relative risks from affected-sibship data. Given a sufficient sample size, a reasonable estimate of relative risk may always be obtained using only affected pairs from sibships with two affected and no unaffected siblings. This estimate, which has a large variance, may then be used in a three-stage procedure (which we call the alpha method) to estimate consistently both the ascertainment probabilities and the relative risks with greater precision. We additionally propose correction factors to eliminate small-sample bias of relative risks and investigate the bias due to error in the estimate of disease locus location.

Bias↗

What is the extent of prokaryotic diversity?

The extent of microbial diversity is an intrinsically fascinating subject of profound practical importance. The term 'diversity' may allude to the number of taxa or species richness as well as their relative abundance. There is uncertainty about both, primarily because sample sizes are too small. Non-parametric diversity estimators make gross underestimates if used with small sample sizes on unevenly distributed communities. One can make richness estimates over many scales using small samples by assuming a species/taxa-abundance distribution. However, no one knows what the underlying taxa-abundance distributions are for bacterial communities. Latterly, diversity has been estimated by fitting data from gene clone libraries and extrapolating from this to taxa-abundance curves to estimate richness. However, since sample sizes are small, we cannot be sure that such samples are representative of the community from which they were drawn. It is however possible to formulate, and calibrate, models that predict the diversity of local communities and of samples drawn from that local community. The calibration of such models suggests that migration rates are small and decrease as the community gets larger. The preliminary predictions of the model are qualitatively consistent with the patterns seen in clone libraries in 'real life'. The validation of this model is also confounded by small sample sizes. However, if such models were properly validated, they could form invaluable tools for the prediction of microbial diversity and a basis for the systematic exploration of microbial diversity on the planet.

Archaea↗

Estimation of energy balance at the individual and herd level using blood and milk traits in high-yielding dairy cows.

This study aimed to estimate individual and herd-level energy balance (EB) using blood and milk traits in 90 multiparous high-yielding Holstein cows, held on a research farm, from wk 1 to 10 postpartum (p.p.) and to investigate the precision of prediction with successively decreased data sets simulating smaller herd sizes and with pooled samples. Dry matter intake, milk yield, and BW were measured daily from parturition through wk 10 p.p. Milk composition was determined 4 times per week, and milk acetone was measured weekly. Blood samples for the determination of metabolites, hormones, electrolytes, and enzyme activities were taken weekly from wk 1 to 10 p.p. between 0730 and 0900. Body condition scores and ultrasonic measurements of backfat thickness and fat depth in the pelvic area were evaluated in wk 1, 4, and 8 p.p. Concentrations of glucose, cholesterol, urea, insulin, insulin-like growth factor-1, triiodothyronine, and thyroxine (T4) in blood plasma and of lactose and urea in milk were positively correlated with EB, whereas concentrations of nonesterified fatty acids (NEFA), creatinine, albumin, beta-hydroxybutyrate, and growth hormone and enzyme activities in blood, and concentrations of fat, protein, fat:lactose ratio, and acetone in milk were negatively correlated with EB. Leptin concentration was not correlated to EB over the first 10 wk p.p. To estimate EB linear mixed-effects, models were developed by backward selection procedures. The most informative traits for estimation of EB were the fat:lactose ratio in milk and NEFA and T4 concentrations in blood. The precision of estimation of EB in individual cows was low. Using blood in addition to milk traits did not result in higher precision of estimation of herd-level EB, and decreasing sample sizes considerably lowered the precision of EB prediction. Estimation of overall mean herd-level EB over the first 10 wk p.p. using pooled samples was precise even with small sample sizes, but does not consider the level of EB in particular weeks. In conclusion, estimation of herd-level EB at individual weeks using milk traits only has practical implication with herd sizes of > or = 100 cows if calving is highly seasonal and of or = 400 cows if calving is uniformly distributed. Using blood in addition to milk traits does not improve precision of estimation of herd-level EB, regardless of sample size.

3-Hydroxybutyric Acid↗

Variation of estimates of SNP and haplotype diversity and linkage disequilibrium in samples from the same population due to experimental and evolutionary sample size.

Studies of genetic polymorphisms and diversity between and within human populations are increasingly characterised by a very large number of genetic markers but using a relatively small number of individuals from which DNA samples were taken. In this report we examine the limitations of a small experimental sample size relative to a large genomic sample size, and quantify the sampling variance of a number of measures of diversity and linkage disequilibrium. The relationship between sample size and observed levels of polymorphism and haplotype diversity at the level of a gene is investigated under a neutral model of sequence evolution, using coalescent simulations. It is shown that the effect of evolutionary sampling, as manifested by differences between samples (genes) in measures of diversity estimated using very large sample sizes, is substantial, with a coefficient of variation of the number of detected polymorphic SNPs or haplotypes in the order of 15%. The effect of experimental design (sample size) is also very large, and a number of 'significant' results reported in the literature can be explained by sampling alone. The expected correlation coefficient of measures of linkage disequilibrium across samples from the same population has been quantified and found to be consistent with empirical estimates from the literature.

Chromosomes, Human↗

Statistical effects of varying sample sizes on the precision of percentile estimates.

The present study evaluates the precision of outlying percentile estimates, with age- and sex-associated variations and facilitates decisions needed to revise the current NCHS 1977 Growth Charts with regard to 1) the inclusion of 3(rd) and 97(th) percentiles and 2) the selection of survey data for the construction of the revised growth charts. Simulation was performed to obtain data with distribution characteristics similar to those of The Third National Health and Nutritional Examination Survey (NHANES III) (1988-1991) data. NHANES III consists of a two-phase, 6-year, complex stratified multistage probability cluster, cross-sectional survey conducted from 1988 through 1994 to represent the US noninstitutionalized population. Phase I of the survey consisted of 679 boys and 622 girls in age groups 3, 8, 13, and 18 years. Weight and stature, the body mass index (BMI) (weight/stature(2); kg/m(2)) was calculated. The results show that 1) the precision of the percentile estimates is greater for stature than for weight and BMI, 2) percentiles during the pubertal period are less precise than those during the prepubertal and postpubertal periods for weight and BMI but there is little difference for stature, and 3) percentile estimates are more precise for girls than boys for weight and BMI, but not for stature. The present findings suggest that pooling of NHANES III and earlier National Center for Health Statistics (NCHS) survey data is necessary to achieve reasonable precision for the 3(rd) and 97(th) percentile estimates. Am. J. Hum. Biol. 12:64-74, 2000. Copyright 2000 Wiley-Liss, Inc.

Journal Article↗

A Monte Carlo investigation of the robustness of the Wald, likelihood Ratio, and Median tests with specified symmetric and asymmetric marginal distributions.

The effect of nonnormality on the Type I (tau) error when comparing two independent binomial proportions (P) or the nonparametric alternatives, the Median (Me), Wald (W), and Likelihood Ratio (LR), has not been investigated. If these selected tests are overly conservative the implied loss of power would moderate their practical use. The purpose of the present study was to investigate the impact of nonnormality on small to moderate sample sizes on the estimated tau for alpha = 0.10, 0.05, and 0.01 for the P, Me, W, and LR tests. Samples were generated from nine long-tailed symmetric and asymmetric distributions using a multiplicative congruential generator. For each marginal distribution and for a variety of sample sizes, the proportion of samples for which the test statistic exceeded the 10, 5, and 1 percentage points was tabulated. For data that mimic a symmetric distribution, the median test uniformly yields an empirical alpha considerably less than tau, while the likelihood ratio test consistently overestimates tau for small samples (n < or = 15) over all symmetric distributions and empirical alpha levels. For asymmetric distributions, the median test again yields an empirical alpha significantly less than tau. Similar underestimates of tau were found for the chi-square (2 df), chi-square (4 df) log normal, and gamma (2, 1) distributions. The likelihood ratio test consistently overestimates tau for small samples (n < or = 15) over all asymmetric distributions and empirical alpha levels. The independent proportions test produces an empirical alpha closest to tau for n = 10 for all asymmetric distributions.(ABSTRACT TRUNCATED AT 250 WORDS)

Data Interpretation, Statistical↗

[A national survey on current status of the important parasitic diseases in human population].

In order to understand the current status and trends of the important parasitic diseases in human population, to evaluate the effect of control activities in the past decade and provide scientific base for further developing control strategies, a national survey was carried out in the country (Taiwan, Hongkong and Macau not included) from June, 2001 to 2004 under the sponsorship of the Ministry of Health. The sample sizes of the nationwide survey and of the survey in each province (autonomous region and municipality, P/A/M) were determined following a calculating formula based on an estimation of the sample size of random sampling to the rate of population. A procedure of stratified cluster random sampling was conducted in each province based on geographical location and economical condition with three strata: county/city, township/town, and spot, each spot covered a sample of 500 people. Parasitological examinations were conducted for the infections of soil-transmitted nematodes, Taenia spp, and Clonorchis sinensis, including Kato-Katz thick smear method, scotch cellulose adhesive tape technique and test tube-filter paper culture (for larvae). At the same time, another sampled investigation for Clonorchis sinensis infection was carried out in the known endemic areas in 27 provinces. Serological tests combined with questionnaire and/or clinical diagnosis were applied for hydatid disease, cysticercosis, paragonimiasis, trichinosis, and toxoplasmosis. A total sampled population of 356 629 from the 31 P/A/M was examined by parasitological methods and 26 species of helminth were recorded. Among these helminth, human infections of Metorchis orientalis and Echinostoma aegypti were detected in Fujian Province which seemed to be the first report in the world, and Haplorchis taichui infection in Guangxi Region was the first human infection record in the country. The overall prevalence of helminth infections was 21.74%. The prevalence of soil-transmitted nematodes was 19.56% (including hookworm infection 6.12%, Ascaris infection 12.72% and Trichuris infection 4.63%), and the estimated number of population infected with soil-transmitted nematodes was 129 million (with 39.3, 85.93 and 29.09 million for hookworm, Ascaris and Trichuris infections respectively). The prevalence of Taenia infection was 0.28% with an infected population of 550 000. The prevalence of Clonorchis sinensis in the national survey was 0.58%. From the survey in the known Clonorchis endemic areas with a sample of 217 829, terobius vermicularis infection in children under 12 years old was 10.28%. The positive rate of serological tests for hydatid disease, cysticercosis, paragonimiasis, trichinosis, and toxoplasmosis was 12.04% (4 796/39 826), 0.58% (553/96 008), 1.71% (1 163/68 209), 3.38% (3 149/93 239) and 7.88% (3 737/47 444) respectively. In comparison to the last national survey in 1990, the prevalence of hookworm, Ascaris lumbricoides and Trichuris trichiura infections has been reduced by 60.72%, 71.29% and 73.60% respectively, and the number of infected people by soil-transmitted nematodes has declined remarkably. However, the prevalence of Clonorchis infection significantly increased in the provinces of Guangdong, Guangxi and Jilin by 182%, 164% and 630% respectively. A remarkable increase of the prevalence of Taenia infection was found in Sichuan and Tibet, by 98% and 97% respectively. Echinococcosis is important in the Western part of China. Many parasitic diseases are still highly prevalent in the rural and pastoral areas with higher prevalence, morbidity and certain case fatality in farmers and herdsmen, especially in women and children.

Adolescent↗

Sample size determination for establishing equivalence/noninferiority via ratio of two proportions in matched-pair design.

In this article, we propose approximate sample size formulas for establishing equivalence or noninferiority of two treatments in match-pairs design. Using the ratio of two proportions as the equivalence measure, we derive sample size formulas based on a score statistic for two types of analyses: hypothesis testing and confidence interval estimation. Depending on the purpose of a study, these formulas can be used to provide a sample size estimate that guarantees a prespecified power of a hypothesis test at a certain significance level or controls the width of a confidence interval with a certain confidence level. Our empirical results confirm that these score methods are reliable in terms of true size, coverage probability, and skewness. A liver scan detection study is used to illustrate the proposed methods.

Biopsy↗

Sequential stopping rules in clinical trials.

This paper reviews aspects of the development of sequential analysis of clinical trial data in medicine and suggests simple strategies for progress. The emphasis is on the pragmatic and ethical requirements of aspects of the design of phase III trials and in circumstances of genuine uncertainty characterized by much clinical experimentation. In particular consideration is given to the consequences of determining sample sizes from incorrect estimates of treatment effects. Armitage's work on sequential trials is traced to simple group sequential procedures based on repeated significance tests to minimize expected sample sizes in a wide class of experimental situations.

Clinical Trials as Topic↗

Interval estimation of the attributable risk for multiple exposure levels in case-control studies with confounders.

The attributable risk (AR) is one of the most important and commonly-used epidemiological indices to assess the public health importance of an association between a risk factor and a disease. When the underlying risk factor has multiple exposure levels in the presence of confounders, we consider the case-control studies using random sampling to collect the cases and controls here. We develop four asymptotic interval estimators for AR, including the interval estimator using Wald's statistic, the interval estimator using the logarithmic transformation, the interval estimator using the logit transformation, and the interval estimator derived from a quadratic equation. We apply Monte Carlo simulation to evaluate the finite-sample performance of these interval estimators in a variety of situations. We demonstrate that given an adequately large sample size, all the estimators developed here can actually perform reasonably well. We note that the interval estimator using the logit transformation may be of limited use when the number of studied subjects is not large. We also note that the interval estimator using the logarithmic transformation can lose efficiency compared to the interval estimator using Wald's statistic or the interval estimator derived from a quadratic equation developed in this paper. Finally, we use the data taken from a case control study of the oral contraceptive use in myocardial infarction patients with various smoking levels to illustrate he practical usefulness of these estimators.

Adult↗

An empirical exploration of data quality in DNA-based population inventories.

I present data from 21 population inventory studies - 20 of them on bears - that relied on the noninvasive collection of hair, and review the methods that were used to prevent genetic errors in these studies. These methods were designed to simultaneously minimize errors (which can bias estimates of abundance) and per-sample analysis effort (which can reduce the precision of estimates by limiting sample size). A variety of approaches were used to probe the reliability of the empirical data, producing a mean, per-study estimate of no more than one undetected error in either direction (too few or too many individuals identified in the laboratory). For the type of samples considered here (plucked hair samples), the gain or loss of individuals in the laboratory can be reduced to a level that is inconsequential relative to the more universal sources of bias and imprecision that can affect mark-recapture studies, assuming that marker systems are selected according to stated guidelines, marginal samples are excluded at an early stage, similar pairs of genotypes are scrutinized, and laboratory work is performed with skill and care.

Animals↗

Polybrominated diphenyl ether (PBDE) levels in an expanded market basket survey of U.S. food and estimated PBDE dietary intake by age and sex.

OBJECTIVES: Our objectives in this study were to expand a previously reported U.S. market basket survey using a larger sample size and to estimate levels of PBDE intake from food for the U.S. general population by sex and age. METHODS: We measured concentrations of 13 polybrominated diphenyl ether (PBDE) congeners in food in 62 food samples. In addition, we estimated levels of PBDE intake from food for the U.S. general population by age (birth through > or = 60 years of age) and sex. RESULTS: In food samples, concentrations of total PBDEs varied from 7.9 pg/g (parts per trillion) in milk to 3,726 pg/g in canned sardines. Fish were highest in PBDEs (mean, 1,120 pg/g; median, 616 pg/g; range, 11.14-3,726 pg/g). This was followed by meat (mean, 383 pg/g; median, 190 pg/g; range, 39-1,426 pg/g) and dairy products (mean, 116 pg/g; median, 32.2 pg/g; range, 7.9-683 pg/g). However, using estimates for food consumption (excluding nursing infants), meat accounted for the highest U.S. dietary PBDE intake, followed by dairy and fish, with almost equal contributions. Adult females had lower dietary intake of PBDEs than did adult males, based on body weight. We estimated PBDE intake from food to be 307 ng/kg/day for nursing infants and from 2 ng/kg/day at 2-5 years of age for both males and females to 0.9 ng/kg/day in adult females. CONCLUSION: Dietary exposure alone does not appear to account for the very high body burdens measured. The indoor environment (dust, air) may play an important role in PBDE body burdens in addition to food.

Adolescent↗

Intra-class correlation estimates for assessment of vitamin A intake in children.

In many community-based surveys, multi-level sampling is inherent in the design. In the design of these studies, especially to calculate the appropriate sample size, investigators need good estimates of intra-class correlation coefficient (ICC), along with the cluster size, to adjust for variation inflation due to clustering at each level. The present study used data on the assessment of clinical vitamin A deficiency and intake of vitamin A-rich food in children in a district in India. For the survey, 16 households were sampled from 200 villages nested within eight randomly-selected blocks of the district. ICCs and components of variances were estimated from a three-level hierarchical random effects analysis of variance model. Estimates of ICCs and variance components were obtained at village and block levels. Between-cluster variation was evident at each level of clustering. In these estimates, ICCs were inversely related to cluster size, but the design effect could be substantial for large clusters. At the block level, most ICC estimates were below 0.07. At the village level, many ICC estimates ranged from 0.014 to 0.45. These estimates may provide useful information for the design of epidemiological studies in which the sampled (or allocated) units range in size from households to large administrative zones.

Analysis of Variance↗

Alternative ways of estimating serological titer reproducibility.

A quantitative measure of the reproducibility of serum antibody titers has recently been proposed (R. J. Wood and T. M. Durham, J. Clin. Microbiol. 11: 541-545, 1980). The measure advocated is "the probability that the maximum ratio of two distinct (integer) titers (obtained in the blind) on the same specimen will not exceed 2." This measure of the reproducibility of serological titers is considered to be a fixed probability for any given specimen and set of test conditions. Although it is a fixed constant during the time period of a study, there are alternative methods one might use to compute an estimate of it, using laboratory data. Four such methods of estimating test reproducibility are discussed and evaluated. The estimates obtained from the two principle methods are evaluated quantitatively by means of Monte Carlo computer simulation. The simulation results show that, from a given sample of replicate integer titers, these two principal methods yield estimates that are highly correlated. In addition, with moderate numbers of replicates (sample size) these methods provide estimates that are on the average properly directed at the true reproducibility values (that is, are essentially, unbiased), particularly when the true reproducibility of the test is at least 0.9. The reliability, or stability, of the alternative estimates is studied for selected sample size.

Analysis of Variance↗

Balanced designs in longitudinal population pharmacokinetic studies.

A simulation study was performed using a balanced design to determine the sample size required for accurate and precise estimation of a parameter at a given level of intersubject variability in a longitudinal population pharmacokinetic study. A two-compartment model parameterized in terms of clearance (Cl), volumes of the central (V1) and peripheral (V2) compartments, and intercompartmental clearance (Q) with multiple intravenous bolus inputs was assumed. Six samples were obtained from each subject using the informative profile (block) randomized design. Variability (in terms of coefficient of variation, CV) in model parameters was varied between 30% and 100%, and residual variability was fixed at 15%. Sample sizes ranging from 30 to 1,000 subjects were studied, and a hundred replicate data sets were generated and analyzed with NONMEM for each sample size at each CV. A sample size of 30 was required for accurate and precise estimation of structural model parameters when CV < or = 75%, except for Cl where it is adequate for CV < or = 100%. A sample size of 80 was required for intersubject variability estimation with CV < or = 60%. Robust estimates of variability in Cl were obtained with sample sizes of 30 (CV < or = 45%), 60 (CV 60-75%), and 100 (CV > or = 75%). Positively biased estimates of residual variability were obtained irrespective of sample size at > or = 60% CV. This indicates that estimates of residual variability obtained in study situations where CV > or = 60% should be interpreted with caution. In such situations model misspecification may not be the issue, because in this simulation study concentration-time profiles were generated and analyzed with the same model. Although these results should be interpreted within the context of the study, they provide a framework for addressing the issue of sample size in longitudinal population pharmacokinetic study with a balanced sampling design. The result of a population pharmacokinetic study can be anticipated by comparing the results of several simulations in which the various input factors have been varied.

Computer Simulation↗

Improved variance estimation of classification performance via reduction of bias caused by small sample size.

BACKGROUND: Supervised learning for classification of cancer employs a set of design examples to learn how to discriminate between tumors. In practice it is crucial to confirm that the classifier is robust with good generalization performance to new examples, or at least that it performs better than random guessing. A suggested alternative is to obtain a confidence interval of the error rate using repeated design and test sets selected from available examples. However, it is known that even in the ideal situation of repeated designs and tests with completely novel samples in each cycle, a small test set size leads to a large bias in the estimate of the true variance between design sets. Therefore different methods for small sample performance estimation such as a recently proposed procedure called Repeated Random Sampling (RSS) is also expected to result in heavily biased estimates, which in turn translates into biased confidence intervals. Here we explore such biases and develop a refined algorithm called Repeated Independent Design and Test (RIDT). RESULTS: Our simulations reveal that repeated designs and tests based on resampling in a fixed bag of samples yield a biased variance estimate. We also demonstrate that it is possible to obtain an improved variance estimate by means of a procedure that explicitly models how this bias depends on the number of samples used for testing. For the special case of repeated designs and tests using new samples for each design and test, we present an exact analytical expression for how the expected value of the bias decreases with the size of the test set. CONCLUSION: We show that via modeling and subsequent reduction of the small sample bias, it is possible to obtain an improved estimate of the variance of classifier performance between design sets. However, the uncertainty of the variance estimate is large in the simulations performed indicating that the method in its present form cannot be directly applied to small data sets.

Analysis of Variance↗