Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Risch's lambda values for human obesity.

OBJECTIVE: Risch's lambda statistic (lambda R) is related to the heritability of traits and can be useful in several contexts, including the conduct of power analyses to determine sample size for gene mapping studies. However, values of lambda R have not been presented for human obesity. DESIGN AND RESULTS: Using both analytic and empirical approaches, the present study calculates estimates of lambda R. Examples are provided to illustrate the use of these estimates for determining sample size for genetic mapping studies.

Adolescent↗

The genetic structure of female life history in D. melanogaster: comparisons among populations.

Two questions were addressed: (1) What is the genetic variance-covariance structure of a suite of four female life history traits in D. melanogaster? and (2) Does the genetic architecture of these traits differ among populations? Three populations of D. melanogaster were studied. Genetic variances and covariances were estimated by sib analysis three times for each population: immediately upon establishment of populations in the laboratory, and subsequently after approximately 6 months and 2 years of laboratory culture. Entire genetic variance-covariance matrices, as well as their individual components, were compared between populations by means of likelihood ratio tests. All traits studied were significantly heritable in at least one-half of estimates. Despite large sample sizes, additive genetic covariances were for the most part not statistically significant, and only two significant negative covariance estimates were obtained throughout the experiments. Therefore, these experiments provide little support for evolutionary life history theories that are based on negative genetic correlations among life history components. Neither do they support the idea that genetic variance for fitness components is maintained by trade-offs. Evidence suggests that the G matrix of one population was initially different from those of the other two populations. Those differences disappeared after 2 years of laboratory culture. At the level of individual (co)variance components, there were relatively few differences among populations, and the overall impression was that the three populations had generally similar genetic architectures for the traits studied.

Animals↗

Performance of generalized estimating equations in practical situations.

Moment methods for analyzing repeated binary responses have been proposed by Liang and Zeger (1986, Biometrika 73, 13-22), and extended by Prentice (1988, Biometrics 44, 1033-1048). In their generalized estimating equations (GEE), both Liang and Zeger (1986) and Prentice (1988) estimate the parameters associated with the expected value of an individual's vector of binary responses as well as the correlations between pairs of binary responses. In this paper, we discuss one-step estimators, i.e., estimators obtained from one step of the generalized estimating equations, and compare their performance to that of the fully iterated estimators in small samples. In simulations, we find the performance of the one-step estimator to be qualitatively similar to that of the fully iterated estimator. When the sample size is small and the association between binary responses is high, we recommend using the one-step estimator to circumvent convergence problems associated with the fully iterated GEE algorithm. Furthermore, we find the GEE methods to be more efficient than ordinary logistic regression with variance correction for estimating the effect of a time-varying covariate.

Air Pollution↗

[An effective method for the estimation and comparison of the ED50 with small sample sizes].

In ED50 experiments the relationship between dose and probability of response is often modelled by the probit function. Standard statistical analysis estimates the parameters of this function by the maximum likelihood principle and derives the ED50 and its fiducial limits from these parameters. Bayesian analysis is more effective in two respects: It optionally includes prior information and in all but very few instances yields confidence intervals, whereas fiducial intervals often cannot be determined. Bayesian analysis of experiments with one substance has been treated in GRIEVE (1988). In the present article the mathematically interested reader is shown how to compare two substances. The probability of higher ED50 in the one substance as well as estimates of the ratio of the ED50's are obtained. The methods are easily extended to the effective dose for any other reasonable percentage of animals, e.g. ED90 or ED25. Experiments concerning lethal doses can be analysed by these methods as well. Both types of analysis are applied in two examples which compare new batches of vaccines with an established standard. In the first example both substances are nearly equivalent, while in the second example the new batch is considerably more efficient. An interactive FORTRAN program for a personal computer is available (cf. last section of 5.). It computes the maximum likelihood and the Bayesian solution, using approximate formulas in the latter case. Due to these approximations it was possible to develop a Bayesian program which is fast enough to run on a PC. Validation procedures have been performed. The output consists of a print file and, optionally, an ASCII file containing the coordinates of the posterior probability density and distribution functions.

Animals↗

Cytomorphometry. A methodologic study of preparation techniques, selection methods and sample sizes.

The influence of methodologic aspects on cytomorphometric features was studied using preparations of hepatoma and/or mastocytoma cells. First, two preparation techniques (smear and oese) were compared. Second, four methods of selecting cells for cytomorphometric analysis (two conventional and two stratified methods) were tested for reproducibility. Third, heterogeneous cell populations were used to estimate the required sample size using the running coefficient of variation (CV), and the results were compared with expected (theoretical) values of the required sample size calculated using the standard error of the mean. The results showed significantly lower CVs for the smear preparation technique. The stratified methods appeared to be superior to the conventional methods for selecting cells for measurement. The experimentally assessed sample sizes were considerably lower than the corresponding theoretical calculations. These findings suggest that morphometric assessments in cytologic smears should utilize a stratified cell selection method. While experimentally assessed sample sizes are relatively small and therefore better routinely applicable, they may yield less reliable results in some cases. The need to test a sample for its reproducibility as well as its discriminatory power is emphasized.

Animals↗

Modelling of mortality data from a multi-centre study in Japan by means of Poisson regression with error in variables.

BACKGROUND: Death rates of particular categories in epidemiological studies are often based on a small number of occurrences which can be well described by a Poisson distribution. METHOD: We applied this model for the analysis of a multi-centre study in five Japanese counties where the death rates of stomach cancer (ICD-9 code 151) in four age groups are known. In our example some covariates of the cases (e.g. plasma lycopene levels) are unknown values and are estimated from a randomly chosen collective. Therefore these values are subject to a sampling error. The inclusion of errors in variables (e-i-v) into the statistical model can adequately describe such a situation. The model is estimated in a Bayesian framework by means of resampling techniques. RESULTS: Based on the posterior distribution of the parameters the relative risk of stomach cancer is 0.46 (95% confidence interval: 0.23-0.79) comparing the maximum of the population medians of lycopene with the minimum. The estimated overdispersion is close to zero indicating only minor interference with other possible explanatory variables. In addition, we show that inclusion of e-i-v can give more accurate estimates of the parameters even from small sample sizes. CONCLUSIONS: Appropriate statistical methods allow the accurate estimation of relative risks from small sample sizes and from low number of cases. Lycopene plasma levels are good predictors for stomach cancer.

Adult↗

Estimation of a parameter and its exact confidence interval following sequential sample size reestimation trials.

For confirmatory trials of regulatory decision making, it is important that adaptive designs under consideration provide inference with the correct nominal level, as well as unbiased estimates, and confidence intervals for the treatment comparisons in the actual trials. However, naive point estimate and its confidence interval are often biased in adaptive sequential designs. We develop a new procedure for estimation following a test from a sample size reestimation design. The method for obtaining an exact confidence interval and point estimate is based on a general distribution property of a pivot function of the Self-designing group sequential clinical trial by Shen and Fisher (1999, Biometrics55, 190-197). A modified estimate is proposed to explicitly account for futility stopping boundary with reduced bias when block sizes are small. The proposed estimates are shown to be consistent. The computation of the estimates is straightforward. We also provide a modified weight function to improve the power of the test. Extensive simulation studies show that the exact confidence intervals have accurate nominal probability of coverage, and the proposed point estimates are nearly unbiased with practical sample sizes.

Biometry↗

Smooth estimation of the reliability function.

Problems with censored data arise quite frequently in reliability applications. Estimation of the reliability function is usually of concern. Reliability function estimators proposed by Kaplan and Meier (1958), Breslow (1972), are generally used when dealing with censored data. These estimators have the known properties of being asymptotically unbiased, uniformly strongly consistent, and weakly convergent to the same Gaussian process, when properly normalized. We study the properties of the smoothed Kaplan-Meier estimator with a suitable kernel function in this paper. The smooth estimator is compared with the Kaplan-Meier and Breslow estimators for large sample sizes giving an exact expression for an appropriately normalized difference of the mean square error (MSE) of the two estimators. This quantifies the deficiency of the Kaplan-Meier estimator in comparison to the smoothed version. We also obtain a non-asymptotic bound on an expected L1-type error under weak conditions. Some simulations are carried out to examine the performance of the suggested method.

Humans↗

The design of observer agreement studies with binary assessments.

We discuss the design of observer agreement studies with binary assessments, with particular emphasis on the need for adequate sample size and the use of replicate observations. First, we present a method and tables for determining the sample size required for ensuring a desired precision for the estimate of the probability of disagreement between two observers. Second, for studies including replicate observations, we present a statistical model that allows estimation of the magnitude of within- and between-observer variation. We then derive sample sizes guaranteeing a specified precision for these estimates, present tables of these sample sizes and give examples of their use.

Confidence Intervals↗

On stability designs in drug shelf-life estimation.

In this paper, various stability designs, including matrixing and bracketing designs for determining drug shelf-life, are considered. We propose a criterion for design selection based on the precision of drug shelf-life estimates. For a fixed sample size, it is recommended that the design with the best precision for estimating the shelf-life should be used. For a fixed desired precision, the design with the smallest sample size is the best choice of design. An example is presented to illustrate the proposed method.

Drug Stability↗

Fluorescent-based typing of the two short tandem repeat loci HUMTH01 and HUMACTBP2: reproducibility of size measurements and genetic variation in the Swedish population.

The aim of this study was to investigate the reproducibility of genetic typing of two tetrameric short tandem repeat (STR) loci and the extent of genetic variation in the Swedish population. An automated, fluorescent-based Applied Biosystems 373A sequencer was used for typing of the HUMTH01 and HUMACTBP2 loci (also named SE33). The former locus has seven alleles in the size range of 154-174 bp, while the latter is a complex locus with more than 32 alleles in the range of 227-316 bp. Using different fluorescent dyes, polymerase chain reaction (PCR) products from the two STR loci were sized in one lane using an internal size standard. In order to compare within- and between-gel reproducibility of fragment size estimates, a control sample was typed three times on each of 20 gels. Within the gel, the standard deviation (SD) of fragment size variability was less than 0.1 bp for four fragment sizes between 158-291 bp. Standard deviations between gels were slightly higher for the two shorter fragment sizes (HUMTH01), while the larger fragments varied between 0.3 and 0.4 bp (HUMACTBP2). The amount of genetic variation was investigated in samples from three Swedish cities (n = 301). Seven alleles were found at HUMTH01 and the observed heterozygosity was 0.77. At the HUMACTBP2 locus more than thirty alleles were found and the observed heterozygosity was 0.96. The observed genotype frequencies at HUMTH01 and HUMACTBP2 did not deviate significantly from Hardy-Weinberg expectations. No indication of a significant excess of homozygotes was found at any of the loci. We conclude that both HUMTH01 and HUMACTBP2 can be reliably typed using the method described. However, the latter locus requires an allelic ladder to be run on each gel.

Alleles↗

Radiography as primary outcome in rheumatoid arthritis: acceptable sample sizes for trials with 3 months' follow up.

OBJECTIVES: To investigate whether plain radiographs can show changes in joint damage due to rheumatoid arthritis (RA) within 3 months. METHODS: 188 film pairs taken with a 3 month interval were evaluated. They were scored with (chronological) and without (paired) knowledge of the sequence of the films according to the Sharp/van der Heijde method. Changes in joint damage were analysed on a group and an individual level for different subsets of patients. Sample sizes required to detect statistically and clinically significant differences were estimated based on the percentages of patients with progression larger than the smallest detectable change (SDC). RESULTS: Changes in joint damage were seen by both the chronological and the paired scoring method. The percentage of patients with progression of joint damage larger than the corresponding SDCs (1.7 and 2.4) varied in the subsets from 18% to 64% if based on the chronological change-scores and from 9% to 36% using paired change-scores. Acceptable sample size estimates were seen in several subsets, depending on (a) how the investigated drug would reduce the individual risk of progression of joint damage (by an absolute or a relative risk reduction model); (b) how damage was scored (chronological or paired); (c) the baseline risk; and (d) whether a two sided or one sided test would be used. CONCLUSIONS: Changes in joint damage due to RA can be detected reliably already within 3 months. This finding can be used to plan short term, randomised controlled trials with radiographic progression as primary outcome.

Adult↗

Estimation of the most likely number of individuals from commingled human skeletal remains.

This study examines quantification techniques applicable to human skeletal remains, and in particular the Lincoln index (LI), the minimum number of individuals (MNI), and what we refer to as the most likely number of individuals (MLNI), which is a modification of the LI by Chapman ([1951] Univ. Calif. Publ. Stat. 1:131-159). As part of the study, a test of pair-matching between commingled homologous elements, e.g., right and left femora, was performed based on gross morphology. The results show that pair-matching can be accurately performed, and that the MLNI is a useful technique for dealing with well-preserved commingled remains recovered from archaeological excavations and/or forensic investigations. Our results show that it is potentially misleading to draw population conclusions based on the MNI, except in instances where recovery is near 100%. The MLNI was found to be the best method to compensate for the potential underestimates of the MNI and potential bias in the original LI estimates resulting from small sample sizes. We demonstrate the use of MLNI in estimating the number of individuals from Lodge 21 at the Larson site, a late protohistoric structure at which the inhabitants were massacred and subsequently had their skeletal elements commingled by further taphonomic processes. We also show how to calculate estimates and standard errors for the recovery probabilities of skeletal elements.

Bone and Bones↗

An empirical evaluation of alternative methods of estimation for confirmatory factor analysis with ordinal data.

Confirmatory factor analysis (CFA) is widely used for examining hypothesized relations among ordinal variables (e.g., Likert-type items). A theoretically appropriate method fits the CFA model to polychoric correlations using either weighted least squares (WLS) or robust WLS. Importantly, this approach assumes that a continuous, normal latent process determines each observed variable. The extent to which violations of this assumption undermine CFA estimation is not well-known. In this article, the authors empirically study this issue using a computer simulation study. The results suggest that estimation of polychoric correlations is robust to modest violations of underlying normality. Further, WLS performed adequately only at the largest sample size but led to substantial estimation difficulties with smaller samples. Finally, robust WLS performed well across all conditions.

Factor Analysis, Statistical↗

Regression modeling of competing crude failure probabilities.

In a randomized trial of tamoxifen therapy for breast cancer, women can experience tumor recurrence or die from competing causes. One goal of analysis is to describe the effect of tamoxifen on the probabilities of recurrence or death from other causes. To this end, we propose a semi-parametric transformation model for the crude failure probabilities of a competing risk, conditional on covariates. The model is developed as an extension of the standard approach to survival data with independent right censoring. Estimation of the regression coefficients is achieved with a rank-based least squares criterion. Simulations show that the procedure works well with practical sample sizes. A separate estimating function is developed for the baseline parameter. Prediction of covariate-adjusted failure probabilities is considered. The methodology is motivated and illustrated with data from the tamoxifen trial.

Journal Article↗

Fine-scale genetic structure and gene dispersal inferences in 10 neotropical tree species.

The extent of gene dispersal is a fundamental factor of the population and evolutionary dynamics of tropical tree species, but directly monitoring seed and pollen movement is a difficult task. However, indirect estimates of historical gene dispersal can be obtained from the fine-scale spatial genetic structure of populations at drift-dispersal equilibrium. Using an approach that is based on the slope of the regression of pairwise kinship coefficients on spatial distance and estimates of the effective population density, we compare indirect gene dispersal estimates of sympatric populations of 10 tropical tree species. We re-analysed 26 data sets consisting of mapped allozyme, SSR (simple sequence repeat), RAPD (random amplified polymorphic DNA) or AFLP (amplified fragment length polymorphism) genotypes from two rainforest sites in French Guiana. Gene dispersal estimates were obtained for at least one marker in each species, although the estimation procedure failed under insufficient marker polymorphism, limited sample size, or inappropriate sampling area. Estimates generally suffered low precision and were affected by assumptions regarding the effective population density. Averaging estimates over data sets, the extent of gene dispersal ranged from 150 m to 1200 m according to species. Smaller gene dispersal estimates were obtained in species with heavy diaspores, which are presumably not well dispersed, and in populations with high local adult density. We suggest that limited seed dispersal could indirectly limit effective pollen dispersal by creating higher local tree densities, thereby increasing the positive correlation between pollen and seed dispersal distances. We discuss the potential and limitations of our indirect estimation procedure and suggest guidelines for future studies.

French Guiana↗

Improving estimates of genetic maps: a maximum likelihood approach.

As a result of previous large, multipoint linkage studies there is a substantial amount of existing marker data. Due to the increased sample size, genetic maps estimated from these data could be more accurate than publicly available maps. However, current methods for map estimation are restricted to data sets containing pedigrees with a small number of individuals, or cannot make full use of marker data that are observed at several loci on members of large, extended pedigrees. In this article, a maximum likelihood (ML) method for map estimation that can make full use of the marker data in a large, multipoint linkage study is described. The method is applied to replicate sets of simulated marker data involving seven linked loci, and pedigree structures based on the real multipoint linkage study of Abkevich et al. (2003, American Journal of Human Genetics 73, 1271-1281). The variance of the ML estimate is accurately estimated, and tests of both simple and composite null hypotheses are performed. An efficient procedure for combining map estimates over data sets is also suggested.

Algorithms↗

Foreign-born emigration: a new approach and estimates based on matched CPS files.

The utility of postcensal population estimates depends on the adequate measurement of four major components of demographic change: fertility, mortality, immigration, and emigration. Of the four components, emigration, especially of the foreign-born, has proved the most difficult to gauge. Without "direct" methods (i.e., methods identifying who emigrates and when), demographers have relied on indirect approaches, such as residual methods. Residual estimates, however are sensitive to inaccuracies in their constituent parts and are particularly ill-suited for measuring the emigration of recent arrivals. Here we introduce a new method for estimating foreign-born emigration that takes advantage of the sample design of the Current Population Survey (CPS): repeated interviews of persons in the same housing units over a period of 16 months. Individuals appearing in a first March Supplement to the CPS but not the next include those who died in the intervening year, those who moved within the country, and those who emigrated. We use statistical methods to estimate the proportion of emigrants among those not present in the follow-up interview. Our method produces emigration estimates that are comparable to those from residual methods in the case of longer-term residents (immigrants who arrived more than 10 years ago), but yields higher--and what appear to be more accurate--estimates for recent arrivals. Although somewhat constrained by sample size, we also generate estimates by age, sex, region of birth, and duration of residence in the United States.

Adolescent↗