Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Prevalence of known diabetes in Chennai City.

AIM: To determine prevalence of known diabetes in those more than 20 years of age in Chennai city. METHODOLOGY: Urban population was selected for the survey. Assuming the prevalence of known diabetes as 5.0% in those aged > 20 years, the cluster sample size calculated to estimate it with 95% CI and +/- 10% precision, was 25800 individuals of all ages. This population obtained from 200 households in each of 30 randomly selected corporation divisions of the city, was surveyed by social workers by house to house enquiry. General information and health status of every member of the household were recorded on prescribed forms. This survey was conducted during January-July, 1998. RESULTS: Among 26,066 individuals of all ages 779 had known diabetes and 99.4% of them had type 2 diabetes. The prevalence of known diabetes was 2.9% for all ages and both sexes combined. Crude and age-standardized prevalence was 4.9% (95% CI 4.6-5.2) for those aged > 20 years. The standardized prevalence was 10.5% (95% CI 9.8 - 11.2) in those aged > or = 40 years. The prevalence was significantly high (P < 0.05) in females. CONCLUSION: The prevalence of known diabetes was low in total population but increased in those aged > 20 and further increased in those aged > or = 40 years. The causes for high prevalence in > or = 40 year age group needs to be explored in this population.

Adolescent↗

The neurons of the retinal ganglion cell layer of the guinea pig: quantitative analysis of their distribution and size.

1. The topographical distribution of ganglion cells and displaced amacrine cells in the guinea pig retina is described. 2. Neurons were counted in the ganglion cell layer of retinal whole mounts stained by the method of Nissl or retrogradely labeled with horseradish peroxidase. Neuronal soma size was estimated from samples taken from different retinal regions. 3. We estimate that a total of 295,000 neurons comprise the guinea pig ganglion cell layer and they consist of 159,000 ganglion cells and 136,000 displaced amacrine cells. 4. The visual streak is poorly differentiated. Ganglion cell density reaches a peak of 2,272 cells/mm2 in a temporal expansion of the visual streak, 4-5 mm toward the optic disk. The visual streak temporal expansion may represent the analogue of the area centralis for this species. The ventral hemi-retina has a higher ganglion cell density than the dorsal hemi-retina. The displaced amacrine cells are more uniformly distributed than the ganglion cells. 5. The present paper provides relevant data concerning the number and distribution of the neurons of the retinal ganglion cell which were not available or were very contradictory in the literature.

Animals↗

[Relation between salt intake and blood pressure].

We investigated in a population group of 78 middle-aged men urinary sodium and potassium excretion per 24 hours under normal living conditions and the relationship to blood pressure. The mean urinary Na excretion in 24 hours was 241.3 mmol which corresponded to an intake of 14 g NaCl per day. The maximum NaCl consumption was as high as 30 g/day. The urinary Na excretion correlated in a linear fashion with the K excretion but not with the blood pressure. The blood pressure in the investigated group correlated with age, body mass index (kg/m2) and serum creatinine level. The body mass index and serum creatinine participated significantly in the blood pressure level, when multiple linear regression was used, where the urinary Na and K excretion per 24 hours were insignificant factors. In the investigated group a direct correlation of Na intake and blood pressure was not found, the investigation provided, however, information on the variability of urinary Na and K excretion/24 hours in the population and enables to estimate the necessary sample size to resolve this problem.

Adult↗

Preparations for AIDS vaccine trials. Retention, behavior change, and HIV-seroconversion among injecting drug users (IDUs) and sexual partners of IDUs.

The likelihood that subjects in human immunodeficiency virus (HIV) vaccine efficacy trials will alter their behavioral risks for HIV infection over time must be considered in evaluating the feasibility of such trials and in estimating the necessary sample sizes to be enrolled. Potential subjects for future vaccine efficacy trials include injecting drug users (IDUs) and others who may be difficult to retain in studies and who may alter HIV-risk-related behaviors substantially over time. We have investigated behavior change, retention, and HIV seroconversion among 577 New York City resident IDUs and sexual partners of IDUs enlisted between July 1 and December 31, 1992. We attempted to see all subjects every 3 months for interviews, blood donation and HIV testing. We were able to retain 68% of subjects in the study through the third scheduled recall at 7.5-10.5 months after enlistment. HIV seroconversion through March 1, 1994, was 1.33/100 person-years at risk. There was a significant inverse relationship between HIV seroconversion and retention at the 9-month recall after adjusting for age, gender, and the amount of locator information provided by subjects at enlistment. Among subjects seen at each of the scheduled visits at 3, 6, and 9 months after enrollment, modest but statistically significant behavior changes that reduced risk were observed in self-reported drug injection frequency, heroin injection frequency, sexual contact with IDUs, and sharing of needles/syringes. The magnitude of these changes in risk, however, was small and may be transient. The behavior changes observed to date do not appear to be large enough to substantially alter calculations of sample sizes needed in future HIV vaccine efficacy trials.

AIDS Vaccines↗

[Critical examination of scoring systems in therapeutic trials].

Scoring systems give a check-list and methodological informations which have to be found in controlled therapeutic trials reports and papers. These systems try to quantify each item to give a global score. The Chalmer's list is the most wellknown. It allows a balance in scoring taking in account the quality of the endpoints. Other lists are more simple. Many check-lists allow the scoring of the methodological design or the statistical analysis. In all systems the major methodological points are: the randomization, the description of the population, the double blind, the estimation of the sample size, the handling of withdrawal and drop out, the major endpoint, the patients follow-up, the statistical analysis and the data presentation. All these scoring systems have several limits: the quantitative evaluation of each item is subjective and the point scoring has never been validated, some scoring systems are old and don't integrate new methodological methods, the scores never included the clinical interest of the trial, some items are questionable, others are forgotten (intention to treat analysis, steering comity...). Scoring systems allow a control of the methodological quality of clinical trials but don't include the clinical or scientific interest of the study. These systems are a useful methodological tool for publication process in medical journals and for new drugs authorization. The evaluation by authors themselves of the quality of their papers using a standardized scoring system could clarify the reviewers decisions.

Clinical Trials as Topic↗

Randomized observational studies on the economics of therapies--biometrical experience of two trials.

Economic studies in medicine are intended to investigate costs, associated with a particular problem dealing with the indication, diagnosis or therapy, for instance, whether the high costs involved in a highly intensive or innovative therapy could be balanced by the eventual savings made, due to the shorter periods of treatment. In such situations a randomized controlled trial is necessary to find out which therapy or which therapeutical strategy is least expensive in the long run. Economic studies do, however, present some specific problems. Making a list of all the cost-relevant treatment items can be very laborious, but the use of flat rates and lump sums alone cannot lead to a complete cost analysis. Often, costs between hospitals vary more than between treatment regimens. Early and sudden deaths incur low costs and may bias the results. Furthermore, costs are distributed with a long and heavy upper tail including extreme outliers. This does, in fact, complicate the estimation of the sample size. In this article, these problems are outlined and, with the help of the data obtained from two randomized economic trials in health care, solutions are proposed and discussed.

Costs and Cost Analysis↗

The prevalence of endometriosis in women with chronic pelvic pain.

BACKGROUND: The 2004 American College of Obstetrics and Gynecology clinical management guideline states that the prevalence of endometriosis is approximately 33% in women with chronic pelvic pain (CPP). This estimate came from a review showing that 28% of adult women with CPP were found to have endometriosis. The prevalence of 28% in adult women was arrived based on a compilation of 11 published studies. Yet even within the 11 studies, the reported prevalence of endometriosis varies wildly, ranging from 2 to 74%. Such an astounding variation or heterogeneity raises the question whether it is appropriate to use a single prevalence of endometriosis for all women with CPP. METHODS: We sought to identify possible sources of heterogeneities in the estimation of prevalence of endometriosis in women with CPP. We included more studies that reported prevalence estimates than the review, and examined the effect of sample size and the year of publication on the heterogeneity. RESULTS: The year of publication is positively associated with the prevalence estimate, which may indicate an increasing awareness of various appearances of endometriosis, or the prevalence of endometriosis may have increased among women with CPP. An alternative analysis with removal of four studies reporting highest prevalence estimates indicated that sample size is negatively associated with the prevalence estimates while the year of publication became only marginally significant. CONCLUSIONS: There are identifiable sources of heterogeneity in prevalence estimates, with the year of publication, sample size, and difference in evaluation of CPP being three apparent sources. Having a single prevalence estimate for all women with CPP may be too simplistic at best. The true prevalence is very likely to be higher than 33%.

Adult↗

Self-designing two-stage trials to minimize expected costs.

In the design of clinical trials, the sample size for the trial is traditionally calculated from estimates of parameters of interest, such as the mean treatment effect, which can often be inaccurate. However, recalculation of the sample size based on an estimate of the parameter of interest that uses accumulating data from the trial can lead to inflation of the overall Type I error rate of the trial. The self-designing method of Fisher, also known as the variance-spending method, allows the use of all accumulating data in a sequential trial (including the estimated treatment effect) in determining the sample size for the next stage of the trial without inflating the Type I error rate. We propose a self-designing group sequential procedure to minimize the expected total cost of a trial. Cost is an important parameter to consider in the statistical design of clinical trials due to limited financial resources. Using Bayesian decision theory on the accumulating data, the design specifies sequentially the optimal sample size and proportion of the test statistic's variance needed for each stage of a trial to minimize the expected cost of the trial. The optimality is with respect to a prior distribution on the parameter of interest. Results are presented for a simple two-stage trial. This method can extend to nonmonetary costs, such as ethical costs or quality-adjusted life years.

Bayes Theorem↗

Trials which randomize practices II: sample size.

BACKGROUND: When practices are randomized in a trial and observations are made on the patients to assess the relative effectiveness of the different interventions, sample size calculations need to estimate the number of practices required, not just the total number of patients. OBJECTIVE: Our aims were to introduce the methodology for appropriate sample size calculation and discuss the implications for power. METHOD: A worked example from general practice is used. DISCUSSION: Designs which randomize practices are less powerful than designs which randomize patients to intervention groups, particularly where a large number of patients is recruited from each practice. Studies which randomize few practices should be avoided if possible, as the loss of power is considerable and simple randomization may not ensure comparability of intervention groups.

Family Practice↗

Sample size calculations for a split-cluster, beta-binomial design in the assessment of toxicity.

Mouse embryo assays are recommended to test materials used for in vitro fertilization for toxicity. In such assays, a number of embryos is divided in a control group, which is exposed to a neutral medium, and a test group, which is exposed to a potentially toxic medium. Inferences on toxicity are based on observed differences in successful embryo development between the two groups. However, mouse embryo assays tend to lack power due to small group sizes. This paper focuses on the sample size calculations for one such assay, the Nijmegen mouse embryo assay (NMEA), in order to obtain an efficient and statistically validated design. The NMEA follows a stratified (mouse), randomized (embryo), balanced design (also known as a split-cluster design). We adopted a beta-binomial approach and obtained a closed sample size formula based on an estimator for the within-cluster variance. Our approach assumes that the average success rate of the mice and the variance thereof, which are breed characteristics that can be easily estimated from historical data, are known. To evaluate the performance of the sample size formula, a simulation study was undertaken which suggested that the predicted sample size was quite accurate. We confirmed that incorporating the a priori knowledge and exploiting the intra-cluster correlations enable a smaller sample size. Also, we explored some departures from the beta-binomial assumption. First, departures from the compound beta-binomial distribution to an arbitrary compound binomial distribution lead to the same formulas, as long as some general assumptions hold. Second, our sample size formula compares to the one derived from a linear mixed model for continuous outcomes in case the compound (beta-)binomial estimator is used for the within-cluster variance.

Animals↗

Relationship between prevalence and intensity of Plasmodium falciparum infection in natural populations of Anopheles mosquitoes.

Wild-caught Anopheles gambiae s. l. and An. funestus were dissected and their midguts were examined for the presence of Plasmodium falciparum oocyst infections. The mean intensity of infection and the prevalence of infected mosquitoes were determined for each sample, with one sample representing the mosquitoes caught in a single house at any given time. The patterns of infection were investigated using the relationships between prevalence, intensity, and variance within samples, and were found to be consistent with laboratory infections. The overall distribution of oocysts is characterized by a mixture of negative binomial distributions with means determined by the infectiousness of the human hosts, and a constant degree of aggregation (k = 0.0767) presumably determined by the development of oocysts within mosquitoes. The prevalence/intensity relationship was treated as a bivariate distribution to ascertain the effect of sample size on the accuracy of estimation, and to allow inference of intensity from prevalence. In mathematical models fitted to the collected data, sample size affected directly the minimum possible prevalence of infection, and the accuracy of both mean and prevalence estimations. Based on minimum possible positive prevalence rates, data from samples of less than 20-25 mosquitoes would provide unacceptable errors in prevalence estimations. However, natural oocyst rates are consistently higher than the minimum prevalence, and it is suggested that any interpretations from samples of less than approximately 40 mosquitoes must be treated with some caution. Such variation in natural samples means that prediction of intensity of infection from prevalence (or vice versa) is extremely inaccurate.

Animals↗

Parameter estimation for the calibration and variance stabilization of microarray data.

We derive and validate an estimator for the parameters of a transformation for the joint calibration (normalization) and variance stabilization of microarray intensity data. With this, the variances of the transformed intensities become approximately independent of their expected values. The transformation is similar to the logarithm in the high intensity range, but has a smaller slope for intensities close to zero. Applications have shown better sensitivity and specificity for the detection of differentially expressed genes. In this paper, we describe the theoretical aspects of the method. We incorporate calibration and variance-mean dependence into a statistical model and use a robust variant of the maximum-likelihood method to estimate the transformation parameters. Using simulations, we investigate the size of the estimation error and its dependence on sample size and the presence of outliers. We find that the error decreases with the square root of the number of probes per array and that the estimation is robust against the presence of differentially expressed genes. Software is publicly available as an R package through the Bioconductor project (http://www.bioconductor.org).

Journal Article↗

Nonparametric analysis of clustered ROC curve data.

Current methods for estimating the accuracy of diagnostic tests require independence of the test results in the sample. However, cases in which there are multiple test results from the same patient are quite common. In such cases, estimation and inference of the accuracy of diagnostic tests must account for intracluster correlation. In the present paper, the structural components method of DeLong, DeLong, and Clarke-Pearson (1988, Biometrics 44, 837-844) is extended to the estimation of the Receiver Operating Characteristics (ROC) curve area for clustered data, incorporating the concepts of design effect and effective sample size used by Rao and Scott (1992, Biometrics 48, 577-585) for clustered binary data. Results of a Monte Carlo simulation study indicate that the size of statistical tests that assume independence is inflated in the presence of intracluster correlation. The proposed method, on the other hand, appropriately handles a wide variety of intracluster correlations, e.g., correlations between true disease statuses and between test results. In addition, the method can be applied to both continuous and ordinal test results. A strategy for estimating sample size requirements for future studies using clustered data is discussed.

Angiography, Digital Subtraction↗

Parameter recovery for the partial credit model using MULTILOG.

This study investigated parameter recovery for the partial credit model using the MULTILOG computer program. Factors studied were the sample size and the number of item parameters, which were manipulated by systematically varying the number of steps per item and the number of items. The findings suggest that the ratio of sample size to number of item parameters being estimated as a "rule of thumb" can be a more complete guideline when the number of steps per item is taken into account. Accurate estimation of ability can be obtained across all conditions, even with sample sizes as small as 250. With regard to estimation of step values, however, more caution is warranted. Accurate estimation of the step values of items which have more categories requires larger sample sizes for a given number of total parameters to be estimated.

Humans↗

Total number and mean size of alveoli in mammalian lung estimated using fractionator sampling and unbiased estimates of the Euler characteristic of alveolar openings.

Estimation of alveolar number in the lung has traditionally been done by assuming a geometric shape and counting alveolar profiles in single, independent sections. In this study, we used the unbiased disector principle to estimate the Euler characteristic (and thereby the number) of alveolar openings in rat lungs and rhesus monkey lung lobes and to obtain robust estimates of average alveolar volume. The estimator of total alveolar number was based on systematic, uniformly random sampling using the fractionator sampling design. The number of alveoli in the rat lung ranged from 17.3 x 10(6) to 24.6 x 10(6), with a mean of 20.1 x 10(6). The average number of alveoli in the two left lung lobes in the monkey ranged from 48.8 x 10(6) to 67.1 x 10(6) with a mean of 57.7 x 10(6). The coefficient of error due to stereological sampling was of the order of 0.06 in both rats and monkeys and the biological variation (coefficient of variance between individuals) was 0.15 in rat and 0.13 in monkey (left lobe, only). Between subdivisions (left/right in rat and cranial/caudal in monkey) there was an increase in variation, most markedly in the rat. With age (2-13 years) the alveolar volume increased 3-fold (as did parenchymal volume) in monkeys, but the alveolar number was unchanged. This study illustrates that use of the Euler characteristic and fractionator sampling is a robust and efficient, unbiased principle for the estimation of total alveolar number in the lung or in well-defined parts of it.

Age Factors↗

Revisiting proportion estimators.

Proportion estimators are quite frequently used in many application areas. The conventional proportion estimator (number of events divided by sample size) encounters a number of problems when the data are sparse as will be demonstrated in various settings. The problem of estimating its variance when sample sizes become small is rarely addressed in a satisfying framework. Specifically, we have in mind applications like the weighted risk difference in multicenter trials or stratifying risk ratio estimators (to adjust for potential confounders) in epidemiological studies. It is suggested to estimate p using the parametric family p(c) and p(1 - p) using p(c)(1 - p(c)), where p(c) = (X + c)/(n + 2c). We investigate the estimation problem of choosing c > or = 0 from various perspectives including minimizing the average mean squared error of p(c), average bias and average mean squared error of p(c)(1 - p(c)). The optimal value of c for minimizing the average mean squared error of p(c) is found to be independent of n and equals c = 1. The optimal value of c for minimizing the average mean squared error of p(c)(1 - p(c)) is found to be dependent of n with limiting value c = 0.833. This might justify to use a near-optimal value of c = 1 in practice which also turns out to be beneficial when constructing confidence intervals of the form p(c)+/-1.96 square root of np(c)(1 - p(c))/(n + 2c).

Bias↗

Estimation of usual intakes: What We Eat in America-NHANES.

Usual intakes of nutrients are reliable indicators for making associations between diet and health or disease risks. Estimates of consumption of specific foods and food groups are also important for evaluating the progress in meeting key objectives in such national public health initiatives as Healthy People 2010. Reliable and valid estimates of intakes of particular foods, food ingredients, dietary supplements and other bioactive substances are also needed for dietary assessment and regulatory purposes. The ability to generate useful estimates of these constituents often requires much larger sample sizes than are needed for estimating nutrient intakes. Statistical methods recommended by the National Academy of Sciences are described that provide estimates of distributions of usual nutrient intakes and permit dietary assessment and planning at the population level. Statistical and modeling approaches for estimating intakes of foods, dietary supplements and other bioactive substances are also summarized. Based on the deliberations of discussion groups consisting of members of key stakeholder groups involved in the planning, implementation and utilization of national survey data, a high priority was placed on the need for more research to determine the best approaches for applying these methods to dietary data in the integrated What We Eat in America-National Health and Nutrition Examination Survey (NHANES).

Aged↗

Impact of misclassification in genotype-exposure interaction studies: example of N-acetyltransferase 2 (NAT2), smoking, and bladder cancer.

Errors in genotype determination can lead to bias in the estimation of genotype effects and gene-environment interactions and increases in the sample size required for molecular epidemiologic studies. We evaluated the effect of genotype misclassification on odds ratio estimates and sample size requirements for a study of NAT2 acetylation status, smoking, and bladder cancer risk. Errors in the assignment of NAT2 acetylation status by a commonly used 3-single nucleotide polymorphism (SNP) genotyping assay, compared with an 11-SNP assay, were relatively small (sensitivity of 94% and specificity of 100%) and resulted in only slight biases of the interaction parameters. However, use of the 11-SNP assay resulted in a substantial decrease in sample size needs to detect a previously reported NAT2-smoking interaction for bladder cancer: 1,121 cases instead of 1,444 cases, assuming a 1:1 case-control ratio. This example illustrates how reducing genotype misclassification can result in substantial decreases in sample size requirements and possibly substantial decreases in the cost of studies to evaluate interactions.

Alleles↗