Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Foreign-born emigration: a new approach and estimates based on matched CPS files.

The utility of postcensal population estimates depends on the adequate measurement of four major components of demographic change: fertility, mortality, immigration, and emigration. Of the four components, emigration, especially of the foreign-born, has proved the most difficult to gauge. Without "direct" methods (i.e., methods identifying who emigrates and when), demographers have relied on indirect approaches, such as residual methods. Residual estimates, however are sensitive to inaccuracies in their constituent parts and are particularly ill-suited for measuring the emigration of recent arrivals. Here we introduce a new method for estimating foreign-born emigration that takes advantage of the sample design of the Current Population Survey (CPS): repeated interviews of persons in the same housing units over a period of 16 months. Individuals appearing in a first March Supplement to the CPS but not the next include those who died in the intervening year, those who moved within the country, and those who emigrated. We use statistical methods to estimate the proportion of emigrants among those not present in the follow-up interview. Our method produces emigration estimates that are comparable to those from residual methods in the case of longer-term residents (immigrants who arrived more than 10 years ago), but yields higher--and what appear to be more accurate--estimates for recent arrivals. Although somewhat constrained by sample size, we also generate estimates by age, sex, region of birth, and duration of residence in the United States.

Adolescent↗

Risk ratio and rate ratio estimation in case-cohort designs: hypertension and cardiovascular mortality.

Multivariate analysis in case-base designs depends on approximate methods. In the present study, new pseudo-likelihood methods are developed for this design. With these methods, the case-cohort risk ratio and rate ratio as well as their standard errors are easily estimated using logistic regression and Poisson regression, respectively. This is illustrated by the association between hypertension and cardiovascular mortality in a cohort, estimated by case-cohort analysis, using samples of several sizes. The estimates are compared with those obtaining in full-cohort and nested case-control designs. The results indicate that these methods, which require nothing but widely available computer software, are valid. The case-cohort design, therefore, is a good, sometimes even advantageous alternative to the nested case-control design, in studying a disease that is not very rare. Application of the risk ratio method to the full cohort, using a 'sample' of 100 per cent follows logically; whenever the true risk ratio is desired instead of the odds ratio, a multivariate model for its estimation is therefore available.

Adult↗

Are sample sizes usually at least an order of magnitude too low for reliable estimates of leaf asymmetry?

Estimates of leaf size and asymmetry for individual trees are often obtained using sample sizes that are too small to take into account the possibility that size and asymmetry may be affected by the position of the leaf on the tree. This issue was addressed by exploring variation in leaf size and asymmetry within an individual of Alder (Alnus glutinosa). We found differences between branches for leaf size and for signed asymmetry but not for unsigned asymmetry. We also found that the size of a leaf was not correlated with its position on a branch and that the asymmetry of a leaf was not correlated with either its position on a branch or with the asymmetry of its neighbour. Repeated subsampling of a sample of 870 leaves showed that a subsample size approaching 500 leaves was required for consistently reliable estimates of the standard deviation of unsigned asymmetry. Smaller subsamples were required for consistently reliable estimates of mean unsigned asymmetry and of the mean and standard deviation of leaf size, but subsamples of less than 100 leaves provided consistently reliable estimates only of mean leaf size. For this species, reliable estimates of an individual's level of asymmetry are obtained only if several hundred leaves are sampled over several branches, but it is not necessary to sample the same sequence of leaves from each branch.

Analysis of Variance↗

Sources of heterogeneities in estimating the prevalence of endometriosis in infertile and previously fertile women.

OBJECTIVE: To identify possible sources of heterogeneity in estimates of the prevalence of endometriosis in previously fertile women and in women with infertility. DESIGN: A pooled analysis of previously published studies reporting prevalence estimates. SETTING: Academic. PATIENTS: None. INTERVENTIONS: None. MAIN OUTCOME MEASURES: None. RESULTS: There were tremendous heterogeneities in prevalence estimates for both the fertile and infertile groups. In addition, the prevalence estimates increased with the year of publication, but decreased with sample size. For previously fertile women, the heterogeneity in prevalence estimates was no longer significant after the effects of sample size and year of publication on estimates were taken into account. In the infertile group, however, there was still a sizeable heterogeneity unaccounted for after both sample size and year of publication were taken into account. CONCLUSIONS: A single prevalence estimate for the entire fertile or infertile group may be too simplistic at best. More precise prevalence estimates, likely to be age-dependent, await carefully designed and executed studies that will also record covariates such as age at surgery and referral patterns.

Adolescent↗

Maximizing the entropy of histogram bar heights to explore neural activity: a simulation study on auditory and tactile fibers.

Neurophysiologists often use histograms to explore patterns of activity in neural spike trains. The bin size selected to construct a histogram is crucial: too large bin widths result in coarse histograms, too small bin widths expand unimportant detail. Peri-stimulus time (PST) histograms of simulated nerve fibers were studied in the current article. This class of histograms gives information about neural activity in the temporal domain and is a density estimate for the spike rate. Scott's rule based on modem statistical theory suggests that the optimal bin size is inversely proportional to the cube root of sample size. However, this estimate requires a priori knowledge about the density function. Moreover, there are no good algorithms for adaptive-mesh histograms, which have variable bin sizes to minimize estimation errors. Therefore, an unconventional technique is proposed here to help experimenters in practice. This novel method maximizes the entropy of histogram-bar heights to find the unique bin size, which generates the highest disorder in a histogram (i.e., the most complex histogram), and is useful as a starting point for neural data mining. Although the proposed method is ad hoc from a density-estimation point of view, it is simple, efficient and more helpful in the experimental setting where no prior statistical information on neural activity is available. The results of simulations based on the entropy method are also discussed in relation to Ellaway's cumulative-sum technique, which can detect subtle changes in neural activity in certain conditions.

Algorithms↗

Dose-effect approaches to risk assessment.

Risk assessment is the attempt to characterize the chance of obtaining an adverse effect after exposure to an agent. Traditionally, high levels of an agent have been used to estimate the likelihood a lower dose might have an effect either by using low-dose extrapolation models or by attempting to establish a dose with no observable effects (NOEL). Low-dose extrapolation models yield estimates for small effects, but these estimates may vary by orders of magnitude depending upon the function chosen to represent the data. NOEL's are imprecise because a true no-effect level is indeterminant and the inability to determine an observable effect depends primarily on background variability. Newer methods use data from portions of the dose-effect function where error is smaller to estimate risks. Risk estimates using two of these approaches are compared for two different sample sizes. Each method produced the same estimate with the larger samples at low risk, but with increasing levels of risk and smaller samples the estimates obtained using these methods diverged.

Animals↗

The effect of partial noncompliance on the power of a clinical trial.

Noncompliance is an important concern in randomized trials and must be taken into account when calculating sample sizes. The effect of noncompliance is usually assessed through a simple binary model that assumes that the patient either does or does not comply with the allocated treatment. In this article I present an example from cancer prevention, and show that different time courses of noncompliance can have different effects on the power, with widely varying estimates of the sample size required. The time course as well as the level of noncompliance should be considered when planning trials with long-term treatments.

Aged↗

The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed.

BACKGROUND AND OBJECTIVE: Publication bias and other sample size effects are issues for meta-analyses of test accuracy, as for randomized trials. We investigate limitations of standard funnel plots and tests when applied to meta-analyses of test accuracy and look for improved methods. METHODS: Type I and type II error rates for existing and alternative tests of sample size effects were estimated and compared in simulated meta-analyses of test accuracy. RESULTS: Type I error rates for the Begg, Egger, and Macaskill tests are inflated for typical diagnostic odds ratios (DOR), when disease prevalence differs from 50% and when thresholds favor sensitivity over specificity or vice versa. Regression and correlation tests based on functions of effective sample size are valid, if occasionally conservative, tests for sample size effects. Empirical evidence suggests that they have adequate power to be useful tests. When DORs are heterogeneous, however, all tests of funnel plot asymmetry have low power. CONCLUSION: Existing tests that use standard errors of odds ratios are likely to be seriously misleading if applied to meta-analyses of test accuracy. The effective sample size funnel plot and associated regression test of asymmetry should be used to detect publication bias and other sample size related effects.

Diagnostic Errors↗

[Clinical trials on glaucoma: differences according to the medical or surgical nature of the treatment being evaluated].

PURPOSE: To compare the quality of clinical trials on glaucoma between those evaluating the effectiveness of medical treatments and those evaluating surgical treatments. METHOD: Clinical trials on glaucoma published in seven international journals between January 1980 and December 1999 were selected. The papers were revised by researchers with a background in epidemiology using a standard qualitative questionnaire. Proportions were compared using Fisher's exact test. RESULTS: Sample size was pre-estimated in 19% of medical treatment trials and 2% of surgical trials (p=0.005); masking (72% vs. 9%; p<0.001) and intention-to-treat analysis (17 vs. 0 papers; p<0.001) were also more frequent in medical trials. Only 50% of the trials correctly described the patient flow. CONCLUSIONS: Quality in clinical trials on glaucoma medical treatment was higher than in surgical trials regarding sample size pre-estimation, masking and intention-to-treat analysis. However, both medical and surgical trials should improve in these aspects and in the patient flow description

Bibliometrics↗

Simple analytic procedures for rapid microcomputer-assisted cluster surveys in developing countries.

Surveys are often deemed necessary in developing countries when routine sources of data are not considered adequate to answer important policy-related questions. Although field work often goes smoothly, many surveys become bogged down in the analysis stage. With the availability of microcomputers and contemporary software, investigators in developing countries can use rapid survey methodology (RSM) to process, analyze, and report survey findings more quickly than ever before. Presented in this paper are three simple analytic procedures for planning and doing two-stage, rapid cluster surveys. All were successfully used in three rapid surveys in rural regions of Burma and Thailand. By use of a spreadsheet and graphics software package, the three procedures (a) derive the first-stage selection of 30 cluster sites with probability proportionate to size, (b) calculate variance estimates and confidence limits for the parameters of interest and graphically present the findings as 90, 95, and 99 percent confidence intervals, and (c) estimate the necessary sample size for planning two-stage, rapid cluster surveys. The procedures can be used both in the field and in teaching workshops or courses on survey methods. Examples are given from three rapid surveys conducted in Hlegu Township, Burma, and Sisaket Province, Thailand. In both countries, local health professionals were first taught the methods in a 1-week workshop before they used the procedure for conducting the rapid computer-assisted surveys.

Analysis of Variance↗

Brain space for a learned task: strong intraspecific evidence for neural correlates of singing behavior in songbirds.

There is a controversial issue in neuroscience whether the expansion of neural network space permits the development of more complex behavior. One of the best-known model systems for studying the relationship between brain space and behavior is song production and the associated song control system in songbirds. Although the neuroanatomical background of song production is well established, the direct link between song nuclei volumes and song traits remains puzzling. Analyses within species have provided conflicting results regarding the association between song nuclei volumes and measures of song complexity and song length. Based on a meta-analysis, we present here the results of the first synthetic review, in which we test for overall intraspecific patterns in relation to bird song and the size of associated neural tissues. We found significant positive relationships between the volume of two important song nuclei (HVC and RA) and repertoire size and song length. We assessed the importance of absolute and relative volumes, and found that a control for the covariation with the telencephalon may be important. By estimating the adequate sample size that would be needed to reach sufficient statistical power in particular studies, we conclude that previous studies finding non-significant associations between song and volumes of brain nuclei were of weak power. When we factored out the covariation between song length and repertoire size, we found that these traits may explain independently significant amount of variations in the relative volume of HVC, but not of RA. The link between the volumes of song nuclei and song features has important theoretical implications with regard to the neurobiology and evolution of bird song.

Animals↗

REPLI: a program in BASIC for determination of approximate sample size.

REPLI, a program written in elementary BASIC, calculates the approximate sample size, which is required to detect a desired difference between any two group means in an experiment with n groups for a given probability and at three significance levels of the means difference. A prior knowledge of the variability of data in the groups is expected in order to base the estimate on a rational footing. If this knowledge does not exist, an educated guess and/or several trials with different assumptions on the most likely variability can be used. The a priori estimate prevents that sample sizes are completely out of a reasonable range. Since the program is applicable for experimental settings where several groups need be investigated, it is particularly interesting to users of ANOVA and/or comparable non-parametric tests.

Analysis of Variance↗

Multilevel modeling and practice-based research.

PURPOSE: The health care system in the United States is inherently hierarchical. Patients are "nested" within physicians who in turn are "nested" within practices. Much of the research data gathered in practice-based research networks (PBRNs) also have similar patterns of nesting (clustering). When research data are nested, statistical approaches to the data must account for the multilevel nature of the data or risk errors in interpretation. We illustrate the concept of multilevel structure and provide examples with implications for practice-based research. METHODS: We present a selection of multilevel (hierarchical) models and contrast them with traditional linear regression models, using an example of a simulated observational study to illustrate increasingly complex statistical approaches, as well as to explore the consequences of ignoring clustering in data. Additionally, we discuss other types of outcome data and designs, and the effects of clustering on sample size and power. RESULTS: Multilevel models demonstrate that the effects of physician-level activities may differ from clinic to clinic as well as between rural and urban settings; this variability would be undetected in traditional linear regression approaches. Study conclusions differed when the data were analyzed with multilevel methods compared with traditional linear regression methods. Clustered data also affected sample size; as the intraclass correlation increased and the patients per cluster increased, the required number of patients increased dramatically. CONCLUSIONS: Recognizing and accounting for multilevel structure when analyzing data from PBRN studies can lead to more accurate conclusions, as well as offer opportunities to explore contextual effects and differences across sites. Accommodating multilevel structure in planning research studies can result in more appropriate estimation of required sample size.

Biomedical Research↗

Conotoxins - new vistas for peptide therapeutics.

There are approximately 500 species of predatory cone snails within the genus Conus. They comprise what is arguably the largest single genus of marine animals alive today. It has been estimated that the venom of each Conus species has between 50 and 200 components. These highly constrained sulfur rich components or conotoxins represent a unique arsenal of neuropharmacologically active peptides that have been evolutionarily tailored to afford unprecedented and exquisite selectivity for a wide variety of ion-channel subtypes. Remarkable divergence occurs when cone snails speciate. Consequently, the complement of venom peptides in any one Conus species is distinct from that of any other species. Hence many thousands of peptides that modulate ion channel function are present within Conus venoms. Evolutionary pressures have afforded a "pre-optimized," structurally sophisticated library that has been "fine tuned" over 50 million years. The statistics associated with sampling such libraries bear testimony to the validity and feasibility of this strategy. Although approximately 100 conotoxin sequences have been published in the scientific literature, representing a mere 0.2 % of the estimated library size, this sample has already afforded a peptide of proven clinical utility and several pre-clinical leads for CNS disorders. Conus libraries represent a rich pharmacopoeia and the potential to "therapeutically mine" such a resource appears limitless. The paucity of synthetic methodologies necessary to achieve the regioisomeric folding patterns present in these native peptides precludes access to synthetic conotoxin libraries, further validating the overall "mining" strategy. In this article, we will present a pragmatic overview of the molecular diversity as well as the neurobiological mechanisms that define each major class of conotoxin.

Amino Acid Motifs↗

Some properties of r equivalent: a simple effect size indicator.

One version of r equivalent, calculated from Fisher's exact test p values and recommended for small samples, is considered "a more realistic . . . [and] a more accurate estimate of the population correlation than . . . the sample correlation, r sample" (R. Rosenthal & D. B. Rubin, 2003, p. 494). Small sample properties of r sample and of two effect size estimators (r equivalent* and r hybrid) that use r equivalent were examined: r sample is preferable to r equivalent* (defined as r equivalent used without restrictions) in terms of bias and mean squared error (MSE); r hybrid (defined as r equivalent only when r sample = 1.0) is generally preferable to r equivalent*, and preferable to r sample in terms of MSEs, except when population correlations are very large. Conditions favoring r sample over r equivalent* and r hybrid in meta-analyses are noted.

Analysis of Variance↗

Within-plant distribution of twospotted spider mites (Acari: Tetranychidae) on impatiens: development of a presence-absence sampling plan.

The twospotted spider mite, Tetranychus urticae Koch, is an important pest of impatiens, a floricultural crop of increasing economic importance in the United States. The large amount of foliage on individual impatiens plants, the small size of mites, and their ability to quickly build high populations make a reliable sampling method essential when developing a pest management program. In our study, we were particularly interested in using spider mite counts as a basis for releasing biological control agents. The within-plant distribution of mites was established in greenhouse experiments and these data were used to identify the sampling unit. Leaves were divided into three zones according to location on the plant: inner, intermediate, and other. On average, 40, 33, and 27% of the leaves belonged to the inner, intermediate, and other leaf zones, respectively. However, because 60% of the mites consistently were found on the intermediate leaves, intermediate leaves were chosen as the sampling unit. These results lead to the development of a presence-absence sampling method for T. urticae by using Taylor coefficients generic for this pest. The accuracy of this method was verified against an independent data set. By determining numerical or binomial sample sizes for consistently estimating twospotted spider mite populations, growers will now be able to determine the number of predatory mites that should be released to control twospotted spider mites on impatiens.

Animals↗

The effect of psychological interventions on anxiety and depression in cancer patients: results of two meta-analyses.

The findings of two meta-analyses of trials of psychological interventions in patients with cancer are presented: the first using anxiety and the second depression, as a main outcome measure. The majority of the trials were preventative, selecting subjects on the basis of a cancer diagnosis rather than on psychological criteria. For anxiety, 25 trials were identified and six were excluded because of missing data. The remaining 19 trials (including five unpublished) had a combined effect size of 0.42 standard deviations in favour of treatment against no-treatment controls (95% confidence interval (CI) 0.08-0.74, total sample size 1023). A most robust estimate is 0.36 which is based on a subset of trials which were randomized, scored well on a rating of study quality, had a sample size > 40 and in which the effect of trials with very large effects were cancelled out. For depression, 30 trials were identified, but ten were excluded because of missing data. The remaining 20 trials (including six unpublished) had a combined effect size of 0.36 standard deviations in favour of treatment against no-treatment controls (95% CI 0.06-0.66, sample size 1101). This estimate was robust for publication bias, but not study quality, and was inflated by three trials with very large effects. A more robust estimate of mean effect is the clinically weak to negligible value of 0.19. Group therapy is at least as effective as individual. Only four trials targeted interventions at those identified as at risk of, or suffering significant psychological distress, these were associated with clinically powerful effects (trend) relative to unscreened subjects. The findings suggest that preventative psychological interventions in cancer patients may have a moderate clinical effect upon anxiety but not depression. There are indications that interventions targeted at those at risk of or suffering significant psychological distress have strong clinical effects. Evidence on the effectiveness of such targeted interventions and of the feasibility and effects of group therapy in a European context is required.

Anxiety↗

Statistics in ophthalmic research: two eyes, one eye or the mean?

BACKGROUND: Ophthalmic data, while different among individuals, are usually similar between fellow eyes of the same individual. This study was designed to illustrate alternative approaches to account for the correlation between fellow eyes. This is important for making inferences using data from both eyes. METHODS: With the use of a real data set from a population-based study, we described the distribution of intraocular pressure (IOP) by estimating the mean and standard deviation (SD) and evaluated the potential risk factors of higher IOP based on the regression method. The units of observation studied were of both eyes, right eye only, left eye only, the eyes with higher IOP and the mean value of both eyes. Furthermore, the generalized estimating equation (GEE) method was used to account for the correlation between fellow eyes in the regression analysis. Results and inferences from the different approaches were compared. RESULTS: The analysis included all the eyes, providing the largest sample size and unbiased estimates of the mean and SDs. There were some discrepancies among different approaches in the regression analysis. The GEE method simultaneously evaluated the effects of both eyes, and increased precision and enhanced inferences. CONCLUSIONS: Inconsistent results among different ophthalmic studies result from variations in not only study design and courses but also statistical methods. Making the best use of appropriate statistical techniques, which account for between eye correlation, provides valid statistical inferences.

Blood Pressure↗