Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Prognostic factors in node-negative breast cancer: a review of studies with sample size more than 200 and follow-up more than 5 years.

OBJECTIVE: To review the published literature on prognostic factors in patients with node-negative breast cancer, focusing principally on recent studies with large sample sizes and extended follow-up periods. SUMMARY BACKGROUND DATA: Although numerous studies have examined prognostic factors in patients with breast cancer, relatively few have dealt specifically with node-negative disease, and interpretation has been limited by small sample size and limited follow-up times. METHODS: A review of the Medline database from 1996 to 2000 was undertaken, with additional papers published before 1996 identified through review articles. For inclusion in the analysis, papers needed to meet the following core criteria: 200 or more node-negative patients with invasive breast carcinoma; median follow-up time at least 5 years; method of testing and cut-off points specified; overall survival and/or disease-free survival specified; and relative risk or statistical probability values given for comparisons. RESULTS: Three or more papers that met the core criteria were retrieved for each of 11 potential prognostic factors. Of these, tumor size, tumor grade, cathepsin-D, Ki-67, S-phase fraction, mitotic index, and vascular invasion showed a significant association with survival outcomes; HER2/neu and DNA ploidy showed no significant association; and estrogen receptor status and p53 showed mixed results. Lack of standardization in measurement techniques for many of the markers, including cathepsin-D, Ki-67, HER2/neu, and p53, limited their current clinical usefulness. CONCLUSIONS: In large studies with extended follow-up periods, tumor size, tumor grade, cathepsin-D, Ki-67, S-phase fraction, mitotic index, and vascular invasion showed a significant association with survival outcome measures in patients with early-stage node-negative breast cancer. Because of technical difficulties and variations in the measurement of many of these factors, tumor size and tumor grade remain the only markers that currently have broad clinical usefulness for this patient group.

Analysis of Variance↗

Power and sample size calculations for clinical trials of myofascial pain of jaw muscles.

When a clinical trial is planned, the approximate number of subjects needed for significant differences between/among groups to be detected must be estimated. Sample-size calculations provide the investigator with this information. This paper discusses the choice of outcome measures and describes the steps used to estimate the numbers of subjects necessary for a study that compares treatments for patients with chronic myofascial pain of jaw muscles. Within- and between-subject variances were estimated for the chosen variables, the subjects' pain ratings on visual analogue scales. Sample sizes were then calculated for theoretical differences between groups by pre-treatment means and overall standard deviations (Cohen, 1977). The results of this analysis can be used by other researchers when planning studies involving these types of patients.

Adolescent↗

Are sample sizes usually at least an order of magnitude too low for reliable estimates of leaf asymmetry?

Estimates of leaf size and asymmetry for individual trees are often obtained using sample sizes that are too small to take into account the possibility that size and asymmetry may be affected by the position of the leaf on the tree. This issue was addressed by exploring variation in leaf size and asymmetry within an individual of Alder (Alnus glutinosa). We found differences between branches for leaf size and for signed asymmetry but not for unsigned asymmetry. We also found that the size of a leaf was not correlated with its position on a branch and that the asymmetry of a leaf was not correlated with either its position on a branch or with the asymmetry of its neighbour. Repeated subsampling of a sample of 870 leaves showed that a subsample size approaching 500 leaves was required for consistently reliable estimates of the standard deviation of unsigned asymmetry. Smaller subsamples were required for consistently reliable estimates of mean unsigned asymmetry and of the mean and standard deviation of leaf size, but subsamples of less than 100 leaves provided consistently reliable estimates only of mean leaf size. For this species, reliable estimates of an individual's level of asymmetry are obtained only if several hundred leaves are sampled over several branches, but it is not necessary to sample the same sequence of leaves from each branch.

Analysis of Variance↗

Sample size and statistical power assessing the effect of interventions in the context of mixture distributions with detection limits.

Often in randomized clinical trials and observational cohort studies, a non-negative continuously distributed response variable is measured in treatment and control groups. In the presence of true zeros for the response variable, a two-part zero-inflated log-normal model (which assumes that the data has a probability mass at zero and a continuous response for values greater than zero) is usually recommended. However, in some environmental health and human immunodeficiency virus (HIV) studies, quantitative assays for metabolites of toxicants, or quantitative HIV RNA measurements are subject to left-censoring due to values falling below the limit of detection (LD). Here, a zero-inflated log-normal mixture model is often suggested since true zeros are indistinguishable from left-censored values due to the LD. When the probabilities of true zeros in the two groups are not restricted to be equal, the information contributed by values falling below LD is used only to estimate the probability of true zeros in the context of mixture distributions. We derived the required sample size to assess the effect of a treatment in the context of mixture models with equal and unequal variances based on the left-truncated log-normal distribution. Methods for calculation of statistical power are also presented. We calculate the required sample size and power for a recent study estimating the effect of oltipraz on reducing urinary levels of the hydroxylated metabolite aflatoxin M(1) (AFM(1)) in a randomized, placebo-controlled, double-blind phase IIa chemoprevention trial in Qidong, China. A Monte Carlo simulation study is conducted to investigate the performance of the proposed methods.

Aflatoxin M1↗

Impact of sample size on the performance of multiple-model pharmacokinetic simulations.

Monte Carlo simulations are increasingly used to predict pharmacokinetic variability of antimicrobials in a population. We investigated the sample size necessary to provide robust pharmacokinetic predictions. To obtain reasonably robust predictions, a nonparametric model derived from a sample population size of >/=50 appears to be necessary as the input information.

Anti-Bacterial Agents↗

Sample size considerations in genetic polymorphism studies.

OBJECTIVES: Molecular studies for genetic polymorphisms are being carried out for a number of different applications, such as genetic disorders in different populations, pharmacogenomics, genetic identification of ethnic groups for forensic and legal applications, genetic identification of breed/stock in animals and plants for commercial applications and conservation of germ plasm. In this paper, for a random sampling scheme, we address two questions: (A) What should be the minimum size of the sample so that, with a prespecified probability, all alleles at a given locus (or haplotypes at a given set of loci) are detected? (B) What should be the sample size so that the allele frequency distribution at a given locus (or haplotype frequency distribution at a given set of loci) is estimated reliably within permissible error limits? METHODS: We have used combinatorial probabilistic arguments and Monte Carlo simulations to answer these questions. RESULTS: We found that the minimum sample size required in case A depends mainly on the prespecified probability of detecting all alleles, while in case B, it varies greatly depending on the permissible error in estimation (which will vary with the application). We have obtained the minimum sample sizes for different degrees of polymorphism at a locus under high stringency, as well as a relaxed level of permissible error. We present a detailed sampling procedure for estimating allele frequencies at a given locus, which will be of use in practical applications. CONCLUSION: Since the sample size required for reliable estimation of allele frequency distribution increases with the number of alleles at the locus, there is a strong case for using biallelic markers (like single nucleotide polymorphisms) when the available sample size is about 800 or less.

Gene Frequency↗

Sample size calculation for planning group sequential longitudinal trials.

Procedures are developed in this paper for sample size calculations for planning a group sequential longitudinal trial with various correlation structures, using a test statistic based on generalized estimating equations (Lee, Kim and Tsiatis) and group sequential boundaries based on type I error spending functions (Lan and DeMets).

Adolescent↗

Test statistics and sample size formulae for comparative binomial trials with null hypothesis of non-zero risk difference or non-unity relative risk.

When it is required to establish a materially significant difference between two treatments, or, alternatively, to show that two treatments are equivalent, standard test statistics and sample size formulae based on a null hypothesis of no difference no longer apply. This paper reviews some of the test statistics and sample size formulae proposed for comparative binomial trials when the null hypothesis is of a specified non-zero difference or non-unity relative risk. Methods based on restricted maximum likelihood estimation are recommended and applied to studies of pertussis vaccine.

Analysis of Variance↗

Sample size calculations for cluster randomised trials. Changing Professional Practice in Europe Group (EU BIOMED II Concerted Action).

OBJECTIVES: Cluster randomised trials, in which groups of individuals are randomised, are increasingly being used in the health field. Adopting a clustered approach has implications for the design of such trials, and sample size calculations need to be inflated to accommodate for the clustering effect. Reliable estimates of intracluster correlation coefficients (ICCs) are required for robust sample size calculations to be made; however, little empirical evidence is available on their likely size, and on factors which influence their magnitude. The aim of this study was to generate empirical estimates of ICCs and to explore factors which may affect their magnitude. METHODS: Empirical estimates of ICCs were calculated for both process variables and patient outcomes from a number of datasets of primary and secondary care implementation studies. RESULTS: Estimates of ICCs varied according to setting and type of outcome. Estimates of ICCs for process variables were higher than those for patient outcomes, and estimates derived from secondary care were higher than those from primary care. ICCs for process variables in primary care were of the order of 0.05-0.15, whilst those in secondary care were of the order of 0.3. Estimates for patient outcomes in primary care were generally lower than 0.05. CONCLUSIONS: Adopting cluster randomisation has implications for the design, size and analysis of clinical trials. This study gives an insight into the potential size of ICCs in primary and secondary care, and provides a practical guide to researchers to aid the planning of future studies in this area.

Cluster Analysis↗

Model selection in covariance structures analysis and the "problem" of sample size: a clarification.

Complex models for covariance matrices are structures that specify many parameters, whereas simple models require only a few. When a set of models of differing complexity is evaluated by means of some goodness of fit indices, structures with many parameters are more likely to be selected when the number of observations is large, regardless of other utility considerations. This is known as the sample size problem in model selection decisions. This article argues that this influence of sample size is not necessarily undesirable. The rationale behind this point of view is described in terms of the relationships among the population covariance matrix and 2 model-based estimates of it. The implications of these relationships for practical use are discussed.

Analysis of Variance↗

On sample sizes to estimate the protective efficacy of a vaccine.

To estimate vaccine protective efficacy, defined as VE = 1 - ARV/ARU where ARV is the disease attack rate in the vaccinated group and ARU is the disease attack rate in the controls, investigators have used both cohort and case-control designs. For each design, we present a method for calculation of the sample size required to provide an approximate confidence interval for VE of predetermined width and probability of coverage. The required sample size is a function of the desired width of the confidence interval, the probability of coverage, the assumed VE, and, for cohort designs, the assumed disease attack rate in the controls, and for case-control designs, the assumed vaccine exposure prevalence for the controls.

Child, Preschool↗

Expanding an existing multiple choice test with a mixed format test: simulation study on sample size and item recovery in concurrent calibration.

When a new set of mixed format items is augmented with a previous old multiple-choice (MC) test, those mixed format items should be linked to the existing old MC test. This study used simulation to investigate sample size effect on recovery of known item parameter from the concurrent calibration in the context of horizontal equating, where the new mixed format tests are equated to the existing MC test which acts as the common linking items. In the partial credit model following the Andrich style parameterization, item location and item step parameters were differentially affected by the sample size. Item location parameters were recovered better than item step parameters at the individual item, the sub-test, and the total test level. This study also shows the outward bias for the item location parameter estimated by the maximum likelihood estimator.

Educational Measurement↗

Trend tests for case-control studies of genetic markers: power, sample size and robustness.

The Cochran-Armitage trend test is commonly used as a genotype-based test for candidate gene association. Corresponding to each underlying genetic model there is a particular set of scores assigned to the genotypes that maximizes its power. When the variance of the test statistic is known, the formulas for approximate power and associated sample size are readily obtained. In practice, however, the variance of the test statistic needs to be estimated. We present formulas for the required sample size to achieve a prespecified power that account for the need to estimate the variance of the test statistic. When the underlying genetic model is unknown one can incur a substantial loss of power when a test suitable for one mode of inheritance is used where another mode is the true one. Thus, tests having good power properties relative to the optimal tests for each model are useful. These tests are called efficiency robust and we study two of them: the maximin efficiency robust test is a linear combination of the standardized optimal tests that has high efficiency and the MAX test, the maximum of the standardized optimal tests. Simulation results of the robustness of these two tests indicate that the more computationally involved MAX test is preferable.

Case-Control Studies↗

Quantification of variability and uncertainty using mixture distributions: evaluation of sample size, mixing weights, and separation between components.

Variability is the heterogeneity of values within a population. Uncertainty refers to lack of knowledge regarding the true value of a quantity. Mixture distributions have the potential to improve the goodness of fit to data sets not adequately described by a single parametric distribution. Uncertainty due to random sampling error in statistics of interests can be estimated based upon bootstrap simulation. In order to evaluate the robustness of using mixture distribution as a basis for estimating both variability and uncertainty, 108 synthetic data sets generated from selected population mixture log-normal distributions were investigated, and properties of variability and uncertainty estimates were evaluated with respect to variation in sample size, mixing weight, and separation between components of mixtures. Furthermore, mixture distributions were compared with single-component distributions. Findings include: (1). mixing weight influences the stability of variability and uncertainty estimates; (2). bootstrap simulation results tend to be more stable for larger sample sizes; (3). when two components are well separated, the stability of bootstrap simulation is improved; however, a larger degree of uncertainty arises regarding the percentiles coinciding with the separated region; (4). when two components are not well separated, a single distribution may often be a better choice because it has fewer parameters and better numerical stability; and (5). dependencies exist in sampling distributions of parameters of mixtures and are influenced by the amount of separation between the components. An emission factor case study based upon NO(x) emissions from coal-fired tangential boilers is used to illustrate the application of the approach.

Journal Article↗

Effect of changing the bioequivalence range from (0.80, 1.20) to (0.80, 1.25) on the power and sample size.

International harmonization of guidelines for bioequivalence assessment has led to a wide acceptance of the multiplicative model for the extent and rate characteristics AUC and Cmax and--in consistency with this--of the bioequivalence range (0.80, 1.25). The effect of this change from (0.80, 1.20) on the power of the two one-sided test procedure and the sample sizes based thereon is investigated as a function of the within-subject coefficient of variation (CV) and the ratio mu T/mu R of expected medians for test and reference. The relative reduction in sample size is practically zero for mu T/mu R < or = 0.9 and then gradually increases as mu T/mu R approaches 1.2. At mu T/mu R = 1, the reduction is up to 20%. For a fixed ratio mu T/mu R this reduction increases with the coefficient of variation, reaching a plateau at a CV of about 25%.

Models, Statistical↗

Sample size in qualitative research.

A common misconception about sampling in qualitative research is that numbers are unimportant in ensuring the adequacy of a sampling strategy. Yet, simple sizes may be too small to support claims of having achieved either informational redundancy or theoretical saturation, or too large to permit the deep, case-oriented analysis that is the raison-d'être of qualitative inquiry. Determining adequate sample size in qualitative research is ultimately a matter of judgment and experience in evaluating the quality of the information collected against the uses to which it will be put, the particular research method and purposeful sampling strategy employed, and the research product intended.

Humans↗

Mounting a community-randomized trial: sample size, matching, selection, and randomization issues in PRISM.

This paper discusses some of the processes for establishing a large cluster-randomized trial of a community and primary care intervention in 16 local government areas in Victoria, Australia. The development of the trial in terms of design factors such as sample size estimates and the selection and randomization of communities to intervention or comparison is described. The intervention program to be implemented in Program of Resources, Information and Support for Mothers (PRISM) was conceived as a whole community approach to improving support for all mothers in the first 12 months after birth. A cluster-randomized trial was thus the design of choice from the outset. With a limited number of communities available, a matched-pair design with eight pairs was chosen. Sample size estimates, adjusting for the cluster randomization and the pair-matched design, showed that with eight pairs, on average, 800 women from each community would need to respond to provide sufficient power to determine a 3% reduction in the prevalence of maternal depression 6 months after birth-a reduction deemed to be a worthwhile impact of the intervention to be reliably detected at 80% power. The process of selecting suitable communities and matching them into pairs required careful collection of data on numbers of births, size of the local government areas (LGAs), and an assessment of the capacity of communities to implement the intervention. Ways of dealing with boundary issues associated with potential contamination are discussed. Methods for the selection of feasible configurations of sets of pairs and the ultimate allocation to intervention or comparison are provided in detail. Ultimately, all such studies are a balancing act between selecting the minimum number of communities to detect a meaningful outcome effect of an intervention and the maximum size budget and other resources allow.

Adult↗

On sample-size and power calculations for studies using confidence intervals.

A recent trend in epidemiologic analysis has been away from significance tests and toward confidence intervals. In accord with this trend, several authors have proposed the use of expected confidence intervals in the design of epidemiologic studies. This paper discusses how expected confidence intervals, if not properly centered, can be misleading indicators of the discriminatory power of a study. To rectify such problems, the study must be designed so that the confidence interval has a high probability of not containing at least one plausible but incorrect parameter value. To achieve this end, conventional formulas for power and sample size may be used. Expected intervals, if properly centered, can be used to design uniformly powerful studies but will yield sample-size requirements far in excess of previously proposed methods.

Epidemiologic Methods↗