Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

On choosing the number of interim analyses in clinical trials.

Small but important therapeutic effects of new treatments can be most efficiently detected through the study of large randomized prospective series of patients. Such large scale clinical trials are nowadays commonplace. The alternative is years of polemic and debate surrounding several trials each too small to detect plausible differences with any certainty. Such trials produce equivocal and contradictory results, which could be predicted from power calculations based upon sensible pre-trial estimates of treatment differences. Unfortunately such calculations often lead to sample sizes of several thousands. It is not surprising that investigators tend to be over-optimistic in their estimation of treatment effects (which are necessarily uncertain) especially when the sample size requirements are so stark. In this paper a method is outlined for incorporating into the sample size calculations the uncertainty of the estimate made at the design stage of a clinical trial. In particular a formal scheme is described for deciding how many interim analyses should be performed to satisfy ethical and pragmatic requirements of large clinical trial design. Although the argument will be 'Bayesian', the criteria for assessment and comparison will be strictly of a Neyman-Pearson (i.e. significance testing) kind.

Clinical Trials as Topic↗

Minimum sample sizes for identifying chromosomal fragile sites from individuals: Monte Carlo estimation.

A Monte Carlo simulation procedure was used to estimate the exact level of the standardized X2 test statistic (Xs2) for randomness in the FSM methodology for the identification of fragile sites from chromosomal breakage data for single individuals. A random-number generator was used to simulate 10,000 chromosomal breakage data sets, each corresponding to the null hypothesis of no fragile sites for numbers of chromosomal breaks (n) from 1 to 2000 and at three levels of chromosomal band resolution (k). The reliability of the test was assessed by comparisons of the empirical and nominal alpha levels for each of the corresponding values of n and k. These analyses indicate that the sparse and discrete nature of chromosomal breakage data results in large and unpredictable discrepancies between the empirical and nominal alpha levels when fragile site identifications are based on small numbers of breaks (n < 0.5 k). With n > or = 0.5 k, the distribution of Xs2 appears to be stable and non-significant differences in the empirical and nominal alpha levels are generally obtained. These results are inherent to the nature of the data and are, therefore, relevant to any statistical model for the identification of fragile sites from chromosomal breakage data. For FSM identification of fragile sites at alpha = 0.05, we suggest that n > or = 0.5 k is the minimum reliable number of mapped chromosomal breaks per individual.

Chromosome Fragile Sites↗

Bayesian methods for multiple capture-recapture surveys.

To estimate the total size of a closed population, a multiple capture-recapture sampling design can be used. This sampling design has been used traditionally to estimate the size of wildlife populations and is becoming more widely used to estimate the size of hard-to-count human populations. This paper presents Bayesian methods for obtaining point and interval estimates from data gathered from capture-recapture surveys. A numerical example involving the estimation of the size of a fish population is given to illustrate the methods.

Animals↗

Qualitative and quantitative histopathology in transitional cell carcinomas of the urinary bladder. An international investigation of intra- and interobserver reproducibility.

BACKGROUND: Histopathologic, prognosis-related grading of malignancy by means of morphologic examination in transitional cell carcinomas of the urinary bladder (TCC) may be subject to observer variation, resulting in a reduced level of reproducibility. This may confound comparisons of treatment results. Using objective, unbiased stereologic techniques and ordinary histomorphometry, such problems may be solved. EXPERIMENTAL DESIGN: A study of 110 patients with papillary or solid transitional cell carcinomas of the urinary bladder in stage Ta through T4 was carried out, addressing reproducibility of both qualitative and quantitative grading methods. Grading of malignancy was performed by one observer in Japan (using the World Health Organization scheme), and by two observers in Denmark (using the Bergkvist system). A "translation" between the systems, grade for grade, and kappa statistics were used in evaluating the reproducibility. Unbiased estimates of nuclear mean volume, nuclear mean profile area, nuclear volume fraction, nuclear profile density index, and mitotic profile density index were obtained twice in 55 of the studied cases by one observer in Japan and one in Denmark, using a random, systematic sampling scheme. RESULTS: The results were compared by bivariate correlation analyses and Kendall's tau. The international interobserver reproducibility of qualitative gradings was rather poor (kappa = 0.51), especially for grade 2 tumors (kappa = 0.28). Likewise, the interobserver agreement on the Bergkvist scheme was poor (kappa = 0.43). On the other hand was the interobserver agreement on invasion high (kappa = 0.75). The intraobserver reproducibility of the quantitative histopathologic variables was excellent in both Japan and Denmark for estimates of nuclear mean volume (r = 0.93), for nuclear mean profile area (0.78 < r < 0.83), and for nuclear profile density index (0.85 < r < 0.89), whereas the reproducibility for nuclear volume fraction was somewhat poorer (0.68 < r < 0.64). The slopes of the correlation lines were not significantly different from unity. Estimates of mitotic profile density index also showed acceptable intraobserver reproducibility (Kendall's tau > 0.53). CONCLUSIONS: The international, interobserver reproducibility of the quantitative estimators yielded similar results for all histopathologic variables investigated, except for nuclear volume fraction (r = 0.54). This can probably be related to the manual design of the sampling scheme and may be solved by introducing a motorized object stage in the systematic selection of fields of vision for quantitative measurements. However, the nuclear mean size estimators are unaffected by such sampling variability. The results obtained in this study stress the need for objective, quantitative histopathologic techniques substituting qualitative, subjective methods in prognosis-related grading of malignancy.

Adult↗

Nuclear size and shape in fine needle aspiration biopsy samples of the prostate.

OBJECTIVE: To study the potential of nuclear size and shape estimates in interpreting fine needle aspiration biopsy (FNAB) samples of the prostate. STUDY DESIGN: Morphometry was used to outline nuclei of prostate cells. Cell groups were selected by an experienced cytologist. RESULTS: The mean area of nuclei in the most atypical cell groups among definitely malignant samples (n = 17) varied from 26.3 to 93.3 micron 2 and in normal prostate cells (n = 10) from 15.6 to 33.7 micron 2. Perfect distinction of definitely benign and slightly atypical samples (n = 13) from definitely malignant samples was possible when the samples were characterized by the weighted means of the mean nuclear areas of the cell groups in the samples. The means of individual cell groups allowed correct distinction in only 84.8% of cell groups. Shape factors did not have any diagnostic value. CONCLUSION: Morphometric nuclear size estimates from ethanol-fixed FNAB samples of the prostate are of diagnostic value and can potentially be used as part of multivariate diagnostic models when selected by an experienced cytologist according to strict criteria. However, measurement should be done from several cell groups (at least three of the most-atypical cell groups) in each sample.

Adenocarcinoma↗

Sample size and power determination for clustered repeated measurements.

It is common in epidemiological and clinical studies that each subject has repeated measurements on a single common variable, while the subjects are also 'clustered'. To compute sample size or power of a test, we have to consider two types of correlation: correlation among repeated measurements within the same subject, and correlation among subjects in the same cluster. We develop, based on generalized estimating equations, procedures for computing sample size and power with clustered repeated measurements. Explicit formulae are derived for comparing two means, two slopes and two proportions, under several simple correlation structures.

Alveolar Bone Loss↗

Optimal sampling strategies for two-stage studies.

The optimal allocation of available resources is the concern of every investigator in choosing a study design. The recent development of statistical methods for the analysis of two-stage data makes these study designs attractive for their economy and efficiency. However, little work has been done on deriving two-stage designs that are optimal under the kinds of constraints encountered in practice. The methods presented in this paper provide a means of deriving designs that will maximize precision for a fixed total budget or minimize the study cost necessary to achieve a desired precision. These optimal designs depend on the relative information content and the relative cost of gathering the first- and second-stage data. In place of the usual sample size calculations, the investigator can use pilot data to estimate the study size and second-stage sampling fractions. The gains in efficiency that can result from such carefully designed studies are illustrated here by deriving and implementing optimal designs using data from the Coronary Artery Surgery Study.

Aged↗

School-level intraclass correlation for physical activity in adolescent girls.

PURPOSE: The Trial for Activity in Adolescent Girls (TAAG) is a multi-center group-randomized trial to reduce the usual decline in moderate to vigorous physical activity (MVPA) among middle-school girls. In group-randomized trials, the group-level intraclass correlation (ICC) has a strong inverse relationship to power and a good estimate of ICC is needed to determine sample size. As a result, we conducted a substudy to estimate the school-level ICC for intensity-weighted minutes of MVPA measured using an accelerometer. METHODS: To estimate the ICC, each of six sites recruited two schools and randomly selected 45 eighth grade girls from each school; 80.7% participated. Each girl wore an Actigraph accelerometer for 7 d. Readings above 1500 counts per half minute were counted as MVPA. These counts were converted into metabolic equivalents (MET) and summed over 6 a.m. to midnight to provide MET-minutes per 18-h day of MVPA. Minutes of MVPA per 18-h day also were calculated ignoring the MET value. RESULTS: The unadjusted school-level ICC for minutes of MVPA was 0.0205 (95%CI: -0.0079, 0.1727) and for MET-minutes of MVPA was 0.0045 (95% CI: -0.0147, 0.1145). Adjustment for age and BMI had no measurable effect, whereas adjustment for ethnicity reduced both ICC; adjusted values were 0.0175 (95% CI: -0.0092, 0.1622) for minutes of MVPA and 0.0000 (95% CI: -0.0166, 0.0968) for MET-minutes of MVPA. This information was used to calculate the number of schools and girls needed for TAAG to have 90% power to detect a 50% reduction in the decline of MET-minutes of MVPA between sixth and eighth grade. CONCLUSIONS: The results called for 36 schools in TAAG, with 120 girls invited for measurements at each school, and a minimum participation rate of 80%.

Acceleration↗

Determination of significant relative risks and optimal sampling procedures in prospective and retrospective comparative studies of various sizes.

Methods are given for determining the relative risks which it is possible to demonstrate as statistically significant with given probability in prospective and retrospective studies of a particular size. Also considered is the ratio of the sample sizes in the two groups being compared which will provide the most precise estimate of relative risk for a given total sample size.

Biometry↗

An E-M algorithm and testing strategy for multiple-locus haplotypes.

This paper gives an expectation maximization (EM) algorithm to obtain allele frequencies, haplotype frequencies, and gametic disequilibrium coefficients for multiple-locus systems. It permits high polymorphism and null alleles at all loci. This approach effectively deals with the primary estimation problems associated with such systems; that is, there is not a one-to-one correspondence between phenotypic and genotypic categories, and sample sizes tend to be much smaller than the number of phenotypic categories. The EM method provides maximum-likelihood estimates and therefore allows hypothesis tests using likelihood ratio statistics that have chi 2 distributions with large sample sizes. We also suggest a data resampling approach to estimate test statistic sampling distributions. The resampling approach is more computer intensive, but it is applicable to all sample sizes. A strategy to test hypotheses about aggregate groups of gametic disequilibrium coefficients is recommended. This strategy minimizes the number of necessary hypothesis tests while at the same time describing the structure of disequilibrium. These methods are applied to three unlinked dinucleotide repeat loci in Navajo Indians and to three linked HLA loci in Gila River (Pima) Indians. The likelihood functions of both data sets are shown to be maximized by the EM estimates, and the testing strategy provides a useful description of the structure of gametic disequilibrium. Following these applications, a number of simulation experiments are performed to test how well the likelihood-ratio statistic distributions are approximated by chi 2 distributions. In most circumstances the chi 2 grossly underestimated the probability of type I errors. However, at times they also overestimated the type 1 error probability. Accordingly, we recommended hypothesis tests that use the resampling method.

Algorithms↗

Mild senile dementia of the Alzheimer type. 4. Evaluation of intervention.

The design of trials of interventions intended to slow or arrest the progression of senile dementia of the Alzheimer type must be based on analysis of the natural history of the disease. Using a random coefficients statistical model, we analyzed the natural history of senile dementia of the Alzheimer type in carefully defined subjects with mild disease (n = 68) for periods of up to 10 years. Subject performance was assessed longitudinally on batteries of clinical and psychometric measures. The characteristics of these measures were analyzed relevant to their utility as outcome measures for long-term trials in patients with senile dementia of the Alzheimer type. Estimates were made of sample sizes required to show arrest, and 50% or 25% slowing in the progression of mild disease. We suggest that a clinically relevant global measure, such as the Sum of Boxes of the Clinical Dementia Rating scale, and a performance-based clinical scale or psychometric measure would be appropriate in a 12- or 24-month trial enrolling subjects with mild senile dementia of the Alzheimer type.

Actuarial Analysis↗

Superficial motor units are larger than deeper motor units in human vastus lateralis muscle.

Previous studies have suggested that regionalization may occur for human motor units, whereby smaller motor units are located in deeper parts of the muscle and larger motor units are located in more superficial portions. We examined this possibility in the human vastus lateralis muscle using macro-EMG (electromyography) to estimate motor unit size. The sample consisted of nine individuals from whom 114 motor units were recorded at forces ranging between 5% and 60% MVC. Peak-to-peak macro-EMG amplitude was well correlated with macro area (Spearman rho = 0.96). There was a statistically significant inverse relationship between recording depth and macro peak-to-peak amplitude (rho = -0.402, p < 0.001). We conclude that there is a nonrandom distribution of motor units in human muscle, with larger motor units located in more superficial regions and smaller units located in deeper regions. Clinicians who monitor motor unit activity need to recognize that a representative sample of motor unit recordings should include motor units from both deeper and more superficial regions of muscle.

Action Potentials↗

The frequency distribution of blood-flow velocities in the extraocular vessels.

PURPOSE: The study was carried out in order to assess the distribution of normal values of the blood-flow velocity in the extraocular vessels. METHODS: In 240 healthy visitors to a public fair, blood-flow characteristics in the extraocular vessels were measured, and the resistivity index was calculated. Blood-flow velocity was measured with a color Doppler imaging device, using a 7.5-mHz linear-array transducer. Peak-systolic and end-diastolic blood-flow velocities in the arteries were measured, and the resistivity index was calculated. In the central retinal vein the minimal and maximal blood-flow velocities were measured. The statistical analysis of the 14 measured and calculated variables included descriptive statistics, frequency distribution, and quantile plots. RESULTS: The quantile plots of the cumulative frequency showed that none of these 14 variables are normally distributed. Also, no normal distribution could be achieved by adjustment of the data by age. CONCLUSIONS: The blood-flow velocities in the extraocular vessels measured are not distributed normally. Therefore, nonparametric tests are to be used for statistical analysis if the sample size is small. The estimation of tolerance intervals has to be based on distribution-free assumptions.

Adolescent↗

Testing equivalence between two laboratories or two methods using paired-sample analysis and interval hypothesis testing.

A modified interval hypothesis testing procedure based on paired-sample analysis is described, as well as its application in testing equivalence between two bioanalytical laboratories or two methods. This testing procedure has the advantage of reducing the risk of wrongly concluding equivalence when in fact two laboratories or two methods are not equivalent. The advantage of using paired-sample analysis is that the test is less confounded by the intersample variability than unpaired-sample analysis when incurred biological samples with a wide range of concentrations are included in the experiments. Practical aspects including experimental design, sample size calculation and power estimation are also discussed through examples.

Laboratories↗

Analyses of nuclear ldhA gene and mtDNA control region sequences of Atlantic northern bluefin tuna populations.

There has been considerable debate about whether the Atlantic northern bluefin tuna exist as a single panmictic unit. We have addressed this issue by examining both mitochondrial DNA control region nucleotide sequences and nuclear gene ldhA allele frequencies in replicate size or year class samples of northern bluefin tuna from the Mediterranean Sea and the northwestern Atlantic Ocean. Pairwise comparisons of multiple year class samples from the 2 regions provided no evidence for population subdivision. Similarly, analyses of molecular variance of both mitochondrial and ldhA data revealed no significant differences among or between samples from the 2 regions. These results demonstrate the importance of analyzing multiple year classes and large sample sizes to obtain accurate estimates when using allele frequencies to characterize a population. It is important to note that the absence of genetic evidence for population substructure does not unilaterally constitute evidence of a single panmictic population, as genetic differentiation can be prevented by large population sizes and by migration.

Journal Article↗

Rationale and design of the Folic Acid for Vascular Outcome Reduction In Transplantation (FAVORIT) trial.

BACKGROUND: Patients with chronic kidney disease, including kidney transplant recipients, are at high risk for cardiovascular disease (CVD). In addition to the constellation of traditional CVD risk factors in chronic kidney disease, elevated total homocysteine (tHcy) is notably more prevalent among the general population. The Folic Acid for Vascular Outcome Reduction In Transplantation (FAVORIT) trial is designed to evaluate whether lowering tHcy using vitamin supplementation reduces CVD events in renal transplant recipients. METHODS: FAVORIT is a multicenter double-blind randomized controlled clinical trial. Participants are clinically stable renal transplant recipients who are 6 months or longer posttransplant with elevated tHcy. Patients are randomized to a multivitamin that includes either a high-dose or low-dose of folic acid (5 or 0 mg), vitamin B6 (50 or 1.4 mg), and vitamin B12 (1000 or 2 microg). The primary end point is a composite of incident or recurrent CVD outcomes, that is, coronary heart, cerebrovascular, or abdominal aortic/lower extremity arterial events. A sample size of 4000 is estimated to provide 87% power to detect a 20% treatment effect. Recruitment is expected to continue until July 2006, with follow-up through June 2010. RESULTS: From August 2002 through December 2004, 2234 of the target 4000 patients were enrolled. In accordance with trial design, mean (SD) screening tHcy was elevated (17.4 +/- 6.2 micromol/L), and mean (SD) estimated creatinine clearance was consistent with stable renal function (58.0 +/- 18.6 mL/min). Evaluating baseline results to date, 42% of the randomized participants had a history of diabetes mellitus, and 21% had prevalent CVD. CONCLUSIONS: The FAVORIT trial is designed with sufficient power and follow-up time to detect a clinically relevant change in CVD risk between renal transplant recipients receiving a high or low tHcy-lowering folic acid multivitamin. Preliminary screening and baseline data support the trial's objectives.

Adult↗

Analgesia and anesthesia for neonates: study design and ethical issues.

OBJECTIVE: The purpose of this article is to summarize the clinical, methodologic, and ethical considerations for researchers interested in designing future trials in neonatal analgesia and anesthesia, hopefully stimulating additional research in this field. METHODS: The MEDLINE, PubMed, EMBASE, and Cochrane register databases were searched using subject headings related to infant, newborn, neonate, analgesia, anesthesia, ethics, and study design. Cross-references and personal files were searched manually. Studies reporting original data or review articles related to these topics were assessed and critically evaluated by experts for each topical area. Data on population demographics, study characteristics, and cognitive and behavioral outcomes were abstracted and synthesized in a systematic manner and refined by group members. Data synthesis and results were reviewed by a panel of independent experts and presented to a wider audience including clinicians, scientists, regulatory personnel, and industry representatives at the Newborn Drug Development Initiative workshop. Recommendations were revised after extensive discussions at the workshop and between committee members. RESULTS: Designing clinical trials to investigate novel or currently available approaches for analgesia and anesthesia in neonates requires consideration of salient study designs and ethical issues. Conditions requiring treatment include pain/stress resulting from invasive procedures, surgical operations, inflammatory conditions, and routine neonatal intensive care. Study design considerations must define the inclusion and exclusion criteria, a rationale for stratification, the confounding effects of comorbid conditions, and other clinical factors. Significant ethical issues include the constraints of studying neonates, obtaining informed consent, making risk-benefit assessments, defining compensation or rewards for participation, safety considerations, the use of placebo controls, and the variability among institutional review boards in interpreting federal guidelines on human research. For optimal study design, investigators must formulate well-defined study questions, choose appropriate trial designs, estimate drug efficacy, calculate sample size, determine the duration of the studies, identify pharmacokinetic and pharmacodynamic parameters, and avoid drug-drug interactions. Specific outcome measures may include scoring on pain assessment scales, various biomarkers and their patterns of response, process outcomes (eg, length of stay, time to extubation), intermediate or long-term outcomes, and safety parameters. CONCLUSIONS: Much more research is needed in this field to formulate a scientifically sound, evidence-based, and clinically useful framework for management of anesthesia and analgesia in neonates. Newer study designs and additional ethical dilemmas may be defined with accumulating data in this field.

Analgesia↗

Statistics in physiology and pharmacology: a slow and erratic learning curve.

1. Learning how to apply statistical analyses to the results of experimental or clinical studies may take a lifetime of trial (and sometimes error), as it has done in the author's case. There is no evidence that biomedical investigators of the present generation are on a steeper learning curve. Gross misunderstandings of the purpose and functions of statistical analysis are apparent in applications to research grant-giving bodies and ethics committees, in manuscripts submitted to journals and sometimes in published papers. 2. Although estimation of minimal group (sample) size for a given power is an essential step in planning clinical studies, it seems to be used rarely in laboratory experimental work. This is despite exhortations to restrict the number of animals used to a minimum. 3. Most investigators use hypothesis testing to analyse their results, but their understanding of the meaning of the resultant P-values is slight. 4. A flaw found almost universally in biomedical manuscripts is to make multiple inferences from the results of a single study. The goal of statistical analysis is to maintain the familywise type I error rate (risk of false-positive inference) at a predetermined level (usually 5%). But, when multiple inferences are made from the same experiment, the risk of false-positive error is inflated. There are two solutions to this problem: (i) use a multiple comparison procedure to control the familywise type I error rate; and (ii) test a single, global hypothesis. 5. Biomedical investigators have been quick to acquire computer statistics software and to use it to analyse their experiments. However, they have been slow to recognize the limitations of this software. These include: (i) inadequate documentation of routines, so that neither the user nor the reader of published papers can be sure how the tests have been executed; (ii) flawed algorithms for the execution of statistical procedures; and (iii) failure to recognize that the best software for their purposes is that which takes them just beyond their statistical horizons. 6. The obvious solution to these difficulties is to recruit a biomedical statistician into every research group, at a relatively trivial cost. However, properly qualified biostatisticians are in desperately short supply in Australia. It follows that research groups, national grant-giving agencies and academic institutions must make provision for the proper training and subsequent employment of biostatisticians.

Data Interpretation, Statistical↗