Confidence intervals and sample sizes.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Many clinical trials yield data on an ordered categorical scale such as very good, good, moderate, poor. Under the assumption of proportional odds, such data can be analysed using techniques of logistic regression. In simple comparisons of two treatments this approach becomes equivalent to the Mann-Whitney test. In this paper sample size formulae consistent with an eventual logistic regression analysis are derived. The influence on efficiency of the number and breadth of categories will be examined. Effects of misclassification and of stratification are discussed, and examples of the calculations are given.
Clinical trials would gain from incorporating 'Phase 0' chronobiologic pilot designs both from the viewpoint of (statistical) power and cost-effectiveness. Herein, this statement is documented by power computations and is further illustrated by clinical examples answering specific questions. Power computations show the merits both of chronobiologic designs (that assign samples at equidistant intervals to cover one full cycle of anticipated pertinent rhythms) and of chronobiologic analyses (the cosinor versus the analysis of variance). Randomized clinical trials would gain from incorporating a concern for timing as well as dosing in all three stages of clinical trials (Phase I, II and III focusing on toxicity, efficacy and a comparison with the current best treatment, respectively) and could be cost-effectively preceded by 'Phase 0' trials so as to detect, sooner and with smaller sample sizes, desired or undesired effects that may otherwise be missed.
For a quantitative laboratory test the 0.975 fractile of the distribution of reference values is commonly used as a discrimination limit, and the sensitivity of the test is the proportion of diseased subjects with values exceeding this limit. A comparison of the estimates of sensitivity between two tests without taking into account the sampling variation of the discrimination limits can increase the type I error to about seven times the nominal value of 0.05. Correct statistical procedures are considered, and the power and required sample size are studied for Gaussian and log-Gaussian distributions of diagnostic test values. The results may be useful for the planning phase of studies to evaluate quantitative diagnostic tests.
Grizzle, Starmer, and Koch (1969, Biometrics 25, 489-503) presented a unified approach for data analysis when the outcome variable is measured on a nominal or ordinal scale. The technique uses a weighted least squares methodology, and hypotheses are tested using asymptotic chi-square statistics. In this paper, we adapt these procedures to the problem of determining the minimum sample size required for an applied research effort, and use the noncentral versions of these chi-square statistics. The results are compared against several procedures widely used in the literature, and are found to concur well with these techniques. As well, some new situations are considered.
Receiver operating characteristic (ROC) curves and their associated indices are valuable tools for the assessment of the accuracy of diagnostic tests. The area under the ROC curve is a popular summary measure of the accuracy of a test. The full area under the ROC curve, however, has been criticized because it gives equal weight to all false positive error rates. Alternative indices include the area under the ROC curve in a particular range of false positive rates ('partial' area) and the sensitivity of the test for a single fixed false positive rate (FPR). We present a unified approach for computing sample size for binormal ROC curves and their indices. Our method uses Taylor series expansions to derive approximate large-sample estimates of the variance and covariance of binormal ROC curve parameters. Several examples from diagnostic radiology illustrate the proposed method.
A direct exposure probe has been used to obtain mass spectra of underivatized guanosine, deoxyguanosine, sucrose and the p-nitrophenyl-beta-D-glucuronide. In all cases, a protonated molecular ion is produced with good relative abundance. The effects of heating rate and sample size on the production of the [MH]+ ION ARE EXAMINED IN DETAIL FROM TOTAL ION ANd single ion currents produced during rapid, repetitive scanning of the spectra after probe insertion. From this data we conclude that protonated molecular ions are produced as a result of the enhanced volatility of neutral molecules on the probe surface, followed by chemical ionization, and not by surface ionization.
When designing a clinical trial to test the equality of survival distributions for two treatment groups, the usual assumptions are exponential survival, uniform patient entry, full compliance, and censoring only administratively at the end of the trial. Various authors have presented methods for estimation of sample size or power under these assumptions, some of which allow for an R-year accrual period with T total years of study, T greater than R. The method of Lachin (1981, Controlled Clinical Trials 2, 93-113) is extended to allow for cases where patients enter the trial in a nonuniform manner over time, patients may exit from the trial due to loss to follow-up (other than administrative), other patients may continue follow-up although failing to comply with the treatment regimen, and a stratified analysis may be planned according to one or more prognostic covariates.
In bioequivalence studies Cmax and AUC serve as the primary pharmacokinetic characteristics of rate and extent of absorption. Based on pharmacokinetic relationships and on empirical evidence, the distribution of these characteristics corresponds to a multiplicative model, which implies a logarithmic normal distribution in the case of a parametric analysis. Hence, consideration is given to exact and approximate formulas of sample sizes in the case of a multiplicative model.
The number of metaphases that need to be examined to detect a Y-chromosome in a lymphocyte of a possible freemartin with a prescribed degree of certainty depends on the distribution of % XY cells among freemartins. To quantify this, the beta-binomial distribution was fit by maximum likelihood to% XY cells observed in blood samples of 70 freemartins (twin births only). The estimate mean and standard deviation of XY cells frequency in 70 freemartins (12,885 metaphases) was 0.48 and 0.30, respectively. These estimates correspond to a U-shaped distribution of XY cell frequency, i.e., one in which either an XX or an XY cell type predominates in most freemartins. These values indicate that sample sized of 26 and 168 are required to by 95% and 99% confident, respectively, that a female co-twin will not be misclassified.
The statistical difficulties of estimating cancer risks from low doses of a carcinogen are illustrated by examples from radiation carcinogenesis. Although more is known about dose-response relationships for ionizing radiation than for any other environmental carcinogen, estimates of cancer risk from low radiation doses have been extremely controversial; disagreements by factors of 100 or more are not uncommon. Direct estimation, based on data from populations exposed to low doses, is usually impracticable because of sample size requirements. Curve-fitting analyses, by which higher dose data determine lower dose risk estimates, require simple dose-response models if the estimates are to be statistically stable. The current level of knowledge about biological mechanisms of carcinogenesis dose not usually permit the confident assumption of a simple model, however; thus frequently the choice is between unstable risk estimates obtained using general models and statistically stable estimates whose stability depends on arbitrary model assumptions.
The fra(X) or Martin-Bell syndrome is the most common cause of inherited mental retardation (MR) in males. It is also associated with a variety of unusual behavioral and developmental disorders. Recent studies found great variability in the estimated strength of association between "autism" and the fra(X) syndrome, but not between MR and fra(X). We examined 31 studies which investigated the association of fra(X) syndrome with either MR or "autism" and found that the conclusion of those researchers could be significantly affected by sample size. Different behavioral and cytogenetic protocols will also influence the strength of association between fra(X) and autism.
The association of a candidate gene with disease can be efficiently evaluated by a case-control study in which allele frequencies are compared for diseased cases and unaffected controls. However, when the distribution of genotypes in the population deviates from Hardy-Weinberg proportions, the frequency of genotypes--rather than alleles--should be compared by the Armitage test for trend. We present formulas for power and sample size for studies that use Armitage's trend test. The formulas make no assumptions about Hardy-Weinberg equilibrium, but do assume random ascertainment of cases and controls, all of whom are independent of one another. We demonstrate the accuracy of the formulas by simulations.
For the two-period crossover design and a multiplicative model (logarithmic normal distribution) the decision procedure of choice is based on the inclusion of the shortest 90%-confidence interval for the ratio of expected medians for test and reference in the equivalence range. This inclusion rule is equivalent to the two one-sided tests procedure. Sample sizes based on the power of the latter have been given by Diletti et al. [1991] for an equivalence range of 0.8 to 1.25. Corresponding tables for the tighter equivalence range of 0.9 to 1.11 as well as for the wider range of 0.7 to 1.43 are given in this amendment.
For the two-period crossover design and a multiplicative model (logarithmic normal distribution) the decision procedure of choice is based on the inclusion of the shortest 90%-confidence interval for the ratio of expected medians for test and reference in the equivalence range. This inclusion rule is equivalent to the two one-sided tests procedure. Sample sizes based on the power of the latter have been given by Diletti et al. [1991] for an equivalence range of 0.8 to 1.25. Corresponding tables for the tighter equivalence range of 0.9 to 1.11 as well as for the wider range of 0.7 to 1.43 are given in this amendment.
The use function approach to group sequential methods has been explored previously. Any group sequential design requires specifying the frequency and times of repeated analyses, but only the use function approach allows deviations from those specified in the design, without affecting the type I error in the analysis. This paper illustrates how the use function provides a simple and flexible design procedure and how the initially projected maximum sample size can be calculated for randomized clinical trials in which responses are known relatively soon after patient entry and for which there is an early stopping rule built into the study protocol. Also the consequence of using the proposed design procedure is investigated in terms of the operating characteristics of the subsequent group sequential analyses.
Erythrocyte sedimentation significantly influenced test results when leucocyte retention in glass bead columns was used as a measure of leucocyte adhesiveness. In blood with elevated ESR, this influence was especially evident. Erythrocyte sedimentation highly increased leucocyte retention in vertical columns with a downward flow, whereas in slightly tilted columns with an upward flow, the retention was reduced. The error in measurements caused by erythrocyte sedimentation could be corrected by means of the haematocrit values in samples and eluates. Correction was better for results obtained with slightly elevated than with vertical columns. The granulocyte retention rates increased with increasing sample size, whereas lymphocyte retention remained constant.
Confidence intervals are important summary measures that provide useful information from clinical investigations, especially when comparing data from different populations or sites. Studies of a diagnostic test should include both point estimates and confidence intervals for the tests' sensitivity and specificity. Equally important measures of a test's efficiency are likelihood ratios at each test outcome level. We present a method for calculating likelihood ratio confidence intervals for tests that have positive or negative results, tests with non-positive/non-negative results, and tests reported on an ordinal outcome scale. In addition, we demonstrate a sample size estimation procedure for diagnostic test studies based on the desired likelihood ratio confidence interval. The renewed interest in confidence intervals in the medical literature is important, and should be extended to studies analyzing diagnostic tests.