Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Sample Size”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Feature selection and classifier performance in computer-aided diagnosis: the effect of finite sample size.

In computer-aided diagnosis (CAD), a frequently used approach for distinguishing normal and abnormal cases is first to extract potentially useful features for the classification task. Effective features are then selected from this entire pool of available features. Finally, a classifier is designed using the selected features. In this study, we investigated the effect of finite sample size on classification accuracy when classifier design involves stepwise feature selection in linear discriminant analysis, which is the most commonly used feature selection algorithm for linear classifiers. The feature selection and the classifier coefficient estimation steps were considered to be cascading stages in the classifier design process. We compared the performance of the classifier when feature selection was performed on the design samples alone and on the entire set of available samples, which consisted of design and test samples. The area Az under the receiver operating characteristic curve was used as our performance measure. After linear classifier coefficient estimation using the design samples, we studied the hold-out and resubstitution performance estimates. The two classes were assumed to have multidimensional Gaussian distributions, with a large number of features available for feature selection. We investigated the dependence of feature selection performance on the covariance matrices and means for the two classes, and examined the effects of sample size, number of available features, and parameters of stepwise feature selection on classifier bias. Our results indicated that the resubstitution estimate was always optimistically biased, except in cases where the parameters of stepwise feature selection were chosen such that too few features were selected by the stepwise procedure. When feature selection was performed using only the design samples, the hold-out estimate was always pessimistically biased. When feature selection was performed using the entire finite sample space, the hold-out estimates could be pessimistically or optimistically biased, depending on the number of features available for selection, the number of available samples, and their statistical distribution. For our simulation conditions, these estimates were always pessimistically (conservatively) biased if the ratio of the total number of available samples per class to the number of available features was greater than five.

Algorithms↗

Subgroup analyses in randomized trials: risks of subgroup-specific analyses; power and sample size for the interaction test.

OBJECTIVE: Despite guidelines recommending the use of formal tests of interaction in subgroup analyses in clinical trials, inappropriate subgroup-specific analyses continue. Moreover, trials designed to detect overall treatment effects have limited power to detect treatment-subgroup interactions. This article quantifies the error rates associated with subgroup analyses. STUDY DESIGN AND SETTING: Simulations quantified the risks of misinterpreting subgroup analyses as evidence of differential subgroup effects and the limited power of the interaction test in trials designed to detect overall treatment effects. RESULTS: Although formal interaction tests performed as expected with respect to false positives, subgroup-specific tests were considerably less reliable: A significant effect in one subgroup only was observed in 7% to 64% of simulations depending on trial characteristics. Regarding power of the interaction test, a trial with 80% power for the overall effect had only 29% power to detect an interaction effect of the same magnitude. For interactions of this size to be detected with the same power as the overall effect, sample sizes should be inflated fourfold, increasing dramatically for interactions smaller than 20% of the overall effect. CONCLUSION: Although it is generally recognized that subgroup analyses can produce spurious results, the extent of the problem may be underestimated.

Data Interpretation, Statistical↗

A sample-size-optimal Bayesian procedure for sequential pharmaceutical trials.

Consider a pharmaceutical trial where the consequences of different decisions are expressed on a financial scale. The efficacy of the new drug under consideration has a prior distribution obtained from the underlying biological process, animal experiments, clinical experience, and so forth. Berry and Ho (Biometrics 44, 219-227) show how these components are used to establish an optimal (Bayes) sequential testing procedure, assuming a known constant sample size at each decision point. We show in this article how it is also possible to optimize further, with respect to the sample-size rule. This component of the design, which is missing from most sequential procedures, has the potential to yield considerably larger expected net gains (equivalently, considerably smaller Bayes risks).

Animals↗

Power and sample size in cost-effectiveness analysis.

For resource allocation under a constrained budget, optimal decision rules for mutually exclusive programs require that the treatment with the highest incremental cost-effectiveness ratio (ICER) below a willingness-to-pay (WTP) criterion be funded. This is equivalent to determining the treatment with the smallest net health cost. The designer of a cost-effectiveness study needs to select a sample size so that the power to reject the null hypothesis, the equality of the net health costs of two treatments, is high. A recently published formula derived under normal distribution theory overstates sample-size requirements. Using net health costs, the authors present simple methods for power analysis based on conventional normal and on nonparametric statistical theory.

Bias↗

The effects of sensitivity and specificity of case selection on validity, sample size, precision, and power in hospital-based case-control studies.

The consequences of imperfect sensitivity and specificity in disease diagnosis in epidemiologic studies have conventionally been assessed by models of misclassification which assume a fixed number of study participants. This assumption is not usually applicable to case-control studies in which disease diagnosis is part of the case selection process and sensitivity and specificity will, for a given time period and source of cases, affect the size of the case group. In this paper, a mathematical model that incorporates this is developed in the framework of a hospital-based case-control study. The separate and combined effects of imperfect sensitivity and specificity of case diagnosis on validity, sample size, precision, and power are assessed. The authors conclude that if several diagnostic procedures are available, specificity of case diagnosis should usually take precedence over sensitivity for the sake of validity. Although increasing specificity and sacrificing sensitivity may compromise precision to some extent, the latter can often be fully compensated for by an increased control:case ratio. Imperfect specificity also compromises power despite increased sample size. Since clinical diagnoses tend to focus on high sensitivity and sacrifice some specificity, their uncritical adoption for case recruitment in case-control studies may compromise their validity.

Case-Control Studies↗

Determining equivalence and the impact of sample size in anti-infective studies: a point to consider.

The problem of establishing the equivalence of an experimental treatment to a control with respect to a binary ("success" or "failure") response variable may be solved using an approximate (1-alpha) 100% confidence interval for the difference in the response rates (i.e., success probabilities). If the goal is to show that the experimental treatment is not sufficiently worse than the control, then a decision rule based on the magnitude of one confidence limit can be used. A procedure suggested by the Food and Drug Administration allows the value to which the confidence limit is to be compared to depend on the data. The consequences of determining the sample size assuming that the aforementioned value is fixed are examined. The probability of declaring equivalence and exact sample sizes are also presented for the procedure.

Anti-Infective Agents↗

The sample size for a clinical trial: a Bayesian-decision theoretic approach.

Using decision theory, what is an appropriate sample size for a clinical trial, with a binary endpoint? We present a program, suitable for actual planning, which, with some extensions, implements Canner's solution to this question. Examples with a discussion are given. Implications of a Bayesian approach are discussed. Bayesian and Neyman--Pearson approaches are compared.

Bayes Theorem↗

Symmetry of the canine femur: implications for experimental sample size requirements.

The purpose of this study was to determine the degree of bilateral variability in cross-sectional geometric properties of the adult canine proximal femur and to use these data to determine minimum detectable treatment effects in paired and independent experimental designs. Thirteen pairs of canine femora were sectioned at nine locations and 16 cross-sectional geometric properties were determined for each section location. The canine femur was found to be bilaterally symmetrical. For a given sample size, the magnitude of the detectable treatment effect was (a) smaller for diameters than for areas and area moments of inertia and (b) smaller within the middiaphysis than proximally. The data from this study can be used to estimate sample size requirements for experiments in which the treatment effect is determined by using the contralateral femur as a control. It was found that an increase from 3 to 7 animals would have a much greater effect on improving the sensitivity of an experiment than would an increase from 7 to 11 animals.

Animals↗

A Bayesian approach to establishing sample size and monitoring criteria for phase II clinical trials.

Thall and Simon propose a Bayesian approach to phase II clinical trials with binary outcomes and continuous monitoring. The efficacy theta E of an experimental treatment E is evaluated relative to that of a standard treatment S based on data from an uncontrolled trial of E, an informative prior for theta S, and a noninformative prior for theta E. The trial continues until E is shown with high posterior probability to be either promising or not promising, or until a predetermined maximum sample size is reached. Operating characteristics are evaluated under fixed values of the success probability of E. In this paper, we propose two extensions of this decision structure, describe sample size and monitoring criteria, and provide numerical guidelines for implementation. The first extension gives criteria from early termination of trials unlikely to yield conclusive results, based on the marginal (predictive) distribution of the observed success rate. The second extension allows early termination only if E is found to be not promising compared to S. Operating characteristics of each of these designs are evaluated numerically over a range of design parameterizations. We also examine the effects of intermittent monitoring on the design's properties. An application of this approach to a leukemia biochemotherapy trial is described.

Antineoplastic Agents↗

The likelihood ratio test for the two-component normal mixture problem: power and sample size analysis.

We find, through simulation and modeling, an approximation to the alternative distribution of the likelihood ratio test for two-component mixtures in which the components have different means but equal variances. We consider the range of mixing proportions from 0.5 through .95. Our simulation results indicate a dependence of power on the mixing proportion when pi less than .2 and pi greater than .80. Our model results indicated that the alternative distribution is approximately noncentral chi-square, possibly with 2 degrees of freedom. Using this model, we estimate a sample of 40 is needed to have 50% power to detect a difference between means equal to 3.0 for mixing proportions between .2 and .8. The sample size increases to 50 when the mixing proportion is .90 (or .1) and 82 when the mixing proportion is .95 (or .05). This paper contains a complete table of sample sizes needed for 50%, 80%, and 90% power.

Alleles↗

The randomized concentration-controlled trial: an evaluation of its sample size efficiency.

A randomized concentration-controlled trial (RCCT) is one in which subjects are randomly assigned to predetermined levels of average plasma drug concentration. These target concentrations can be achieved (within reasonable ranges) by an individualized pharmacokinetically controlled dosing scheme. The RCCT is designed to minimize the interindividual pharmacokinetic (PK) variability within comparison groups and consequently decrease the variability in clinical response within these groups. In this paper, we investigate the extent of improvement in sample size efficiency that can be gained from the RCCT design in comparison to the traditional randomized dose-controlled trial (RDCT) design. Our investigations involve both theoretical arguments and simulation studies, illustrated with data on PK and pharmacodynamic (PD) characteristics of the antiasthma drug theophylline. Aside from safety concerns that strongly suggest the use of RCCT for drugs with narrow therapeutic windows, sample size considerations favor the choice of RCCT in many situations, as shown in this paper.

Administration, Oral↗

The impact of dietary measurement error on planning sample size required in a cohort study.

Dietary measurement error has two consequences relevant to epidemiologic studies: first, a proportion of subjects are misclassified into the wrong groups, and second, the distribution of reported intakes is wider than the distribution of true intakes. While the first effect has been dealt with by several other authors, the second effect has not received as much attention. Using a simple errors-in-measurement model, the authors investigate the implications of measurement error for the distribution of fat intake. They then show how the inference of a more narrow distribution of true intakes affects the calculation of sample size for a cohort study. The authors give an example of the calculation for a cohort study investigating dietary fat and colorectal cancer. This shows that measurement error has a profound effect on sample size, requiring a six- to eightfold increase over the number required in the absence of error, if the correlation coefficient between reported and true intakes is 0.65. Reliable detection of a relative risk of 1.36 between a true intake of greater than 47.5% calories from fat and less than 25% calories from fat would require approximately one million subjects.

Bias↗

Sample size and type-token ratios for oral language of preschool children.

This study investigated the stability of five type-token ratios (TTRs) in 50-utterance oral language samples segmented into nine lengths. The samples were obtained from 83 children, 3, 4, and 5 years of age. The five TTRs included the basic type-token ratio, the corrected type-token ratio, the root type-token ratio, the bilogarithmic type-token ratio, and the Characteristic K. The sample segment sizes consisted of the first, second, third, and fourth 50-word segments; the first and second 100-word segments; the first 150-word segment; the first 200-word segment; and the total 50-utterance sample. Based on the results of this study, TTR measures on the language of young children should not be compared for samples that differ in number of words. Further, TTR measures for sample size of 50 and 100 words have reliabilities that are judged to be inadequate for research or clinical purposes. The size of the language sample needed for minimum reliability of .70 is 350 words. Greater reliability requires larger word-sample size.

Child Language↗

Modification of the computational procedure in Parker and Bregman's method of calculating sample size from matched case-control studies with a dichotomous exposure.

Several authors have addressed the problem of calculating sample size for a matched case-control study with a dichotomous exposure. The approach of Parker and Bregman (1986, Biometrics42, 919-926) is, in our view, one of the most satisfactory, since it requires specification of quantities that are often easily available to the investigator. However, its recommended implementation involves a computational approximation. We show here that the approximation performs poorly in extreme situations and can be easily replaced with a more exact calculation.

Case-Control Studies↗

A note on sample size determination for bioequivalence studies with high-order crossover designs.

Similar to Liu and Chow, approximate formulas for sample size determination are derived based on Schuirmann's two one-sided tests procedure for bioequivalence studies for the additive and the multiplicative models under various higher order crossover designs for comparing two formulations of a drug product. The higher order crossover designs under study include Balaam's design, the two-sequence dual design, and two four-period designs (with two and four sequences), which are commonly used for assessment of bioequivalence between formulations. The derived formulas are simple enough to be carried out with a pocket calculator. The number of subjects required for each of the four higher order designs are tabulated for selected powers and various parameter values.

Clinical Trials as Topic↗

Comparison of tests and sample size formulae for proving therapeutic equivalence based on the difference of binomial probabilities.

To prove the hypothesis that a new treatment is as effective as a standard one, a possibility is to test the one-sided null hypothesis of a clear inferiority of the new treatment against the alternative hypothesis that, if at all, it is only negligibly inferior. Such a problem is of clinical relevance if, for instance, a new treatment with an effectiveness which is comparable to that of the standard one would be preferred if it is less toxic. If the difference between the two treatments is measured by the difference of failure rates, approximate statistical tests and sample size formulae have to be used. This paper reports the results of an extensive empirical investigation comparing the well known calculations proposed by Blackwelder, Rodary, ComNougue and Tournade, which propose a sample size formula for the test of Dunnett and Gent, and Farrington and Manning. The investigation was conducted in order to allow more comprehensive conclusions than those which may be drawn from the limited examples given by the last authors. For the usual settings in clinical trials, the formulae of Farrington and Manning are recommended. However, there are combinations of statistical parameters for which they are not preferred.

Binomial Distribution↗

A Bayesian decision approach for sample size determination in phase II trials.

Stallard (1998, Biometrics 54, 279-294) recently used Bayesian decision theory for sample-size determination in phase II trials. His design maximizes the expected financial gains in the development of a new treatment. However, it results in a very high probability (0.65) of recommending an ineffective treatment for phase III testing. On the other hand, the expected gain using his design is more than 10 times that of a design that tightly controls the false positive error (Thall and Simon, 1994, Biometrics 50, 337-349). Stallard's design maximizes the expected gain per phase II trial, but it does not maximize the rate of gain or total gain for a fixed length of time because the rate of gain depends on the proportion of treatments forwarding to the phase III study. We suggest maximizing the rate of gain, and the resulting optimal one-stage design becomes twice as efficient as Stallard's one-stage design. Furthermore, the new design has a probability of only 0.12 of passing an ineffective treatment to phase III study.

Bayes Theorem↗

What sample sizes are required for pooling surgical case durations among facilities to decrease the incidence of procedures with little historical data?

BACKGROUND: Better predictions of each case's duration would reduce operating room labor costs and patient waiting times. A barrier to using historical case duration data to predict the duration of future cases is the absence for some cases of previous data for the same scheduled procedure from the same facility. The authors examined sample size requirements for pooling case duration data from several facilities to create a 90% chance of having case duration data for almost all procedures. METHODS: Four academic medical centers provided data, totaling 200,401 cases classified by the scheduled Current Procedural Terminology codes. RESULTS: The 12% of cases in which procedures occurred once or twice accounted for 79% of procedures or combinations of procedures. When a procedure was being performed for the first time at a facility, that same procedure had been performed previously at least once at one or more of the other three facilities only 13-25% of the time. More than 1 million cases would be needed to have a 90% chance of having at least 3 cases for each procedure observed in the original 200,401 cases. However, with N = 200,401 cases in our initial data set, we observed less than one third of the estimated total number of possible procedures. CONCLUSIONS: The lack of historical case duration data for scheduled procedures is an important cause of inaccuracy in predicting case durations. However, millions of cases probably would be required to provide historical case duration data for almost all procedures.

Academic Medical Centers↗