Search PubMed⌕ Search

Biomedical subjects

H J Keselman

Publications and source records attributed to H J Keselman.

At least 19 recordsLinked to original sources

Robust tests for the multivariate Behrens-Fisher problem.

Hotelling's T2 procedure is used to test the equality of means in two-group multivariate designs when covariances are homogeneous. A number of alternatives to T2, which are robust to covariance heterogeneity, have been proposed in the literature. However, all are sensitive to departures from multivariate normality. We demonstrate how to obtain multivariate tests that are robust to covariance heterogeneity and non-normality with estimators of location and scale based on trimming and Winsorizing. The performance of six alternatives to T2 was examined via Monte Carlo methods when characteristics of the research design, degree of covariance heterogeneity, and degree of non-normality were manipulated. We have recently developed a program written in the SAS/IML language that can be used to implement these robust multivariate tests. Recommendations are provided on the specific data-analytic conditions under which these tests should be adopted.

Analysis of Variance↗

An alternative to Cohen's standardized mean difference effect size: a robust parameter and confidence interval in the two independent groups case.

The authors argue that a robust version of Cohen's effect size constructed by replacing population means with 20% trimmed means and the population standard deviation with the square root of a 20% Winsorized variance is a better measure of population separation than is Cohen's effect size. The authors investigated coverage probability for confidence intervals for the new effect size measure. The confidence intervals were constructed by using the noncentral t distribution and the percentile bootstrap. Over the range of distributions and effect sizes investigated in the study, coverage probability was better for the percentile bootstrap confidence interval.

Analysis of Variance↗

The new and improved two-sample T test.

This article considers the problem of comparing two independent groups in terms of some measure of location. It is well known that with Student's two-independent-sample t test, the actual level of significance can be well above or below the nominal level, confidence intervals can have inaccurate probability coverage, and power can be low relative to other methods. A solution to deal with heterogeneity is Welch's (1938) test. Welch's test deals with heteroscedasticity but can have poor power under arbitrarily small departures from normality. Yuen (1974) generalized Welch's test to trimmed means; her method provides improved control over the probability of a Type I error, but problems remain. Transformations for skewness improve matters, but the probability of a Type I error remains unsatisfactory in some situations. We find that a transformation for skewness combined with a bootstrap method improves Type I error control and probability coverage even if sample sizes are small.

Humans↗

Multivariate tests of means in independent groups designs. Effects of covariance heterogeneity and nonnormality.

Health evaluation research often employs multivariate designs in which data on several outcome variables are obtained for independent groups of subjects. This article examines statistical procedures for testing hypotheses of multivariate mean equality in two-group designs. The conventional test for multivariate means, Hotelling's T2, rests on certain assumptions about the distribution of the data and the population variances and covariances. When these assumptions are violated, which is often the case in applied health research, T2 will result in invalid conclusions about the null hypothesis. This article describes parametric procedures that are robust, or insensitive, to assumption violations. A numeric example illustrates the statistical concepts that are presented and a computer program to implement these robust solutions is introduced.

Algorithms↗

Pairwise multiple comparison test procedures: an update for clinical child and adolescent psychologists.

Locating pairwise differences among treatment groups is a common practice of applied researchers. Articles published in this journal have addressed the issue of statistical inference within the context of an analysis of variance (ANOVA) framework, describing procedures for comparing means, among other issues. In particular, 1 article (Jaccard & Guilamo-Ramos, 2002b) presented some new methods of performing contrasts of means whereas another presented a framework for obtaining robust tests within this same context (Jaccard & Guilamo-Ramos, 2002a). The purpose of this article is to add to these contributions by presenting some newer methods for conducting pairwise comparisons of means, that is by extending the contributions of the first article and applying the framework of the second article to pairwise multiple comparisons. The newer methods are intended to provide additional sensitivity to detect treatment group differences and provide tests that are robust to the effects of variance heterogeneity, nonnormality, or both.

Adolescent↗

Comparing measures of the 'typical' score across treatment groups.

Researchers can adopt one of many different measures of central tendency to examine the effect of a treatment variable across groups. These include least squares means, trimmed means, M-estimators and medians. In addition, some methods begin with a preliminary test to determine the shapes of distributions before adopting a particular estimator of the typical score. We compared a number of recently developed adaptive robust methods with respect to their ability to control Type I error and their sensitivity to detect differences between the groups when data were non-normal and heterogeneous, and the design was unbalanced. In particular, two new approaches to comparing the typical score across treatment groups, due to Babu, Padmanabhan, and Puri, were compared to two new methods presented by Wilcox and by Keselman, Wilcox, Othman, and Fradette. The procedures examined generally resulted in good Type I error control and therefore, on the basis of this critetion, it would be difficult to recommend one method over the other. However, the power results clearly favour one of the methods presented by Wilcox and Keselman; indeed, in the vast majority of the cases investigated, this most favoured approach had substantially larger power values than the other procedures, particularly when there were more than two treatment groups.

Humans↗

Modern robust data analysis methods: measures of central tendency.

Various statistical methods, developed after 1970, offer the opportunity to substantially improve upon the power and accuracy of the conventional t test and analysis of variance methods for a wide range of commonly occurring situations. The authors briefly review some of the more fundamental problems with conventional methods based on means; provide some indication of why recent advances, based on robust measures of location (or central tendency), have practical value; and describe why modern investigations dealing with nonnormality find practical problems when comparing means, in contrast to earlier studies. Some suggestions are made about how to proceed when using modern methods.

Humans↗

A generally robust approach to hypothesis testing in independent and correlated groups designs.

Standard least squares analysis of variance methods suffer from poor power under arbitrarily small departures from normality and fail to control the probability of a Type I error when standard assumptions are violated. These problems are vastly reduced when using a robust measure of location; incorporating bootstrap methods can result in additional benefits. This paper illustrates the use of trimmed means with an approximate degrees of freedom heteroskedastic statistic for independent and correlated groups designs in order to achieve robustness to the biasing effects of nonnormality and variance heterogeneity. As well, we indicate when a boostrap methodology can be effectively employed to provide improved Type I error control. We also illustrate, with examples from the psychophysiological literature, the use of a new computer program to obtain numerical results for these solutions.

Algorithms↗

Repeated measures one-way ANOVA based on a modified one-step M-estimator.

Wilcox, Keselman, Muska and Cribbie (2000) found a method for comparing the trimmed means of dependent groups that performed well in simulations, in terms of Type I errors, with a sample size as small as 21. Theory and simulations indicate that little power is lost under normality when using trimmed means rather than untrimmed means, and trimmed means can result in substantially higher power when sampling from a heavy-tailed distribution. However, trimmed means suffer from two practical concerns described in this paper. Replacing trimmed means with a robust M-estimator addresses these concerns, but control over the probability of a Type I error can be unsatisfactory when the sample size is small. Methods based on a simple modification of a one-step M-estimator that address the problems with trimmed means are examined. Several omnibus tests are compared, one of which performed well in simulations, even with a sample size of 11.

Analysis of Variance↗

Pairwise multiple comparisons: a model comparison approach versus stepwise procedures.

Researchers in the behavioural sciences have been presented with a host of pairwise multiple comparison procedures that attempt to obtain an optimal combination of Type I error control, power, and ease of application. However, these procedures share one important limitation: intransitive decisions. Moreover, they can be characterized as a piecemeal approach to the problem rather than a holistic approach. Dayton has recently proposed a new approach to pairwise multiple comparisons testing that eliminates intransitivity through a model selection procedure. The present study compared the model selection approach (and a protected version) with three powerful and easy-to-use stepwise multiple comparison procedures in terms of the proportion of times that the procedure identified the true pattern of differences among a set of means across several one-way layouts. The protected version of the model selection approach selected the true model a significantly greater proportion of times than the stepwise procedures and, in most cases, was not affected by variance heterogeneity and non-normality.

Humans↗

Controlling the rate of Type I error over a large set of statistical tests.

When many tests of significance are examined in a research investigation with procedures that limit the probability of making at least one Type I error--the so-called familywise techniques of control--the likelihood of detecting effects can be very low. That is, when familywise error controlling methods are adopted to assess statistical significance, the size of the critical value that must be exceeded in order to obtain statistical significance can be extremely large when the number of tests to be examined is also very large. In our investigation we examined three methods for increasing the sensitivity to detect effects when family size is large: the false discovery rate of error control presented by Benjamini and Hochberg (1995), a modified false discovery rate presented by Benjamini and Hochberg (2000) which estimates the number of true null hypotheses prior to adopting false discovery rate control, and a familywise method modified to control the probability of committing two or more Type I errors in the family of tests examined--not one, as is the case with the usual familywise techniques. Our results indicated that the level of significance for the two or more familywise method of Type I error control varied with the testing scenario and needed to be set on occasion at values in excess of 0.15 in order to control the two or more rate at a reasonable value of 0.01. In addition, the false discovery rate methods typically resulted in substantially greater power to detect non-null effects even though their levels of significance were set at the standard 0.05 value. Accordingly, we recommend the Benjamini and Hochberg (1995, 2000) methods of Type I error control when the number of tests in the family is large.

Achievement↗

Mixed-model pairwise multiple comparisons of repeated measures means.

One approach to the analysis of repeated measures data allows researchers to model the covariance structure of the data rather than presume a certain structure, as is the case with conventional univariate and multivariate test statistics. This mixed-model approach was evaluated for testing all possible pairwise differences among repeated measures marginal means in a Between-Subjects x Within-Subjects design. Specifically, the authors investigated Type I error and power rates for a number of simultaneous and stepwise multiple comparison procedures using SAS (1999) PROC MIXED in unbalanced designs when normality and covariance homogeneity assumptions did not hold. J. P. Shaffer's (1986) sequentially rejective step-down and Y. Hochberg's (1988) sequentially acceptive step-up Bonferroni procedures, based on an unstructured covariance structure, had superior Type I error control and power to detect true pairwise differences across the investigated conditions.

Humans↗

The analysis of repeated measures designs: a review.

Repeated measures ANOVA can refer to many different types of analysis. Specifically, this vague term can refer to conventional tests of significance, one of three univariate solutions with adjusted degrees of freedom, two different types of multivariate statistic, or approaches that combine univariate and multivariate tests. Accordingly, it is argued that, by only reporting probability values and referring to statistical analyses as repeated measures ANOVA, authors convey neither the type of analysis that was used nor the validity of the reported probability value, since each of these approaches has its own strengths and weaknesses. The various approaches are presented with a discussion of their strengths and weaknesses, and recommendations are made regarding the 'best' choice of analysis. Additional topics discussed include analyses for missing data and tests of linear contrasts.

Analysis of Variance↗

An examination of the robustness of the empirical Bayes and other approaches for testing main and interaction effects in repeated measures designs.

In a previous paper, Boik presented an empirical Bayes (EB) approach to the analysis of repeated measurements. The EB approach is a blend of the conventional univariate and multivariate approaches. Specifically, in the EB approach, the underlying covariance matrix is estimated by a weighted sum of the univariate and multivariate estimators. In addition to demonstrating that his approach controls test size and frequently is more powerful than either the epsilon-adjusted univariate or multivariate approaches, Boik showed how conventional multivariate software can be used to conduct EB analyses. Our investigation examined the Type I error properties of the EB approach when its derivational assumptions were not satisfied as well as when other factors known to affect the conventional tests of significance were varied. For comparative purposes we also investigated procedures presented by Huynh and by Keselman, Carriere, and Lix, procedures designed for non-spherical data and covariance heterogeneity, as well as an adjusted univariate and multivariate test statistic. Our results indicate that when the response variable is normally distributed and group sizes are equal, the EB approach was robust to violations of its derivational assumptions and therefore is recommended due to the power findings reported by Boik. However, we also found that both the EB approach and the adjusted univariate and multivariate procedures were prone to depressed or elevated rates of Type I error when data were non-normally distributed and covariance matrices and group sizes were either positively or negatively paired with one another. On the other hand, the Huynh and Keselman et al. procedures were generally robust to these same pairings of covariance matrices and group sizes.

Bayes Theorem↗

Repeated measures ANOVA: some new results on comparing trimmed means and means.

This paper considers the common problem of testing the equality of means in a repeated measures design. Recent results indicate that practical problems can arise when computing confidence intervals for all pairwise differences of the means in conjunction with the Bonferroni inequality. This suggests, and is confirmed here, that a problem might occur when performing an omnibus test of equal means. The problem is that the probability of rejecting is not minimized when the means are equal and the usual univariate F test is used with the Huynh-Feldt correction (epsilon) for the degrees of freedom. That is, power can actually decrease as the mean of one group is lowered, although eventually it increases. A similar problem is found when using a multivariate method (Hotelling's T2). Moreover, the probability of a Type I error can exceed the nominal level by a large amount. The paper considers methods for correcting this problem, and new results on comparing trimmed means are reported as well. In terms of both Type I errors and power, simulations reported here suggest that a percentile t bootstrap used with 20% trimmed means and an analogue of the epsilon-adjusted F gives the best results. This is consistent with extant theoretical results comparing methods based on means with trimmed means.

Analysis of Variance↗

Testing treatment effects in repeated measures designs: trimmed means and bootstrapping.

Non-normality and covariance heterogeneity between groups affect the validity of the traditional repeated measures methods of analysis, particularly when group sizes are unequal. A non-pooled Welch-type statistic (WJ) and the Huynh Improved General Approximation (IGA) test generally have been found to be effective in controlling rates of Type I error in unbalanced non-spherical repeated measures designs even though data are non-normal in form and covariance matrices are heterogeneous. However, under some conditions of departure from multisample sphericity and multivariate normality their rates of Type I error have been found to be elevated. Westfall and Young's results suggest that Type I error control could be improved by combining bootstrap methods with methods based on trimmed means. Accordingly, in our investigation we examined four methods for testing for main and interaction effects in a between- by within-subjects repeated measures design: (a) the IGA and WJ tests with least squares estimators based on theoretically determined critical values; (b) the IGA and WJ tests with least squares estimators based on empirically determined critical values; (c) the IGA and WJ tests with robust estimators based on theoretically determined critical values; and (d) the IGA and WJ tests with robust estimators based on empirically determined critical values. We found that the IGA tests were always robust to assumption violations whether based on least squares or robust estimators or whether critical values were obtained through theoretical or empirical methods. The WJ procedure, however, occasionally resulted in liberal rates of error when based on least squares estimators but always proved robust when applied with robust estimators. Neither approach particularly benefited from adopting bootstrapped critical values. Recommendations are provided to researchers regarding when each approach is best.

Humans↗

Testing treatment effects in repeated measures designs: an update for psychophysiological researchers.

In 1987, Jennings enumerated data analysis procedures that authors must follow for analyzing effects in repeated measures designs when submitting papers to Psychophysiology. These prescriptions were intended to counteract the effects of nonspherical data, a condition know to produce biased tests of significance. Since this editorial policy was established, additional refinements to the analysis of these designs have appeared in print in a number of sources that are not likely to be routinely read by psychophysiological researchers. Accordingly, this paper includes additional procedures not previously enumerated in the editorial policy that can be used to analyze repeated measurements. Furthermore, I indicate how numerical solutions can easily be obtained.

Bias↗