Search PubMed⌕ Search

Biomedical subjects

Michael J Brusco

Publications and source records attributed to Michael J Brusco.

6 recordsLinked to original sources

Clustering, seriation, and subset extraction of confusion data.

The study of confusion data is a well established practice in psychology. Although many types of analytical approaches for confusion data are available, among the most common methods are the extraction of 1 or more subsets of stimuli, the partitioning of the complete stimulus set into distinct groups, and the ordering of the stimulus set. Although standard commercial software packages can sometimes facilitate these types of analyses, they are not guaranteed to produce optimal solutions. The authors present a MATLAB *.m file for preprocessing confusion matrices, which includes fitting of the similarity-choice model. Two additional MATLAB programs are available for optimally clustering stimuli on the basis of confusion data. The authors also developed programs for optimally ordering stimuli and extracting subsets of stimuli using information from confusion matrices. Together, these programs provide several pragmatic alternatives for the applied researcher when analyzing confusion data. Although the programs are described within the context of confusion data, they are also amenable to other types of proximity data.

Cluster Analysis↗

Bicriterion methods for partitioning dissimilarity matrices.

Partitioning indices associated with the within-cluster sums of pairwise dissimilarities often exhibit a systematic bias towards clusters of a particular size, whereas minimization of the partition diameter (i.e. the maximum dissimilarity element across all pairs of objects within the same cluster) does not typically have this problem. However, when the partition-diameter criterion is used, there is often a myriad of alternative optimal solutions that can vary significantly with respect to their substantive interpretation. We propose a bicriterion partitioning approach that considers both diameter and within-cluster sums in the optimization problem and facilitates selection from among the alternative optima. We developed several MATLAB-based exchange algorithms that rapidly provide excellent solutions to bicriterion partitioning problems. These algorithms were evaluated using synthetic data sets, as well as an empirical dissimilarity matrix.

Cluster Analysis↗

Bicriterion seriation methods for skew-symmetric matrices.

The decomposition of an asymmetric proximity matrix into its symmetric and skew-symmetric components is a well-known principle in combinatorial data analysis. The seriation of the skew-symmetric component can emphasize information corresponding to the sign or absolute magnitude of the matrix elements, and the choice of objective criterion can have a profound impact on the ordering. In this research note, we propose a bicriterion approach for seriation of a skew-symmetric matrix incorporating both sign and magnitude information. Two numerical demonstrations reveal that the bicriterion procedure is an effective alternative to direct seriation of the skew-symmetric matrix, facilitating favourable trade-offs among sign and magnitude information.

Humans↗

Clustering binary data in the presence of masking variables.

A number of important applications require the clustering of binary data sets. Traditional nonhierarchical cluster analysis techniques, such as the popular K-means algorithm, can often be successfully applied to these data sets. However, the presence of masking variables in a data set can impede the ability of the K-means algorithm to recover the true cluster structure. The author presents a heuristic procedure that selects an appropriate subset from among the set of all candidate clustering variables. Specifically, this procedure attempts to select only those variables that contribute to the definition of true cluster structure while eliminating variables that can hide (or mask) that true structure. Experimental testing of the proposed variable-selection procedure reveals that it is extremely successful at accomplishing this goal.

Cluster Analysis↗

On the concordance among empirical confusion matrices for visual and tactual letter recognition.

In this article, we examine the concordance among 19 empirical confusion matrices for visual and tactual recognition of capital letters of the alphabet. As a measure of concordance, we employed an index based on within-stimulus triads of letters. Unlike correlation measures of agreement that are based on a one-to-one matching of matrix elements, the selected index directly captures the internal structures of the confusion matrices prior to the comparison. Permutation tests revealed statistically significant concordance among 166 of 171 pairs of matrices in the study. Concordance of confusion structure among tactual matrices tended to be somewhat stronger than concordance among the visual matrices.

Empirical Research↗

An enhanced branch-and-bound algorithm for a partitioning problem.

This paper focuses on the problem of developing a partition of n objects based on the information in a symmetric, non-negative dissimilarity matrix. The goal is to partition the objects into a set of non-overlapping subsets with the objective of minimizing the sum of the within-subset dissimilarities. Optimal solutions to this problem can be obtained using dynamic programming, branch-and-bound and other mathematical programming methods. An improved branch-and-bound algorithm is shown to be particularly efficient. The improvements include better upper bounds that are obtained via a fast exchange algorithm and, more importantly, sharper lower bounds obtained through sequential solution of submatrices. A modified version of the branch-and-bound algorithm for minimizing the diameter of a partition is also presented. Computational results for both synthetic and empirical dissimilarity matrices reveal the effectiveness of the branch-and-bound methodology.

Algorithms↗