Search PubMed⌕ Search

PubMed · 16026595

Reproducible clusters from microarray research: whither?

Abstract

MOTIVATION: In cluster analysis, the validity of specific solutions, algorithms, and procedures present significant challenges because there is no null hypothesis to test and no 'right answer'. It has been noted that a replicable classification is not necessarily a useful one, but a useful one that characterizes some aspect of the population must be replicable. By replicable we mean reproducible across multiple samplings from the same population. Methodologists have suggested that the validity of clustering methods should be based on classifications that yield reproducible findings beyond chance levels. We used this approach to determine the performance of commonly used clustering algorithms and the degree of replicability achieved using several microarray datasets. METHODS: We considered four commonly used iterative partitioning algorithms (Self Organizing Maps (SOM), K-means, Clutsering LARge Applications (CLARA), and Fuzzy C-means) and evaluated their performances on 37 microarray datasets, with sample sizes ranging from 12 to 172. We assessed reproducibility of the clustering algorithm by measuring the strength of relationship between clustering outputs of subsamples of 37 datasets. Cluster stability was quantified using Cramer's v2 from a kXk table. Cramer's v2 is equivalent to the squared canonical correlation coefficient between two sets of nominal variables. Potential scores range from 0 to 1, with 1 denoting perfect reproducibility. RESULTS: All four clustering routines show increased stability with larger sample sizes. K-means and SOM showed a gradual increase in stability with increasing sample size. CLARA and Fuzzy C-means, however, yielded low stability scores until sample sizes approached 30 and then gradually increased thereafter. Average stability never exceeded 0.55 for the four clustering routines, even at a sample size of 50. These findings suggest several plausible scenarios: (1) microarray datasets lack natural clustering structure thereby producing low stability scores on all four methods; (2) the algorithms studied do not produce reliable results and/or (3) sample sizes typically used in microarray research may be too small to support derivation of reliable clustering results. Further research should be directed towards evaluating stability performances of more clustering algorithms on more datasets specially having larger sample sizes with larger numbers of clusters considered.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nikhil R Garge, Grier P Page, Alan P Sprague, Bernard S Gorman, David B Allison. 2005-07-15. Reproducible clusters from microarray research: whither?. https://doi.org/10.1186/1471-2105-6-s2-s10

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

A spatial scan statistic for ordinal data.

Spatial scan statistics are widely used for count data to detect geographical disease clusters of high or low incidence, mortality or prevalence and to evaluate their statistical significance. Some data are ordinal or continuous in nature, however, so that it is necessary to dichotomize the data to use a traditional scan statistic for count data. There is then a loss of information and the choice of cut-off point is often arbitrary. In this paper, we propose a spatial scan statistic for ordinal data, which allows us to analyse such data incorporating the ordinal structure without making any further assumptions. The test statistic is based on a likelihood ratio test and evaluated using Monte Carlo hypothesis testing. The proposed method is illustrated using prostate cancer grade and stage data from the Maryland Cancer Registry. The statistical power, sensitivity and positive predicted value of the test are examined through a simulation study.

Cluster Analysis↗

Analysis of cluster randomized cross-over trial data: a comparison of methods.

In a cluster randomized cross-over trial, all participating clusters receive both intervention and control treatments consecutively, in separate time periods. Patients recruited by each cluster within the same time period receive the same intervention, and randomization determines order of treatment within a cluster. Such a design has been used on a number of occasions. For analysis of the trial data, the approach of analysing cluster-level summary measures is appealing on the grounds of simplicity, while hierarchical modelling allows for the correlation of patients within periods within clusters and offers flexibility in the model assumptions. We consider several cluster-level approaches and hierarchical models and make comparison in terms of empirical precision, coverage, and practical considerations. The motivation for a cluster randomized trial to employ cross-over of trial arms is particularly strong when the number of clusters available is small, so we examine performance of the methods under small, medium and large (6, 18, 30) numbers of clusters. One hierarchical model and two cluster-level methods were found to perform consistently well across the designs considered. These three methods are efficient, provide appropriate standard errors and coverage, and continue to perform well when incorporating adjustment for an individual-level covariate. We conclude that choice between hierarchical models and cluster-level methods should be influenced by the extent of complexity in the planned analysis.

Cluster Analysis↗