Search PubMed⌕ Search

PubMed · 14581611

Robust singular value decomposition analysis of microarray data.

Abstract

In microarray data there are a number of biological samples, each assessed for the level of gene expression for a typically large number of genes. There is a need to examine these data with statistical techniques to help discern possible patterns in the data. Our technique applies a combination of mathematical and statistical methods to progressively take the data set apart so that different aspects can be examined for both general patterns and very specific effects. Unfortunately, these data tables are often corrupted with extreme values (outliers), missing values, and non-normal distributions that preclude standard analysis. We develop a robust analysis method to address these problems. The benefits of this robust analysis will be both the understanding of large-scale shifts in gene effects and the isolation of particular sample-by-gene effects that might be either unusual interactions or the result of experimental flaws. Our method requires a single pass and does not resort to complex "cleaning" or imputation of the data table before analysis. We illustrate the method with a commercial data set.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Li Liu, Douglas M Hawkins, Sujoy Ghosh, S Stanley Young. 2003-10-27. Robust singular value decomposition analysis of microarray data.. https://doi.org/10.1073/pnas.1733249100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Clustering individuals using INMTD: a novel versatile multi-view embedding framework integrating omics and imaging data.

MOTIVATION: Combining omics and images can lead to a more comprehensive clustering of individuals than classic single-view approaches. Among the various approaches for multi-view clustering, nonnegative matrix tri-factorization (NMTF) and nonnegative Tucker decomposition (NTD) are advantageous in learning low-rank embeddings with promising interpretability. Besides, there is a need to handle unwanted drivers of clusterings (i.e. confounders). RESULTS: In this work, we introduce a novel multi-view clustering method based on NMTF and NTD, named INMTD, which integrates omics and 3D imaging data to derive unconfounded subgroups of individuals. According to the adjusted Rand index, INMTD outperformed other clustering methods on a synthetic dataset with known clusters. In the application to real-life facial-genomic data, INMTD generated biologically relevant embeddings for individuals, genetics, and facial morphology. By removing confounded embedding vectors, we derived an unconfounded clustering with better internal and external quality; the genetic and facial annotations of each derived subgroup highlighted distinctive characteristics. In conclusion, INMTD can effectively integrate omics data and 3D images for unconfounded clustering with biologically meaningful interpretation. AVAILABILITY AND IMPLEMENTATION: INMTD is freely available at https://github.com/ZuqiLi/INMTD.

Cluster Analysis↗

Focused library design in GPCR projects on the example of 5-HT(2c) agonists: comparison of structure-based virtual screening with ligand-based search methods.

The aim of this study was to investigate the usefulness of structure-based virtual screening (VS) for focused library design in G protein-coupled receptors (GPCR) projects on the example of 5-HT(2c) agonists. We compared the performance of structure-based VS against two different homology models using FRED for docking and ScreenScore, FlexX, and PMF for rescoring with the results of 12 ligand-based similarity searches using four different query compounds and three different similarity metrics (Daylight, FTree, Phacir). The result of the similarity search showed much variation, from an enrichment factor up to 3.2 to worse than random, whereas the structure-based VS gave a more stable result with a constant enrichment factor around 2. Additionally, actives retrieved by the structure-based approach were more diverse than the actives among the top scorers of the similarity searches. Based on these results, we suggest basing a focused library design for a GPCR project on a combination of a ligand-based similarity search and structure-based docking.

Cluster Analysis↗