Search PubMed⌕ Search

Biomedical subjects

Mattias Wahde

Publications and source records attributed to Mattias Wahde.

4 recordsLinked to original sources

A survey of methods for classification of gene expression data using evolutionary algorithms.

The rapid increase in the quantity of available biologic data over the last decade, brought about by the introduction of massively parallel methods for gene expression measurements, has highlighted the need for more efficient computational techniques for analysis. This paper reviews the use of evolutionary algorithms (EAs) in connection with classification based on gene expression data matrices. Brief introductions to data classification methods and EAs are given, followed by a survey of studies dealing with the application of evolutionary algorithms to various (cancer related) data sets. The general conclusion, based on the published results surveyed here, is that EAs may constitute an efficient method for optimal gene selection, and can also help in reducing the size (number of features used) of classifiers. In many cases, the classification accuracy obtained using EAs, often in conjunction with other methods, represents a significant improvement over results obtained without the use of EAs. However, long-term, independent clinical follow-up studies will be essential to validate prognostic markers identified by the use of EA-based methods.

Algorithms↗

Effective dimensionality for principal component analysis of time series expression data.

Large-scale expression data are today measured for thousands of genes simultaneously. This development has been followed by an exploration of theoretical tools to get as much information out of these data as possible. Several groups have used principal component analysis (PCA) for this task. However, since this approach is data-driven, care must be taken in order not to analyze the noise instead of the data. As a strong warning towards uncritical use of the output from a PCA, we employ a newly developed procedure to judge the effective dimensionality of a specific data set. Although this data set is obtained during the development of rat central nervous system, our finding is a general property of noisy time series data. Based on knowledge of the noise-level for the data, we find that the effective number of dimensions that are meaningful to use in a PCA is much lower than what could be expected from the number of measurements. We attribute this fact both to effects of noise and the lack of independence of the expression levels. Finally, we explore the possibility to increase the dimensionality by performing more measurements within one time series, and conclude that this is not a fruitful approach.

Algorithms↗

Assessing the significance of consistently mis-regulated genes in cancer associated gene expression matrices.

MOTIVATION: The simplest level of statistical analysis of cancer associated gene expression matrices is aimed at finding consistently up- or down-regulated genes within a given set of tumor samples. Considering the high level of gene expression diversity detected in cancer, one needs to assess the probability that the consistent mis-regulation of a given gene is due to chance. Furthermore, it is important to determine the required sample number that will ensure the meaningful statistical analysis of massively parallel gene expression measurements. RESULTS: The probability of consistent mis-regulation is calculated in this paper for binarized gene expression data, using combinatorial considerations. For practical purposes, we also provide a set of accurate approximate formulas for determining the same probability in a computationally less intensive way. When the pool of mis-regulatable genes is restricted, the probability of consistent mis-regulation can be overestimated. We show, however, that this effect has little practical consequences for cancer associated gene expression measurements published in the literature. Finally, in order to aid experimental design, we have provided estimates on the required sample number that will ensure that the detected consistent mis-regulation is not due to chance. Our results suggest that less than 20 sufficiently diverse tumor samples may be enough to identify consistently mis-regulated genes in a statistically significant manner. AVAILABILITY: An implementation using Mathematica (tm) of the main equation of the paper, (4), is available at www.me.chalmers.se/~mwahde/bioinfo.html.

Algorithms↗

Effective dimensionality of large-scale expression data using principal component analysis.

Large-scale expression data are today measured for thousands of genes simultaneously. This development is followed by an exploration of theoretical tools to get as much information out of these data as possible. One line is to try to extract the underlying regulatory network. The models used thus far, however, contain many parameters, and a careful investigation is necessary in order not to over-fit the models. We employ principal component analysis to show how, in the context of linear additive models, one can get a rough estimate of the effective dimensionality (the number of information-carrying dimensions) of large-scale gene expression datasets. We treat both the lack of independence of different measurements in a time series and the fact that that measurements are subject to some level of noise, both of which reduce the effective dimensionality and thereby constrain the complexity of models which can be built from the data.

Gene Expression Profiling↗