Search PubMed⌕ Search

Biomedical subjects

Michael Hörnquist

Publications and source records attributed to Michael Hörnquist.

6 recordsLinked to original sources

Reverse engineering galactose regulation in yeast through model selection.

We examine the application of statistical model selection methods to reverse-engineering the control of galactose utilization in yeast from DNA microarray experiment data. In these experiments, relationships among gene expression values are revealed through modifications of galactose sugar level and genetic perturbations through knockouts. For each gene variable, we select predictors using a variety of methods, taking into account the variance in each measurement. These methods include maximization of log-likelihood with Cp, AIC, and BIC penalties, bootstrap and cross-validation error estimation, and coefficient shrinkage via the Lasso.

Journal Article↗

Visualization of large-scale correlations in gene expressions.

Large-scale expression data are today measured for several thousands of genes simultaneously. Furthermore, most genes are being categorized according to their properties. This development has been followed by an exploration of theoretical tools to integrate these diverse data types. A key problem is the large noise-level in the data. Here, we investigate ways to extract the remaining signals within these noisy data sets. We find large-scale correlations within data from Saccharomyces cerevisiae with respect to properties of the encoded proteins. These correlations are visualized in a way that is robust to the underlying noise in the measurement of the individual gene expressions. In particular, for S. cerevisiae we observe that the proteins corresponding to the 400 highest expressed genes typically are localized to the cytoplasm. These most expressed genes are not essential for cell survival.

Cell Nucleus↗

Effective dimensionality for principal component analysis of time series expression data.

Large-scale expression data are today measured for thousands of genes simultaneously. This development has been followed by an exploration of theoretical tools to get as much information out of these data as possible. Several groups have used principal component analysis (PCA) for this task. However, since this approach is data-driven, care must be taken in order not to analyze the noise instead of the data. As a strong warning towards uncritical use of the output from a PCA, we employ a newly developed procedure to judge the effective dimensionality of a specific data set. Although this data set is obtained during the development of rat central nervous system, our finding is a general property of noisy time series data. Based on knowledge of the noise-level for the data, we find that the effective number of dimensions that are meaningful to use in a PCA is much lower than what could be expected from the number of measurements. We attribute this fact both to effects of noise and the lack of independence of the expression levels. Finally, we explore the possibility to increase the dimensionality by performing more measurements within one time series, and conclude that this is not a fruitful approach.

Algorithms↗

Scale-free growing networks imply linear preferential attachment.

It has been recognized for some time that a network grown by the addition of nodes with linear preferential attachment will possess a scale-free distribution of connectivities. Here we prove by some analytical arguments that the linearity is a necessary component to obtain this kind of distribution. However, the preferential linking rate does not necessarily apply to single nodes, but to groups of nodes of the same connectivity. We also point out that for a time-varying mean connectivity the linking rate will deviate from a linear expression by an extra asymptotically logarithmic term.

Journal Article↗

Effective dimensionality of large-scale expression data using principal component analysis.

Large-scale expression data are today measured for thousands of genes simultaneously. This development is followed by an exploration of theoretical tools to get as much information out of these data as possible. One line is to try to extract the underlying regulatory network. The models used thus far, however, contain many parameters, and a careful investigation is necessary in order not to over-fit the models. We employ principal component analysis to show how, in the context of linear additive models, one can get a rough estimate of the effective dimensionality (the number of information-carrying dimensions) of large-scale gene expression datasets. We treat both the lack of independence of different measurements in a time series and the fact that that measurements are subject to some level of noise, both of which reduce the effective dimensionality and thereby constrain the complexity of models which can be built from the data.

Gene Expression Profiling↗

Constructing and analyzing a large-scale gene-to-gene regulatory network--lasso-constrained inference and biological validation.

We construct a gene-to-gene regulatory network from time-series data of expression levels for the whole genome of the yeast Saccharomyces cerevisae, in a case where the number of measurements is much smaller than the number of genes in the network. This network is analyzed with respect to present biological knowledge of all genes (according to the Gene Ontology database), and we find some of its large-scale properties to be in accordance with known facts about the organism. The linear modeling employed here has been explored several times, but due to lack of any validation beyond investigating individual genes, it has been seriously questioned with respect to its applicability to biological systems. Our results show the adequacy of the approach and make further investigations of the model meaningful.

Algorithms↗