Search PubMed⌕ Search

Biomedical subjects

Janne Nikkilä

Publications and source records attributed to Janne Nikkilä.

5 recordsLinked to original sources

Differential endocrine regulation of genes enriched in initial segment and distal caput of the mouse epididymis as revealed by genome-wide expression profiling.

We have performed genome-wide expression profiling of endocrine regulation of genes expressed in the mouse initial segment (IS) and distal caput of the epididymis by using Affymetrix microarrays. The data revealed that of the 15 020 genes expressed in the epididymis, 35% were enriched in one of the two regions studied, indicating that differential functions can be attributed to the IS and the more distal caput regions. The data, furthermore, showed that 27% of the genes expressed in the IS and/or distal caput epididymidis are under the regulation of testicular factors present in the duct fluid, while bloodborne androgens can regulate for 14% of them. This is in line with the high testis dependency of epididymal physiology. We then focused on genes with moderate or strong expression, showing strict segment enrichment and strong dependency on testicular factors. Analyses of the 59 genes, including upregulated and downregulated genes, fulfilling the criteria indicated that the expression of 18 (17 downregulated genes; 1 upregulated gene) of 19 gonadectomy-responsive genes enriched in the IS was not maintained by the androgen treatment, whereas the expression of all six downregulated genes enriched in the distal caput and the majority of those with no strict segment enrichment of expression (28 of 34; consisting of 23 downregulated and 5 upregulated genes) were maintained by androgens. Hence, it is evident that testicular factors other than androgens are important for the expression of IS-enriched genes, whereas the expression of distal caput-enriched genes is typically regulated by androgens. Identical data were obtained by independent clustering analyses performed for the expression data of 3626 epididymal genes. Several novel genes with putative involvement in epididymal sperm maturation, such as a disintegrin and metallopeptidase domain 28 (Adam28) and a solute carrier organic anion transporter family, member 4C1 (Slco4c1), were identified, indicating that this approach is successful for identifying novel epididymal genes.

ADAM Proteins↗

Exploratory modeling of yeast stress response and its regulation with gCCA and associative clustering.

We model dependencies between m multivariate continuous-valued information sources by a combination of (i) a generalized canonical correlations analysis (gCCA) to reduce dimensionality while preserving dependencies in m - 1 of them, and (ii) summarizing dependencies with the remaining one by associative clustering. This new combination of methods avoids multiway associative clustering which would require a multiway contingency table and hence suffer from curse of dimensionality of the table. The method is applied to summarizing properties of yeast stress by searching for dependencies (commonalities) between expression of genes of baker's yeast Saccharomyces cerevisiae in various stressful treatments, and summarizing stress regulation by finally adding data about transcription factor binding sites.

Animals↗

Trustworthiness and metrics in visualizing similarity of gene expression.

BACKGROUND: Conventionally, the first step in analyzing the large and high-dimensional data sets measured by microarrays is visual exploration. Dendrograms of hierarchical clustering, self-organizing maps (SOMs), and multidimensional scaling have been used to visualize similarity relationships of data samples. We address two central properties of the methods: (i) Are the visualizations trustworthy, i.e., if two samples are visualized to be similar, are they really similar? (ii) The metric. The measure of similarity determines the result; we propose using a new learning metrics principle to derive a metric from interrelationships among data sets. RESULTS: The trustworthiness of hierarchical clustering, multidimensional scaling, and the self-organizing map were compared in visualizing similarity relationships among gene expression profiles. The self-organizing map was the best except that hierarchical clustering was the most trustworthy for the most similar profiles. Trustworthiness can be further increased by treating separately those genes for which the visualization is least trustworthy. We then proceed to improve the metric. The distance measure between the expression profiles is adjusted to measure differences relevant to functional classes of the genes. The genes for which the new metric is the most different from the usual correlation metric are listed and visualized with one of the visualization methods, the self-organizing map, computed in the new metric. CONCLUSIONS: The conjecture from the methodological results is that the self-organizing map can be recommended to complement the usual hierarchical clustering for visualizing and exploring gene expression data. Discarding the least trustworthy samples and improving the metric still improves it.

Animals↗

Analysis and visualization of gene expression data using self-organizing maps.

Cluster structure of gene expression data obtained from DNA microarrays is analyzed and visualized with the Self-Organizing Map (SOM) algorithm. The SOM forms a non-linear mapping of the data to a two-dimensional map grid that can be used as an exploratory data analysis tool for generating hypotheses on the relationships, and ultimately of the function of the genes. Similarity relationships within the data and cluster structures can be visualized and interpreted. The methods are demonstrated by computing a SOM of yeast genes. The relationships of known functional classes of genes are investigated by analyzing their distribution on the SOM, the cluster structure is visualized by the U-matrix method, and the clusters are characterized in terms of the properties of the expression profiles of the genes. Finally, it is shown that the SOM visualizes the similarity of genes in a more trustworthy way than two alternative methods, multidimensional scaling and hierarchical clustering.

Cluster Analysis↗

Associative clustering for exploring dependencies between functional genomics data sets.

High-throughput genomic measurements, interpreted as cooccurring data samples from multiple sources, open up a fresh problem for machine learning: What is in common in the different data sets, that is, what kind of statistical dependencies are there between the paired samples from the different sets? We introduce a clustering algorithm for exploring the dependencies. Samples within each data set are grouped such that the dependencies between groups of different sets capture as much of pairwise dependencies between the samples as possible. We formalize this problem in a novel probabilistic way, as optimization of a Bayes factor. The method is applied to reveal commonalities and exceptions in gene expression between organisms and to suggest regulatory interactions in the form of dependencies between gene expression profiles and regulator binding patterns.

Algorithms↗