Search PubMed⌕ Search

Biomedical subjects

Jaakko Astola

Publications and source records attributed to Jaakko Astola.

7 recordsLinked to original sources

Optimized LOWESS normalization parameter selection for DNA microarray data.

BACKGROUND: Microarray data normalization is an important step for obtaining data that are reliable and usable for subsequent analysis. One of the most commonly utilized normalization techniques is the locally weighted scatterplot smoothing (LOWESS) algorithm. However, a much overlooked concern with the LOWESS normalization strategy deals with choosing the appropriate parameters. Parameters are usually chosen arbitrarily, which may reduce the efficiency of the normalization and result in non-optimally normalized data. Thus, there is a need to explore LOWESS parameter selection in greater detail. RESULTS AND DISCUSSION: In this work, we discuss how to choose parameters for the LOWESS method. Moreover, we present an optimization approach for obtaining the fraction of data points utilized in the local regression and analyze results for local print-tip normalization. The optimization procedure determines the bandwidth parameter for the local regression by minimizing a cost function that represents the mean-squared difference between the LOWESS estimates and the normalization reference level. We demonstrate the utility of the systematic parameter selection using two publicly available data sets. The first data set consists of three self versus self hybridizations, which allow for a quantitative study of the optimization method. The second data set contains a collection of DNA microarray data from a breast cancer study utilizing four breast cancer cell lines. Our results show that different parameter choices for the bandwidth window yield dramatically different calibration results in both studies. CONCLUSIONS: Results derived from the self versus self experiment indicate that the proposed optimization approach is a plausible solution for estimating the LOWESS parameters, while results from the breast cancer experiment show that the optimization procedure is readily applicable to real-life microarray data normalization. In summary, the systematic approach to obtain critical parameters in the LOWESS technique is likely to produce data that optimally meets assumptions made in the data preprocessing step and thereby makes studies utilizing the LOWESS method unambiguous and easier to repeat.

Algorithms↗

Effects of Herceptin treatment on global gene expression patterns in HER2-amplified and nonamplified breast cancer cell lines.

Herceptin is a humanized monoclonal antibody targeted against the extracellular domain of the HER2 oncogene, which is amplified and overexpressed in 10-34% of breast cancers. Herceptin therapy provides effective treatment in HER2-positive metastatic breast cancer, although a favorable treatment response is not achieved in all cases. Here, we show that Herceptin treatment induces a dose-dependent growth reduction in breast cancer cell lines with HER2 amplification, whereas nonamplified cell lines are practically resistant. Time-course analysis of global gene expression patterns in amplified and nonamplified cell lines indicated a major change in transcript levels between 24 and 48 h of Herceptin treatment. A step-wise gene selection algorithm revealed a set of 439 genes whose temporal expression profiles differed most between the amplified and nonamplified cell lines. The discriminatory power of these genes was confirmed by both hierarchical clustering and self-organizing map analyses. In the amplified cell lines, the Herceptin treatment induced the expression of several genes involved in RNA processing and DNA repair, while cell adhesion mediators and known oncogenes, such as c-FOS and c-KIT, were downregulated. These results provide additional clues to the downstream effects of blocking the HER2 pathway in breast cancer and may provide new targets for more effective treatment.

Antibodies, Monoclonal↗

Fast iterative gene clustering based on information theoretic criteria for selecting the cluster structure.

Grouping of genes into clusters according to their expression levels is important for deriving biological information, e.g., on gene functions based on microarray and other related analyses. The paper introduces the selection of the number of clusters based on the minimum description length (MDL) principle for the selection of the number of clusters in gene expression data. The main feature of the new method is the ability to evaluate in a fast way the number of clusters according to the sound MDL principle, without exhaustive evaluations over all possible partitions of the gene set. The estimation method can be used in conjunction with various clustering algorithms. A recent clustering algorithm using principal component analysis, the "gene shaving" (GS) procedure, can be modified to make use of the new MDL estimation method, replacing the Gap statistics originally used in GS algorithm. The resulting clustering algorithm is shown to perform better than GS-Gap and CEM (classification expectation maximization), in the simulations using artificial data. The proposed method is applied to B-cell differentiation data, and the resulting clusters are compared with those found by self-organizing maps (SOM).

Algorithms↗

A novel strategy for microarray quality control using Bayesian networks.

MOTIVATION: High-throughput microarray technologies enable measurements of the expression levels of thousands of genes in parallel. However, microarray printing, hybridization and washing may create substantial variability in the quality of the data. As erroneous measurements may have a drastic impact on the results by disturbing the normalization schemes and by introducing expression patterns that lead to incorrect conclusions, it is crucial to discard low quality observations in the early phases of a microarray experiment. A typical microarray experiment consists of tens of thousands of spots on a microarray, making manual extraction of poor quality spots impossible. Thus, there is a need for a reliable and general microarray spot quality control strategy. RESULTS: We suggest a novel strategy for spot quality control by using Bayesian networks, which contain many appealing properties in the spot quality control context. We illustrate how a non-linear least squares based Gaussian fitting procedure can be used in order to extract features for a spot on a microarray. The features we used in this study are: spot intensity, size of the spot, roundness of the spot, alignment error, background intensity, background noise, and bleeding. We conclude that Bayesian networks are a reliable and useful model for microarray spot quality assessment. SUPPLEMENTARY INFORMATION: http://sigwww.cs.tut.fi/TICSP/SpotQuality/.

Algorithms↗

The role of certain Post classes in Boolean network models of genetic networks.

A topic of great interest and debate concerns the source of order and remarkable robustness observed in genetic regulatory networks. The study of the generic properties of Boolean networks has proven to be useful for gaining insight into such phenomena. The main focus, as regards ordered behavior in networks, has been on canalizing functions, internal homogeneity or bias, and network connectivity. Here we examine the role that certain classes of Boolean functions that are closed under composition play in the emergence of order in Boolean networks. The closure property implies that any gene at any number of steps in the future is guaranteed to be governed by a function from the same class. By means of Derrida curves on random Boolean networks and percolation simulations on square lattices, we demonstrate that networks constructed from functions belonging to these classes have a tendency toward ordered behavior. Thus they are not overly sensitive to initial conditions, and damage does not readily spread throughout the network. In addition, the considered classes are significantly larger than the class of canalizing functions as the connectivity increases. The functions in these classes exhibit the same kind of preference toward biased functions as do canalizing functions, meaning that functions from this class are likely to be biased. Finally, functions from this class have a natural way of ensuring robustness against noise and perturbations, thus representing plausible evolutionarily selected candidates for regulatory rules in genetic networks.

Computational Biology↗

CGH-Plotter: MATLAB toolbox for CGH-data analysis.

CGH-Plotter is a MATLAB toolbox with a graphical user interface for the analysis of comparative genomic hybridization (CGH) microarray data. CGH-Plotter provides a tool for rapid visualization of CGH-data according to the locations of the genes along the genome. In addition, the CGH-Plotter identifies regions of amplifications and deletions, using k-means clustering and dynamic programming. The application offers a convenient way to analyze CGH-data and can also be applied for the analysis of cDNA microarray expression data. CGH-Plotter toolbox is platform independent and requires MATLAB 6.1 or higher to operate.

Cluster Analysis↗

Data extraction from composite oligonucleotide microarrays.

Microarray or DNA chip technology is revolutionizing biology by empowering researchers in the collection of broad-scope gene information. It is well known that microarray-based measurements exhibit a substantial amount of variability due to a number of possible sources, ranging from hybridization conditions to image capture and analysis. In order to make reliable inferences and carry out quantitative analysis with microarray data, it is generally advisable to have more than one measurement of each gene. The availability of both between-array and within-array replicate measurements is essential for this purpose. Although statistical considerations call for increasing the number of replicates of both types, the latter is particularly challenging in practice due to a number of limiting factors, especially for in-house spotting facilities. We propose a novel approach to design so-called composite microarrays, which allow more replicates to be obtained without increasing the number of printed spots.

Gene Expression Profiling↗