Search PubMed⌕ Search

Biomedical subjects

Peter Bühlmann

Publications and source records attributed to Peter Bühlmann.

9 recordsLinked to original sources

Model-based boosting in high dimensions.

SUMMARY: The R add-on package mboost implements functional gradient descent algorithms (boosting) for optimizing general loss functions utilizing componentwise least squares, either of parametric linear form or smoothing splines, or regression trees as base learners for fitting generalized linear, additive and interaction models to potentially high-dimensional data. AVAILABILITY: Package mboost is available from the Comprehensive R Archive Network (http://CRAN.R-project.org) under the terms of the General Public Licence (GPL).

Algorithms↗

A systematic comparison and evaluation of biclustering methods for gene expression data.

MOTIVATION: In recent years, there have been various efforts to overcome the limitations of standard clustering approaches for the analysis of gene expression data by grouping genes and samples simultaneously. The underlying concept, which is often referred to as biclustering, allows to identify sets of genes sharing compatible expression patterns across subsets of samples, and its usefulness has been demonstrated for different organisms and datasets. Several biclustering methods have been proposed in the literature; however, it is not clear how the different techniques compare with each other with respect to the biological relevance of the clusters as well as with other characteristics such as robustness and sensitivity to noise. Accordingly, no guidelines concerning the choice of the biclustering method are currently available. RESULTS: First, this paper provides a methodology for comparing and validating biclustering methods that includes a simple binary reference model. Although this model captures the essential features of most biclustering approaches, it is still simple enough to exactly determine all optimal groupings; to this end, we propose a fast divide-and-conquer algorithm (Bimax). Second, we evaluate the performance of five salient biclustering algorithms together with the reference model and a hierarchical clustering method on various synthetic and real datasets for Saccharomyces cerevisiae and Arabidopsis thaliana. The comparison reveals that (1) biclustering in general has advantages over a conventional hierarchical clustering approach, (2) there are considerable performance differences between the tested methods and (3) already the simple reference model delivers relevant patterns within all considered settings.

Algorithms↗

Low-order conditional independence graphs for inferring genetic networks.

As a powerful tool for analyzing full conditional (in-)dependencies between random variables, graphical models have become increasingly popular to infer genetic networks based on gene expression data. However, full (unconstrained) conditional relationships between random variables can be only estimated accurately if the number of observations is relatively large in comparison to the number of variables, which is usually not fulfilled for high-throughput genomic data. Recently, simplified graphical modeling approaches have been proposed to determine dependencies between gene expression profiles. For sparse graphical models such as genetic networks, it is assumed that the zero- and first-order conditional independencies still reflect reasonably well the full conditional independence structure between variables. Moreover, low-order conditional independencies have the advantage that they can be accurately estimated even when having only a small number of observations. Therefore, using only zero- and first-order conditional dependencies to infer the complete graphical model can be very useful. Here, we analyze the statistical and probabilistic properties of these low-order conditional independence graphs (called 0-1 graphs). We find that for faithful graphical models, the 0-1 graph contains at least all edges of the full conditional independence graph (concentration graph). For simple structures such as Markov trees, the 0-1 graph even coincides with the concentration graph. Furthermore, we present some asymptotic results and we demonstrate in a simulation study that despite their simplicity, 0-1 graphs are generally good estimators of sparse graphical models. Finally, the biological relevance of some applications is summarized.

Algorithms↗

Survival ensembles.

We propose a unified and flexible framework for ensemble learning in the presence of censoring. For right-censored data, we introduce a random forest algorithm and a generic gradient boosting algorithm for the construction of prognostic and diagnostic models. The methodology is utilized for predicting the survival time of patients suffering from acute myeloid leukemia based on clinical and genetic covariates. Furthermore, we compare the diagnostic capabilities of the proposed censored data random forest and boosting methods, applied to the recurrence-free survival time of node-positive breast cancer patients, with previously published findings.

Algorithms↗

Sparse graphical Gaussian modeling of the isoprenoid gene network in Arabidopsis thaliana.

We present a novel graphical Gaussian modeling approach for reverse engineering of genetic regulatory networks with many genes and few observations. When applying our approach to infer a gene network for isoprenoid biosynthesis in Arabidopsis thaliana, we detect modules of closely connected genes and candidate genes for possible cross-talk between the isoprenoid pathways. Genes of downstream pathways also fit well into the network. We evaluate our approach in a simulation study and using the yeast galactose network.

Arabidopsis↗

Gene expression signatures identify rhabdomyosarcoma subtypes and detect a novel t(2;2)(q35;p23) translocation fusing PAX3 to NCOA1.

Rhabdomyosarcoma is a pediatric tumor type, which is classified based on histological criteria into two major subgroups, namely embryonal rhabdomyosarcoma and alveolar rhabdomyosarcoma. The majority, but not all, alveolar rhabdomyosarcoma carry the specific PAX3(7)/FKHR-translocation, whereas there is no consistent genetic abnormality recognized in embryonal rhabdomyosarcoma. To gain additional insight into the genetic characteristics of these subtypes, we used oligonucleotide microarrays to measure the expression profiles of a group of 29 rhabdomyosarcoma biopsy samples (15 embryonal rhabdomyosarcoma, and 10 translocation-positive and 4 translocation-negative alveolar rhabdomyosarcoma). Hierarchical clustering revealed expression signatures clearly discriminating all three of the subgroups. Differentially expressed genes included several tyrosine kinases and G protein-coupled receptors, which might be amenable to pharmacological intervention. In addition, the alveolar rhabdomyosarcoma signature was used to classify an additional alveolar rhabdomyosarcoma case lacking any known PAX3 or PAX7 fusion as belonging to the translocation-positive group, leading to the identification of a novel translocation t(2;2)(q35;p23), which generates a fusion protein composed of PAX3 and the nuclear receptor coactivator NCOA1, having similar transactivation properties as PAX3/FKHR. These experiments demonstrate for the first time that gene expression profiling is capable of identifying novel chromosomal translocations.

Base Sequence↗

Gene expression profiles and risk stratification in childhood acute lymphoblastic leukemia.

BACKGROUND AND OBJECTIVES: Childhood acute lymphoblastic leukemia (ALL) is a heterogeneous disease. There are several distinct genetic subtypes, characterized by typical changes in gene expression pattern. In addition to cytogenetic markers, the in vivo response to treatment is an emerging prognostic marker for risk stratification. However, it has not yet been reported whether gene expression profiles can predict risk group stratification already at the time of diagnosis. DESIGN AND METHODS: We analyzed bone marrow samples of 31 ALL patients to identify changes in gene expression that are associated with the current risk assignment, irrespective of the genetic subtype. Gene expression profiles were established using oligonucleotide microarrays. RESULTS: Considering all low- and high-risk patients, no gene was capable of predicting the risk assignment already at time of diagnosis. However, screening for risk group associated genes using more homogeneous subsets of patients revealed 10(6) discriminatory probe sets. The prognostic significance of these probe sets was subsequently determined for the entire series of patients. Using the selected subgroups as the training set and the remaining samples as an independent test set, logistic regression using 3 predictor variables could accurately predict current risk assignment for 10 out of 12 patients. INTERPRETATION AND CONCLUSIONS: Gene expression profiles established from a cytogenetically heterogeneous study group are not, as yet, sufficiently accurate to be used prognostically in a clinical setting. Additional risk-associated gene expression analyses need to be performed in more homogeneous sets of patients.

Child↗

Boosting for tumor classification with gene expression data.

MOTIVATION: Microarray experiments generate large datasets with expression values for thousands of genes but not more than a few dozens of samples. Accurate supervised classification of tissue samples in such high-dimensional problems is difficult but often crucial for successful diagnosis and treatment. A promising way to meet this challenge is by using boosting in conjunction with decision trees. RESULTS: We demonstrate that the generic boosting algorithm needs some modification to become an accurate classifier in the context of gene expression data. In particular, we present a feature preselection method, a more robust boosting procedure and a new approach for multi-categorical problems. This allows for slight to drastic increase in performance and yields competitive results on several publicly available datasets. AVAILABILITY: Software for the modified boosting algorithms as well as for decision trees is available for free in R at http://stat.ethz.ch/~dettling/boosting.html.

Algorithms↗

Supervised clustering of genes.

BACKGROUND: We focus on microarray data where experiments monitor gene expression in different tissues and where each experiment is equipped with an additional response variable such as a cancer type. Although the number of measured genes is in the thousands, it is assumed that only a few marker components of gene subsets determine the type of a tissue. Here we present a new method for finding such groups of genes by directly incorporating the response variables into the grouping process, yielding a supervised clustering algorithm for genes. RESULTS: An empirical study on eight publicly available microarray datasets shows that our algorithm identifies gene clusters with excellent predictive potential, often superior to classification with state-of-the-art methods based on single genes. Permutation tests and bootstrapping provide evidence that the output is reasonably stable and more than a noise artifact. CONCLUSIONS: In contrast to other methods such as hierarchical clustering, our algorithm identifies several gene clusters whose expression levels clearly distinguish the different tissue types. The identification of such gene clusters is potentially useful for medical diagnostics and may at the same time reveal insights into functional genomics.

Algorithms↗