Search PubMed⌕ Search

Biomedical subjects

Werner Dubitzky

Publications and source records attributed to Werner Dubitzky.

10 recordsLinked to original sources

Avoiding model selection bias in small-sample genomic datasets.

MOTIVATION: Genomic datasets generated by high-throughput technologies are typically characterized by a moderate number of samples and a large number of measurements per sample. As a consequence, classification models are commonly compared based on resampling techniques. This investigation discusses the conceptual difficulties involved in comparative classification studies. Conclusions derived from such studies are often optimistically biased, because the apparent differences in performance are usually not controlled in a statistically stringent framework taking into account the adopted sampling strategy. We investigate this problem by means of a comparison of various classifiers in the context of multiclass microarray data. RESULTS: Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.

Algorithms↗

Instance-based concept learning from multiclass DNA microarray data.

BACKGROUND: Various statistical and machine learning methods have been successfully applied to the classification of DNA microarray data. Simple instance-based classifiers such as nearest neighbor (NN) approaches perform remarkably well in comparison to more complex models, and are currently experiencing a renaissance in the analysis of data sets from biology and biotechnology. While binary classification of microarray data has been extensively investigated, studies involving multiclass data are rare. The question remains open whether there exists a significant difference in performance between NN approaches and more complex multiclass methods. Comparative studies in this field commonly assess different models based on their classification accuracy only; however, this approach lacks the rigor needed to draw reliable conclusions and is inadequate for testing the null hypothesis of equal performance. Comparing novel classification models to existing approaches requires focusing on the significance of differences in performance. RESULTS: We investigated the performance of instance-based classifiers, including a NN classifier able to assign a degree of class membership to each sample. This model alleviates a major problem of conventional instance-based learners, namely the lack of confidence values for predictions. The model translates the distances to the nearest neighbors into 'confidence scores'; the higher the confidence score, the closer is the considered instance to a pre-defined class. We applied the models to three real gene expression data sets and compared them with state-of-the-art methods for classifying microarray data of multiple classes, assessing performance using a statistical significance test that took into account the data resampling strategy. Simple NN classifiers performed as well as, or significantly better than, their more intricate competitors. CONCLUSION: Given its highly intuitive underlying principles--simplicity, ease-of-use, and robustness--the k-NN classifier complemented by a suitable distance-weighting regime constitutes an excellent alternative to more complex models for multiclass microarray data sets. Instance-based classifiers using weighted distances are not limited to microarray data sets, but are likely to perform competitively in classifications of high-dimensional biological data sets such as those generated by high-throughput mass spectrometry.

Algorithms↗

Towards data warehousing and mining of protein unfolding simulation data.

OBJECTIVES: The prediction of protein structure and the precise understanding of protein folding and unfolding processes remains one of the greatest challenges in structural biology and bioinformatics. Computer simulations based on molecular dynamics (MD) are at the forefront of the effort to gain a deeper understanding of these complex processes. Currently, these MD simulations are usually on the order of tens of nanoseconds, generate a large amount of conformational data and are computationally expensive. More and more groups run such simulations and generate a myriad of data, which raises new challenges in managing and analyzing these data. Because the vast range of proteins researchers want to study and simulate, the computational effort needed to generate data, the large data volumes involved, and the different types of analyses scientists need to perform, it is desirable to provide a public repository allowing researchers to pool and share protein unfolding data. METHODS: To adequately organize, manage, and analyze the data generated by unfolding simulation studies, we designed a data warehouse system that is embedded in a grid environment to facilitate the seamless sharing of available computer resources and thus enable many groups to share complex molecular dynamics simulations on a more regular basis. RESULTS: To gain insight into the conformational fluctuations and stability of the monomeric forms of the amyloidogenic protein transthyretin (TTR), molecular dynamics unfolding simulations of the monomer of human TTR have been conducted. Trajectory data and meta-data of the wild-type (WT) protein and the highly amyloidogenic variant L55P-TTR represent the test case for the data warehouse. CONCLUSIONS: Web and grid services, especially pre-defined data mining services that can run on or 'near' the data repository of the data warehouse, are likely to play a pivotal role in the analysis of molecular dynamics unfolding data.

Computational Biology↗

Reverse-engineering gene-regulatory networks using evolutionary algorithms and grid computing.

OBJECTIVE: Living organisms regulate the expression of genes using complex interactions of transcription factors, messenger RNA and active protein products. Due to their complexity, gene-regulatory networks are not fully understood.However, by building computational models it is possible to gain insight into their function and operation. METHODS: Evolutionary algorithms are used to create computational models of gene-regulatory networks based on observed microarray data. These algorithms can be computationally intensive. They will be implemented within an existing grid computing infrastructure, that has been developed for data mining purposes, and which is able to deliver the required compute power. RESULTS: We discuss how models can built achieved using distributed and grid computing technology. In particular we investigate how Condor and JavaSpaces technology is suited to the requirements of our modeling approach. CONCLUSIONS: Determining network models of gene-regulatory networks using evolutionary algorithms not only requires considerable computational power, but also a modeling formalism that can explain the underlying dynamics.

Algorithms↗

Survival trees for analyzing clinical outcome in lung adenocarcinomas based on gene expression profiles: identification of neogenin and diacylglycerol kinase alpha expression as critical factors.

We present survival trees as an exploratory tool for revealing new insights into gene expression profiles in combination with clinical patient data. Survival trees partition the patient data studied into groups with similar survival outcomes and identify characteristic genetic profiles within these groups. We demonstrate the application of survival trees in a study involving the expression profiles of 3,588 genes in 211 lung adenocarcinoma patients. The survival tree identified a group of early-stage cancer patients with relatively low survival rates and another group of advanced-stage patients with remarkably good survival outcome. For both groups, the tree identified characteristic expression profiles of genes that might play a role in cancerogenesis and disease progression, notably the genes for the netrin receptor neogenin and the Ras/Rho kinase modulator diacylglycerol kinase alpha.

Adenocarcinoma↗

Mathematical models of cell cycle regulation.

The cell division cycle is a fundamental process of cell biology and a detailed understanding of its function, regulation and other underlying mechanisms is critical to many applications in biotechnology and medicine. Since a comprehensive analysis of the molecular mechanisms involved is too complex to be performed intuitively, mathematical and computational modelling techniques are essential. This paper is a review and analysis of recent approaches attempting to model cell cycle regulation by means of protein-protein interaction networks.

Algorithms↗

Protein folding and unfolding simulations: a new challenge for data mining.

One of the unsolved paradigms in molecular biology is the protein folding problem. In recent years, with the identification of several diseases as protein folding disorders and with the explosion of genome information and the need for efficient ways to predict protein structure, protein folding became a central issue in molecular sciences research. Using molecular dynamics unfolding simulations of an amyloidogenic protein--transthyretin--as an example, we put forward a series of ideas on how simulations of this type may be used to infer rules and unfolding behavior in amyloidogenic proteins, and to extrapolate rules for protein folding in different structural classes of proteins. These, in turn, could help in the development of protein structure prediction methods. The need to analyse different proteins and to run multiple simulations creates a huge amount of data which has to be stored, managed, analyzed and shared (database and Grid technology; data mining). Once the data is captured, the next challenge is to find meaningful patterns (associations, correlations, clusters, rules, relationships) among molecular properties, or their relative importance at different stages of the folding or unfolding processes. This clearly puts new and interesting challenges to the bioinformatics community.

Computational Biology↗

Representing bioinformatics causality.

This paper reviews a variety of different graphical notations currently in active use for modelling dynamic processes in bioinformatics and biotechnology, and crystallises from these notations a set of properties essential to any proposal for a modelling language seeking to provide an adequate systemic description of biological processes.

Algorithms↗

Multiclass cancer classification using gene expression profiling and probabilistic neural networks.

Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 or more. In addition, microarray data exhibit a high degree of noise. Most of the discussed methods do not adequately address the problem of dimensionality and noise. Furthermore, although machine learning and data mining methods are based on statistics, most such techniques do not address the biologist's requirement for sound mathematical confidence measures. Finally, most machine learning and data mining classification methods fail to incorporate misclassification costs, i.e. they are indifferent to the costs associated with false positive and false negative classifications. In this paper, we present a probabilistic neural network (PNN) model that addresses all these issues. The PNN model provides sound statistical confidences for its decisions, and it is able to model asymmetrical misclassification costs. Furthermore, we demonstrate the performance of the PNN for multiclass gene expression data sets. Here, we compare the performance of the PNN with two machine learning methods, a decision tree and a neural network. To assess and evaluate the performance of the classifiers, we use a lift-based scoring system that allows a fair comparison of different models. The PNN clearly outperformed the other models. The results demonstrate the successful application of the PNN model for multiclass cancer classification.

Artificial Intelligence↗