Search PubMed⌕ Search

Biomedical subjects

Mahlet G Tadesse

Publications and source records attributed to Mahlet G Tadesse.

8 recordsLinked to original sources

Bayesian variable selection for the analysis of microarray data with censored outcomes.

MOTIVATION: A common task in microarray data analysis consists of identifying genes associated with a phenotype. When the outcomes of interest are censored time-to-event data, standard approaches assess the effect of genes by fitting univariate survival models. In this paper, we propose a Bayesian variable selection approach, which allows the identification of relevant markers by jointly assessing sets of genes. We consider accelerated failure time (AFT) models with log-normal and log-t distributional assumptions. A data augmentation approach is used to impute the failure times of censored observations and mixture priors are used for the regression coefficients to identify promising subsets of variables. The proposed method provides a unified procedure for the selection of relevant genes and the prediction of survivor functions. RESULTS: We demonstrate the performance of the method on simulated examples and on several microarray datasets. For the simulation study, we consider scenarios with large number of noisy variables and different degrees of correlation between the relevant and non-relevant (noisy) variables. We are able to identify the correct covariates and obtain good prediction of the survivor functions. For the microarray applications, some of our selected genes are known to be related to the diseases under study and a few are in agreement with findings from other researchers. AVAILABILITY: The Matlab code for implementing the Bayesian variable selection method may be obtained from the corresponding author. CONTACT: mvannucci@stat.tamu.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Algorithms↗

Wavelet thresholding with bayesian false discovery rate control.

The false discovery rate (FDR) procedure has become a popular method for handling multiplicity in high-dimensional data. The definition of FDR has a natural Bayesian interpretation; it is the expected proportion of null hypotheses mistakenly rejected given a measure of evidence for their truth. In this article, we propose controlling the positive FDR using a Bayesian approach where the rejection rule is based on the posterior probabilities of the null hypotheses. Correspondence between Bayesian and frequentist measures of evidence in hypothesis testing has been studied in several contexts. Here we extend the comparison to multiple testing with control of the FDR and illustrate the procedure with an application to wavelet thresholding. The problem consists of recovering signal from noisy measurements. This involves extracting wavelet coefficients that result from true signal and can be formulated as a multiple hypotheses-testing problem. We use simulated examples to compare the performance of our approach to the Benjamini and Hochberg (1995, Journal of the Royal Statistical Society, Series B57, 289-300) procedure. We also illustrate the method with nuclear magnetic resonance spectral data from human brain.

Bayes Theorem↗

Bayesian error-in-variable survival model for the analysis of GeneChip arrays.

DNA microarrays in conjunction with statistical models may help gain a deeper understanding of the molecular basis for specific diseases. An intense area of research is concerned with the identification of genes related to particular phenotypes. The technology, however, is subject to various sources of error that may lead to expression readings that are substantially different from the true transcript levels. Few methods for microarray data analysis have accounted for measurement error in a substantial way and that is the purpose of this investigation. We describe a Bayesian error-in-variable model for the analysis of microarray data from a clinical study of patients with acute lymphoblastic leukemia. We focus in particular on the problem of identifying genes whose expression patterns are associated with duration of remission. This is a question of great practical interest since relapse is a major concern in the treatment of this disease. We explore the effects of ignoring the uncertainty in the expression estimates on the selection and ranking of genes.

Bayes Theorem↗

Identification of DNA regulatory motifs using Bayesian variable selection.

MOTIVATION: Understanding the mechanisms that determine gene expression regulation is an important and challenging problem. A common approach consists of identifying DNA-binding sites from a collection of co-regulated genes and their nearby non-coding DNA sequences. Here, we consider a regression model that linearly relates gene expression levels to a sequence matching score of nucleotide patterns. We use Bayesian models and stochastic search techniques to select transcription factor binding site candidates, as an alternative to stepwise regression procedures used by other investigators. RESULTS: We demonstrate through simulated data the improved performance of the Bayesian variable selection method compared to the stepwise procedure. We then analyze and discuss the results from experiments involving well-studied pathways of Saccharomyces cerevisiae and Schizosaccharomyces pombe. We identify regulatory motifs known to be related to the experimental conditions considered. Some of our selected motifs are also in agreement with recent findings by other researchers. In addition, our results include novel motifs that constitute promising sets for further assessment. AVAILABILITY: The Matlab code for implementing the Bayesian variable selection method may be obtained from the corresponding author.

Algorithms↗

Bayesian variable selection in multinomial probit models to identify molecular signatures of disease stage.

Here we focus on discrimination problems where the number of predictors substantially exceeds the sample size and we propose a Bayesian variable selection approach to multinomial probit models. Our method makes use of mixture priors and Markov chain Monte Carlo techniques to select sets of variables that differ among the classes. We apply our methodology to a problem in functional genomics using gene expression profiling data. The aim of the analysis is to identify molecular signatures that characterize two different stages of rheumatoid arthritis.

Arthritis, Rheumatoid↗

A bayesian hierarchical model for the analysis of Affymetrix arrays.

An area of active research in DNA microarray analysis focuses on identifying differentially expressed genes between normal and malignant tissues. The analysis is complicated by the presence of several unreliable expression readings. Here, we illustrate a methodology where the expression estimates are modeled as censored data and discriminating genes are selected using ANOVA-based criteria.

Analysis of Variance↗

Unraveling gene-gene interactions regulated by ligands of the aryl hydrocarbon receptor.

The co-expression of genes coupled to additive probabilistic relationships was used to identify gene sets predictive of the complex biological interactions regulated by ligands of the aryl hydrocarbon receptor ((Italic)Ahr(/Italic)). To maximize the number of possible gene-gene combinations, data sets from murine embryonic kidney, fetal heart, and vascular smooth muscle cells challenged (Italic)in vitro(/Italic) with ligands of the (Italic)Ahr(/Italic) were used to create predictor/training data sets. Biologically relevant gene predictor sets were calculated for (Italic)Ahr(/Italic), cytochrome P450 1B1, insulin-like growth factor-binding protein-5, lysyl oxidase, and osteopontin. Transcript levels were categorized into ternary expressions and target genes selected from the data set and tested for all possible combinations using three gene sets as predictors of transitional level. The goodness of prediction for each set was quantified using a multivariate nonlinear coefficient of determination. Evidence is presented that predictor gene combinations can be effectively used to resolve gene-gene interactions regulated by (Italic)Ahr(/Italic) ligands. (Italic)Key words:(/Italic) aryl hydrocarbon receptor, bioinformatics, gene networks, genomics. (Italic)Environ Health Perspect (/Italic)112:403-412 (2004). [Online 14 January 2004]

Animals↗

Identification of differentially expressed genes in high-density oligonucleotide arrays accounting for the quantification limits of the technology.

In DNA microarray analysis, there is often interest in isolating a few genes that best discriminate between tissue types. This is especially important in cancer, where different clinicopathologic groups are known to vary in their outcomes and response to therapy. The identification of a small subset of gene expression patterns distinctive for tumor subtypes can help design treatment strategies and improve diagnosis. Toward this goal, we propose a methodology for the analysis of high-density oligonucleotide arrays. The gene expression measures are modeled as censored data to account for the quantification limits of the technology, and two gene selection criteria based on contrasts from an analysis of covariance (ANCOVA) model are presented. The model is formulated in a hierarchical Bayesian framework, which in addition to making the fit of the model straightforward and computationally efficient, allows us to borrow strength across genes. The elicitation of hierarchical priors, as well as issues related to parameter identifiability and posterior propriety, are discussed in detail. We examine the performance of our proposed method on simulated data, then present a detailed case study of an endometrial cancer dataset.

Biometry↗