Search PubMed⌕ Search

Biomedical subjects

Amos Tanay

Publications and source records attributed to Amos Tanay.

15 recordsLinked to original sources

Reciprocal, methylation-dependent binding of Zfp57 and Gzf1 safeguards Dlk1-Dio3 imprinting during developmental reprogramming.

Genomic imprinting secures parent-specific gene expression through differential DNA methylation at imprinted control regions (ICRs). However, how unmethylated alleles resist de novo methylation remains unclear. Using an allelic Dlk1-Dio3 ICR methylation reporter and genome-wide loss-of-function screening, we identify the zinc finger protein GZF1 that binds the unmethylated maternal ICR and protects it from de novo methylation via a regulatory element containing GZF1 and ZFP57 motifs that mediates mutually exclusive, methylation-dependent binding. Loss of either factor causes reciprocal imprinting failure: Gzf1 loss induces maternal allele methylation, H3K4me3 depletion, and silencing of maternal transcripts, whereas Zfp57 loss results in maternalization. Remarkably, GZF1 protects the unmethylated ICR from de novo methylation in both oocytes and embryos, and its loss leads to perinatal death consistent with paternalization of the maternal allele. Together, our findings establish a reciprocal mechanism that maintains parental epigenetic asymmetry across both imprint establishment and embryonic reprogramming.

Animals↗

Constitutive nucleosome depletion and ordered factor assembly at the GRP78 promoter revealed by single molecule footprinting.

Chromatin organization and transcriptional regulation are interrelated processes. A shortcoming of current experimental approaches to these complex events is the lack of methods that can capture the activation process on single promoters. We have recently described a method that combines methyltransferase M.SssI treatment of intact nuclei and bisulfite sequencing allowing the representation of replicas of single promoters in terms of protected and unprotected footprint modules. Here we combine this method with computational analysis to study single molecule dynamics of transcriptional activation in the stress inducible GRP78 promoter. We show that a 350-base pair region upstream of the transcription initiation site is constitutively depleted of nucleosomes, regardless of the induction state of the promoter, providing one of the first examples for such a promoter in mammals. The 350-base pair nucleosome-free region can be dissected into modules, identifying transcription factor binding sites and their combinatorial organization during endoplasmic reticulum stress. The interaction of the transcriptional machinery with the GRP78 core promoter is highly organized, represented by six major combinatorial states. We show that the TATA box is frequently occupied in the noninduced state, that stress induction results in sequential loading of the endoplasmic reticulum stress response elements, and that a substantial portion of these elements is no longer occupied following recruitment of factors to the transcription initiation site. Studying the positioning of nucleosomes and transcription factors at the single promoter level provides a powerful tool to gain novel insights into the transcriptional process in eukaryotes.

Base Pairing↗

Quantification of protein half-lives in the budding yeast proteome.

A complete description of protein metabolism requires knowledge of the rates of protein production and destruction within cells. Using an epitope-tagged strain collection, we measured the half-life of >3,750 proteins in the yeast proteome after inhibition of translation. By integrating our data with previous measurements of protein and mRNA abundance and translation rate, we provide evidence that many proteins partition into one of two regimes for protein metabolism: one optimized for efficient production or a second optimized for regulatory efficiency. Incorporation of protein half-life information into a simple quantitative model for protein production improves our ability to predict steady-state protein abundance values. Analysis of a simple dynamic protein production model reveals a remarkable correlation between transcriptional regulation and protein half-life within some groups of coregulated genes, suggesting that cells coordinate these two processes to achieve uniform effects on protein abundances. Our experimental data and theoretical analysis underscore the importance of an integrative approach to the complex interplay between protein degradation, transcriptional regulation, and other determinants of protein metabolism.

Fungal Proteins↗

Extensive low-affinity transcriptional interactions in the yeast genome.

Major experimental and computational efforts are targeted at the characterization of transcriptional networks on a genomic scale. The ultimate goal of many of these studies is to construct networks associating transcription factors with genes via well-defined binding sites. Weaker regulatory interactions other than those occurring at high-affinity binding sites are largely ignored and are not well understood. Here I show that low-affinity interactions are abundant in vivo and quantifiable from current high-throughput ChIP experiments. I develop algorithms that predict DNA-binding energies from sequences and ChIP data across a wide dynamic range of affinities and use them to reveal widespread functionality of low-affinity transcription factor binding. Evolutionary analysis suggests that binding energies of many transcription factors are conserved even in promoters lacking classical binding sites. Gene expression analysis shows that such promoters can generate significant expression. I estimate that while only a small percentage of the genome is strongly regulated by a typical transcription factor, up to an order of magnitude more may be involved in weaker interactions. Low-affinity transcription factor-DNA interaction may therefore be important both evolutionarily and functionally.

Algorithms↗

A probabilistic methodology for integrating knowledge and experiments on biological networks.

Biological systems are traditionally studied by focusing on a specific subsystem, building an intuitive model for it, and refining the model using results from carefully designed experiments. Modern experimental techniques provide massive data on the global behavior of biological systems, and systematically using these large datasets for refining existing knowledge is a major challenge. Here we introduce an extended computational framework that combines formalization of existing qualitative models, probabilistic modeling, and integration of high-throughput experimental data. Using our methods, it is possible to interpret genomewide measurements in the context of prior knowledge on the system, to assign statistical meaning to the accuracy of such knowledge, and to learn refined models with improved fit to the experiments. Our model is represented as a probabilistic factor graph, and the framework accommodates partial measurements of diverse biological elements. We study the performance of several probabilistic inference algorithms and show that hidden model variables can be reliably inferred even in the presence of feedback loops and complex logic. We show how to refine prior knowledge on combinatorial regulatory relations using hypothesis testing and derive p-values for learned model features. We test our methodology and algorithms on a simulated model and on two real yeast models. In particular, we use our method to explore uncharacterized relations among regulators in the yeast response to hyper-osmotic shock and in the yeast lysine biosynthesis system. Our integrative approach to the analysis of biological regulation is demonstrated to synergistically combine qualitative and quantitative evidence into concrete biological predictions.

Cell Physiological Phenomena↗

EXPANDER--an integrative program suite for microarray data analysis.

BACKGROUND: Gene expression microarrays are a prominent experimental tool in functional genomics which has opened the opportunity for gaining global, systems-level understanding of transcriptional networks. Experiments that apply this technology typically generate overwhelming volumes of data, unprecedented in biological research. Therefore the task of mining meaningful biological knowledge out of the raw data is a major challenge in bioinformatics. Of special need are integrative packages that provide biologist users with advanced but yet easy to use, set of algorithms, together covering the whole range of steps in microarray data analysis. RESULTS: Here we present the EXPANDER 2.0 (EXPression ANalyzer and DisplayER) software package. EXPANDER 2.0 is an integrative package for the analysis of gene expression data, designed as a 'one-stop shop' tool that implements various data analysis algorithms ranging from the initial steps of normalization and filtering, through clustering and biclustering, to high-level functional enrichment analysis that points to biological processes that are active in the examined conditions, and to promoter cis-regulatory elements analysis that elucidates transcription factors that control the observed transcriptional response. EXPANDER is available with pre-compiled functional Gene Ontology (GO) and promoter sequence-derived data files for yeast, worm, fly, rat, mouse and human, supporting high-level analysis applied to data obtained from these six organisms. CONCLUSION: EXPANDER integrated capabilities and its built-in support of multiple organisms make it a very powerful tool for analysis of microarray data. The package is freely available for academic users at http://www.cs.tau.ac.il/~rshamir/expander.

Algorithms↗

Conservation and evolvability in regulatory networks: the evolution of ribosomal regulation in yeast.

Transcriptional modules of coregulated genes play a key role in regulatory networks. Comparative studies show that modules of coexpressed genes are conserved across taxa. However, little is known about the mechanisms underlying the evolution of module regulation. Here, we explore the evolution of cis-regulatory programs associated with conserved modules by integrating expression profiles for two yeast species and sequence data for a total of 17 fungal genomes. We show that although the cis-elements accompanying certain conserved modules are strictly conserved, those of other conserved modules are remarkably diverged. In particular, we infer the evolutionary history of the regulatory program governing ribosomal modules. We show how a cis-element emerged concurrently in dozens of promoters of ribosomal protein genes, followed by the loss of a more ancient cis-element. We suggest that this formation of an intermediate redundant regulatory program allows conserved transcriptional modules to gradually switch from one regulatory mechanism to another while maintaining their functionality. Our work provides a general framework for the study of the dynamics of promoter evolution at the level of transcriptional modules and may help in understanding the evolvability and increased redundancy of transcriptional regulation in higher organisms.

Computational Biology↗

A global view of pleiotropy and phenotypically derived gene function in yeast.

Pleiotropy, the ability of a single mutant gene to cause multiple mutant phenotypes, is a relatively common but poorly understood phenomenon in biology. Perhaps the greatest challenge in the analysis of pleiotropic genes is determining whether phenotypes associated with a mutation result from the loss of a single function or of multiple functions encoded by the same gene. Here we estimate the degree of pleiotropy in yeast by measuring the phenotypes of 4710 mutants under 21 environmental conditions, finding that it is significantly higher than predicted by chance. We use a biclustering algorithm to group pleiotropic genes by common phenotype profiles. Comparisons of these clusters to biological process classifications, synthetic lethal interactions, and protein complex data support the hypothesis that this method can be used to genetically define cellular functions. Applying these functional classifications to pleiotropic genes, we are able to dissect phenotypes into groups associated with specific gene functions.

Algorithms↗

Integrative analysis of genome-wide experiments in the context of a large high-throughput data compendium.

Biological systems are orchestrated by heterogeneous regulatory programs that control complex processes and adapt to a dynamic environment. Recent advances in high-throughput experimental methods provide genome-wide perspectives on such regulatory programs. A considerable amount of data on the behavior of model systems in a variety of conditions is rapidly accumulating. Still, the dominant paradigm is to analyze new genome-wide experiments separately from any other extant data, for example, by clustering the new data alone. Here we introduce a new methodology for analyzing the results of a new functional genomic study vis-à-vis a large compendium of previously published results from heterogeneous experimental techniques. We demonstrate our methodology on Saccharomyces cerevisiae, using a compendium of some 2000 experiments from 60 different publications. Most importantly, we show how the integrated analysis reveals unexpected connections among biological processes, and differentiates between novel and known effects in the analyzed experiments. Such characterization is impossible when new data sets are studied in isolation. Our results exemplify the power of the integrative approach in the analysis of genomic scale data sets and call for a paradigm shift in their study.

Algorithms↗

Revealing modularity and organization in the yeast molecular network by integrated analysis of highly heterogeneous genomewide data.

The dissection of complex biological systems is a challenging task, made difficult by the size of the underlying molecular network and the heterogeneous nature of the control mechanisms involved. Novel high-throughput techniques are generating massive data sets on various aspects of such systems. Here, we perform analysis of a highly diverse collection of genomewide data sets, including gene expression, protein interactions, growth phenotype data, and transcription factor binding, to reveal the modular organization of the yeast system. By integrating experimental data of heterogeneous sources and types, we are able to perform analysis on a much broader scope than previous studies. At the core of our methodology is the ability to identify modules, namely, groups of genes with statistically significant correlated behavior across diverse data sources. Numerous biological processes are revealed through these modules, which also obey global hierarchical organization. We use the identified modules to study the yeast transcriptional network and predict the function of >800 uncharacterized genes. Our analysis framework, SAMBA (Statistical-Algorithmic Method for Bicluster Analysis), enables the processing of current and future sources of biological information and is readily extendable to experimental techniques and higher organisms.

Amino Acids↗

Multilevel modeling and inference of transcription regulation.

The understanding of transcription regulation is a major goal of today's biology. The challenge is to utilize diverse high-throughput data in order to infer mechanistic models of transcription control. We propose a new model which integrates transcription factor-gene affinity, protein abundance, and gene expression profiles. The model provides a detailed, yet computationally tractable description of the relations between transcription factors, their binding sites at gene promoters, and the combinatorial regulation of transcription. At the core, our model manipulates dose-affinity-response functions that associate transcription factor concentrations and transcription factor-DNA affinities to determine the rate of transcription factor-DNA reactions. We study computational problems that arise in optimizing such models and develop polynomial algorithms for certain problems. We show how to assess missing values (notably protein abundance) and describe a novel framework to infer models from currently available data. On budding yeast carbohydrate metabolism data, our results demonstrate the sensitivity and specificity of the approach. They also suggest new active binding sites and a regulation model for the transcription program of the galactose system.

Algorithms↗

Modeling and analysis of heterogeneous regulation in biological networks.

In this study, we propose a novel model for the representation of biological networks and provide algorithms for learning model parameters from experimental data. Our approach is to build an initial model based on extant biological knowledge and refine it to increase the consistency between model predictions and experimental data. Our model encompasses networks which contain heterogeneous biological entities (mRNA, proteins, metabolites) and aims to capture diverse regulatory circuitry on several levels (metabolism, transcription, translation, post-translation and feedback loops, among them). Algorithmically, the study raises two basic questions: how to use the model for predictions and inference of hidden variables states, and how to extend and rectify model components. We show that these problems are hard in the biologically relevant case where the network contains cycles. We provide a prediction methodology in the presence of cycles and a polynomial time, constant factor approximation for learning the regulation of a single entity. A key feature of our approach is the ability to utilize both high-throughput experimental data, which measure many model entities in a single experiment, as well as specific experimental measurements of few entities or even a single one. In particular, we use together gene expression, growth phenotypes, and proteomics data. We tested our strategy on the lysine biosynthesis pathway in yeast. We constructed a model of more than 150 variables based on an extensive literature survey and evaluated it with diverse experimental data. We used our learning algorithms to propose novel regulatory hypotheses in several cases where the literature-based model was inconsistent with the experiments. We showed that our approach has better accuracy than extant methods of learning regulation.

Algorithms↗

A global view of the selection forces in the evolution of yeast cis-regulation.

The interaction between transcription factors and their DNA binding sites is key to understanding gene regulation. By performing a genome-wide study of the evolutionary dynamics in yeast promoters, we provide a first global view of the network of selection forces in the evolution of transcription factor binding sites. This analysis gives rise to new models for binding site activity, identifies families of related binding sites, and characterizes the functional similarities among them. We discovered rich and highly optimized selective pressures operating inside and around these families. In several cases, this organization reveals that a single transcription factor has multiple functional modes. We demonstrate how such functional heterogeneity is related to the binding site's affinity and how it is exploited in transcription programs.

Binding Sites↗

Discovering statistically significant biclusters in gene expression data.

In gene expression data, a bicluster is a subset of the genes exhibiting consistent patterns over a subset of the conditions. We propose a new method to detect significant biclusters in large expression datasets. Our approach is graph theoretic coupled with statistical modelling of the data. Under plausible assumptions, our algorithm is polynomial and is guaranteed to find the most significant biclusters. We tested our method on a collection of yeast expression profiles and on a human cancer dataset. Cross validation results show high specificity in assigning function to genes based on their biclusters, and we are able to annotate in this way 196 uncharacterized yeast genes. We also demonstrate how the biclusters lead to detecting new concrete biological associations. In cancer data we are able to detect and relate finer tissue types than was previously possible. We also show that the method outperforms the biclustering algorithm of Cheng and Church (2000).

Algorithms↗

Minreg: inferring an active regulator set.

Regulatory relations between genes are an important component of molecular pathways. Here, we devise a novel global method that uses a set of gene expression profiles to find a small set of relevant active regulators, identify the genes that they regulate, and automatically annotate them. We show that our algorithm is capable of handling a large number of genes in a short time and is robust to a wide range of parameters. We apply our method to a combined dataset of S. cerevisiae expression profiles, and validate the resulting model of regulation by cross-validation and extensive biological analysis of the selected regulators and their derived annotations.

Algorithms↗