Search PubMed⌕ Search

Biomedical subjects

Yonatan Bilu

Publications and source records attributed to Yonatan Bilu.

5 recordsLinked to original sources

Conservation of expression and sequence of metabolic genes is reflected by activity across metabolic states.

Variation in gene expression levels on a genomic scale has been detected among different strains, among closely related species, and within populations of genetically identical cells. What are the driving forces that lead to expression divergence in some genes and conserved expression in others? Here we employ flux balance analysis to address this question for metabolic genes. We consider the genome-scale metabolic model of Saccharomyces cerevisiae, and its entire space of optimal and near-optimal flux distributions. We show that this space reveals underlying evolutionary constraints on expression regulation, as well as on the conservation of the underlying gene sequences. Genes that have a high range of optimal flux levels tend to display divergent expression levels among different yeast strains and species. This suggests that gene regulation has diverged in those parts of the metabolic network that are less constrained. In addition, we show that genes that are active in a large fraction of the space of optimal solutions tend to have conserved sequences. This supports the possibility that there is less selective pressure to maintain genes that are relevant for only a small number of metabolic states.

Cell Proliferation↗

The design of transcription-factor binding sites is affected by combinatorial regulation.

BACKGROUND: Transcription factors regulate gene expression by binding to specific cis-regulatory elements in gene promoters. Although DNA sequences that serve as transcription-factor binding sites have been characterized and associated with the regulation of numerous genes, the principles that govern the design and evolution of such sites are poorly understood. RESULTS: Using the comprehensive mapping of binding-site locations available in Saccharomyces cerevisiae, we examined possible factors that may have an impact on binding-site design. We found that binding sites tend to be shorter and fuzzier when they appear in promoter regions that bind multiple transcription factors. We further found that essential genes bind relatively fewer transcription factors, as do divergent promoters. We provide evidence that novel binding sites tend to appear in specific promoters that are already associated with multiple sites. CONCLUSION: Two principal models may account for the observed correlations. First, it may be that the interaction between multiple factors compensates for the decreased specificity of each specific binding sequence. In such a scenario, binding-site fuzziness is a consequence of the presence of multiple binding sites. Second, binding sites may tend to appear in promoter regions that are subject to low selective pressure, which also allows for fuzzier motifs. The latter possibility may account for the relatively low number of binding sites found in promoters of essential genes and in divergent promoters.

Binding Sites↗

ProtoNet: hierarchical classification of the protein space.

The ProtoNet site provides an automatic hierarchical clustering of the SWISS-PROT protein database. The clustering is based on an all-against-all BLAST similarity search. The similarities' E-score is used to perform a continuous bottom-up clustering process by applying alternative rules for merging clusters. The outcome of this clustering process is a classification of the input proteins into a hierarchy of clusters of varying degrees of granularity. ProtoNet (version 1.3) is accessible in the form of an interactive web site at http://www.protonet.cs.huji.ac.il. ProtoNet provides navigation tools for monitoring the clustering process with a vertical and horizontal view. Each cluster at any level of the hierarchy is assigned with a statistical index, indicating the level of purity based on biological keywords such as those provided by SWISS-PROT and InterPro. ProtoNet can be used for function prediction, for defining superfamilies and subfamilies and for large-scale protein annotation purposes.

Animals↗

The advantage of functional prediction based on clustering of yeast genes and its correlation with non-sequence based classifications.

Sequence similarity is probably the most widely used tool to infer functional linkage between proteins. The fully sequenced, much researched, genome of Saccharomyces cerevisiae gives us on opportunity to compare and statistically quantify computational methods based on sequence similarity, which aim to detect such linkage. In addition, the amount of data regarding Saccharomyces Cerevisiae genes and proteins, which is not directly based on sequence is rapidly increasing. Consequently, it allows investigation of the connections and correlation between classification based on these types of data and that based solely on sequence similarity. In this work we start with a simple clustering algorithm to cluster genes based on the BLAST E-score of their similarity. We analyze how well one can infer function from these clusters and for how many of the genes that are currently unknown one can suggest a prediction. Given these parameters, we show that even a simple algorithm achieves better results than simply considering the BLAST output of matching genes. In the second part of the paper, we show that there is a highly significant correlation (p-value < 10(-4) for the vast majority of the experiments) between the aforementioned clusters and other types of classifications. Namely, we show that a pair of genes being clustered together is correlated with these genes having similar expression patterns in DNA array experiments and with the encoded proteins being involved in protein-protein interactions. Although this correlation is highly significant, it is, of course, not strong enough to be, by itself, a tool for predicting co-regulation of genes or interaction of proteins. We discuss possible explanations for this correlation. Furthermore, the statistical evaluation of these results should be considered when developing tools that are aimed at making such predictions.

Algorithms↗

Faster algorithms for optimal multiple sequence alignment based on pairwise comparisons.

Multiple Sequence Alignment (MSA) is one of the most fundamental problems in computational molecular biology. The running time of the best known scheme for finding an optimal alignment, based on dynamic programming, increases exponentially with the number of input sequences. Hence, many heuristics were suggested for the problem. We consider a version of the MSA problem where the goal is to find an optimal alignment in which matches are restricted to positions in predefined matching segments. We present several techniques for making the dynamic programming algorithm more efficient, while still finding an optimal solution under these restrictions. We prove that it suffices to find an optimal alignment of the predefined sequence segments, rather than single letters, thereby reducing the input size and thus improving the running time. We also identify "shortcuts" that expedite the dynamic programming scheme. Empirical study shows that, taken together, these observations lead to an improved running time over the basic dynamic programming algorithm by 4 to 12 orders of magnitude, while still obtaining an optimal solution. Under the additional assumption that matches between segments are transitive, we further improve the running time for finding the optimal solution by restricting the search space of the dynamic programming algorithm.

Algorithms↗