Search PubMed⌕ Search

Biomedical subjects

E Domany

Publications and source records attributed to E Domany.

At least 19 recordsLinked to original sources

Gene expression profiles of AML derived stem cells; similarity to hematopoietic stem cells.

Tumors contain a fraction of cancer stem cells that maintain the propagation of the disease. The CD34(+)CD38(-) cells, isolated from acute myeloid leukemia (AML), were shown to be enriched leukemic stem cells (LSC). We isolated the CD34(+)CD38(-) cell fraction from AML and compared their gene expression profiles to the CD34(+)CD38(+) cell fraction, using microarrays. We found 409 genes that were at least twofold over- or underexpressed between the two cell populations. These include underexpression of DNA repair, signal transduction and cell cycle genes, consistent with the relative quiescence of stem cells, and chromosomal aberrations and mutations of leukemic cells. Comparison of the LSC expression data to that of normal hematopoietic stem cells (HSC) revealed that 34% of the modulated genes are shared by both LSC and HSC, supporting the suggestion that the LSC originated within the HSC progenitors. We focused on the Notch pathway since Jagged-2, a Notch ligand was found to be overexpressed in the LSC samples. We show that DAPT, an inhibitor of gamma-secretase, a protease that is involved in Jagged and Notch signaling, inhibits LSC growth in colony formation assays. Identification of additional genes that regulate LSC self-renewal may provide new targets for therapy.

Cell Cycle↗

Antigen-chip technology for accessing global information about the state of the body.

Traditionally, immunologic diagnosis has been based on an attempt to correlate each disease with a specific immune reactivity, such as an antibody or a T-cell response to a single antigen specific for the disease entity. The state of the body, however, appears to be encoded by the immune system in collectives of reactivities and not by single reactivities. Here we describe our use of microarray technology and informatics to develop an antigen chip capable of detecting global patterns of antibodies binding to hundreds of antigens simultaneously. The patterns fashion diagnostic signatures.

Animals↗

Sorting points into neighborhoods (SPIN): data analysis and visualization by ordering distance matrices.

SUMMARY: We introduce a novel unsupervised approach for the organization and visualization of multidimensional data. At the heart of the method is a presentation of the full pairwise distance matrix of the data points, viewed in pseudocolor. The ordering of points is iteratively permuted in search of a linear ordering, which can be used to study embedded shapes. Several examples indicate how the shapes of certain structures in the data (elongated, circular and compact) manifest themselves visually in our permuted distance matrix. It is important to identify the elongated objects since they are often associated with a set of hidden variables, underlying continuous variation in the data. The problem of determining an optimal linear ordering is shown to be NP-Complete, and therefore an iterative search algorithm with O(n3) step-complexity is suggested. By using sorting points into neighborhoods, i.e. SPIN to analyze colon cancer expression data we were able to address the serious problem of sample heterogeneity, which hinders identification of metastasis related genes in our data. Our methodology brings to light the continuous variation of heterogeneity--starting with homogeneous tumor samples and gradually increasing the amount of another tissue. Ordering the samples according to their degree of contamination by unrelated tissue allows the separation of genes associated with irrelevant contamination from those related to cancer progression. AVAILABILITY: Software package will be available for academic users upon request.

Algorithms↗

Expression profiles of acute lymphoblastic and myeloblastic leukemias with ALL-1 rearrangements.

The ALL-1 gene is directly involved in 5-10% of acute lymphoblastic leukemias (ALLs) and acute myeloid leukemias (AMLs) by fusion to other genes or through internal rearrangements. DNA microarrays were used to determine expression profiles of ALLs and AMLs with ALL-1 rearrangements. These profiles distinguish those tumors from other ALLs and AMLs. The expression patterns of ALL-1-associated tumors, in particular ALLs, involve oncogenes, tumor suppressors, antiapoptotic genes, drug-resistance genes, etc., and correlate with the aggressive nature of the tumors. The genes whose expression differentiates between ALLs with and without ALL-1 rearrangement were further divided into several groups, enabling separation of ALL-1-associated ALLs into two subclasses. One of the groups included 43 genes that exhibited expression profiles closely linked to ALLs with ALL-1 rearrangements. Further, there were evident differences between the expression profiles of AMLs in which ALL-1 had undergone fusion to other genes and AMLs with partial duplication of ALL-1. The extensive analysis described here pinpointed genes that might have a direct role in pathogenesis.

Chromosome Aberrations↗

Statistical properties of contact vectors.

We study the statistical properties of contact vectors, a construct to characterize a protein's structure. The contact vector of an N-residue protein is a list of N integers n(i), representing the number of residues in contact with residue i. We study analytically (at mean-field level) and numerically the amount of structural information contained in a contact vector. Analytical calculations reveal that a large variance in the contact numbers reduces the degeneracy of the mapping between contact vectors and structures. Exact enumeration for lengths up to N=16 on the three-dimensional cubic lattice indicates that the growth rate of number of contact vectors as a function of N is only 3% less than that for contact maps. In particular, for compact structures we present numerical evidence that, practically, each contact vector corresponds to only a handful of structures. We discuss how this information can be used for better structure prediction.

Computational Biology↗

Computational capacity of an odorant discriminator: the linear separability of curves.

We introduce and study an artificial neural network inspired by the probabilistic receptor affinity distribution model of olfaction. Our system consists of N sensory neurons whose outputs converge on a single processing linear threshold element. The system's aim is to model discrimination of a single target odorant from a large number p of background odorants within a range of odorant concentrations. We show that this is possible provided p does not exceed a critical value p(c) and calculate the critical capacity alpha(c) = p(c)/N. The critical capacity depends on the range of concentrations in which the discrimination is to be accomplished. If the olfactory bulb may be thought of as a collection of such processing elements, each responsible for the discrimination of a single odorant, our study provides a quantitative analysis of the potential computational properties of the olfactory bulb. The mathematical formulation of the problem we consider is one of determining the capacity for linear separability of continuous curves, embedded in a large-dimensional space. This is accomplished here by a numerical study, using a method that signals whether the discrimination task is realizable, together with a finite-size scaling analysis.

Algorithms↗

DNA microarrays identification of primary and secondary target genes regulated by p53.

The transcriptional program regulated by the tumor suppressor p53 was analysed using oligonucleotide microarrays. A human lung cancer cell line that expresses the temperature sensitive murine p53 was utilized to quantitate mRNA levels of various genes at different time points after shifting the temperature to 32 degrees C. Inhibition of protein synthesis by cycloheximide (CHX) was used to distinguish between primary and secondary target genes regulated by p53. In the absence of CHX, 259 and 125 genes were up or down-regulated respectively; only 38 and 24 of these genes were up and down-regulated by p53 also in the presence of CHX and are considered primary targets in this cell line. Cluster analysis of these data using the super paramagnetic clustering (SPC) algorithm demonstrate that the primary genes can be distinguished as a single cluster among a large pool of p53 regulated genes. This procedure identified additional genes that co-cluster with the primary targets and can also be classified as such genes. In addition to cell cycle (e.g. p21, TGF-beta, Cyclin E) and apoptosis (e.g. Fas, Bak, IAP) related genes, the primary targets of p53 include genes involved in many aspects of cell function, including cell adhesion (e.g. Thymosin, Smoothelin), signaling (e.g. H-Ras, Diacylglycerol kinase), transcription (e.g. ATF3, LISCH7), neuronal growth (e.g. Ninjurin, NSCL2) and DNA repair (e.g. BTG2, DDB2). The results suggest that p53 activates concerted opposing signals and exerts its effect through a diverse network of transcriptional changes that collectively alter the cell phenotype in response to stress.

Animals↗

Resampling method for unsupervised estimation of cluster validity.

We introduce a method for validation of results obtained by clustering analysis of data. The method is based on resampling the available data. A figure of merit that measures the stability of clustering solutions against resampling is introduced. Clusters that are stable against resampling give rise to local maxima of this figure of merit. This is presented first for a one-dimensional data set, for which an analytic approximation for the figure of merit is derived and compared with numerical measurements. Next, the applicability of the method is demonstrated for higher-dimensional data, including gene microarray expression data.

Cluster Analysis↗

Comparison of two optimization methods to derive energy parameters for protein folding: perceptron and Z score.

Two methods were proposed recently to derive energy parameters from known native protein conformations and corresponding sets of decoys. One is based on finding, by means of a perceptron learning scheme, energy parameters such that the native conformations have lower energies than the decoys. The second method maximizes the difference between the native energy and the average energy of the decoys, measured in terms of the width of the decoys' energy distribution (Z-score). Whereas the perceptron method is sensitive mainly to "outlier" (i.e., extremal) decoys, the Z-score optimization is governed by the high density regions in decoy-space. We compare the two methods by deriving contact energies for two very different sets of decoys: the first obtained for model lattice proteins and the second by threading. We find that the potentials derived by the two methods are of similar quality and fairly closely related. This finding indicates that standard, naturally occurring sets of decoys are distributed in a way that yields robust energy parameters (that are quite insensitive to the particular method used to derive them). The main practical implication of this finding is that it is not necessary to fine-tune the potential search method to the particular set of decoys used.

Algorithms↗

Coupled two-way clustering analysis of gene microarray data.

We present a coupled two-way clustering approach to gene microarray data analysis. The main idea is to identify subsets of the genes and samples, such that when one of these is used to cluster the other, stable and significant partitions emerge. The search for such subsets is a computationally complex task. We present an algorithm, based on iterative clustering, that performs such a search. This analysis is especially suitable for gene microarray data, where the contributions of a variety of biological mechanisms to the gene expression levels are entangled in a large body of experimental data. The method was applied to two gene microarray data sets, on colon cancer and leukemia. By identifying relevant subsets of the data and focusing on them we were able to discover partitions and correlations that were masked and hidden when the full dataset was used in the analysis. Some of these partitions have clear biological interpretation; others can serve to identify possible directions for future research.

Cluster Analysis↗

Toward an energy function for the contact map representation of proteins.

We analyzed several energy functions for predicting the native state of proteins from an energy minimization procedure. We derived the parameters of a given energy function by imposing the basic requirement that the energy of the native conformation of a protein is lower than that of any conformation chosen from a set of decoys. Our work is motivated by a recent result which proved that the simple pairwise contact approximation of the energy is insufficient to satisfy simultaneously such a basic requirement for all the proteins in a database. Here, we investigate the reasons of such negative results and show how to improve the predictive power of methods based on energy minimization. We generated decoys by gapless threading, and we derive energy parameters by perceptron learning. We first considered hydrophobic contributions to the energy, defined in several ways, and showed that the additional hydrophobic terms enlarge slightly the number of proteins that can be stabilized together. Next, we performed various modifications of the pairwise energy term. We introduced (1) a distinction between inter-residue contacts on the surface and in the core of a protein and (2) a simple distance-dependent pairwise interaction in which a two-tier definition of contact replaces the original (single-tier) one. Our results suggest that a detailed treatment of the pairwise potential is likely to be more relevant than the consideration of other forces.

Algorithms↗

Can a pairwise contact potential stabilize native protein folds against decoys obtained by threading?

We present a method to derive contact energy parameters from large sets of proteins. The basic requirement on which our method is based is that for each protein in the database the native contact map has lower energy than all its decoy conformations that are obtained by threading. Only when this condition is satisfied one can use the proposed energy function for fold identification. Such a set of parameters can be found (by perceptron learning) if Mp, the number of proteins in the database, is not too large. Other aspects that influence the existence of such a solution are the exact definition of contact and the value of the critical distance Rc, below which two residues are considered to be in contact. Another important novel feature of our approach is its ability to determine whether an energy function of some suitable proposed form can or cannot be parameterized in a way that satisfies our basic requirement. As a demonstration of this, we determine the region in the (Rc, Mp) plane in which the problem is solvable, i.e., we can find a set of contact parameters that stabilize simultaneously all the native conformations. We show that for large enough databases the contact approximation to the energy cannot stabilize all the native folds even against the decoys obtained by gapless threading.

Algorithms↗

Protein folding using contact maps.

We discuss the problem of representations of protein structure and give the definition of contact maps. We present a method to obtain a three-dimensional polypeptide conformation from a contact map. We also explain how to deal with the case of nonphysical contact maps. We describe a stochastic method to perform dynamics in contact map space. We explain how the motion is restricted to physical regions of the space. First, we introduce the exact free energy of a contact map and discuss two simple approximations to it. Second, we present a method to derive energy parameters based on perception learning. We prove in an extensive number of situations that the pairwise contact approximation both when alone and when supplemented with a hydrophobic term is unsuitable for stabilizing proteins' native states.

Aprotinin↗

Folding Lennard-Jones proteins by a contact potential.

We studied the possibility to approximate a Lennard-Jones interaction by a pairwise contact potential. First we used a Lennard-Jones potential to design off-lattice, protein-like heteropolymer sequences, whose lowest energy (native) conformations were then identified by molecular dynamics. Then we turned to investigate whether one can find a pairwise contact potential, whose ground states are the contact maps associated with these native conformations. We show that such a requirement cannot be satisfied exactly, i.e., no such contact parameters exist. Nevertheless, we found that one can find contact energy parameters for which an energy minimization procedure, acting in the space of contact maps, yields maps whose corresponding structures are close to the native ones. Finally, we show that when these structures are used as the initial point of a molecular dynamics energy minimization process, the correct native folds are recovered with high probability.

Drug Design↗

Efficient dynamics in the space of contact maps.

BACKGROUND: Two problems are of major importance in protein fold prediction: how to generate plausible conformations, and how to choose an energy function to identify the native state. Contact maps are a simple representation of protein structure and offer a promising framework to address these two issues. RESULTS: In this work we develop Monte Carlo dynamics in contact map space. The procedure is divided into four steps: non-local dynamics, in which large-scale "cluster" moves are performed (clusters are in approximate correspondence with secondary structure elements); local dynamics, in which secondary structure location is optimized; reconstruction, in which the physicality of the contact map is restored; and refinement, which consists of a further Monte Carlo energy minimization in real space. We demonstrate that such a dynamical procedure is effective in producing uncorrelated low-energy states. CONCLUSIONS: The procedure introduced in this paper very effectively generates a representative ensemble of conformations. We are able to show that existing sets of pairwise contact energy parameters are not suitable to single out the native state within this ensemble. The remaining outstanding issue in protein folding is to find an energy function that can discriminate the native state from decoys.

Models, Chemical↗

Recovery of protein structure from contact maps.

BACKGROUND: Prediction of a protein's structure from its amino acid sequence is a key issue in molecular biology. While dynamics, performed in the space of two-dimensional contact maps, eases the necessary conformational search, it may also lead to maps that do not correspond to any real three-dimensional structure. To remedy this, an efficient procedure is needed to reconstruct three-dimensional conformations from their contact maps. RESULTS: We present an efficient algorithm to recover the three-dimensional structure of a protein from its contact map representation. We show that when a physically realizable map is used as target, our method generates a structure whose contact map is essentially similar to the target. furthermore, the reconstructed and original structures are similar up to the resolution of the contact map representation. Next, we use nonphysical target maps, obtained by corrupting a physical one; in this case, our method essentially recovers the underlying physical map and structure. Hence, our algorithm will help to fold proteins, using dynamics in the space of contact maps. Finally, we investigate the manner in which the quality of the recovered structure degrades when the number of contacts is reduced. CONCLUSIONS: The procedure is capable of assigning quickly and reliably a three-dimensional structure to a given contact map. It is well suited for use in parallel with dynamics in contact map space to project a contact map onto its closest physically allowed structural counterpart.

Algorithms↗