Search PubMed⌕ Search

Biomedical subjects

Shankar Subramaniam

Publications and source records attributed to Shankar Subramaniam.

At least 19 recordsLinked to original sources

Structure-centric searching enables global mapping of the public metabolome.

Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.

Journal Article↗

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article↗

Mitoproteome: human heart mitochondrial protein sequence database.

The human mitochondrial proteome database has been developed by deriving data from a combination of public repositories and experimental and computational prediction methods. The experimental data is derived from highly purified mitochondria from human heart tissue, whereas predictions have been performed by MITOPRED, a genome-scale method for the prediction of nucleus-encoded mitochondrial proteins. Mitochondrial protein sequences from different sources have been clustered to generate a nonredundant dataset. Annotations related to the protein function, structure, disease association, pathways, and so on are collected from a number of public databases using commonly used UNIX and Perl scripts. This chapter provides a detailed description of various data sources and methods used to download, curate, parse, and generate meaningful annotations from primary as well as derived databases.

Computational Biology↗

The alliance for cellular signaling plasmid collection: a flexible resource for protein localization studies and signaling pathway analysis.

Cellular responses to inputs that vary both temporally and spatially are determined by complex relationships between the components of cell signaling networks. Analysis of these relationships requires access to a wide range of experimental reagents and techniques, including the ability to express the protein components of the model cells in a variety of contexts. As part of the Alliance for Cellular Signaling, we developed a robust method for cloning large numbers of signaling ORFs into Gateway entry vectors, and we created a wide range of compatible expression platforms for proteomics applications. To date, we have generated over 3000 plasmids that are available to the scientific community via the American Type Culture Collection. We have established a website at www.signaling-gateway.org/data/plasmid/ that allows users to browse, search, and blast Alliance for Cellular Signaling plasmids. The collection primarily contains murine signaling ORFs with an emphasis on kinases and G protein signaling genes. Here we describe the cloning, databasing, and application of this proteomics resource for large scale subcellular localization screens in mammalian cell lines.

Animals↗

LMSD: LIPID MAPS structure database.

The LIPID MAPS Structure Database (LMSD) is a relational database encompassing structures and annotations of biologically relevant lipids. Structures of lipids in the database come from four sources: (i) LIPID MAPS Consortium's core laboratories and partners; (ii) lipids identified by LIPID MAPS experiments; (iii) computationally generated structures for appropriate lipid classes; (iv) biologically relevant lipids manually curated from LIPID BANK, LIPIDAT and other public sources. All the lipid structures in LMSD are drawn in a consistent fashion. In addition to a classification-based retrieval of lipids, users can search LMSD using either text-based or structure-based search options. The text-based search implementation supports data retrieval by any combination of these data fields: LIPID MAPS ID, systematic or common name, mass, formula, category, main class, and subclass data fields. The structure-based search, in conjunction with optional data fields, provides the capability to perform a substructure search or exact match for the structure drawn by the user. Search results, in addition to structure and annotations, also include relevant links to external databases. The LMSD is publicly available at www.lipidmaps.org/data/structure/.

Databases, Factual↗

Components of the antigen processing and presentation pathway revealed by gene expression microarray analysis following B cell antigen receptor (BCR) stimulation.

BACKGROUND: Activation of naïve B lymphocytes by extracellular ligands, e.g. antigen, lipopolysaccharide (LPS) and CD40 ligand, induces a combination of common and ligand-specific phenotypic changes through complex signal transduction pathways. For example, although all three of these ligands induce proliferation, only stimulation through the B cell antigen receptor (BCR) induces apoptosis in resting splenic B cells. In order to define the common and unique biological responses to ligand stimulation, we compared the gene expression changes induced in normal primary B cells by a panel of ligands using cDNA microarrays and a statistical approach, CLASSIFI (Cluster Assignment for Biological Inference), which identifies significant co-clustering of genes with similar Gene Ontology annotation. RESULTS: CLASSIFI analysis revealed an overrepresentation of genes involved in ion and vesicle transport, including multiple components of the proton pump, in the BCR-specific gene cluster, suggesting that activation of antigen processing and presentation pathways is a major biological response to antigen receptor stimulation. Proton pump components that were not included in the initial microarray data set were also upregulated in response to BCR stimulation in follow up experiments. MHC Class II expression was found to be maintained specifically in response to BCR stimulation. Furthermore, ligand-specific internalization of the BCR, a first step in B cell antigen processing and presentation, was demonstrated. CONCLUSION: These observations provide experimental validation of the computational approach implemented in CLASSIFI, demonstrating that CLASSIFI-based gene expression cluster analysis is an effective data mining tool to identify biological processes that correlate with the experimental conditional variables. Furthermore, this analysis has identified at least thirty-eight candidate components of the B cell antigen processing and presentation pathway and sets the stage for future studies focused on a better understanding of the components involved in and unique to B cell antigen processing and presentation.

Algorithms↗

Locally defined protein phylogenetic profiles reveal previously missed protein interactions and functional relationships.

Phylogenetic profiles encode patterns of presence or absence of genes across genomes, and these profiles can be used to assign functional relationships to nonhomologous pairs of proteins (Pellegrini et al., Proc Natl Acad Sci USA 1999;96:4284-4288). Although it is well known that many proteins were created from combinations of domains, most of the existing implementations of phylogenetic profiles do not consider this fact. Here, we introduce an extension that considers the multidomain nature of proteins and test the method against the known interaction data sets. Whereas earlier implementations associated one entire sequence with one protein phylogenetic profile (Single-Profile), our method instead breaks the sequence into a set of segments of predetermined size and constructs a separate profile for each segment (Multiple-Profile). The results show that the Multiple-Profile method performs as well as the Single-Profile method. However, the two methods share, surprisingly, a small fraction of their predictions, indicating that the Multiple-Profile method can detect known interactions missed by the Single-Profile method. Thus, the Multiple-Profile method can be used with other methods to determine functional relationships on a genome scale with wider coverage.

Binding Sites↗

Identification of signaling components required for the prediction of cytokine release in RAW 264.7 macrophages.

BACKGROUND: Release of immuno-regulatory cytokines and chemokines during inflammatory response is mediated by a complex signaling network. Multiple stimuli produce different signals that generate different cytokine responses. Current knowledge does not provide a complete picture of these signaling pathways. However, using specific markers of signaling pathways, such as signaling proteins, it is possible to develop a 'coarse-grained network' map that can help understand common regulatory modules for various cytokine responses and help differentiate between the causes of their release. RESULTS: Using a systematic profiling of signaling responses and cytokine release in RAW 264.7 macrophages made available by the Alliance for Cellular Signaling, an analysis strategy is presented that integrates principal component regression and exhaustive search-based model reduction to identify required signaling factors necessary and sufficient to predict the release of seven cytokines (G-CSF, IL-1alpha, IL-6, IL-10, MIP-1alpha, RANTES, and TNFalpha) in response to selected ligands. This study provides a model-based quantitative estimate of cytokine release and identifies ten signaling components involved in cytokine production. The models identified capture many of the known signaling pathways involved in cytokine release and predict potentially important novel signaling components, like p38 MAPK for G-CSF release, IFNgamma- and IL-4-specific pathways for IL-1a release, and an M-CSF-specific pathway for TNFalpha release. CONCLUSION: Using an integrative approach, we have identified the pathways responsible for the differential regulation of cytokine release in RAW 264.7 macrophages. Our results demonstrate the power of using heterogeneous cellular data to qualitatively and quantitatively map intermediate cellular phenotypes.

Analysis of Variance↗

Kdo2-Lipid A of Escherichia coli, a defined endotoxin that activates macrophages via TLR-4.

The LIPID MAPS Consortium (www.lipidmaps.org) is developing comprehensive procedures for identifying all lipids of the macrophage, following activation by endotoxin. The goal is to quantify temporal and spatial changes in lipids that occur with cellular metabolism and to develop bioinformatic approaches that establish dynamic lipid networks. To achieve these aims, an endotoxin of the highest possible analytical specification is crucial. We now report a large-scale preparation of 3-deoxy-D-manno-octulosonic acid (Kdo)(2)-Lipid A, a nearly homogeneous Re lipopolysaccharide (LPS) sub-structure with endotoxin activity equal to LPS. Kdo(2)-Lipid A was extracted from 2 kg cell paste of a heptose-deficient Escherichia coli mutant. It was purified by chromatography on silica, DEAE-cellulose, and C18 reverse-phase resin. Structure and purity were evaluated by electrospray ionization/mass spectrometry, liquid chromatography/mass spectrometry and (1)H-NMR. Its bioactivity was compared with LPS in RAW 264.7 cells and bone marrow macrophages from wild-type and toll-like receptor 4 (TLR-4)-deficient mice. Cytokine and eicosanoid production, in conjunction with gene expression profiling, were employed as readouts. Kdo(2)-Lipid A is comparable to LPS by these criteria. Its activity is reduced by >10(3) in cells from TLR-4-deficient mice. The purity of Kdo(2)-Lipid A should facilitate structural analysis of complexes with receptors like TLR-4/MD2.

Animals↗

LMPD: LIPID MAPS proteome database.

The LIPID MAPS Proteome Database (LMPD) is an object-relational database of lipid-associated protein sequences and annotations. The initial release contains 2959 records, representing human and mouse proteins involved in lipid metabolism. UniProt IDs were obtained based on keyword search of KEGG and GO databases, and this LMPD protein list was then enhanced with annotations from UniProt, EntrezGene, ENZYME, GO, KEGG and other public resources. We also assigned associations with general lipid categories, based on GO and KEGG annotations. Users may search LMPD by database ID or keyword, and filter by species and/or lipid class associations; from the search results, one can then access a compilation of data relevant to each protein of interest, cross-linked to external databases. The LIPID MAPS Proteome Database (LMPD) is publicly available from the LIPID MAPS Consortium website (http://www.lipidmaps.org/). The direct URL is http://www.lipidmaps.org/data/proteome/index.cgi.

Animals↗

Detecting conserved interaction patterns in biological networks.

Molecular interaction data plays an important role in understanding biological processes at a modular level by providing a framework for understanding cellular organization, functional hierarchy, and evolutionary conservation. As the quality and quantity of network and interaction data increases rapidly, the problem of effectively analyzing this data becomes significant. Graph theoretic formalisms, commonly used for these analysis tasks, often lead to computationally hard problems due to their relation to subgraph isomorphism. This paper presents an innovative new algorithm, MULE, for detecting frequently occurring patterns and modules in biological networks. Using an innovative graph simplification technique based on ortholog contraction, which is ideally suited to biological networks, our algorithm renders these problems computationally tractable and scalable to large numbers of networks. We show, experimentally, that our algorithm can extract frequently occurring patterns in metabolic pathways and protein interaction networks from the KEGG, DIP, and BIND databases within seconds. When compared to existing approaches, our graph simplification technique can be viewed either as a pruning heuristic, or a closely related, but computationally simpler task. When used as a pruning heuristic, we show that our technique reduces effective graph sizes significantly, accelerating existing techniques by several orders of magnitude! Indeed, for most of the test cases, existing techniques could not even be applied without our pruning step. When used as a stand-alone analysis technique, MULE is shown to convey significant biological insights at near-interactive rates. The software, sample input graphs, and detailed results for comprehensive analysis of nine eukaryotic PPI networks are available at www.cs.purdue.edu/homes/koyuturk/mule.

Algorithms↗

Pairwise alignment of protein interaction networks.

With an ever-increasing amount of available data on protein-protein interaction (PPI) networks and research revealing that these networks evolve at a modular level, discovery of conserved patterns in these networks becomes an important problem. Although available data on protein-protein interactions is currently limited, recently developed algorithms have been shown to convey novel biological insights through employment of elegant mathematical models. The main challenge in aligning PPI networks is to define a graph theoretical measure of similarity between graph structures that captures underlying biological phenomena accurately. In this respect, modeling of conservation and divergence of interactions, as well as the interpretation of resulting alignments, are important design parameters. In this paper, we develop a framework for comprehensive alignment of PPI networks, which is inspired by duplication/divergence models that focus on understanding the evolution of protein interactions. We propose a mathematical model that extends the concepts of match, mismatch, and gap in sequence alignment to that of match, mismatch, and duplication in network alignment and evaluates similarity between graph structures through a scoring function that accounts for evolutionary events. By relying on evolutionary models, the proposed framework facilitates interpretation of resulting alignments in terms of not only conservation but also divergence of modularity in PPI networks. Furthermore, as in the case of sequence alignment, our model allows flexibility in adjusting parameters to quantify underlying evolutionary relationships. Based on the proposed model, we formulate PPI network alignment as an optimization problem and present fast algorithms to solve this problem. Detailed experimental results from an implementation of the proposed framework show that our algorithm is able to discover conserved interaction patterns very effectively, in terms of both accuracies and computational cost.

Algorithms↗

Inferring functional information from domain co-evolution.

MOTIVATION: Co-evolution is a powerful mechanism for understanding protein function. Prior work in this area has shown that co-evolving proteins are more likely to share the same function than those that do not because of functional constraints. Many of the efforts founded on this observation, however, are at the level of entire sequences, implicitly assuming that the complete protein sequence follows a single evolutionary trajectory. Since it is well known that a domain can exist in various contexts, this assumption is not valid for numerous multi-domain proteins. Motivated by these observations, we introduce a novel technique called Coevolutionary-Matrix that captures co-evolution between regions of two proteins. Instead of using existing domain information, the method exploits residue-level conservation to identify co-evolving regions that might correspond to domains. RESULTS: We show that the Coevolutionary-Matrix method can detect greater number of known functional associations for the Escherichia coli proteins when compared with earlier implementations of phylogenetic profiles. Furthermore, co-evolving regions of proteins detected by our method enable us to make hypotheses about their specific functions, many of which are supported by existing biochemical studies.

Algorithms↗

Molecular determinants of crosstalk between nuclear receptors and toll-like receptors.

Nuclear receptors (NRs) repress transcriptional responses to diverse signaling pathways as an essential aspect of their biological activities, but mechanisms determining the specificity and functional consequences of transrepression remain poorly understood. Here, we report signal- and gene-specific repression of transcriptional responses initiated by engagement of toll-like receptors (TLR) 3, 4, and 9 in macrophages. The glucocorticoid receptor (GR) represses a large set of functionally related inflammatory response genes by disrupting p65/interferon regulatory factor (IRF) complexes required for TLR4- or TLR9-dependent, but not TLR3-dependent, transcriptional activation. This mechanism requires signaling through MyD88 and enables the GR to differentially regulate pathogen-specific programs of gene expression. PPARgamma and LXRs repress overlapping transcriptional targets by p65/IRF3-independent mechanisms and cooperate with the GR to synergistically transrepress distinct subsets of TLR-responsive genes. These findings reveal combinatorial control of homeostasis and immune responses by nuclear receptors and suggest new approaches for treatment of inflammatory diseases.

Animals↗

pTARGET [corrected] a new method for predicting protein subcellular localization in eukaryotes.

MOTIVATION: There is a scarcity of efficient computational methods for predicting protein subcellular localization in eukaryotes. Currently available methods are inadequate for genome-scale predictions with several limitations. Here, we present a new prediction method, pTARGET that can predict proteins targeted to nine different subcellular locations in the eukaryotic animal species. RESULTS: The nine subcellular locations predicted by pTARGET include cytoplasm, endoplasmic reticulum, extracellular/secretory, golgi, lysosomes, mitochondria, nucleus, plasma membrane and peroxisomes. Predictions are based on the location-specific protein functional domains and the amino acid compositional differences across different subcellular locations. Overall, this method can predict 68-87% of the true positives at accuracy rates of 96-99%. Comparison of the prediction performance against PSORT showed that pTARGET prediction rates are higher by 11-60% in 6 of the 8 locations tested. Besides, the pTARGET method is robust enough for genome-scale prediction of protein subcellular localizations since, it does not rely on the presence of signal or target peptides. AVAILABILITY: A public web server based on the pTARGET method is accessible at the URL http://bioinformatics.albany.edu/~ptarget. Datasets used for developing pTARGET can be downloaded from this web server. Source code will be available on request from the corresponding author.

Algorithms↗

VAMPIRE microarray suite: a web-based platform for the interpretation of gene expression data.

Microarrays are invaluable high-throughput tools used to snapshot the gene expression profiles of cells and tissues. Among the most basic and fundamental questions asked of microarray data is whether individual genes are significantly activated or repressed by a particular stimulus. We have previously presented two Bayesian statistical methods for this level of analysis, collectively known as variance-modeled posterior inference with regional exponentials (VAMPIRE). These methods each require a sophisticated modeling step followed by integration of a posterior probability density. We present here a publicly available, web-based platform that allows users to easily load data, associate related samples and identify differentially expressed features using the VAMPIRE statistical framework. In addition, this suite of tools seamlessly integrates a novel gene annotation tool, known as GOby, which identifies statistically overrepresented gene groups. Unlike other tools in this genre, GOby can localize enrichment while respecting the hierarchical structure of annotation systems like Gene Ontology (GO). By identifying statistically significant enrichment of GO terms, Kyoto Encyclopedia of Genes and Genomes pathways, and TRANSFAC transcription factor binding sites, users can gain substantial insight into the physiological significance of sets of differentially expressed genes. The VAMPIRE microarray suite can be accessed at http://genome.ucsd.edu/microarray.

Bayes Theorem↗

A comprehensive classification system for lipids.

Lipids are produced, transported, and recognized by the concerted actions of numerous enzymes, binding proteins, and receptors. A comprehensive analysis of lipid molecules, "lipidomics," in the context of genomics and proteomics is crucial to understanding cellular physiology and pathology; consequently, lipid biology has become a major research target of the postgenomic revolution and systems biology. To facilitate international communication about lipids, a comprehensive classification of lipids with a common platform that is compatible with informatics requirements has been developed to deal with the massive amounts of data that will be generated by our lipid community. As an initial step in this development, we divide lipids into eight categories (fatty acyls, glycerolipids, glycerophospholipids, sphingolipids, sterol lipids, prenol lipids, saccharolipids, and polyketides) containing distinct classes and subclasses of molecules, devise a common manner of representing the chemical structures of individual lipids and their derivatives, and provide a 12 digit identifier for each unique lipid molecule. The lipid classification scheme is chemically based and driven by the distinct hydrophobic and hydrophilic elements that compose the lipid. This structured vocabulary will facilitate the systematization of lipid biology and enable the cataloging of lipids and their properties in a way that is compatible with other macromolecular databases.

Database Management Systems↗

Reconstruction of cellular signalling networks and analysis of their properties.

The study of cellular signalling over the past 20 years and the advent of high-throughput technologies are enabling the reconstruction of large-scale signalling networks. After careful reconstruction of signalling networks, their properties must be described within an integrative framework that accounts for the complexity of the cellular signalling network and that is amenable to quantitative modelling.

Animals↗