Search PubMed⌕ Search

Biomedical subjects

Shankar Subramaniam

Publications and source records attributed to Shankar Subramaniam.

At least 37 records · Page 2Linked to original sources

Analysis of the major patterns of B cell gene expression changes in response to short-term stimulation with 33 single ligands.

We examined the major patterns of changes in gene expression in mouse splenic B cells in response to stimulation with 33 single ligands for 0.5, 1, 2, and 4 h. We found that ligands known to directly induce or costimulate proliferation, namely, anti-IgM (anti-Ig), anti-CD40 (CD40L), LPS, and, to a lesser extent, IL-4 and CpG-oligodeoxynucleotide (CpG), induced significant expression changes in a large number of genes. The remaining 28 single ligands produced changes in relatively few genes, even though they elicited measurable elevations in intracellular Ca(2+) and cAMP concentration and/or protein phosphorylation, including cytokines, chemokines, and other ligands that interact with G protein-coupled receptors. A detailed comparison of gene expression responses to anti-Ig, CD40L, LPS, IL-4, and CpG indicates that while many genes had similar temporal patterns of change in expression in response to these ligands, subsets of genes showed unique expression patterns in response to IL-4, anti-Ig, and CD40L.

Animals↗

Computational modeling reveals how interplay between components of a GTPase-cycle module regulates signal transduction.

Heterotrimeric G protein signaling is regulated by signaling modules composed of heterotrimeric G proteins, active G protein-coupled receptors (Rs), which activate G proteins, and GTPase-activating proteins (GAPs), which deactivate G proteins. We term these modules GTPase-cycle modules. The local concentrations of these proteins are spatially regulated between plasma membrane microdomains and between the plasma membrane and cytosol, but no data or models are available that quantitatively explain the effect of such regulation on signaling. We present a computational model of the GTPase-cycle module that predicts that the interplay of local G protein, R, and GAP concentrations gives rise to 16 distinct signaling regimes and numerous intermediate signaling phenomena. The regimes suggest alternative modes of the GTPase-cycle module that occur based on defined local concentrations of the component proteins. In one mode, signaling occurs while G protein and receptor are unclustered and GAP eliminates signaling; in another, G protein and receptor are clustered and GAP can rapidly modulate signaling but does not eliminate it. Experimental data from multiple GTPase-cycle modules is interpreted in light of these predictions. The latter mode explains previously paradoxical data in which GAP does not alter maximal current amplitude of G protein-activated ion channels, but hastens signaling. The predictions indicate how variations in local concentrations of the component proteins create GTPase-cycle modules with distinctive phenotypes. They provide a quantitative framework for investigating how regulation of local concentrations of components of the GTPase-cycle module affects signaling.

GTP Phosphohydrolases↗

Conserved sequence and structure association motifs in antibody-protein and antibody-hapten complexes.

In this paper, we present the association requirements across a wide variety of antibody-antigen complexes. Phylogenetic analysis clearly indicates the representative nature of our structural dataset. Antigen molecules range from small-molecule haptens to complete protein structures. Common association motifs identified include five conserved tyrosine residues and a single conserved arginine residue from CDR-H3. Further, specificity is refined by a diverse array of antibody-antigen electrostatic interactions that maximize complex specificity. Through analysis of calculated pKa shifts on antigen binding, we find that these interactions are conserved at 23 alignment 'hot-spot' positions. Despite consistent roles in defining substrate specificity, 16 hot-spot positions are conserved less than 50% of the time. On the other hand, because of the conserved functional role of these positions, mutant screening at hot-spots is more likely to result in increased antigen specificity than elsewhere. Therefore, we believe these results should facilitate subsequent antibody design experimentation.

Amino Acid Sequence↗

MITOPRED: a web server for the prediction of mitochondrial proteins.

MITOPRED web server enables prediction of nucleus-encoded mitochondrial proteins in all eukaryotic species. Predictions are made using a new algorithm based primarily on Pfam domain occurrence patterns in mitochondrial and non-mitochondrial locations. Pre-calculated predictions are instantly accessible for proteomes of Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila, Homo sapiens, Mus musculus and Arabidopsis species as well as all the eukaryotic sequences in the Swiss-Prot and TrEMBL databases. Queries, at different confidence levels, can be made through four distinct options: (i) entering Swiss-Prot/TrEMBL accession numbers; (ii) uploading a local file with such accession numbers; (iii) entering protein sequences; (iv) uploading a local file containing protein sequences in FASTA format. Automated updates are scheduled for the pre-calculated prediction database so as to provide access to the most current data. The server, its documentation and the data are available from http://mitopred.sdsc.edu.

Algorithms↗

SledgeHMMER: a web server for batch searching the Pfam database.

The SledgeHMMER web server is intended for genome-scale searching of the Pfam database without having to install this database and the HMMER software locally. The server implements a parallelized version of hmmpfam, the program used for searching the Pfam HMM database. Pfam search results have been calculated for the entire Swiss-Prot and TrEmbl database sequences (approximately 1.2 million) on 256 processors of IA64-based teragrid machines. The Pfam database can be searched in local, glocal or merged mode, using either gathering or E-value thresholds. Query sequences are first matched against the pre-calculated entries to retrieve results, and those without matches are processed through a new search process. Results are emailed in a space-delimited tabular format upon completion of the search. While most other Pfam-searching web servers set a limit of one sequence per query, this server processes batch sequences with no limit on the number of input sequences. The web server and downloadable data are accessible from http://SledgeHmmer.sdsc.edu.

Algorithms↗

Decoding cilia function: defining specialized genes required for compartmentalized cilia biogenesis.

The evolution of the ancestral eukaryotic flagellum is an example of a cellular organelle that became dispensable in some modern eukaryotes while remaining an essential motile and sensory apparatus in others. To help define the repertoire of specialized proteins needed for the formation and function of cilia, we used comparative genomics to analyze the genomes of organisms with prototypical cilia, modified cilia, or no cilia and identified approximately 200 genes that are absent in the genomes of nonciliated eukaryotes but are conserved in ciliated organisms. Importantly, over 80% of the known ancestral proteins involved in cilia function are included in this small collection. Using Drosophila as a model system, we then characterized a novel family of proteins (OSEGs: outer segment) essential for ciliogenesis. We show that osegs encode components of a specialized transport pathway unique to the cilia compartment and are related to prototypical intracellular transport proteins.

Animals↗

MITOPRED: a genome-scale method for prediction of nucleus-encoded mitochondrial proteins.

MOTIVATION: Currently available methods for the prediction of subcellular location of mitochondrial proteins rely largely on the presence of mitochondrial targeting signals in the protein sequences. However, a large fraction of mitochondrial proteins lack such signals, making those tools ineffective for genome-scale prediction of mitochondria-targeted proteins. Here, we propose a method for genome-scale prediction of nucleus-encoded mitochondrial proteins. The new method, MITOPRED, is based on the Pfam domain occurrence patterns and the amino acid compositional differences between mitochondrial and non-mitochondrial proteins. RESULTS: MITOPRED could predict mitochondrial proteins with 100% specificity at a 44% sensitivity rate and with 67% specificity at 99% sensitivity. Additionally, it was sufficiently robust to predict mitochondrial proteins across different eukaryotic species with similar accuracy. Based on Matthews correlation coefficient measure, the prediction performance of MITOPRED is clearly superior (0.73) to those of the two popular methods TargetP (0.51) and PSORT (0.53). Using this method, we predicted the nucleus-encoded mitochondrial proteins from six complete genomes (three invertebrate, two vertebrate and one plant species) and estimated the total number in each genome. In human, our method estimated the existence of 1362 mitochondrial proteins corresponding to 4.8% of the total proteome. AVAILABILITY: MITOPRED program is freely accessible at http://mitopred.sdsc.edu. Source code is available on request from the authors. SUPPLEMENTARY INFORMATION: Training data sets are also available at http://mitopred.sdsc.edu

Algorithms↗

MitoProteome: mitochondrial protein sequence database and annotation system.

MitoProteome is an object-relational mitochondrial protein sequence database and annotation system. The initial release contains 847 human mitochondrial protein sequences, derived from public sequence databases and mass spectrometric analysis of highly purified human heart mitochondria. Each sequence is manually annotated with primary function, subfunction and subcellular location, and extensively annotated in an automated process with data extracted from external databases, including gene information from LocusLink and Ensembl; disease information from OMIM; protein-protein interaction data from MINT and DIP; functional domain information from Pfam; protein fingerprints from PRINTS; protein family and family-specific signatures from InterPro; structure data from PDB; mutation data from PMD; BLAST homology data from NCBI NR; and proteins found to be related based on LocusLink and SWISS-PROT references and sequence and taxonomy data. By highly automating the processes of maintaining the MitoProteome Protein List and extracting relevant data from external databases, we are able to present a dynamic database, updated frequently to reflect changes in public resources. The MitoProteome database is publicly available at http://www. mitoproteome.org/. Users may browse and search MitoProteome, and access a complete compilation of data relevant to each protein of interest, cross-linked to external databases.

Computational Biology↗

Bioinformatics and cellular signaling.

The understanding of cellular function requires an integrated analysis of context-specific, spatiotemporal data from diverse sources. Recent advances in describing the genomic and proteomic 'parts list' of the cell and deciphering the interrelationship of these parts are described, including genome-wide location analysis, standards for microarray data analysis, and two-hybrid and mass spectrometry approaches. This information is being collected and curated in databases such as the Alliance for Cellular Signaling (AfCS) Molecule Pages, which will serve as vital tools for the reconstruction and analysis of cellular signaling networks.

Cell Physiological Phenomena↗

Conservation of electrostatic properties within enzyme families and superfamilies.

Electrostatic interactions play a key role in enzyme catalytic function. At long range, electrostatics steer the incoming ligand/substrate to the active site, and at short distances, electrostatics provide the specific local interactions for catalysis. In cases in which electrostatics determine enzyme function, orthologs should share the electrostatic properties to maintain function. Often, electrostatic potential maps are employed to depict how conserved surface electrostatics preserve function. We expand on previous efforts to explain conservation of function, using novel electrostatic sequence and structure analyses of four enzyme families and one enzyme superfamily. We show that the spatial charge distribution is conserved within each family and superfamily. Conversely, phylogenetic analysis of key electrostatic residues provide the evolutionary origins of functionality.

Amino Acid Sequence↗

Protein fragment clustering and canonical local shapes.

A novel clustering method is used to cluster protein fragments by shape. The centroids (mean fragments from each cluster) form a basis set of structural motifs. A database of 156,643 seven-residue fragments is used, and eight different basis sets with varying levels of resolution are generated. Coarse basis sets contain tens of centroids and provide meaningful local shapes, which are more detailed than the traditional secondary structure categories. High-resolution basis sets contain thousands of centroids and can be used to model tertiary structure of longer segments. The basis sets generated fit nontraining set proteins with the expected accuracy.

Animals↗

Protein local structure prediction from sequence.

A basis set of protein canonical fragments, or centroids, represents the range of local structure found in globular proteins. We develop a methodology to predict centroids from the amino acid sequence. The predictor gives the probability of each centroid in the basis set, at each loci along the backbone. The predictor selects the best-fit centroid at about 40% of the loci. The predicted probabilities are accurate and can be used to judge the confidence of each centroid prediction. For example, when filtering out centroids with <0.50 probability, the predictor is 65% accurate, although such high-probability centroids occur at only 28% of the loci. Centroids with high probability can be interpreted as segments that are highly influenced by the amino acid sequence, whereas centroids with low probability can be interpreted as segments that are more likely influenced by tertiary contacts. Low-resolution, starting point structures, can be generated by fitting the predicted centroids together.

Animals↗

Neuroscience data and tool sharing: a legal and policy framework for neuroinformatics.

The requirements for neuroinformatics to make a significant impact on neuroscience are not simply technical--the hardware, software, and protocols for collaborative research--they also include the legal and policy frameworks within which projects operate. This is not least because the creation of large collaborative scientific databases amplifies the complicated interactions between proprietary, for-profit R&D and public "open science." In this paper, we draw on experiences from the field of genomics to examine some of the likely consequences of these interactions in neuroscience. Facilitating the widespread sharing of data and tools for neuroscientific research will accelerate the development of neuroinformatics. We propose approaches to overcome the cultural and legal barriers that have slowed these developments to date. We also draw on legal strategies employed by the Free Software community, in suggesting frameworks neuroinformatics might adopt to reinforce the role of public-science databases, and propose a mechanism for identifying and allowing "open science" uses for data whilst still permitting flexible licensing for secondary commercial research.

Computational Biology↗

Sequence-function analysis of the K+-selective family of ion channels using a comprehensive alignment and the KcsA channel structure.

Sequence-function analysis of K(+)-selective channels was carried out in the context of the 3.2 A crystal structure of a K(+) channel (KcsA) from Streptomyces lividans (Doyle et al., 1998). The first step was the construction of an alignment of a comprehensive set of K(+)-selective channel sequences forming the putative permeation path. This pathway consists of two transmembrane segments plus an extracellular linker. Included in the alignment are channels from the eight major classes of K(+)-selective channels from a wide variety of species, displaying varied rectification, gating, and activation properties. Segments of the alignment were assigned to structural motifs based on the KcsA structure. The alignment's accuracy was verified by two observations on these motifs: 1), the most variability is shown in the turret region, which functionally is strongly implicated in susceptibility to toxin binding; and 2), the selectivity filter and pore helix are the most highly conserved regions. This alignment combined with the KcsA structure was used to assess whether clusters of contiguous residues linked by hydrophobic or electrostatic interactions in KcsA are conserved in the K(+)-selective channel family. Analysis of sequence conservation patterns in the alignment suggests that a cluster of conserved residues is critical for determining the degree of K(+) selectivity. The alignment also supports the near-universality of the "glycine hinge" mechanism at the center of the inner helix for opening K channels. This mechanism has been suggested by the recent crystallization of a K channel in the open state. Further, the alignment reveals a second highly conserved glycine near the extracellular end of the inner helix, which may be important in minimizing deformation of the extracellular vestibule as the channel opens. These and other sequence-function relationships found in this analysis suggest that much of the permeation path architecture in KcsA is present in most K(+)-selective channels. Because of this finding, the alignment provides a robust starting point for homology modeling of the permeation paths of other K(+)-selective channel classes and elucidation of sequence-function relationships therein. To assay these applications, a homology model of the Shaker A channel permeation path was constructed using the alignment and KcsA as the template, and its structure evaluated in light of established structural criteria.

Bacterial Proteins↗

Overview of the Alliance for Cellular Signaling.

The Alliance for Cellular Signaling is a large-scale collaboration designed to answer global questions about signalling networks. Pathways will be studied intensively in two cells--B lymphocytes (the cells of the immune system) and cardiac myocytes--to facilitate quantitative modelling. One goal is to catalyse complementary research in individual laboratories; to facilitate this, all alliance data are freely available for use by the entire research community.

B-Lymphocytes↗

The Molecule Pages database.

The Alliance for Cellular Signaling (AfCS)-Nature Molecule Pages will be a comprehensive database of key facts about more than 3,000 proteins involved in cell signalling. Each entry will be created by invited experts and be peer-reviewed. Alongside the large-scale experiments being conducted by the AfCS scientists, the wealth of information contained in this database offers the potential of accelerating the pace of discovery in signal transduction research.

Automation↗

Natural coordinate representation for the protein backbone structure.

A new model for describing the geometry of the C(alpha) backbone atoms in protein molecules is derived. This model uses one continuous variable per amino acid. This is half the number of degrees-of-freedom used in traditional backbone models. The new model was tested on 721 PDB structures and its average accuracy was determined to be 1.14 A cRMSD. This model can be used as a description of local structure that provides higher resolution than the traditional secondary structure categories. Also, because this structure description is one-dimensional, it can be used to align structures with the same efficiency and convergence properties available in the popular sequence alignment tools. Furthermore, the 1:1 correspondence with the amino acid sequence has implications for combined sequence/structure alignment. Conventional secondary structure prediction was used to further reduce the number of degrees-of-freedom in 16 test proteins. In those cases, the average cRMSD degraded from 0.96 to 2.33 A while the number of degrees-of-freedom improved (reduced) by more than 30%.

Amino Acids↗

Bioinformatics of cellular signalling.

The completion of the human genome sequencing provides a unique opportunity to understand the complex functioning of cells in terms of myriad biochemical pathways. Of special significance are pathways involved in cellular signalling. Understanding how signal transduction occurs in cells is of paramount importance to medicine and pharmacology. The major steps involved in deciphering signalling pathways are: (a) identifying the molecules involved in signalling; (b) figuring out who talks to whom, i.e. deciphering molecular interactions in a context specific manner; (c) obtaining the spatiotemporal location of the signalling events; (d) reconstructing signalling modules and networks evoked in specific response to input; (e) correlating the signalling response to different cellular inputs; and (f) deciphering cross-talk between signalling modules in response to single and multiple inputs. High-throughput experimental investigations offer the promise of providing data pertaining to the above steps. A major challenge, then, is the organization of this data into knowledge in the form of hypothesis, models and context-specific understanding. The Alliance for Cellular Signaling (AfCS) is a multi-institution, multidisciplinary project and its primary objective is to utilize a multitude of high throughput approaches to obtain context-specific knowledge of cellular response to input. It is anticipated that the AfCS experimental data in combination with curated gene and protein annotations, available from public repositories, will serve as a basis for reconstruction of signalling networks. It will then be possible to model the networks mathematically to obtain quantitative measures of cellular response. In this paper we describe some of the bioinformatics strategies employed in the AfCS.

Animals↗