Search PubMed⌕ Search

Biomedical subjects

M S Madhusudhan

Publications and source records attributed to M S Madhusudhan.

At least 19 recordsLinked to original sources

Gene expression profiling of the human maternal-fetal interface reveals dramatic changes between midgestation and term.

Human placentation entails the remarkable integration of fetal and maternal cells into a single functional unit. In the basal plate region (the maternal-fetal interface) of the placenta, fetal cytotrophoblasts from the placenta invade the uterus and remodel the resident vasculature and avoid maternal immune rejection. Knowing the molecular bases for these unique cell-cell interactions is important for understanding how this specialized region functions during normal pregnancy with implications for tumor biology and transplantation immunology. Therefore, we undertook a global analysis of the gene expression profiles at the maternal-fetal interface. Basal plate biopsy specimens were obtained from 36 placentas (14-40 wk) at the conclusion of normal pregnancies. RNA was isolated, processed, and hybridized to HG-U133A&B Affymetrix GeneChips. Surprisingly, there was little change in gene expression during the 14- to 24-wk interval. In contrast, 418 genes were differentially expressed at term (37-40 wk) as compared with midgestation (14-24 wk). Subsequent analyses using quantitative PCR and immunolocalization approaches validated a portion of these results. Many of the differentially expressed genes are known in other contexts to be involved in differentiation, motility, transcription, immunity, angiogenesis, extracellular matrix dissolution, or lipid metabolism. One sixth were nonannotated or encoded hypothetical proteins. Modeling based on structural homology revealed potential functions for 31 of these proteins. These data provide a reference set for understanding the molecular components of the dialogue taking place between maternal and fetal cells in the basal plate as well as for future comparisons of alterations in this region that occur in obstetric complications.

Female↗

Protein complex compositions predicted by structural similarity.

Proteins function through interactions with other molecules. Thus, the network of physical interactions among proteins is of great interest to both experimental and computational biologists. Here we present structure-based predictions of 3387 binary and 1234 higher order protein complexes in Saccharomyces cerevisiae involving 924 and 195 proteins, respectively. To generate candidate complexes, comparative models of individual proteins were built and combined together using complexes of known structure as templates. These candidate complexes were then assessed using a statistical potential, derived from binary domain interfaces in PIBASE (http://salilab.org/pibase). The statistical potential discriminated a benchmark set of 100 interface structures from a set of sequence-randomized negative examples with a false positive rate of 3% and a true positive rate of 97%. Moreover, the predicted complexes were also filtered using functional annotation and sub-cellular localization data. The ability of the method to select the correct binding mode among alternates is demonstrated for three camelid VHH domain-porcine alpha-amylase interactions. We also highlight the prediction of co-complexed domain superfamilies that are not present in template complexes. Through integration with MODBASE, the application of the method to proteomes that are less well characterized than that of S.cerevisiae will contribute to expansion of the structural and functional coverage of protein interaction space. The predicted complexes are deposited in MODBASE (http://salilab.org/modbase).

Algorithms↗

Variable gap penalty for protein sequence-structure alignment.

The penalty for inserting gaps into an alignment between two protein sequences is a major determinant of the alignment accuracy. Here, we present an algorithm for finding a globally optimal alignment by dynamic programming that can use a variable gap penalty (VGP) function of any form. We also describe a specific function that depends on the structural context of an insertion or deletion. It penalizes gaps that are introduced within regions of regular secondary structure, buried regions, straight segments and also between two spatially distant residues. The parameters of the penalty function were optimized on a set of 240 sequence pairs of known structure, spanning the sequence identity range of 20-40%. We then tested the algorithm on another set of 238 sequence pairs of known structures. The use of the VGP function increases the number of correctly aligned residues from 81.0 to 84.5% in comparison with the optimized affine gap penalty function; this difference is statistically significant according to Student's t-test. We estimate that the new algorithm allows us to produce comparative models with an additional approximately 7 million accurately modeled residues in the approximately 1.1 million proteins that are detectably related to a known structure.

Algorithms↗

MODBASE: a database of annotated comparative protein structure models and associated resources.

MODBASE (http://salilab.org/modbase) is a database of annotated comparative protein structure models for all available protein sequences that can be matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on MODELLER for fold assignment, sequence-structure alignment, model building and model assessment (http:/salilab.org/modeller). MODBASE is updated regularly to reflect the growth in protein sequence and structure databases, and improvements in the software for calculating the models. MODBASE currently contains 3 094 524 reliable models for domains in 1 094 750 out of 1 817 889 unique protein sequences in the UniProt database (July 5, 2005); only models based on statistically significant alignments and models assessed to have the correct fold despite insignificant alignments are included. MODBASE also allows users to generate comparative models for proteins of interest with the automated modeling server MODWEB (http://salilab.org/modweb). Our other resources integrated with MODBASE include comprehensive databases of multiple protein structure alignments (DBAli, http://salilab.org/dbali), structurally defined ligand binding sites and structurally defined binary domain interfaces (PIBASE, http://salilab.org/pibase) as well as predictions of ligand binding sites, interactions between yeast proteins, and functional consequences of human nsSNPs (LS-SNP, http://salilab.org/LS-SNP).

Binding Sites↗

MODBASE, a database of annotated comparative protein structure models, and associated resources.

MODBASE (http://salilab.org/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on the MODELLER package for fold assignment, sequence-structure alignment, model building and model assessment (http:/salilab.org/modeller). MODBASE uses the MySQL relational database management system for flexible querying and CHIMERA for viewing the sequences and structures (http://www.cgl.ucsf.edu/chimera/). MODBASE is updated regularly to reflect the growth in protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different data sets. The largest data set contains 1,26,629 models for domains in 659,495 out of 1,182,126 unique protein sequences in the complete Swiss-Prot/TrEMBL database (August 25, 2003); only models based on alignments with significant similarity scores and models assessed to have the correct fold despite insignificant alignments are included. Another model data set supports target selection and structure-based annotation by the New York Structural Genomics Research Consortium; e.g. the 53 new structures produced by the consortium allowed us to characterize structurally 24,113 sequences. MODBASE also contains binding site predictions for small ligands and a set of predicted interactions between pairs of modeled sequences from the same genome. Our other resources associated with MODBASE include a comprehensive database of multiple protein structure alignments (DBALI, http://salilab.org/dbali) as well as web servers for automated comparative modeling with MODPIPE (MODWEB, http://salilab. org/modweb), modeling of loops in protein structures (MODLOOP, http://salilab.org/modloop) and predicting functional consequences of single nucleotide polymorphisms (SNPWEB, http://salilab. org/snpweb).

Amino Acid Sequence↗

Alignment of protein sequences by their profiles.

The accuracy of an alignment between two protein sequences can be improved by including other detectably related sequences in the comparison. We optimize and benchmark such an approach that relies on aligning two multiple sequence alignments, each one including one of the two protein sequences. Thirteen different protocols for creating and comparing profiles corresponding to the multiple sequence alignments are implemented in the SALIGN command of MODELLER. A test set of 200 pairwise, structure-based alignments with sequence identities below 40% is used to benchmark the 13 protocols as well as a number of previously described sequence alignment methods, including heuristic pairwise sequence alignment by BLAST, pairwise sequence alignment by global dynamic programming with an affine gap penalty function by the ALIGN command of MODELLER, sequence-profile alignment by PSI-BLAST, Hidden Markov Model methods implemented in SAM and LOBSTER, pairwise sequence alignment relying on predicted local structure by SEA, and multiple sequence alignment by CLUSTALW and COMPASS. The alignment accuracies of the best new protocols were significantly better than those of the other tested methods. For example, the fraction of the correctly aligned residues relative to the structure-based alignment by the best protocol is 56%, which can be compared with the accuracies of 26%, 42%, 43%, 48%, 50%, 49%, 43%, and 43% for the other methods, respectively. The new method is currently applied to large-scale comparative protein structure modeling of all known sequences.

Algorithms↗

Tools for comparative protein structure modeling and analysis.

The following resources for comparative protein structure modeling and analysis are described (http://salilab.org): MODELLER, a program for comparative modeling by satisfaction of spatial restraints; MODWEB, a web server for automated comparative modeling that relies on PSI-BLAST, IMPALA and MODELLER; MODLOOP, a web server for automated loop modeling that relies on MODELLER; MOULDER, a CPU intensive protocol of MODWEB for building comparative models based on distant known structures; MODBASE, a comprehensive database of annotated comparative models for all sequences detectably related to a known structure; MODVIEW, a Netscape plugin for Linux that integrates viewing of multiple sequences and structures; and SNPWEB, a web server for structure-based prediction of the functional impact of a single amino acid substitution.

Internet↗

RasGRP4, a new mast cell-restricted Ras guanine nucleotide-releasing protein with calcium- and diacylglycerol-binding motifs. Identification of defective variants of this signaling protein in asthma, mastocytosis, and mast cell leukemia patients and demonstration of the importance of RasGRP4 in mast cell development and function.

A cDNA was isolated from interleukin 3-developed, mouse bone marrow-derived mast cells (MCs) that contained an insert (designated mRasGRP4) that had not been identified in any species at the gene, mRNA, or protein level. By using a homology-based cloning approach, the approximately 2.6-kb hRasGRP4 transcript was also isolated from the mononuclear progenitors residing in the peripheral blood of normal individuals. This transcript information was then used to locate the RasGRP4 gene in the mouse and human genomes, to deduce its exon/intron organization, and then to identify 10 single nucleotide polymorphisms in the human gene that result in 5 amino acid differences. The >15-kb hRasGRP4 gene consists of 18 exons and resides on a region of chromosome 19q13.1 that had not been sequenced by the Human Genome Project. Human and mouse MCs and their progenitors selectively express RasGRP4, and this new intracellular protein contains all of the domains present in the RasGRP family of guanine nucleotide exchange factors even though it is <50% identical to its closest homolog. Recombinant RasGRP4 can activate H-Ras in a cation-dependent manner. Transfection experiments also suggest that RasGRP4 is a diacylglycerol/phorbol ester receptor. Transcript analysis of an asthma patient, a mastocytosis patient, and the HMC-1 cell line derived from a MC leukemia patient revealed the presence of substantial amounts of non-functional forms of hRasGRP4 due to an inability to remove intron 5 in the precursor transcript. Because only abnormal forms of hRasGRP4 were identified in the HMC-1 cell line, this immature MC progenitor was used to address the function of RasGRP4 in MCs. HMC-1 leukemia cells differentiated and underwent granule maturation when induced to express a normal form of RasGRP4. Thus, RasGRP4 plays an important role in the final stages of MC development.

Amino Acid Sequence↗

Reliability of assessment of protein structure prediction methods.

The reliability of ranking of protein structure modeling methods is assessed. The assessment is based on the parametric Student's t test and the nonparametric Wilcox signed rank test of statistical significance of the difference between paired samples. The approach is applied to the ranking of the comparative modeling methods tested at the fourth meeting on Critical Assessment of Techniques for Protein Structure Prediction (CASP). It is shown that the 14 CASP4 test sequences may not be sufficient to reliably distinguish between the top eight methods, given the model quality differences and their standard deviations. We suggest that CASP needs to be supplemented by an assessment of protein structure prediction methods that is automated, continuous in time, based on several criteria applied to a large number of models, and with quantitative statistical reliability assigned to each characterization.

Computer Simulation↗

Human tryptase epsilon (PRSS22), a new member of the chromosome 16p13.3 family of human serine proteases expressed in airway epithelial cells.

Probing of the GenBank expressed sequence tag (EST) data base with varied human tryptase cDNAs identified two truncated ESTs that subsequently were found to encode overlapping portions of a novel human serine protease (designated tryptase epsilon or protease, serine S1 family member 22 (PRSS22)). The tryptase epsilon gene resides on chromosome 16p13.3 within a 2.5-Mb complex of serine protease genes. Although at least 7 of the 14 genes in this complex encode enzymatically active proteases, only one tryptase epsilon-like gene was identified. The trachea and esophagus were found to contain the highest steady-state levels of the tryptase epsilon transcript in adult humans. Although the tryptase epsilon transcript was scarce in adult human lung, it was present in abundance in fetal lung. Thus, the tryptase epsilon gene is expressed in the airways in a developmentally regulated manner that is different from that of other human tryptase genes. At the cellular level, tryptase epsilon is a major product of normal pulmonary epithelial cells, as well as varied transformed epithelial cell lines. Enzymatically active tryptase epsilon is also constitutively secreted from these cells. The amino acid sequence of human tryptase epsilon is 38-44% identical to those of human tryptase alpha, tryptase beta I, tryptase beta II, tryptase beta III, transmembrane tryptase/tryptase gamma, marapsin, and Esp-1/testisin. Nevertheless, comparative protein structure modeling and functional studies using recombinant material revealed that tryptase epsilon has a substrate preference distinct from that of its other family members. These data indicate that the products of the chromosome 16p13.3 complex of tryptase genes evolved to carry out varied functions in humans.

Adult↗

Computer modeling and molecular dynamics simulations of ligand bound complexes of bovine angiogenin: dinucleotide topology at the active site of RNase a family proteins.

We have undertaken the modeling of substrate-bound structures of angiogenin. In our recent study, we modeled the dinucleotide ligand binding to human angiogenin. In the present study, the substrates CpG, UpG, and CpA were docked onto bovine angiogenin. This was achieved by overcoming the problem of an obstruction to the B1 site by the C-terminus and identifying residues that bind to the second base. The modeled complexes retain biochemically important interactions. The docked models were subjected to 1 ns of molecular dynamics, and structures from the simulation were refined by using simulated annealing. Our models explained the enzyme's specificity for both B1 and B2 bases as observed experimentally. The nature of binding of the dinucleotide substrate was compared with that of the mononucleotide product. The models of these complexes were also compared with those obtained earlier with human angiogenin. On the basis of the simulations and annealed structures, we came up with a consensus topology of dinucleotide ligands that binds to human and bovine angiogenins. This dinucleotide conformation can serve as a starting model for ligand-bound complex structures for RNase A family of proteins. We demonstrated this capability by generating the complex structure of CpA bound to eosinophil-derived neurotoxin (EDN) by fitting the consensus topology of CpA to the crystal structure of native EDN.

Animals↗

Tryptase 4, a new member of the chromosome 17 family of mouse serine proteases.

Genomic blot analysis raised the possibility that uncharacterized tryptase genes reside on chromosome 17 at the complex containing the three genes that encode mouse mast cell protease (mMCP) 6, mMCP-7, and transmembrane tryptase (mTMT). Probing of GenBank's expressed sequence tag data base with these three tryptase cDNAs resulted in the identification of an expressed sequence tag that encodes a portion of a novel mouse serine protease (now designated mouse tryptase 4 (mT4) because it is the fourth member of this family). 5'- and 3'-rapid amplification of cDNA ends approaches were carried out to deduce the nucleotide sequence of the full-length mT4 transcript. This information was then used to clone its approximately 5.0-kilobase pair gene. Chromosome mapping analysis of its gene, sequence analysis of its transcript, and comparative protein structure modeling of its translated product revealed that mT4 is a new member of the chromosome 17 family of mouse tryptases. mT4 is 40-44% identical to mMCP-6, mMCP-7, and mTMT, and this new serine protease has all of the structural features of a functional tryptase. Moreover, mT4 is enzymatically active when expressed in insect cells. Due to its 17-mer hydrophobic domain at its C terminus, mT4 is a membrane-anchored tryptase more analogous to mTMT than the other members of its family. As assessed by RNA blot, reverse transcriptase-polymerase chain reaction, and/or in situ hybridization analysis, mT4 is expressed in interleukin-5-dependent mouse eosinophils, as well as in ovaries and testes. The observation that recombinant mT4 is preferentially retained in the endoplasmic reticulum of transiently transfected COS-7 cells suggests a convertase-like role for this integral membrane serine protease.

Amino Acid Sequence↗

Short-strong hydrogen bonds and a low barrier transition state for the proton transfer reaction in RNase A catalysis: a quantum chemical study.

There is growing evidence that some enzymes catalyze reactions through the formation of short-strong hydrogen bonds as first suggested by Gerlt and Gassman. Support comes from several experimental and quantum chemical studies that include correlation energies on model systems. In the present study, the process of proton transfer between hydroxyl and imidazole groups, a model of the crucial step in the hydrolysis of RNA by the enzymes of the RNase A family, is investigated at the quantum mechanical level of density functional theory and perturbation theory at the MP2 level. The model focuses on the nature of the formation of a complex between the important residues of the protein and the hydroxyl group of the substrate. We have also investigated different configurations of the ground state that are important in the proton transfer reaction. The nature of bonding between the catalytic unit of the enzyme and the substrate in the model is investigated by Bader's atoms in molecule theory. The contributions of solvation and vibrational energies corresponding to the reactant, the transition state and the product configurations are also evaluated. Furthermore, the effect of protein environment is investigated by considering the catalytic unit surrounded by complete proteins--RNase A and Angiogenin. The results, in general, indicate the formation of a short-strong hydrogen bond and the formation of a low barrier transition state for the proton transfer model of the enzyme.

Catalysis↗

Computer modeling of human angiogenin-dinucleotide substrate interaction.

Structures of substrate bound human angiogenin complexes have been obtained for the first time by computer modeling. The dinucleotides CpA and UpA have been docked onto human angiogenin using a systematic grid search procedure in torsion and Eulerian angle space. The docking was guided throughout by the similarity of angiogenin-substrate interactions with interactions of RNase A and its substrate. The models were subjected to 1 nanosecond of molecular dynamics to access their stability. Structures extracted from MD simulations were refined by simulated annealing. Stable hydrogen bonds that bridged protein and ligand residues during the MD simulations were taken as restraints for simulated annealing. Our analysis on the MD structures and annealed models explains the substrate specificity of human angiogenin and is in agreement with experimental results. This study also predicts the B2 binding site residues of angiogenin, for which no experimental information is available so far. In the case of one of the substrates, CpA, we have also identified the presence of a water molecule that invariantly bridges the B2 base with the protein. We have compared our results to the RNase A-substrate complex and highlight the similarities and differences.

Amino Acid Sequence↗

Deducing hydration sites of a protein from molecular dynamics simulations.

Invariant water molecules that are of structural or functional importance to proteins are detected from their presence in the same location in different crystal structures of the same protein or closely related proteins. In this study we have investigated the location of invariant water molecules from MD simulations of ribonuclease A, HIV1-protease and Hen egg white lysozyme. Snapshots of MD trajectories represent the structure of a dynamic protein molecule in a solvated environment as opposed to the static picture provided by crystallography. The MD results are compared to an analysis on crystal structures. A good correlation is observed between the two methods with more than half the hydration sites identified as invariant from crystal structures featuring as invariant in the MD simulations which include most of the functionally or structurally important residues. It is also seen that the propensities of occupying the various hydration sites on a protein for structures obtained from MD and crystallographic studies are different. In general MD simulations can be used to predict invariant hydration sites when there is a paucity of crystallographic data or to complement crystallographic results.

Animals↗

EVA: continuous automatic evaluation of protein structure prediction servers.

UNLABELLED: Evaluation of protein structure prediction methods is difficult and time-consuming. Here, we describe EVA, a web server for assessing protein structure prediction methods, in an automated, continuous and large-scale fashion. Currently, EVA evaluates the performance of a variety of prediction methods available through the internet. Every week, the sequences of the latest experimentally determined protein structures are sent to prediction servers, results are collected, performance is evaluated, and a summary is published on the web. EVA has so far collected data for more than 3000 protein chains. These results may provide valuable insight to both developers and users of prediction methods. AVAILABILITY: http://cubic.bioc.columbia.edu/eva. CONTACT: eva@cubic.bioc.columbia.edu

Automation↗

Protein structure modeling for structural genomics.

The shapes of most protein sequences will be modeled based on their similarity to experimentally determined protein structures. The current role, limitations, challenges and prospects for protein structure modeling (using information about genes and genomes) are discussed in the context of structural genomics.

Computer Simulation↗

Comparison of the dynamics of bovine and human angiogenin: a molecular dynamics study.

Molecular dynamics simulations have been carried out for 1 ns on human and bovine angiogenin systems in an effort to compare and contrast their dynamics. An analysis of their dynamics is done by examining the rms deviations, following hydrogen-bonding interactions and looking at the role of water in and around the protein. The C-terminus of bovine angiogenin moves appreciably during dynamics suggesting a better structure for ligand binding. However, we do not find any evidence of a conformation where the glutamate residue that obstructs the active site takes on a different conformation. We observe a differential hydrogen-bonding pattern in the active site regions of bovine and human angiogenins, which could have a bearing on the different catalytic activities of the proteins. We also propose that the differential binding of the monoclonal antibody toward the two proteins might be due sequential and not conformational differences. Water molecules might play an important functional role in both proteins given their subtle functional differences. A simple computation on the molecular dynamics data has been carried out to identify locations in and around the protein that are invariably occupied by water. The locations of nearly half the waters we have identified from the simulation as being invariant in bovine angiogenin occupy similar locations in the bovine angiogenin crystal structure. The positions of the waters identified in human angiogenin differ considerably from that of bovine angiogenin.

Animals↗