Search PubMed⌕ Search

Biomedical subjects

Andrej Sali

Publications and source records attributed to Andrej Sali.

53 records · Page 3Linked to original sources

Tools for comparative protein structure modeling and analysis.

The following resources for comparative protein structure modeling and analysis are described (http://salilab.org): MODELLER, a program for comparative modeling by satisfaction of spatial restraints; MODWEB, a web server for automated comparative modeling that relies on PSI-BLAST, IMPALA and MODELLER; MODLOOP, a web server for automated loop modeling that relies on MODELLER; MOULDER, a CPU intensive protocol of MODWEB for building comparative models based on distant known structures; MODBASE, a comprehensive database of annotated comparative models for all sequences detectably related to a known structure; MODVIEW, a Netscape plugin for Linux that integrates viewing of multiple sequences and structures; and SNPWEB, a web server for structure-based prediction of the functional impact of a single amino acid substitution.

Internet↗

EVA: Evaluation of protein structure prediction servers.

EVA (http://cubic.bioc.columbia.edu/eva/) is a web server for evaluation of the accuracy of automated protein structure prediction methods. The evaluation is updated automatically each week, to cope with the large number of existing prediction servers and the constant changes in the prediction methods. EVA currently assesses servers for secondary structure prediction, contact prediction, comparative protein structure modelling and threading/fold recognition. Every day, sequences of newly available protein structures in the Protein Data Bank (PDB) are sent to the servers and their predictions are collected. The predictions are then compared to the experimental structures once a week; the results are published on the EVA web pages. Over time, EVA has accumulated prediction results for a large number of proteins, ranging from hundreds to thousands, depending on the prediction method. This large sample assures that methods are compared reliably. As a result, EVA provides useful information to developers as well as users of prediction methods.

Automation↗

Study of the structural dynamics of the E coli 70S ribosome using real-space refinement.

Cryo-EM density maps showing the 70S ribosome of E. coli in two different functional states related by a ratchet-like motion were analyzed using real-space refinement. Comparison of the two resulting atomic models shows that the ribosome changes from a compact structure to a looser one, coupled with the rearrangement of many of the proteins. Furthermore, in contrast to the unchanged inter-subunit bridges formed wholly by RNA, the bridges involving proteins undergo large conformational changes following the ratchet-like motion, suggesting an important role of ribosomal proteins in facilitating the dynamics of translation.

Bacterial Proteins↗

From words to literature in structural proteomics.

Technical advances on several frontiers have expanded the applicability of existing methods in structural biology and helped close the resolution gaps between them. As a result, we are now poised to integrate structural information gathered at multiple levels of the biological hierarchy - from atoms to cells - into a common framework. The goal is a comprehensive description of the multitude of interactions between molecular entities, which in turn is a prerequisite for the discovery of general structural principles that underlie all cellular processes.

Animals↗

NIH workshop on structural proteomics of biological complexes.

Recently, some 50 biologists and officials from government funding agencies met at the NIH campus in Bethesda, MD to explore the interdisciplinary science and organization of the emerging field of structural proteomics. Structural proteomics aims to discover most macromolecular complexes and characterize their three-dimensional structures and functional mechanisms in space and time. The goal seems daunting, but the consensus was that the prize would be commensurate with the effort invested, given the importance of molecular machines and functional networks in biology and medicine. Identification of assemblies and transient complexes combined with their structural and functional characterization will allow us to understand, control, design, and change the functioning of larger biological systems as well as to contribute to drug target discovery, lead discovery, and lead optimization for treatment of human disease.

Macromolecular Substances↗

ModView, visualization of multiple protein sequences and structures.

SUMMARY: We describe ModView, a web application for visualization of multiple protein sequences and structures. ModView integrates a multiple structure viewer, a multiple sequence alignment editor, and a database querying engine. It is possible to interactively manipulate hundreds of proteins, to visualize conservative and variable residues, active and binding sites, fragments, and domains in protein families, as well as to display large macromolecular complexes such as ribosomes or viruses. As a Netscape plug-in, ModView can be included in HTML pages along with text and figures, which makes it useful for teaching and presentations. ModView is also suitable as a graphical interface to various databases because it can be controlled through JavaScript commands and called from CGI scripts. AVAILABILITY: ModView is available at http://guitar.rockefeller.edu/modview.

Database Management Systems↗

The zebrafish forkhead transcription factor Foxi1 specifies epibranchial placode-derived sensory neurons.

Vertebrate epibranchial placodes give rise to visceral sensory neurons that transmit vital information such as heart rate, blood pressure and visceral distension. Despite the pivotal roles they play, the molecular program underlying their development is not well understood. Here we report that the zebrafish mutation no soul, in which epibranchial placodes are defective, disrupts the fork headrelated, winged helix domain-containing protein Foxi1. Foxi1 is expressed in lateral placodal progenitor cells. In the absence of foxi1 activity, progenitor cells fail to express the basic helix-loop-helix gene neurogenin that is essential for the formation of neuronal precursors, and the paired homeodomain containing gene phox2a that is essential for neuronal differentiation and maintenance. Consequently, increased cell death is detected indicating that the placodal progenitor cells take on an apoptotic pathway. Furthermore, ectopic expression of foxi1 is sufficient to induce phox2a-positive and neurogenin-positive cells. Taken together, these findings suggest that Foxi1 is an important determination factor for epibranchial placodal progenitor cells to acquire both neuronal fate and subtype visceral sensory identity.

Amino Acid Sequence↗

Use of single point mutations in domain I of beta 2-glycoprotein I to determine fine antigenic specificity of antiphospholipid autoantibodies.

Autoantibodies against beta(2)-glycoprotein I (beta(2)GPI) appear to be a critical feature of the antiphospholipid syndrome (APS). As determined using domain deletion mutants, human autoantibodies bind to the first of five domains present in beta(2)GPI. In this study the fine detail of the domain I epitope has been examined using 10 selected mutants of whole beta(2)GPI containing single point mutations in the first domain. The binding to beta(2)GPI was significantly affected by a number of single point mutations in domain I, particularly by mutations in the region of aa 40-43. Molecular modeling predicted these mutations to affect the surface shape and electrostatic charge of a facet of domain I. Mutation K19E also had an effect, albeit one less severe and involving fewer patients. Similar results were obtained in two different laboratories using affinity-purified anti-beta(2)GPI in a competitive inhibition ELISA and with whole serum in a direct binding ELISA. This study confirms that anti-beta(2)GPI autoantibodies bind to domain I, and that the charged surface patch defined by residues 40-43 contributes to a dominant target epitope.

Amino Acid Substitution↗

Probing the specificity of a trypanosomal aromatic alpha-hydroxy acid dehydrogenase by site-directed mutagenesis.

The aromatic l-alpha-hydroxy acid dehydrogenase (AHDAH) from Trypanosoma cruzi has over 50% sequence identity with cytosolic malate dehydrogenases (cMDHs), yet it is unable to reduce oxaloacetate. Molecular modeling of the three-dimensional structure of AHADH using the pig cMDH as template directed the construction of several mutants. AHADH shares with MDHs the essential catalytic residues H195 and R171 (using Eventoff's numbering). The AHADH A102R mutant became able to reduce oxaloacetate, while remaining fully active towards aromatic alpha-oxoacids. The Y237G mutant diminished its affinity for all of the natural substrates, whereas the double mutant A102R/Y237G was more active than Y237G and had similar activity with oxaloacetate and with aromatic substrates. The present results reinforce our proposal that AHADH arose by a moderate number of point mutations from a cMDH no longer present in the parasite.

Alcohol Oxidoreductases↗

RasGRP4, a new mast cell-restricted Ras guanine nucleotide-releasing protein with calcium- and diacylglycerol-binding motifs. Identification of defective variants of this signaling protein in asthma, mastocytosis, and mast cell leukemia patients and demonstration of the importance of RasGRP4 in mast cell development and function.

A cDNA was isolated from interleukin 3-developed, mouse bone marrow-derived mast cells (MCs) that contained an insert (designated mRasGRP4) that had not been identified in any species at the gene, mRNA, or protein level. By using a homology-based cloning approach, the approximately 2.6-kb hRasGRP4 transcript was also isolated from the mononuclear progenitors residing in the peripheral blood of normal individuals. This transcript information was then used to locate the RasGRP4 gene in the mouse and human genomes, to deduce its exon/intron organization, and then to identify 10 single nucleotide polymorphisms in the human gene that result in 5 amino acid differences. The >15-kb hRasGRP4 gene consists of 18 exons and resides on a region of chromosome 19q13.1 that had not been sequenced by the Human Genome Project. Human and mouse MCs and their progenitors selectively express RasGRP4, and this new intracellular protein contains all of the domains present in the RasGRP family of guanine nucleotide exchange factors even though it is <50% identical to its closest homolog. Recombinant RasGRP4 can activate H-Ras in a cation-dependent manner. Transfection experiments also suggest that RasGRP4 is a diacylglycerol/phorbol ester receptor. Transcript analysis of an asthma patient, a mastocytosis patient, and the HMC-1 cell line derived from a MC leukemia patient revealed the presence of substantial amounts of non-functional forms of hRasGRP4 due to an inability to remove intron 5 in the precursor transcript. Because only abnormal forms of hRasGRP4 were identified in the HMC-1 cell line, this immature MC progenitor was used to address the function of RasGRP4 in MCs. HMC-1 leukemia cells differentiated and underwent granule maturation when induced to express a normal form of RasGRP4. Thus, RasGRP4 plays an important role in the final stages of MC development.

Amino Acid Sequence↗

MODBASE, a database of annotated comparative protein structure models.

MODBASE (http://guitar.rockefeller.edu/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on PSI-BLAST, IMPALA and MODELLER. MODBASE uses the MySQL relational database management system for flexible and efficient querying, and the MODVIEW Netscape plugin for viewing and manipulating multiple sequences and structures. It is updated regularly to reflect the growth of the protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different datasets. The largest dataset contains models for domains in 304 517 out of 539 171 unique protein sequences in the complete TrEMBL database (23 March 2001); only models based on significant alignments (PSI-BLAST E-value < 10(-4)) and models assessed to have the correct fold are included. Other datasets include models for target selection and structure-based annotation by the New York Structural Genomics Research Consortium, models for prediction of genes in the Drosophila melanogaster genome, models for structure determination of several ribosomal particles and models calculated by the MODWEB comparative modeling web server.

Animals↗

Statistical potentials for fold assessment.

A protein structure model generally needs to be evaluated to assess whether or not it has the correct fold. To improve fold assessment, four types of a residue-level statistical potential were optimized, including distance-dependent, contact, Phi/Psi dihedral angle, and accessible surface statistical potentials. Approximately 10,000 test models with the correct and incorrect folds were built by automated comparative modeling of protein sequences of known structure. The criterion used to discriminate between the correct and incorrect models was the Z-score of the model energy. The performance of a Z-score was determined as a function of many variables in the derivation and use of the corresponding statistical potential. The performance was measured by the fractions of the correctly and incorrectly assessed test models. The most discriminating combination of any one of the four tested potentials is the sum of the normalized distance-dependent and accessible surface potentials. The distance-dependent potential that is optimal for assessing models of all sizes uses both C(alpha) and C(beta) atoms as interaction centers, distinguishes between all 20 standard residue types, has the distance range of 30 A, and is derived and used by taking into account the sequence separation of the interacting atom pairs. The terms for the sequentially local interactions are significantly less informative than those for the sequentially nonlocal interactions. The accessible surface potential that is optimal for assessing models of all sizes uses C(beta) atoms as interaction centers and distinguishes between all 20 standard residue types. The performance of the tested statistical potentials is not likely to improve significantly with an increase in the number of known protein structures used in their derivation. The parameters of fold assessment whose optimal values vary significantly with model size include the size of the known protein structures used to derive the potential and the distance range of the accessible surface potential. Fold assessment by statistical potentials is most difficult for the very small models. This difficulty presents a challenge to fold assessment in large-scale comparative modeling, which produces many small and incomplete models. The results described in this study provide a basis for an optimal use of statistical potentials in fold assessment.

Algorithms↗

Reliability of assessment of protein structure prediction methods.

The reliability of ranking of protein structure modeling methods is assessed. The assessment is based on the parametric Student's t test and the nonparametric Wilcox signed rank test of statistical significance of the difference between paired samples. The approach is applied to the ranking of the comparative modeling methods tested at the fourth meeting on Critical Assessment of Techniques for Protein Structure Prediction (CASP). It is shown that the 14 CASP4 test sequences may not be sufficient to reliably distinguish between the top eight methods, given the model quality differences and their standard deviations. We suggest that CASP needs to be supplemented by an assessment of protein structure prediction methods that is automated, continuous in time, based on several criteria applied to a large number of models, and with quantitative statistical reliability assigned to each characterization.

Computer Simulation↗

Evolution and physics in comparative protein structure modeling.

From a physical perspective, the native structure of a protein is a consequence of physical forces acting on the protein and solvent atoms during the folding process. From a biological perspective, the native structure of proteins is a result of evolution over millions of years. Correspondingly, there are two types of protein structure prediction methods, de novo prediction and comparative modeling. We review comparative protein structure modeling and discuss the incorporation of physical considerations into the modeling process. A good starting point for achieving this aim is provided by comparative modeling by satisfaction of spatial restraints. Incorporation of physical considerations is illustrated by an inclusion of solvation effects into the modeling of loops.

Biophysical Phenomena↗

LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures.

A database comprising all ligand-binding sites of known structure aligned with all related protein sequences and structures is described. Currently, the database contains approximately 50000 ligand-binding sites for small molecules found in the Protein Data Bank (PDB). The structure-structure alignments are obtained by the Combinatorial Extension (CE) program (Shindyalov and Bourne, Protein Eng., 11, 739-747, 1998) and sequence-structure alignments are extracted from the ModBase database of comparative protein structure models for all known protein sequences (Sanchez et al., Nucleic Acids Res., 28, 250-253, 2000). It is possible to search for binding sites in LigBase by a variety of criteria. LigBase reports summarize ligand data including relevant structural information from the PDB file, such as ligand type and size, and contain links to all related protein sequences in the TrEMBL database. Residues in the binding sites are graphically depicted for comparison with other structurally defined family members. LigBase provides a resource for the analysis of families of related binding sites.

Binding Sites↗