Search PubMed⌕ Search

Biomedical subjects

Andrej Sali

Publications and source records attributed to Andrej Sali.

At least 37 records · Page 2Linked to original sources

Urokinase-type plasminogen activator is a preferred substrate of the human epithelium serine protease tryptase epsilon/PRSS22.

Tryptase epsilon is a member of the chromosome 16p13.3 family of human serine proteases that is preferentially expressed by epithelial cells. Recombinant pro-tryptase epsilon was generated to understand how the exocytosed zymogen might be activated outside of the epithelial cell, as well as to address its possible role in normal and diseased states. Using expression/site-directed mutagenesis approaches, we now show that Lys20, Cys90, and Asp92 in the protease's substrate-binding cleft regulate its enzymatic activity. We also show that Arg(-1) in the propeptide domain controls its ability to autoactivate. In vitro studies revealed that recombinant tryptase epsilon possesses a restricted substrate specificity. Once activated, tryptase epsilon cannot be inhibited effectively by the diverse array of protease inhibitors present in normal human plasma. Moreover, this epithelium protease is not highly susceptible to alpha1-antitrypsin or secretory leukocyte protease inhibitor, which are present in the lung. Recombinant tryptase epsilon could not cleave fibronectin, vitronectin, laminin, single-chain tissue-type plasminogen activator, plasminogen, or any prominent serum protein. Nevertheless, tryptase epsilon readily converted single-chain pro-urokinase-type plasminogen activator (pro-uPA/scuPA) into its mature, enzymatically active protease. Tryptase epsilon also was able to induce pro-uPA-expressing smooth muscle cells to increase their migration through a basement membrane-like extracellular matrix. The ability to activate uPA in the presence of varied protease inhibitors suggests that tryptase epsilon plays a prominent role in fibrinolysis and other uPA-dependent reactions in the lung.

Animals↗

PIBASE: a comprehensive database of structurally defined protein interfaces.

MOTIVATION: In recent years, the Protein Data Bank (PDB) has experienced rapid growth. To maximize the utility of the high resolution protein-protein interaction data stored in the PDB, we have developed PIBASE, a comprehensive relational database of structurally defined interfaces between pairs of protein domains. It is composed of binary interfaces extracted from structures in the PDB and the Probable Quaternary Structure server using domain assignments from the Structural Classification of Proteins and CATH fold classification systems. RESULTS: PIBASE currently contains 158,915 interacting domain pairs between 105,061 domains from 2125 SCOP families. A diverse set of geometric, physiochemical and topologic properties are calculated for each complex, its domains, interfaces and binding sites. A subset of the interface properties are used to remove interface redundancy within PDB entries, resulting in 20,912 distinct domain-domain interfaces. The complexes are grouped into 989 topological classes based on their patterns of domain-domain contacts. The binary interfaces and their corresponding binding sites are categorized into 18,755 and 30,975 topological classes, respectively, based on the topology of secondary structure elements. The utility of the database is illustrated by outlining several current applications. AVAILABILITY: The database is accessible via the world wide web at http://salilab.org/pibase SUPPLEMENTARY INFORMATION: http://salilab.org/pibase/suppinfo.html.

Algorithms↗

New York-Structural GenomiX Research Consortium (NYSGXRC): a large scale center for the protein structure initiative.

Structural GenomiX, Inc. (SGX), four New York area institutions, and two University of California schools have formed the New York Structural GenomiX Research Consortium (NYSGXRC), an industrial/academic Research Consortium that exploits individual core competencies to support all aspects of the NIH-NIGMS funded Protein Structure Initiative (PSI), including protein family classification and target selection, generation of protein for biophysical analyses, sample preparation for structural studies, structure determination and analyses, and dissemination of results. At the end of the PSI Pilot Study Phase (PSI-1), the NYSGXRC will be capable of producing 100-200 experimentally determined protein structures annually. All Consortium activities can be scaled to increase production capacity significantly during the Production Phase of the PSI (PSI-2). The Consortium utilizes both centralized and de-centralized production teams with clearly defined deliverables and hand-off procedures that are supported by a web-based target/sample tracking system (SGX Laboratory Information Data Management System, LIMS, and NYSGXRC Internal Consortium Experimental Database, ICE-DB). Consortium management is provided by an Executive Committee, which is composed of the PI and all Co-PIs. Progress to date is tracked on a publicly available Consortium web site (http://www.nysgxrc.org) and all DNA/protein reagents and experimental protocols are distributed freely from the New York City Area institutions. In addition to meeting the requirements of the Pilot Study Phase and preparing for the Production Phase of the PSI, the NYSGXRC aims to develop modular technologies that are transferable to structural biology laboratories in both academe and industry. The NYSGXRC PI and Co-PIs intend the PSI to have a transforming effect on the disciplines of X-ray crystallography and NMR spectroscopy of biological macromolecules. Working with other PSI-funded Centers, the NYSGXRC seeks to create the structural biology laboratory of the future. Herein, we present an overview of the organization of the NYSGXRC and describe progress toward development of a high-throughput Gene-->Structure platform. An analysis of current and projected consortium metrics reflects progress to date and delineates opportunities for further technology development.

Cloning, Molecular↗

Multiple cathepsin B isoforms in schistosomula of Trichobilharzia regenti: identification, characterisation and putative role in migration and nutrition.

Among schistosomatids, Trichobilharzia regenti, displays an unusual migration through the peripheral and central nervous system prior to residence in the nasal cavity of the definitive avian host. Migration causes tissue degradation and neuromotor dysfunction both in birds and experimentally infected mice. Although schistosomula have a well-developed gut, the peptidases elaborated that might facilitate nutrition and migration are unknown. This is, in large part, due to the difficulty in isolating large numbers of migrating larvae. We have identified and characterised the major 33 kDa cathepsin B-like cysteine endopeptidase in extracts of migrating schistosomula using fluorogenic peptidyl substrates with high extinction coefficients and irreversible affinity-labels. From first strand schistosomula cDNA, degenerate PCR and Rapid Amplification of cDNA End protocols were used to identify peptidase isoforms termed TrCB1.1-TrCB1.6. Highest sequence homology is to the described Schistosoma mansoni and Schistosoma japonicum cathepsins B1. Two isoforms (TrCB1.5 and 1.6) encode putatively inactive enzymes as the catalytic cysteine is substituted by glycine. Two other isoforms, TrCB1.1 and 1.4, were functionally expressed as zymogens in Pichia pastoris. Specific polyclonal antibodies localised the peptidases exclusively in the gut of schistosomula and reacted with a 33kDa protein in worm extracts. TrCB1.1 zymogen was unable to catalyse its own activation, but was trans-processed and activated by S. mansoni asparaginyl endopeptidase (SmAE aka. S. mansoni legumain). In contrast, TrCB1.4 zymogen auto-activated, but was resistant to the action of SmAE. Both activated isoforms displayed different pH-dependent specificity profiles with peptidyl substrates. Also, both isoforms degraded myelin basic protein, the major protein component of nervous tissue, but were inefficient against hemoglobin, thus supporting the adaptation of T. regenti gut peptidases to parasitism of host nervous tissue.

Animals↗

Structural characterization of components of protein assemblies by comparative modeling and electron cryo-microscopy.

We explore structural characterization of protein assemblies by a combination of electron cryo-microscopy (cryoEM) and comparative protein structure modeling. Specifically, our method finds an optimal atomic model of a given assembly subunit and its position within an assembly by fitting alternative comparative models into a cryoEM map. The alternative models are calculated by MODELLER [J. Mol. Biol. 234 (1993) 313] from different sequence alignments between the modeled protein and its template structures. The fitting of these models into a cryoEM density map is performed either by FOLDHUNTER [J. Mol. Biol. 308 (2001) 1033] or by a new density fitting module of MODELLER (Mod-EM). Identification of the most accurate model is based on the correlation between the model accuracy and the quality of fit into the cryoEM density map. To quantify this correlation, we created a benchmark consisting of eight proteins of different structural folds with corresponding density maps simulated at five resolutions from 5 to 15 angstroms, with three noise levels each. Each of the proteins in the set was modeled based on 300 different alignments to their remotely related templates (12-32% sequence identity), spanning the range from entirely inaccurate to essentially accurate alignments. The benchmark revealed that one of the most accurate models can usually be identified by the quality of its fit into the cryoEM density map, even for noisy maps at 15 angstroms resolution. Therefore, a cryoEM density map can be helpful in improving the accuracy of a comparative model. Moreover, a pseudo-atomic model of a component in an assembly may be built better with comparative models of the native subunit sequences than with experimentally determined structures of their homologs.

Cryoelectron Microscopy↗

Combining electron microscopy and comparative protein structure modeling.

Recently, advances have been made in methods and applications that integrate electron microscopy density maps and comparative modeling to produce atomic structures of macromolecular assemblies. Electron microscopy can benefit from comparative modeling through the fitting of comparative models into electron microscopy density maps. Also, comparative modeling can benefit from electron microscopy through the use of intermediate-resolution density maps in fold recognition, template selection and sequence-structure alignment.

Cryoelectron Microscopy↗

Structural characterization of assemblies from overall shape and subcomplex compositions.

We suggest structure characterization of macromolecular assemblies by combining assembly shape determined by electron cyromicroscopy with information about subunit proximity determined by affinity purification. To achieve this aim, structure characterization is expressed as a problem in satisfaction of spatial restraints that (1) represents subunits as spheres, (2) encodes information about the subunit excluded volume, assembly shape, and pulldowns in a scoring function, and (3) finds subunit configurations that satisfy the input restraints by an optimization of the scoring function. Testing of the approach with model systems suggests its feasibility.

Computational Biology↗

Components of coated vesicles and nuclear pore complexes share a common molecular architecture.

Numerous features distinguish prokaryotes from eukaryotes, chief among which are the distinctive internal membrane systems of eukaryotic cells. These membrane systems form elaborate compartments and vesicular trafficking pathways, and sequester the chromatin within the nuclear envelope. The nuclear pore complex is the portal that specifically mediates macromolecular trafficking across the nuclear envelope. Although it is generally understood that these internal membrane systems evolved from specialized invaginations of the prokaryotic plasma membrane, it is not clear how the nuclear pore complex could have evolved from organisms with no analogous transport system. Here we use computational and biochemical methods to perform a structural analysis of the seven proteins comprising the yNup84/vNup107-160 subcomplex, a core building block of the nuclear pore complex. Our analysis indicates that all seven proteins contain either a beta-propeller fold, an alpha-solenoid fold, or a distinctive arrangement of both, revealing close similarities between the structures comprising the yNup84/vNup107-160 subcomplex and those comprising the major types of vesicle coating complexes that maintain vesicular trafficking pathways. These similarities suggest a common evolutionary origin for nuclear pore complexes and coated vesicles in an early membrane-curving module that led to the formation of the internal membrane systems in modern eukaryotes.

Biochemistry↗

Structure-based assessment of missense mutations in human BRCA1: implications for breast and ovarian cancer predisposition.

The BRCA1 gene from individuals at risk of breast and ovarian cancers can be screened for the presence of mutations. However, the cancer association of most alleles carrying missense mutations is unknown, thus creating significant problems for genetic counseling. To increase our ability to identify cancer-associated mutations in BRCA1, we set out to use the principles of protein three-dimensional structure as well as the correlation between the cancer-associated mutations and those that abolish transcriptional activation. Thirty-one of 37 missense mutations of known impact on the transcriptional activation function of BRCA1 are readily rationalized in structural terms. Loss-of-function mutations involve nonconservative changes in the core of the BRCA1 C-terminus (BRCT) fold or are localized in a groove that presumably forms a binding site involved in the transcriptional activation by BRCA1; mutations that do not abolish transcriptional activation are either conservative changes in the core or are on the surface outside of the putative binding site. Next, structure-based rules for predicting functional consequences of a given missense mutation were applied to 57 germ-line BRCA1 variants of unknown cancer association. Such a structure-based approach may be helpful in an integrated effort to identify mutations that predispose individuals to cancer.

BRCA1 Protein↗

MODBASE, a database of annotated comparative protein structure models, and associated resources.

MODBASE (http://salilab.org/modbase) is a relational database of annotated comparative protein structure models for all available protein sequences matched to at least one known protein structure. The models are calculated by MODPIPE, an automated modeling pipeline that relies on the MODELLER package for fold assignment, sequence-structure alignment, model building and model assessment (http:/salilab.org/modeller). MODBASE uses the MySQL relational database management system for flexible querying and CHIMERA for viewing the sequences and structures (http://www.cgl.ucsf.edu/chimera/). MODBASE is updated regularly to reflect the growth in protein sequence and structure databases, as well as improvements in the software for calculating the models. For ease of access, MODBASE is organized into different data sets. The largest data set contains 1,26,629 models for domains in 659,495 out of 1,182,126 unique protein sequences in the complete Swiss-Prot/TrEMBL database (August 25, 2003); only models based on alignments with significant similarity scores and models assessed to have the correct fold despite insignificant alignments are included. Another model data set supports target selection and structure-based annotation by the New York Structural Genomics Research Consortium; e.g. the 53 new structures produced by the consortium allowed us to characterize structurally 24,113 sequences. MODBASE also contains binding site predictions for small ligands and a set of predicted interactions between pairs of modeled sequences from the same genome. Our other resources associated with MODBASE include a comprehensive database of multiple protein structure alignments (DBALI, http://salilab.org/dbali) as well as web servers for automated comparative modeling with MODPIPE (MODWEB, http://salilab. org/modweb), modeling of loops in protein structures (MODLOOP, http://salilab.org/modloop) and predicting functional consequences of single nucleotide polymorphisms (SNPWEB, http://salilab. org/snpweb).

Amino Acid Sequence↗

A structural perspective on protein-protein interactions.

Structures of macromolecular complexes are necessary for a mechanistic description of biochemical and cellular processes. They can be solved by experimental methods, such as X-ray crystallography, NMR spectroscopy and electron microscopy, as well as by computational protein structure prediction, docking and bioinformatics. Recent advances and applications of these methods emphasize the need for hybrid approaches that combine a variety of data to achieve better efficiency, accuracy, resolution and completeness.

Animals↗

High-throughput computational and experimental techniques in structural genomics.

Structural genomics has as its goal the provision of structural information for all possible ORF sequences through a combination of experimental and computational approaches. The access to genome sequences and cloning resources from an ever-widening array of organisms is driving high-throughput structural studies by the New York Structural Genomics Research Consortium. In this report, we outline the progress of the Consortium in establishing its pipeline for structural genomics, and some of the experimental and bioinformatics efforts leading to structural annotation of proteins. The Consortium has established a pipeline for structural biology studies, automated modeling of ORF sequences using solved (template) structures, and a novel high-throughput approach (metallomics) to examining the metal binding to purified protein targets. The Consortium has so far produced 493 purified proteins from >1077 expression vectors. A total of 95 have resulted in crystal structures, and 81 are deposited in the Protein Data Bank (PDB). Comparative modeling of these structures has generated >40,000 structural models. We also initiated a high-throughput metal analysis of the purified proteins; this has determined that 10%-15% of the targets contain a stoichiometric structural or catalytic transition metal atom. The progress of the structural genomics centers in the U.S. and around the world suggests that the goal of providing useful structural information on most all ORF domains will be realized. This projected resource will provide structural biology information important to understanding the function of most proteins of the cell.

Algorithms↗

Detection of homologous proteins by an intermediate sequence search.

We developed a variant of the intermediate sequence search method (ISS(new)) for detection and alignment of weakly similar pairs of protein sequences. ISS(new) relates two query sequences by an intermediate sequence that is potentially homologous to both queries. The improvement was achieved by a more robust overlap score for a match between the queries through an intermediate. The approach was benchmarked on a data set of 2369 sequences of known structure with insignificant sequence similarity to each other (BLAST E-value larger than 0.001); 2050 of these sequences had a related structure in the set. ISS(new) performed significantly better than both PSI-BLAST and a previously described intermediate sequence search method. PSI-BLAST could not detect correct homologs for 1619 of the 2369 sequences. In contrast, ISS(new) assigned a correct homolog as the top hit for 121 of these 1619 sequences, while incorrectly assigning homologs for only nine targets; it did not assign homologs for the remainder of the sequences. By estimate, ISS(new) may be able to assign the folds of domains in approximately 29,000 of the approximately 500,000 sequences unassigned by PSI-BLAST, with 90% specificity (1 - false positives fraction). In addition, we show that the 15 alignments with the most significant BLAST E-values include the nearly best alignments constructed by ISS(new).

Algorithms↗

Alignment of protein sequences by their profiles.

The accuracy of an alignment between two protein sequences can be improved by including other detectably related sequences in the comparison. We optimize and benchmark such an approach that relies on aligning two multiple sequence alignments, each one including one of the two protein sequences. Thirteen different protocols for creating and comparing profiles corresponding to the multiple sequence alignments are implemented in the SALIGN command of MODELLER. A test set of 200 pairwise, structure-based alignments with sequence identities below 40% is used to benchmark the 13 protocols as well as a number of previously described sequence alignment methods, including heuristic pairwise sequence alignment by BLAST, pairwise sequence alignment by global dynamic programming with an affine gap penalty function by the ALIGN command of MODELLER, sequence-profile alignment by PSI-BLAST, Hidden Markov Model methods implemented in SAM and LOBSTER, pairwise sequence alignment relying on predicted local structure by SEA, and multiple sequence alignment by CLUSTALW and COMPASS. The alignment accuracies of the best new protocols were significantly better than those of the other tested methods. For example, the fraction of the correctly aligned residues relative to the structure-based alignment by the best protocol is 56%, which can be compared with the accuracies of 26%, 42%, 43%, 48%, 50%, 49%, 43%, and 43% for the other methods, respectively. The new method is currently applied to large-scale comparative protein structure modeling of all known sequences.

Algorithms↗

Comprehensive search for cysteine cathepsins in the human genome.

Our study was aimed at examinating whether or not the human genome encodes for previously unreported cysteine cathepsins. To this end, we used analyses of the genome sequence and mRNA expression levels. The program TBLASTN was employed to scan the draft sequence of the human genome for the 11 known cysteine cathepsins. The cathepsin-like segments in the genome were inspected, filtered, and annotated. In addition to the known cysteine cathepsins, the scan identified three pseudogenes, closely related to cathepsin L, on chromosome 10, as well as two remote homologs, tubulointerstitial protein antigen and tubulointerstitial protein antigen-related protein. No new members of the family were identified. mRNA expression profiles for 10 known human cysteine cathepsins showed varying expression levels in 46 different human tissues and cell lines. No expression of any of the three cathepsin L-like pseudogenes was found. Based on these results, it is likely that to date all human cysteine cathepsins are known.

Cathepsins↗

ModLoop: automated modeling of loops in protein structures.

SUMMARY: ModLoop is a web server for automated modeling of loops in protein structures. The input is the atomic coordinates of the protein structure in the Protein Data Bank format, and the specification of the starting and ending residues of one or more segments to be modeled, containing no more than 20 residues in total. The output is the coordinates of the non-hydrogen atoms in the modeled segments. A user provides the input to the server via a simple web interface, and receives the output by e-mail. The server relies on the loop modeling routine in MODELLER that predicts the loop conformations by satisfaction of spatial restraints, without relying on a database of known protein structures. For a rapid response, ModLoop runs on a cluster of Linux PC computers. AVAILABILITY: The server is freely accessible to academic users at http://salilab.org/modloop

Amino Acid Sequence↗

Comparative protein structure modeling by iterative alignment, model building and model assessment.

Comparative or homology protein structure modeling is severely limited by errors in the alignment of a modeled sequence with related proteins of known three-dimensional structure. To ameliorate this problem, we have developed an automated method that optimizes both the alignment and the model implied by it. This task is achieved by a genetic algorithm protocol that starts with a set of initial alignments and then iterates through re-alignment, model building and model assessment to optimize a model assessment score. During this iterative process: (i) new alignments are constructed by application of a number of operators, such as alignment mutations and cross-overs; (ii) comparative models corresponding to these alignments are built by satisfaction of spatial restraints, as implemented in our program MODELLER; (iii) the models are assessed by a variety of criteria, partly depending on an atomic statistical potential. When testing the procedure on a very difficult set of 19 modeling targets sharing only 4-27% sequence identity with their template structures, the average final alignment accuracy increased from 37 to 45% relative to the initial alignment (the alignment accuracy was measured as the percentage of positions in the tested alignment that were identical to the reference structure-based alignment). Correspondingly, the average model accuracy increased from 43 to 54% (the model accuracy was measured as the percentage of the C(alpha) atoms of the model that were within 5 A of the corresponding C(alpha) atoms in the superposed native structure). The present method also compares favorably with two of the most successful previously described methods, PSI-BLAST and SAM. The accuracy of the final models would be increased further if a better method for ranking of the models were available.

Algorithms↗