Search PubMed⌕ Search

Biomedical subjects

Sandor Vajda

Publications and source records attributed to Sandor Vajda.

At least 19 recordsLinked to original sources

PIPER: an FFT-based protein docking program with pairwise potentials.

The Fast Fourier Transform (FFT) correlation approach to protein-protein docking can evaluate the energies of billions of docked conformations on a grid if the energy is described in the form of a correlation function. Here, this restriction is removed, and the approach is efficiently used with pairwise interaction potentials that substantially improve the docking results. The basic idea is approximating the interaction matrix by its eigenvectors corresponding to the few dominant eigenvalues, resulting in an energy expression written as the sum of a few correlation functions, and solving the problem by repeated FFT calculations. In addition to describing how the method is implemented, we present a novel class of structure-based pairwise intermolecular potentials. The DARS (Decoys As the Reference State) potentials are extracted from structures of protein-protein complexes and use large sets of docked conformations as decoys to derive atom pair distributions in the reference state. The current version of the DARS potential works well for enzyme-inhibitor complexes. With the new FFT-based program, DARS provides much better docking results than the earlier approaches, in many cases generating 50% more near-native docked conformations. Although the potential is far from optimal for antibody-antigen pairs, the results are still slightly better than those given by an earlier FFT method. The docking program PIPER is freely available for noncommercial applications.

Antibodies↗

Computational solvent mapping reveals the importance of local conformational changes for broad substrate specificity in mammalian cytochromes P450.

Computational solvent mapping moves small organic molecules as probes around a protein surface, finds favorable binding positions, clusters the conformations, and ranks the clusters on the basis of their average free energy. Prior mapping studies of enzymes, crystallized in either substrate-free or substrate-bound form, have shown that the largest number of solvent probe clusters invariably overlaps in the active site. We have applied this method to five cytochromes P450. As expected, the mapping of two bacterial P450s, P450 cam (CYP101) and P450 BM-3 (CYP102), identified the substrate-binding sites in both ligand-bound and ligand-free P450 structures. However, the mapping finds the active site only in the ligand-bound structures of the three mammalian P450s, 2C5, 2C9, and 2B4. Thus, despite the large cavities seen in the unbound structures of these enzymes, the features required for binding small molecules are formed only in the process of substrate binding. The ability of adjusting their binding sites to substrates that differ in size, shape, and polarity is likely to be responsible for the broad substrate specificity of these mammalian P450s. Similar behavior was seen at "hot spots" of protein-protein interfaces that can also bind small molecules in grooves created by induced fit. In addition, the binding of S-warfarin to P450 2C9 creates a high-affinity site for a second ligand, which may help to explain the prevalence of drug-drug interactions involving this and other mammalian P450s.

Animals↗

Computational screening of phthalate monoesters for binding to PPARgamma.

Phthalate esters are ubiquitous environmental contaminants that interact with peroxisome proliferator-activated receptors (PPARs), a family of nuclear receptors. Molecular docking and free energy calculations were performed in an effort to identify novel phthalate ligands of PPARgamma, a subtype expressed in a wide range of human tissues. The method was validated using several agonists and partial agonists of PPARgamma, whose binding orientations were correctly reproduced; however, reduced accuracy in docking was observed with ligands of increasing size and flexibility. Improved results were obtained by introduction of a more accurate scoring function based on the all-atom molecular mechanics potential CHARMM and a generalized Born/surface area solvation term ACE (analytical continuum electrostatics). Comparison of the lowest CHARMM/ACE energy of each phthalate vs the logarithm of the experimentally determined EC(50) value for PPARgamma trans-activation yielded a good correlation (R(2) = 0.82). Thus, we can reliably distinguish phthalates that bind and activate PPARgamma from those that do not, with the computational method predicting relative PPARgamma binding activities with some degree of accuracy. We have applied this method to screen a series of 73 mono-ortho-phthalate esters listed in the Available Chemicals Directory. Several putative PPARgamma binding phthalates were identified, including compounds that are known PPARgamma agonists. These findings support the use of computational methods to identify environmental chemicals that warrant further experimental evaluation for PPAR binding and trans-activation potential in cell-based models.

Binding, Competitive↗

Characterization of protein-ligand interaction sites using experimental and computational methods.

The ability to identify the sites of a protein that can bind with high affinity to small, drug-like compounds has been an important goal in drug design. Accurate prediction of druggable sites and the identification of small compounds binding in those sites have provided the input for fragment-based combinatorial approaches that allow for a more thorough exploration of the chemical space, and that have the potential to yield molecules that are more lead-like than those found using traditional high-throughput screening. Current progress in experimental and computational methods for identifying and characterizing druggable ligand binding sites on protein targets is reviewed herein, including a discussion of successful nuclear magnetic resonance, X-ray crystallography and tethering technologies. Classical geometric and energy-based computational methods are also discussed, with particular focus on two powerful technologies, that is, computational solvent mapping and grand canonical Monte Carlo simulations (as used by Locus Pharmaceuticals Inc). Both methods can be used to reliably identify druggable sites on proteins and to facilitate the design of novel, low-nanomolar-affinity ligands.

Animals↗

Clustering of domains of functionally related enzymes in the interaction database PRECISE by the generation of primary sequence patterns.

The PRECISE database was developed by our laboratory to allow for the systematic study of the ligand interactions common to a set of functionally related enzymes, where an interaction site is defined broadly as any residue(s) that interact with a ligand. During the construction of PRECISE, enzyme chains are extracted from the protein data bank (PDB) and clustered according to functional homology as defined by the enzyme commission (EC) nomenclature system. A sequence representative is chosen from each cluster based on the criterion set forth by the non-redundant PDB set, and pair-wise alignments of each cluster member to the representative are performed. Atom-based residue-ligand interactions are calculated for each cluster member, and the summation of ligand interactions for all cluster members at each aligned position is determined. Although we were able to successfully align most clusters using a simple dynamic programming algorithm, several cluster created exhibited poor pair-wise alignments of each cluster member to its sequence representative. We hypothesized that the observed alignment problems were, in most cases, due to the incorrect separation and alignment of different domains in multi-domain proteins, a mistake that frequently causes error proliferation in functional annotation. Here we present the results of generating primary sequence patterns for each poorly aligned cluster in PRECISE to assess the extent to which multi-domain proteins that are incorrectly aligned contributes to poor pair-wise alignments of each cluster member to its representative. This requires the use of an iterative locally optimal pair-wise alignment algorithm to build a hierarchical similarity-based sequence pattern for a set of functionally related enzymes. Our results show that poor alignments in PRECISE are caused most frequently by the misalignment of multi-domain proteins, and that the generation of primary sequence patterns for the assignment of sequence family membership yields better alignments for the functionally related enzyme clusters in PRECISE than our original alignment algorithm.

Amino Acid Sequence↗

Classification of protein complexes based on docking difficulty.

Based on the results of several groups using different docking methods, the key properties that determine the expected success rate in protein-protein docking calculations are measures of conformational change, interface area, and hydrophobicity. A classification of protein complexes in terms of these measures provides a prediction of docking difficulty. This classification is used to study the targets of the CAPRI docking experiment. Results show that targets with a moderate expected difficulty were indeed predicted well by a number of groups, whereas the use of additional a priori information was necessary to obtain good results for some very difficult targets. The analysis indicates that CAPRI and other relatively large-scale docking studies represent very important steps toward understanding the capabilities and limitations of current protein-protein docking methods.

Algorithms↗

Performance of the first protein docking server ClusPro in CAPRI rounds 3-5.

To evaluate the current status of the protein-protein docking field, the CAPRI experiment came to life. Researchers are given the receptor and ligand 3-dimensional (3D) coordinates before the cocrystallized complex is published. Human predictions of the complex structure are supposed to be submitted within 3 weeks, whereas the server ClusPro has only 24 h and does not make use of any biochemical information. From the 10 targets analyzed in the second evaluation meeting of CAPRI, ClusPro was able to predict meaningful models for 5 targets using only empirical free energy estimates. For two of the targets, the server predictions were assessed to be among the best in the field. Namely, for Targets 8 and 12, ClusPro predicted the model with the most accurate binding-site interface and the model with the highest percentage of nativelike contacts, among 180 and 230 submissions, respectively. After CAPRI, the server has been further developed to predict oligomeric assemblies, and new tools now allow the user to restrict the search for the complex to specific regions on the protein surface, significantly enhancing the predictive capabilities of the server. The performance of ClusPro in CAPRI Rounds 3-5 suggests that clustering the low free energy (i.e., desolvation and electrostatic energy) conformations of a homogeneous conformational sampling of the binding interface is a fast and reliable procedure to detect protein-protein interactions and eliminate false positives. Not including targets that had a significant structural rearrangement upon binding, the success rate of ClusPro was found to be around 71%.

Algorithms↗

Optimal clustering for detecting near-native conformations in protein docking.

Clustering is one of the most powerful tools in computational biology. The conventional wisdom is that events that occur in clusters are probably not random. In protein docking, the underlying principle is that clustering occurs because long-range electrostatic and/or desolvation forces steer the proteins to a low free-energy attractor at the binding region. Something similar occurs in the docking of small molecules, although in this case shorter-range van der Waals forces play a more critical role. Based on the above, we have developed two different clustering strategies to predict docked conformations based on the clustering properties of a uniform sampling of low free-energy protein-protein and protein-small molecule complexes. We report on significant improvements in the automated prediction and discrimination of docked conformations by using the cluster size and consensus as a ranking criterion. We show that the success of clustering depends on identifying the appropriate clustering radius of the system. The clustering radius for protein-protein complexes is consistent with the range of the electrostatics and desolvation free energies (i.e., between 4 and 9 Angstroms); for protein-small molecule docking, the radius is set by van der Waals interactions (i.e., at approximately 2 Angstroms). Without any a priori information, a simple analysis of the histogram of distance separations between the set of docked conformations can evaluate the clustering properties of the data set. Clustering is observed when the histogram is bimodal. Data clustering is optimal if one chooses the clustering radius to be the minimum after the first peak of the bimodal distribution. We show that using this optimal radius further improves the discrimination of near-native complex structures.

Amino Acid Sequence↗

Exploring the binding site structure of the PPAR gamma ligand-binding domain by computational solvent mapping.

Solvent mapping moves molecular probes, small organic molecules containing various functional groups, around the protein surface, finds favorable positions, clusters the conformations, and ranks the clusters based on the average free energy. Using at least six different solvents as probes, the probes cluster in major pockets of the functional site, providing detailed and reliable information on the amino acid residues that are important for ligand binding. Solvent mapping was applied to 12 structures of the peroxisome proliferator activated receptor gamma (PPARgamma) ligand-binding domain (LBD), including 2 structures without a ligand, 2 structures with a partial agonist, and 8 structures with a PPAR agonist bound. The analysis revealed 10 binding "hot spots", 4 in the ligand-binding pocket, 2 in the coactivator-binding region, 1 in the dimerization domain, 2 around the ligand entrance site, and 1 minor site without a known function. Mapping is a major source of information on the role and cooperativity of these sites. It shows that large portions of the ligand-binding site are already formed in the PPARgamma apostructure, but an important pocket near the AF-2 transactivation domain becomes accessible only in structures that are cocrystallized with strong agonists. Conformational changes were seen in several other sites, including one involved in the stabilization of the LBD and two others at the region of the coactivator binding. The number of probe clusters retained by these sites depends on the properties of the bound agonist, providing information on the origin of correlations between ligand and coactivator binding.

Alkanesulfonates↗

PRECISE: a Database of Predicted and Consensus Interaction Sites in Enzymes.

PRECISE (Predicted and Consensus Interaction Sites in Enzymes) is a database of interactions between the amino acid residues of an enzyme and its ligands (substrate and transition state analogs, cofactors, inhibitors and products). It is available online at http://precise.bu.edu/. In the current version, all information on interactions is extracted from the enzyme-ligand complexes in the Protein Data Bank (PDB) by performing the following steps: (i) clustering homologous enzyme chains such that, in each cluster, the proteins have the same EC number and all sequences are similar; (ii) selecting a representative chain for each cluster; (iii) selecting ligand types; (iv) finding non-bonded interactions and hydrogen bonds; and (v) summing the interactions for all chains within the cluster. The output of the search is the color-coded sequence of the representative. The colors indicate the total number of interactions found at each amino acid position in all chains of the cluster. Clicking on a residue displays a detailed list of interactions for that residue. Optional filters allow restricting the output to selected chains in the cluster, to non-bonded or hydrogen bonding interactions, and to selected ligand types. The binding site information is essential for understanding and altering substrate specificity and for the design of enzyme inhibitors.

Amino Acid Sequence↗

Deletion of Ser-171 causes inactivation, proteasome-mediated degradation and complete deficiency of human transaldolase.

Homozygous deletion of three nucleotides coding for Ser-171 (S171) of TAL-H (human transaldolase) has been identified in a female patient with liver cirrhosis. Accumulation of sedoheptulose 7-phosphate raised the possibility of TAL (transaldolase) deficiency in this patient. In the present study, we show that the mutant TAL-H gene was effectively transcribed into mRNA, whereas no expression of the TALDeltaS171 protein or enzyme activity was detected in TALDeltaS171 fibroblasts or lymphoblasts. Unlike wild-type TAL-H-GST fusion protein (where GST stands for glutathione S-transferase), TALDeltaS171-GST was solubilized only in the presence of detergents, suggesting that deletion of Ser-171 caused conformational changes. Recombinant TALDeltaS171 had no enzymic activity. TALDeltaS171 was effectively translated in vitro using rabbit reticulocyte lysates, indicating that the absence of TAL-H protein in TALDeltaS171 fibroblasts and lymphoblasts may be attributed primarily to rapid degradation. Treatment with cell-permeable proteasome inhibitors led to the accumulation of TALDeltaS171 in whole cell lysates and cytosolic extracts of patient lymphoblasts, suggesting that deletion of Ser-171 led to rapid degradation by the proteasome. Although the TALDeltaS171 protein became readily detectable in proteasome inhibitor-treated cells, it displayed no appreciable enzymic activity. The results suggest that deletion of Ser-171 leads to inactivation and proteasome-mediated degradation of TAL-H. Since TAL-H is a regulator of apoptosis signal processing, complete deficiency of TAL-H may be relevant for the pathogenesis of liver cirrhosis.

Cells, Cultured↗

Anchor residues in protein-protein interactions.

We show that the mechanism for molecular recognition requires one of the interacting proteins, usually the smaller of the two, to anchor a specific side chain in a structurally constrained binding groove of the other protein, providing a steric constraint that helps to stabilize a native-like bound intermediate. We identify the anchor residues in 39 protein-protein complexes and verify that, even in the absence of their interacting partners, the anchor side chains are found in conformations similar to those observed in the bound complex. These ready-made recognition motifs correspond to surface side chains that bury the largest solvent-accessible surface area after forming the complex (> or =100 A2). The existence of such anchors implies that binding pathways can avoid kinetically costly structural rearrangements at the core of the binding interface, allowing for a relatively smooth recognition process. Once anchors are docked, an induced fit process further contributes to forming the final high-affinity complex. This later stage involves flexible (solvent-exposed) side chains that latch to the encounter complex in the periphery of the binding pocket. Our results suggest that the evolutionary conservation of anchor side chains applies to the actual structure that these residues assume before the encounter complex and not just to their loci. Implications for protein docking are also discussed.

Antigen-Antibody Complex↗

ClusPro: a fully automated algorithm for protein-protein docking.

ClusPro (http://nrc.bu.edu/cluster) represents the first fully automated, web-based program for the computational docking of protein structures. Users may upload the coordinate files of two protein structures through ClusPro's web interface, or enter the PDB codes of the respective structures, which ClusPro will then download from the PDB server (http://www.rcsb.org/pdb/). The docking algorithms evaluate billions of putative complexes, retaining a preset number with favorable surface complementarities. A filtering method is then applied to this set of structures, selecting those with good electrostatic and desolvation free energies for further clustering. The program output is a short list of putative complexes ranked according to their clustering properties, which is automatically sent back to the user via email.

Algorithms↗

Consensus alignment server for reliable comparative modeling with distant templates.

Consensus is a server developed to produce high-quality alignments for comparative modeling, and to identify the alignment regions reliable for copying from a given template. This is accomplished even when target-template sequence identity is as low as 5%. Combining the output from five different alignment methods, the server produces a consensus alignment, with a reliability measure indicated for each position and a prediction of the regions suitable for modeling. Models built using the server predictions are typically within 3 A rms deviations from the crystal structure. Users can upload a target protein sequence and specify a template (PDB code); if no template is given, the server will search for one. The method has been validated on a large set of homologous protein structure pairs. The Consensus server should prove useful for modelers for whom the structural reliability of the model is critical in their applications. It is currently available at http://structure.bu.edu/cgi-bin/consensus/consensus.cgi.

Algorithms↗

ClusPro: an automated docking and discrimination method for the prediction of protein complexes.

MOTIVATION: Predicting protein interactions is one of the most challenging problems in functional genomics. Given two proteins known to interact, current docking methods evaluate billions of docked conformations by simple scoring functions, and in addition to near-native structures yield many false positives, i.e. structures with good surface complementarity but far from the native. RESULTS: We have developed a fast algorithm for filtering docked conformations with good surface complementarity, and ranking them based on their clustering properties. The free energy filters select complexes with lowest desolvation and electrostatic energies. Clustering is then used to smooth the local minima and to select the ones with the broadest energy wells-a property associated with the free energy at the binding site. The robustness of the method was tested on sets of 2000 docked conformations generated for 48 pairs of interacting proteins. In 31 of these cases, the top 10 predictions include at least one near-native complex, with an average RMSD of 5 A from the native structure. The docking and discrimination method also provides good results for a number of complexes that were used as targets in the Critical Assessment of PRedictions of Interactions experiment. AVAILABILITY: The fully automated docking and discrimination server ClusPro can be found at http://structure.bu.edu

Algorithms↗

Protein-protein docking: is the glass half-full or half-empty?

Are current docking methods capable of building complexes from putative component protein structures? Results of recent computational studies, including those of the CAPRI (Critical Assessment of Protein Interactions) competition, were used to determine the key properties for successful docking and introduce a classification of protein complexes based on docking difficulty. Enzyme-inhibitor complexes could be determined with reasonable accuracy - possibly to within a few alternative structures. Results for antigen-antibody pairs are less predictable, and data for small signaling complexes are generally poor. However, moderate amounts of experimental data can remove uncertainty and the methodology is rapidly improving. Transient complexes with large interface areas undergo substantial conformational change and are beyond the reach of current docking methods. The docking of such complexes might therefore require fundamentally new approaches.

Algorithms↗

Combination of scoring functions improves discrimination in protein-protein docking.

Two structure-based potentials are used for both filtering (i.e., selecting a subset of conformations generated by rigid-body docking), and rescoring and ranking the selected conformations. ACP (atomic contact potential) is an atom-level extension of the Miyazawa-Jernigan potential parameterized on protein structures, whereas RPScore (residue pair potential score) is a residue-level potential, based on interactions in protein-protein complexes. These potentials are combined with other energy terms and applied to 13 sets of protein decoys, as well as to the results of docking 10 pairs of unbound proteins. For both potentials, the ability to discriminate between near-native and non-native docked structures is substantially improved by refining the structures and by adding a van der Waals energy term. It is observed that ACP and RPScore complement each other in a number of ways (e.g., although RPScore yields more hits than ACP, mainly as a result of its better performance for charged complexes, ACP usually ranks the near-native complexes better). As a general solution to the protein-docking problem, we have found that the best discrimination strategies combine either an RPScore filter with an ACP-based scoring function, or an ACP-based filter with an RPScore-based scoring function. Thus, ACP and RPScore capture complementary structural information, and combining them in a multistage postprocessing protocol provides substantially better discrimination than the use of the same potential for both filtering and ranking the docked conformations.

Algorithms↗

Identification of substrate binding sites in enzymes by computational solvent mapping.

Enzyme structures determined in organic solvents show that most organic molecules cluster in the active site, delineating the binding pocket. We have developed algorithms to perform solvent mapping computationally, rather than experimentally, by placing molecular probes (small molecules or functional groups) on a protein surface, and finding the regions with the most favorable binding free energy. The method then finds the consensus site that binds the highest number of different probes. The probe-protein interactions at this site are compared to the intermolecular interactions seen in the known complexes of the enzyme with various ligands (substrate analogs, products, and inhibitors). We have mapped thermolysin, for which experimental mapping results are also available, and six further enzymes that have no experimental mapping data, but whose binding sites are well characterized. With the exception of haloalkane dehalogenase, which binds very small substrates in a narrow channel, the consensus site found by the mapping is always a major subsite of the substrate-binding site. Furthermore, the probes at this location form hydrogen bonds and non-bonded interactions with the same residues that interact with the specific ligands of the enzyme. Thus, once the structure of an enzyme is known, computational solvent mapping can provide detailed and reliable information on its substrate-binding site. Calculations on ligand-bound and apo structures of enzymes show that the mapping results are not very sensitive to moderate variations in the protein coordinates.

Algorithms↗