Search PubMed⌕ Search

PubMed · 10794420

Analysis of knowledge-based protein-ligand potentials using a self-consistent method.

Abstract

We propose a self-consistent approach to analyze knowledge-based atom-atom potentials used to calculate protein-ligand binding energies. Ligands complexed to actual protein structures were first built using the SMoG growth procedure (DeWitte & Shakhnovich, 1996) with a chosen input potential. These model protein-ligand complexes were used to construct databases from which knowledge-based protein-ligand potentials were derived. We then tested several different modifications to such potentials and evaluated their performance on their ability to reconstruct the input potential using the statistical information available from a database composed of model complexes. Our data indicate that the most significant improvement resulted from properly accounting for the following key issues when estimating the reference state: (1) the presence of significant nonenergetic effects that influence the contact frequencies and (2) the presence of correlations in contact patterns due to chemical structure. The most successful procedure was applied to derive an atom-atom potential for real protein-ligand complexes. Despite the simplicity of the model (pairwise contact potential with a single interaction distance), the derived binding free energies showed a statistically significant correlation (approximately 0.65) with experimental binding scores for a diverse set of complexes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J Shimada, A V Ishchenko, E I Shakhnovich. 2000. Analysis of knowledge-based protein-ligand potentials using a self-consistent method.. https://doi.org/10.1110/ps.9.4.765

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Identification of MHC Ligands Through Allele-Guided Isolation Combined With Machine Learning for Improved MHC Assignment Using ARDisplay-I.

The isolation of major histocompatibility complex (MHC) ligands and subsequent analysis by mass spectrometry is considered the gold standard for defining targets for T cell-based immunotherapies. However, as many targets of high tumor specificity are only presented at low abundance on the cell surface of tumor cells, the efficient isolation of these peptides is crucial for their successful detection. Here, we demonstrate how optimizing the MHC ligand isolation strategy, based on both the presenting MHC alleles and the individual peptide level, enhances the identification of specific MHC ligands. This ideally acknowledges not only the hydrophobicity but also the post-translational modifications of the respective MHC ligands. To further improve the identification and characterization of MHC ligands, we developed an MHC class I ligand prediction algorithm (ARDisplay-I) that outperforms current state-of-the-art tools when benchmarked against competitors such as netMHCpan 4.1, MixMHCpred, or MHCflurry. Implementing these strategies can augment the development of T cell receptor-based therapies by improving the identification of novel immunotherapy targets and enriching the resources available in the computational immunology field through a superior MHC presentation prediction algorithm.

Ligands↗

An image-based protein-ligand binding representation learning framework via multi-level flexible dynamics trajectory pre-training.

MOTIVATION: Accurate prediction of protein-ligand binding (PLB) relationships plays a crucial role in drug discovery, which helps identify drugs that modulate the activity of specific targets. Traditional biological assays for measuring PLB relationships are time consuming and costly. In addition, models for predicting PLB relationships have been developed and widely used in drug discovery tasks. However, learning more accurate PLB representations is essential to meet the stringent standards required for drug discovery. RESULTS: We propose an image-based PLB representation learning framework, called ImagePLB, which equips ligand representation learner (LRL) and protein representation learner (PRL) to accept 3D multi-view ligand images and protein graphs as input, respectively, and learns rich interaction information between ligand and protein through a binding representation learner (BRL). Considering the scarcity of protein-ligand pairs, we further propose a multi-level next trajectory prediction (MLNTP) task to pre-train ImagePLB on the 4D flexible dynamics trajectory of 16 972 complexes, including ligand level, protein level, and complex level, to learn information related to trajectories. Besides, by introducing trajectory regularization (TR), we effectively alleviate the problem of high (even almost identical) feature similarity caused by adjacent trajectories. Compared with the current state-of-the-art methods, ImagePLB has achieved competitive improvements on PLB-related prediction tasks, including protein-ligand affinity and efficacy prediction tasks. This study opens the door to the image-based PLB learning paradigm. AVAILABILITY AND IMPLEMENTATION: All data and implementation details of code can be obtained from https://github.com/HongxinXiang/ImagePLB.

Ligands↗

A mathematical analysis of SELEX.

Systematic evolution of ligands by exponential enrichment (SELEX) is a procedure by which a mixture of nucleic acids that vary in sequence can be separated into pure components with the goal of isolating those with specific biochemical activities. The basic idea is to combine the mixture with a specific target molecule and then separate the target-NA complex from the resulting reaction. The target-NA complex is then separated by mechanical means (for example by filtration), the NA is then eluted from the complex, amplified by polymerase chain reaction (PCR) and the process repeated. After several rounds, one should be left with a pool of [NA] that consists mostly of the species in the original pool that best binds to the target. In Irvine et al. [Irvine, D., Tuerk, C., Gold, L., 1991. SELEXION, systematic evolution of nucleic acids by exponential enrichment with integrated optimization by non-linear analysis. J. Mol. Biol. 222, 739-761] a mathematical analysis of this process was given. In this paper we revisit Irvine et al. [Ibid]. By rewriting the equations for the SELEX process, we considerably reduce the labor of computing the round to round distribution of nucleic acid fractions. We also establish necessary and sufficient conditions for the SELEX process to converge to a pool consisting solely of the best binding nucleic acid to a fixed target in a manner that maximizes the percentage of bound target. The assumption is that there is a single nucleic acid binding site on the target that permits occupation by not more than one nucleic acid. We analyze the case for which there is no background loss (no support losses and no free [NA] left on the support). We then examine the case in which such there are such losses. The significance of the analysis is that it suggests an experimental approach for the SELEX process as defined in Irvine et al. [Ibid] to converge to a pool consisting of a single best binding nucleic acid without recourse to any a priori information about the nature of the binding constants or the distribution of the individual nucleic acid fragments.

Ligands↗