Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “virtual screening”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

A method for including protein flexibility in protein-ligand docking: improving tools for database mining and virtual screening.

Second-generation methods for docking ligands into their biological receptors, such as FLOG, provide for flexibility of the ligand but not of the receptor. Molecular dynamics based methods, such as free energy perturbation, account for flexibility, solvent effects, etc., but are very time consuming. We combined the use of statistical analysis of conformational samples from short-run protein molecular dynamics with grid-based docking protocols and demonstrated improved performance in two test cases. Our statistical analysis explores the importance of the average strength of a potential interaction with the biological target and optionally applies a weighting depending on the variability in the strength of the interaction seen during dynamics simulation. Using these methods, we improved the num-top-ranked 10% of a database of drug-like molecules, in searches based on the three-dimensional structure of the protein. These methods are able to match the ability of manual docking to assess likely inactivity on steric grounds and indeed to rank order ligands from a homologous series of cyclooxygenase-2 inhibitors with good correlation to their true activity. Furthermore, these methods reduce the need for human intervention in setting up molecular docking experiments.

Combinatorial Chemistry Techniques↗

Evaluation of docking strategies for virtual screening of compound databases: cAMP-dependent serine/threonine kinase as an example.

In an effort to establish efficient docking routines for computational screening of compound databases on protein structures, cAMP-dependent protein kinase has been selected as a test case and a variety of docking options and scoring functions were compared. These included rigid-body and flexible docking and scoring based on surface complementarity and/or force field energy. Inhibitors were removed from complex crystal structures and added to compound libraries in their binding conformations and, in addition, deliberately modified conformations. Rigid-body docking and contact scoring well reproduced two of three experimental enzyme-inhibitor complexes. Ligand docking with flexible torsional angles failed to do so but anchored search of some inhibitors converged near to experimental structures, however, only when energy scoring was applied.

1-(5-Isoquinolinesulfonyl)-2-Methylpiperazine↗

Virtual screening against metalloenzymes for inhibitors and substrates.

Molecular docking uses the three-dimensional structure of a receptor to screen databases of small molecules for potential ligands, often based on energetic complementarity. For many docking scoring functions, which calculate nonbonded interactions, metalloenzymes are challenging because of the partial covalent nature of metal-ligand interactions. To investigate how well molecular docking can identify potential ligands of metalloenzymes using a "standard" scoring function, we have docked the MDL Drug Data Report (MDDR), a functionally annotated database of 95,000 small molecules, against the X-ray crystal structures of five metalloenzymes. These enzymes included three zinc proteases, the nickel analogue of an iron enzyme, and a molybdenum metalloenzyme. The ability of the docking program to retrospectively enrich the annotated ligands as high-scoring hits for each enzyme and to calculate proper geometries was evaluated. In all five systems, the annotated ligands within the MDDR were enriched at least 20 times over random. To test the approach prospectively, a sixth target, the zinc beta-lactamase from Bacteroides fragilis, was screened against the fragment-like subset of the ZINC database. We purchased and tested 15 compounds from among the top 50 top-ranked ligands from docking, and found 5 inhibitors with apparent K(i) values less than 120 microM, the best of which was 2 microM. A more ambitious test still was predicting actual substrates for a seventh target, a Zn-dependent phosphotriesterase from Pseudomonas diminuta. Screening the Available Chemicals Directory (ACD) identified 25 thiophosphate esters as potential substrates within the top 100 ranked compounds. Eight of these, all previously uncharacterized for this enzyme, were acquired and tested, and all were confirmed experimentally as substrates. These results suggest that a simple, noncovalent scoring function may be used to identify inhibitors of at least some metalloenzymes.

Crystallography, X-Ray↗

Virtual screening for SARS-CoV protease based on KZ7088 pharmacophore points.

Pharmacophore modeling can provide valuable insight into ligand-receptor interactions. It can also be used in 3D (dimensional) database searching for potentially finding biologically active compounds and providing new research ideas and directions for drug-discovery projects. To stimulate the structure-based drug design against SARS (severe acute respiratory syndrome), a pharmacophore search was conducted over 3.6 millions of compounds based on the atomic coordinates of the complex obtained by docking KZ7088 (a derivative of AG7088) to SARS CoV M(pro) (coronavirus main proteinase), as reportedly recently (Chou, K. C.; Wei, D. Q.; Zhong, W. Z. Biochem. Biophys. Res. Commun. 2003, 308, 148-151). It has been found that, of the 3.6 millions of compounds screened, 0.07% are with the score satisfying five of the six pharmacophore points. Moreover, each of the hit compounds has been evaluated for druggability according to 13 metrics based on physical, chemical, and structural properties. Of the 0.07% compounds thus retrieved, 17% have a perfect score of 1.0; while 23% with one druggable rule violation, 13% two violations, and 47% more than two violations. If the criterion for druggability is set at a maximum allowance of two rule violations, we obtain that only about 0.03% of the compounds screened are worthy of further tests by experiments. These findings will significantly narrow down the search scope for potential compounds, saving substantial time and money. Finally, the featured templates derived from the current study will also be very useful for guiding the design and synthesis of effective drugs for SARS therapy.

Coronavirus 3C Proteases↗

Analysis of data fusion methods in virtual screening: theoretical model.

This paper presents a theoretical model of how data fusion can be used to combine the results of multiple similarity searches of chemical databases. The model is based on frequency distributions of similarity values that are fused using a multiple integration over regions defined by the particular fusion rule that is being applied. For pairwise fusion, the resulting double integrals are straightforward to evaluate for simple model distributions. Similarity values for recovered-active and recovered-nonactive frequency distributions are independently modeled using a constant background, linearly biased terms, and a first-order correlated term. The model shows that two standard fusion rules can give performance enhancements in some cases but that the results of fusion are dependent on many factors that, taken together, can lead to seemingly inconsistent levels of enhancement.

Algorithms↗

LigandScout: 3-D pharmacophores derived from protein-bound ligands and their use as virtual screening filters.

From the historically grown archive of protein-ligand complexes in the Protein Data Bank small organic ligands are extracted and interpreted in terms of their chemical characteristics and features. Subsequently, pharmacophores representing ligand-receptor interaction are derived from each of these small molecules and its surrounding amino acids. Based on a defined set of only six types of chemical features and volume constraints, three-dimensional pharmacophore models are constructed, which are sufficiently selective to identify the described binding mode and are thus a useful tool for in-silico screening of large compound databases. The algorithms for ligand extraction and interpretation as well as the pharmacophore creation technique from the automatically interpreted data are presented and applied to a rhinovirus capsid complex as application example.

Antineoplastic Agents↗

A virtual screening approach for thymidine monophosphate kinase inhibitors as antitubercular agents based on docking and pharmacophore models.

Docking and pharmacophore screening tools were used to examine the binding of ligands in the active site of thymidine monophosphate kinase of Mycobacterium tuberculosis. Docking analysis of deoxythymidine monophosphate (dTMP) analogues suggests the role of hydrogen bonding and other weak interactions in enzyme selectivity. Water-mediated hydrogen-bond networks and a halogen-bond interaction seem to stabilize the molecular recognition. A pharmacophore model was developed using 20 dTMP analogues. The pharmacophoric features were complementary to the active site residues involved in the ligand recognition. On the basis of these studies, a composite screening model that combines the features from both the docking analysis and the pharmacophore model was developed. The composite model was validated by screening a database spiked with 47 known inhibitors. The model picked up 42 of these, giving an enrichment factor of 17. The validated model was used to successfully screen an in-house database of about 500,000 compounds. Subsequent screening with other filters gave 186 hit molecules.

Antitubercular Agents↗

New methods for ligand-based virtual screening: use of data fusion and machine learning to enhance the effectiveness of similarity searching.

Similarity searching using a single bioactive reference structure is a well-established technique for accessing chemical structure databases. This paper describes two extensions of the basic approach. First, we discuss the use of group fusion to combine the results of similarity searches when multiple reference structures are available. We demonstrate that this technique is notably more effective than conventional similarity searching in scaffold-hopping searches for structurally diverse sets of active molecules; conversely, the technique will do little to improve the search performance if the actives are structurally homogeneous. Second, we make the assumption that the nearest neighbors resulting from a similarity search, using a single bioactive reference structure, are also active and use this assumption to implement approximate forms of group fusion, substructural analysis, and binary kernel discrimination. This approach, called turbo similarity searching, is notably more effective than conventional similarity searching.

Artificial Intelligence↗

Assessing different classification methods for virtual screening.

How well do different classification methods perform in selecting the ligands of a protein target out of large compound collections not used to train the model? Support vector machines, random forest, artificial neural networks, k-nearest-neighbor classification with genetic-algorithm-optimized feature selection, trend vectors, naïve Bayesian classification, and decision tree were used to divide databases into molecules predicted to be active and those predicted to be inactive. Training and predicted activities were treated as binary. The database was generated for the ligands of five different biological targets which have been the object of intense drug discovery efforts: HIV-reverse transcriptase, COX2, dihydrofolate reductase, estrogen receptor, and thrombin. We report significant differences in the performance of the methods independent of the biological target and compound class. Different methods can have different applications; some provide particularly high enrichment, others are strong in retrieving the maximum number of actives. We also show that these methods do surprisingly well in predicting recently published ligands of a target on the basis of initial leads and that a combination of the results of different methods in certain cases can improve results compared to the most consistent method.

Algorithms↗

Virtual screening using binary kernel discrimination: effect of noisy training data and the optimization of performance.

Binary kernel discrimination (BKD) uses a training set of compounds, for which structural and qualitative activity data are available, to produce a model that can then be applied to the structures of other compounds in order to predict their likely activity. Experiments with the MDL Drug Data Report database show that the optimal value of the smoothing parameter, and hence the predictive power of BKD, is crucially dependent on the number of false positives in the training set. It is also shown that the best results for BKD are achieved using one particular optimization method for the determination of the smoothing parameter that lies at the heart of the method and using the Jaccard/Tanimoto coefficient in the kernel function that is used to compute the similarity between a test set molecule and the members of the training set.

Algorithms↗

Applications of self-organizing neural networks in virtual screening and diversity selection.

Artificial neural networks provide a powerful technique for the analysis and modeling of nonlinear relationships between molecular structures and pharmacological activity. Many network types, including Kohonen and counterpropagation, also provide an intuitive method for the visual assessment of correspondence between the input and output data. This work shows how a combination of neural networks and radial distribution function molecular descriptors can be applied in various areas of industrial pharmaceutical research. These applications include the prediction of biological activity, the selection of screening candidates (cherry picking), and the extraction of representative subsets from large compound collections such as combinatorial libraries. The methods described have also been implemented as an easy-to-use Web tool, allowing chemists to perform interactive neural network experiments on the Novartis intranet.

Algorithms↗

The impact of tautomer forms on pharmacophore-based virtual screening.

In the field of in silico screening, many applications do not automatically consider possible tautomeric states of molecules. However, the detection of new compound candidates might rely on correct structural description, which is important for the perfect fit toward the biologically relevant interactions. In this paper, we present a new exhaustive tautomer enumeration approach implemented by means of the CACTVS software package. The approach contains a set of 21 predefined SMIRKS-based transforms and a powerful transformation engine that is capable of generating most tautomers described comprehensively in the literature or found in databases in the field of medicinal chemistry. User-defined tautomer rules applied to specific structural databases or scientific issues can be implemented easily and used instead of the predefined rules. In addition, we describe the impact of tautomer-enriched databases on pharmacophore screening approaches for human matrix metalloproteinase 8 as an example of a protein-based pharmacophore screening scenario and for human cyclin-dependent kinases as an example of a ligand-based pharmacophore screening approach. In both test cases, as a preprocessing step, we have used our new tautomer enumerator tool for the tautomer enrichment of the screening data sets and have used it as a postprocessing step to remove tautomeric duplicates from the results. We could demonstrate that the tautomer-enriched screening data sets show significant advantages compared to their non-enhanced counterparts. The discrimination between hits and nonhits was significantly better in the case of tautomer-enriched databases. Moreover, it has been proved that tautomer-enhanced databases will lead to a higher number of potential hits.

CDC2 Protein Kinase↗

A novel automated lazy learning QSAR (ALL-QSAR) approach: method development, applications, and virtual screening of chemical databases using validated ALL-QSAR models.

A novel automated lazy learning quantitative structure-activity relationship (ALL-QSAR) modeling approach has been developed on the basis of the lazy learning theory. The activity of a test compound is predicted from a locally weighted linear regression model using chemical descriptors and the biological activity of the training set compounds most chemically similar to this test compound. The weights with which training set compounds are included in the regression depend on the similarity of those compounds to a test compound. We have applied the ALL-QSAR method to several experimental chemical data sets including 48 anticonvulsant agents with known ED50 values, 48 dopamine D1-receptor antagonists with known competitive binding affinities (Ki), and a Tetrahymena pyriformis data set containing 250 phenolic compounds with toxicity IGC50 values. When applied to database screening, models developed for anticonvulsant agents identified several known anticonvulsant compounds that were not only absent in the training set but highly chemically dissimilar to the training set compounds. This initial success indicates that ALL-QSAR can be further exploited as a general tool for accurate bioactivity prediction and database screening in drug design and discovery. Because of its local nature, the ALL-QSAR approach appears to be especially well-suited for the development of highly predictive models for the sparse or unevenly distributed data sets.

Database Management Systems↗

"Bayes affinity fingerprints" improve retrieval rates in virtual screening and define orthogonal bioactivity space: when are multitarget drugs a feasible concept?

Conventional similarity searching of molecules compares single (or multiple) active query structures to each other in a relative framework, by means of a structural descriptor and a similarity measure. While this often works well, depending on the target, we show here that retrieval rates can be improved considerably by incorporating an external framework describing ligand bioactivity space for comparisons ("Bayes affinity fingerprints"). Structures are described by Bayes scores for a ligand panel comprising about 1000 activity classes extracted from the WOMBAT database. The comparison of structures is performed via the Pearson correlation coefficient of activity classes, that is, the order in which two structures are similar to the panel activity classes. Compound retrieval on a recently published data set could be improved by as much as 24% relative (9% absolute). Knowledge about the shape of the "bioactive chemical universe" is thus beneficial to identifying similar bioactivities. Principal component analysis was employed to further analyze activity space with the objective to define orthogonal ligand bioactive chemical space, leading to nine major (roughly orthogonal) activity axes. Employing only those nine activity classes, retrieval rates are still comparable to original Bayes affinity fingerprints; thus, the concept of orthogonal bioactive ligand chemical space was validated as being an information-rich but low-dimensional representation of bioactivity space. Correlations between activity classes are a major determinant to gauge whether the desired multitarget activity of drugs is (on the basis of current knowledge) a feasible concept because it measures the extent to which activities can be optimized independently, or only by strongly influencing one another.

Algorithms↗

Identification of structurally diverse growth hormone secretagogue agonists by virtual screening and structure-activity relationship analysis of 2-formylaminoacetamide derivatives.

Two molecules with known growth hormone secretagogue (GHS) agonist activity were used as templates to computationally screen approximately 80000 compounds. A total of 108 candidate compounds were selected, and five of them were found to be active in the low-micromolar range in both cell-based and direct binding assays. These compounds were structurally diverse and significantly differed from known GHS agonists. The most active compound was subjected to SAR evaluation, which slightly increased its potency and identified molecular regions important for specific GHS agonist activity.

Acetamides↗

Soft docking and multiple receptor conformations in virtual screening.

Protein conformational change is an important consideration in ligand-docking screens, but it is difficult to predict. A simple way to account for protein flexibility is to soften the criterion for steric fit between ligand and receptor. A more comprehensive but more expensive method would be to sample multiple receptor conformations explicitly. Here, these two approaches are compared. A "soft" scoring function was created by attenuating the repulsive term in the Lennard-Jones potential, allowing for a closer approach between ligand and protein. The standard, "hard" Lennard-Jones potential was used for docking to multiple receptor conformations. The Available Chemicals Directory (ACD) was screened against two cavity sites in the T4 lysozyme. These sites undergo small but significant conformational changes on ligand binding, making them good systems for soft docking. The ACD was also screened against the drug target aldose reductase, which can undergo large conformational changes on ligand binding. We evaluated the ability of the scoring functions to identify known ligands from among the over 200 000 decoy molecules in the database. The soft potential was always better at identifying known ligands than the hard scoring function when only a single receptor conformation was used. Conversely, the soft function was worse at identifying known leads than the hard function when multiple receptor conformations were used. This was true even for the cavity sites and was especially true for aldose reductase. To test the multiple-conformation method predictively, we screened the ACD for molecules that preferentially docked to the expanded conformation of aldose reductase, known to bind larger ligands. Six novel molecules that ranked among the top 0.66% of hits from the multiple-conformation calculation, but ranked relatively poorly in the soft docking calculation, were tested experimentally for enzyme inhibition. Four of these six inhibited the enzyme, the best with an IC(50) of 8 microM. Although ligands can get better scores in soft docking, the same is also true for decoys. The improved ranking of such decoys can come at the expense of true ligands.

Aldehyde Reductase↗