Search PubMed⌕ Search

Biomedical subjects

Santosh Putta

Publications and source records attributed to Santosh Putta.

9 recordsLinked to original sources

Feature-map vectors: a new class of informative descriptors for computational drug discovery.

In order to develop robust machine-learning or statistical models for predicting biological activity, descriptors that capture the essence of the protein-ligand interaction are required. In the absence of structural information from X-ray or NMR experiments, deriving informative descriptors can be difficult. We have developed feature-map vectors (FMVs), a new class of descriptors based on chemical features, to address this challenge. FMVs, which are derived from the conformational models of a few actives, are low dimensional, problem specific, and highly interpretable. By using shape-based alignments and scoring with chemical features, FMVs can combine information about a molecule's shape and the pharmacophores it can match. In five validation studies, bag classifiers built using FMVs have shown high enrichments for identifying actives for five diverse targets: CDK2, 5-HT(3), DHFR, thrombin, and ACE. The interpretability of these descriptors has been demonstrated for CDK2 and 5-HT(3), where the method automatically discovers the standard literature pharmacophore.

Algorithms↗

Conformation mining: an algorithm for finding biologically relevant conformations.

Discovering essential features shared by active compounds, an important step in drug-design, is complicated by conformational flexibility. We present a new algorithm to efficiently mine the conformational space of multiple actives and find small subsets of conformations likely to be biologically relevant. The approach identifies chemical and steric similarities between actives, providing insight into features important for binding when structural data are absent. Validation studies (thrombin and CDK2 data) produce alignments similar to protein-based alignments.

Algorithms↗

Performance of 3D-database molecular docking studies into homology models.

The performance of docking studies into protein active sites constructed by homology model building was investigated using CDK2 and factor VIIa screening data sets. When the sequence identity between model and template near the binding site area is greater than approximately 50%, roughly 5 times more active compounds are identified than would be found randomly. This performance is comparable to docking to crystal structures.

Binding Sites↗

Building predictive ADMET models for early decisions in drug discovery.

This review discusses the current challenges facing researchers developing computational models to predict absorption, distribution, metabolism, excretion and toxicity (ADMET) for early drug discovery. The strengths and weaknesses of different modeling approaches are reviewed and a survey of recent strategies to model several key ADMET parameters, including intestinal permeability, blood-brain barrier penetration, metabolism, bioavailability and drug toxicities, is presented.

Biological Availability↗

Evaluation of a novel shape-based computational filter for lead evolution: application to thrombin inhibitors.

A novel shape-feature-based computational method is described and used to rapidly filter compound libraries. The computational model, built using three-dimensional conformations of active and inactive molecules, consists of a collection of whole molecule shapes and chemical feature positions that are ranked according to their correlation with activity. A small ensemble of these shapes and features is used to filter virtual compound libraries. The method is applied to two thrombin data sets and is shown to be efficient in identifying novel scaffolds with enhanced hit rates.

Combinatorial Chemistry Techniques↗

A novel shape-feature based approach to virtual library screening.

The shape of and the chemical features of a ligand are both critical for biological activity. This paper presents a strategy that uses these descriptors to build a computational model for virtual screening of bioactive compounds. Molecules are represented in a binary shape-feature descriptor space as bit-strings, and their relative activities are used to identify the subset of the bit-string that is most relevant to bioactivity. This subset is used to score virtual libraries. We describe the computational details of the method and present an example validation experiment on thrombin inhibitors.

Computer Simulation↗

Active learning with support vector machines in the drug discovery process.

We investigate the following data mining problem from computer-aided drug design: From a large collection of compounds, find those that bind to a target molecule in as few iterations of biochemical testing as possible. In each iteration a comparatively small batch of compounds is screened for binding activity toward this target. We employed the so-called "active learning paradigm" from Machine Learning for selecting the successive batches. Our main selection strategy is based on the maximum margin hyperplane-generated by "Support Vector Machines". This hyperplane separates the current set of active from the inactive compounds and has the largest possible distance from any labeled compound. We perform a thorough comparative study of various other selection strategies on data sets provided by DuPont Pharmaceuticals and show that the strategies based on the maximum margin hyperplane clearly outperform the simpler ones.

Computer-Aided Design↗

A novel subshape molecular descriptor.

Molecules with similar shapes and features often have similar biological activity. Several computational approaches search chemical databases for new leads or templates based on overall molecular shape similarity. However, active molecules often present critical subshapes that are required for binding, which may be missed by comparing overall shape similarity. We present a new approach to compare molecular shapes of different sizes and to calculate subshape similarity. We developed a skeletal representation of the shape which is topologically unrelated to covalent chemical connectivity. This simplifies rotational and translational sampling. We test initial possible alignments by matching similar triangles. This triangle-matching filter rapidly eliminates most geometrically impossible matches. Surviving matches are filtered further in successive stages. These stages involve direction, feature, and shape matching procedures. Our approach is applied to several situations demonstrating lead discovery and evolution.

Journal Article↗

Using ensembles to classify compounds for drug discovery.

This paper introduces Signal, a novel method for classifying activity against a small molecule drug target. Signal creates an ensemble, or collection, of meaningful descriptors chosen from a much larger property space. The method works with a variety of descriptor types, including fingerprints that represent four-point pharmacophores or shape descriptors. It also exploits information from both active and inactive compounds and generates predictive models suitable for high throughput screening data analysis. Given the fingerprints and activity data for a set of compounds, Signal is a two step process. The first step is to Evaluate the Descriptors: for each descriptor in the fingerprint, quantify and rank the correlation between the activity of the compounds and the presence of that descriptor. The second step is to Create an Ensemble Model: use the high ranking descriptors to create a model of activity against the biological target. For the first step, two possible ranking strategies were investigated: mutual information and chi-square. For the second step, two types of ensemble models were investigated: high ranking and a novel method called high ranking set cover. Of the four possible pairings, the combination of chi-square and high ranking set cover performed the best on a Thrombin data set.

Algorithms↗