Search PubMed⌕ Search

Biomedical subjects

Robert P Sheridan

Publications and source records attributed to Robert P Sheridan.

15 recordsLinked to original sources

Web enabling technology for the design, enumeration, optimization and tracking of compound libraries.

Motivated by the need to augment Merck's in-house small molecule collection, web-based tools for designing, enumerating, optimizing and tracking compound libraries have been developed. The path leading to the current version of this Virtual Library Tool Kit (VLTK) is discussed in context of the (then) available commercial offerings and the constraints and requirements imposed by the end users. Though the effort was initiated to simplify the tasks of designing novel, drug-like and diverse compound libraries containing between 2K-10K unique entities, it has also evolved into a powerful tool for outsourcing syntheses as well as lead identification and optimization. The web tool includes components that select reagents, analyze synthons, identify backup reagents, enumerate libraries, calculate properties, optimize libraries and finally track the synthesized compounds through biological assays. In addition to accommodating project specific designs and virtual 3D library scanning, the application includes tools for parallel synthesis, laboratory automation and compound registration.

Combinatorial Chemistry Techniques↗

A model for predicting likely sites of CYP3A4-mediated metabolism on drug-like molecules.

We have developed a rapid semiquantitative model for evaluating the relative susceptibilities of different sites on drug molecules to metabolism by cytochrome P450 3A4. The model is based on the energy necessary to remove a hydrogen radical from each site, plus the surface area exposure of the hydrogen atom. The energy of hydrogen radical abstraction is conventionally measured by AM1 semiempirical molecular orbital calculations. AM1 calculations show the following order of radical stabilities for the hydrogen atom abstractions: sp2 centers > heteroatom sp3 centers > carbon sp3 centers. Since AM1 calculations are too time intensive for routine work, we developed a statistical trend vector model, which is used to estimate the AM1 abstraction energy of a hydrogen atom from its local atomic environment. We carried out AM1 and trend vector calculations on 50 CYP3A4 substrates whose major sites of metabolism are known in the literature. A plot of the lowest hydrogen radical formation energy versus its sterically accessible surface area exposure for these 50 substrates shows that only those hydrogen atoms with solvent accessible surface area exposure > or = 8.0 A(2) are susceptible to CYP3A4-mediated metabolism. This approach forms the basis for our general model, which predicts sites on drugs that are susceptible to cytochrome P450 3A4-mediated hydrogen radical abstraction followed by a hydroxylation reaction. This model, in conjunction with specific enzyme site binding requirements, can aid in identifying possible sites of metabolism catalyzed by other cytochrome P450 enzymes.

Binding Sites↗

Amino acid substitution of arginine 80 in 17beta-hydroxysteroid dehydrogenase type 3 and its effect on NADPH cofactor binding and oxidation/reduction kinetics.

17beta-Hydroxysteroid dehydrogenase type 3 (17beta-HSD-3) is a member of the short-chain dehydrogenase/reductase (SDR) family and is essential for the reductive conversion of inactive C(19)-steroid, androstenedione, to the biologically active androgen, testosterone, which plays a central role in the development of the male phenotype. Mutations that inactivate this enzyme give rise to a rare form of male pseudohermaphroditism, referred to as 17beta-HSD-3 deficiency. One such mutation is the replacement of arginine at position 80 with glutamine, compromising enzyme activity by increasing the cofactor binding constant 60-fold. In the absence of a 17beta-HSD-3 crystal structure, we have grafted its amino acid sequence for the NADPH binding site on the X-ray crystal structures of glutathione reductase (Protein Data Bank code 1gra) and 17beta-HSD type 1 (Protein Data Bank codes 1fdv and 1fdu) where we find the trunk of the arginine 80 side chain forms part of the hydrophobic pocket for the purine ring of adenosine while its guanidinium moiety interacts with the 2'-phosphate to both stabilize cofactor binding and neutralize its intrinsic negative charge through two hydrogen bonds. To qualitatively assess the role arginine 80 plays in both selecting and stabilizing NADPH binding, it was replaced with each amino acid and the mutant enzymes subjected to enzymatic analysis. There are only seven enzymes exhibiting any measurable enzymatic activity with arginine approximately lysine>leucine>glutamine>methionine>tyrosine>isoleucine. With an aspartic acid at position 58 in 17beta-HSD-3 occupying the equivalent space in the cofactor binding pocket as arginine 224 in glutathione reductase or serine 12 in 17beta-HSD-1, there was an expectation that some of the mutants might use NADH as a cofactor. In no case was NADH found to substitute for NADPH.

17-Hydroxysteroid Dehydrogenases↗

Why do we need so many chemical similarity search methods?

Computational tools to search chemical structure databases are essential to finding leads early in a drug discovery project. Similarity methods are among the most diverse and most useful. We will present some lessons we have gathered over many years experience with in-house methods on several therapeutic problems. The effectiveness of any similarity method can vary greatly from one biological activity to another in a way that is difficult to predict. Also, any two methods tend to select different subsets of actives from a database, so it is advisable to use several search methods where possible.

Decision Support Techniques↗

A simple method for visualizing the differences between related receptor sites.

Pastor and Cruciani [J. Med. Chem. 38 (1995) 4637] and Kastenholz et al. [J. Med. Chem. 43 (2000) 3033] pioneered methods for comparing related receptors, with the ultimate goal of designing selective ligands. Such methods start with a reasonable superposition of high-resolution three-dimensional (3D) structures of the receptors. Next, molecular field maps are calculated for each receptor. Then the maps are analyzed to determine which map features are correlated with a particular subset of receptors. We present a method FLOGTV, based on the trend vector paradigm [J. Chem. Inf. Comput. Sci. 25 (1985) 64] to perform the analysis. This is mathematically simpler than the GRID/CPCA method of Kastenholz et al. and allows for the simultaneous comparison of many receptor structures. Also, the trend vector paradigm provides a method of selecting isopotential contours that are well above "noise". We demonstrate the method on four examples: HIV proteases versus two-domain acid proteases, thrombin versus trypsin and factor Xa, bacterial dihydrofolate reductases (DHFRs) versus vertebrate DHFRs, and P38 versus ERK protein kinases.

Animals↗

A simple method for visualizing the differences between related receptor sites.

Pastor and Cruciani [J. Med. Chem. 38 (1995) 4637] and Kastenholz et al. [J. Med. Chem. 43 (2000) 3033] pioneered methods for comparing related receptors, with the ultimate goal of designing selective ligands. Such methods start with a reasonable superposition of high-resolution three-dimensional (3D) structures of the receptors. Next, molecular field maps are calculated for each receptor. Then the maps are analyzed to determine which map features are correlated with a particular subset of receptors. We present a method FLOGTV, based on the trend vector paradigm [J. Chem. Inf. Comput. Sci. 25 (1985) 64] to perform the analysis. This is mathematically simpler than the GRID/CPCA method of Kastenholz et al. and allows for the simultaneous comparison of many receptor structures. Also, the trend vector paradigm provides a method of selecting isopotential contours that are well above "noise". We demonstrate the method on four examples: HIV proteases versus two-domain acid proteases, thrombin versus trypsin and factor Xa, bacterial dihydrofolate reductases (DHFRs) versus vertebrate DHFRs, and P38 versus ERK protein kinases.

Algorithms↗

The most common chemical replacements in drug-like compounds.

We have written a method that extracts one-to-one replacements of chemical groups in pairs of drug-like molecules with the same biological activity and counts the frequency of the replacements in a large collection of such molecules. There are two variations on the method that differ in their treatment of replacements in rings. This method is one possible approach to systematically identify candidate bioisosteres. Here we look at the MDDR database because it has a large diversity of drug-like compounds in a large number of therapeutic areas. The most frequent replacements in MDDR seem generally consistent with medicinal chemistry intuition about what chemical groups are equivalent or with groups that are easily converted by synthetic or metabolic pathways. This method can be applied to any set of molecules wherein the molecules can be paired by similar biological activity.

Algorithms↗

Finding multiactivity substructures by mining databases of drug-like compounds.

We have developed a method, given a database of molecules and associated activities, to identify molecular substructures that are associated with many different biological activities. These may be therapeutic areas (e.g. antihypertensive) and/or mechanism-based activities (e.g. renin inhibitor). This information helps us avoid chemical classes that are likely to have unanticipated side effects and also can suggest combinatorial libraries that might have activity on a variety of receptor targets. The method was applied to the USPDI and MDDR databases. There are clearly substructures in each database that occur in many compounds and span a variety of therapeutic categories. Some of these are expected, but some are not.

Combinatorial Chemistry Techniques↗

Random forest: a classification and regression tool for compound classification and QSAR modeling.

A new classification and regression tool, Random Forest, is introduced and investigated for predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Random Forest is an ensemble of unpruned classification or regression trees created by using bootstrap samples of the training data and random feature selection in tree induction. Prediction is made by aggregating (majority vote or averaging) the predictions of the ensemble. We built predictive models for six cheminformatics data sets. Our analysis demonstrates that Random Forest is a powerful tool capable of delivering performance that is among the most accurate methods to date. We also present three additional features of Random Forest: built-in performance assessment, a measure of relative importance of descriptors, and a measure of compound similarity that is weighted by the relative importance of descriptors. It is the combination of relatively high prediction accuracy and its collection of desired features that makes Random Forest uniquely suited for modeling in cheminformatics.

Journal Article↗

Calculating similarities between biological activities in the MDL Drug Data Report database.

There are a number of licensed databases that assign biological activities to druglike compounds. The MDL Drug Data Report (MDDR), compiled from the patent literature, is a popular example. It contains several hundred distinct activities, some of which are therapeutic areas (e.g., Antihypertensive) and some of which are related to specific enzymes or receptors (e.g., ACE inhibitor). There are several data mining applications where it would be useful to calculate a similarity between any two activities. Two distinct activity labels can have a significant similarity for a number of reasons: two activities can be nearly synonymous (e.g., CCK B antagonist vs Gastrin antagonist), one activity may be a subset of another (e.g., Dopamine (D2) agonist vs Dopamine agonist), or an activity can be the mechanism by which another activity works (e.g., ACE inhibitor vs Antihypertensive), etc. In an ideal world, similarities for two activities could be calculated simply by comparing the compounds they have in common, but in hand-curated databases such as the MDDR the assignment of activities to compounds are inevitably inconsistent and incomplete. We propose a number of methods of calculating activity-activity similarities that hopefully compensate for errors in hand-curation. Two of these, TIMI and trend vector, show promise. Soft clustering of the activities using a union of similarity methods shows a reasonable association of therapeutic areas with their mechanisms.

Algorithms↗

Similarity to molecules in the training set is a good discriminator for prediction accuracy in QSAR.

How well can a QSAR model predict the activity of a molecule not in the training set used to create the model? A set of retrospective cross-validation experiments using 20 diverse in-house activity sets were done to find a good discriminator of prediction accuracy as measured by root-mean-square difference between observed and predicted activity. Among the measures we tested, two seem useful: the similarity of the molecule to be predicted to the nearest molecule in the training set and/or the number of neighbors in the training set, where neighbors are those more similar than a user-chosen cutoff. The molecules with the highest similarity and/or the most neighbors are the best-predicted. This trend holds true for narrow training sets and, to a lesser degree, for many diverse training sets and does not depend on which QSAR method or descriptor is used. One may define the similarity using a different descriptor than that used for the QSAR model. The similarity dependence for diverse training sets is somewhat unexpected. It appears to be greater for those data sets where the association of similar activities vs similar structures (as encoded in the Patterson plot) is stronger. We propose a way to estimate the reliability of the prediction of an arbitrary chemical structure on a given QSAR model, given the training set from which the model was derived.

Journal Article↗

Boosting: an ensemble learning tool for compound classification and QSAR modeling.

A classification and regression tool, J. H. Friedman's Stochastic Gradient Boosting (SGB), is applied to predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Stochastic Gradient Boosting is a procedure for building a sequence of models, for instance regression trees (as in this paper), whose outputs are combined to form a predicted quantity, either an estimate of the biological activity, or a class label to which a molecule belongs. In particular, the SGB procedure builds a model in a stage-wise manner by fitting each tree to the gradient of a loss function: e.g., squared error for regression and binomial log-likelihood for classification. The values of the gradient are computed for each sample in the training set, but only a random sample of these gradients is used at each stage. (Friedman showed that the well-known boosting algorithm, AdaBoost of Freund and Schapire, could be considered as a particular case of SGB.) The SGB method is used to analyze 10 cheminformatics data sets, most of which are publicly available. The results show that SGB's performance is comparable to that of Random Forest, another ensemble learning method, and are generally competitive with or superior to those of other QSAR methods. The use of SGB's variable importance with partial dependence plots for model interpretation is also illustrated.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Enhanced virtual screening by combined use of two docking methods: getting the most on a limited budget.

Flexible ligand docking is a routine part of a modern structure-based lead discovery process. As of today, there are quite a number of commercial docking programs that can be used to screen large databases (hundreds of thousands to millions of compounds). However, limiting factors such as the number of commercial software licenses needed to perform docking simultaneously on multiple processors ("software cost") and the relatively long time required per molecule to get good results ("quality-to-speed") should be taken into account when planning a large docking run. How can we optimize the efficiency of selecting lead candidates by docking, in respect to the quality of the results, search speed, and software cost? We present a combination of two methods, our "fast-free-approximate" in-house docking program and the "slow-costly-accurate" ICM-Dock, as an example of one solution to the problem. Our proposed protocol is illustrated by a series of virtual screening experiments aimed at identifying active compounds in the MDL Drug Data Report database. In more than half of the 20 cases examined, at least several actives per protein target were identified in approximately 24 hours per target.

Algorithms↗

Reagent Selector: using Synthon Analysis to visualize reagent properties and assist in combinatorial library design.

Reagent Selector is an intranet-based tool that aids in the selection of reagents for use in combinatorial library construction. The user selects an appropriate reagent group as a query, for example, primary amines, and further refines it on the basis of various physicochemical properties, resulting in a list of potential reagents. The results of this selection process are, in turn, converted into synthons: the fragments or R-groups that are to be incorporated into the combinatorial library. The Synthon Analysis interface graphically depicts the chemical properties for each synthon as a function of the topological bond distance from the scaffold attachment point. Displayed in this fashion, the user is able to visualize the property space for the universe of synthons as well as that of the synthons selected. Ultimately, the reagent list that embodies the selected synthons is made available to the user for reagent procurement. Application of the approach to a sample reagent list for a G-protein coupled receptor targeted library is described.

Combinatorial Chemistry Techniques↗

Molecular transformations as a way of finding and exploiting consistent local QSAR.

The idea of a "transformation", making a small change to a chemical structure, for instance removing or replacing a substituent, is familiar to chemists. We suggest two ways of representing a transformation in silico, as a substructure descriptor difference vector, and as the set of atoms remaining once a maximum common substructure is eliminated. Such transformations can be filtered sensibly, and it is easy to compare one transformation to another. These representations have two applications. First, we can use these methods to automatically organize and display sets of closely related compounds such that any consistent local QSAR in a data set can be easily seen, the T-ANALYZE application. Second, we can suggest to a chemist how to change a molecule "on hand" to a more active one based on local QSAR for that activity, the T-MORPH application.

Journal Article↗