Search PubMed⌕ Search

Biomedical subjects

Alexander Golbraikh

Publications and source records attributed to Alexander Golbraikh.

17 recordsLinked to original sources

Novel ligands for the human histamine H1 receptor: synthesis, pharmacology, and comparative molecular field analysis studies of 2-dimethylamino-5-(6)-phenyl-1,2,3,4-tetrahydronaphthalenes.

This paper reports the synthesis of a novel series of (+/-)-2-dimethylamino- 5- and 6-phenyl-1,2,3,4-tetrahydronaphthalene derivatives (5- and 6-APTs), and, corresponding affinity, functional activity, and, molecular modeling studies with regard to drug design targeting the human histamine H1 receptor. The 5-APTs have 2- to 4-fold higher H1 receptor affinity than the endogenous agonist histamine. The chemical nature of a meta-substituent on the 5-APT pendant phenyl moiety does not significantly affect H1 affinity. In contrast, analogous meta-substitution for the 6-APTs increases H1 affinity up to 100-fold. The new APTs do not activate H1 receptor-linked intracellular signaling and apparently are competitive H1 antagonists. A new model that establishes structural parameters for binding to the human H1 receptor by APTs and other ligands was developed using 3-D QSAR (CoMFA). The model predicts H1 ligand binding with a higher degree of external predictability compared to a previously reported model. The APTs also were examined for activity at human serotonin 5-HT2A and 5-HT2C receptors, which are phylogenetically closely related to the H1 receptor. 5-APT and m-Cl-6-APT were identified as novel agonists that selectively activate 5-HT2C receptors. It is concluded that the lipophilic (brain-penetrating) APT molecular scaffold may have pharmacotherapeutic potential in neuropsychiatric diseases.

Animals↗

Development of quantitative structure-binding affinity relationship models based on novel geometrical chemical descriptors of the protein-ligand interfaces.

Novel geometrical chemical descriptors have been derived on the basis of the computational geometry of protein-ligand interfaces and Pauling atomic electronegativities (EN). Delaunay tessellation has been applied to a diverse set of 517 X-ray characterized protein-ligand complexes yielding a unique collection of interfacial nearest neighbor atomic quadruplets for each complex. Each quadruplet composition was characterized by a single descriptor calculated as the sum of the EN values for the four participating atom types. We termed these simple descriptors generated from atomic EN values and derived with the Delaunay Tessellation the ENTess descriptors and used them in the variable selection k-nearest neighbor quantitative structure-binding affinity relationship (QSBR) studies of 264 diverse protein-ligand complexes with known binding constants. Twenty-four complexes with chemically dissimilar ligands were set aside as an independent validation set, and the remaining dataset of 240 complexes was divided into multiple training and test sets. The best models were characterized by the leave-one-out cross-validated correlation coefficient q(2) as high as 0.66 for the training set and the correlation coefficient R(2) as high as 0.83 for the test set. The high predictive power of these models was confirmed independently by applying them to the validation set of 24 complexes yielding R(2) as high as 0.85. We conclude that QSBR models built with the ENTess descriptors can be instrumental for predicting the binding affinity of receptor-ligand complexes.

Computer Simulation↗

Quantitative structure-activity relationship analysis of pyridinone HIV-1 reverse transcriptase inhibitors using the k nearest neighbor method and QSAR-based database mining.

We have developed quantitative structure-activity relationship (QSAR) models for 44 non-nucleoside HIV-1 reverse transcriptase inhibitors (NNRTIs) of the pyridinone derivative type. The k nearest neighbor (kNN) variable selection approach was used. This method utilizes multiple descriptors such as molecular connectivity indices, which are derived from two-dimensional molecular topology. The modeling process entailed extensive validation including the randomization of the target property (Y-randomization) test and the division of the dataset into multiple training and test sets to establish the external predictive power of the training set models. QSAR models with high internal and external accuracy were generated, with leave-one-out cross-validated R2 (q2) values ranging between 0.5 and 0.8 for the training sets and R2 values exceeding 0.6 for the test sets. The best models with the highest internal and external predictive power were used to search the National Cancer Institute database. Derivatives of the pyrazolo[3,4-d]pyrimidine and phenothiazine type were identified as promising novel NNRTIs leads. Several candidates were docked into the binding pocket of nevirapine with the AutoDock (version 3.0) software. Docking results suggested that these types of compounds could be binding in the NNRTI binding site in a similar mode to a known non-nucleoside inhibitor nevirapine.

Database Management Systems↗

Application of predictive QSAR models to database mining: identification and experimental validation of novel anticonvulsant compounds.

We have developed a drug discovery strategy that employs variable selection quantitative structure-activity relationship (QSAR) models for chemical database mining. The approach starts with the development of rigorously validated QSAR models obtained with the variable selection k nearest neighbor (kNN) method (or, in principle, with any other robust model-building technique). Model validation is based on several statistical criteria, including the randomization of the target property (Y-randomization), independent assessment of the training set model's predictive power using external test sets, and the establishment of the model's applicability domain. All successful models are employed in database mining concurrently; in each case, only variables selected as a result of model building (termed descriptor pharmacophore) are used in chemical similarity searches comparing active compounds of the training set (queries) with those in chemical databases. Specific biological activity (characteristic of the training set compounds) of external database entries found to be within a predefined similarity threshold of the training set molecules is predicted on the basis of the validated QSAR models using the applicability domain criteria. Compounds judged to have high predicted activities by all or the majority of all models are considered as consensus hits. We report on the application of this computational strategy for the first time for the discovery of anticonvulsant agents in the Maybridge and National Cancer Institute (NCI) databases containing ca. 250,000 compounds combined. Forty-eight anticonvulsant agents of the functionalized amino acid (FAA) series were used to build kNN variable selection QSAR models. The 10 best models were applied to mining chemical databases, and 22 compounds were selected as consensus hits. Nine compounds were synthesized and tested at the NIH Epilepsy Branch, Rockville, MD using the same biological test that was employed to assess the anticonvulsant activity of the training set compounds; of these nine, four were exact database hits and five were derived from the hits by minor chemical modifications. Seven of these nine compounds were confirmed to be active, indicating an exceptionally high hit rate. The approach described in this report can be used as a general rational drug discovery tool.

Amides↗

Development and validation of k-nearest-neighbor QSPR models of metabolic stability of drug candidates.

Computational ADME (absorption, distribution, metabolism, and excretion) models may be used early in the drug discovery process in order to flag drug candidates with potentially problematic ADME profiles. We report the development, validation, and application of quantitative structure-property relationship (QSPR) models of metabolic turnover rate for compounds in human S9 homogenate. Biological data were obtained from uniform bioassays of 631 diverse chemicals proprietary to GlaxoSmithKline (GSK). The models were built with topological molecular descriptors such as molecular connectivity indices or atom pairs using the k-nearest neighbor variable selection optimization method developed at the University of North Carolina (Zheng, W.; Tropsha, A. A novel variable selection QSAR approach based on the k-nearest neighbor principle. J. Chem. Inf. Comput. Sci., 2000, 40, 185-194.). For the purpose of validation, the whole data set was divided into training and test sets. The training set QSPR models were characterized by high internal accuracy with leave-one-out cross-validated R(2) (q(2)) values ranging between 0.5 and 0.6. The test set compounds were correctly classified as stable or unstable in S9 assay with an accuracy above 85%. These models were additionally validated by in silico metabolic stability screening of 107 new chemicals under development in several drug discovery programs at GSK. One representative model generated with MolConnZ descriptors predicted 40 compounds to be metabolically stable (turnover rate less than 25%), and 33 of them were indeed found to be stable experimentally. This success (83% concordance) in correctly picking chemicals that are metabolically stable in the human S9 homogenate spells a rapid, computational screen for generating components of the ADME profile in a drug discovery process.

Algorithms↗

Quantitative structure-activity relationship analysis of functionalized amino acid anticonvulsant agents using k nearest neighbor and simulated annealing PLS methods.

We report the development of rigorously validated quantitative structure-activity relationship (QSAR) models for 48 chemically diverse functionalized amino acids with anticonvulsant activity. Two variable selection approaches, simulated annealing partial least squares (SA-PLS) and k nearest neighbor (kNN), were employed. Both methods utilize multiple descriptors such as molecular connectivity indices or atom pair descriptors, which are derived from two-dimensional molecular topology. QSAR models with high internal accuracy were generated, with leave-one-out cross-validated R(2) (q(2)) values ranging between 0.6 and 0.8. The q(2) values for the actual dataset were significantly higher than those obtained for the same dataset with randomly shuffled activity values, indicating that models were statistically significant. The original dataset was further divided into several training and test sets, with highly predictive models providing q(2) values greater than 0.5 for the training sets and R(2) values greater than 0.6 for the test sets. These models were capable of predicting with reasonable accuracy the activity of 13 novel compounds not included in the original dataset. The successful development of highly predictive QSAR models affords further design and discovery of novel anticonvulsant agents.

Amino Acids↗

Antitumor agents. 213. Modeling of epipodophyllotoxin derivatives using variable selection k nearest neighbor QSAR method.

We have applied a variable selection k nearest neighbor quantitative structure-activity relationship (kNN QSAR) method to develop predictive QSAR models for 157 epipodophyllotoxins synthesized previously in our ongoing effort to develop potential anticancer agents. QSAR models were generated using multiple topological descriptors of chemical structures, including molecular connectivity indices (MCI) and molecular operating environment descriptors. The 157 compounds were separated into several training and test sets. The robustness of QSAR models was characterized by the values of the internal leave one out cross-validated R2 (q2) for the training set and external predictive R2 for the test set. The significance of the training set models was confirmed by statistically higher values of q2 for the original data set as compared to q2 values for the same data set with randomly shuffled activities. kNN QSAR models were compared with those obtained with the comparative molecular field analysis method; the kNN QSAR approach afforded models with higher values of both q2 and predictive R2. One of the best models obtained from kNN analysis using MCI as descriptors provided q2 and predictive R2 values of 0.60 and 0.62, respectively. QSAR models developed in these studies shall aid in future design of novel potent epipodophyllotoxin derivatives.

Antineoplastic Agents↗

Beware of q2!

Validation is a crucial aspect of any quantitative structure-activity relationship (QSAR) modeling. This paper examines one of the most popular validation criteria, leave-one-out cross-validated R2 (LOO q2). Often, a high value of this statistical characteristic (q2 > 0.5) is considered as a proof of the high predictive ability of the model. In this paper, we show that this assumption is generally incorrect. In the case of 3D QSAR, the lack of the correlation between the high LOO q2 and the high predictive ability of a QSAR model has been established earlier [Pharm. Acta Helv. 70 (1995) 149; J. Chemomet. 10(1996)95; J. Med. Chem. 41 (1998) 2553]. In this paper, we use two-dimensional (2D) molecular descriptors and k nearest neighbors (kNN) QSAR method for the analysis of several datasets. No correlation between the values of q2 for the training set and predictive ability for the test set was found for any of the datasets. Thus, the high value of LOO q2 appears to be the necessary but not the sufficient condition for the model to have a high predictive power. We argue that this is the general property of QSAR models developed using LOO cross-validation. We emphasize that the external validation is the only way to establish a reliable QSAR model. We formulate a set of criteria for evaluation of predictive ability of QSAR models.

Data Interpretation, Statistical↗

Predictive QSAR modeling based on diversity sampling of experimental datasets for the training and test set selection.

One of the most important characteristics of Quantitative Structure Activity Relashionships (QSAR) models is their predictive power. The latter can be defined as the ability of a model to predict accurately the target property (e.g., biological activity) of compounds that were not used for model development. We suggest that this goal can be achieved by rational division of an experimental SAR dataset into the training and test set, which are used for model development and validation, respectively. Given that all compounds are represented by points in multidimensional descriptor space, we argue that training and test sets must satisfy the following criteria: (i) Representative points of the test set must be close to those of the training set; (ii) Representative points of the training set must be close to representative points of the test set; (iii) Training set must be diverse. For quantitative description of these criteria, we use molecular dataset diversity indices introduced recently (Golbraikh, A., J. Chem. Inf. Comput. Sci., 40 (2000) 414-425). For rational division of a dataset into the training and test sets, we use three closely related sphere-exclusion algorithms. Using several experimental datasets, we demonstrate that QSAR models built and validated with our approach have statistically better predictive power than models generated with either random or activity ranking based selection of the training and test sets. We suggest that rational approaches to the selection of training and test sets based on diversity principles should be used routinely in all QSAR modeling research.

Algorithms↗

Novel ZE-isomerism descriptors derived from molecular topology and their application to QSAR analysis.

We introduce several series of novel ZE-isomerism descriptors derived directly from two-dimensional molecular topology. These descriptors make use of a quantity named ZE-isomerism correction, which is added to the vertex degrees of atoms connected by double bonds in Z and E configurations. This approach is similar to the one described previously for topological chirality descriptors (Golbraikh, A., et al. J. Chem. Inf. Comput. Sci. 2001, 41, 147-158). The ZE-isomerism descriptors include modified molecular connectivity indices, overall Zagreb indices, extended connectivity, overall connectivity, and topological charge indices. They can be either real or complex numbers. Mathematical properties of different subgroups of ZE-isomerism descriptors are discussed. These descriptors circumvent the inability of conventional topological indices to distinguish between Z and E isomers. The applicability of ZE-isomerism descriptors to QSAR analysis is demonstrated in the studies of a series of 131 anticancer agents inhibiting tubulin polymerization.

Antineoplastic Agents↗

QSAR modeling of alpha-campholenic derivatives with sandalwood odor.

Three-dimensional quantitative structure-activity relationship (3D-QSAR) models were developed for a series of 44 synthetic alpha-campholenic derivatives with sandalwood odor. These compounds have complex stereochemistry as they contain up to five chiral atoms. To address stereospecificity of odor intensity, a 3D-QSAR method was developed, which does not require spatial alignment of molecules. In this method, compounds are represented as derivatives of several common structural templates with several substituents, which are numbered according to their relative spatial positions in the molecule. Both wholistic and substituent descriptors calculated with the TSAR software were used as independent variables. Based on published experimental data of sandalwood odor intensities, two discrete scales of the odor intensity with equal or unequal intervals between the threshold values were developed. The data set was divided into a training set of 38 compounds and a test set of six compounds. To build QSAR models, a stepwise multiple linear regression method was used. The best model was obtained using the unequal scale of odor intensity: for the training set, the leave one out cross-validated R(2) (q(2)) was 0.80, the correlation coefficient R between actual and predicted odor intensities was 0.93, and the correlation coefficient for the test set was 0.95. The QSAR models developed in this study contribute to the better understanding of structural, electronic, and lipophilic properties responsible for sandalwood odor. Furthermore, the QSAR approach reported herein can be applied to other data sets that include compounds with complex stereochemistry.

Journal Article↗

QSAR modeling using chirality descriptors derived from molecular topology.

Topological descriptors of chemical structures (such as molecular connectivity indices) are widely used in Quantitative Structure-Activity Relationships (QSAR) studies. Unfortunately, these descriptors lack the ability to discriminate between stereoisomers, which limits their application in QSAR. To circumvent this problem, we recently introduced chirality descriptors derived from molecular graphs and applied them in QSAR studies of ecdysteroids (Golbraikh A.; Bonchev, D.; Tropsha, A. J. Chem. Inf. Comput. Sci. 2001,41, 147-158). In this paper, we extend our earlier work by applying chirality descriptors to four data sets containing chiral compounds. All models were derived with the k-nearest neighbors (kNN) QSAR method developed in our laboratory (Zheng, W.; Tropsha, A. J. Chem. Inf. Comput. Sci. 2000, 40, 185-194). They were validated using the same training and test sets that were employed in various, mostly 3D-QSAR, investigations published by other authors. We show that for all data sets 2D-QSAR models that use a combination of chirality descriptors with conventional (chirality insensitive) topological descriptors afford better or similar predictive ability as compared to models generated with 3D-QSAR approaches. The results presented in this paper reassure that 2D-QSAR modeling provides a powerful alternative to 3D-QSAR.

Databases, Factual↗

Combinatorial QSAR of ambergris fragrance compounds.

A combinatorial quantitative structure-activity relationships (Combi-QSAR) approach has been developed and applied to a data set of 98 ambergris fragrance compounds with complex stereochemistry. The Combi-QSAR approach explores all possible combinations of different independent descriptor collections and various individual correlation methods to obtain statistically significant models with high internal (for the training set) and external (for the test set) accuracy. Seven different descriptor collections were generated with commercially available MOE, CoMFA, CoMMA, Dragon, VolSurf, and MolconnZ programs; we also included chirality topological descriptors recently developed in our laboratory (Golbraikh, A.; Bonchev, D.; Tropsha, A. J. Chem. Inf. Comput. Sci. 2001, 41, 147-158). CoMMA descriptors were used in combination with MOE descriptors. MolconnZ descriptors were used in combination with chirality descriptors. Each descriptor collection was combined individually with four correlation methods, including k-nearest neighbors (kNN) classification, Support Vector Machines (SVM), decision trees, and binary QSAR, giving rise to 28 different types of QSAR models. Multiple diverse and representative training and test sets were generated by the divisions of the original data set in two. Each model with high values of leave-one-out cross-validated correct classification rate for the training set was subjected to extensive internal and external validation to avoid overfitting and achieve reliable predictive power. Two validation techniques were employed, i.e., the randomization of the target property (in this case, odor intensity) also known as the Y-randomization test and the assessment of external prediction accuracy using test sets. We demonstrate that not every combination of the data modeling technique and the descriptor collection yields a validated and predictive QSAR model. kNN classification in combination with CoMFA descriptors was found to be the best QSAR approach overall since predictive models with correct classification rates for both training and test sets of 0.7 and higher were obtained for all divisions of the ambergris data set into the training and test sets. Many predictive QSAR models were also found using a combination of kNN classification method with other collections of descriptors. The combinatorial QSAR affords automation, computational efficiency, and higher probability of identifying significant QSAR models for experimental data sets than the traditional approaches that rely on a single QSAR method.

Algorithms↗

Combinatorial QSAR modeling of P-glycoprotein substrates.

Quantitative structure-activity (property) relationship (QSAR/QSPR) models are typically generated with a single modeling technique using one type of molecular descriptors. Recently, we have begun to explore a combinatorial QSAR approach which employs various combinations of optimization methods and descriptor types and includes rigorous and consistent model validation (Kovatcheva, A.; Golbraikh, A.; Oloff, S.; Xiao, Y.; Zheng, W.; Wolschann, P.; Buchbauer, G.; Tropsha, A. Combinatorial QSAR of Ambergris Fragrance Compounds. J. Chem. Inf. Comput. Sci. 2004, 44, 582-95). Herein, we have applied this approach to a data set of 195 diverse substrates and nonsubstrates of P-glycoprotein (P-gp) that plays a crucial role in drug resistance. Modeling methods included k-nearest neighbors classification, decision tree, binary QSAR, and support vector machines (SVM). Descriptor sets included molecular connectivity indices, atom pair (AP) descriptors, VolSurf descriptors, and molecular operation environment descriptors. Each descriptor type was used with every QSAR modeling technique; so, in total, 16 combinations of techniques and descriptor types have been considered. Although all combinations resulted in models with a high correct classification rate for the training set (CCR(train)), not all of them had high classification accuracy for the test set (CCR(test)). Thus, predictive models have been generated only for some combinations of the methods and descriptor types, and the best models were obtained using SVM classification with either AP or VolSurf descriptors; they were characterized by CCR(train) = 0.94 and 0.88 and CCR(test) = 0.81 and 0.81, respectively. The combinatorial QSAR approach identified models with higher predictive accuracy than those reported previously for the same data set. We suggest that, in the absence of any universally applicable "one-for-all" QSAR methodology, the combinatorial QSAR approach should become the standard practice in QSPR/QSAR modeling.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

A novel automated lazy learning QSAR (ALL-QSAR) approach: method development, applications, and virtual screening of chemical databases using validated ALL-QSAR models.

A novel automated lazy learning quantitative structure-activity relationship (ALL-QSAR) modeling approach has been developed on the basis of the lazy learning theory. The activity of a test compound is predicted from a locally weighted linear regression model using chemical descriptors and the biological activity of the training set compounds most chemically similar to this test compound. The weights with which training set compounds are included in the regression depend on the similarity of those compounds to a test compound. We have applied the ALL-QSAR method to several experimental chemical data sets including 48 anticonvulsant agents with known ED50 values, 48 dopamine D1-receptor antagonists with known competitive binding affinities (Ki), and a Tetrahymena pyriformis data set containing 250 phenolic compounds with toxicity IGC50 values. When applied to database screening, models developed for anticonvulsant agents identified several known anticonvulsant compounds that were not only absent in the training set but highly chemically dissimilar to the training set compounds. This initial success indicates that ALL-QSAR can be further exploited as a general tool for accurate bioactivity prediction and database screening in drug design and discovery. Because of its local nature, the ALL-QSAR approach appears to be especially well-suited for the development of highly predictive models for the sparse or unevenly distributed data sets.

Database Management Systems↗

Predictive QSAR modeling based on diversity sampling of experimental datasets for the training and test set selection.

One of the most important characteristics of Quantitative Structure Activity Relashionships (QSAR) models is their predictive power. The latter can be defined as the ability of a model to predict accurately the target property (e.g., biological activity) of compounds that were not used for model development. We suggest that this goal can be achieved by rational division of an experimental SAR dataset into the training and test set, which are used for model development and validation, respectively. Given that all compounds are represented by points in multidimensional descriptor space, we argue that training and test sets must satisfy the following criteria: (i) Representative points of the test set must be close to those of the training set; (ii) Representative points of the training set must be close to representative points of the test set; (iii) Training set must be diverse. For quantitative description of these criteria, we use molecular dataset diversity indices introduced recently (Golbraikh, A., J. Chem. Inf. Comput. Sci., 40 (2000) 414-425). For rational division of a dataset into the training and test sets, we use three closely related sphere-exclusion algorithms. Using several experimental datasets, we demonstrate that QSAR models built and validated with our approach have statistically better predictive power than models generated with either random or activity ranking based selection of the training and test sets. We suggest that rational approaches to the selection of training and test sets based on diversity principles should be used routinely in all QSAR modeling research.

Algorithms↗

Rational selection of training and test sets for the development of validated QSAR models.

Quantitative Structure-Activity Relationship (QSAR) models are used increasingly to screen chemical databases and/or virtual chemical libraries for potentially bioactive molecules. These developments emphasize the importance of rigorous model validation to ensure that the models have acceptable predictive power. Using k nearest neighbors (kNN) variable selection QSAR method for the analysis of several datasets, we have demonstrated recently that the widely accepted leave-one-out (LOO) cross-validated R2 (q2) is an inadequate characteristic to assess the predictive ability of the models [Golbraikh, A., Tropsha, A. Beware of q2! J. Mol. Graphics Mod. 20, 269-276, (2002)]. Herein, we provide additional evidence that there exists no correlation between the values of q2 for the training set and accuracy of prediction (R2) for the test set and argue that this observation is a general property of any QSAR model developed with LOO cross-validation. We suggest that external validation using rationally selected training and test sets provides a means to establish a reliable QSAR model. We propose several approaches to the division of experimental datasets into training and test sets and apply them in QSAR studies of 48 functionalized amino acid anticonvulsants and a series of 157 epipodophyllotoxin derivatives with antitumor activity. We formulate a set of general criteria for the evaluation of predictive power of QSAR models.

Algorithms↗