Search PubMed⌕ Search

Biomedical subjects

Alexander Tropsha

Publications and source records attributed to Alexander Tropsha.

At least 19 recordsLinked to original sources

Conserved Filovirus Proteins as Targets of Broad-Spectrum Antivirals.

Filoviruses are enveloped, non-segmented, negative-strand RNA viruses belonging to the Filoviridae family, which includes five genera: Ebolavirus, Marburgvirus, Cuevavirus, Striavirus, and Thamnovirus. Members of this family cause severe and, often, fatal hemorrhagic fevers in humans and non-human primates, with high mortality rates. To date, only two filoviruses, Ebola virus (EBOV) and Marburg virus (MARV), are known to infect humans and are listed as priority pathogens by the World Health Organization due to their potential for re-emergence and the current lack of effective vaccines and antiviral treatments. In this study, we identify and characterize conserved binding sites within key filoviral proteins to support the development of broad-spectrum, direct-acting antiviral agents. We validated the significance of these conserved regions for drug discovery using existing experimental data. Our analysis revealed notably high sequence similarity among proteins from filoviruses capable of infecting humans (EBOV, TAFV, BDBV, SUDV, MARV, and RAVV) compared to those from non-zoonotic species, with the highest conservation observed in the L and VP40 proteins-both critical for viral genome transcription and replication. Furthermore, we compiled and analyzed available experimental data on known antiviral compounds targeting these proteins, identifying several agents with cross-filovirus activity, including Galidesivir, Remdesivir, and Favipiravir. The integrated approach described here-combining sequence and structural conservation analysis with chemical structure and antiviral activity data-demonstrates a strategy that could be extended to the development of broad-spectrum therapeutics across multiple viral families.

Broad Spectrum Antiviral↗

QSAR modeling of human serum protein binding with several modeling techniques utilizing structure-information representation.

Four modeling techniques, using topological descriptors to represent molecular structure, were employed to produce models of human serum protein binding (% bound) on a data set of 1008 experimental values, carefully screened from publicly available sources. To our knowledge, this data is the largest set on human serum protein binding reported for QSAR modeling. The data was partitioned into a training set of 808 compounds and an external validation test set of 200 compounds. Partitioning was accomplished by clustering the compounds in a structure descriptor space so that random sampling of 20% of the whole data set produced an external test set that is a good representative of the training set with respect to both structure and protein binding values. The four modeling techniques include multiple linear regression (MLR), artificial neural networks (ANN), k-nearest neighbors (kNN), and support vector machines (SVM). With the exception of the MLR model, the ANN, kNN, and SVM QSARs were ensemble models. Training set correlation coefficients and mean absolute error ranged from r2=0.90 and MAE=7.6 for ANN to r2=0.61 and MAE=16.2 for MLR. Prediction results from the validation set yielded correlation coefficients and mean absolute errors which ranged from r2=0.70 and MAE=14.1 for ANN to a low of r2=0.59 and MAE=18.3 for the SVM model. Structure descriptors that contribute significantly to the models are discussed and compared with those found in other published models. For the ANN model, structure descriptor trends with respect to their affects on predicted protein binding can assist the chemist in structure modification during the drug design process.

Blood Proteins↗

Novel ligands for the human histamine H1 receptor: synthesis, pharmacology, and comparative molecular field analysis studies of 2-dimethylamino-5-(6)-phenyl-1,2,3,4-tetrahydronaphthalenes.

This paper reports the synthesis of a novel series of (+/-)-2-dimethylamino- 5- and 6-phenyl-1,2,3,4-tetrahydronaphthalene derivatives (5- and 6-APTs), and, corresponding affinity, functional activity, and, molecular modeling studies with regard to drug design targeting the human histamine H1 receptor. The 5-APTs have 2- to 4-fold higher H1 receptor affinity than the endogenous agonist histamine. The chemical nature of a meta-substituent on the 5-APT pendant phenyl moiety does not significantly affect H1 affinity. In contrast, analogous meta-substitution for the 6-APTs increases H1 affinity up to 100-fold. The new APTs do not activate H1 receptor-linked intracellular signaling and apparently are competitive H1 antagonists. A new model that establishes structural parameters for binding to the human H1 receptor by APTs and other ligands was developed using 3-D QSAR (CoMFA). The model predicts H1 ligand binding with a higher degree of external predictability compared to a previously reported model. The APTs also were examined for activity at human serotonin 5-HT2A and 5-HT2C receptors, which are phylogenetically closely related to the H1 receptor. 5-APT and m-Cl-6-APT were identified as novel agonists that selectively activate 5-HT2C receptors. It is concluded that the lipophilic (brain-penetrating) APT molecular scaffold may have pharmacotherapeutic potential in neuropsychiatric diseases.

Animals↗

Development of quantitative structure-binding affinity relationship models based on novel geometrical chemical descriptors of the protein-ligand interfaces.

Novel geometrical chemical descriptors have been derived on the basis of the computational geometry of protein-ligand interfaces and Pauling atomic electronegativities (EN). Delaunay tessellation has been applied to a diverse set of 517 X-ray characterized protein-ligand complexes yielding a unique collection of interfacial nearest neighbor atomic quadruplets for each complex. Each quadruplet composition was characterized by a single descriptor calculated as the sum of the EN values for the four participating atom types. We termed these simple descriptors generated from atomic EN values and derived with the Delaunay Tessellation the ENTess descriptors and used them in the variable selection k-nearest neighbor quantitative structure-binding affinity relationship (QSBR) studies of 264 diverse protein-ligand complexes with known binding constants. Twenty-four complexes with chemically dissimilar ligands were set aside as an independent validation set, and the remaining dataset of 240 complexes was divided into multiple training and test sets. The best models were characterized by the leave-one-out cross-validated correlation coefficient q(2) as high as 0.66 for the training set and the correlation coefficient R(2) as high as 0.83 for the test set. The high predictive power of these models was confirmed independently by applying them to the validation set of 24 complexes yielding R(2) as high as 0.85. We conclude that QSBR models built with the ENTess descriptors can be instrumental for predicting the binding affinity of receptor-ligand complexes.

Computer Simulation↗

Structure-based function inference using protein family-specific fingerprints.

We describe a method to assign a protein structure to a functional family using family-specific fingerprints. Fingerprints represent amino acid packing patterns that occur in most members of a family but are rare in the background, a nonredundant subset of PDB; their information is additional to sequence alignments, sequence patterns, structural superposition, and active-site templates. Fingerprints were derived for 120 families in SCOP using Frequent Subgraph Mining. For a new structure, all occurrences of these family-specific fingerprints may be found by a fast algorithm for subgraph isomorphism; the structure can then be assigned to a family with a confidence value derived from the number of fingerprints found and their distribution in background proteins. In validation experiments, we infer the function of new members added to SCOP families and we discriminate between structurally similar, but functionally divergent TIM barrel families. We then apply our method to predict function for several structural genomics proteins, including orphan structures. Some predictions have been corroborated by other computational methods and some validated by subsequent functional characterization.

Bacterial Proteins↗

Application of validated QSAR models of D1 dopaminergic antagonists for database mining.

Rigorously validated quantitative structure-activity relationship (QSAR) models have been developed for 48 antagonists of the dopamine D1 receptor and applied to mining chemical datasets to discover novel potential antagonists. Several QSAR methods have been employed, including comparative molecular field analysis (CoMFA), simulated annealing-partial least squares (SA-PLS), k-nearest neighbor (kNN), and support vector machines (SVM). With the exception of CoMFA, these approaches employed 2D topological descriptors generated with the MolConnZ software package (EduSoft, LLC. MolconnZ, version 4.05; http://www.eslc.vabiotech.com/ [4.05], 2003). The original dataset was split into training and test sets to allow for external validation of each training set model. The resulting models were characterized by cross-validated R2 (q2) for the training set and predictive R2 values for the test set of (q2/R2) 0.51/0.47 for CoMFA, 0.7/0.76 for kNN, R2 for the training and test sets of 0.74/0.71 for SVM, and training set fitness and test set R2 values of 0.68/0.63 for SA-PLS. Validated QSAR models with R2 > 0.7, (i.e., kNN and SVM) were used to mine three publicly available chemical databases: the National Cancer Institute (NCI) database of ca. 250,000 compounds, the Maybridge Database of ca. 56,000 compounds, and the ChemDiv Database of ca. 450,000 compounds. These searches resulted in only 54 consensus hits (i.e., predicted active by all models); five of them were previously characterized as dopamine D1 ligands, but were not present in the original dataset. A small fraction of the purported D1 ligands did not contain a catechol ring found in all known dopamine full agonist ligands, suggesting that they may be novel structural antagonist leads. This study illustrates that the combined application of predictive QSAR modeling and database mining may provide an important avenue for rational computer-aided drug discovery.

Databases, Factual↗

Quantitative structure-activity relationship analysis of pyridinone HIV-1 reverse transcriptase inhibitors using the k nearest neighbor method and QSAR-based database mining.

We have developed quantitative structure-activity relationship (QSAR) models for 44 non-nucleoside HIV-1 reverse transcriptase inhibitors (NNRTIs) of the pyridinone derivative type. The k nearest neighbor (kNN) variable selection approach was used. This method utilizes multiple descriptors such as molecular connectivity indices, which are derived from two-dimensional molecular topology. The modeling process entailed extensive validation including the randomization of the target property (Y-randomization) test and the division of the dataset into multiple training and test sets to establish the external predictive power of the training set models. QSAR models with high internal and external accuracy were generated, with leave-one-out cross-validated R2 (q2) values ranging between 0.5 and 0.8 for the training sets and R2 values exceeding 0.6 for the test sets. The best models with the highest internal and external predictive power were used to search the National Cancer Institute database. Derivatives of the pyrazolo[3,4-d]pyrimidine and phenothiazine type were identified as promising novel NNRTIs leads. Several candidates were docked into the binding pocket of nevirapine with the AutoDock (version 3.0) software. Docking results suggested that these types of compounds could be binding in the NNRTI binding site in a similar mode to a known non-nucleoside inhibitor nevirapine.

Database Management Systems↗

Evaluation of the relative stability of liganded versus ligand-free protein conformations using Simplicial Neighborhood Analysis of Protein Packing (SNAPP) method.

Many proteins change their conformation upon ligand binding. For instance, bacterial periplasmic binding proteins (bPBPs), which transport nutrients into the cytoplasm, generally consist of two globular domains connected by strands, forming a hinge. During ligand binding, hinge motion changes the conformation from the open to the closed form. Both forms can be crystallized without a ligand, suggesting that the energy difference between them is small. We applied Simplicial Neighborhood Analysis of Protein Packing (SNAPP) as a method to evaluate the relative stability of open and closed forms in bPBPs. Using united residue representation of amino acids, SNAPP performs Delaunay tessellation of the protein, producing an aggregate of space-filling, irregular tetrahedra with nearest neighbor residues at the vertices. The SNAPP statistical scoring function is derived from log-likelihood scores for all possible quadruplet compositions of amino acids found in a representative subset of the Protein Data Bank, and the sum of the scores for a given protein provides the total SNAPP score. Results of scoring for bPBPs suggest that in most cases, the unliganded form is more stable than the liganded form, and this conclusion is corroborated by similar observations of other proteins undergoing conformation changes upon binding their ligands. The results of these studies suggest that the SNAPP method can be used to predict the relative stability of accessible protein conformations. Furthermore, the SNAPP method allows delineation of the role of individual residues in protein stabilization, thereby providing new testable hypotheses for rational site-directed mutagenesis in the context of protein engineering.

Amino Acids↗

A conformational change in heparan sulfate 3-O-sulfotransferase-1 is induced by binding to heparan sulfate.

The 3-O-sulfation of glucosamine by heparan sulfate 3-O-sulfotransferase-1 (3-OST-1) is a key modification step during the biosynthesis of anticoagulant heparan sulfate (HS). In this paper, we present evidence of a conformational change that occurs in 3-OST-1 upon binding to heparan sulfate. The intrinsic fluorescence of 3-OST-1 was increased in the presence of HS, suggesting a conformational change. This apparent conformational change was further investigated using differential chemical modification of 3-OST-1 to measure the solvent accessibility of the lysine residues. 3-OST-1 was treated with acetic anhydride in either the presence or absence of HS using both acetic anhydride and hexadeuterioacetic anhydride under nondenaturing and denaturing conditions, respectively. The relative reactivity of the lysine residues to acetylation and [2H] acetylation in the presence or absence of HS was analyzed by measuring the ratio of acetylated and deuterioacetylated peptides using matrix-assisted laser desorption ionization mass spectrometry. The solvent accessibilities of the lysine residues were altered differentially depending on their location. In particular, we observed a group of lysine residues in the C-terminus of 3-OST-1 that become more solvent accessible when 3-OST-1 binds to HS. This observation indicates that a conformational change could be occurring during substrate binding. A truncated mutant of 3-OST-1 that lacked this C-terminal region was expressed and found to exhibit a 200-fold reduction in sulfotransferase activity. The results from this study will contribute to our understanding of the interactions between 3-OSTs and HS.

Amino Acid Substitution↗

Application of predictive QSAR models to database mining: identification and experimental validation of novel anticonvulsant compounds.

We have developed a drug discovery strategy that employs variable selection quantitative structure-activity relationship (QSAR) models for chemical database mining. The approach starts with the development of rigorously validated QSAR models obtained with the variable selection k nearest neighbor (kNN) method (or, in principle, with any other robust model-building technique). Model validation is based on several statistical criteria, including the randomization of the target property (Y-randomization), independent assessment of the training set model's predictive power using external test sets, and the establishment of the model's applicability domain. All successful models are employed in database mining concurrently; in each case, only variables selected as a result of model building (termed descriptor pharmacophore) are used in chemical similarity searches comparing active compounds of the training set (queries) with those in chemical databases. Specific biological activity (characteristic of the training set compounds) of external database entries found to be within a predefined similarity threshold of the training set molecules is predicted on the basis of the validated QSAR models using the applicability domain criteria. Compounds judged to have high predicted activities by all or the majority of all models are considered as consensus hits. We report on the application of this computational strategy for the first time for the discovery of anticonvulsant agents in the Maybridge and National Cancer Institute (NCI) databases containing ca. 250,000 compounds combined. Forty-eight anticonvulsant agents of the functionalized amino acid (FAA) series were used to build kNN variable selection QSAR models. The 10 best models were applied to mining chemical databases, and 22 compounds were selected as consensus hits. Nine compounds were synthesized and tested at the NIH Epilepsy Branch, Rockville, MD using the same biological test that was employed to assess the anticonvulsant activity of the training set compounds; of these nine, four were exact database hits and five were derived from the hits by minor chemical modifications. Seven of these nine compounds were confirmed to be active, indicating an exceptionally high hit rate. The approach described in this report can be used as a general rational drug discovery tool.

Amides↗

Quantitative structure-pharmacokinetic parameters relationships (QSPKR) analysis of antimicrobial agents in humans using simulated annealing k-nearest-neighbor and partial least-square analysis methods.

We have developed quantitative structure-pharmacokinetic parameters relationship (QSPKR) models using k-nearest-neighbor (k-NN) and partial least-square (PLS) methods to predict the volume of distribution at steady state (Vss) and clearance (CL) of 44 antimicrobial agents in humans. The performance of QSPKR was determined by the values of the internal leave-one-out, crossvalidated coefficient of determination q(2) for the training set and external predictive r(2) for the test set. The best simulated annealing (SA)-kNN model was highly predictive for Vss and provided q(2) and r(2) values of 0.93 and 0.80, respectively. For all compounds, the model produced average fold error values for Vss of 1.00 and for 93% of the compounds provided predictions that were within a twofold error of actual values. The best SA-kNN model for prediction of CL yielded q(2) and r(2) values of 0.77 and 0.94, respectively, and had an average fold rror of 1.05. Use of PLS methods resulted in inferior QSPKR models. The SA-kNN QSPKR approach has utility in drug discovery and development in the identification of compounds that possess appropriate pharmacokinetic characteristics in humans, and will assist in the selection of a suitable starting dose for Phase I, first-time-in-man studies.

Anti-Infective Agents↗

Modeling of p38 mitogen-activated protein kinase inhibitors using the Catalyst HypoGen and k-nearest neighbor QSAR methods.

We have employed in parallel the Catalyst HypoGen pharmacophore modeling approach and the variable selection k-nearest neighbor quantitative structure-activity relationship (kNN QSAR) method to model a diverse data set of p38 mitogen-activated protein (MAP) kinase inhibitors. The HypoGen pharmacophore model, developed from a novel automated training set selection protocol, identified chemical functional features that were characteristic of the active compounds and differentiated the active from the inactive inhibitors. The kNN QSAR modeling employed topological descriptors and afforded predictive QSAR models with consistently high values of both leave-one-out cross-validated R2 for the training set and predictive R2 for the test set. The results of both modeling approaches were sensitive to the selection of the training and test sets used for model development and validation. The resulting Catalyst pharmacophore and kNN QSAR models can be used concurrently for rapid virtual screening of chemical databases to identify novel p38 MAP kinase inhibitors.

Algorithms↗

Three new consensus QSAR models for the prediction of Ames genotoxicity.

Three QSAR methods, artificial neural net (ANN), k-nearest neighbors (kNN), and Decision Forest (DF), were applied to 3363 diverse compounds tested for their Ames genotoxicity. The ratio of mutagens to non-mutagens was 60/40 for this dataset. This group of compounds includes >300 therapeutic drugs. All models were developed using the same initial set of 148 topological indices: molecular connectivity chi indices and electrotopological state indices (atom-type, bond-type and group-type E-state), as well as binary indicators. While previous studies have found logP to be a determining factor in genotoxicity, it was not found to be important by any modeling method employed in this study. The three models yielded an average training/test concordance value of 88%, with a low percentage of false positives and false negatives. External validation testing on 400 compounds not used for QSAR model development gave an average concordance of 82%. This value increased to 92% upon removal of less reliable outcomes, as determined by a reliability criterion used within each model. The ANN model showed the best performance in predicting drug compounds, yielding 97% concordance (34/35 drugs) after the removal of less reliable predictions. The appreciable commonality found among the top 10 ranked descriptors from each model is of particular interest because of the diversity in the learning algorithms and descriptor selection techniques employed in this study. Forty percent of the most important descriptors in any one model are found in one or two other models. Fourteen of the most important descriptors relate directly to known toxicophores involved in potent genotoxic responses in Salmonella typhimurium. A comparison of the validation results with those of MULTICASE and DEREK indicated that the new models presented in this work perform substantially better than the former models in predicting genotoxicity of therapeutic drugs. Substantially higher specificity was achieved with these new models as compared with MULTICASE or DEREK with comparable sensitivities among all models.

Algorithms↗

Autoimmunity is triggered by cPR-3(105-201), a protein complementary to human autoantigen proteinase-3.

It remains unclear how and why autoimmunity occurs. Here we show evidence for a previously unrecognized and possibly general mechanism of autoimmunity. This new finding was discovered serendipitously using material from patients with inflammatory vascular disease caused by antineutrophil cytoplasmic autoantibodies (ANCA) with specificity for proteinase-3 (PR-3). Such patients harbor not only antibodies to the autoantigen (PR-3), but also antibodies to a peptide translated from the antisense DNA strand of PR-3 (complementary PR-3, cPR-3) or to a mimic of this peptide. Immunization of mice with the middle region of cPR-3 resulted in production of antibodies not only to cPR-3, but also to the immunogen's sense peptide counterpart, PR-3. Both human and mouse antibodies to PR-3 and cPR-3 bound to each other, indicating idiotypic relationships. These findings indicate that autoimmunity can be initiated through an immune response against a peptide that is antisense or complementary to the autoantigen, which then induces anti-idiotypic antibodies (autoantibodies) that cross-react with the autoantigen.

Amino Acid Sequence↗

Development of a four-body statistical pseudo-potential to discriminate native from non-native protein conformations.

MOTIVATION: Most scoring functions used in protein fold recognition employ two-body (pseudo) potential energies. The use of higher-order terms may improve the performance of current algorithms. METHODS: Proteins are represented by the side chain centroids of amino acids. Delaunay tessellation of this representation defines all sets of nearest neighbor quadruplets of amino acids. Four-body contact scoring function (log likelihoods of residue quadruplet compositions) is derived by the analysis of a diverse set of proteins with known structures. A test protein is characterized by the total score calculated as the sum of the individual log likelihoods of composing amino acid quadruplets. RESULTS: The scoring function distinguishes native from partially unfolded or deliberately misfolded structures. It also discriminates between pre- and post-transition state and native structures in the folding simulations trajectory of Chymotrypsin Inhibitor 2 (CI2).

Algorithms↗

Development and validation of k-nearest-neighbor QSPR models of metabolic stability of drug candidates.

Computational ADME (absorption, distribution, metabolism, and excretion) models may be used early in the drug discovery process in order to flag drug candidates with potentially problematic ADME profiles. We report the development, validation, and application of quantitative structure-property relationship (QSPR) models of metabolic turnover rate for compounds in human S9 homogenate. Biological data were obtained from uniform bioassays of 631 diverse chemicals proprietary to GlaxoSmithKline (GSK). The models were built with topological molecular descriptors such as molecular connectivity indices or atom pairs using the k-nearest neighbor variable selection optimization method developed at the University of North Carolina (Zheng, W.; Tropsha, A. A novel variable selection QSAR approach based on the k-nearest neighbor principle. J. Chem. Inf. Comput. Sci., 2000, 40, 185-194.). For the purpose of validation, the whole data set was divided into training and test sets. The training set QSPR models were characterized by high internal accuracy with leave-one-out cross-validated R(2) (q(2)) values ranging between 0.5 and 0.6. The test set compounds were correctly classified as stable or unstable in S9 assay with an accuracy above 85%. These models were additionally validated by in silico metabolic stability screening of 107 new chemicals under development in several drug discovery programs at GSK. One representative model generated with MolConnZ descriptors predicted 40 compounds to be metabolically stable (turnover rate less than 25%), and 33 of them were indeed found to be stable experimentally. This success (83% concordance) in correctly picking chemicals that are metabolically stable in the human S9 homogenate spells a rapid, computational screen for generating components of the ADME profile in a drug discovery process.

Algorithms↗