Biomedical subjects
Johann Gasteiger
Publications and source records attributed to Johann Gasteiger.
A novel workflow for the inverse QSPR problem using multiobjective optimization.
A workflow for the inverse quantitative structure-property relationship (QSPR) problem is reported in this paper for the de novo design of novel chemical entities (NCE) in silico through the application of existing QSPR models to calculate multiple objectives, including prediction confidence measures, to be optimized during the de novo design process. Two physical property datasets are applied as case studies of the inverse QSPR workflow (IQW): mean molecular polarizability and aqueous solubility. The case studies demonstrate the optimization of molecular structures to within a property range of interest; the optimized structures are then validated against QSPR models that are generated from sets of alternative descriptors to those used in the IQW. The paper concludes with a discussion of the results from the case studies.
Development of a structural model for NF-kappaB inhibition of sesquiterpene lactones using self-organizing neural networks.
A variety of sesquiterpene lactones (SLs) possess considerable anti-inflammatory activity. Several studies have shown that they exert this effect in part by inhibiting the activation of the transcription factor NF-kappaB. In the present study we elaborated on the investigation of a data set of 103 structurally diverse SLs for which we had previously developed several different QSAR equations dependent on the skeletal type. Use of 3D structure descriptors resulted in a single model for the entire data set. In particular, local radial distribution functions (L-RDF) were used that centered on the methylene-carbonyl substructure believed to be the site of attack of cysteine-38 of the p65/NF-kappaB subunit. The model was developed by using a counterpropagation neural network (CPGNN), attesting to the power of this method for establishing structure-activity-relationships. The investigations shed more light onto the influence of the chemical structure on NF-kappaB inhibitory activity.
Comparative molecular surface analysis (CoMSA) for virtual combinatorial library screening of styrylquinoline HIV-1 blocking agents.
We used comparative molecular surface analysis to design molecules for the synthesis as part of the search for new HIV-1 integrase inhibitors. We analyzed the virtual combinatorial library (VCL) constituted from various moieties of styrylquinoline and styrylquinazoline inhibitors. Since imines can be applied in a strategy of dynamic combinatorial chemistry (DCC), we also tested similar compounds in which the -C=N- or -N=C- linker connected the heteroaromatic and aromatic moieties. We then used principal component analysis (PCA) or self-organizing maps (SOM), namely, the Kohonen neural networks to obtain a clustering plot analyzing the diversity of the VCL formed. Previously synthesized compounds of known activity, used as molecular probes, were projected onto this plot, which provided a set of promising virtual drugs. Moreover, we further modified the above mentioned VCL to include the single bond linker -C-N- or -N-C-. This allowed increasing compound stability but expanded also the diversity between the available molecular probes and virtual targets. The application of the CoMSA with SOM indicated important differences between such compounds and active molecular probes. We synthesized such compounds to verify the computational predictions.
Chemoinformatics: a new field with a long tradition.
Chemoinformatics is the application of informatics methods to solve chemical problems. Although this term was introduced only a few years ago, this field has a long history with its roots going back more than 40 years. Work on chemical structure representation and searching, quantitative structure-activity relationships, chemometrics, molecular modeling as well as computer-assisted structure elucidation and synthesis design was initiated in the 1960s. These different origins have now merged into a discipline of its own that is in full bloom. All areas of chemistry from analytical chemistry to drug design can benefit from chemoinformatics methods. And there are still many challenging chemical problems waiting for solutions through the further development of chemoinformatics.
The de novo design of median molecules within a property range of interest.
In this paper an application is presented of the median molecule workflow to the de novo design of novel molecular entities with a property profile of interest. Median molecules are structures that are optimised to be similar to a set of existing molecules of interest as an approach for lead exploration and hopping. An overview of this workflow is provided together with an example of an instance using the similarity to camphor and menthol as objectives. The methodology of the experiments is defined and the workflow is applied to designing novel molecules for two physical property datasets: mean molecular polarisability and aqueous solubility. This paper concludes with a discussion of the characteristics of this method.
Virtual computational chemistry laboratory--design and description.
Internet technology offers an excellent opportunity for the development of tools by the cooperative effort of various groups and institutions. We have developed a multi-platform software system, Virtual Computational Chemistry Laboratory, http://www.vcclab.org, allowing the computational chemist to perform a comprehensive series of molecular indices/properties calculations and data analysis. The implemented software is based on a three-tier architecture that is one of the standard technologies to provide client-server services on the Internet. The developed software includes several popular programs, including the indices generation program, DRAGON, a 3D structure generator, CORINA, a program to predict lipophilicity and aqueous solubility of chemicals, ALOGPS and others. All these programs are running at the host institutes located in five countries over Europe. In this article we review the main features and statistics of the developed system that can be used as a prototype for academic and industry models.
Sesquiterpene lactone-based classification of three Asteraceae tribes: a study based on self-organizing neural networks applied to chemosystematics.
This work describes an application of artificial neural networks on a small data set of sesquiterpene lactones (STLs) of three tribes of the family Asteraceae. Structurally different types of representative STLs from seven subtribes of the tribes Eupatorieae, Heliantheae and Vernonieae were selected as input data for self-organizing neural networks. Encoding the 3D molecular structures of STLs and their projection onto Kohonen maps allowed the classification of Asteraceae into tribes and subtribes. This approach allowed the evaluation of structural similarities among different sets of 3D structures of sesquiterpene lactones and their correlation with the current taxonomic classification of the family. Predictions of the occurrence of STLs from a plant species according to the taxa they belong to were also performed by the networks. The methodology used in this work can be applied to chemosystematic or chemotaxonomic studies of Asteraceae.
A quantitative structure--activity relationship model for the intrinsic activity of uncouplers of oxidative phosphorylation.
A quantitative structure-activity relationship (QSAR) has been derived for the prediction of the activity of phenols in uncoupling oxidative and photophosphorylation. Twenty-one compounds with experimental data for uncoupling activity as well as for the acid dissociation constant, pKa, and for partitioning constants of the neutral and the charged species into model membranes were analyzed. From these measured data, the effective concentration in the membrane was derived, which allowed the study of the intrinsic activity of uncouplers within the membrane. A linear regression model for the intrinsic activity could be established using the following three descriptors: solvation free energies of the anions, an estimate for heterodimer formation describing transport processes, and pKa values describing the speciation of the phenols. In a next step, the aqueous effect concentrations were modeled by combining the model for the intrinsic uncoupling activity with descriptors accounting for the uptake into membranes. Results obtained with experimental membrane-water partitioning data were compared with the results obtained with experimental octanol-water partition coefficients, log Kow, and with calculated log Kow values. The properties of these different measures of lipophilicity were critically discussed.
Enabling the exploration of biochemical pathways.
The Biochemical Pathways Wall Chart (http://www.expasy.org/tools/pathways/ref.1) has been converted into a molecule and reaction database. Major features of this database are that each molecule is represented by lists of all atoms and bonds (as connection tables), and in the reactions the reaction centre, the atoms and bonds directly involved in the bond rearrangement process, are marked. The information in the database has been enriched by a set of diverse 3D structure conformations generated by the programs CORINA and ROTATE. The web-based structure and reaction retrieval system C@ROL provides a wide range of search methods to mine this rich database. The database is accessible at http://www2.chemie.uni-erlangen.de/services/biopath/index.html and http://www.mol-net.de/databases/biopath.html .
Linear and nonlinear functions on modeling of aqueous solubility of organic compounds by two structure representation methods.
Several quantitative models for the prediction of aqueous solubility of organic compounds were developed based on a diverse dataset with 2084 compounds by using multi-linear regression analysis and backpropagation neural networks. The compounds were described by two different structure representation methods: (1) with 18 topological descriptors; and (2) with 32 radial distribution function codes representing the 3D structure of a molecule and eight additional descriptors. The dataset was divided into a training and a test set based on Kohonen's self-organizing neural network. Good prediction results were obtained for backpropagation neural network models: with 18 topological descriptors, for the 936 compounds in the test set, a correlation coefficient of 0.92, and a standard deviation of 0.62 were achieved; with 3D descriptors, for the 866 compounds in the test set, a correlation coefficient of 0.90, and a standard deviation of 0.73 were achieved. The models were also tested by using another dataset, and the relationship of the two datasets was examined by Kohonen's self-organizing neural network.
Physicochemical effects in the representation of molecular structures for drug designing.
After the identification of a biological target, drug design is to analyze the relationships between the structure of potential ligands and their biological activity. A hierarchy of structure representation is presented here considering either the constitution of a molecule, its 3D structure, or the molecular surface. At each level, a variety of physicochemical effects can be accounted for. Furthermore, the special requirements of learning algorithm, such as neural networks, are taken into consideration. Application to problems from combinatorial chemistry, lead identification, high-throughput screening, and prediction of ADME-Tox properties are given.
Use of the Kohonen neural network for rapid screening of ex vivo anti-HIV activity of styrylquinolines.
Using the Kohonen neural network, the electrostatic potentials on the molecular surfaces of 14 styrylquinoline derivatives were drawn as comparative two-dimensional maps and compared with their known human immunodeficiency virus (HIV)-1 replication blocking potency in cells. A feature of the potential map was discovered to be related with the HIV-1 blocking activity and was used to unmask the activity of further five analogues, previously described but whose cytotoxicity precluded an estimation of their activity, and to predict the activity of 10 new compounds while the experimental data were unknown. The measurements performed later turned out to agree with the predictions.
Prediction of 1H NMR chemical shifts using neural networks.
Counterpropagation neural networks were applied to the fast prediction of 1H NMR chemical shifts of CHn groups in organic compounds. The training set consisted of 744 examples of protons that were represented by physicochemical, topological, and geometric descriptors. The selection of descriptors was performed by genetic algorithms, and the models obtained were compared to those containing all the descriptors. The best models yielded very good predictions for an independent prediction set of 259 cases (mean absolute error for whole set, 0.25 ppm; mean absolute error for 90% of cases, 0.19 ppm) and for application cases consisting of four natural products recently described. Some stereochemical effects could be correctly predicted. A useful feature of the system resides in its ability to be retrained with a specific data set of compounds if improved predictions for related structures are required.
Prediction of enantiomeric selectivity in chromatography. Application of conformation-dependent and conformation-independent descriptors of molecular chirality.
In order to process molecular chirality by computational methods and to obtain predictions for properties that are influenced by chirality, a fixed-length conformation-dependent chirality code is introduced. The code consists of a set of molecular descriptors representing the chirality of a 3D molecular structure. It includes information about molecular geometry and atomic properties, and can distinguish between enantiomers, even if chirality does not result from chiral centers. The new molecular transform was applied to two datasets of chiral compounds, each of them containing pairs of enantiomers that had been separated by chiral chromatography. The elution order within each pair of isomers was predicted by means of Kohonen neural networks (NN) using the chirality codes as input. A previously described conformation-independent chirality code was also applied and the results were compared. In both applications clustering of the two classes of enantiomers (first eluted and last eluted enantiomers) could be successfully achieved by NN and accurate predictions could be obtained for independent test sets. The chirality code described here has a potential for a broad range of applications from stereoselective reactions to analytical chemistry and to the study of biological activity of chiral compounds.
Prediction of enantiomeric excess in a combinatorial library of catalytic enantioselective reactions.
A quantitative structure-enantioselectivity relationship was established for a combinatorial library of enantioselective reactions performed by addition of diethyl zinc to benzaldehyde. Chiral catalysts and additives were encoded by their chirality codes and presented as input to neural networks. The networks were trained to predict the enantiomeric excess. With independent test sets, predictions of enantiomeric excess could be made with an average error as low as 6% ee. Multilinear regression, perceptrons, and support vector machines were also evaluated as modeling tools. The method is of interest for the computer-aided design of combinatorial libraries involving chiral compounds or enantioselective reactions. This is the first example of a quantitative structure-property relationship based on chirality codes.
Design and analysis of a combinatorial library of HEPT analogues: comparison of selection methodologies and inspection of the actually covered chemical space.
A large virtual library of 125 396 HEPT analogues, built by combining all fragments present in the published 180-compound HEPT family, has been studied in terms of diversity criteria and the goodness of the 11 available standard diversity selection methods analyzed. All the algorithms under study, except Cell-based Density, have rank above a random selection of compounds, with Optimum and Standard Deviation based Binning and Cell-based Fraction algorithms being the best choices. Furthermore, analysis of the actually tested compounds has been performed to compare the traditional drug discovery methodology versus a rational selection of combinatorial libraries approach.
Prediction of aqueous solubility of organic compounds based on a 3D structure representation.
Two quantitative models for the prediction of aqueous solubility of 1293 organic compounds were developed by a Multilinear Regression (MLR) analysis and a Back-Propagation (BPG) neural network. The molecules were described by a set of 32 values of a Radial Distribution Function (RDF) code representing the 3D structure and eight additional descriptors. The 1293 compounds were divided into a training set of 797 compounds and a test set of 496 compounds based on a Kohonen self-organizing neural network map. The obtained models show a good predictive power: for the test set, a correlation coefficient of 0.96 and a standard deviation of 0.59 were achieved by the back-propagation neural network approach.