Search PubMed⌕ Search

Biomedical subjects

Andreas Bender

Publications and source records attributed to Andreas Bender.

26 records · Page 2Linked to original sources

General melting point prediction based on a diverse compound data set and artificial neural networks.

We report the development of a robust and general model for the prediction of melting points. It is based on a diverse data set of 4173 compounds and employs a large number of 2D and 3D descriptors to capture molecular physicochemical and other graph-based properties. Dimensionality reduction is performed by principal component analysis, while a fully connected feed-forward back-propagation artificial neural network is employed for model generation. The melting point is a fundamental physicochemical property of a molecule that is controlled by both single-molecule properties and intermolecular interactions due to packing in the solid state. Thus, it is difficult to predict, and previously only melting point models for clearly defined and smaller compound sets have been developed. Here we derive the first general model that covers a comparatively large and relevant part of organic chemical space. The final model is based on 2D descriptors, which are found to contain more relevant information than the 3D descriptors calculated. Internal random validation of the model achieves a correlation coefficient of R(2) = 0.661 with an average absolute error of 37.6 degrees C. The model is internally consistent with a correlation coefficient of the test set of Q(2) = 0.658 (average absolute error 38.2 degrees C) and a correlation coefficient of the internal validation set of Q(2) = 0.645 (average absolute error 39.8 degrees C). Additional validation was performed on an external drug data set consisting of 277 compounds. On this external data set a correlation coefficient of Q(2) = 0.662 (average absolute error 32.6 degrees C) was achieved, showing ability of the model to generalize. Compared to an earlier model for the prediction of melting points of druglike compounds our model exhibits slightly improved performance, despite the much larger chemical space covered. The remaining model error is due to molecular properties that are not captured using single-molecule based descriptors, namely both inter- and intramolecular interactions and crystal packing, for which examples of and reasons for outliers are given.

Journal Article↗

A discussion of measures of enrichment in virtual screening: comparing the information content of descriptors with increasing levels of sophistication.

We have performed virtual screening using some very simple features, by employing the number of atoms per element as molecular descriptors but without regard to any structural information whatsoever. Surprisingly, these atom counts are able to outperform virtual-affinity-based fingerprints and Unity fingerprints in some activity classes. Although molecular weight and other biases were known in target-based virtual screening settings (docking), we report the effect of using very simple descriptors for ligand-based virtual screening, by using clearly defined biological targets and employing a large data set (>100,000 compounds) containing multiple (11) activity classes. Structure-unaware atom count vectors as descriptors in combination with the Euclidean distance measure are able to achieve "enrichment factors" over random selection of around 4 (depending on the particular class of active compounds), putting the enrichment factors reported for more sophisticated virtual screening methods in a different light. They are also able to retrieve active compounds with novel scaffolds instead of merely the expected structural analogues. The added value of many currently used virtual screening methods (calculated as enrichment factors) drops down to a factor of between 1 and 2, instead of often reported double-digit figures. The observed effect is much less profound for simple descriptors such as molecular weight and is only present in cases of atypical (larger) ligands. The current state of virtual screening is not as sophisticated as might be expected, which is due to descriptors still not being able to capture structural properties relevant to binding. This fact can partly be explained by highly nonlinear structure-activity relationships, which represent a severe limitation of the "similar property principle" in the context of bioactivity.

Computer Simulation↗

Analysis of activity space by fragment fingerprints, 2D descriptors, and multitarget dependent transformation of 2D descriptors.

The effect of multitarget dependent descriptor transformation on classification performance is explored in this work. To this end decision trees as well as neural net QSAR in combination with PLS were applied to predict the activity class of 5HT3 ligands, angiotensin converting enzyme inhibitors, 3-hydroxyl-3-methyl glutaryl coenzyme A reductase inhibitors, platelet activating factor antagonists, and thromboxane A2 antagonists. Physicochemical descriptors calculated by MOE and fragment-based descriptors (MOLPRINT 2D) were employed to generate descriptor vectors. In a subsequent step the physicochemical descriptor vectors were transformed to a lower dimensional space using multitarget dependent descriptor transformation. Cross-validation of the original physicochemical descriptors in combination with decision trees and neural net QSAR as well as cross-validation of PLS multitarget transformed descriptors with neural net QSAR were performed. For comparison this was repeated using fragment-based descriptors in combination with decision trees.

Angiotensin-Converting Enzyme Inhibitors↗

Harvesting chemical information from the Internet using a distributed approach: ChemXtreme.

The Internet is a comprehensive resource of chemical information which is at the same time largely unstructured. It provides a wealth of scientific information such as experimental data and requires a suitable automated data mining and analysis tool for its meaningful exploration. The Java based software presented here, ChemXtreme, is developed for harvesting chemical information from the Internet employing the Google API in combination with a distributed client/server text analysis architecture based on JavaRMI. It represents the first and until now the only toolkit for automated structured data retrieval from the Internet which is itself open source. ChemXtreme employs the "search the search engine" strategy, where the URLs returned from the search engine are analyzed further via textual pattern analysis. This process resembles the manual analysis of the hit list, where relevant data are captured and, by means of human intervention, are mined into a format suitable for further analysis. ChemXtreme on the other hand transforms chemical information automatically into a structured format suitable for storage in databases and further analysis and also provides links to the original information source. The query data retrieved from the search engine by the server is encoded, encrypted, and compressed and then sent to all the participating active clients in the network for parsing. Relevant information identified by the clients on the retrieved Web sites is sent back to the server, verified, and added to the database for data mining and further analysis. The distributed further analysis of URLs in a client/server architecture scales very favorably, thus producing only minimal overhead.

Chemistry↗

Characterizing bitterness: identification of key structural features and development of a classification model.

This work describes the first approach in the development of a comprehensive classification method for bitterness of small molecules. The data set comprises 649 bitter and 13 530 randomly selected molecules from the MDL Drug Data Repository (MDDR) which are analyzed by circular fingerprints (MOLPRINT 2D) and information-gain feature selection. The feature selection proposes substructural features which are statistically correlated to bitterness. Classification is performed on the selected features via a naïve Bayes classifier. The substructural features upon which the classification is based are able to discriminate between bitter and random compounds, and thus we propose they are also functionally responsible for causing the bitter taste. Such substructures include various sugar moieties as well as highly branched carbon scaffolds. Cynaropicrine contains a number of the substructural features found to be statistically associated with bitterness and thus was correctly predicted to be bitter by our model. Alternatively, both promethazine and saccharin contain fewer of these substructural features, and thus the bitterness in these compounds was not identified. Two different classes of bitter compounds were identified, namely those which are larger and contain mainly oxygen and carbon and often sugar moieties, and those which are rather smaller and contain additional nitrogen and/or sulfur fragments. The classifier is able to predict 72.1% of the bitter compounds. Feature selection reduces the number of false-positives while also increasing the number of false negatives to 69.5% of bitter compounds correctly predicted. Overall, the method presented here presents both one of the largest databases of bitter compounds presently available as well as a relatively reliable classification method.

Algorithms↗

Chemoinformatics-based classification of prohibited substances employed for doping in sport.

Representative molecules from 10 classes of prohibited substances were taken from the World Anti-Doping Agency (WADA) list, augmented by molecules from corresponding activity classes found in the MDDR database. Together with some explicitly allowed compounds, these formed a set of 5245 molecules. Five types of fingerprints were calculated for these substances. The random forest classification method was used to predict membership of each prohibited class on the basis of each type of fingerprint, using 5-fold cross-validation. We also used a k-nearest neighbors (kNN) approach, which worked well for the smallest values of k. The most successful classifiers are based on Unity 2D fingerprints and give very similar Matthews correlation coefficients of 0.836 (kNN) and 0.829 (random forest). The kNN classifiers tend to give a higher recall of positives at the expense of lower precision. A naïve Bayesian classifier, however, lies much further toward the extreme of high recall and low precision. Our results suggest that it will be possible to produce a reliable and quantitative assignment of membership or otherwise of each class of prohibited substances. This should aid the fight against the use of bioactive novel compounds as doping agents, while also protecting athletes against unjust disqualification.

Algorithms↗

Melting point prediction employing k-nearest neighbor algorithms and genetic parameter optimization.

We have applied the k-nearest neighbor (kNN) modeling technique to the prediction of melting points. A data set of 4119 diverse organic molecules (data set 1) and an additional set of 277 drugs (data set 2) were used to compare performance in different regions of chemical space, and we investigated the influence of the number of nearest neighbors using different types of molecular descriptors. To compute the prediction on the basis of the melting temperatures of the nearest neighbors, we used four different methods (arithmetic and geometric average, inverse distance weighting, and exponential weighting), of which the exponential weighting scheme yielded the best results. We assessed our model via a 25-fold Monte Carlo cross-validation (with approximately 30% of the total data as a test set) and optimized it using a genetic algorithm. Predictions for drugs based on drugs (separate training and test sets each taken from data set 2) were found to be considerably better [root-mean-squared error (RMSE)=46.3 degrees C, r2=0.30] than those based on nondrugs (prediction of data set 2 based on the training set from data set 1, RMSE=50.3 degrees C, r2=0.20). The optimized model yields an average RMSE as low as 46.2 degrees C (r2=0.49) for data set 1, and an average RMSE of 42.2 degrees C (r2=0.42) for data set 2. It is shown that the kNN method inherently introduces a systematic error in melting point prediction. Much of the remaining error can be attributed to the lack of information about interactions in the liquid state, which are not well-captured by molecular descriptors.

Algorithms↗

"Bayes affinity fingerprints" improve retrieval rates in virtual screening and define orthogonal bioactivity space: when are multitarget drugs a feasible concept?

Conventional similarity searching of molecules compares single (or multiple) active query structures to each other in a relative framework, by means of a structural descriptor and a similarity measure. While this often works well, depending on the target, we show here that retrieval rates can be improved considerably by incorporating an external framework describing ligand bioactivity space for comparisons ("Bayes affinity fingerprints"). Structures are described by Bayes scores for a ligand panel comprising about 1000 activity classes extracted from the WOMBAT database. The comparison of structures is performed via the Pearson correlation coefficient of activity classes, that is, the order in which two structures are similar to the panel activity classes. Compound retrieval on a recently published data set could be improved by as much as 24% relative (9% absolute). Knowledge about the shape of the "bioactive chemical universe" is thus beneficial to identifying similar bioactivities. Principal component analysis was employed to further analyze activity space with the objective to define orthogonal ligand bioactive chemical space, leading to nine major (roughly orthogonal) activity axes. Employing only those nine activity classes, retrieval rates are still comparable to original Bayes affinity fingerprints; thus, the concept of orthogonal bioactive ligand chemical space was validated as being an information-rich but low-dimensional representation of bioactivity space. Correlations between activity classes are a major determinant to gauge whether the desired multitarget activity of drugs is (on the basis of current knowledge) a feasible concept because it measures the extent to which activities can be optimized independently, or only by strongly influencing one another.

Algorithms↗