Search PubMed⌕ Search

Biomedical subjects

Harald Mauser

Publications and source records attributed to Harald Mauser.

5 recordsLinked to original sources

A robust clustering method for chemical structures.

A clustering method based on finding the largest set of disconnected fragments that two chemical compounds have in common is shown to be able to group structures in a way that is ideally suited to medicinal chemistry programs. We describe how markedly improved results can be obtained by using a similarity metric that accounts not just for the size of the shared fragments but also on their relative arrangement in the two parent compounds. The use of a physiochemical atom typing scheme is also shown to provide significant contributions. Results from calculations using a test set consisting of actives from nine different important biological target proteins demonstrate the strengths of our clustering method and the advantages over other approaches that are widely used throughout the pharmaceutical industry.

Chemistry, Organic↗

A validation study on the practical use of automated de novo design.

The de novo design program Skelgen has been used to design inhibitor structures for four targets of pharmaceutical interest. The designed structures are compared to modeled binding modes of known inhibitors (i) visually and (ii) by means of a novel similarity measure considering the size and spatial proximity of the maximum common substructure of two small molecules. It is shown that the Skelgen algorithm generates representatives of many inhibitor classes within a very short time and that the new similarity measure is useful for comparing and clustering designed structures. The results demonstrate the necessity of properly defining search constraints in practical applications of de novo design.

Algorithms↗

Prediction of UV and ESI-MS signal intensities.

All major pharmaceutical companies maintain large collections of compounds that are used either for screening against biological targets or as synthetic precursors. The quality assessment of these compounds is typically done by liquid chromatography combined with mass spectroscopy (LC/MS) and UV purity control. To facilitate the analysis of the analytical data, we have built computational models to predict UV and MS signal intensities under experimental LC/MS conditions. The discriminant partial-least-squares technique was used for classifying compounds into those most likely to yield a MS signal and others where the signal is below the detection limit (94% and 88% correct predictions, respectively). In the case of UV prediction, we compared this statistical linear-regression technique to a knowledge-based approach. A combination of both techniques proved to be the most reliable (96/98% correct predictions of UV-active/ UV-inactive compounds). Both models have been incorporated into the automated compound integrity profiling at F. Hoffmann-La Roche.

Chromogenic Compounds↗

Ensemble methods for classification in cheminformatics.

We describe the application of ensemble methods to binary classification problems on two pharmaceutical compound data sets. Several variants of single and ensembles models of k-nearest neighbors classifiers, support vector machines (SVMs), and single ridge regression models are compared. All methods exhibit robust classification even when more features are given than observations. On two data sets dealing with specific properties of drug-like substances (cytochrome P450 inhibition and "Frequent Hitters", i.e., unspecific protein inhibition), we achieve classification rates above 90%. We are able to reduce the cross-validated misclassification rate for the Frequent Hitters problem by a factor of 2 compared to previous results obtained for the same data set with different modeling techniques.

Journal Article↗

Database clustering with a combination of fingerprint and maximum common substructure methods.

We present an efficient method to cluster large chemical databases in a stepwise manner. Databases are first clustered with an extended exclusion sphere algorithm based on Tanimoto coefficients calculated from Daylight fingerprints. Substructures are then extracted from clusters by iterative application of a maximum common substructure algorithm. Clusters with common substructures are merged through a second application of an exclusion sphere algorithm. In a separate step, singletons are compared to cluster substructures and added to a cluster if similarity is sufficiently high. The method identifies tight clusters with conserved substructures and generates singletons only if structures are truly distinct from all other library members. The method has successfully been applied to identify the most frequently occurring scaffolds in databases, for the selection of analogues of screening hits and in the prioritization of chemical libraries offered by commercial vendors.

Journal Article↗