Search PubMed⌕ Search

Biomedical subjects

Douglas M Hawkins

Publications and source records attributed to Douglas M Hawkins.

8 recordsLinked to original sources

Using recursive partitioning analysis to evaluate compound selection methods.

The design and analysis of a screening set for high throughput screening is complex. We examine three statistical strategies for compound selection, random, clustering, and space-filling. We examine two types of chemical descriptors, BCUTs and principal components of Dragon Constitutional descriptors. Based on the predictive power of multiple tree recursive partitioning, we reached the following tentative conclusions. Random designs appear to be as good as clustering and space-filling designs. For analysis, BCUTs appear to be better than principal components scores based upon Constitutional Descriptors. We confirm previous results that model-based selection of compounds can lead to improved screening hit rates.

Decision Trees↗

Robust singular value decomposition analysis of microarray data.

In microarray data there are a number of biological samples, each assessed for the level of gene expression for a typically large number of genes. There is a need to examine these data with statistical techniques to help discern possible patterns in the data. Our technique applies a combination of mathematical and statistical methods to progressively take the data set apart so that different aspects can be examined for both general patterns and very specific effects. Unfortunately, these data tables are often corrupted with extreme values (outliers), missing values, and non-normal distributions that preclude standard analysis. We develop a robust analysis method to address these problems. The benefits of this robust analysis will be both the understanding of large-scale shifts in gene effects and the isolation of particular sample-by-gene effects that might be either unusual interactions or the result of experimental flaws. Our method requires a single pass and does not resort to complex "cleaning" or imputation of the data table before analysis. We illustrate the method with a commercial data set.

Cluster Analysis↗

Prediction of human blood: air partition coefficient: a comparison of structure-based and property-based methods.

In recent years, there has been increased interest in the development and use of quantitative structure-activity/property relationship (QSAR/QSPR) models. For the most part, this is due to the fact that experimental data is sparse and obtaining such data is costly, while theoretical structural descriptors can be obtained quickly and inexpensively. In this study, three linear regression methods, viz. principal component regression (PCR), partial least squares (PLS), and ridge regression (RR), were used to develop QSPR models for the estimation of human blood:air partition coefficient (logPblood:air) for a group of 31 diverse low-molecular weight volatile chemicals from their computed molecular descriptors. In general, RR was found to be superior to PCR or PLS. Comparisons were made between models developed using parameters based solely on molecular structure and linear regression (LR) models developed using experimental properties, including saline:air partition coefficient (logPsaline:air) and olive oil:air partition coefficient (logPolive oil:air), as independent variables, indicating that the structure-property correlations are comparable to the property-property correlations. The best models, however, were those that used rat logPblood:air as the independent variable. Haloalkane subgroups were modeled separately for comparative purposes and, although models based on the congeneric compounds were superior, the models developed on the complete set of diverse compounds were of acceptable quality. The structural descriptors were placed into one of three classes based on level of complexity: topostructural (TS), topochemical (TC), or three-dimensional/geometrical (3D). Modeling was performed using the structural descriptor classes both in a hierarchical fashion and separately. The results indicate that highest quality structure-based models, in terms of descriptor classes, were those derived using TC descriptors.

Animals↗

Diagnostics for conformity of paired quantitative measurements.

Matched pairs data arise in many contexts - in case-control clinical trials, for example, and from cross-over designs. They also arise in experiments to verify the equivalence of quantitative assays. This latter use (which is the main focus of this paper) raises difficulties not always seen in other matched pairs applications. Since the designs deliberately vary the analyte levels over a wide range, issues of variance dependent on mean, calibrations of differing slopes, and curvature all need to be added to the usual model assumptions such as normality. Violations in any of these assumptions invalidate the conventional matched pairs analysis. A graphical method, due to Bland and Altman, of looking at the relationship between the average and the difference of the members of the pairs is shown to correspond to a formal testable regression model. Using standard regression diagnostics, one may detect and diagnose departures from the model assumptions and remedy them - for example using variable transformations. Examples of different common scenarios and possible approaches to handling them are shown.

Chemistry Techniques, Analytical↗

Assessing model fit by cross-validation.

When QSAR models are fitted, it is important to validate any fitted model-to check that it is plausible that its predictions will carry over to fresh data not used in the model fitting exercise. There are two standard ways of doing this-using a separate hold-out test sample and the computationally much more burdensome leave-one-out cross-validation in which the entire pool of available compounds is used both to fit the model and to assess its validity. We show by theoretical argument and empiric study of a large QSAR data set that when the available sample size is small-in the dozens or scores rather than the hundreds, holding a portion of it back for testing is wasteful, and that it is much better to use cross-validation, but ensure that this is done properly.

Journal Article↗

Quantitative structure-activity relationship modeling of juvenile hormone mimetic compounds for Culex pipiens larvae, with a discussion of descriptor-thinning methods.

Quantitative structure-activity relationship (QSAR) modelers often encounter the problem of multicollinearity owing to the availability of large numbers of computable molecular descriptors. Sparsity of the variables while using descriptors such as atom pairs increases the complexity. Three different predictor-thinning methods, namely, a modified Gram-Schmidt algorithm, a marginal soft thresholding algorithm, and LASSO (least absolute shrinkage and selection operator), were utilized to reduce the number of descriptors prior to developing linear models. Juvenile hormone (JH) activity of 304 compounds on Culex pipiens larvae was taken as the model data set, and predictor trimming of a large number of diverse descriptors comprising 268 global molecular descriptors (topostructural, topochemical, and geometrical), 13 quantum chemical descriptors, and 915 atom pairs (substructural counts) was applied prior to linear regression by the ridge regression method. The data set (N = 304) was split into five calibration data sets of random samples of sizes 60/110/160/210/260, and the remaining 244/194/144/94/44 compounds were used for validations. LASSO was not found to be a very effective method in handling a large set of descriptors because the number of predictors retained could not exceed the number of observations. The results indicated that the modified Gram-Schmidt algorithm could be used to trim the number of predictors in the global molecular descriptor set where collinearity of the descriptors was the major concern. On the contrary, the soft thresholding approach was found to be an effective tool in subset selection from a diverse set of descriptors having both sparsity and multicollinearity, as in the case of the combined set of atom pairs and global molecular descriptors. The final model developed after variable selection was dominated more by atom pairs, which indicated the important structural moieties that affect JH activity of the compounds. The success of the method reiterates the fact that QSAR or quantitative structure-property relationship (QSPR) models can be developed for a diverse set of compounds using properly parametrized and diverse sets of descriptors, of course, with the selection of the appropriate statistical tools.

Animals↗

Combining chemodescriptors and biodescriptors in quantitative structure-activity relationship modeling.

In view of the wide distribution of halocarbons in our world, their toxicity is a public health concern. Previous work has shown that various measures of toxicity can be predicted with standard molecular descriptors. In our work, biodescriptors of the effect of halocarbons on the liver were obtained by exposing hepatocytes to 14 halocarbons and a control and by producing two-dimensional electrophoresis gels to assess the expressed proteome. The resulting spot abundances provide additional biological information that might be used in toxicity prediction. QSAR models were fitted via ridge regression to predict eight dependent toxicity measures: d37, arr, EC50MTT, EC50LDH, EC20SH, LECLP, LECROS, and LECCAT. Three predictor sets were used for each-the chemodescriptors alone, the biodescriptors alone, and the combined set of both chemo- and biodescriptors. The results differed somewhat from one dependent to another, but overall it was shown that better results could be obtained by using both chemo- and biodescriptors in the model than by using either chemo- or biodescriptors alone. The library of compounds used was small and quite homogeneous, so our immediate conclusions are correspondingly limited in scope, but we believe the underlying methodologies have broad applicability at the interface of chemical and biological descriptors.

Animals↗