Search PubMed⌕ Search

PubMed · 15729854

Validation tools for variable subset regression.

Abstract

Variable selection is applied frequently in QSAR research. Since the selection process influences the characteristics of the finally chosen model, thorough validation of the selection technique is very important. Here, a validation protocol is presented briefly and two of the tools which are part of this protocol are introduced in more detail. The first tool, which is based on permutation testing, allows to assess the inflation of internal figures of merit (such as the cross-validated prediction error). The other tool, based on noise addition, can be used to determine the complexity and with it the stability of models generated by variable selection. The obtained statistical information is important in deciding whether or not to trust the predictive abilities of a specific model. The graphical output of the validation tools is easily accessible and provides a reliable impression of model performance. Among others, the tools were employed to study the influence of leave-one-out and leave-multiple-out cross-validation on model characteristics. Here, it was confirmed that leave-multiple-out cross-validation yields more stable models. To study the performance of the entire validation protocol, it was applied to eight different QSAR data sets with default settings. In all cases internal and external model performance was good, indicating that the protocol serves its purpose quite well.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Knut Baumann, Nikolaus Stiefl. Validation tools for variable subset regression.. https://doi.org/10.1007/s10822-004-4071-5

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship↗

QSAR study on pK(a) vis-à-vis physiological activity of sulfonamides: a dominating role of surface tension (inverse steric parameter).

The paper describes the dominating role of surface tension (ST) on the modeling, monitoring, and estimating pK(a) for a large series of 43 substituted sulfonamides. Because of the direct correlation of ST with parachor (Pc) vis-a-vis molecular volume (MV), ST is considered as a steric parameter. Single as well as multi-parametric regressions have indicated that ST has a dominating role in QSAR of the set of sulfonamides used and that excellent results are obtained in multi-parametric regression analysis. The results are discussed critically on the basis of statistical parameters.

Quantitative Structure-Activity Relationship↗

Derivation and applications of molecular descriptors based on approximate surface area.

Three sets of molecular descriptors that can be computed from a molecular connection table are defined. The descriptors are based on the subdivision and classification of the molecular surface area according to atomic properties (such as contribution to logP, molar refractivity, and partial charge). The resulting 32 descriptors are shown (a) to be weakly correlated with each other; (b) to encode many traditional molecular descriptors; and (c) to be useful for QSAR, QSPAR, and compound classification.

Quantitative Structure-Activity Relationship↗