Search PubMed⌕ Search

Biomedical subjects

S C Basak

Publications and source records attributed to S C Basak.

43 records · Page 3Linked to original sources

QSAR with few compounds and many features.

Fitting quantitative structure-activity relationships (QSAR) requires different statistical methodologies and, to some degree, philosophies depending on the "shape" of the data matrix. When few features are used and there are many compounds, it is a reasonable expectation that good feature subset selection may be made and that nonlinearities and nonadditivities can be detected and diagnosed. Where there are many features and few compounds, this is unrealistic. Methods such as ridge regression RR, PLS, and principal component regression PCR, which abjure feature selection and rely on linearity may provide good predictions and fair understanding. We report a development of ridge regression for the underdetermined case by using generalized cross-validation to choose the ridge constant and perform F-tests for additional information. Conventional regression diagnostics can be used in followup to identify nonlinearities and other departures from model. We illustrate the approach with QSAR models of four data sets using calculated molecular descriptors.

Algorithms↗

On the characterization of DNA primary sequences by triplet of nucleic acid bases.

We consider construction of a set of smaller 4 x 4 matrices to represent DNA primary sequences which are based on enumeration of all 64 triplets of nucleic acids bases. The leading eigenvalue from the constructed matrices has been selected as an invariant for construction of a vector to characterize DNA. Additional invariants considered of the derived condensed matrices of DNA include a 64-component vector, the components of which consist of ordered triplets XYZ, with X, Y, Z = A, C, G, T. Construction of similarity/dissimilarity tables based on different invariants for a set of sequences of DNA belonging to the first exon of the beta-globin gene of eight species illustrates the utility of newly formulated invariants for DNA.

Animals↗

Prediction of mutagenicity of aromatic and heteroaromatic amines from structure: a hierarchical QSAR approach.

Due to the lack of experimental data, there has been increasing use of theoretical structural descriptors in the hazard assessment of chemicals. We have used a hierarchical approach to develop class-specific quantitative structure-activity relationship (QSAR) models for the prediction of mutagenicity of a set of 95 aromatic and heteroaromatic amines. The hierarchical approach begins with the simplest molecular descriptors, the topostructural, which encode limited chemical information. The complexity is then increased, adding topochemical, geometric, and finally quantum chemical parameters. We have also added log P to the set of independent variables. The results indicate that the topological parameters, i.e., the topostructural and topochemical indices, explain the majority of the variance, and that the inclusion of log P, geometric, and quantum chemical parameters does not result in significantly improved predictive models.

Algorithms↗

Quantitative structure-property relationships (QSPRs) for the estimation of vapor pressure: a hierarchical approach using mathematical structural descriptors.

A set of 379 molecular descriptors was calculated for use in hierarchical quantitative structure-property relationship (QSPR) modeling of vapor pressure for a structurally diverse database consisting of 469 chemicals. The hierarchical approach utilizes topostructural, topochemical, geometrical, and quantum chemical descriptors in a stepwise fashion to develop QSPR models. In this way, the relative roles of the various levels of descriptors can be examined. The results show that the easily calculated topological descriptors explain the majority of the variance and that the addition of geometrical and quantum chemical descriptors does not result in a significantly improved model.

Journal Article↗

Prediction of complement-inhibitory activity of benzamidines using topological and geometric parameters.

A hierarchical approach to quantitative structure-activity relationship (QSAR) modeling has been used to estimate the complement-inhibitory potency of 105 benzamidines. This hierarchical approach uses topostructural, topochemical, and geometric parameters in a stepwise fashion to build increasingly more complex models. The results show that topostructural indices alone, specifically I(D), predict inhibitory potency reasonably well. The addition of topochemical and geometrical parameters to the set of descriptors provides only marginal improvement in predictive power. However, when taken alone, the geometric parameter (3D)W provides a more stable model than the topostructural one.

Animals↗

Use of statistical and neural net approaches in predicting toxicity of chemicals.

Hierarchical quantitative structure-activity relationships (H-QSAR) have been developed as a new approach in constructing models for estimating physicochemical, biomedicinal, and toxicological properties of interest. This approach uses increasingly more complex molecular descriptors in a graduated approach to model building. In this study, statistical and neural network methods have been applied to the development of H-QSAR models for estimating the acute aquatic toxicity (LC50) of 69 benzene derivatives to Pimephales promelas (fathead minnow). Topostructural, topochemical, geometrical, and quantum chemical indices were used as the four levels of the hierarchical method. It is clear from both the statistical and neural network models that topostructural indices alone cannot adequately model this set of congeneric chemicals. Not surprisingly, topochemical indices greatly increase the predictive power of both statistical and neural network models. Quantum chemical indices also add significantly to the modeling of this set of acute aquatic toxicity data.

Animals↗

Simple numerical descriptor for quantifying effect of toxic substances on DNA sequences.

Many chemicals are known to be toxic to living organisms, inducing mutations and deletions at the chromosomal and genetic level. One of the tasks in risk assessment of genotoxic chemicals is to devise a simple numerical descriptor which may be used to quantify the relationship between chemical dose and the effect on the genetic sequences. We have developed numerical descriptors to characterize different DNA sequences which are especially useful in sequence comparisons. These descriptors have been developed from a graphical representational technique that enables easy visualization of changes in base distributions arising from evolutionary or other effects. In this paper we propose a scheme to use these descriptors as a label to help quantify the potential risk hazard of chemicals inducing mutations and deletions in DNA sequences.

Base Sequence↗