Search PubMed⌕ Search

Biomedical subjects

Subhash C Basak

Publications and source records attributed to Subhash C Basak.

At least 19 recordsLinked to original sources

Quantitative structure-activity relationship (QSAR) studies of quinolone antibacterials against M. fortuitum and M. smegmatis using theoretical molecular descriptors.

The incidence of tuberculosis infections that are resistant to conventional drug therapy has risen steadily in the last decade. Several of the quinolone antibacterials have been examined as inhibitors of M. tuberculosis infection as well as other mycobacterial infections. However, not much has been done to examine specific structure-activity relationships of the quinolone antibacterials against mycobacteria. The present paper describes quantitative structure-activity relationship modeling for a series of antimycobacterial compounds. Most of the antimycobacterial compounds do not have sufficient physicochemical data, and thus predictive methods based on experimental data are of limited use in this situation. Hence, there is a need for the development of quantitative structure-activity relationship (QSAR) models utilizing theoretical molecular descriptors that can be calculated directly from molecular structures. Descriptors associated with chemical structures of N-1 and C-7 substituted quinolone derivatives as well as 8-substituted quinolone derivatives with good antimycobacterial activities against M. fortuitum and M. smegmatis have been evaluated. Ridge regression (RR), Principal component regression (PCR), and partial least squares (PLS) regression were used, comparatively, to develop predictive models for antibacterial activity, based on the activities of the above compounds. The independent variables include topostructural, topochemical and 3-D geometrical indices, which were used in a hierarchical fashion in the model-development process. The predictive ability of the models was assessed by the cross-validated R2. Comparison of the relative effectiveness of the various classes of molecular descriptors in the regression models shows that the easily calculable topological indices explain most of the variance in the data.

Anti-Bacterial Agents↗

Complex graph matrix representations and characterizations of proteomic maps and chemically induced changes to proteomes.

We have presented a complex graph matrix representation to characterize proteomics maps obtained from 2D-gel electrophoresis. In this method, each bubble in a 2D-gel proteomics map is represented by a complex number with components which are charge and mass. Then, a graph with complex weights is constructed by connecting the vertices in the relative order of abundance. This yields adjacency matrices and distance matrices of the proteomics graph with complex weights. We have computed the spectra, eigenvectors, and other properties of complex graphs and the Euclidian/graph distance obtained from the complex graphs. The leading eigenvalues and eigenvectors and, likewise, the smallest eigenvalues and eigenvectors, and the entire graph spectral patterns of the complex matrices derived from them yield novel weighted biodescriptors that characterize proteomics maps with information of charge and masses of proteins. We have also applied these eigenvector and eigenvalue maps to contrast the normal cells and cells exposed to four peroxisome proliferators, namely, clofibrate, diethylhexyl phthalate (DEHP), perfluorodecanoic acid (PFDA), and perfluoroctanoic acid (PFOA). Our complex eigenspectra show that the proteomic response induced by DEHP differs from the corresponding responses of other three chemicals consistent with their chemical structures and properties.

Algorithms↗

Chirality index, molecular overlay and biological activity of diastereoisomeric mosquito repellents.

Both 1-methylisopropyl 2-(2-hydroxyethyl)piperidine-1-carboxylate, (Picaridin((R))) and cyclohex-3-enyl 2-methylpiperidin-1-yl ketone (AI3-37220; 220) have two asymmetric centers, and the four diastereoisomers of each compound are known to have differing degrees of mosquito-repellent activity according to quantitative behavioral assays conducted at the United States Department of Agriculture. Computational chemistry was used to identify the structural and configurational basis for repellent activity. Molecular overlay of the optimized geometries of the lowest energy conformers of the diastereoisomers was investigated to elucidate the role of chiral centers in 220 and Picaridin. It was found that the presence of a chiral carbon alpha to the nitrogen with the S configuration in the piperidine ring is essential to the three-dimensional arrangement of the atoms of the pharmacophore for effective repellent activity.

Animals↗

Counter-propagation artificial neural network as a tool for the independent variable selection: structure-mutagenicity study on aromatic amines.

The counter-propagation artificial neural network (CP ANN) technique was applied for the independent variable selection and for structure-mutagenic potency modeling on a set of 95 aromatic and heteroaromatic amines with biological activity investigated experimentally by an in vitro assay. The molecular structures were represented by 275 independent variables classified as topostructural, topochemical, geometrical and quantum-chemical descriptors. As a result of the neural network modeling, the following descriptors were found to be the most important for structure-activity relationship: 5 chi -path connectivity index of order h = 5, 3chibC-bond cluster connectivity index of order h = 3, J(B)-Balaban's J index based on bond types, SHSNH2-electrotopological state index values for atoms, phia-flexibility index (kappa p1 x kappa p2/nvx), IC0-mean information content or complexity of a graph based on the 0 order neighborhood of vertices in a hydrogen-filled graph and ELUMO. The leave one out (LOO) method was used in order to test and select the models for mutagenicity prediction. The statistical parameters for the 7-descriptors model are R(Model) = 0.96 and Rcv = 0.85, respectively. In the next step, the number of variables was reduced and the 4-descriptors model was found (R(Model) = 0.95 and Rcv = 0.85) and classified as the best one.

Algorithms↗

QSAR study using topological indices for inhibition of carbonic anhydrase II by sulfanilamides and Schiff bases.

A QSAR study of two sets of carbonic anhydrase inhibitors is presented using a variety of molecular descriptors including topological indices. The first set consists of 29 benzenesulphonamides, and the second set includes 35 sulphanilamide Schiff bases. Two regression methodologies have been used involving ridge regression and the CODESSA program, and their results are compared with those of previous QSAR studies. Good correlations were found for the former set, and less satisfactory results for the latter set when the number of molecular descriptors is kept below five.

Animals↗

Usefulness of graphical invariants in quantitative structure-activity correlations of tuberculostatic drugs of the isonicotinic acid hydrazide type.

Quantitative structure-activity relationship (QSAR) studies have been performed for a series of 2-substituted isonicotinic acid hydrazides utilizing theoretical molecular descriptors. 223 topological (topostructural and topochemical) indices along with seven geometrical descriptors were computed for the prediction of antibacterial activity against Mycobacterium tuberculosis. Ridge-regression models assessed by cross-validated R2 have been formulated, and a comparative study on the relative effectiveness of physicochemical vis-à-vis theoretical molecular descriptors performed. The models developed clearly indicate the supremacy of structure-activity over property-activity relationships in the current study and can be used to evaluate the potential tuberculostatic activity of other INH derivatives, real or hypothetical.

Antitubercular Agents↗

Novel map descriptors for characterization of toxic effects in proteomics maps.

We consider a novel numerical characterization of proteomics maps based on the construction of a graph obtained by connecting all protein spots in a proteomics map that are at distance equal to, or smaller than, a critical distance D(c). We refer to the so constructed graph as a cluster graph and we calculate four associated characteristic matrices, previously considered in the literature: (1) the Euclidean-distance matrix ED; (2) the neighborhood-distance matrix ND; (3) the path-distance matrix based on the shortest paths between connected spots PD; and (4) the quotient matrix Q, the elements of which are given as the quotient of the corresponding elements of ED and ND matrices. Numerical descriptors for proteomics maps include in particular the leading eigenvalue of the Q matrix and the family of associated "higher order" matrices defined as powers of Q. These map descriptors show considerable sensitivity to perturbations of proteomics maps by toxicants.

Animals↗

Prediction of human blood: air partition coefficient: a comparison of structure-based and property-based methods.

In recent years, there has been increased interest in the development and use of quantitative structure-activity/property relationship (QSAR/QSPR) models. For the most part, this is due to the fact that experimental data is sparse and obtaining such data is costly, while theoretical structural descriptors can be obtained quickly and inexpensively. In this study, three linear regression methods, viz. principal component regression (PCR), partial least squares (PLS), and ridge regression (RR), were used to develop QSPR models for the estimation of human blood:air partition coefficient (logPblood:air) for a group of 31 diverse low-molecular weight volatile chemicals from their computed molecular descriptors. In general, RR was found to be superior to PCR or PLS. Comparisons were made between models developed using parameters based solely on molecular structure and linear regression (LR) models developed using experimental properties, including saline:air partition coefficient (logPsaline:air) and olive oil:air partition coefficient (logPolive oil:air), as independent variables, indicating that the structure-property correlations are comparable to the property-property correlations. The best models, however, were those that used rat logPblood:air as the independent variable. Haloalkane subgroups were modeled separately for comparative purposes and, although models based on the congeneric compounds were superior, the models developed on the complete set of diverse compounds were of acceptable quality. The structural descriptors were placed into one of three classes based on level of complexity: topostructural (TS), topochemical (TC), or three-dimensional/geometrical (3D). Modeling was performed using the structural descriptor classes both in a hierarchical fashion and separately. The results indicate that highest quality structure-based models, in terms of descriptor classes, were those derived using TC descriptors.

Animals↗

A comparative study of proteomics maps using graph theoretical biodescriptors.

This paper reports the development of new methods for mathematical characterization of effects of different toxic agents on the cellular proteome. We describe numerical characterization of proteomics maps based on mathematical invariants. A graph is first associated with a proteomics map by considering partial ordering of spots on 2-D gels by ordering proteins with respect to the mass and the charge, the two properties by which proteins are separated. The graph is then embedded over the map, and several graph theoretical invariants have been constructed. In particular we consider invariants that can be extracted from the Euclidean distance-adjacency matrix of the embedded graph, in which only Euclidean distances between adjacent vertices of a graph are considered. The approach is illustrated using proteomics patterns of normal liver cells of rats and those derived from liver cells of animals exposed to four peroxisome proliferators. In contrast to direct comparison of spot abundance our approach incorporates information on spots locations. The difference between the two approaches is that in the first case only changes in abundances are considered as a measure of perturbation of the proteome map, but in the second case not only the charge but also the mass of proteins are used for ordering protein spots.

Animals↗

Prediction of cellular toxicity of halocarbons from computed chemodescriptors: a hierarchical QSAR approach.

A hierarchical quantitative structure-activity relationship (HiQSAR) approach was used to estimate toxicity and genetic toxicity for a set of 55 halocarbons using computed chemodescriptors. The descriptors consisted of topostructural (TS), topochemical (TC), geometrical, semiempirical (AM1) quantum chemical, and ab initio (STO-3G, 6-31G(d), 6-311G, 6-311G(d), and aug-cc-pVTZ) quantum chemical indices. For the two toxicity endpoints investigated, ARR and D(37), the TC indices gave the best cross-validated R(2) values. The 3-D indices also performed either as well as or slightly superior to the TC indices. For the four categories of quantum chemical indices used for the development of predictive models, the AM1 parameters gave the worst performance, and the most advanced ab initio (B3LYP/aug-CC-pVTZ) parameters gave the best results when used alone. This was also the case when the quantum chemical indices were used in the hierarchical QSAR approach for both of the toxicity endpoints, ARR and D(37). The models resulting from HiQSAR are of sufficiently good quality to estimate toxicity of halocarbons from structure.

Aspergillus niger↗

QSAR modeling of flotation collectors using principal components extracted from topological indices.

Several topological indices were calculated for substituted-cupferrons that were tested as collectors for the froth flotation of uranium. The principal component analysis (PCA) was used for data reduction. Seven principal components (PC) were found to account for 98.6% of the variance among the computed indices. The principal components thus extracted were used in stepwise regression analyses to construct regression models for the prediction of separation efficiencies (Es) of the collectors. A two-parameter model with a correlation coefficient of 0.889 and a three-parameter model with a correlation coefficient of 0.913 were formed. PCs were found to be better than partition coefficient to form regression equations, and inclusion of an electronic parameter such as Hammett sigma or quantum mechanically derived electronic charges on the chelating atoms did not improve the correlation coefficient significantly. The method was extended to model the separation efficiencies of mercaptobenzothiazoles (MBT) and aminothiophenols (ATP) used in the flotation of lead and zinc ores, respectively. Five principal components were found to explain 99% of the data variability in each series. A three-parameter equation with correlation coefficient of 0.985 and a two-parameter equation with correlation coefficient of 0.926 were obtained for MBT and ATP, respectively. The amenability of separation efficiencies of chelating collectors to QSAR modeling using PCs based on topological indices might lead to the selection of collectors for synthesis and testing from a virtual database.

Journal Article↗

Assessing model fit by cross-validation.

When QSAR models are fitted, it is important to validate any fitted model-to check that it is plausible that its predictions will carry over to fresh data not used in the model fitting exercise. There are two standard ways of doing this-using a separate hold-out test sample and the computationally much more burdensome leave-one-out cross-validation in which the entire pool of available compounds is used both to fit the model and to assess its validity. We show by theoretical argument and empiric study of a large QSAR data set that when the available sample size is small-in the dozens or scores rather than the hundreds, holding a portion of it back for testing is wasteful, and that it is much better to use cross-validation, but ensure that this is done properly.

Journal Article↗

On the dependence of a characterization of proteomics maps on the number of protein spots considered.

We have reexamined the numerical characterization of proteomics maps based on the construction of novel distance matrices associated with the nearest neighbor graph for the protein spots. In particular we consider dependence of a characterization of proteomics map on the number of proteins considered in the analysis. We examined a collection of proteomics maps in which we approximately doubled the number of spots to be used for quantitative analysis, considering cases of maps having 30, 50, 100, 250, 500, and 1054 protein spots. For each case we have compared the similarity-dissimilarity results for five proteomics maps of rat liver cells associated with the control group and four proliferators administrated by intraperitoneal injection. We found that proteins maps based on a set of about the 250 most abundant proteins spots suffice for a satisfactory numerical characterization of such maps.

Animals↗

Quantitative structure-activity relationship modeling of juvenile hormone mimetic compounds for Culex pipiens larvae, with a discussion of descriptor-thinning methods.

Quantitative structure-activity relationship (QSAR) modelers often encounter the problem of multicollinearity owing to the availability of large numbers of computable molecular descriptors. Sparsity of the variables while using descriptors such as atom pairs increases the complexity. Three different predictor-thinning methods, namely, a modified Gram-Schmidt algorithm, a marginal soft thresholding algorithm, and LASSO (least absolute shrinkage and selection operator), were utilized to reduce the number of descriptors prior to developing linear models. Juvenile hormone (JH) activity of 304 compounds on Culex pipiens larvae was taken as the model data set, and predictor trimming of a large number of diverse descriptors comprising 268 global molecular descriptors (topostructural, topochemical, and geometrical), 13 quantum chemical descriptors, and 915 atom pairs (substructural counts) was applied prior to linear regression by the ridge regression method. The data set (N = 304) was split into five calibration data sets of random samples of sizes 60/110/160/210/260, and the remaining 244/194/144/94/44 compounds were used for validations. LASSO was not found to be a very effective method in handling a large set of descriptors because the number of predictors retained could not exceed the number of observations. The results indicated that the modified Gram-Schmidt algorithm could be used to trim the number of predictors in the global molecular descriptor set where collinearity of the descriptors was the major concern. On the contrary, the soft thresholding approach was found to be an effective tool in subset selection from a diverse set of descriptors having both sparsity and multicollinearity, as in the case of the combined set of atom pairs and global molecular descriptors. The final model developed after variable selection was dominated more by atom pairs, which indicated the important structural moieties that affect JH activity of the compounds. The success of the method reiterates the fact that QSAR or quantitative structure-property relationship (QSPR) models can be developed for a diverse set of compounds using properly parametrized and diverse sets of descriptors, of course, with the selection of the appropriate statistical tools.

Animals↗

Combining chemodescriptors and biodescriptors in quantitative structure-activity relationship modeling.

In view of the wide distribution of halocarbons in our world, their toxicity is a public health concern. Previous work has shown that various measures of toxicity can be predicted with standard molecular descriptors. In our work, biodescriptors of the effect of halocarbons on the liver were obtained by exposing hepatocytes to 14 halocarbons and a control and by producing two-dimensional electrophoresis gels to assess the expressed proteome. The resulting spot abundances provide additional biological information that might be used in toxicity prediction. QSAR models were fitted via ridge regression to predict eight dependent toxicity measures: d37, arr, EC50MTT, EC50LDH, EC20SH, LECLP, LECROS, and LECCAT. Three predictor sets were used for each-the chemodescriptors alone, the biodescriptors alone, and the combined set of both chemo- and biodescriptors. The results differed somewhat from one dependent to another, but overall it was shown that better results could be obtained by using both chemo- and biodescriptors in the model than by using either chemo- or biodescriptors alone. The library of compounds used was small and quite homogeneous, so our immediate conclusions are correspondingly limited in scope, but we believe the underlying methodologies have broad applicability at the interface of chemical and biological descriptors.

Animals↗

Proteomic maps-toxicity relationship of halocarbons studied with similarity index and genetic algorithm.

In this work we analyzed proteomic maps obtained from hepatocytes, which were treated with 14 halocarbons. A similarity index was introduced as a robust measure of similarity between two maps or between two selections of spots within the maps. A searching algorithm was used to identify the spots that may play an important role in toxicity mechanism. The highest correlation coefficients obtained between the similarity index and biological parameter were larger than 0.9.

Algorithms↗

On invariants of a 2-D proteome map derived from neighborhood graphs.

We consider the problem of the construction of invariants for characterization of 2-D maps, such as 2-D proteome maps, 2-D NMR spectral maps, etc., that in addition to facilitating cataloguing such maps, can be used for comparison of maps and numerical evaluation of their degree of similarity. A novel approach, based on the concept that the nearest neighborhood of points (spots) on a map are sufficiently flexible to allow one not only to vary the number of points used for characterization of the map but also the density of information on their relative positions, is put forward. The method is illustrated with the Coomassie brilliant blue stained 2-D gel electrophoresis patterns of the proteomes from liver cells of healthy male Fisher F344 rats and the rats treated with four peroxisome proliferators.

Animals↗