Search PubMed⌕ Search

Biomedical subjects

Igor V Tetko

Publications and source records attributed to Igor V Tetko.

24 records · Page 2Linked to original sources

Application of ALOGPS to predict 1-octanol/water distribution coefficients, logP, and logD, of AstraZeneca in-house database.

The ALOGPS 2.1 was developed to predict 1-octanol/water partition coefficients, logP, and aqueous solubility of neutral compounds. An exclusive feature of this program is its ability to incorporate new user-provided data by means of self-learning properties of Associative Neural Networks. Using this feature, it calculated a similar performance, RMSE = 0.7 and mean average error 0.5, for 2569 neutral logP, and 8122 pH-dependent logD(7.4), distribution coefficients from the AstraZeneca "in-house" database. The high performance of the program for the logD(7.4) prediction looks surprising, because this property also depends on ionization constants pKa. Therefore, logD(7.4) is considered to be more difficult to predict than its neutral analog. We explain and illustrate this result and, moreover, discuss a possible application of the approach to calculate other pharmacokinetic and biological activities of chemicals important for drug development.

1-Octanol↗

An unsupervised automatic method for sorting neuronal spike waveforms in awake and freely moving animals.

The present study introduces an approach to automatic classification of extracellularly recorded action potentials of neurons. The classification of spike waveform is considered a pattern recognition problem of special segments of signal that correspond to the appearance of spikes. The spikes generated by one neuron should be recognized as members of the same class. The spike waveforms are described by the nonlinear oscillating model as an ordinary differential equation with perturbation, thus characterizing the signal distortions in both amplitude and phase. It is shown that the use of local variables reduces the problem of spike recognition to the separation of a mixture of normal distributions in the transformed feature space. We have developed an unsupervised iteration-learning algorithm that estimates the number of classes and their centers according to the distance between spike trajectories in phase space. This algorithm scans the learning set to evaluate spike trajectories with maximal probability density in their neighborhood. Following the learning, the procedure of minimal distance is used to perform spike recognition. Estimation of trajectories in phase space requires calculation of the first- and second-order derivatives, and integral operators with piecewise polynomial kernels were used. This provided the computational efficiency of the developed approach for real-time application as required by recordings in behaving animals and in human neurosurgical operations. The new method of spike sorting was tested on simulated and real data and performed better than other approaches currently used in neurophysiology.

Action Potentials↗

The WWW as a tool to obtain molecular parameters.

This article analyses molecular property calculation resources available on the Internet. The first section summarizes the on-line database resources that could be useful to search molecular and biological properties of chemicals, and indicates some principal databases with physicochemical, thermochemical, toxicity, cancer and HIV data. The second section overviews popular standalone programs for calculation of molecular descriptors. Some of these programs can be downloaded for free and used as standalone applications for calculation of molecular descriptors. The third section describes on-line tools for the prediction of molecular properties, activities and calculation of molecular descriptors. Analysis of emerging tools that can be useful to developing new on-line servers for the prediction of molecular parameters and properties is also given.

Chemistry, Pharmaceutical↗

Neural network studies. 4. Introduction to associative neural networks.

Associative neural network (ASNN) represents a combination of an ensemble of feed-forward neural networks and the k-nearest neighbor technique. This method uses the correlation between ensemble responses as a measure of distance amid the analyzed cases for the nearest neighbor technique. This provides an improved prediction by the bias correction of the neural network ensemble. An associative neural network has a memory that can coincide with the training set. If new data becomes available, the network further improves its predictive ability and provides a reasonable approximation of the unknown function without a need to retrain the neural network ensemble. This feature of the method dramatically improves its predictive ability over traditional neural networks and k-nearest neighbor techniques, as demonstrated using several artificial data sets and a program to predict lipophilicity of chemical compounds. Another important feature of ASNN is the possibility to interpret neural network results by analysis of correlations between data cases in the space of models. It is shown that analysis of such correlations makes it possible to provide "property-targeted" clustering of data. The possible applications and importance of ASNN in drug design and medicinal and combinatorial chemistry are discussed. The method is available on-line at http://www.vcclab.org/lab/asnn.

Neural Networks, Computer↗

Application of associative neural networks for prediction of lipophilicity in ALOGPS 2.1 program.

This article provides a systematic study of several important parameters of the Associative Neural Network (ASNN), such as the number of networks in the ensemble, distance measures, neighbor functions, selection of smoothing parameters, and strategies for the user-training feature of the algorithm. The performance of the different methods is assessed with several training/test sets used to predict lipophilicity of chemical compounds. The Spearman rank-order correlation coefficient and Parzen-window regression methods provide the best performance of the algorithm. If additional user data is available, an improved prediction of lipophilicity of chemicals up to 2-5 times can be calculated when the appropriate smoothing parameters for the neural network are selected. The detected best combinations of parameters and strategies are implemented in the ALOGPS 2.1 program that is publicly available at http://www.vcclab.org/lab/alogps.

Journal Article↗

Benchmarking of linear and nonlinear approaches for quantitative structure-property relationship studies of metal complexation with ionophores.

A benchmark of several popular methods, Associative Neural Networks (ANN), Support Vector Machines (SVM), k Nearest Neighbors (kNN), Maximal Margin Linear Programming (MMLP), Radial Basis Function Neural Network (RBFNN), and Multiple Linear Regression (MLR), is reported for quantitative-structure property relationships (QSPR) of stability constants logK1 for the 1:1 (M:L) and logbeta2 for 1:2 complexes of metal cations Ag+ and Eu3+ with diverse sets of organic molecules in water at 298 K and ionic strength 0.1 M. The methods were tested on three types of descriptors: molecular descriptors including E-state values, counts of atoms determined for E-state atom types, and substructural molecular fragments (SMF). Comparison of the models was performed using a 5-fold external cross-validation procedure. Robust statistical tests (bootstrap and Kolmogorov-Smirnov statistics) were employed to evaluate the significance of calculated models. The Wilcoxon signed-rank test was used to compare the performance of methods. Individual structure-complexation property models obtained with nonlinear methods demonstrated a significantly better performance than the models built using multilinear regression analysis (MLRA). However, the averaging of several MLRA models based on SMF descriptors provided as good of a prediction as the most efficient nonlinear techniques. Support Vector Machines and Associative Neural Networks contributed in the largest number of significant models. Models based on fragments (SMF descriptors and E-state counts) had higher prediction ability than those based on E-state indices. The use of SMF descriptors and E-state counts provided similar results, whereas E-state indices lead to less significant models. The current study illustrates the difficulties of quantitative comparison of different methods: conclusions based only on one data set without appropriate statistical tests could be wrong.

Algorithms↗