Search PubMed⌕ Search

Biomedical subjects

P C Jurs

Publications and source records attributed to P C Jurs.

At least 19 recordsLinked to original sources

Development of binary classification of structural chromosome aberrations for a diverse set of organic compounds from molecular structure.

Classification models are generated to predict in vitro cytogenetic results for a diverse set of 383 organic compounds. Both k-nearest neighbor and support vector machine models are developed. They are based on calculated molecular structure descriptors. Endpoints used are the labels clastogenic or nonclastogenic according to an in vitro chromosomal aberration assay with Chinese hamster lung cells. Compounds that were tested with both a 24 and 48 h exposure are included. Each compound is represented by calculated molecular structure descriptors encoding the topological, electronic, geometrical, or polar surface area aspects of the structure. Subsets of informative descriptors are identified with genetic algorithm feature selection coupled to the appropriate classification algorithm. The overall classification success rate for a k-nearest neighbor classifier built with just six topological descriptors is 81.2% for the training set and 86.5% for an external prediction set. The overall classification success rate for a three-descriptor support vector machine model is 99.7% for the training set, 92.1% for the cross-validation set, and 83.8% for an external prediction set.

Algorithms↗

Linear regression and computational neural network prediction of tetrahymena acute toxicity for aromatic compounds from molecular structure.

A quantitative structure toxicity relationship (QSTR) has been derived for a diverse set of 448 industrially important aromatic solvents. Toxicity was expressed as the 50% growth impairment concentration (ICG(50)) for the ciliated protozoa Tetrahymena and spans the range -1.46 to 3.36 log units. Molecular descriptors that encode topological, geometrical, electronic, and hybrid geometrical-electronic structural features were calculated for each compound. Subsets of molecular descriptors were selected via a simulated annealing technique and a genetic algorithm. From this reduced pool of descriptors, multiple linear regression models and nonlinear models using computational neural networks (CNNs) were derived and then used to predict the ICG(50) values for an external set of representative compounds. An average of 10 nonlinear CNN models with 11-5-1 architecture was found to best describe the system with root-mean-square errors of 0.28, 0.29, and 0.34 log units for the training, cross validation, and prediction sets, respectively.

Animals↗

Classification of multidrug-resistance reversal agents using structure-based descriptors and linear discriminant analysis.

Linear discriminant analysis is used to generate models to classify multidrug-resistance reversal agents based on activity. Models are generated and evaluated using multidrug-resistance reversal activity values for 609 compounds measured using adriamycin-resistant P388 murine leukemia cells. Structure-based descriptors numerically encode molecular features which are used in model formation. Two types of models are generated: one type to classify compounds as inactive, moderately active, and active (three-class problem) and one type to classify compounds as inactive or active without considering the moderately active class (two-class problem). Two activity distributions are considered, where the separation between inactive and active compounds is different. When the separation between inactive and active classes is small, a model based on nine topological descriptors is developed that produces a classification rate of 83.1% correct for an external prediction set. Larger separation between active and inactive classes raises the prediction set classification rate to 92.0% correct using a model with six topological descriptors. Models are further validated through Monte Carlo experiments in which models are generated after class labels have been scrambled. The classification rates achieved demonstrate that the models developed could serve as a screening mechanism to identify potentially useful MDRR agents from large libraries of compounds.

Animals↗

Prediction of fathead minnow acute toxicity of organic compounds from molecular structure.

Interest in the prediction of toxicity without the use of experimental data is growing, and quantitative structure-activity relationship (QSAR) methods are valuable for such predictions. A QSAR study of acute aqueous toxicity of 375 diverse organic compounds has been developed using only calculated structural features as independent variables. Toxicity is expressed as -log(LD(50)) with the units -log(millimoles per liter) and ranges from -3 to 6. Multiple linear regression and computational neural networks (CNNs) are utilized for model building. The best model is a nonlinear CNN model based on eight calculated molecular structure descriptors. The root-mean-square log(LD(50)) errors for the training, cross-validation, and prediction sets of this CNN model are 0.71, 0.77, and 0.74 -log(mmol/L), respectively. These results are compared to a previous study with the same data set which included many more descriptors and used experimental data in the descriptor pool.

Animals↗

Prediction of acute mammalian toxicity of organophosphorus pesticide compounds from molecular structure.

A quantitative structure-activity relationship (QSAR) investigation was done for the acute oral mammalian toxicity (LD50) of a set of 54 organophosphorus pesticide compounds. The compounds were represented with calculated molecular structure descriptors, which encoded their topological, electronic, and geometrical features. Feature selection was done with a genetic algorithm to find subsets of descriptors that would support a high quality computational neural network (CNN) model to link the structural descriptors to the -log(mmol/kg) values for the compounds. The best seven-descriptor non-linear CNN model found had an rms error of 0.22 log units for the training set compounds and 0.25 log units for the prediction set compounds.

Animals↗

Neural network classification and quantification of organic vapors based on fluorescence data from a fiber-optic sensor array.

Computational neural networks have been developed to classify and quantify nine organic vapors. The neural network analyses used data that consisted of the change in fluorescence from a sensor array that consisted of 19 fiber optics with immobilized dye in polymer matrices. Plots of change in fluorescence intensity versus time were measured as pulses of analyte were presented to the sensor array. Descriptors were calculated from the intensity vs time plots, and they were used to build neural network models that accurately classified and quantified each of the nine analytes. Most of the data were used to train the neural networks (training set members), some were used to assist termination of training (cross-validation set members), and some were used to validate the models (prediction set members). Classification rates approaching 100% were achieved for the training set data, and 90% of the members in the prediction set were correctly classified. In addition, 97% of the prediction set observations were assigned a correct relative concentration.

Acetates↗

Simulation of the 13C nuclear magnetic resonance spectra of trisaccharides using multiple linear regression analysis and neural networks.

Predictive models are developed for the 13C NMR chemical shifts of the carbon atoms comprising the central rings of 46 trisaccharide compounds. Thirty-nine trisaccharides are used as a training set for development of models using regression analysis and computational neural networks, and seven compounds are used as an external prediction set. The descriptors used in the models are developed directly from the molecular structures of the trisaccharides. Three different methods of descriptor selection are compared. The dependence of the models on the geometries of the trisaccharides is explored. The models developed with geometric descriptors are better than those developed without geometric descriptors, although the latter models are still of a comparable quality. Overall, the best model found is a neural network based on descriptors selected by multiple linear regression.

Algorithms↗

Quantitative structure-retention and structure-odor intensity relationships for a diverse group of odor-active compounds.

Numerical representations of structure-based features are used to estimate both the retention indexes and sweetnesses of a diverse set of industrially important fragrance compounds. Retention indexes measured on nonpolar as well as polar stationary phases are modeled with accuracies of 3.6% and 5.6% at the mean of the respective retention ranges. Similar success was achieved when the developed equations were applied to predict the retention indexes of external data set compounds. Finally, the implications of using strictly 2-D structural information versus incorporating geometrical information are explored and discussed. The intensity of sweetness attributed to each compound is quantitatively predicted using identical multiple linear regression techniques. Difficulties encountered in this portion of the study warranted a critique of the procedures used to gain access to the odor data. As a consequence, the limited control exerted over several experimental variables is questioned.

Odorants↗

Prediction of gas chromatographic relative retention times of stimulants and narcotics.

The ADAPT software system was used to create models for the prediction of gas chromatographic relative retention times (RRTs) of stimulants and narcotics that are analyzed in doping control of athletes. The two main methods that were followed for building the models were the quantitative structure-retention relationship (QSRR) and multiple linear regression analysis. The main proposed model for the entire data set had a multiple correlation coefficient R = 0.991 and standard error s = 0.046 or approximately 4.5%. Because of the relatively high standard error of the main model, a second model was built on a subset of compounds with R = 0.982 and s = 0.027 or approximately 2.5%.

Central Nervous System Stimulants↗

Prediction of gas chromatographic relative retention times of anabolic steroids.

The prediction of gas chromatographic relative retention times (RRTs) of anabolic steroids, used in the doping control of athletes, was performed by a quantitative structure-retention relationship (QSRR) and multiple linear regression analysis study. A nine-variable model was generated with a multiple correlation coefficient R = 0.991 and relative standard error of less than 3%. Preliminary results indicated that the application of the model, especially in the prediction of RRTs of metabolites of the anabolic steroids, will be helpful.

Anabolic Agents↗

Quantitative structure-retention relationship studies of odor-active aliphatic compounds with oxygen-containing functional groups.

High-quality regression equations (R greater than 0.996) modeling the gas chromatographic retention indices of 115 odor-active compounds on stationary phases of different polarity are generated by using the ADAPT software system. Multiple linear regression techniques are used to describe the statistical relationship between the Kováts retention indices and structure-based molecular descriptors. These descriptors encode topological, geometrical, and electronic features of the molecules. The utility of several new descriptors encoding functions of partial atomic charge and solvent accessible surface area is demonstrated. Quantitative predictions of odor threshold values for a subset of these compounds containing the alcohol functional group are also calculated by using a similar methodology.

Hydrocarbons↗

Cluster analysis of acrylates to guide sampling for toxicity testing.

A set of 143 acrylates drawn from the TSCA inventory have been investigated for structurally defined clusters of compounds to simplify sampling for future toxicity screening. Each acrylate was represented by eight descriptors calculated from the molecular structure. Several standard clustering methods have been used to find five natural clusters of compounds. These five clusters are largely populated by compounds with similar chemical attributes with separate clusters formed for compounds with high absolute partial atomic charges, hydrophobic compounds, small compounds, halogenated compounds, and large or oligomeric compounds.

Acrylates↗