Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

SVM-based feature selection for characterization of focused compound collections.

Artificial neural networks, the support vector machine (SVM), and other machine learning methods for the classification of molecules are often considered as a "black box", since the molecular features that are most relevant for a given classifier are usually not presented in a human-interpretable form. We report on an SVM-based algorithm for the selection of relevant molecular features from a trained classifier that might be important for an understanding of ligand-receptor interactions. The original SVM approach was extended to allow for feature selection. The method was applied to characterize focused libraries of enzyme inhibitors. A comparison with classical Kolmogorov-Smirnov (KS)-based feature selection was performed. In most of the applications the SVM method showed sustained classification accuracy, thereby relying on a smaller number of molecular features than KS-based classifiers. In one case both methods produced comparable results. Limiting the calculation of descriptors to only the most relevant ones for a certain biological activity can also be used to speed up high-throughput virtual screening.

Algorithms↗

Using surrogate modeling in the prediction of fibrinogen adsorption onto polymer surfaces.

We present a Surrogate (semiempirical) Model for prediction of protein adsorption onto the surfaces of biodegradable polymers that have been designed for tissue engineering applications. The protein used in these studies, fibrinogen, is known to play a key role in blood clotting. Therefore, fibrinogen adsorption dictates the performance of implants exposed to blood. The Surrogate Model combines molecular modeling, machine learning and an Artificial Neural Network. This novel approach includes an accounting for experimental error using a Monte Carlo analysis. Briefly, measurements of human fibrinogen adsorption were obtained for 45 polymers. A total of 106 molecular descriptors were generated for each polymer. Of these, 102 descriptors were computed using the Molecular Operating Environment (MOE) software based upon the polymer chemical structures, two represented different monomer types, and two were measured experimentally. The Surrogate Model was developed in two stages. In the first stage, the three descriptors with the highest correlation to adsorption were determined by calculating the information gain of each descriptor. Here a Monte Carlo approach enabled a direct assessment of the effect of the experimental uncertainty on the results. The three highest-ranking descriptors, defined as those with the highest information gain for the sample set, were then selected as the input variables for the second stage, an Artificial Neural Network (ANN) to predict fibrinogen adsorption. The ANN was trained using one-half of the experimental data set (the training set) selected at random. The effect of experimental error on predictive capability was again explored using a Monte Carlo analysis. The accuracy of the ANN was assessed by comparison of the predicted values for fibrinogen adsorption with the experimental data for the remaining polymers (the validation set). The mean value of the Pearson correlation coefficient for the validation data sets was 0.54 +/- 0.12. The average root-mean-square (relative) error in prediction for the validation data sets is 38%. This is an order of magnitude less than the range of experimental values (i.e., 366%) and compares favorably with the average percent relative standard deviation of the experimental measurements (i.e., 17.9%). The effects of each of the user-defined parameters in the ANN were explored. None were observed to have a significant effect on the results. Thus, the Surrogate Model can be used to accurately and unambiguously identify polymers whose fibrinogen absorption is at the limits of the range (i.e., low or high) which is an essential requirement for assessing polymers for regenerative tissue applications.

Adsorption↗

Noise reduction method for molecular interaction energy: application to in silico drug screening and in silico target protein screening.

We developed a new method to improve the accuracy of molecular interaction data using a molecular interaction matrix. This method was applied to enhance the database enrichment of in silico drug screening and in silico target protein screening using a protein-compound affinity matrix calculated by a protein-compound docking software. Our assumption was that the protein-compound binding free energy of a compound could be improved by a linear combination of its docking scores with many different proteins. We proposed two approaches to determine the coefficients of the linear combination. The first approach is based on similarity among the proteins, and the second is a machine-learning approach based on the known active compounds. These methods were applied to in silico screening of the active compounds of several target proteins and in silico target protein screening.

Pharmaceutical Preparations↗

SMIREP: predicting chemical activity from SMILES.

Most approaches to structure-activity-relationship (SAR) prediction proceed in two steps. In the first step, a typically large set of fingerprints, or fragments of interest, is constructed (either by hand or by some recent data mining techniques). In the second step, machine learning techniques are applied to obtain a predictive model. The result is often not only a highly accurate but also hard to interpret model. In this paper, we demonstrate the capabilities of a novel SAR algorithm, SMIREP, which tightly integrates the fragment and model generation steps and which yields simple models in the form of a small set of IF-THEN rules. These rules contain SMILES fragments, which are easy to understand to the computational chemist. SMIREP combines ideas from the well-known IREP rule learner with a novel fragmentation algorithm for SMILES strings. SMIREP has been evaluated on three problems: the prediction of binding activities for the estrogen receptor (Environmental Protection Agency's (EPA's) Distributed Structure-Searchable Toxicity (DSSTox) National Center for Toxicological Research estrogen receptor (NCTRER) Database), the prediction of mutagenicity using the carcinogenic potency database (CPDB), and the prediction of biodegradability on a subset of the Environmental Fate Database (EFDB). In these applications, SMIREP has the advantage of producing easily interpretable rules while having predictive accuracies that are comparable to those of alternative state-of-the-art techniques.

Algorithms↗

An efficient in silico screening method based on the protein-compound affinity matrix and its application to the design of a focused library for cytochrome P450 (CYP) ligands.

A new method has been developed to design a focused library based on available active compounds using protein-compound docking simulations. This method was applied to the design of a focused library for cytochrome P450 (CYP) ligands, not only to distinguish CYP ligands from other compounds but also to identify the putative ligands for a particular CYP. Principal component analysis (PCA) was applied to the protein-compound affinity matrix, which was obtained by thorough docking calculations between a large set of protein pockets and chemical compounds. Each compound was depicted as a point in the PCA space. Compounds that were close to the known active compounds were selected as candidate hit compounds. A machine-learning technique optimized the docking scores of the protein-compound affinity matrix to maximize the database enrichment of the known active compounds, providing an optimized focused library.

Artificial Intelligence↗

Pattern recognitiion and structure-activity relationship studies. Computer-assisted prediction of antitumor activity in structurally diverse drugs in an experimental mouse brain tumor system.

This paper reports the application of pattern recognition and substructural analysis to the problem of predicting the antineoplastic activity of 24 test compounds in an experimental mouse brain tumor system based on 138 structurally diverse compounds tested in this tumor system. The molecules were represented by three types of substructural fragments, the augmented atom, the heteropath, and the ring fragments. Of the two pattern recognition methods used to predict the activity of the test compounds the nearest neighbor method predicted 83% correctly while the learning machine method predicted 92% correctly. The test structures and the important substructural fragments used in this study are given and the implications of these results are discussed.

Animals↗

A consideration for structure-taste correlations of perillartines using pattern-recognition techniques.

The relationships between molecular structure and taste quality (sweet or bitter) or several perillartine derivatives were investigated using pattern-recognition techniques. For the classification of these compounds into two classes (sweet or bitter), a significant discriminant function was developed by the use of linear learning machine. All the compounds were assigned correctly to their observed taste classes by the function involving three parameters (one hydrophobic and two steric). In addition, the K-L transformation technique was used for examination of classification results.

Cyclohexenes↗

New approach to pharmacophore mapping and QSAR analysis using inductive logic programming. Application to thermolysin inhibitors and glycogen phosphorylase B inhibitors.

A key problem in QSAR is the selection of appropriate descriptors to form accurate regression equations for the compounds under study. Inductive logic programming (ILP) algorithms are a class of machine-learning algorithms that have been successfully applied to a number of SAR problems. Unlike other QSAR methods, which use attributes to describe chemical structure, ILP uses relations. This gives ILP the advantages of not requiring explicit superimposition of individual compounds in a dataset, of dealing naturally with multiple conformations, and of using a language much closer to that used normally by chemists. We unify ILP and standard regression techniques to give a QSAR method that has the strength of ILP at describing steric structure with the familiarity and power of regression methods. Complex pharmacophores, correlating with activity, were identified and used as new indicator variables, along with the comparative molecular field analysis (CoMFA) prediction, to form predictive regression equations. We compared the formation of 3D-QSARs using standard CoMFA with the use of ILP on the well-studied thermolysin zinc protease inhibitor dataset and a glycogen phosphorylase inhibitor dataset. In each case the addition of ILP variables produced statistically better results (P < 0.01 for thermolysin and P < 0.05 for GP datasets) than the CoMFA analysis. Moreover, the new ILP variables were not found to increase the complexity of the final QSAR equations and gave possible insight into the binding mechanism of the ligand-protein complex under study.

Algorithms↗

Finding more needles in the haystack: A simple and efficient method for improving high-throughput docking results.

The technology underpinning high-throughput docking (HTD) has developed over the past few years to where it has become a vital tool in modern drug discovery. Although the performance of various docking algorithms is adequate, the ability to accurately and consistently rank compounds using a scoring function remains problematic. We show that by employing a simple machine learning method (naïve Bayes) it is possible to significantly overcome this deficiency. Compounds from the Available Chemical Directory (ACD), along with known active compounds, were docked into two protein targets using three software packages. In cases where HTD alone was able to show some enrichment, the application of naïve Bayes was able to improve upon the enrichment. The application of this methodology to enrich HTD results can be carried out without a priori knowledge of the activity of compounds and results in superior enrichment of known actives compared to the use of scoring methods alone.

Bayes Theorem↗

Combination of a naive Bayes classifier with consensus scoring improves enrichment of high-throughput docking results.

We have previously shown that a machine learning technique can improve the enrichment of high-throughput docking (HTD) results. In the previous cases studied, however, the application of a naive Bayes classifier failed to improve enrichment for instances where HTD alone was unable to generate an acceptable enrichment. We present here a protocol to rescue poor docking results a priori using a combination of rank-by-median consensus scoring and naive Bayesian categorization.

Algorithms↗

Nonlinear quantitative structure-activity relationship for the inhibition of dihydrofolate reductase by pyrimidines.

A novel method for quantitative structure-activity relationship (QSAR) analysis is presented. The method, which does not assume any particular functional form for the QSAR, develops nonlinear relationships between parameters describing a set of molecules and the activity of the molecules. For the QSAR of the inhibition of Escherichia coli dihydrofolate reductase by 2,4-diamino-5-(substituted benzyl)pyrimidines, the method compares favorably to other nonlinear methods. Cross-validation trials demonstrate that the predictive ability is as accurate as other methods, and the method is simpler and faster than neural network and machine-learning methods. Consequently, its implementation is much easier, and interpretation of the generated QSAR is more straightforward.

Escherichia coli↗

Accurate quantitative structure-property relationship model to predict the solubility of C60 in various solvents based on a novel approach using a least-squares support vector machine.

A least-squares support vector machine (LSSVM) was used for the first time as a novel machine-learning technique for the prediction of the solubility of C60 in a large number of diverse solvents using calculated molecular descriptors from the molecular structure alone and on the basis of the software CODESSA as inputs. The heuristic method of CODESSA was used to select the correlated descriptors and build the linear model. Both the linear and the nonlinear models can give very satisfactory prediction results: the square of the correlation coefficient R(2) was 0.892 and 0.903, and the root-mean-square error was 0.126 and 0.116, respectively, for the whole data set. The prediction result of the LSSVM model is better than that obtained by the heuristic method and the reference, which proved LSSVM was a useful tool in the prediction of the solubility of C60. In addition, this paper provided a new and effective method for predicting the solubility of C60 from its structures and gave some insight into the structural features related to the solubility of C60 in different solvents.

Electrochemistry↗

QSAR and classification study of 1,4-dihydropyridine calcium channel antagonists based on least squares support vector machines.

The least squares support vector machine (LSSVM), as a novel machine learning algorithm, was used to develop quantitative and classification models as a potential screening mechanism for a novel series of 1,4-dihydropyridine calcium channel antagonists for the first time. Each compound was represented by calculated structural descriptors that encode constitutional, topological, geometrical, electrostatic, quantum-chemical features. The heuristic method was then used to search the descriptor space and select the descriptors responsible for activity. Quantitative modeling results in a nonlinear, seven-descriptor model based on LSSVM with mean-square errors 0.2593, a predicted correlation coefficient (R(2)) 0.8696, and a cross-validated correlation coefficient (R(cv)(2)) 0.8167. The best classification results are found using LSSVM: the percentage (%) of correct prediction based on leave one out cross-validation was 91.1%. This paper provides a new and effective method for drug design and screening.

Algorithms↗

Warmr: a data mining tool for chemical data.

Data mining techniques are becoming increasingly important in chemistry as databases become too large to examine manually. Data mining methods from the field of Inductive Logic Programming (ILP) have potential advantages for structural chemical data. In this paper we present Warmr, the first ILP data mining algorithm to be applied to chemoinformatic data. We illustrate the value of Warmr by applying it to a well studied database of chemical compounds tested for carcinogenicity in rodents. Data mining was used to find all frequent substructures in the database, and knowledge of these frequent substructures is shown to add value to the database. One use of the frequent substructures was to convert them into probabilistic prediction rules relating compound description to carcinogenesis. These rules were found to be accurate on test data, and to give some insight into the relationship between structure and activity in carcinogenesis. The substructures were also used to prove that there existed no accurate rule, based purely on atom-bond substructure with less than seven conditions, that could predict carcinogenicity. This results put a lower bound on the complexity of the relationship between chemical structure and carcinogenicity. Only by using a data mining algorithm, and by doing a complete search, is it possible to prove such a result. Finally the frequent substructures were shown to add value by increasing the accuracy of statistical and machine learning programs that were trained to predict chemical carcinogenicity. We conclude that Warmr, and ILP data mining methods generally, are an important new tool for analysing chemical databases.

Algorithms↗

Does size really matter--using a decision tree approach for comparison of three different databases from the medical field of acute appendicitis.

Decision trees have been successfully used for years in many medical decision making applications. Transparent representation of acquired knowledge and fast algorithms made decision trees one of the most often used symbolic machine learning approaches. This paper concentrates on the problem of separating acute appendicitis, which is a special problem of acute abdominal pain, from other diseases that cause acute abdominal pain by use of an decision tree approach. Early and accurate diagnosing of acute appendicitis is still a difficult and challenging problem in everyday clinical routine. An important factor in the error rate is poor discrimination between acute appendicitis and other diseases that cause acute abdominal pain. This error rate is still high, despite considerable improvements in history-taking and clinical examination, computer-aided decision-support, and special investigation such as ultrasound. We investigated three databases of different size with cases of acute abdominal pain to complete this task as successful as possible. The results show that the size of the database does not necessary directly influence the success of the decision tree built on it. Surprisingly we got the best results from the decision trees built on the smallest and the biggest database, where the database with medium size (relative to the other two) was not so successful. Despite this we were able to produce decision tree classifiers that were capable of producing correct decisions on test data sets with accuracy up to 84%, sensitivity to acute appendicitis up to 90%, and specificity up to 80% on the same test set.

Acute Disease↗

Generating decision trees from otoneurological data with a variable grouping method.

When medical data sets are modelled by machine learning methods, wealth of variables may be available. This paper deals with variable selection for decision tree induction in the context of two otoneurological data sets: vertigo data, and postoperative nausea and vomiting data. First, a variable grouping method based on measures of association and graph theoretic techniques was used to gain insight into data. Then, representations of learning data were defined using the information from discovered variable groups, and decision trees were generated. The use of variable grouping method was beneficial by revealing interesting associations between variables and enabling generation of accurate and reasonable decision trees that modelled the application areas from different viewpoints.

Data Collection↗

Genetic algorithm based system for patient scheduling in highly constrained situations.

In medicine and health care there are a lot of situations when patients have to be scheduled on different devices and/or with different physicians or therapists. It may concern preventive examinations, laboratory tests or convalescent therapies, therefore we are always looking for an optimal schedule that would result in finishing all the activities scheduled as soon as possible, with the least patient waiting time and maximum device utilization. Since patient scheduling is a highly complex problem, it is impossible to make a qualitative schedule by hand or even with exact heuristic methods. Therefore we developed a powerful automated scheduling method for highly constrained situations based on genetic algorithms and machine learning. In this paper we present the method, together with the whole process of schedule generation, the important parameters to direct the evolution and how the algorithm is guaranteed to produce only feasible solutions, not breaking any of the required constraints. We applied the described method to a problem of scheduling patients with different therapy needs to a limited number of therapeutic devices, but the algorithm can be easily modified for use in similar situations. The results are quite encouraging and since all the solutions are feasible, the method can be easily incorporated into an interactive user interface, which can be of major importance when scheduling patients, and human resources in general, is considered.

Algorithms↗

Global landscape of protein complexes in the yeast Saccharomyces cerevisiae.

Identification of protein-protein interactions often provides insight into protein function, and many cellular processes are performed by stable protein complexes. We used tandem affinity purification to process 4,562 different tagged proteins of the yeast Saccharomyces cerevisiae. Each preparation was analysed by both matrix-assisted laser desorption/ionization-time of flight mass spectrometry and liquid chromatography tandem mass spectrometry to increase coverage and accuracy. Machine learning was used to integrate the mass spectrometry scores and assign probabilities to the protein-protein interactions. Among 4,087 different proteins identified with high confidence by mass spectrometry from 2,357 successful purifications, our core data set (median precision of 0.69) comprises 7,123 protein-protein interactions involving 2,708 proteins. A Markov clustering algorithm organized these interactions into 547 protein complexes averaging 4.9 subunits per complex, about half of them absent from the MIPS database, as well as 429 additional interactions between pairs of complexes. The data (all of which are available online) will help future studies on individual proteins as well as functional genomics and systems biology.

Biological Evolution↗