Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Identifying interaction sites in "recalcitrant" proteins: predicted protein and RNA binding sites in rev proteins of HIV-1 and EIAV agree with experimental data.

Protein-protein and protein nucleic acid interactions are vitally important for a wide range of biological processes, including regulation of gene expression, protein synthesis, and replication and assembly of many viruses. We have developed machine learning approaches for predicting which amino acids of a protein participate in its interactions with other proteins and/or nucleic acids, using only the protein sequence as input. In this paper, we describe an application of classifiers trained on datasets of well-characterized protein-protein and protein-RNA complexes for which experimental structures are available. We apply these classifiers to the problem of predicting protein and RNA binding sites in the sequence of a clinically important protein for which the structure is not known: the regulatory protein Rev, essential for the replication of HIV-1 and other lentiviruses. We compare our predictions with published biochemical, genetic and partial structural information for HIV-1 and EIAV Rev and with our own published experimental mapping of RNA binding sites in EIAV Rev. The predicted and experimentally determined binding sites are in very good agreement. The ability to predict reliably the residues of a protein that directly contribute to specific binding events--without the requirement for structural information regarding either the protein or complexes in which it participates--can potentially generate new disease intervention strategies.

Amino Acid Sequence↗

Integrated approach for designing medical decision support systems with knowledge extracted from clinical databases by statistical methods.

In clinical research data is often studied by a particular method without previous analysis of quality or semantic contents which could link clinical database and data analytical (e.g. statistical) procedures. In order to avoid bias caused by this situation, we propose that the analysis of medical data should be divided into two main steps. In the first one we concentrate on conducting the quality, semantic and structure analyses. In the second step our aim is to build an appropriate dictionary of data analysis methods for further knowledge extraction. Methods like robust statistical techniques, procedures for mixed continuous and discrete data, fuzzy linguistic approach, machine learning and neural networks can be included. The results may be evaluated both using test samples and applying other relevant data-analytical techniques to the particular problem under the study.

Artificial Intelligence↗

A diagnostic expert system for colonic lesions.

The diagnostic expert system for colonic lesions (DESCL) was designed to discriminate colonic adenoma and adenocarcinoma from normal colonic tissue. Although it was originally developed for use in conjunction with a machine vision analytic system, the DESCL has evolved into a teaching tool and a model for conceptual machine learning. The expert system is table driven and consists of a shell and a knowledge base. The latter comprises a series of architectural and cytologic observations and a quantitative estimate of diagnostic importance relating these observations to diagnostic outcome. In a validation study of 100 colonic lesions, the expert system achieved a success rate of 98%. It has the flexibility to allow individual pathologists to "customize" the knowledge base to suit their diagnostic criteria.

Adenocarcinoma↗

Evidence for synaptic plasticity in the cerebellar cortex.

The learning machine model of the cerebellum by Marr and Albus contains a special type of synaptic plasticity. Experimental evidence for this synaptic plasticity has been meager, but very recently positive evidence has become available. Ito, Sakurai and Tongroach (7) demonstrated the occurrence of a long-lasting depression in mossy fiber responsiveness of Purkinje cells subsequent to conjuctive stimulation of mossy fibers and climbing fibers. A similar long-lasting depression was shown to occurred in sensitivity of Purkinje cell dendrites to a putative neurotransmitter of parallel fibers, i.e., glutamate. Furthermore, Ito and Kano (6) produced a long-lasting depression in the molecular layer of the cerebellar cortex by simultaneous direct stimulation of parallel fibers and climbing fibers. These long-lasting depressions appear to represent a synaptic plasticity of the form proposed by Albus.

Animals↗

Constructive induction and protein tertiary structure prediction.

To date, the only methods that have been used successfully to predict protein structures have been based on identifying homologous proteins whose structures are known. However, such methods are limited by the fact that some proteins have similar structure but no significant sequence homology. We consider two ways of applying machine learning to facilitate protein structure prediction. We argue that a straightforward approach will not be able to improve the accuracy of classification achieved by clustering by alignment scores alone. In contrast, we present a novel constructive induction approach that learns better representations of amino acid sequences in terms of physical and chemical properties. Our learning method combines knowledge and search to shift the representation of sequences so that semantic similarity is more easily recognized by syntactic matching. Our approach promises not only to find new structural relationships among protein sequences, but also expands our understanding of the roles knowledge can play in learning via experience in this challenging domain.

Artificial Intelligence↗

Transmembrane segment prediction from protein sequence data.

We consider the automated identification of transmembrane domains in membrane protein sequences. 324 proteins (containing 1585 segments) were examined, representing every protein in the PIR database having the transmembrane domain feature annotation. Machine learning techniques were used to evaluate the efficacy of alternative hydrophobicity measures and windowing techniques. We describe a simpler measure of hydrophobicity and a new variable window size concept. We demonstrate that these techniques are superior to some previous techniques in minimizing the segment error rate. Using these new techniques, we describe an algorithm that has a 7.9% segment error rate on the sampled proteins, while classifying 16.7% of the amino acid residues as transmembrane.

Algorithms↗

Inductive logic programming used to discover topological constraints in protein structures.

This paper describes the application of the Inductive Logic Programming (ILP) program GOLEM to the discovery of constraints in the packing of beta-sheets in alpha/beta proteins. These constraints (rules) have a role in understanding the protein folding problem. Constraints were learnt for four features of beta-sheet packing: the winding direction of two sequential strands, whether two consecutive strands pack parallel or anti-parallel, whether two strands pack adjacently, and whether a beta-strand is at an edge. Investigation of the learnt constraints revealed interesting patterns, some of which were previously known, others that were novel. Novel features include the discovery: that the relationship between pairs of sequential strands is in general one of decreasing size, and that more sequential pairs of strands wind in the direction out than the direction in. We conclude that machine learning has a useful place in molecular biology as a pattern discovery tool.

Animals↗

Predicting protein folding classes without overly relying on homology.

An important open problem in molecular biology is how to use computational methods to understand the structure and function of proteins given only their primary sequences. We describe and evaluate an original machine-learning approach to classifying protein sequences according to their structural folding class. Our work is novel in several respects: we use a set of protein classes that previously have not been used for classifying primary sequences, and we use a unique set of attributes to represent protein sequences to the learners. We evaluate our approach by measuring its ability to correctly classify proteins that were not in its training set. We compare our input representation to a commonly used input representation--amino acid composition--and show that our approach more accurately classifies proteins that have very limited homology to the sequences on which the systems are trained.

Algorithms↗

An inductive algorithm approach to knowledge acquisition for expert system development. A pilot study.

Knowledge acquisition, which consists of knowledge elicitation and knowledge representation, often is considered the weakest link in the design of expert systems. Systems frequently are built on the knowledge of one expert and require extensive use of knowledge engineering techniques to elicit this knowledge from the expert. Inductive algorithms are a potential alternative method of knowledge acquisition for expert system development. The aim of this pilot study was to examine the feasibility of applying machine learning techniques, specifically, inductive algorithms, to an existing research database as a method for knowledge elicitation and knowledge representation for expert system development. Two inductive algorithms (C4 and Classification and Regression Trees [CART]) that generate decision trees were selected for the analysis using a data set of 201 patients hospitalized for Pneumocystis carinii pneumonia. Neither C4 nor CART produced trees with an accuracy that was significantly better than the baseline accuracy (71.3%) for prediction of outcome in the data set. The mean accuracy of the C4 decision trees was below baseline and the mean accuracy of CART decision trees was 74.6%. The experts found both algorithms comprehensible, but not adequate, and identified important missing predictor variables. The study findings suggest that additional research is needed to examine the appropriate use of inductive algorithms in the transformation of nursing data and information into nursing knowledge.

Algorithms↗

Artificial intelligence within the chemical laboratory.

Various techniques within the area of artificial intelligence such as expert systems and neural networks may play a role during the problem-solving processes within the clinical biochemical laboratory. Neural network analysis provides a non-algorithmic approach to information processing, which results in the ability of the computer to form associations and to recognize patterns or classes among data. It belongs to the machine learning techniques which also include probabilistic techniques such as discriminant function analysis and logistic regression and information theoretical techniques. These techniques may be used to extract knowledge from example patients to optimize decision limits and identify clinically important laboratory quantities. An expert system may be defined as a computer program that can give advice in a well-defined area of expertise and is able to explain its reasoning. Declarative knowledge consists of statements about logical or empirical relationships between things. Expert systems typically separate declarative knowledge residing in a knowledge base from the inference engine: an algorithm that dynamically directs and controls the system when it searches its knowledge base. A tool is an expert system without a knowledge base. The developer of an expert system uses a tool by entering knowledge into the system. Many, if not the majority of problems encountered at the laboratory level are procedural. A problem is procedural if it is possible to write up a step-by-step description of the expert's work or if it can be represented by a decision tree. To solve problems of this type only small expert system tools and/or conventional programming are required.(ABSTRACT TRUNCATED AT 250 WORDS)

Artificial Intelligence↗

Induction of medical expert system rules based on rough sets and resampling methods.

Automated knowledge acquisition is an important research issue in improving the efficiency of medical expert systems. Rules for medical expert systems consists of two parts: one is a proposition part, which represent a if-then rule, and the other is probabilistic measures, which represents reliability of that rule. Therefore, acquisition of both knowledge is very important for application of machine learning methods to medical domains. Extending concepts of rough set theory to probabilistic domain, we introduce a new approach to knowledge acquisition, which induces probabilistic rules based on rough set theory (PRIMEROSE) and develop a program that extracts rules for an expert system from clinical database, using this method. The results show that the derived rules almost correspond to those of medical experts.

Artificial Intelligence↗

Improving prediction of preterm birth using a new classification scheme and rule induction.

Prediction of preterm birth is a poorly understood domain. The existing manual methods of assessment of preterm birth are 17%-38% accurate. The machine learning system LERS was used for three different datasets about pregnant women. Rules induced by LERS were used in conjunction with a classification scheme of LERS, based on "bucket brigade algorithm" of genetic algorithms and enhanced by partial matching. The resulting prediction of preterm birth in new, unseen cases is much more accurate (68%-90%).

Algorithms↗

Knowledge discovery in clinical databases based on variable precision rough set model.

Since a large amount of clinical data are being stored electronically, discovery of knowledge from such clinical databases is one of the important growing research area in medical informatics. For this purpose, we develop KDD-R (a system for Knowledge Discovery in Databases using Rough sets), an experimental system for knowledge discovery and machine learning research using variable precision rough sets (VPRS) model, which is an extension of original rough set model. This system works in the following steps. First, it preprocesses databases and translates continuous data into discretized ones. Second, KDD-R checks dependencies between attributes and reduces spurious data. Third, the system computes rules from reduced databases. Finally, fourth, it evaluates decision making. For evaluation, this system is applied to a clinical database of meningoencephalitis, whose computational results show that several new findings are obtained.

Artificial Intelligence↗

Why Johnny can't reengineer health care processes with information technology.

Many educational institutions are developing curricula that integrate computer and business knowledge and skills concerning a specific industry, such as banking or health care. We have developed a curriculum that emphasizes, equally, medical, computer, and business management concepts. Along the way we confronted a formidable obstacle, namely the domain specificity of the reference disciplines. Knowledge within each domain is sufficiently different from other domains that it reduces the leverage of building on preexisting knowledge and skills. We review this problem from the point of view of cognitive science (in particular, knowledge representation and machine learning) to suggest strategies for coping with incommensurate domain ontologies. These strategies include reflective judgment, implicit learning, abstraction, generalization, analogy, multiple inheritance, project-orientation, selectivity, goal- and failure-driven learning, and case- and story-based learning.

Commerce↗

Information, intelligence, and interface: the pillars of a successful medical information system.

This paper addresses three key issues facing developers of clinical and/or research medical information systems. 1. INFORMATION. The basic function of every database is to store information about the phenomenon under investigation. There are many ways to organize information in a computer; however only a few will prove optimal for any real life situation. Computer Science theory has developed several approaches to database structure, with relational theory leading in popularity among end users [8]. Strict conformance to the rules of relational database design rewards the user with consistent data and flexible access to that data. A properly defined database structure minimizes redundancy i.e.,multiple storage of the same information. Redundancy introduces problems when updating a database, since the repeated value has to be updated in all locations--missing even a single value corrupts the whole database, and incorrect reports are produced [8]. To avoid such problems, relational theory offers a formal mechanism for determining the number and content of data files. These files not only preserve the conceptual schema of the application domain, but allow a virtually unlimited number of reports to be efficiently generated. 2. INTELLIGENCE. Flexible access enables the user to harvest additional value from collected data. This value is usually gained via reports defined at the time of database design. Although these reports are indispensable, with proper tools more information can be extracted from the database. For example, machine learning, a sub-discipline of artificial intelligence, has been successfully used to extract knowledge from databases of varying size by uncovering a correlation among fields and records[1-6, 9]. This knowledge, represented in the form of decision trees, production rules, and probabilistic networks, clearly adds a flavor of intelligence to the data collection and manipulation system. 3. INTERFACE. Despite the obvious importance of collecting data and extracting knowledge, current systems often impede these processes. Problems stem from the lack of user friendliness and functionality. To overcome these problems, several features of a successful human-computer interface have been identified [7], including the following "golden" rules of dialog design [7]: consistency, use of shortcuts for frequent users, informative feedback, organized sequence of actions, simple error handling, easy reversal of actions, user-oriented focus of control, and reduced short-term memory load. To this list of rules, we added visual representation of both data and query results, since our experience has demonstrated that users react much more positively to visual rather than textual information. In our design of the Orthopaedic Trauma Registry--under development at the Carolinas Medical Center--we have made every effort to follow the above rules. The results were rewarding--the end users actually not only want to use the product, but also to participate in its development.

Artificial Intelligence↗

A generalized hidden Markov model for the recognition of human genes in DNA.

We present a statistical model of genes in DNA. A Generalized Hidden Markov Model (GHMM) provides the framework for describing the grammar of a legal parse of a DNA sequence (Stormo & Haussler 1994). Probabilities are assigned to transitions between states in the GHMM and to the generation of each nucleotide base given a particular state. Machine learning techniques are applied to optimize these probabilities using a standardized training set. Given a new candidate sequence, the best parse is deduced from the model using a dynamic programming algorithm to identify the path through the model with maximum probability. The GHMM is flexible and modular, so new sensors and additional states can be inserted easily. In addition, it provides simple solutions for integrating cardinality constraints, reading frame constraints, "indels", and homology searching. The description and results of an implementation of such a gene-finding model, called Genie, is presented. The exon sensor is a codon frequency model conditioned on windowed nucleotide frequency and the preceding codon. Two neural networks are used, as in (Brunak, Engelbrecht, & Knudsen 1991), for splice site prediction. We show that this simple model performs quite well. For a cross-validated standard test set of 304 genes [ftp:@www-hgc.lbl.gov/pub/genesets] in human DNA, our gene-finding system identified up to 85% of protein-coding bases correctly with a specificity of 80%. 58% of exons were exactly identified with a specificity of 51%. Genie is shown to perform favorably compared with several other gene-finding systems.

Chromosomes, Human↗

Prediction of enzyme classification from protein sequence without the use of sequence similarity.

We describe a novel approach for predicting the function of a protein from its amino-acid sequence. Given features that can be computed from the amino-acid sequence in a straightforward fashion (such as pI, molecular weight, and amino-acid composition), the technique allows us to answer questions such as: Is the protein an enzyme? If so, in which Enzyme Commission (EC) class does it belong? Our approach uses machine learning (ML) techniques to induce classifiers that predict the EC class of an enzyme from features extracted from its primary sequence. We report on a variety of experiments in which we explored the use of three different ML techniques in conjunction with training datasets derived from PDB and from Swiss-Prot. We also explored the use of several different feature sets. Our method is able to predict the first EC number of an enzyme with 74% accuracy (thereby assigning the enzyme to one of six broad categories of enzyme function), and to predict the second EC number of an enzyme with 68% accuracy (thereby assigning the enzyme to one of 57 subcategories of enzyme function). This technique could be a valuable complement to sequence-similarity searches and to pathway-analysis methods.

Algorithms↗