Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Ecologic niche modeling and differentiation of populations of Triatoma brasiliensis neiva, 1911, the most important Chagas' disease vector in northeastern Brazil (hemiptera, reduviidae, triatominae).

Ecologic niche modeling has allowed numerous advances in understanding the geographic ecology of species, including distributional predictions, distributional change and invasion, and assessment of ecologic differences. We used this tool to characterize ecologic differentiation of Triatoma brasiliensis populations, the most important Chagas' disease vector in northeastern Brazil. The species' ecologic niche was modeled based on data from the Fundação Nacional de Saúde of Brazil (1997-1999) with the Genetic Algorithm for Rule-Set Prediction (GARP). This method involves a machine-learning approach to detecting associations between occurrence points and ecologic characteristics of regions. Four independent "ecologic niche models" were developed and used to test for ecologic differences among T. brasiliensis populations. These models confirmed four ecologically distinct and differentiated populations, and allowed characterization of dimensions of niche differentiation. Patterns of ecologic similarity matched patterns of molecular differentiation, suggesting that T. brasiliensis is a complex of distinct populations at various points in the process of speciation.

Animals↗

POISE: Spectral Inference of Parent-of-Origin Effects in Unlabeled Genomic Data.

MOTIVATION: Parent of Origin Effects (POEs), where the effect of an an allele on a phenotype differs based on maternal or paternal inheritance implicated in growth, metabolism, and neurodevelopment. Traditional tests for POEs require family data to determine parental origins of transmitted alleles. Given that such studies are expensive and time consuming compared to genome-wide association studies (GWAS), tests that function absent inheritance information are highly desirable. We develop a method, based on community detection from machine learning, that infers POEs via a spectral decomposition, obtains confidence intervals via a non-parametric bootstrap, and safeguards against confounding by non POE sources of variation. We refer to our method as Parent of Origin Inference via Spectral Estimation (POISE). RESULTS: We demonstrate that POISE is well-calibrated under both Gaussian and heavy-tailed noise in simulation studies, with improved robustness to true POEs compared to existing covariance-based tests. POISE provides per-trait effect estimates with bias-corrected bootstrap confidence intervals and incorporates an information-theoretic minimum detectable effect size that filters unreliable estimates, conferring robustness to covariance-deflating variance QTL. We then apply POISE to GWAS data from the UK Biobank using BMI, LDL cholesterol, and HDL cholesterol. POISE recovers established POE loci and identifies 134 additional variants at genes implicated in lipid metabolism, immune regulation, and growth. AVAILABILITY AND IMPLEMENTATION: The code for this method in Python is available at https://github.com/bystrogenomics/POISE.

Community Detection↗

An intelligent system for the tracking--localization of changes & exploratory analysis--of long-term ECG.

An interactive methodology for the analysis of long-term ECG is introduced. It is an anthropomimetic technique and consists of three parts. At the first stage, the clinicians' scan of the ECG traces is imitated and changes in the shape of QRS are quantified. In the sequel, a clinician involves in the interpretation of the most prominent changes providing the patient-dependent prototypes for the subsequent machine learning procedure. Finally, a classification scheme incorporates the portion of medical knowledge needed to explore the whole patient's ECG. This scheme, being very robust to noise, presents excellent generalization properties and can serve as a reliable automation in a future examination of the certain subject.

Arrhythmias, Cardiac↗

Using classification tree and logistic regression methods to diagnose myocardial infarction.

Early and accurate diagnosis of myocardial infarction (MI) in patients who present to the Emergency Room (ER) complaining of chest pain is an important problem in emergency medicine. A number of decision aids have been developed to assist with this problem but have not achieved general use. Machine learning techniques, including classification tree and logistic regression (LR) methods, have the potential to create simple but accurate decision aids. Both a classification tree (FT Tree) and an LR model (FT LR) have been developed to predict the probability that a patient with chest pain is having an MI based solely upon data available at time of presentation to the ER. Training data came from a data set collected in Edinburgh, Scotland. Each model was then tested on a separate Edinburgh data set, as well as on a data set from a different hospital in Sheffield, England. Previously published models, the Goldman classification tree[1] and Kennedy LR equation[2], were evaluated on the same test data sets. On the Edinburgh test set, results showed that the FT Tree, FT LR, and Kennedy LR performed equally well, with ROC curve areas of 94.04%, 94.28%, and 94.30%, respectively, while the Goldman Tree's performance was significantly poorer, with an area of 84.03%. The difference in ROC areas between the first three models and the Goldman model is significant beyond the 0.0001 level. On the Sheffield test set, results showed that the FT Tree, FT LR, and Kennedy LR ROC areas were not significantly different (p > = 0.17), while the FT Tree again outperformed the Goldman Tree (p = 0.006). Unlike previous work[3], this study indicates that classification trees, which have certain advantages over LR models, may perform as well as LR models in the diagnosis of patients with MI.

Algorithms↗

Automatic prediction of trauma registry procedure codes from emergency room dictations.

Current natural language processing techniques for recognition of concepts in the electronic medical record have been insufficient to allow their broad use for coding information automatically. We have undertaken a preliminary investigation into the use of machine learning methods to recognize procedure codes from emergency room dictations for a trauma registry. Our preliminary results indicate moderate success, and we believe future enhancements with additional learning techniques and selected natural language processing approaches will be fruitful.

Artificial Intelligence↗

Evaluating variable selection methods for diagnosis of myocardial infarction.

This paper evaluates the variable selection performed by several machine-learning techniques on a myocardial infarction data set. The focus of this work is to determine which of 43 input variables are considered relevant for prediction of myocardial infarction. The algorithms investigated were logistic regression (with stepwise, forward, and backward selection), backpropagation for multilayer perceptrons (input relevance determination), Bayesian neural networks (automatic relevance determination), and rough sets. An independent method (self-organizing maps) was then used to evaluate and visualize the different subsets of predictor variables. Results show good agreement on some predictors, but also variability among different methods; only one variable was selected by all models.

Algorithms↗

WWW search engine for Slovenian and English medical documents.

The information tool for the organization and searching of Slovenian and English medical documents is presented. The tool, partly still in development phases, performs automatic subject description of documents, searching with natural language queries and rankig of search hits according to their relevance. The search engine allows the searcher to use relevance feedback in order to perform incremental improvement of search results. The machine learning system TILDE for learning user profiles was also applied. Documents marked by the user as relevant or non-relevant are used to find characteristics that distinguish relevant documents from non-relevant ones.

Abstracting and Indexing↗

Constructing biological knowledge bases by extracting information from text sources.

Recently, there has been much effort in making databases for molecular biology more accessible and interoperable. However, information in text form, such as MEDLINE records, remains a greatly underutilized source of biological information. We have begun a research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases. Our approach to this task is to use machine-learning methods to induce routines for extracting facts from text. We describe two learning methods that we have applied to this task--a statistical text classification method, and a relational learning method--and our initial experiments in learning such information-extraction routines. We also present an approach to decreasing the cost of learning information-extraction routines by learning from "weakly" labeled training data.

Artificial Intelligence↗

Protein fold class prediction: new methods of statistical classification.

Feed forward neural networks are compared with standard and new statistical classification procedures for the classification of proteins. We applied logistic regression, an additive model and projection pursuit regression from the methods based on a posterior probabilities; linear, quadratic and a flexible discriminant analysis from the methods based on class conditional probabilities, and the K-nearest-neighbors classification rule. Both, the apparent error rate obtained with the training sample (n = 143) and the test error rate obtained with the test sample (n = 125) and the 10-fold cross validation error were calculated. We conclude that some of the standard statistical methods are potent competitors to the more flexible tools of machine learning.

Algorithms↗

Combinatorial approaches to finding subtle signals in DNA sequences.

Signal finding (pattern discovery in unaligned DNA sequences) is a fundamental problem in both computer science and molecular biology with important applications in locating regulatory sites and drug target identification. Despite many studies, this problem is far from being resolved: most signals in DNA sequences are so complicated that we don't yet have good models or reliable algorithms for their recognition. We complement existing statistical and machine learning approaches to this problem by a combinatorial approach that proved to be successful in identifying very subtle signals.

Algorithms↗

Models to predict cardiovascular risk: comparison of CART, multilayer perceptron and logistic regression.

The estimate of a multivariate risk is now required in guidelines for cardiovascular prevention. Limitations of existing statistical risk models lead to explore machine-learning methods. This study evaluates the implementation and performance of a decision tree (CART) and a multilayer perceptron (MLP) to predict cardiovascular risk from real data. The study population was randomly splitted in a learning set (n = 10,296) and a test set (n = 5,148). CART and the MLP were implemented at their best performance on the learning set and applied on the test set and compared to a logistic model. Implementation, explicative and discriminative performance criteria are considered, based on ROC analysis. Areas under ROC curves and their 95% confidence interval are 0.78 (0.75-0.81), 0.78 (0.75-0.80) and 0.76 (0.73-0.79) respectively for logistic regression, MLP and CART. Given their implementation and explicative characteristics, these methods can complement existing statistical models and contribute to the interpretation of risk.

Artificial Intelligence↗

Event discovery in medical time-series data.

Vast amounts of clinical information are generated daily on patients in the health care setting. Increasingly, this information is collected and stored for its potential utility in advancing health care. Knowledge-based systems, for example, might be able to apply rules to the collected data to determine whether a patient has a certain condition. Often, however, the underlying knowledge needed to write such rules is not well understood. How could these clinical data be useful then? Use of machine learning is one answer. We present a pipeline for discovering the knowledge needed for event detection in medical time-series data. We demonstrate how this process can be applied in the development of intelligent patient monitoring for the intensive care unit (ICU). Specifically, we develop a system for detecting Otrue alarmO situations in the ICU, where currently as many as 86% of bedside monitor alarms are false.

Artificial Intelligence↗

On prognostic models, artificial intelligence and censored observations.

The development of prognostic models for assisting medical practitioners with decision making is not a trivial task. Models need to possess a number of desirable characteristics and few, if any, current modelling approaches based on statistical or artificial intelligence can produce models that display all these characteristics. The inability of modelling techniques to provide truly useful models has led to interest in these models being purely academic in nature. This in turn has resulted in only a very small percentage of models that have been developed being deployed in practice. On the other hand, new modelling paradigms are being proposed continuously within the machine learning and statistical community and claims, often based on inadequate evaluation, being made on their superiority over traditional modelling methods. We believe that for new modelling approaches to deliver true net benefits over traditional techniques, an evaluation centric approach to their development is essential. In this paper we present such an evaluation centric approach to developing extensions to the basic k-nearest neighbour (k-NN) paradigm. We use standard statistical techniques to enhance the distance metric used and a framework based on evidence theory to obtain a prediction for the target example from the outcome of the retrieved exemplars. We refer to this new k-NN algorithm as Censored k-NN (Ck-NN). This reflects the enhancements made to k-NN that are aimed at providing a means for handling censored observations within k-NN.

Algorithms↗

A missing data treatment for data mining applications in medical information systems.

To apply user-friendly, easily operated and accessible tools to handle missing data resulting from an auto-stored medical information system, these tools are applied to satisfy general users from different disciplines (i.e. statistics and machine-learning), followed by medical information system development. This study attempts to develop a new logic separation inference method applied to a database with a format like most real-world medical records containing many missing data and miscellaneous variables. It is expected that this method should have better performance than currently accessible methods. The newly developed logic separation inference method shows a classification power of 0.997 (elimination method is 1), which is better than the simple replacing method (replaced by mode shows 0.974). Both inference methods (mode and mean) have superior classification power to the simple replacing method. The missing data treatment processes introduced in this study can be completed on a MS Excel spreadsheet without any complicated calculation; therefore, they can satisfy general users. This new missing data treatment method is only applied up to 60% of the missing data (missing at random). However, when there is large amount of data, it is expected that this method also can be applied to a database missing more than 60%.

Humans↗

Data mining of spectroscopic data for biomarker discovery.

The goals of precise diagnosis, prevention and treatment of disease can be realized through the discovery of biological markers. Spectroscopic tools can simultaneously detect and quantify multiple small molecule and macromolecular components of biological samples, and are therefore ideal methods for the discovery of previously uncharacterized markers. However, the identification of meaningful spectral features is complicated by the lack of foreknowledge of the molecular nature of a disease, spectral noise and biological variability that is uncorrelated with the disease state. Pattern recognition techniques, both statistical and machine-learning, have been increasingly used in recent years with spectroscopic data to identify markers and classify patients into disease subsets. This review summarizes recent developments, limitations and future prospects in the use of data mining techniques with magnetic resonance spectroscopy, mass spectrometry and optical spectroscopy for the discovery of biomarkers.

Biomarkers↗

Logistic regression model: an assessment of variability of predictions.

Risk prediction models available for cardiovascular prevention are statistical or based on machine learning methods. This paper investigates whether the logistic regression method can be considered as reference for validation of other methods. In order to test the stability of the predictions using this method, we performed two types of analyses on 50 random training and test samples drawn from the same database. In first analyses three models were obtained by forced entry of different sets of four variables. In second analyses, models were built with increasing number of predictive variables. The predictive performance was assessed by the area under the ROC curve. Although across-samples variability is low for a given model, it is large enough to lead to wrong conclusions when comparing different prediction methods. We also suggest that a low events-per-variable ratio alters the stability of a model's coefficients but does not affect the variability of prediction performance.

Area Under Curve↗

Comparison of three databases with a decision tree approach in the medical field of acute appendicitis.

Decision trees have been successfully used for years in many medical decision making applications. Transparent representation of acquired knowledge and fast algorithms made decision trees one of the most often used symbolic machine learning approaches. This paper concentrates on the problem of separating acute appendicitis, which is a special problem of acute abdominal pain from other diseases that cause acute abdominal pain by use of an decision tree approach. Early and accurate diagnosing of acute appendicitis is still a difficult and challenging problem in everyday clinical routine. An important factor in the error rate is poor discrimination between acute appendicitis and other diseases that cause acute abdominal pain. This error rate is still high, despite considerable improvements in history-taking and clinical examination, computer-aided decision-support and special investigation, such as ultrasound. We investigated three different large databases with cases of acute abdominal pain to complete this task as successful as possible. The results show that the size of the database does not necessary directly influence the success of the decision tree built on it. Surprisingly we got the best results from the decision trees built on the smallest and the biggest database, where the database with medium size (relative to the other two) was not so successful. Despite that we were able to produce decision tree classifiers that were capable of producing correct decisions on test data sets with accuracy up to 84%, sensitivity to acute appendicitis up to 90%, and specificity up to 80% on the same test set.

Abdomen, Acute↗

Molecular classification of human carcinomas by use of gene expression signatures.

Classification of human tumors according to their primary anatomical site of origin is fundamental for the optimal treatment of patients with cancer. Here we describe the use of large-scale RNA profiling and supervised machine learning algorithms to construct a first-generation molecular classification scheme for carcinomas of the prostate, breast, lung, ovary, colorectum, kidney, liver, pancreas, bladder/ureter, and gastroesophagus, which collectively account for approximately 70% of all cancer-related deaths in the United States. The classification scheme was based on identifying gene subsets whose expression typifies each cancer class, and we quantified the extent to which these genes are characteristic of a specific tumor type by accurately and confidently predicting the anatomical site of tumor origin for 90% of 175 carcinomas, including 9 of 12 metastatic lesions. The predictor gene subsets include those whose expression is typical of specific types of normal epithelial differentiation, as well as other genes whose expression is elevated in cancer. This study demonstrates the feasibility of predicting the tissue origin of a carcinoma in the context of multiple cancer classes.

Carcinoma↗