Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

Artificial neural network technologies to identify biomarkers for therapeutic intervention.

High-throughput technologies such as DNA/RNA microarrays, mass spectrometry and protein chips are creating unprecedented opportunities to accelerate towards the understanding of living systems and the identification of target genes and pathways for drug development and therapeutic intervention. However, the increasing volumes of data generated by molecular profiling experiments pose formidable challenges to investigate an overwhelming mass of information and turn it into predictive, deployable markers. Advanced biostatistics and machine learning methods from computer science have been applied to analyze and correlate numerical values of profiling intensities to physiological states. This article reviews the application of artificial neural networks, an information-processing tool, to the identification of sets of diagnostic/prognostic biomarkers from high-throughput profiling data.

Animals↗

A bi-recursive neural network architecture for the prediction of protein coarse contact maps.

Prediction of contact maps may be seen as a strategic step towards the solution of fundamental open problems in structural genomics. In this paper we focus on coarse grained maps that describe the spatial neighborhood relation between secondary structure elements (helices, strands, and coils) of a protein. We introduce a new machine learning approach for scoring candidate contact maps. The method combines a specialized noncausal recursive connectionist architecture and a heuristic graph search algorithm. The network is trained using candidate graphs generated during search. We show how the process of selecting and generating training examples is important for tuning the precision of the predictor.

Algorithms↗

[Microarrays: technologies overview and data analysis].

DNA microarrays are a powerful tool to investigate differential gene expression for thousands of genes simultaneously. In this review, recent advances in DNA microarray technologies and their applications are examined. Various DNA microarray platforms are described along with their methods for fabrication and their use. In addition some algorithms and tools for the analysis of microarray expression data, including clustering methods, partitioning and machine learning methods are discussed.

Gene Expression Profiling↗

Knowledge discovery and system biology in molecular medicine: an application on neurodegenerative diseases.

The possibility to study an organism in terms of system theory has been proposed in the past, but only the advancement of molecular biology techniques allow us to investigate the dynamical properties of a biological system in a more quantitative and rational way than before . These new techniques can gave only the basic level view of an organisms functionality. The comprehension of its dynamical behaviour depends on the possibility to perform a multiple level analysis. Functional genomics has stimulated the interest in the investigation the dynamical behaviour of an organism as a whole. These activities are commonly known as System Biology, and its interests ranges from molecules to organs. One of the more promising applications is the 'disease modeling'. The use of experimental models is a common procedure in pharmacological and clinical researches; today this approach is supported by 'in silico' predictive methods. This investigation can be improved by a combination of experimental and computational tools. The Machine Learning (ML) tools are able to process different heterogeneous data sources, taking into account this peculiarity, they could be fruitfully applied to support a multilevel data processing (molecular, cellular and morphological) that is the prerequisite for the formal model design; these techniques can allow us to extract the knowledge for mathematical model development. The aim of our work is the development and implementation of a system that combines ML and dynamical models simulations. The program is addressed to the virtual analysis of the pathways involved in neurodegenerative diseases. These pathologies are multifactorial diseases and the relevance of the different factors has not yet been well elucidated. This is a very complex task; in order to test the integrative approach our program has been limited to the analysis of the effects of a specific protein, the Cyclin dependent kinase 5 (CDK5) which relies on the induction of neuronal apoptosis. The system has a modular structure centred on a textual knowledge discovery approach. The text mining is the only way to enhance the capability to extract ,from multiple data sources, the information required for the dynamical simulator. The user may access the publically available modules through the following site: http://biocomp.ge.ismac.cnr.it.

Biomedical Research↗

Health. Care. Anywhere. Today.

What if clinical quality medical equipment were available to every consumer in a form factor that was inexpensive, accurate, and easy to use? What if this equipment provided information that previously was un-measurable or very difficult to measure? What if the physiological state of individuals, at resolutions measured in thousandths of a second instead of in visits per year, could be measured easily, making it possible to ascertain caloric intake and expenditure, patterns of sleep, contextual activities such as working-out and driving, even parameters of mental state and health. What aspect of healthcare would not change? We present a system that is available today that enables this vision. This award-winning multi-channel wearable physiological monitor has enabled the collection of more than 90 million minutes of data in natural settings from thousands of subjects engaged in diverse activities. Data modeling efforts are resulting in applications that present meaningful and actionable information in real-time to users and their designated collaborators (physicians, family members, counselors, coaches, etc.) We describe the SenseWear system, its design, and a summary of validation studies, current commercial applications, and ongoing research. This discussion will show how the convergence of design for wearability, advances in machine learning, and improvements in wireless technology will manifest the future of health care as personal, ubiquitous, and collaborative.

Clothing↗

Robust diagnosis of non-Hodgkin lymphoma phenotypes validated on gene expression data from different laboratories.

A major challenge in cancer diagnosis from microarray data is the need for robust, accurate, classification models which are independent of the analysis techniques used and can combine data from different laboratories. We propose such a classification scheme originally developed for phenotype identification from mass spectrometry data. The method uses a robust multivariate gene selection procedure and combines the results of several machine learning tools trained on raw and pattern data to produce an accurate meta-classifier. We illustrate and validate our method by applying it to gene expression datasets: the oligonucleotide HuGeneFL microarray dataset of Shipp et al. (www.genome.wi.mit.du/MPR/lymphoma) and the Hu95Av2 Affymetrix dataset (DallaFavera's laboratory, Columbia University). Our pattern-based meta-classification technique achieves higher predictive accuracies than each of the individual classifiers , is robust against data perturbations and provides subsets of related predictive genes. Our techniques predict that combinations of some genes in the p53 pathway are highly predictive of phenotype. In particular, we find that in 80% of DLBCL cases the mRNA level of at least one of the three genes p53, PLK1 and CDK2 is elevated, while in 80% of FL cases, the mRNA level of at most one of them is elevated.

Biomarkers, Tumor↗

Clustering time-varying gene expression profiles using scale-space signals.

The functional state of an organism is determined largely by the pattern of expression of its genes. The analysis of gene expression data from gene chips has primarily revolved around clustering and classification of the data using machine learning techniques based on the intensity of expression alone with the time-varying pattern mostly ignored. In this paper, we present a pattern recognition-based approach to capturing similarity by finding salient changes in the time-varying expression patterns of genes. Such changes can give clues about important events, such as gene regulation by cell-cycle phases, or even signal the onset of a disease. Specifically, we observe that dissimilarity between time series is revealed by the sharp twists and bends produced in a higher-dimensional curve formed from the constituent signals. Scale-space analysis is used to detect the sharp twists and turns and their relative strength with respect to the component signals is estimated to form a shape similarity measure between time profiles. A clustering algorithm is presented to cluster gene profiles using the scale-space distance as a similarity metric. Multi-dimensional curves formed from time series within clusters are used as cluster prototypes or indexes to the gene expression database, and are used to retrieve the functionally similar genes to a query gene profile. Extensive comparison of clustering using scale-space distance in comparison to traditional Euclidean distance is presented on the yeast genome database.

Algorithms↗

[Fish-based assessment methods for the ecological status of aquatic systems].

A short overview of fish-based assessment methods for aquatic systems is presented. Multimetric indices, as, e.g., the index of biotic integrity (IBI), firstly developed in USA and later adapted for European river basins and other countries, are shortly described. Non-multimetric indices are also discussed, e.g. the ichthyological index (II) and the index of ecological status of fish communities (ISECI), both proposed for monitoring Italian rivers. Moreover, statistical and predicting methods based on machine learning techniques are described. Finally, a new approach for developing standardised fish-based methods useful to assess ecological status of Italian rivers is proposed. Although rather complex, the use of bony fish in biomonitoring is promising and requires a multidisciplinary approach to be adopted.

Animals↗

Use of an electronic nose to diagnose bacterial sinusitis.

BACKGROUND: Having previously established that an electronic nose (enose) can distinguish among bacteria samples, between cerebrospinal fluid leak and serum, and can identify patients with ventilator-associated pneumonia, we hypothesized that bacterial sinusitis could be diagnosed by sampling exhaled gas with an enose. METHODS: Using a nasal continuous positive airway pressure mask, we sampled gas exhaled through the nose of patients with sinusitis and compared them with controls. Data were first projected onto the principal components and then classified by support vector machine (SVM), a machine learning algorithm for pattern recognition. RESULTS: SVM analysis showed good discrimination using three approaches. First, 11 samples were used to create a training set that was used to predict whether individual samples from each set were a member of the control or infected sets. The enose was correct 98.4% of the time. Second, one-half of the samples from each of the same 11 control and infected groups were used to construct a training set, which was used to predict infection in the remaining samples. The enose was correct 82% of the time. Finally, 68 samples (34 positive and 34 controls) were analyzed using a leave-one-out scheme for creating training sets and testing sets. This method, designed to reflect the generalization property of the SVM classifier, scored a classification rate of 72%. CONCLUSION: Using the enose to sample nasal exhalation from patients with suspected sinusitis, we were able to predict correctly the diagnosis of sinusitis in at least 72% of the samples. The next step will be to do forward prediction using this model.

Algorithms↗

Deriving the expected utility of a predictive model when the utilities are uncertain.

Predictive models are often constructed from clinical databases with the goal of eventually helping make better clinical decisions. Evaluating models using decision theory is therefore natural. When constructing a model using statistical and machine learning methods, however, we are often uncertain about precisely how the model will be used. Thus, decision-independent measures of classification performance, such as the area under an ROC curve, are popular. As a complementary method of evaluation, we investigate techniques for deriving the expected utility of a model under uncertainty about the model's utilities. We demonstrate an example of the application of this approach to the evaluation of two models that diagnose coronary artery disease.

Artificial Intelligence↗

Automatic processing of spoken dialogue in the home hemodialysis domain.

Spoken medical dialogue is a valuable source of information, and it forms a foundation for diagnosis, prevention and therapeutic management. However, understanding even a perfect transcript of spoken dialogue is challenging for humans because of the lack of structure and the verbosity of dialogues. This work presents a first step towards automatic analysis of spoken medical dialogue. The backbone of our approach is an abstraction of a dialogue into a sequence of semantic categories. This abstraction uncovers structure in informal, verbose conversation between a caregiver and a patient, thereby facilitating automatic processing of dialogue content. Our method induces this structure based on a range of linguistic and contextual features that are integrated in a supervised machine-learning framework. Our model has a classification accuracy of 73%, compared to 33% achieved by a majority baseline (p<0.01). This work demonstrates the feasibility of automatically processing spoken medical dialogue.

Algorithms↗

Analysis of polarity information in medical text.

Knowing the polarity of clinical outcomes is important in answering questions posed by clinicians in patient treatment. We treat analysis of this information as a classification problem. Natural language processing and machine learning techniques are applied to detect four possibilities in medical text: no outcome, positive outcome, negative outcome, and neutral outcome. A supervised learning method is used to perform the classification at the sentence level. Five feature sets are constructed: unigrams, bigrams, change phrases, negations, and categories. The performance of different combinations of feature sets is compared. The results show that generalization using the category information in the domain knowledge base Unified Medical Language System is effective in the task. The effect of context information is significant. Combining linguistic features and domain knowledge leads to the highest accuracy.

Artificial Intelligence↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗

Question analysis for biomedical question answering.

We are developing a biomedical question answering system. This paper describes our system's architecture and our question analysis component. Specifically, we have explored the use of various supervised machine learning approaches to filter out unanswerable questions based on physicians' annotations.

Artificial Intelligence↗

Supporting the curation of biological databases with reusable text mining.

Curators of biological databases transfer knowledge from scientific publications, a laborious and expensive manual process. Machine learning algorithms can reduce the workload of curators by filtering relevant biomedical literature, though their widespread adoption will depend on the availability of intuitive tools that can be configured for a variety of tasks. We propose a new method for supporting curators by means of document categorization, and describe the architecture of a curator-oriented tool implementing this method using techniques that require no computational linguistic or programming expertise. To demonstrate the feasibility of this approach, we prototyped an application of this method to support a real curation task: identifying PubMed abstracts that contain allergen cross-reactivity information. We tested the performance of two different classifier algorithms (CART and ANN), applied to both composite and single-word features, using several feature scoring functions. Both classifiers exceeded our performance targets, the ANN classifier yielding the best results. These results show that the method we propose can deliver the level of performance needed to assist database curation.

Allergens↗

[Early detection of pancreatic cancer by novel proteomic technique].

Pancreatic cancer is the fifth leading cause of cancer-related mortality in Japan. Early detection in pancreatic cancer is one of the most feasible strategies to improve outcome. We compared plasma proteome between pancreatic cancer patients and healthy controls using surface-enhanced laser desorption/ionization coupled with hybrid quadrupole time-of-flight mass spectrometry. Proteomic spectra were generated from a total of 245 plasma samples obtained from two institutes. A discriminating proteomic pattern was built from training cohort using machine learning algorithm and was applied two validation cohorts. This set discriminating cancer patients in the first validation cohort with a sensitivity of 90.9% and a specificity of 91.9%, and was further validated in an independent cohort at a second institution. When combined with CA19-9, 100% tumor of pancreatic cancers, including early stage tumors, were detected. In this report, we describe a possible detection of early pancreatic cancer using novel proteomic technique.

Biomarkers, Tumor↗

Two models for outcome prediction - a comparison of logistic regression and neural networks.

OBJECTIVES: Accurately predicting disease progress from a set of predictive variables is an important aspect of clinical work. For binary outcomes, the classical approach is to develop prognostic logistic regression (LR) models. Alternatively, machine learning algorithms were proposed with artificial neural networks (ANN) having become popular over the last decades. Although some studies have compared predictive accuracies of LR and ANN models, some concerns regarding their methodological quality have been voiced. Our comparison has the advantage of being based on two large independent data sets allowing for elaborate model development and independent validation. METHODS: From the German Stroke Database, a learning data set including 1754 prospectively recruited patients with acute ischemic stroke was used. Utilizing LR and ANN, two prognostic models were developed predicting restitution of functional independence and survival after 100 days. The resulting models were applied to classify 1470 patients with acute ischemic stroke; this test data set was collected independently from the learning data. Error fractions in the test data were determined, and differences in error fractions between the algorithms were calculated with 95% confidence intervals. RESULTS: For most prognostic models, error fractions in the test data were below 40%. There was no difference between the algorithms except for the model predicting completely versus incompletely restituted or deceased patients (difference in error fractions = 4.01% [2.10-5.96%], p = 0.0001). CONCLUSIONS: The conscientiously applied LR remains the gold standard for prognostic modelling; however, ANN can be an alternative automated "quick and easy" multivariate analysis.

Aged↗

Bioinformatics and its impact on clinical research methods. Findings from the Section on Bioinformatics.

OBJECTIVES: To summarize current excellent research in the field of bioinformatics. METHOD: Synopsis of the articles selected for the IMIA Yearbook 2006. RESULTS: Current research in the field of bioinformatics clearly shows ongoing unification of experimental findings and clinical outcomes. Microarray data, gene sequences and clinical data are more and more perceived as different but related facets of one entity. Significant work is done in the area of text and data mining in order to bring together patient data and biochemical phenomena by means of ontologies. A strong trend in the clinical field is performance of exhaustive studies on DNA material derived from patients that suffer from diseases that are already known to be inherited. Examination of appropriate methods covers data and text mining, ontologies as well as machine learning and classification. CONCLUSIONS: The best paper selection of articles on bioinformatics shows examples of excellent research on methods used for studying inherited diseases and their underlying genetic dispositions. Clinical studies, inclusion of experimental findings like microarray data, and of knowledge representation formats all lead to a better understanding the linkage between gene sequences, biological functions and clinical findings in the form of healthy state or physiological disorders.

Awards and Prizes↗