Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

[DNA chip data mining].

DNA chip data routinely contain gene expression levels of thousands of genes and the analysis should be supported by various computational tools. To be brief, the analysis procedure consists of four steps including image scanning, image processing, mathematical interpretation and biological interpretation. In image processing step, we should detect the spots and measure the signals of the spots and the background. In mathematical interpretation step, first of all we should massage the measured signals to make them appropriate for further mathematical analysis. The massaged data could be analyzed by various computational methods especially when the data were generated for multiple samples comparisons. The clustering techniques including hierarchical clustering, k-means clustering, SOTA, SOM are the most popular methods in this step. Various other multivariate statistics and related machine learning techniques are being introduced and applied to DNA chip data analysis recently. And finally the most important step we should tackle is the biological interpretation task. Although the depth of the domain knowledge about the biological situation under which the data were generated is the most important factor to elucidate the biological context, it could be supported by various bioinformatics tools including MEDLINE abstract processing by NLP techniques or genetic network models constructed by Boolean networks algorithms.

Algorithms↗

Selecting informative genes with parallel genetic algorithms in tissue classification.

Recent advances in biotechnology offer the ability to measure the levels of expression of thousands of genes in parallel. Analysis of such data can provide understanding and insight into gene function and regulatory mechanisms. Several machine learning approaches have been used to aid to understand the functions of genes. However, these tasks are made more difficult due to the noisy nature of array data and the overwhelming number of gene features. In this paper, we use the parallel genetic algorithm to filter out the informative genes relative to classification. By combing with the classification method proposed by Golub et al. and Slonim et al., we classify the data sets with tissues of different classes, and the preliminary results are presented in this paper.

Algorithms↗

A knowledge model for the interpretation and visualization of NLP-parsed discharged summaries.

At our institution, a Natural Language Processing (NLP) tool called MedLEE is used on a daily basis to parse medical texts including complete discharge summaries. MedLEE transforms written text into a generic structured format, which preserves the richness of the underlying natural language expressions by the use of concept modifiers (like change, certainty, degree and status). As a tradeoff, extraction of application-specific medical information is difficult without a clear understanding of how these modifiers combine. We report on a knowledge model for MedLEE modifiers that is helpful for a high level interpretation of NLP data and is used for the generation of two distinct views on NLP-parsed discharge summaries: A physician view offering a condensed overview of the severity of patient problems and a data mining view featuring binary problem states useful for machine learning.

Artificial Intelligence↗

Global analysis of large-scale chemical and biological experiments.

Research in the life sciences is increasingly dominated by high-throughput data collection methods that benefit from a global approach to data analysis. Recent innovations that facilitate such comprehensive analyses are highlighted. Several developments enable the study of the relationships between newly derived experimental information, such as biological activity in chemical screens or gene expression studies, and prior information, such as physical descriptors for small molecules or functional annotation for genes. The way in which global analyses can be applied to both chemical screens and transcription profiling experiments using a set of common machine learning tools is discussed.

Animals↗

Individuality of handwriting.

Motivated by several rulings in United States courts concerning expert testimony in general, and handwriting testimony in particular, we undertook a study to objectively validate the hypothesis that handwriting is individual. Handwriting samples of 1,500 individuals, representative of the U.S. population with respect to gender, age, ethnic groups, etc., were obtained. Analyzing differences in handwriting was done by using computer algorithms for extracting features from scanned images of handwriting. Attributes characteristic of the handwriting were obtained, e.g., line separation, slant, character shapes, etc. These attributes, which are a subset of attributes used by forensic document examiners (FDEs), were used to quantitatively establish individuality by using machine learning approaches. Using global attributes of handwriting and very few characters in the writing, the ability to determine the writer with a high degree of confidence was established. The work is a step towards providing scientific support for admitting handwriting evidence in court. The mathematical approach and the resulting software also have the promise of aiding the FDE.

Adolescent↗

DNA splice site detection: a comparison of specific and general methods.

In an era when whole organism genomes are being routinely sequenced, the problem of gene finding has become a key issue on the road to understanding. For eukaryotic organisms a large part of locating the genes is accomplished by predicting the likely location of splice sites on a DNA strand. This problem of splice site location has been ap- proached using a number of machine learning or statistical methods tailored more or less specifically to the nature of the problem. Recently large margin classifiers and boosting methods have been found to give improvements over more traditional methods in a number of areas. Here we compare large margin classifiers (SVM and CMLS) and boosted decision trees with the three most common models used for splice site detection (WMM, WAM, and MDT). We find that the newer methods compare favorably in all cases and can yield significant improvement in some cases.

Algorithms↗

Maximum entropy modeling for mining patient medication status from free text.

Using a classification scheme of patient medication status we sought to recognize and categorize medications mentioned in the unrestricted text of clinical documents generated in clinical practice. The categories refer to the patient's status with respect to the medication such as discontinuation, start or initiation, and continuation of a given medication. This categorization is performed with a machine learning technique, Maximum Entropy (ME), that is well suited to incorporating heterogeneous sources of information necessary for classifying patient's medication status. We use hand labeled training data to generate ME models and test 5 different training feature sets. Our results show that the most optimal feature set includes a combination of the following: two words preceding and following the mention of the drug, the subject of the sentence in which the drug mention occurs, the 2 words following the subject, and a binary feature vector of lexicalized semantic cues indicative of medication status or its change. The average predictive power of a model trained on these features is approximately 89%.

Artificial Intelligence↗

[Value of combined detection of tumor markers for the prediction of small cell and non-small cell lung cancer].

To evaluate the value of detection of 4 tumor markers(CEA, CA125, gastrin, and NSE) for histological types in patients with lung cancer and to improve the predicted efficiency of tumor markers for distinguishing between small cell lung cancer(SCLC) and non-small cell lung cancer (NSCLC), these 4 tumor markers in serum were determined in 51 patients (21 cases with SCLC, 30 cases with NSCLC) with confirmed primary diagnosis of lung cancer of different histology by radioimmunoassay. Linear learning machine method, PRIMA method and KNN method were used to classify SCLC and NSCLC. The levels of gastrin and NSE in SCLC were apparently higher than those of gastrin and NSE in NSCLC, but the levels of CEA and CA125 in SCLC were significantly lower than those in NSCLC. Smoking had an effect on the levels of CEA and CA125, but had little effect on those of gastrin and NSE. The total accuracy of the three methods was over 85% in distinguishing SCLC from NSCLC. So combined detection of the four tumor markers in serum might be useful in the prediction of histological types in patients with lung cancer.

Adenocarcinoma↗

Identification of osteopontin as a prognostic plasma marker for head and neck squamous cell carcinomas.

PURPOSE: Tumor hypoxia modifies treatment efficacy and promotes tumor progression. Here, we investigated the relationship between osteopontin (OPN), tumor pO(2), and prognosis in patients with head and neck squamous cell carcinomas (HNSCC). EXPERIMENTAL DESIGN: We performed linear discriminant analysis, a machine learning algorithm, on the NCI-60 cancer cell line microarray expression database to identify a gene profile that best distinguish cell lines with high Von-Hippel Lindau (VHL) gene expression, an important regulator of hypoxia-related genes, from those with low expression. Plasma OPN levels in 15 volunteers, 31 VHL patients, and 54 HNSCC patients were quantitatively measured by ELISA. The relationships between plasma OPN levels, tumor pO(2) as measured by the Eppendorf microelectrode, freedom from relapse (FFR), and survival in HNSCC patients were evaluated. RESULTS: Microarray analysis indicated that OPN gene expression inversely correlated with that of VHL. These findings were confirmed by Northern blot analysis. ELISA studies and Western blot in a HNSCC cell line demonstrated that hypoxia exposure resulted in increased OPN secretion. Patients with VHL syndrome had significantly higher plasma OPN levels than healthy volunteers. Plasma OPN level inversely correlated with tumor pO(2) (P = 0.003, r = -0.42). OPN levels correlated with clinical outcomes. The 1-year FFR and survival rates were 80 and 100%, respectively, for patients with OPN levels 450 ng/ml (P = 0.002 and 0.0005). Multivariate analysis revealed that OPN was an independent predictor for FFR and survival. CONCLUSIONS: Plasma OPN levels appeared to correlate with tumor hypoxia in HNSCC patients and may serve as noninvasive tests to identify patients at high risk for tumor recurrence.

Adult↗

Playing biology's name game: identifying protein names in scientific text.

A growing body of work is devoted to the extraction of protein or gene interaction information from the scientific literature. Yet, the basis for most extraction algorithms, i.e. the specific and sensitive recognition of protein and gene names and their numerous synonyms, has not been adequately addressed. Here we describe the construction of a comprehensive general purpose name dictionary and an accompanying automatic curation procedure based on a simple token model of protein names. We designed an efficient search algorithm to analyze all abstracts in MEDLINE in a reasonable amount of time on standard computers. The parameters of our method are optimized using machine learning techniques. Used in conjunction, these ingredients lead to good search performance. A supplementary web page is available at http://cartan.gmd.de/ProMiner/.

Abstracting and Indexing↗

Functional discrimination of gene expression patterns in terms of the gene ontology.

The ever-growing amount of experimental data in molecular biology and genetics requires its automated analysis, by employing sophisticated knowledge discovery tools. We use an Inductive Logic Programming (ILP) learner to induce functional discrimination rules between genes studied using microarrays and found to be differentially expressed in three recently discovered subtypes of adenocarcinoma of the lung. The discrimination rules involve functional annotations from the Proteome HumanPSD database in terms of the Gene Ontology, whose hierarchical structure is essential for this task. While most of the lower levels of gene expression data (pre)processing have been automated, our work can be seen as a step toward automating the higher level functional analysis of the data. We view our application not just as a prototypical example of applying more sophisticated machine learning techniques to the functional analysis of genes, but also as an incentive for developing increasingly more sophisticated functional annotations and ontologies, that can be automatically processed by such learning algorithms.

Adenocarcinoma↗

A probabilistic information retrieval approach to medical annotation in SWISS-PROT.

The goal of medical annotation of human proteins in Swiss-Prot is to add features specifically intended for researchers working on genetic diseases and polymorphisms. For this purpose, it is necessary to search through a vast number of publications containing relevant information. Promising results have been obtained by applying natural language processing and machine learning techniques to solve this problem. By using the Probabilistic Latent Categorizer on representative query sets, 69% recall and 59% precision was achieved for relevant documents. This classifier also rejected irrelevant abstracts with more than 96% precision. Better linguistic pre-processing of source documents can further improve such computer approach.

Databases, Protein↗

Text categorization models for retrieval of high quality articles in internal medicine.

The discipline of Evidence Based Medicine (EBM) studies formal and quasi-formal methods for identifying high quality medical information and abstracting it in useful forms so that patients receive the best customized care possible [1]. Current computer-based methods for finding high quality information in PubMed and similar bibliographic resources utilize search tools that employ preconstructed Boolean queries. These clinical queries are derived from a combined application of (a) user interviews, (b) ad-hoc manual document quality review, and (c) search over a constrained space of disjunctive Boolean queries. The present research explores the use of powerful text categorization (machine learning) methods to identify content-specific and high-quality PubMed articles. Our results show that models built with the proposed approach outperform the Boolean based PubMed clinical query filters in discriminatory power.

Area Under Curve↗

Classifying instantaneous cognitive states from FMRI data.

We consider the problem of detecting the instantaneous cognitive state of a human subject based on their observed functional Magnetic Resonance Imaging (fMRI) data. Whereas fMRI has been widely used to determine average activation in different brain regions, our problem of automatically decoding instantaneous cognitive states has received little attention. This problem is relevant to diagnosing cognitive processes in neurologically normal and abnormal subjects. We describe a machine learning approach to this problem, and report on its successful use for discriminating cognitive states such as observing a picture versus reading a sentence, and reading a word about people versus reading a word about buildings.

Artificial Intelligence↗

Using rough sets, neural networks, and logistic regression to predict compliance with cholesterol guidelines goals in patients with coronary artery disease.

Coronary artery disease is a leading cause of death and disability in the United States and throughout the developed world. Results from large randomized, blinded, placebo-controlled trials have demonstrated clearly the benefit of lowering LDL cholesterol in lowering the risk for coronary artery disease. Unfortunately, despite the quantity of evidence, and the availability of medications that can efficiently lower LDL cholesterol with few side effects, not everyone who could benefit from cholesterol lowering interventions actually receives them. Despite the dissemination of national care guidelines for the evaluation and treatment of cholesterol levels (NCEP - National Cholesterol Education Program), compliance with such guidelines is suboptimal. There clearly is room for improvement in narrowing the gap between evidence based guidelines and actual clinical practice. The ability to classify those patients who are or will likely to be noncompliant on the basis of patient data routinely collected during patient care could be potentially useful by enabling the focusing of limited health care resources to those who are or will be at high risk of being under treated. In order to explore this possibility further, we attempted to create such classifiers of cholesterol guideline compliance. To do this, we obtained data from an ambulatory electronic medical record system at use at the MGH adult primary care practices for over 20 years. We obtained the data from this hierarchically-structured EMR using its own native query language, called MQL (Medical Query Language). Next, we applied to the collected data the machine learning techniques of rough set theory, neural networks (feed forward backpropagation nets), and logistic regression. We did this by using commonly available software that for the most part is freely available via the internet. We then compared the accuracy of the classifier models using the receiver operating characteristic (ROC) area and C-index summary metrics.

Cholesterol↗

Artificial intelligence techniques for bioinformatics.

This review provides an overview of the ways in which techniques from artificial intelligence (AI) can be usefully employed in bioinformatics, both for modelling biological data and for making new discoveries. The paper covers three techniques: symbolic machine learning approaches (nearest neighbour and identification tree techniques), artificial neural networks and genetic algorithms. Each technique is introduced and supported with examples taken from the bioinformatics literature. These examples include folding prediction, viral protease cleavage prediction, classification, multiple sequence alignment and microarray gene expression analysis.

Algorithms↗

Using nuclear morphometry to discriminate the tumorigenic potential of cells: a comparison of statistical methods.

Despite interest in the use of nuclear morphometry for cancer diagnosis and prognosis as well as to monitor changes in cancer risk, no generally accepted statistical method has emerged for the analysis of these data. To evaluate different statistical approaches, Feulgen-stained nuclei from a human lung epithelial cell line, BEAS-2B, and a human lung adenocarcinoma (non-small cell) cancer cell line, NCI-H522, were subjected to morphometric analysis using a CAS-200 imaging system. The morphometric characteristics of these two cell lines differed significantly. Therefore, we proceeded to address the question of which statistical approach was most effective in classifying individual cells into the cell lines from which they were derived. The statistical techniques evaluated ranged from simple, traditional, parametric approaches to newer machine learning techniques. The multivariate techniques were compared based on a systematic cross-validation approach using 10 fixed partitions of the data to compute the misclassification rate for each method. For comparisons across cell lines at the level of each morphometric feature, we found little to distinguish nonparametric from parametric approaches. Among the linear models applied, logistic regression had the highest percentage of correct classifications; among the nonlinear and nonparametric methods applied, the Classification and Regression Trees model provided the highest percentage of correct classifications. Classification and Regression Trees has appealing characteristics: there are no assumptions about the distribution of the variables to be used, there is no need to specify which interactions to test, and there is no difficulty in handling complex, high-dimensional data sets containing mixed data types.

Adenocarcinoma↗

A knowledge based approach for automated signal generation in pharmacovigilance.

BACKGROUND: Pharmacovigilance experts detect new adverse drug reactions (ADR) by manually reviewing spontaneous reporting systems. Automated signal generation aims to focus the attention of experts on drug-adverse event associations which are disproportionally present in the database. Although adverse events are coded by means of controlled vocabularies such as the MedDRA dictionary, this semantic information is not taken into account for signal generation. OBJECTIVE: To improve the performance of current signal detection algorithms using knowledge based approach. METHOD: We developed a formal ontology of ADRs and built a data mining tool that uses description logic representations of MedDRA terms to group medically related case reports. RESULTS: This knowledge based approach increased the sensitivity of signal detection with no decrease in specificity. DISCUSSION: A knowledge based approach improved the performance of signal detection tools. However, the huge work-load involved in the knowledge engineering step limits the use of this approach for machine learning.

Adverse Drug Reaction Reporting Systems↗