Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Mining of biological data II: assessing data structure and class homogeneity by cluster analysis.

An important step in data analysis is class assignment which is usually done on the basis of a macroscopic phenotypic or bioprocess characteristic, such as high vs low growth, healthy vs diseased state, or high vs. low productivity. Unfortunately, such an assignment may lump together samples, which when derived from a more detailed phenotypic or bioprocess description are dissimilar, giving rise to models of lower quality and predictive power. In this paper we present a clustering algorithm for data preprocessing which involves the identification of fundamentally similar lots on the basis of the extent of similarity among the system variables. The algorithm combines aspects of cluster analysis and principal component analysis by applying agglomerative clustering methods to the first principal component of the system data matrix. As part of a rational strategy for developing empirical models, this technique selects lots (samples) which are most appropriate for inclusion in a training set by analyzing multivariate data homogeneity. Samples with similar data structures are identified and grouped together into distinct clusters. This knowledge is used in the formation of potential training sets. Additionally, this technique can identify atypical lots, i.e., samples that are not simply outliers but exhibit the general properties of one class but have been given the assignment of the other. The method is presented along with examples from its application to fermentation data sets.

Algorithms↗

Mining complex clinical data for patient safety research: a framework for event discovery.

Successfully addressing patient safety requires detecting medical events effectively. Given the volume of patients seen at medical centers, detecting events automatically from data that are already available electronically would greatly facilitate patient safety work. We have created a framework for electronic detection. Key steps include: selecting target events, assessing what information is available electronically, transforming raw data such as narrative notes into a coded format, querying the transformed data, verifying the accuracy of event detection, characterizing the events using systems and cognitive approaches, and using what is learned to improve detection.

Database Management Systems↗

Bayesian model selection for mining mass spectrometry data.

A procedure for learning a probabilistic model from mass spectrometry data that accounts for domain specific noise and mitigates the complexity of Bayesian structure learning is presented. We evaluate the algorithm by applying the learned probabilistic model to microorganism detection from mass spectrometry data.

Algorithms↗

Mining DNA microarray data using a novel approach based on graph theory.

The recent demonstration that biochemical pathways from diverse organisms are arranged in scale-free, rather than random, systems [Jeong et al., Nature 407 (2000) 651-654], emphasizes the importance of developing methods for the identification of biochemical nexuses--the nodes within biochemical pathways that serve as the major input/output hubs, and therefore represent potentially important targets for modulation. Here we describe a bioinformatics approach that identifies candidate nexuses for biochemical pathways without requiring functional gene annotation; we also provide proof-of-principle experiments to support this technique. This approach, called Nexxus, may lead to the identification of new signal transduction pathways and targets for drug design.

Apoptosis↗

Mining gene expression data using a novel approach based on hidden Markov models.

In this work we have developed a new framework for microarray gene expression data analysis. This framework is based on hidden Markov models. We have benchmarked the performance of this probability model-based clustering algorithm on several gene expression datasets for which external evaluation criteria were available. The results showed that this approach could produce clusters of quality comparable to two prevalent clustering algorithms, but with the major advantage of determining the number of clusters. We have also applied this algorithm to analyze published data of yeast cell cycle gene expression and found it able to successfully dig out biologically meaningful gene groups. In addition, this algorithm can also find correlation between different functional groups and distinguish between function genes and regulation genes, which is helpful to construct a network describing particular biological associations. Currently, this method is limited to time series data. Supplementary materials are available at http://www.bioinfo.tsinghua.edu.cn/~rich/hmmgep_supp/.

Algorithms↗

Genome analysis with gene-indexing databases.

The recent release of the draft sequence and the eventual completion of the human genome present the scientific community with a rich source of data to mine. Yet, these data are content poor in the absence of additional correlative information. Expressed sequence tag (EST) datasets and their associated gene indices have existed for many years, and represent the first attempt at understanding the complexity of the genome. These datasets remain extremely important as information sources and, in particular, as tools for analyzing the completed genomes. Here, we discuss the nature of ESTs and their associated tools and gene-indexing databases. In particular, we will compare three EST gene indices (UNIGENE, Merck Gene Index Version 2.0 and Doubletwist CAT), discuss how these gene indices are applied for both genome analysis and drug discovery, and demonstrate their importance as a complementary dataset to the annotated human genome.

Databases, Factual↗

Multi-sensors acquisition, data fusion, knowledge mining and alarm triggering in health smart homes for elderly people.

We deal in this paper with the concept of health smart home (HSH) designed to follow dependent people at home in order to avoid the hospitalisation, limiting hospital sojourns to short acute care or fast specific diagnostic investigations. For elderly people the project of such a HSH has been called AISLE (Apartment with Intelligent Sensors for Longevity Effectiveness). For this purpose, system having three levels of automatic measuring (1) the circadian activity, (2) the vegetative state, and (3) some state variables specific of certain organs involved in precise diseases, has been developed within the framework of a 'Health Integrated Smart Home Information System' (HIS2). HIS2 is an experimental platform for technologic development and clinical evaluation, in order to ensure the medical security and quality of life for patients who need home based medical monitoring. Location sensors are placed in each room of the HIS2, allowing the monitoring of patient's successive daily activity phases within the patient's home environment. We proceed with a sampling in an hourly schedule to detect weak variations of the nycthemeral rhythms. Based on numerous measurements, we establish a mean value with confidence limits of activity variables in normal behaviour permitting to detect for example a sudden abnormal event (like a fall) as well as a chronic pathologic activity (like a pollakiuria), allowing us to define a canonical domain within which the patient's activity is qualified to be 'predictable'. Alerts are set off if the patient's activity deviates from a predictable canonical domain. Moreover, we can follow the cardio-respiratory state by measuring the intensity of the respiratory sinusal arrhythmia in order to quantify the integrity of the bulbar vegetative system, and we finally propose to carefully watch abnormal symptoms like arterial pressure or presence of plasma proteins in the expired air flow for early detecting respectively hypertension or pulmonary oedema.

Aged↗

Mining viral protease data to extract cleavage knowledge.

MOTIVATION: The motivation is to identify, through machine learning techniques, specific patterns in HIV and HCV viral polyprotein amino acid residues where viral protease cleaves the polyprotein as it leaves the ribosome. An understanding of viral protease specificity may help the development of future anti-viral drugs involving protease inhibitors by identifying specific features of protease activity for further experimental investigation. While viral sequence information is growing at a fast rate, there is still comparatively little understanding of how viral polyproteins are cut into their functional unit lengths. The aim of the work reported here is to investigate whether it is possible to generalise from known cleavage sites to unknown cleavage sites for two specific viruses-HIV and HCV. An understanding of proteolytic activity for specific viruses will contribute to our understanding of viral protease function in general, thereby leading to a greater understanding of protease families and their substrate characteristics. RESULTS: Our results show that artificial neural networks and symbolic learning techniques (See5) capture some fundamental and new substrate attributes, but neural networks outperform their symbolic counterpart.

Algorithms↗

Mining gene expression data based on template theory.

MOTIVATION: It is understood that clustering genes are useful for exploring scientific knowledge from DNA microarray gene expression data. The explored knowledge can be finally used for annotating biological function for novel genes. Representing the explored knowledge in an efficient manner is then closely related to the classification accuracy. However, this issue has not yet been paid the attention it deserves. RESULT: A novel method based on template theory in cognitive psychology and pattern recognition is developed in this study for representing knowledge extracted from cluster analysis effectively. The basic principle is to represent knowledge according to the relationship between genes and a found cluster structure. Based on this novel knowledge representation method, a pattern recognition algorithm (the decision tree algorithm C4.5) is then used to construct a classifier for annotating biological functions of novel genes. The experiments on five published datasets show that this method has improved the classification performance compared with the conventional method. The statistical tests indicate that this improvement is significant. AVAILABILITY: The software package can be obtained upon request from the author.

Algorithms↗

EXCAVATOR: a computer program for efficiently mining gene expression data.

Massive amounts of gene expression data are generated using microarrays for functional studies of genes and gene expression data clustering is a useful tool for studying the functional relationship among genes in a biological process. We have developed a computer package EXCAVATOR for clustering gene expression profiles based on our new framework for representing gene expression data as a minimum spanning tree. EXCAVATOR uses a number of rigorous and efficient clustering algorithms. This program has a number of unique features, including capabilities for: (i) data- constrained clustering; (ii) identification of genes with similar expression profiles to pre-specified seed genes; (iii) cluster identification from a noisy background; (iv) computational comparison between different clustering results of the same data set. EXCAVATOR can be run from a Unix/Linux/DOS shell, from a Java interface or from a Web server. The clustering results can be visualized as colored figures and 2-dimensional plots. Moreover, EXCAVATOR provides a wide range of options for data formats, distance measures, objective functions, clustering algorithms, methods to choose number of clusters, etc. The effectiveness of EXCAVATOR has been demonstrated on several experimental data sets. Its performance compares favorably against the popular K-means clustering method in terms of clustering quality and computing time.

Algorithms↗

Topology-based cancer classification and related pathway mining using microarray data.

Cancer classification is the critical basis for patient-tailored therapy, while pathway analysis is a promising method to discover the underlying molecular mechanisms related to cancer development by using microarray data. However, linking the molecular classification and pathway analysis with gene network approach has not been discussed yet. In this study, we developed a novel framework based on cancer class-specific gene networks for classification and pathway analysis. This framework involves a novel gene network construction, named ordering network, which exhibits the power-law node-degree distribution as seen in correlation networks. The results obtained from five public cancer datasets showed that the gene networks with ordering relationship are better than those with correlation relationship in terms of accuracy and stability of the classification performance. Furthermore, we integrated the ordering networks, classification information and pathway database to develop the topology-based pathway analysis for identifying cancer class-specific pathways, which might be essential in the biological significance of cancer. Our results suggest that the topology-based classification technology can precisely distinguish cancer subclasses and the topology-based pathway analysis can characterize the correspondent biochemical pathways even if there are subtle, but consistent, changes in gene expression, which may provide new insights into the underlying molecular mechanisms of tumorigenesis.

Gene Expression Profiling↗

Efficiently mining gene expression data via a novel parameterless clustering method.

Clustering analysis has been an important research topic in the machine learning field due to the wide applications. In recent years, it has even become a valuable and useful tool for in-silico analysis of microarray or gene expression data. Although a number of clustering methods have been proposed, they are confronted with difficulties in meeting the requirements of automation, high quality, and high efficiency at the same time. In this paper, we propose a novel, parameterless and efficient clustering algorithm, namely, Correlation Search Technique (CST), which fits for analysis of gene expression data. The unique feature of CST is it incorporates the validation techniques into the clustering process so that high quality clustering results can be produced on the fly. Through experimental evaluation, CST is shown to outperform other clustering methods greatly in terms of clustering quality, efficiency, and automation on both of synthetic and real data sets.

Algorithms↗

Mining mouse microarray data.

Microarrays of mouse genes are now available from several sources, and they have so far given new insights into gene expression in embryonic development, regions of the brain and during apoptosis. Microarray data posted on the internet can be reanalyzed to study a range of questions.

Animals↗

Mining treatment termination data in an adolescent mental health service: a quantitative study.

This study utilizes available clinical information from client records to explore patterns of termination from mental health treatment among adolescents at an urban outpatient mental health center. The analysis focuses on how and why adolescents terminate from treatment and identifies variables associated with "acknowledged" and "unacknowledged" terminations. Findings indicate that termination was acknowledged infrequently, often a brief process that occurred almost as frequently by telephone as in the context of treatment. Contrary to "practice wisdom" concerning treatment termination, adolescents who "dropped out" without a "clinical process" reported considerably more engagement in treatment than those who acknowledged the termination of treatment. Recommendations for a more "open door" policy and a more flexible practice with adolescents are discussed.

Adolescent↗