Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

The G-algorithm for extraction of robust decision rules--children's postoperative intra-atrial arrhythmia case study.

Clinical medicine is facing a challenge of knowledge discovery from the growing volume of data. In this paper, a data mining algorithm (G-algorithm) is proposed for extraction of robust rules that can be used in clinical practice for better understanding and prevention of unwanted medical events. The G-algorithm is applied to the data set obtained for children born with a malformation of the heart (univentricular heart). As the result of the Fontan surgical procedure, designed to palliate the children, 10%-35% of patients postoperatively develop an arrhythmia known as the intra-atrial reentrant tachycardia. There is an obvious need to identify the children that may develop the tachycardia before the surgery is performed. Prior attempts to identify such children with statistical techniques have been unrewarding. The G-algorithm discussed in this paper shows that there exists an unambiguous relationship between measurable features and the tachycardia. The data set used in this study shows that, for 78.08% of infants, the occurrence of tachycardia can be accurately predicted. The authors' prior computational experience with diverse medical data sets indicates that the percentage of accurate predictions may become even higher if data on additional features is collected for a larger data set.

Algorithms↗

Mining multilevel and location-aware service patterns in mobile web environments.

In this correspondence, we address the issue of efficiently mining multilevel and location-aware associated service patterns in a mobile web environment. In terms of multilevel concept, we consider the complex problem that locations and services are of hierarchical structures. We propose a new data mining method named two-dimensional multilevel (2-DML) association rules mining, which can efficiently discover the associated service request patterns by taking into account the multilevel properties of locations and services. The discovered patterns can be effectively utilized in real applications like location-based and personalized services. To the best of our knowledge, this is the first work addressing this research issue. Some variations of the 2-DML method with different properties in terms of execution efficiency and memory efficiency were also developed. Through empirical evaluation, the proposed methods are shown to deliver good performance in terms of efficiency and scalability under various system conditions.

Algorithms↗

maxdLoad2 and maxdBrowse: standards-compliant tools for microarray experimental annotation, data management and dissemination.

BACKGROUND: maxdLoad2 is a relational database schema and Java application for microarray experimental annotation and storage. It is compliant with all standards for microarray meta-data capture; including the specification of what data should be recorded, extensive use of standard ontologies and support for data exchange formats. The output from maxdLoad2 is of a form acceptable for submission to the ArrayExpress microarray repository at the European Bioinformatics Institute. maxdBrowse is a PHP web-application that makes contents of maxdLoad2 databases accessible via web-browser, the command-line and web-service environments. It thus acts as both a dissemination and data-mining tool. RESULTS: maxdLoad2 presents an easy-to-use interface to an underlying relational database and provides a full complement of facilities for browsing, searching and editing. There is a tree-based visualization of data connectivity and the ability to explore the links between any pair of data elements, irrespective of how many intermediate links lie between them. Its principle novel features are: the flexibility of the meta-data that can be captured, the tools provided for importing data from spreadsheets and other tabular representations, the tools provided for the automatic creation of structured documents, the ability to browse and access the data via web and web-services interfaces. Within maxdLoad2 it is very straightforward to customise the meta-data that is being captured or change the definitions of the meta-data. These meta-data definitions are stored within the database itself allowing client software to connect properly to a modified database without having to be specially configured. The meta-data definitions (configuration file) can also be centralized allowing changes made in response to revisions of standards or terminologies to be propagated to clients without user intervention.maxdBrowse is hosted on a web-server and presents multiple interfaces to the contents of maxd databases. maxdBrowse emulates many of the browse and search features available in the maxdLoad2 application via a web-browser. This allows users who are not familiar with maxdLoad2 to browse and export microarray data from the database for their own analysis. The same browse and search features are also available via command-line and SOAP server interfaces. This both enables scripting of data export for use embedded in data repositories and analysis environments, and allows access to the maxd databases via web-service architectures. CONCLUSION: maxdLoad2 http://www.bioinf.man.ac.uk/microarray/maxd/ and maxdBrowse http://dbk.ch.umist.ac.uk/maxdBrowse are portable and compatible with all common operating systems and major database servers. They provide a powerful, flexible package for annotation of microarray experiments and a convenient dissemination environment. They are available for download and open sourced under the Artistic License.

Data Interpretation, Statistical↗

angaGEDUCI: Anopheles gambiae gene expression database with integrated comparative algorithms for identifying conserved DNA motifs in promoter sequences.

BACKGROUND: The completed sequence of the Anopheles gambiae genome has enabled genome-wide analyses of gene expression and regulation in this principal vector of human malaria. These investigations have created a demand for efficient methods of cataloguing and analyzing the large quantities of data that have been produced. The organization of genome-wide data into one unified database makes possible the efficient identification of spatial and temporal patterns of gene expression, and by pairing these findings with comparative algorithms, may offer a tool to gain insight into the molecular mechanisms that regulate these expression patterns. DESCRIPTION: We provide a publicly-accessible database and integrated data-mining tool, angaGEDUCI, that unifies 1) stage- and tissue-specific microarray analyses of gene expression in An. gambiae at different developmental stages and temporal separations following a bloodmeal, 2) functional gene annotation, 3) genomic sequence data, and 4) promoter sequence comparison algorithms. The database can be used to study genes expressed in particular stages, tissues, and patterns of interest, and to identify conserved promoter sequence motifs that may play a role in the regulation of such expression. The database is accessible from the address http://www.angaged.bio.uci.edu. CONCLUSION: By combining gene expression, function, and sequence data with integrated sequence comparison algorithms, angaGEDUCI streamlines spatial and temporal pattern-finding and produces a straightforward means of developing predictions and designing experiments to assess how gene expression may be controlled at the molecular level.

Algorithms↗

Prediction in medicine by integrating regression trees into regression analysis with optimal scaling.

OBJECTIVES: A new data-analysis strategy is proposed to solve the problems of selecting interaction terms in linear regression on the one hand, and of statistically testing the significance of regression trees on the other hand. METHODS: The proposed strategy combines two data mining techniques: regression trees and regression analysis with optimal scaling (CATREG). The method traces small regression trees using the bootstrap and integrates the results as interaction variables (called "trunk variables") into CATREG. RESULTS: An application to data from cardiac patients shows a relative increase of 19% variance accounted for (16% cross-validated variance), by the CATREG model including the trunk variables compared to the model excluding these variables. CONCLUSIONS: This study indicates that trunk variables can be useful to model interaction effects in prediction problems.

Artificial Intelligence↗

Marker identification and classification of cancer types using gene expression data and SIMCA.

OBJECTIVES: High-throughput technologies are radically boosting the understanding of living systems, thus creating enormous opportunities to elucidate the biological processes of cells in different physiological states. In particular, the application of DNA micro-arrays to monitor expression profiles from tumor cells is improving cancer analysis to levels that classical methods have been unable to reach. However, molecular diagnostics based on expression profiling requires addressing computational issues as the overwhelming number of variables and the complex, multi-class nature of tumor samples. Thus, the objective of the present research has been the development of a computational procedure for feature extraction and classification of gene expression data. METHODS: The Soft Independent Modeling of Class Analogy (SIMCA) approach has been implemented in a data mining scheme, which allows the identification of those genes that are most likely to confer robust and accurate classification of samples from multiple tumor types. RESULTS: The proposed method has been tested on two different microarray data sets, namely Golub's analysis of acute human leukemia and the small round blue cell tumors study presented by Khan et al.. The identified features represent a rational and dimensionally reduced base for understanding the biology of diseases, defining targets of therapeutic intervention, and developing diagnostic tools for classification of pathological states. CONCLUSIONS: The analysis of the SIMCA model residuals allows the identification of specific phenotype markers. At the same time, the class analogy approach provides the assignment to multiple classes, such as different pathological conditions or tissue samples, for previously unseen instances.

Biomarkers, Tumor↗

ATP-binding cassette protein E is involved in gene transcription and translation in Caenorhabditis elegans.

ATP-binding cassette protein E (ABCE) gene has been annotated as an RNase L inhibitor in eukaryotes. All eukaryotic species show the ubiquitous presence and high degree of conservation of ABCEs, however, RNase L is present only in mammals. This indicates that ABCEs may function not only as RNase L inhibitors, but also may have other functions that have yet to be determined. As an initial investigation into the novel functions of ABCE, we characterized the gene (Y39E4B.1) in Caenorhabditis elegans by a combination of data mining and functional assays. ABCE promoters drove GFP expressions in hypoderm, pharynx, vulvae, head, and tail neurons at all developmental stages. Three genes, rpl-4, nhr-91, and C07B5.3, were previously found to interact with ABCE. Our expression data showed overlapping expression patterns of ABCE and rpl-4 and nhr-91, but not C07B5.3. RNAi against ABCE resulted in embryonic lethality and slow growth. These data suggest that ABCE protein might be involved in the control of translation and transcription, work as shuttle protein between cytoplasm and nucleus, and possibly as a nucleocytoplasmic transporter. In addition, RNAi data suggest that ABCE and NHR-91 may function in vulvae development and molting pathways in C. elegans. Furthermore, our data suggest that ABCE, along with its interacting components, functions in a well-conserved pathway.

ATP-Binding Cassette Transporters↗

The effective use of a summary table and decision tree methodology to analyze very large healthcare datasets.

Very large datasets typically consists of millions of records, with many variables. Such datasets are stored and maintained by organizations because of the perceived potential information they contain. However, the problem with very large datasets is that traditional methods of data mining are not capable of retrieving this information because the software may be overwhelmed by the memory or computing requirements. In this article we outline a method that can analyze very large datasets. The method initially performs a data reduction step through the use of a summary table, which is then used as a reference dataset that is recursively partitioned to grow a decision tree.

Data Interpretation, Statistical↗

Frequency Finder: a multi-source web application for collection of public allele frequencies of SNP markers.

Publicly available single nucleotide polymorphism (SNP) allele frequencies are an important resource for the selection of genetic markers that may be most useful for gene mapping and association studies. Data mining these allele frequencies through disparate public databases and Websites is time consuming and can result in inconsistent findings. We have developed a web-based software tool, Frequency Finder, to acquire SNP allele frequencies from multiple public data sources and return a summarized result to the user. Our software optimizes and automates the search of candidate markers, decreasing the amount of time it would take to extract pertinent data manually. We have included several methods to output the data, including on-screen and as a compressed text file. We show that Frequency Finder accurately retrieves available frequency data from the available sources. Using this tool, we detect significant differences between Asian, African and Caucasian populations in the allele frequency spectra of 246 097 SNPs. While limited to public databases that provide web-based access to allele frequencies, Frequency Finder provides a single, user-friendly interface for retrieving allele frequencies for large batches of SNPs from multiple data sources.

Algorithms↗

Designing for social data analysis.

The NameVoyager, a Web-based visualization of historical trends in baby naming, has proven remarkably popular. We describe design decisions behind the application and lessons learned in creating an application that makes do-it-yourself data mining popular. The prime lesson, it is hypothesized, is that an information visualization tool may be fruitfully viewed not as a tool but as part of an online social environment. In other words, to design a successful exploratory data analysis tool, one good strategy is to create a system that enables "social" data analysis. We end by discussing the design of an extension of the NameVoyager to a more complex data set, in which the principles of social data analysis played a guiding role.

Computer Graphics↗

Molecular anatomy of an intracranial aneurysm: coordinated expression of genes involved in wound healing and tissue remodeling.

BACKGROUND AND PURPOSE: Approximately 6% of human beings harbor an unruptured intracranial aneurysm. Each year in the United States, >30 000 people suffer a ruptured intracranial aneurysm, resulting in subarachnoid hemorrhage. Despite the high incidence and catastrophic consequences of a ruptured intracranial aneurysm and the fact that there is considerable evidence that predisposition to intracranial aneurysm has a strong genetic component, very little is understood with regard to the pathology and pathogenesis of this disease. METHODS: To begin characterizing the molecular pathology of intracranial aneurysm, we used a global gene expression analysis approach (SAGE-Lite) in combination with a novel data-mining approach to perform a high-resolution transcript analysis of a single intracranial aneurysm, obtained from a 3-year-old girl. RESULTS: SAGE-Lite provides a detailed molecular snapshot of a single intracranial aneurysm. These data suggest that, at least in this specific case, aneurysmal dilation results in a highly dynamic cellular environment in which extensive wound healing and tissue/extracellular matrix remodeling are taking place. Specifically, we observed significant overexpression of genes encoding extracellular matrix components (eg, COL3A1, COL1A1, COL1A2, COL6A1, COL6A2, elastin) and genes involved in extracellular matrix turnover (TIMP-3, OSF-2), cell adhesion and antiadhesion (SPARC, hevin), cytokinesis (PNUTL2), and cell migration (tetraspanin-5). CONCLUSIONS: Although these are preliminary data, representing analysis of only one individual, we present a unique first insight into the molecular basis of aneurysmal disease and define numerous candidate markers for future biochemical, physiological, and genetic studies of intracranial aneurysm. Products of these genes will be the focus of future studies in wider sample sets.

Calcium-Binding Proteins↗

Applications of the double-barreled data in whole-genome shotgun sequence assembly and analysis.

Double-barreled (DB) data have been widely used for the assembly of large genomes. Based on the experience of building the whole-genome working draft of Oryza sativa L. ssp. Indica, we present here the prevailing and improved uses of DB data in the assembly procedure and report on novel applications during the following data-mining processes such as acquiring precise insert fragment information of each clone across the genome, and a new kind of low-cost whole-genome microarray. With the increasing number of organisms being sequenced, we believe that DB data will play an important role both in other assembly procedures and in future genomic studies.

Cloning, Molecular↗

Graphical exploratory data analysis of RNA secondary structure dynamics predicted by the massively parallel genetic algorithm.

Studies indicate that RNA may enter intermediate and multiple conformational states, which may impact gene expression and molecular function. It is known that the biologically functional states of RNA molecules may not correspond to their minimum energy conformations, that kinetic barriers may trap the molecule in a local minimum, that folding often occurs during transcription, and that cases exist in which a molecule will transition between one or more functional conformations. Thus, methods for simulating the folding pathway and dynamic behavior of an RNA molecule are important for the prediction of RNA structure and its associated functions. We have developed several data mining techniques guided by interactive visualization tools associated with our massively parallel genetic algorithm for RNA/DNA secondary structure prediction, MPGAfold, and StructureLab analysis workbench. Most of the methods and tools are also applicable to dynamic programming algorithm (DPA) folding data analysis. When applied to MPGAfold results these methodologies are used to determine the significant intermediate and final structures associated with co-transcriptional and full length RNA folding. Since the genetic algorithm is essentially stochastic, multiple runs are required to develop a consensus understanding of an RNA structure. The interactive visualizations facilitate interpretation of results from sequential or full length individual MPGAfold runs, final results of multiple folding runs, including multiple population sizes, and the results from multiple RNA sequences of one family. This paper describes several of these techniques and shows how they are used to help solve this highly combinatoric problem.

Algorithms↗

Festina lente: late-night thoughts on high-throughput screening of mouse behavior.

A recent perspective discussed high-throughput behavioral analysis using mice, giving the overall impression that this area is lagging behind in neuroscience and biomedical research. Not only are we more optimistic about the current state of the art in behavioral neuroscience and its promise, but we also have reservations about whether high-throughput analysis is always an appropriate goal for most behavioral studies. We argue that behavioral studies should be carried out with clear goals and more regard to the intellectual context in which they have developed. In addition, behavioral studies can be performed quite easily, but this does not ensure the required validity or reliability of the particular tests used. Finally, high throughput may not always be an appropriate goal. We discuss the role of automated data collection and unique data-mining algorithms, and the question of the ethological relevance of behavioral tests.

Algorithms↗

DASS: efficient discovery and p-value calculation of substructures in unordered data.

MOTIVATION: Pattern identification in biological sequence data is one of the main objectives of bioinformatics research. However, few methods are available for detecting patterns (substructures) in unordered datasets. Data mining algorithms mainly developed outside the realm of bioinformatics have been adapted for that purpose, but typically do not determine the statistical significance of the identified patterns. Moreover, these algorithms do not exploit the often modular structure of biological data. RESULTS: We present the algorithm DASS (Discovery of All Significant Substructures) that first identifies all substructures in unordered data (DASS(Sub)) in a manner that is especially efficient for modular data. In addition, DASS calculates the statistical significance of the identified substructures, for sets with at most one element of each type (DASS(P(set))), or for sets with multiple occurrence of elements (DASS(P(mset))). The power and versatility of DASS is demonstrated by four examples: combinations of protein domains in multi-domain proteins, combinations of proteins in protein complexes (protein subcomplexes), combinations of transcription factor target sites in promoter regions and evolutionarily conserved protein interaction subnetworks. AVAILABILITY: The program code and additional data are available at http://www.fli-leibniz.de/tsb/DASS

Algorithms↗

PRIME: a graphical interface for integrating genomic/proteomic databases.

Data mining, finding and integration of information about proteins of interest, is an essential component in modern biological and biomedical research. Even when focusing on a single organism and only on a small number of proteins, there are often dozens fo data sources containing relevant information. We are developing PRIME, a protein information environment, to serve as a virtual central database which integrates distributed heterogeneous information about proteins (linked by common identifier). PRIME has powerful capabilities to visualize all kinds of protein annotation in specialized views. These views can be displayed side by side at the same time and can be synchronized in order to show simultaneously different aspects of identical proteins. These features allow a quick and comprehensive overview of properties of single proteins or protein sets.

Computational Biology↗

Scalable model-based clustering for large databases based on data summarization.

The scalability problem in data mining involves the development of methods for handling large databases with limited computational resources such as memory and computation time. In this paper, two scalable clustering algorithms, bEMADS and gEMADS, are presented based on the Gaussian mixture model. Both summarize data into subclusters and then generate Gaussian mixtures from their data summaries. Their core algorithm, EMADS, is defined on data summaries and approximates the aggregate behavior of each subcluster of data under the Gaussian mixture model. EMADS is provably convergent. Experimental results substantiate that both algorithms can run several orders of magnitude faster than expectation-maximization with little loss of accuracy.

Algorithms↗

An XML Gateway to Patient Data for Medical Research Applications.

As the medical environment becomes increasingly electronic, clinical databases are continually growing, accruing masses of patient information. This wealth of data is an invaluable source of information to researchers, serving as a testbed for the development of new information technologies and as a repository of real-world data for data mining and population-based studies. However, the true utility of this information is not fulfilled, in part because of issues pertaining to security and patient confidentiality, but also due to the lack of an effective infrastructure to access the data. This paper describes a system, DataServer, that permits researchers to query and retrieve data from multiple clinical data sources, automatically deidentifying patient data so that it can be used for research purposes. DataServer functions as an application framework, enabling extensible markup language (XML)-based querying of existing medical databases. Key aspects of DataServer include ready inclusion of new information resources, minimal processing impact on existing clinical systems via a distributed cache, and flexible output representation via XSL (eXtensible Style Language) transforms.

Databases, Factual↗