Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Information Management System for Site Remediation Efforts.

/ Environmental regulatory agencies are responsible for protecting human health and the environment in their constituencies. Their responsibilities include the identification, evaluation, and cleanup of contaminated sites. Leaking underground storage tanks (USTs) constitute a major source of subsurface and groundwater contamination. A significant portion of a regulatory body's efforts may be directed toward the management of UST-contaminated sites. In order to manage remedial sites effectively, vast quantities of information must be maintained, including analytical dataon chemical contaminants, remedial design features, and performance details. Currently, most regulatory agencies maintain such information manually. This makes it difficult to manage the data effectively. Some agencies have introduced automated record-keeping systems. However, the ad hoc approach in these endeavors makes it difficult to efficiently analyze, disseminate, and utilize the data. This paper identifies the information requirements for UST-contaminated site management at the Waste Cleanup Section of the Department of Environmental Resources Management in Dade County, Florida. It presents a viable design for an information management system to meet these requirements. The proposed solution is based on a back-end relational database management system with relevant tools for sophisticated data analysis and data mining. The database is designed with all tables in the third normal form to ensure data integrity, flexible access, and efficient query processing. In addition to all standard reports required by the agency, the system provides answers to ad hoc queries that are typically difficult to answer under the existing system. The database also serves as a repository of information for a decision support system to aid engineering design and risk analysis. The system may be integrated with a geographic information system for effective presentation and dissemination of spatial data.

Journal Article↗

How bioinformatics can help reverse engineer human aging.

To study human aging is an enormous challenge. The complexity of the aging phenotype and the near impossibility of studying aging directly in humans oblige researchers to resort to models and extrapolations. Computational approaches offer a powerful set of tools to study human aging. In one direction we have data-mining methods, from comparative genomics to DNA microarrays, to retrieve information in large amounts of data. Afterwards, tools from systems biology to reverse engineering algorithms allow researchers to integrate different types of information to increase our knowledge about human aging. Computer methodologies will play a crucial role to reconstruct the genetic network of human aging and the associated regulatory mechanisms.

Aging↗

Extended SQL for manipulating clinical warehouse data.

Health care institutions are beginning to collect large amounts of clinical data through patient care applications. Clinical data warehouses make these data available for complex analysis across patient records, benefiting administrative reporting, patient care and clinical research. Data gathered for patient care purposes are difficult to manipulate for analytic tasks; the schema presents conceptual difficulties for the analyst, and many queries perform poorly. An extension to SQL is presented that enables the analyst to designate groups of rows. These groups can then be manipulated and aggregated in various ways to solve a number of useful analytic problems. The extended SQL is concise and runs in linear time, while standard SQL requires multiple statements with polynomial performance. The extensions are extremely powerful for performing aggregations on large amounts of data, which is useful in clinical data mining applications.

Clinical Laboratory Techniques↗

Metabolite profiling for plant functional genomics.

Multiparallel analyses of mRNA and proteins are central to today's functional genomics initiatives. We describe here the use of metabolite profiling as a new tool for a comparative display of gene function. It has the potential not only to provide deeper insight into complex regulatory processes but also to determine phenotype directly. Using gas chromatography/mass spectrometry (GC/MS), we automatically quantified 326 distinct compounds from Arabidopsis thaliana leaf extracts. It was possible to assign a chemical structure to approximately half of these compounds. Comparison of four Arabidopsis genotypes (two homozygous ecotypes and a mutant of each ecotype) showed that each genotype possesses a distinct metabolic profile. Data mining tools such as principal component analysis enabled the assignment of "metabolic phenotypes" using these large data sets. The metabolic phenotypes of the two ecotypes were more divergent than were the metabolic phenotypes of the single-loci mutant and their parental ecotypes. These results demonstrate the use of metabolite profiling as a tool to significantly extend and enhance the power of existing functional genomics approaches.

Arabidopsis↗

An integrated tool for microarray data clustering and cluster validity assessment.

UNLABELLED: In this paper we present a data mining system, which allows the application of different clustering and cluster validity algorithms for DNA microarray data. This tool may improve the quality of the data analysis results, and may support the prediction of the number of relevant clusters in the microarray datasets. This systematic evaluation approach may significantly aid genome expression analyses for knowledge discovery applications. The developed software system may be effectively used for clustering and validating not only DNA microarray expression analysis applications but also other biomedical and physical data with no limitations. AVAILABILITY: The program is freely available for non-profit use on request at http://www.cs.tcd.ie/Nadia.Bolshakova/Machaon.html CONTACT: Nadia.Bolshakova@cs.tcd.ie.

Algorithms↗

Implementing a data warehouse at Inglis Innovative Services.

Data warehouses, data marts, and data mining have been hot topics in the 1990s, offering the promise of a vault of corporate data ripe for decision making. As is true with all promising technologies, the key issue is how to get started. Implementation of a corporate data warehouse involves a lot more than spending a huge amount of money on hardware, software, and consultants. Successful implementation of a data warehouse involves a corporate treasure hunt--identifying and cataloging data. It involves data ownership, data integrity, and business process analysis to determine what the data are, who owns them, how reliable they are, and how they are processed. Finally, implementation of the warehouse drives the issue of how good the decisions are that are based on the information in the warehouse. This article presents a case study of how one healthcare facility dealt with the challenges of implementing a data warehouse.

Computer Communication Networks↗

What is bioinformatics? A proposed definition and overview of the field.

BACKGROUND: The recent flood of data from genome sequences and functional genomics has given rise to new field, bioinformatics, which combines elements of biology and computer science. OBJECTIVES: Here we propose a definition for this new field and review some of the research that is being pursued, particularly in relation to transcriptional regulatory systems. METHODS: Our definition is as follows: Bioinformatics is conceptualizing biology in terms of macromolecules (in the sense of physical-chemistry) and then applying "informatics" techniques (derived from disciplines such as applied maths, computer science, and statistics) to understand and organize the information associated with these molecules, on a large-scale. RESULTS AND CONCLUSIONS: Analyses in bioinformatics predominantly focus on three types of large datasets available in molecular biology: macromolecular structures, genome sequences, and the results of functional genomics experiments (e.g. expression data). Additional information includes the text of scientific papers and "relationship data" from metabolic pathways, taxonomy trees, and protein-protein interaction networks. Bioinformatics employs a wide range of computational techniques including sequence and structural alignment, database design and data mining, macromolecular geometry, phylogenetic tree construction, prediction of protein structure and function, gene finding, and expression data clustering. The emphasis is on approaches integrating a variety of computational methods and heterogeneous data sources. Finally, bioinformatics is a practical discipline. We survey some representative applications, such as finding homologues, designing drugs, and performing large-scale censuses. Additional information pertinent to the review is available over the web at http://bioinfo.mbb.yale.edu/what-is-it.

Computational Biology↗

Cancer patient flows discovery in DRG databases.

In France, cancer care is evolving to the design of regional networks, so as to coordinate expertise, services and resources allocation. Existing information systems along with data-mining tools can provide better knowledge on the distribution of patient flows. We used one year data of the French Diagnosis Related Groups (DRGs) based system to perform our analysis. Formal Concept Analysis has been used to build Iceberg Lattices of cancer patient flows in the French region of Lorraine. This unsupervised conceptual clustering method allowed us to describe patients flows with an easily understandable visual representation.

Continuity of Patient Care↗

Syndromic surveillance using automated collection of computerized discharge diagnoses.

The Syndromic Surveillance Information Collection (SSIC) system aims to facilitate early detection of bioterrorism attacks (with such agents as anthrax, brucellosis, plague, Q fever, tularemia, smallpox, viral encephalitides, hemorrhagic fever, botulism toxins, staphylococcal enterotoxin B, etc.) and early detection of naturally occurring disease outbreaks, including large foodborne disease outbreaks, emerging infections, and pandemic influenza. This is accomplished using automated data collection of visit-level discharge diagnoses from heterogeneous clinical information systems, integrating those data into a common XML (Extensible Markup Language) form, and monitoring the results to detect unusual patterns of illness in the population. The system, operational since January 2001, collects, integrates, and displays data from three emergency department and urgent care (ED/UC) departments and nine primary care clinics by automatically mining data from the information systems of those facilities. With continued development, this system will constitute the foundation of a population-based surveillance system that will facilitate targeted investigation of clinical syndromes under surveillance and allow early detection of unusual clusters of illness compatible with bioterrorism or disease outbreaks.

Bioterrorism↗

Observing and interpreting correlations in metabolomic networks.

MOTIVATION: Metabolite profiling aims at an unbiased identification and quantification of all the metabolites present in a biological sample. Based on their pair-wise correlations, the data obtained from metabolomic experiments are organized into metabolic correlation networks and the key challenge is to deduce unknown pathways based on the observed correlations. However, the data generated is fundamentally different from traditional biological measurements and thus the analysis is often restricted to rather pragmatic approaches, such as data mining tools, to discriminate between different metabolic phenotypes. METHODS AND RESULTS: We investigate to what extent the data generated networks reflect the structure of the underlying biochemical pathways. The purpose of this work is 2-fold: Based on the theory of stochastic systems, we first introduce a framework which shows that the emergent correlations can be interpreted as a 'fingerprint' of the underlying biophysical system. This result leads to a systematic relationship between observed correlation networks and the underlying biochemical pathways. In a second step, we investigate to what extent our result is applicable to the problem of reverse engineering, i.e. to recover the underlying enzymatic reaction network from data. The implications of our findings for other bioinformatics approaches are discussed.

Algorithms↗

Enabling proteomics discovery through visual analysis. The peptide permutation and protein prediction tool.

Proteins play a key role in cellular processes, making proteomics central to understanding systems biology. MS techniques provide a means to observe entire proteomes at a global level. Yet, high-throughput MS proteomics techniques generate data faster than it can currently be analyzed. The success of proteomics depends on high-throughput experimental techniques coupled with sophisticated visual analysis and data-mining methods. Visual analysis has been applied successfully in a number of fields plagued with huge, complex data sets and will likely be an important tool in proteomics discovery. PQuad, a novel visualization of MS proteomics data, provides powerful analysis capabilities that support a number of proteomic data applications. In particular, PQuad supports differential proteomics by simplifying the comparison of peptide sets from different experimental conditions as well as different protein identification or confidence scoring techniques. Finally, PQuad supports data validation and quality control by providing a variety of resolutions for huge amounts of data to reveal errors undetected by other methods.

Algorithms↗

Building manageable rough set classifiers.

An interesting aspect of techniques for data mining and knowledge discovery is their potential for generating hypotheses by discovering underlying relationships buried in the data. However, the set of possible hypotheses is often very large and the extracted models may become prohibitively complex. It is therefore typically desirable to only consider the "strongest" hypotheses, so that smaller models can be obtained that also retain good classificatory capabilities. This paper outlines how rule-based classifiers based on rough set theory and Boolean reasoning that are both small and perform well can be developed. Applied to a real-world medical dataset, the final models are shown to exhibit good performance using only a subset of the available information. Furthermore, the number of resulting rules is low and enables practical a posteriori inspection and interpretation of the models.

Classification↗

SCI-Base: an open-source spinal cord injury animal experimentation database.

To capture all pertinent data during spinal cord injury animal experimentation, the authors have designed and implemented a database that is available for use under a public license. Their goals were to record all daily medical care of paraplegic animals, including unexpected complications; to store all injury parameters and/or therapeutic procedures; to track locomotor scores and other measures of functional recovery; to allow planning and management of experiments; and to serve as an externally linkable, SQL-queryable data mining source. Ultimately, the use of databases such as this will allow multiple neurotrauma laboratories to compare animal data through web meta-analysis.

Animal Husbandry↗

Some statistical and regulatory issues in the evaluation of genetic and genomic tests.

The genomics revolution is reverberating throughout the worlds of pharmaceutical drugs, genetic testing and statistical science. This revolution, which uses single nucleotide polymorphisms (SNPs) and gene expression technology, including cDNA and oligonucleotide microarrays, for a range of tests from home-brews to high-complexity lab kits, can allow the selection or exclusion of patients for therapy (responders or poor metabolizers). The wide variety of US regulatory mechanisms for these tests is discussed. Clinical studies to evaluate the performance of such tests need to follow statistical principles for sound diagnostic test design. Statistical methodology to evaluate such studies can be wide ranging, including receiver operating characteristic (ROC) methodology, logistic regression, discriminant analysis, multiple comparison procedures resampling, Bayesian hierarchical modeling, recursive partitioning, as well as exploratory techniques such as data mining. Recent examples of approved genetic tests are discussed.

Animals↗

Quantitative quality control in microarray experiments and the application in data filtering, normalization and false positive rate prediction.

Data preprocessing including proper normalization and adequate quality control before complex data mining is crucial for studies using the cDNA microarray technology. We have developed a simple procedure that integrates data filtering and normalization with quantitative quality control of microarray experiments. Previously we have shown that data variability in a microarray experiment can be very well captured by a quality score q(com) that is defined for every spot, and the ratio distribution depends on q(com). Utilizing this knowledge, our data-filtering scheme allows the investigator to decide on the filtering stringency according to desired data variability, and our normalization procedure corrects the q(com)-dependent dye biases in terms of both the location and the spread of the ratio distribution. In addition, we propose a statistical model for false positive rate determination based on the design and the quality of a microarray experiment. The model predicts that a lower limit of 0.5 for the replicate concordance rate is needed in order to be certain of true positives. Our work demonstrates the importance and advantages of having a quantitative quality control scheme for microarrays.

Algorithms↗

Effects of replacing the unreliable cDNA microarray measurements on the disease classification based on gene expression profiles and functional modules.

MOTIVATION: Microarrays datasets frequently contain a large number of missing values (MVs), which need to be estimated and replaced for subsequent data mining. The focus of the paper is to study the effects of different MV treatments for cDNA microarray data on disease classification analysis. RESULTS: By analyzing five datasets, we demonstrate that among three kinds of classifiers evaluated in this study, support vector machine (SVM) classifiers are robust to varied MV imputation methods [e.g. replacing MVs by zero, K nearest-neighbor (KNN) imputation algorithm, local least square imputation and Bayesian principal component analysis], while the classification and regression tree classifiers are sensitive in terms of classification accuracy. The KNNclassifiers built on differentially expressed genes (DEGs) are robust to the varied MV treatments, but the performances of the KNN classifiers based on all measured genes can be significantly deteriorated when imputing MVs for genes with larger missing rate (MR) (e.g. MR > 5%). Generally, while replacing MVs by zero performs relatively poor, the other imputation algorithms have little difference in affecting classification performances of the SVM or KNN classifiers. We further demonstrate the power and feasibility of our recently proposed functional expression profile (FEP) approach as means to handle microarray data with MVs. The FEPs, which are derived from the functional modules that are enriched with sets of DEGs and thus can be consistently identified under varied MV treatments, achieve precise disease classification with better biological interpretation. We conclude that the choice of MV treatments should be determined in context of the later approaches used for disease classification. The suggested exclusion criterion of ignoring the genes with larger MR (e.g. >5%), while justifiable for some classifiers such as KNN classifiers, might not be considered as a general rule for all classifiers.

Algorithms↗

Local dimensionality reduction and supervised learning within natural clusters for biomedical data analysis.

Inductive learning systems were successfully applied in a number of medical domains. Nevertheless, the effective use of these systems often requires data preprocessing before applying a learning algorithm. This is especially important for multidimensional heterogeneous data presented by a large number of features of different types. Dimensionality reduction (DR) is one commonly applied approach. The goal of this paper is to study the impact of natural clustering--clustering according to expert domain knowledge--on DR for supervised learning (SL) in the area of antibiotic resistance. We compare several data-mining strategies that apply DR by means of feature extraction or feature selection with subsequent SL on microbiological data. The results of our study show that local DR within natural clusters may result in better representation for SL in comparison with the global DR on the whole data.

Algorithms↗

Applying database technology to clinical and basic research bioinformatics projects.

This paper describes the application of database technology to medical information with the goal of providing medical and clinical researchers with the tools necessary to plan bioinformatics projects. Commercial database management systems were utilized, standard database design practices were applied, a user interface was created, data entered, and the development of analysis tools, including data mining technologies is underway. Databases were constructed based on animal and cell culture models of diabetes and clinical data. Bioinformatics is a useful tool in both basic research and clinical settings. The advantages of relational databases and an approach to managing bioinformatics projects are discussed.

Animals↗