Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Universal electronic health record MUDR.

One of the important research tasks of the European Centre for Medical Informatics, Statistics and Epidemiology - Cardio (EuroMISE Centre - Cardio) is the applied research in the field of electronic health record design including electronic medical guidelines and intelligent systems for data mining and decision support. The research in the field of data storage and data acquisition was inspired by several European projects and standards, mostly by the I4C and TripleC projects. Based on experience gathered during cooperation in the TripleC project we have proposed a description of a flexible information storage model. The motivation for this effort was the large variability of the set of collected features in different departments - including temporal variability. Therefore, a dynamically extensible and modifiable structure of items is needed. In our model we use two basic structures called the knowledge base and data files. The main function of the knowledge base is to express the hierarchy of collectable features - medical concepts, their characteristics and relations among them. The data files structure is used to store the patient's data itself. These two structures can be described using graph theory expressions. Based on this model, a three-layer system architecture named "Multimedia Distributed Record" (MUDR) has been proposed and implemented. During the implementation, modern technologies such as Web Services, SOAP and XML were used. For the practical usage of EHR MUDR, an intelligent application called MUDRc (MUDR Client) was created. It enables physicians to use EHR MUDR in a flexible way. During the development process, maximum emphasis was placed on user-friendliness and comfortable usage of this application. Several methods of data entry can be used: pre-defined forms, direct entry into the tree data structure of the EHR MUDR, or automatic unstructured free-text report parsing and data retrieval. The system enables fast and simple importing and exporting of data as well. The system integrates modern multimedia formats (X-ray photos, sonography and other pictures, video-sequences, audio records) as well as progressive methods of decision support systems realized by medical guidelines and other modules.

Artificial Intelligence↗

Gene expression informatics--it's all in your mine.

Technologies for whole-genome RNA expression studies are becoming increasingly reliable and accessible. However, universal standards to make the data more suitable for comparative analysis and for inter-operability with other information resources have yet to emerge. Improved access to large electronic data sets, reliable and consistent annotation and effective tools for 'data mining' are critical. Analysis methods that exploit large data warehouses of gene expression experiments will be necessary to realize the full potential of this technology.

Animals↗

Linking tumor cell cytotoxicity to mechanism of drug action: an integrated analysis of gene expression, small-molecule screening and structural databases.

An integrated, bioinformatic analysis of three databases comprising tumor-cell-based small molecule screening data, gene expression measurements, and PDB (Protein Data Bank) ligand-target structures has been developed for probing mechanism of drug action (MOA). Clustering analysis of GI50 profiles for the NCI's database of compounds screened across a panel of tumor cells (NCI60) was used to select a subset of unique cytotoxic responses for about 4000 small molecules. Drug-gene-PDB relationships for this test set were examined by correlative analysis of cytotoxic response and differential gene expression profiles within the NCI60 and structural comparisons with known ligand-target crystallographic complexes. A survey of molecular features within these compounds finds thirteen conserved Compound Classes, each class exhibiting chemical features important for interactions with a variety of biological targets. Protein targets for an additional twelve Compound Classes could be directly assigned using drug-protein interactions observed in the crystallographic database. Results from the analysis of constitutive gene expressions established a clear connection between chemo-resistance and overexpression of gene families associated with the extracellular matrix, cytoskeletal organization, and xenobiotic metabolism. Conversely, chemo-sensitivity implicated overexpression of gene families involved in homeostatic functions of nucleic acid repair, aryl hydrocarbon metabolism, heat shock response, proteasome degradation and apoptosis. Correlations between chemo-responsiveness and differential gene expressions identified chemotypes with nonselective (i.e., many) molecular targets from those likely to have selective (i.e., few) molecular targets. Applications of data mining strategies that jointly utilize tumor cell screening, genomic, and structural data are presented for hypotheses generation and identifying novel anticancer candidates.

Antineoplastic Agents↗

Influence of statins on postoperative wound complications after inguinal or ventral herniorrhaphy.

The lipid-lowering agents, statins, are the most commonly prescribed class of drugs in the western world. Because of their widespread use, many patients undergo surgical procedures while on statins. Statins, in addition to cholesterol-lowering effects, also have anticoagulant, immunosuppressive, and antiproliferative properties that may affect the risk of local wound complications. This study investigated the relationship between statins and postoperative wound complications in a large cohort of patients undergoing inguinal or ventral hernia repair. Data mining was performed in the Veterans Integrated Service Network (VISN)16 Data Warehouse. This database contains clinical and demographic information about all veterans cared for at the ten VA Medical Centers that comprise the South Central VA Healthcare Network in the mid-south region of the US. Aggregate data (age, body mass index, smoking history, gender, race, history of diabetes, statin use, and postoperative wound complications) were obtained for all patients who underwent inguinal or ventral hernia repair during the period October 1, 1996-November 30, 2004. During the period of the query, 10,782 patients (10,676 male, 106 female), 1,242 (11.5%) of whom received statins, underwent herniorrhaphy. Statin use did not affect the risk of wound infection or delayed wound healing. Statin use was, however, associated with an increased rate of local postoperative bleeding complications (P=0.01). When the type of hernia, age, smoking, diabetes, and body mass index were included in a multivariate analysis, statins remained borderline significant as an independent predictor of wound hematoma/postoperative bleeding (P=0.04), odds ratio 1.6 (95% CI 1.03-2.44). Patients who undergo inguinal herniorrhaphy while on statins have an increased risk of postoperative wound hematoma/hemorrhage. Focus on additional factors that may affect the propensity to postoperative bleeding and on meticulous intraoperative hemostasis are particularly important in such patients.

Diabetes Mellitus↗

Designing a decision support system for existing clinical organizational structures: considerations from a rheumatology clinic.

The aim of this study was to identify the social and organizational requirements for a decision support system (DSS) to be implemented in a clinical rheumatology setting, utilizing data-mining techniques. Field observations and focus group interviews were used for data collection. The decision-making was found to be situated, patient-focused, and long-term in nature. At the same time, the main part of peer-to-peer communication was informal. Patient records were involved in almost every decision. The conclusion is that the main challenges, when introducing a DSS at a rheumatology unit, are adapting the system to informal communication structures and integrating it with patient records. Considering incentive structures, understanding workflow and incorporating awareness are relevant issues when addressing these issues in future studies.

Adolescent↗

Classification of smoking cessation status with a backpropagation neural network.

This study examined the ability of a backpropagation neural network (BPNN) classifier to distinguish between current and former smokers in the 2000 National Health Interview Survey (NHIS) sample adult file. The BPNN classifier performance exceeded that of random chance, with asymmetric 95% confidence intervals for A(z) (area under receiver operating characteristic curve)=(0.7532, 0.7790). Separation of current and former smokers was imperfect, as illustrated by the receiver operating characteristic (ROC) curve. Additionally, performance did not exceed that of a comparison classifier created using logistic regression. Attribute subset selection identified three novel attributes related to smoking cessation status. This study establishes the ability of backpropagation neural networks to classify a complex health behavior, smoking cessation. It also illustrates the hypothesis-generating capacity of data mining methods when applied to large population-based health survey data. Ultimately, BPNN classifiers of smoking cessation status may be useful in decision support systems for smoking cessation interventions.

Algorithms↗

Data storage and analysis in ArrayExpress.

ArrayExpress is a public resource for microarray data that has two major goals: to serve as an archive providing access to microarray data supporting publications and to build a knowledge base of gene expression profiles. ArrayExpress consists of two tightly integrated databases: ArrayExpress repository, which is an archive, and ArrayExpress data warehouse, which contains reannotated data and is optimized for queries. As of December 2005, ArrayExpress contains gene expression and other microarray data from almost 35,000 hybridizations, comprising over 1200 studies, covering 70 different species. Most data are related to peer-reviewed publications. Password-protected access to prepublication data is provided for reviewers and authors. Data in the repository can be queried by various parameters such as species, authors, or words used in the experiment description. The data warehouse provides a wide range of queries, including ones based on gene and sample properties, and provides capabilities to retrieve data combined from different studies. The ArrayExpress resource also includes Expression Profiler (EP)-a microarray data mining, analysis, and visualization tool-and MIAMExpress-an online data submission tool. This chapter describes all major ArrayExpress components from the user perspective: how to submit to, retrieve from, and analyze data in ArrayExpress.

Animals↗

Visualization and interactive analysis of blood parameters with InfoZoom.

This paper describes the application of the data analysis tool InfoZoom to a database containing the results of blood examinations for about 400 patients with a suspect of thrombosis. The main goal was to find correlations between the measurements and the occurrence of a thrombosis. No automatic method for data mining is used. Instead, InfoZoom uses a novel technique to display data sets as highly compressed tables which always fit completely onto the screen. The user can interactively explore animated tabular views of the data. In this way, the user gets a feeling of the data, detects interesting knowledge, and gains a deep understanding of the data set.

Databases, Factual↗

Alcohol ignition interlock programs.

The alcohol ignition interlock is an in-vehicle DWI control device that prevents a car from starting until the operator provides a breath alcohol concentration (BAC) test below a set level, usually .02% (20 mg/dl) to .04% (40 mg/dl). The first interlock program was begun as a pilot test in California 18 years ago; today all but a few US states, and Canadian provinces have interlock enabling legislation. Sweden has recently implemented a nationwide interlock program. Other nations of the European Union and as well as several Australian states are testing it on a small scale or through pilot research. This article describes the interlock device and reviews the development and current status of interlock programs including their public safety benefit and the public practice impediments to more widespread adoption of these DWI control devices. Included in this review are (1) a discussion of the technological breakthroughs and certification standards that gave rise to the design features of equipment that is in widespread use today; (2) a commentary on the growing level of adoption of interlocks by governments despite the judicial and legislative practices that prevent more widespread use of them; (3) a brief overview of the extant literature documenting a high degree of interlock efficacy while installed, and the rapid loss of their preventative effect on repeat DWI once they are removed from the vehicles; (4) a discussion of the representativeness of subjects in the current research studies; (5) a discussion of research innovations, including motivational intervention efforts that may extend the controlling effect of the interlock, and data mining research that has uncovered ways to use the stored interlock data record of BAC tests in order to predict high risk drivers; and (6) a discussion of communication barriers and conceptual rigidities that may be preventing the alcohol ignition interlock from taking a more prominent role in the arsenal of tools used to control DWI. Whether interlock programs can help public policymakers achieve their expressed goals of substantially reducing the level of impaired driving will remain uncertain until procedural barriers and intransigent judiciary practices can be overcome that provide for more systematic routine use of interlock programs. Despite strong effectiveness evidence in all studies to date, the real potential of this technology to reduce the road toll cannot be estimated until they are more widely adopted.

Alcohol Drinking↗

Advanced database methodology for the Collation of Connectivity data on the Macaque brain (CoCoMac).

The need to integrate massively increasing amounts of data on the mammalian brain has driven several ambitious neuroscientific database projects that were started during the last decade. Databasing the brain's anatomical connectivity as delivered by tracing studies is of particular importance as these data characterize fundamental structural constraints of the complex and poorly understood functional interactions between the components of real neural systems. Previous connectivity databases have been crucial for analysing anatomical brain circuitry in various species and have opened exciting new ways to interpret functional data, both from electrophysiological and from functional imaging studies. The eventual impact and success of connectivity databases, however, will require the resolution of several methodological problems that currently limit their use. These problems comprise four main points: (i) objective representation of coordinate-free, parcellation-based data, (ii) assessment of the reliability and precision of individual data, especially in the presence of contradictory reports, (iii) data mining and integration of large sets of partially redundant and contradictory data, and (iv) automatic and reproducible transformation of data between incongruent brain maps. Here, we present the specific implementation of the 'collation of connectivity data on the macaque brain' (CoCoMac) database (http://www.cocomac.org). The design of this database addresses the methodological challenges listed above, and focuses on experimental and computational neuroscientists' needs to flexibly analyse and process the large amount of published experimental data from tracing studies. In this article, we explain step-by-step the conceptual rationale and methodology of CoCoMac and demonstrate its practical use by an analysis of connectivity in the prefrontal cortex.

Animals↗

Optimization of neural network architecture using genetic programming improves detection and modeling of gene-gene interactions in studies of human diseases.

BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.

Algorithms↗

The bovine QTL viewer: a web accessible database of bovine Quantitative Trait Loci.

BACKGROUND: Many important agricultural traits such as weight gain, milk fat content and intramuscular fat (marbling) in cattle are quantitative traits. Most of the information on these traits has not previously been integrated into a genomic context. Without such integration application of these data to agricultural enterprises will remain slow and inefficient. Our goal was to populate a genomic database with data mined from the bovine quantitative trait literature and to make these data available in a genomic context to researchers via a user friendly query interface. DESCRIPTION: The QTL (Quantitative Trait Locus) data and related information for bovine QTL are gathered from published work and from existing databases. An integrated database schema was designed and the database (MySQL) populated with the gathered data. The bovine QTL Viewer was developed for the integration of QTL data available for cattle. The tool consists of an integrated database of bovine QTL and the QTL viewer to display QTL and their chromosomal position. CONCLUSION: We present a web accessible, integrated database of bovine (dairy and beef cattle) QTL for use by animal geneticists. The viewer and database are of general applicability to any livestock species for which there are public QTL data. The viewer can be accessed at http://bovineqtl.tamu.edu.

Animals↗

Quantitative collagen as a golden standard in differential diagnosing of fibrotic changes in liver tissue.

Determining a presence and degree of liver fibrosis provides means for diagnosing disease related processes. We have used two data mining methods, discriminant and regression analyses, to acquire knowledge from the data of 211 patients. We have shown and discussed that quantitative collagen has a distinguished discriminating power and can serve as a golden standard. We have additionally succeeded to obtain a formula consisting of standardised blood tests that can replace quantitative collagen. Practical implications of this is a non-invasive and cost efficient patient examination. All the results are now left for clinical evaluation and so is the current way of histopathological classifications.

Biopsy↗

Liver guide for monitoring of chronic hepatitis C.

The severity of chronic hepatitis C infection in the individual patient is monitored using blood laboratory findings and liver biopsy. If blood test results could be shown to provide sufficient information concerning the disease, the invasive procedure of liver biopsy could perhaps be avoided in some instances. This study assessed the clinical relevance of blood laboratory tests for detecting disease-related changes in the liver. Histopathological classification was used to assign class membership of the patients and data mining operations were performed in an elaborate way on 19 different data sets. Disease activity could be detected by a small set of blood tests. Extended sets could identify more severe changes, but failed to distinguish them. The extracted rules are implemented as a part of the knowledge base of a corresponding decision support system aimed at specialists and general practitioners.

Analysis of Variance↗

Generalizable mass spectrometry mining used to identify disease state biomarkers from blood serum.

We bring a "spectrum" of classical data mining and statistical analysis methods to bear on discrimination of two groups of spectra from 24 diseased and 17 normal patients. Our primary goal is to accurately estimate the generalizability of this small dataset. After an aggressive preprocessing step that reduces consideration to only 55 peaks, we conduct over 35 out-of-sample cross-validation simulations of logistic regression, binary decision trees, and linear discriminant analysis. Misclassification rates grow worse as the size of the holdout sample increases, with many exceeding 30 percent. The ability to generalize is clearly tempered by the statistical, instrumentation, and biophysical characteristics of the study.

Biomarkers↗

Bayesian graphical models for genomewide association studies.

As the extent of human genetic variation becomes more fully characterized, the research community is faced with the challenging task of using this information to dissect the heritable components of complex traits. Genomewide association studies offer great promise in this respect, but their analysis poses formidable difficulties. In this article, we describe a computationally efficient approach to mining genotype-phenotype associations that scales to the size of the data sets currently being collected in such studies. We use discrete graphical models as a data-mining tool, searching for single- or multilocus patterns of association around a causative site. The approach is fully Bayesian, allowing us to incorporate prior knowledge on the spatial dependencies around each marker due to linkage disequilibrium, which reduces considerably the number of possible graphical structures. A Markov chain-Monte Carlo scheme is developed that yields samples from the posterior distribution of graphs conditional on the data from which probabilistic statements about the strength of any genotype-phenotype association can be made. Using data simulated under scenarios that vary in marker density, genotype relative risk of a causative allele, and mode of inheritance, we show that the proposed approach has better localization properties and leads to lower false-positive rates than do single-locus analyses. Finally, we present an application of our method to a quasi-synthetic data set in which data from the CYP2D6 region are embedded within simulated data on 100K single-nucleotide polymorphisms. Analysis is quick (<5 min), and we are able to localize the causative site to a very short interval.

Bayes Theorem↗

Childhood acute lymphoblastic leukemia in the age of genomics.

The recent sequencing of the human genome and technical breakthroughs now make it possible to simultaneously determine mRNA expression levels of almost all of the identified genes in the human genome. DNA "chip" or microarray technology holds great promise for the development of more refined, biologically-based classification systems for childhood ALL, as well as the identification of new targets for novel therapy. To date gene expression profiles have been described that correlate with subtypes of ALL defined by morphology, immunophenotype, cytogenetic alterations, and response to therapy. Mechanistic insights into treatment failure have come from the definition of mRNA signatures that predict in vitro chemoresistance, as well as differences between blasts at relapse and new diagnosis. New bioinformatics tools optimize data mining, but validation of findings is essential since "over-fitting" the data is a common danger. In the future, genomic analysis will be complemented by evaluation of the cancer proteome.

Child↗

Accurate mass filtering of ion chromatograms for metabolite identification using a unit mass resolution liquid chromatography/mass spectrometry system.

Acceleration of liquid chromatography/mass spectrometric (LC/MS) analysis for metabolite identification critically relies on effective data processing since the rate of data acquisition is much faster than the rate of data mining. The rapid and accurate identification of metabolite peaks from complex LC/MS data is a key component to speeding up the process. Current approaches routinely use selected ion chromatograms that can suffer severely from matrix effects. This paper describes a new method to automatically extract and filter metabolite-related information from LC/MS data obtained at unit mass resolution in the presence of complex biological matrices. This approach is illustrated by LC/MS analysis of the metabolites of verapamil from a rat microsome incubation spiked with biological matrix (bile). MS data were acquired in profile mode on a unit mass resolution triple-quadrupole instrument, externally calibrated using a unique procedure that corrects for both mass axis and mass spectral peak shape to facilitate metabolite identification with high mass accuracy. Through the double-filtering effects of accurate mass and isotope profile, conventional extracted ion chromatograms corresponding to the parent drug (verapamil at m/z 455), demethylated verapamil (m/z 441), and dealkylated verapamil (m/z 291), that contained substantial false-positive peaks, were simplified into chromatograms that are substantially free from matrix interferences. These filtered chromatograms approach what would have been obtained by using a radioactivity detector to detect radio-labeled metabolites of interest.

Animals↗