Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

The University of Minnesota Biocatalysis/Biodegradation Database: post-genomic data mining.

The University of Minnesota Biocatalysis/Biodegradation Database (UM-BBD, http://umbbd.ahc.umn.edu/) provides curated information on microbial catabolism and related biotransformations, primarily for environmental pollutants. Currently, it contains information on over 130 metabolic pathways, 800 reactions, 750 compounds and 500 enzymes. In the past two years, it has increased its breath to include more examples of microbial metabolism of metals and metalloids; and expanded the types of information it includes to contain microbial biotransformations of, and binding interactions with many chemical elements. It has also increased the ways in which this data can be accessed (mined). Structure-based searching was added, for exact matches, similarity, or substructures. Analysis of UM-BBD reactions has lead to a prototype, guided, pathway prediction system. Guided prediction means that the user is shown all possible biotransformations at each step and guides the process to its conclusion. Mining the UM-BBD's data provides a unique view into how the microbial world recycles organic functional groups. UM-BBD users are encouraged to comment on all aspects of the database, including the information it contains and the tools by which it can be mined. The database and prediction system develop under the direction of the scientific community.

Biodegradation, Environmental↗

The genomic response of skeletal muscle to methylprednisolone using microarrays: tailoring data mining to the structure of the pharmacogenomic time series.

High-throughput data collection using gene microarrays has great potential as a method for addressing the pharmacogenomics of complex biological systems. Similarly, mechanism-based pharmacokinetic/pharmacodynamic modeling provides a tool for formulating quantitative testable hypotheses concerning the responses of complex biological systems. As the response of such systems to drugs generally entails cascades of molecular events in time, a time series design provides the best approach to capturing the full scope of drug effects. A major problem in using microarrays for high-throughput data collection is sorting through the massive amount of data in order to identify probe sets and genes of interest. Due to its inherent redundancy, a rich time series containing many time points and multiple samples per time point allows for the use of less stringent criteria of expression, expression change and data quality for initial filtering of unwanted probe sets. The remaining probe sets can then become the focus of more intense scrutiny by other methods, including temporal clustering, functional clustering and pharmacokinetic/pharmacodynamic modeling, which provide additional ways of identifying the probes and genes of pharmacological interest.

Adrenalectomy↗

Haplotype-based linkage disequilibrium mapping via direct data mining.

MOTIVATION: With the availability of large-scale, high-density single-nucleotide polymorphism markers and information on haplotype structures and frequencies, a great challenge is how to take advantage of haplotype information in the association mapping of complex diseases in case-control studies. RESULTS: We present a novel approach for association mapping based on directly mining haplotypes (i.e. phased genotype pairs) produced from case-control data or case-parent data via a density-based clustering algorithm, which can be applied to whole-genome screens as well as candidate-gene studies in small genomic regions. The method directly explores the sharing of haplotype segments in affected individuals that are rarely present in normal individuals. The measure of sharing between two haplotypes is defined by a new similarity metric that combines the length of the shared segments and the number of common alleles around any marker position of the haplotypes, which is robust against recent mutations/genotype errors and recombination events. The effectiveness of the approach is demonstrated by using both simulated datasets and real datasets. The results show that the algorithm is accurate for different population models and for different disease models, even for genes with small effects, and it outperforms some recently developed methods.

Algorithms↗

A novel data mining method to identify assay-specific signatures in functional genomic studies.

BACKGROUND: The highly dimensional data produced by functional genomic (FG) studies makes it difficult to visualize relationships between gene products and experimental conditions (i.e., assays). Although dimensionality reduction methods such as principal component analysis (PCA) have been very useful, their application to identify assay-specific signatures has been limited by the lack of appropriate methodologies. This article proposes a new and powerful PCA-based method for the identification of assay-specific gene signatures in FG studies. RESULTS: The proposed method (PM) is unique for several reasons. First, it is the only one, to our knowledge, that uses gene contribution, a product of the loading and expression level, to obtain assay signatures. The PM develops and exploits two types of assay-specific contribution plots, which are new to the application of PCA in the FG area. The first type plots the assay-specific gene contribution against the given order of the genes and reveals variations in distribution between assay-specific gene signatures as well as outliers within assay groups indicating the degree of importance of the most dominant genes. The second type plots the contribution of each gene in ascending or descending order against a constantly increasing index. This type of plots reveals assay-specific gene signatures defined by the inflection points in the curve. In addition, sharp regions within the signature define the genes that contribute the most to the signature. We proposed and used the curvature as an appropriate metric to characterize these sharp regions, thus identifying the subset of genes contributing the most to the signature. Finally, the PM uses the full dataset to determine the final gene signature, thus eliminating the chance of gene exclusion by poor screening in earlier steps. The strengths of the PM are demonstrated using a simulation study, and two studies of real DNA microarray data--a study of classification of human tissue samples and a study of E. coli cultures with different medium formulations. CONCLUSION: We have developed a PCA-based method that effectively identifies assay-specific signatures in ranked groups of genes from the full data set in a more efficient and simplistic procedure than current approaches. Although this work demonstrates the ability of the PM to identify assay-specific signatures in DNA microarray experiments, this approach could be useful in areas such as proteomics and metabolomics.

Algorithms↗

Visualization of multiple influences on ocellar flight control in giant honeybees with the data-mining tool Viscovery SOMine.

Viscovery SOMine is a software tool for advanced analysis and monitoring of numerical data sets. It was developed for professional use in business, industry, and science and to support dependency analysis, deviation detection, unsupervised clustering, nonlinear regression, data association, pattern recognition, and animated monitoring. Based on the concept of self-organizing maps (SOMs), it employs a robust variant of unsupervised neural networks--namely, Kohonen's Batch-SOM, which is further enhanced with a new scaling technique for speeding up the learning process. This tool provides a powerful means by which to analyze complex data sets without prior statistical knowledge. The data representation contained in the trained SOM is systematically converted to be used in a spectrum of visualization techniques, such as evaluating dependencies between components, investigating geometric properties of the data distribution, searching for clusters, or monitoring new data. We have used this software tool to analyze and visualize multiple influences of the ocellar system on free-flight behavior in giant honeybees. Occlusion of ocelli will affect orienting reactivities in relation to flight target, level of disturbance, and position of the bee in the flight chamber; it will induce phototaxis and make orienting imprecise and dependent on motivational settings. Ocelli permit the adjustment of orienting strategies to environmental demands by enforcing abilities such as centering or flight kinetics and by providing independent control of posture and flight course.

Animals↗

Discovery of predictive models in an injury surveillance database: an application of data mining in clinical research.

A new, evolutionary computation-based approach to discovering prediction models in surveillance data was developed and evaluated. This approach was operationalized in EpiCS, a type of learning classifier system specially adapted to model clinical data. In applying EpiCS to a large, prospective injury surveillance database, EpiCS was found to create accurate predictive models quickly that were highly robust, being able to classify > 99% of cases early during training. After training, EpiCS classified novel data more accurately (p < 0.001) than either logistic regression or decision tree induction (C4.5), two traditional methods for discovering or building predictive models.

Artificial Intelligence↗

Analysis of guideline compliance--a data mining approach.

While guideline-based decision support is safety-critical and typically requires human interaction, offline analysis of guideline compliance can be performed to large extent automatically. We examine the possibility of automatic detection of potential non-compliance followed up with (statistical) association mining. Only frequent associations of non-compliance patterns with various patient data are submitted to medical expert for interpretation. The initial experiment was carried out in the domain of hypertension management.

Decision Making, Computer-Assisted↗

A data mining approach for signal detection and analysis.

The WHO database contains over 2.5 million case reports, analysis of this data set is performed with the intention of signal detection. This paper presents an overview of the quantitative method used to highlight dependencies in this data set. The method Bayesian confidence propagation neural network (BCPNN) is used to highlight dependencies in the data set. The method uses Bayesian statistics implemented in a neural network architecture to analyse all reported drug adverse reaction combinations. This method is now in routine use for drug adverse reaction signal detection. Also this approach has been extended to highlight drug group effects and look for higher order dependencies in the WHO data. Quantitatively unexpectedly strong relationships in the data are highlighted relative to general reporting of suspected adverse effects; these associations are then clinically assessed.

Adverse Drug Reaction Reporting Systems↗

A missing data treatment for data mining applications in medical information systems.

To apply user-friendly, easily operated and accessible tools to handle missing data resulting from an auto-stored medical information system, these tools are applied to satisfy general users from different disciplines (i.e. statistics and machine-learning), followed by medical information system development. This study attempts to develop a new logic separation inference method applied to a database with a format like most real-world medical records containing many missing data and miscellaneous variables. It is expected that this method should have better performance than currently accessible methods. The newly developed logic separation inference method shows a classification power of 0.997 (elimination method is 1), which is better than the simple replacing method (replaced by mode shows 0.974). Both inference methods (mode and mean) have superior classification power to the simple replacing method. The missing data treatment processes introduced in this study can be completed on a MS Excel spreadsheet without any complicated calculation; therefore, they can satisfy general users. This new missing data treatment method is only applied up to 60% of the missing data (missing at random). However, when there is large amount of data, it is expected that this method also can be applied to a database missing more than 60%.

Humans↗

yMGV: helping biologists with yeast microarray data mining.

yMGV (yeast Microarray Global Viewer) was designed to provide biologists with meaningful information from genome-wide yeast expression data. The database includes most of the available expression data published on yeast microarrays over the last 4 years. It provides customizable tools for the rapid visualization of expression profiles associated with a set of genes from all published experiments. It also allows users to compare the results from different publications so that they can identify genes with common expression profiles. We used yMGV to perform global analyses to find a gene expression profile specific for given biological conditions and to locate functional gene clusters on chromosomes. Other organisms will be added to this database. yMGV is accessible on the web at http://transcriptome.ens.fr/ymgv.

Computer Graphics↗

Data mining and enantiophore studies on chiral stationary phases used in HPLC separation.

ChirBase database has been employed to mine the chemical structures of compounds resolved on common commercial chiral stationary phases (CSP). Different data sets were produced. The molecular fingerprint (enantiophore) strategy was then applied over these data sets. Enantiophores are identified by analyzing and mapping the three-dimensional common structural features shared by the ligand molecules imported from ChirBase. Such lists of encoded ligand enantiophores allowed us to generate 3D maps of the molecular interacting fragments, which are supposed to act reciprocally with each CSP. Results show that each CSP is combined with different preferential enantiophore counterparts in the ligand. These differences may be well related to the particular behavior of a given CSP to separate different families of compounds. As expected, supramolecular cellulosic or amylosic CSPs show generalist behavior by resolving a wide range of racemates. On the other hand, molecular CSPs based on a well-defined chiral receptor (such as Whelk-O1 or Chirobiotic T) appear to be more specific and thus specially adapted for more restrictive families of compounds. In addition, enantiophore analyses confirm that the supramolecular CSPs can combine multiple potential binding sites and so offer numerous enantioselective mechanisms toward a ligand. Inversely, on molecular CSPs, the ligand must respect strict geometrical constraints to make possible the chiral discrimination. Additional applications of the methodology are discussed as well.

Journal Article↗