Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

A data mining approach for signal detection and analysis.

The WHO database contains over 2.5 million case reports, analysis of this data set is performed with the intention of signal detection. This paper presents an overview of the quantitative method used to highlight dependencies in this data set. The method Bayesian confidence propagation neural network (BCPNN) is used to highlight dependencies in the data set. The method uses Bayesian statistics implemented in a neural network architecture to analyse all reported drug adverse reaction combinations. This method is now in routine use for drug adverse reaction signal detection. Also this approach has been extended to highlight drug group effects and look for higher order dependencies in the WHO data. Quantitatively unexpectedly strong relationships in the data are highlighted relative to general reporting of suspected adverse effects; these associations are then clinically assessed.

Adverse Drug Reaction Reporting Systems↗

A missing data treatment for data mining applications in medical information systems.

To apply user-friendly, easily operated and accessible tools to handle missing data resulting from an auto-stored medical information system, these tools are applied to satisfy general users from different disciplines (i.e. statistics and machine-learning), followed by medical information system development. This study attempts to develop a new logic separation inference method applied to a database with a format like most real-world medical records containing many missing data and miscellaneous variables. It is expected that this method should have better performance than currently accessible methods. The newly developed logic separation inference method shows a classification power of 0.997 (elimination method is 1), which is better than the simple replacing method (replaced by mode shows 0.974). Both inference methods (mode and mean) have superior classification power to the simple replacing method. The missing data treatment processes introduced in this study can be completed on a MS Excel spreadsheet without any complicated calculation; therefore, they can satisfy general users. This new missing data treatment method is only applied up to 60% of the missing data (missing at random). However, when there is large amount of data, it is expected that this method also can be applied to a database missing more than 60%.

Humans↗

yMGV: helping biologists with yeast microarray data mining.

yMGV (yeast Microarray Global Viewer) was designed to provide biologists with meaningful information from genome-wide yeast expression data. The database includes most of the available expression data published on yeast microarrays over the last 4 years. It provides customizable tools for the rapid visualization of expression profiles associated with a set of genes from all published experiments. It also allows users to compare the results from different publications so that they can identify genes with common expression profiles. We used yMGV to perform global analyses to find a gene expression profile specific for given biological conditions and to locate functional gene clusters on chromosomes. Other organisms will be added to this database. yMGV is accessible on the web at http://transcriptome.ens.fr/ymgv.

Computer Graphics↗

Data mining for on-line support of general practice.

Statistical relationships among symptoms, diagnoses and treatments can be inferred from large data-bases of health records. We investigate how these "empirical norms" can be utilized to improve the efficiency and quality assurance capability of on-line systems in General Practice medicine. Using a survey-based database of General Practice records, we assess hotlists (case sensitive menus) of diagnoses to speed data entry. We also explore norm violation as an indicator of poor quality in practice or data recording. We find that efficient hotlists of diagnoses can be generated from symptoms. Also, we find that less frequently used hypertension treatments are assessed as of a lower quality than more common ones. The results support the hypothesis that empirical norms have a role to play in improved General Practice systems.

Australia↗

Data mining by clinicians.

Clinical databases are becoming commonplace in healthcare environments. However, clinicians have been unable to readily explore these information sources, because currently available data retrieval tools require substantial technical skill, as well as knowledge of the underlying database structures. To address this, we have defined a group of "atomic" queries, including both population-based and temporal predicates, to enable the extraction of clinically meaningful information form these databases. DXtractor is an application that incorporates this functionality, and allows clinicians to simply combine these atomic queries. In doing so, arbitrarily complex data retrieval and exploration becomes possible for the non-programming clinician.

Databases as Topic↗

Construction of a generic reaction knowledge base by reaction data mining.

As synthesis by combinatorial chemistry and high throughput screening have become well-established strategies in the drug discovery process, chemists face increased challenges in managing large amounts of data and using these data to design more diverse and focused libraries. As synthesis is an intuitive and empirical process, however, the classical approaches to computer-assisted synthesis planning do not fully satisfy the needs of the synthetic chemist. We describe a novel computational technique for extracting reaction data and building a generic reaction knowledge base (GRKB) to provide chemists with useful and well-organized knowledge. The method consists of three key steps: (1) the automatic recognition of reaction centers, (2) the definition of a hierarchy of reaction patterns, and (3) the organization of the generic reaction knowledge. Significant reaction knowledge has been discovered via mining a subset of the InfoChem Reaction database. A frame system has been constructed to store and retrieve the GRKB. Applications of this GRKB to synthesis planning are illustrated.

Artificial Intelligence↗

A new web-based data mining tool for the identification of candidate genes for human genetic disorders.

To identify the gene underlying a human genetic disorder can be difficult and time-consuming. Typically, positional data delimit a chromosomal region that contains between 20 and 200 genes. The choice then lies between sequencing large numbers of genes, or setting priorities by combining positional data with available expression and phenotype data, contained in different internet databases. This process of examining positional candidates for possible functional clues may be performed in many different ways, depending on the investigator's knowledge and experience. Here, we report on a new tool called the GeneSeeker, which gathers and combines positional data and expression/phenotypic data in an automated way from nine different web-based databases. This results in a quick overview of interesting candidate genes in the region of interest. The GeneSeeker system is built in a modular fashion allowing for easy addition or removal of databases if required. Databases are searched directly through the web, which obviates the need for data warehousing. In order to evaluate the GeneSeeker tool, we analysed syndromes with known genesis. For each of 10 syndromes the GeneSeeker programme generated a shortlist that contained a significantly reduced number of candidate genes from the critical region, yet still contained the causative gene. On average, a list of 163 genes based on position alone was reduced to a more manageable list of 22 genes based on position and expression or phenotype information. We are currently expanding the tool by adding other databases. The GeneSeeker is available via the web-interface (http://www.cmbi.kun.nl/GeneSeeker/).

Computational Biology↗

Predicting patient's long-term clinical status after hip arthroplasty using hierarchical decision modelling and data mining.

Construction of a prognostic model is presented for the long-term outcome after femoral neck fracture treatment with implantation of hip endoprosthesis. While the model is induced from the follow-up data, we show that the use of additional expert knowledge is absolutely crucial to obtain good predictive accuracy. A schema is proposed where domain knowledge is encoded as a hierarchical decision model of which only a part is induced from the data while the rest is specified by the expert. Although applied to hip endoprosthesis domain, the proposed schema is general and can be used for the construction of other prognostic models where both follow-up data and human expertise is available.

Aged↗

Chem-tox informatics: data mining using a medicinal chemistry building block approach.

Relating chemical structure to biological activity is not a new endeavor, however, the ability to do this on large datasets is just emerging. To cope with the enormous amounts of data being generated, an assortment of computational methods has been developed in the fields of chemoinformatics and computational toxicology. Many of the molecular descriptors used in these approaches are abstract, theoretical constructs that are difficult to understand and visualize. Having easily recognized chemical features, such as those in several new programs, will allow chemists to use toxicological information (or any biological information) when designing new libraries. These improved chem-tox informatics systems will have an impact on library design, hit and lead optimization, development candidate testing and regulatory review.

Animals↗

GeneCards: a novel functional genomics compendium with automated data mining and query reformulation support.

MOTIVATION: Modern biology is shifting from the 'one gene one postdoc' approach to genomic analyses that include the simultaneous monitoring of thousands of genes. The importance of efficient access to concise and integrated biomedical information to support data analysis and decision making is therefore increasing rapidly, in both academic and industrial research. However, knowledge discovery in the widely scattered resources relevant for biomedical research is often a cumbersome and non-trivial task, one that requires a significant amount of training and effort. RESULTS: To develop a model for a new type of topic-specific overview resource that provides efficient access to distributed information, we designed a database called 'GeneCards'. It is a freely accessible Web resource that offers one hypertext 'card' for each of the more than 7000 human genes that currently have an approved gene symbol published by the HUGO/GDB nomenclature committee. The presented information aims at giving immediate insight into current knowledge about the respective gene, including a focus on its functions in health and disease. It is compiled by Perl scripts that automatically extract relevant information from several databases, including SWISS-PROT, OMIM, Genatlas and GDB. Analyses of the interactions of users with the Web interface of GeneCards triggered development of easy-to-scan displays optimized for human browsing. Also, we developed algorithms that offer 'ready-to-click' query reformulation support, to facilitate information retrieval and exploration. Many of the long-term users turn to GeneCards to quickly access information about the function of very large sets of genes, for example in the realm of large-scale expression studies using 'DNA chip' technology or two-dimensional protein electrophoresis. AVAILABILITY: Freely available at http://bioinformatics.weizmann.ac.il/cards/ CONTACT: cards@bioinformatics.weizmann.ac.il

Algorithms↗

Clinical and pharmacogenomic data mining: 1. Generalized theory of expected information and application to the development of tools.

New scientific problems, arising from the human genome project, are challenging the classical means of using statistics. Yet quantified knowledge in the form of rules and rule strengths based on real relationships in data, as opposed to expert opinion, is urgently required for researcher and physician decision support. The problem is that with many parameters, the space to be analyzed is highly dimensional. That is, the combinations of data to examine are subject to a combinatorial explosion as the number of possible events (entries, items, sub-records) (a),(b),(c),... per record (a,b,c,..) increases, and hence much of the space is sparsely populated. These combinatorial considerations are particularly problematic for identifying those associations called "Unicorn Events" which occur significantly less than expected to the extent that they are never seen to be counted. To cope with the combinatorial explosion, a novel numerical "book keeping" approach is taken to generate information terms relating to the combinatorial subsets of events (a,b,c,..), and, most importantly, the zeta (Zeta) function is employed. The incomplete Zeta function zeta(s,n) with s = 1, in which frequencies of occurrence such as n = n(a,b,c,...) determine the range of summation n, is argued to be the natural choice of information function. It emerges from Bayesian integration, taken over the distribution of possible values of information measures for sparse and ample data alike. Expected mutual information l(a;b;c) in nats (i.e., natural units analogous to bits but based on the natural logarithm), such as is available to the observer, is measured as e.g., the difference zeta(s,o(a,b,c..)) - zeta(s,e(a,b,c..)) where o(a,b,c,..) and e(a,b,c,..) are, or relate to, the observed and expected frequencies of occurrence, respectively. For real values of s > 1 the qualitative impact of strongly (positively or negatively) ranked data is preserved despite several numerical approximations. As real s increases, and the output of the information functions converge into three values +1, 0, and -1 nats representing a trinary logic system. For quantitative data, a useful ad hoc method, to report sigma-normalized covariations in an analogous manner to mutual information for significance comparison purposes, is demonstrated. Finally, the potential ability to make use of mutual information in a complex biomedical study, and to include Bayesian prior information derived from statistical, tabular, anecdotal, and expert opinion is briefly illustrated.

Clinical Trials as Topic↗

Dynamic and static approaches to clinical data mining.

In sequential diagnosis, the usefulness of a test can be assessed only in the context of a chosen diagnostic strategy, and depends on the evidence provided by previous test results. Choosing the most useful test at each stage of the evidence-gathering process therefore requires a dynamic approach to data analysis. An implementation of such an approach in an intelligent program for sequential diagnosis based on the evidence-gathering strategies used by doctors is described. On the other hand, a static approach to data analysis is appropriate in the discovery of knowledge required, for example, to explain or justify a diagnosis by identifying the most important findings, both positive and negative, on which the diagnosis is based. An algorithm for the discovery of features which always provide evidence in favour of, or against, a diagnosis selected by the data miner is presented. Dominance relationships among features in the data set are also discovered such that if one feature dominates another, it always provides more evidence in favour of the diagnosis, or less evidence against it.

Abdominal Pain↗

Gene expression databases and data mining.

The DNA microarray technology has arguably caught the attention of the worldwide life science community and is now systematically supporting major discoveries in many fields of study. The majority of the initial technical challenges of conducting experiments are being resolved, only to be replaced with new informatics hurdles, including statistical analysis, data visualization, interpretation, and storage. Two systems of databases, one containing expression data and one containing annotation data are quickly becoming essential knowledge repositories of the research community. This present paper surveys several databases, which are considered "pillars" of research and important nodes in the network. This paper focuses on a generalized workflow scheme typical for microarray experiments using two examples related to cancer research. The workflow is used to reference appropriate databases and tools for each step in the process of array experimentation. Additionally, benefits and drawbacks of current array databases are addressed, and suggestions are made for their improvement.

Breast Neoplasms↗

A data-mining approach to spacer oligonucleotide typing of Mycobacterium tuberculosis.

MOTIVATION: The Direct Repeat (DR) locus of Mycobacterium tuberculosis is a suitable model to study (i) molecular epidemiology and (ii) the evolutionary genetics of tuberculosis. This is achieved by a DNA analysis technique (genotyping), called sp acer oligo nucleotide typing (spoligotyping ). In this paper, we investigated data analysis methods to discover intelligible knowledge rules from spoligotyping, that has not yet been applied on such representation. This processing was achieved by applying the C4.5 induction algorithm and knowledge rules were produced. Finally, a Prototype Selection (PS) procedure was applied to eliminate noisy data. This both simplified decision rules, as well as the number of spacers to be tested to solve classification tasks. In the second part of this paper, the contribution of 25 new additional spacers and the knowledge rules inferred were studied from a machine learning point of view. From a statistical point of view, the correlations between spacers were analyzed and suggested that both negative and positive ones may be related to potential structural constraints within the DR locus that may shape its evolution directly or indirectly. RESULTS: By generating knowledge rules induced from decision trees, it was shown that not only the expert knowledge may be modeled but also improved and simplified to solve automatic classification tasks on unknown patterns. A practical consequence of this study may be a simplification of the spoligotyping technique, resulting in a reduction of the experimental constraints and an increase in the number of samples processed.

Algorithms↗