Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Data mining for on-line support of general practice.

Statistical relationships among symptoms, diagnoses and treatments can be inferred from large data-bases of health records. We investigate how these "empirical norms" can be utilized to improve the efficiency and quality assurance capability of on-line systems in General Practice medicine. Using a survey-based database of General Practice records, we assess hotlists (case sensitive menus) of diagnoses to speed data entry. We also explore norm violation as an indicator of poor quality in practice or data recording. We find that efficient hotlists of diagnoses can be generated from symptoms. Also, we find that less frequently used hypertension treatments are assessed as of a lower quality than more common ones. The results support the hypothesis that empirical norms have a role to play in improved General Practice systems.

Australia↗

Using the gene ontology for microarray data mining: a comparison of methods and application to age effects in human prefrontal cortex.

One of the challenges in the analysis of gene expression data is placing the results in the context of other data available about genes and their relationships to each other. Here, we approach this problem in the study of gene expression changes associated with age in two areas of the human prefrontal cortex, comparing two computational methods. The first method, "overrepresentation analysis" (ORA), is based on statistically evaluating the fraction of genes in a particular gene ontology class found among the set of genes showing age-related changes in expression. The second method, "functional class scoring" (FCS), examines the statistical distribution of individual gene scores among all genes in the gene ontology class and does not involve an initial gene selection step. We find that FCS yields more consistent results than ORA, and the results of ORA depended strongly on the gene selection threshold. Our findings highlight the utility of functional class scoring for the analysis of complex expression data sets and emphasize the advantage of considering all available genomic information rather than sets of genes that pass a predetermined "threshold of significance."

Adolescent↗

Data mining challenges in the design of telemedicine platforms.

The evolution of telemedicine information systems involves the general processes of acquiring useful knowledge from medical data sets for diagnosis, intelligent efficient patient record transmission and autonomous adaptation of biomedical devices and their related software environments using quality of service attributes. Knowledge engineering concepts and methods allow application to the design of intelligent telemedicine platforms, satisfying the performance requirements and the quality assurance criteria of each specialized telemedicine application.

Artificial Intelligence↗

Data mining by clinicians.

Clinical databases are becoming commonplace in healthcare environments. However, clinicians have been unable to readily explore these information sources, because currently available data retrieval tools require substantial technical skill, as well as knowledge of the underlying database structures. To address this, we have defined a group of "atomic" queries, including both population-based and temporal predicates, to enable the extraction of clinically meaningful information form these databases. DXtractor is an application that incorporates this functionality, and allows clinicians to simply combine these atomic queries. In doing so, arbitrarily complex data retrieval and exploration becomes possible for the non-programming clinician.

Databases as Topic↗

Construction of a generic reaction knowledge base by reaction data mining.

As synthesis by combinatorial chemistry and high throughput screening have become well-established strategies in the drug discovery process, chemists face increased challenges in managing large amounts of data and using these data to design more diverse and focused libraries. As synthesis is an intuitive and empirical process, however, the classical approaches to computer-assisted synthesis planning do not fully satisfy the needs of the synthetic chemist. We describe a novel computational technique for extracting reaction data and building a generic reaction knowledge base (GRKB) to provide chemists with useful and well-organized knowledge. The method consists of three key steps: (1) the automatic recognition of reaction centers, (2) the definition of a hierarchy of reaction patterns, and (3) the organization of the generic reaction knowledge. Significant reaction knowledge has been discovered via mining a subset of the InfoChem Reaction database. A frame system has been constructed to store and retrieve the GRKB. Applications of this GRKB to synthesis planning are illustrated.

Artificial Intelligence↗

Hypoplastic left heart syndrome: knowledge discovery with a data mining approach.

Hypoplastic left heart syndrome (HLHS) affects infants and is uniformly fatal without surgical palliation. Post-surgery mortality rates are highly variable and dependent on postoperative management. A data acquisition system was developed for collection of 73 physiologic, laboratory, and nurse-assessed parameters. The acquisition system was designed for the collection on numerous patients. Data records were created at 30s intervals. An expert-validated wellness score was computed for each data record. To efficiently analyze the data, a new metric for assessment of data utility, the combined classification quality measure, was developed. This measure assesses the impact of a feature on classification accuracy without performing computationally expensive cross-validation. The proposed measure can be also used to derive new features that enhance classification accuracy. The knowledge discovery approach allows for instantaneous prediction of interventions for the patient in an intensive care unit. The discovered knowledge can improve care of complex to manage infants by the development of an intelligent bedside advisory system.

Algorithms↗

A new web-based data mining tool for the identification of candidate genes for human genetic disorders.

To identify the gene underlying a human genetic disorder can be difficult and time-consuming. Typically, positional data delimit a chromosomal region that contains between 20 and 200 genes. The choice then lies between sequencing large numbers of genes, or setting priorities by combining positional data with available expression and phenotype data, contained in different internet databases. This process of examining positional candidates for possible functional clues may be performed in many different ways, depending on the investigator's knowledge and experience. Here, we report on a new tool called the GeneSeeker, which gathers and combines positional data and expression/phenotypic data in an automated way from nine different web-based databases. This results in a quick overview of interesting candidate genes in the region of interest. The GeneSeeker system is built in a modular fashion allowing for easy addition or removal of databases if required. Databases are searched directly through the web, which obviates the need for data warehousing. In order to evaluate the GeneSeeker tool, we analysed syndromes with known genesis. For each of 10 syndromes the GeneSeeker programme generated a shortlist that contained a significantly reduced number of candidate genes from the critical region, yet still contained the causative gene. On average, a list of 163 genes based on position alone was reduced to a more manageable list of 22 genes based on position and expression or phenotype information. We are currently expanding the tool by adding other databases. The GeneSeeker is available via the web-interface (http://www.cmbi.kun.nl/GeneSeeker/).

Computational Biology↗

Automating genomic data mining via a sequence-based matrix format and associative rule set.

There is an enormous amount of information encoded in each genome--enough to create living, responsive and adaptive organisms. Raw sequence data alone is not enough to understand function, mechanisms or interactions. Changes in a single base pair can lead to disease, such as sickle-cell anemia, while some large megabase deletions have no apparent phenotypic effect. Genomic features are varied in their data types and annotation of these features is spread across multiple databases. Herein, we develop a method to automate exploration of genomes by iteratively exploring sequence data for correlations and building upon them. First, to integrate and compare different annotation sources, a sequence matrix (SM) is developed to contain position-dependant information. Second, a classification tree is developed for matrix row types, specifying how each data type is to be treated with respect to other data types for analysis purposes. Third, correlative analyses are developed to analyze features of each matrix row in terms of the other rows, guided by the classification tree as to which analyses are appropriate. A prototype was developed and successful in detecting coinciding genomic features among genes, exons, repetitive elements and CpG islands.

Base Sequence↗

Predicting patient's long-term clinical status after hip arthroplasty using hierarchical decision modelling and data mining.

Construction of a prognostic model is presented for the long-term outcome after femoral neck fracture treatment with implantation of hip endoprosthesis. While the model is induced from the follow-up data, we show that the use of additional expert knowledge is absolutely crucial to obtain good predictive accuracy. A schema is proposed where domain knowledge is encoded as a hierarchical decision model of which only a part is induced from the data while the rest is specified by the expert. Although applied to hip endoprosthesis domain, the proposed schema is general and can be used for the construction of other prognostic models where both follow-up data and human expertise is available.

Aged↗

Microarray data mining with visual programming.

UNLABELLED: Visual programming offers an intuitive means of combining known analysis and visualization methods into powerful applications. The system presented here enables users who are not programmers to manage microarray and genomic data flow and to customize their analyses by combining common data analysis tools to fit their needs. AVAILABILITY: http://www.ailab.si/supp/bi-visprog SUPPLEMENTARY INFORMATION: http://www.ailab.si/supp/bi-visprog.

Chromosome Mapping↗

The dragon on the gold: myths and realities for data mining in biomedicine and biotechnology using digital and molecular libraries.

To develop bioscience and personalized medicine in the post-genomic era, the biggest problem may be how to extract knowledge from the rich libraries of biomedical data. A particular dragon protects the gold therein: the dragon is the "curse of dimensionality" and its formidable fire weapon, which is burning researchers, is the "combinatorial explosion". This arises because many genomic, proteomic, clinical, and lifestyle factors may interact that cannot necessarily be considered on a simple pairwise or additive basis. A suggested theoretical solution--or at least "road map" that ameliorates management of these problems--borrows from several disciplines. It is undertaken also in the hope might also lead to research with broader impact on several unresolved issues in biotechnology: conversely, mathematical understanding of processes involving molecular libraries, such as cDNA libraries and DNA in the living cell itself, may open the opportunities to use biotechnology to construct nanotechnological storage and query systems.

Biomedical Research↗

Chem-tox informatics: data mining using a medicinal chemistry building block approach.

Relating chemical structure to biological activity is not a new endeavor, however, the ability to do this on large datasets is just emerging. To cope with the enormous amounts of data being generated, an assortment of computational methods has been developed in the fields of chemoinformatics and computational toxicology. Many of the molecular descriptors used in these approaches are abstract, theoretical constructs that are difficult to understand and visualize. Having easily recognized chemical features, such as those in several new programs, will allow chemists to use toxicological information (or any biological information) when designing new libraries. These improved chem-tox informatics systems will have an impact on library design, hit and lead optimization, development candidate testing and regulatory review.

Animals↗

Socioeconomic inequality of cancer mortality in the United States: a spatial data mining approach.

BACKGROUND: The objective of this study was to demonstrate the use of an association rule mining approach to discover associations between selected socioeconomic variables and the four most leading causes of cancer mortality in the United States. An association rule mining algorithm was applied to extract associations between the 1988-1992 cancer mortality rates for colorectal, lung, breast, and prostate cancers defined at the Health Service Area level and selected socioeconomic variables from the 1990 United States census. Geographic information system technology was used to integrate these data which were defined at different spatial resolutions, and to visualize and analyze the results from the association rule mining process. RESULTS: Health Service Areas with high rates of low education, high unemployment, and low paying jobs were found to associate with higher rates of cancer mortality. CONCLUSION: Association rule mining with geographic information technology helps reveal the spatial patterns of socioeconomic inequality in cancer mortality in the United States and identify regions that need further attention.

Algorithms↗

GeneCards: a novel functional genomics compendium with automated data mining and query reformulation support.

MOTIVATION: Modern biology is shifting from the 'one gene one postdoc' approach to genomic analyses that include the simultaneous monitoring of thousands of genes. The importance of efficient access to concise and integrated biomedical information to support data analysis and decision making is therefore increasing rapidly, in both academic and industrial research. However, knowledge discovery in the widely scattered resources relevant for biomedical research is often a cumbersome and non-trivial task, one that requires a significant amount of training and effort. RESULTS: To develop a model for a new type of topic-specific overview resource that provides efficient access to distributed information, we designed a database called 'GeneCards'. It is a freely accessible Web resource that offers one hypertext 'card' for each of the more than 7000 human genes that currently have an approved gene symbol published by the HUGO/GDB nomenclature committee. The presented information aims at giving immediate insight into current knowledge about the respective gene, including a focus on its functions in health and disease. It is compiled by Perl scripts that automatically extract relevant information from several databases, including SWISS-PROT, OMIM, Genatlas and GDB. Analyses of the interactions of users with the Web interface of GeneCards triggered development of easy-to-scan displays optimized for human browsing. Also, we developed algorithms that offer 'ready-to-click' query reformulation support, to facilitate information retrieval and exploration. Many of the long-term users turn to GeneCards to quickly access information about the function of very large sets of genes, for example in the realm of large-scale expression studies using 'DNA chip' technology or two-dimensional protein electrophoresis. AVAILABILITY: Freely available at http://bioinformatics.weizmann.ac.il/cards/ CONTACT: cards@bioinformatics.weizmann.ac.il

Algorithms↗