Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Mining a clinical data warehouse to discover disease-finding associations using co-occurrence statistics.

This paper applies co-occurrence statistics to discover disease-finding associations in a clinical data warehouse. We used two methods, chi2 statistics and the proportion confidence interval (PCI) method, to measure the dependence of pairs of diseases and findings, and then used heuristic cutoff values for association selection. An intrinsic evaluation showed that 94 percent of disease-finding associations obtained by chi2 statistics and 76.8 percent obtained by the PCI method were true associations. The selected associations were used to construct knowledge bases of disease-finding relations (KB-chi2, KB-PCI). An extrinsic evaluation showed that both KB-chi2 and KB-PCI could assist in eliminating clinically non-informative and redundant findings from problem lists generated by our automated problem list summarization system.

Chi-Square Distribution↗

Applying hybrid reasoning to mine for associative features in biological data.

We develop the means to mine for associative features in biological data. The hybrid reasoning schema for deterministic machine learning and its implementation via logic programming is presented. The methodology of mining for correlation between features is illustrated by the prediction tasks for protein secondary structure and phylogenetic profiles. The suggested methodology leads to a clearer approach to hierarchical classification of proteins and a novel way to represent evolutionary relationships. Comparative analysis of Jasmine and other statistical and deterministic systems (including Explanation-Based Learning and Inductive Logic Programming) are outlined. Advantages of using deterministic versus statistical data mining approaches for high-level exploration of correlation structure are analyzed.

Algorithms↗

Bioinformatic analysis of neuropeptide and receptor expression profiles during midgut metamorphosis in Drosophila melanogaster.

Neuropeptides are important messenger molecules in invertebrates, serving as neuromodulators in the nervous system and as regulatory hormones released into the circulation. Understanding the function of neuropeptides will require the integration of genetic, biochemical, physiological and behavioral information. The advent of DNA microarrays and bioinformatic databases provides a wealth of data describing the expression profiles of thousands of genes during biological processes. One such array catalogs the developmental patterns of gene expression during the metamorphic transformation of the Drosophila midgut. We have mined the data from this experiment to explore changes of expression in genes coding for known neuropeptides, peptide hormones, and their receptors during the metamorphosis of the midgut. We found small but significant changes in the expression of the peptides diuretic hormone, FGLa-type allatostatins, myoinhibiting peptide, ecdysis-triggering hormone, drosokinin and the burs subunit of bursicon, as well as the receptors DAR-2, NPFR1, ALCR-2, Lkr and DH-R. Just as advances have been made in understanding the molecular basis of invertebrate neuropeptide action by analysis of genome projects, data mining of gene expression databases can help to integrate molecular, biochemical and physiological knowledge of biological processes.

Animals↗

Enrichment of high-throughput screening data with increasing levels of noise using support vector machines, recursive partitioning, and laplacian-modified naive bayesian classifiers.

High-throughput screening (HTS) plays a pivotal role in lead discovery for the pharmaceutical industry. In tandem, cheminformatics approaches are employed to increase the probability of the identification of novel biologically active compounds by mining the HTS data. HTS data is notoriously noisy, and therefore, the selection of the optimal data mining method is important for the success of such an analysis. Here, we describe a retrospective analysis of four HTS data sets using three mining approaches: Laplacian-modified naive Bayes, recursive partitioning, and support vector machine (SVM) classifiers with increasing stochastic noise in the form of false positives and false negatives. All three of the data mining methods at hand tolerated increasing levels of false positives even when the ratio of misclassified compounds to true active compounds was 5:1 in the training set. False negatives in the ratio of 1:1 were tolerated as well. SVM outperformed the other two methods in capturing active compounds and scaffolds in the top 1%. A Murcko scaffold analysis could explain the differences in enrichments among the four data sets. This study demonstrates that data mining methods can add a true value to the screen even when the data is contaminated with a high level of stochastic noise.

Journal Article↗

Attribute clustering for grouping, selection, and classification of gene expression data.

This paper presents an attribute clustering method which is able to group genes based on their interdependence so as to mine meaningful patterns from the gene expression data. It can be used for gene grouping, selection, and classification. The partitioning of a relational table into attribute subgroups allows a small number of attributes within or across the groups to be selected for analysis. By clustering attributes, the search dimension of a data mining algorithm is reduced. The reduction of search dimension is especially important to data mining in gene expression data because such data typically consist of a huge number of genes (attributes) and a small number of gene expression profiles (tuples). Most data mining algorithms are typically developed and optimized to scale to the number of tuples instead of the number of attributes. The situation becomes even worse when the number of attributes overwhelms the number of tuples, in which case, the likelihood of reporting patterns that are actually irrelevant due to chances becomes rather high. It is for the aforementioned reasons that gene grouping and selection are important preprocessing steps for many data mining algorithms to be effective when applied to gene expression data. This paper defines the problem of attribute clustering and introduces a methodology to solving it. Our proposed method groups interdependent attributes into clusters by optimizing a criterion function derived from an information measure that reflects the interdependence between attributes. By applying our algorithm to gene expression data, meaningful clusters of genes are discovered. The grouping of genes based on attribute interdependence within group helps to capture different aspects of gene association patterns in each group. Significant genes selected from each group then contain useful information for gene expression classification and identification. To evaluate the performance of the proposed approach, we applied it to two well-known gene expression data sets and compared our results with those obtained by other methods. Our experiments show that the proposed method is able to find the meaningful clusters of genes. By selecting a subset of genes which have high multiple-interdependence with others within clusters, significant classification information can be obtained. Thus, a small pool of selected genes can be used to build classifiers with very high classification rate. From the pool, gene expressions of different categories can be identified.

Algorithms↗

Analysis of genomic and proteomic data using advanced literature mining.

High-throughput technologies, such as proteomic screening and DNA micro-arrays, produce vast amounts of data requiring comprehensive analytical methods to decipher the biologically relevant results. One approach would be to manually search the biomedical literature; however, this would be an arduous task. We developed an automated literature-mining tool, termed MedGene, which comprehensively summarizes and estimates the relative strengths of all human gene-disease relationships in Medline. Using MedGene, we analyzed a novel micro-array expression dataset comparing breast cancer and normal breast tissue in the context of existing knowledge. We found no correlation between the strength of the literature association and the magnitude of the difference in expression level when considering changes as high as 5-fold; however, a significant correlation was observed (r = 0.41; p = 0.05) among genes showing an expression difference of 10-fold or more. Interestingly, this only held true for estrogen receptor (ER) positive tumors, not ER negative. MedGene identified a set of relatively understudied, yet highly expressed genes in ER negative tumors worthy of further examination.

Abstracting and Indexing↗

Automatic appropriateness-evaluation and consultation-suggestion of antibiotics usage via mining of previous prescription data in hospital information system.

Inadequate infection subspecialty is main problem in most local hospital of Taiwan. Appropriate prescription of antimicrobial agents is an important issue in daily clinical practice. In order to avoid possible side effect and cost after inappropriate usage of antibiotics, we developed a system for automatic appropriateness evaluation and consultation suggestion of antibiotics usage via mining of previous prescription data in hospital information system.

Anti-Bacterial Agents↗

Atmospheric monitoring at abandoned mercury mine sites in Asturias (NW Spain).

Mercury concentrations are usually significant in historic Hg mining districts all over the world, so the atmospheric environment is potentially affected. In Asturias, northern Spain, past mining operations have left a legacy of ruins and Hg-rich wastes, soils and sediments in abandoned sites. Total Hg concentrations in the ambient air of these abandoned mine sites have been investigated to evaluate the impact of the Hg emissions. This paper presents the synthesis of current knowledge about atmospheric Hg contents in the area of the abandoned Hg mining and smelting works at 'La Peña-El Terronal' and La Soterraña, located in Mieres and Pola de Lena districts, respectively, both within the Caudal River basin. It was found that average atmospheric Hg concentrations are higher than the background level in the area (0.1 microg Nm(-3)), reaching up to 203.7 microg Nm(-3) at 0.2 m above the ground level, close to the old smelting chimney at El Terronal mine site. Data suggest that past Hg mining activities have big influences on the increased Hg concentrations around abandoned sites and that atmospheric transfer is a major pathway for Hg cycling in these environments.

Air Pollutants↗

Interpreter of maladies: redescription mining applied to biomedical data analysis.

Comprehensive, systematic and integrated data-centric statistical approaches to disease modeling can provide powerful frameworks for understanding disease etiology. Here, one such computational framework based on redescription mining in both its incarnations, static and dynamic, is discussed. The static framework provides bioinformatic tools applicable to multifaceted datasets, containing genetic, transcriptomic, proteomic, and clinical data for diseased patients and normal subjects. The dynamic redescription framework provides systems biology tools to model complex sets of regulatory, metabolic and signaling pathways in the initiation and progression of a disease. As an example, the case of chronic fatigue syndrome (CFS) is considered, which has so far remained intractable and unpredictable in its etiology and nosology. The redescription mining approaches can be applied to the Centers for Disease Control and Prevention's Wichita (KS, USA) dataset, integrating transcriptomic, epidemiological and clinical data, and can also be used to study how pathways in the hypothalamic-pituitary-adrenal axis affect CFS patients.

Algorithms↗

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models↗

Database design and implementation for quantitative image analysis research.

Quantitative image analysis (QIA) goes beyond subjective visual assessment to provide computer measurements of the image content, typically following image segmentation to identify anatomical regions of interest (ROIs). Commercially available picture archiving and communication systems focus on storage of image data. They are not well suited to efficient storage and mining of new types of quantitative data. In this paper, we present a system that integrates image segmentation, quantitation, and characterization with database and data mining facilities. The paper includes generic process and data models for QIA in medicine and describes their practical use. The data model is based upon the Digital Imaging and Communications in Medicine (DICOM) data hierarchy, which is augmented with tables to store segmentation results (ROIs) and quantitative data from multiple experiments. Data mining for statistical analysis of the quantitative data is described along with example queries. The database is implemented in PostgreSQL on a UNIX server. Database requirements and capabilities are illustrated through two quantitative imaging experiments related to lung cancer screening and assessment of emphysema lung disease. The system can manage the large amounts of quantitative data necessary for research, development, and deployment of computer-aided diagnosis tools.

Algorithms↗

Challenges of target/compound data integration from disease to chemistry: a case study of dihydrofolate reductase inhibitors.

Despite the improvements in informatics associated with initiatives in the structure-based design and genomics fields, no straight-forward links are available between a given disease class and drug chemistry. This involves effective linking of disease to protein targets, and then mapping these targets to drug chemistry. In practice, protein-ligand structural analyses and high-throughput screening experiments generate the links between targets implicated in disease and chemical leads. Additionally, large volumes of relevant data are also being produced by high-throughput X-ray crystallography and in-silico docking initiatives. Each of these efforts takes a distinctly different approach to how data is managed and mined, resulting in difficulties in sharing data across each area. This review discusses the diverse approaches taken to data management in these areas, and the challenges associated with the construction of a data warehouse that meets all of the needs of each data type. Using the current work available for dihydrofolate reductase inhibitors, we demonstrate the challenges and opportunities associated with data mining from disease to drug chemistry.

Animals↗

Storing, linking, and mining microarray databases using SRS.

BACKGROUND: SRS (Sequence Retrieval System) has proven to be a valuable platform for storing, linking, and querying biological databases. Due to the availability of a broad range of different scientific databases in SRS, it has become a useful platform to incorporate and mine microarray data to facilitate the analyses of biological questions and non-hypothesis driven quests. Here we report various solutions and tools for integrating and mining annotated expression data in SRS. RESULTS: We devised an Auto-Upload Tool by which microarray data can be automatically imported into SRS. The dataset can be linked to other databases and user access can be set. The linkage comprehensiveness of microarray platforms to other platforms and biological databases was examined in a network of scientific databases. The stored microarray data can also be made accessible to external programs for further processing. For example, we built an interface to a program called Venn Mapper, which collects its microarray data from SRS, processes the data by creating Venn diagrams, and saves the data for interpretation. CONCLUSION: SRS is a useful database system to store, link and query various scientific datasets, including microarray data. The user-friendly Auto-Upload Tool makes SRS accessible to biologists for linking and mining user-owned databases.

Computational Biology↗

Future of toxicology--predictive toxicology: An expanded view of "chemical toxicity".

A chemistry approach to predictive toxicology relies on structure-activity relationship (SAR) modeling to predict biological activity from chemical structure. Such approaches have proven capabilities when applied to well-defined toxicity end points or regions of chemical space. These approaches are less well-suited, however, to the challenges of global toxicity prediction, i.e., to predicting the potential toxicity of structurally diverse chemicals across a wide range of end points of regulatory and pharmaceutical concern. New approaches that have the potential to significantly improve capabilities in predictive toxicology are elaborating the "activity" portion of the SAR paradigm. Recent advances in two areas of endeavor are particularly promising. Toxicity data informatics relies on standardized data schema, developed for particular areas of toxicological study, to facilitate data integration and enable relational exploration and mining of data across both historical and new areas of toxicological investigation. Bioassay profiling refers to large-scale high-throughput screening approaches that use chemicals as probes to broadly characterize biological response space, extending the concept of chemical "properties" to the biological activity domain. The effective capture and representation of legacy and new toxicity data into mineable form and the large-scale generation of new bioassay data in relation to chemical toxicity, both employing chemical structure information to inform and integrate diverse biological data, are opening exciting new horizons in predictive toxicology.

Animals↗

Legal issues to address when managing clinical information across Europe: the ECIT case study (www.ECIT.info).

This paper identifies issues which will need to be addressed in pursuing the aims and objectives of the European Classification of Infertility Taskforce (ECIT), namely: to establish classification codes for infertility management; to improve the consistency of infertility information collection by specialist centres, particularly but not exclusively by computerised systems; to use these codes to enable the transfer of infertility information from specialist centres to national infertility data registries; to develop a Grid linking the data held in European infertility data registries; to use Grid processing to mine the data in the European infertility data registries to optimise patient management improving the effectiveness of treatment and reducing the risk.

Commodification↗

An integrative and interactive framework for improving biomedical pattern discovery and visualization.

Recent progress in medical sciences has led to an explosive growth of data. Due to its inherent complexity and diversity, mining such volumes of data to extract relevant knowledge represents an enormous challenge and opportunity. Interactive pattern discovery and visualization systems for biomedical data mining have received relatively little attention. Emphasis has been traditionally placed on automation and supervised classification problems. Based on self-adaptive neural networks and pattern-validation statistical tools, this paper presents a user-friendly platform to support biomedical pattern discovery and visualization. It has been tested on several types of biomedical data, such as dermatology and cardiology data sets. The results indicate that in comparison to traditional techniques, such as Kohonen Maps, this platform may significantly improve the effectiveness and efficiency of pattern discovery and classification tasks, including problems described by several classes. Furthermore, this study shows how the combination of graphical and statistical tools may make these patterns more meaningful.

Algorithms↗

Pattern recognition techniques in microarray data analysis: a survey.

Recent development of technologies (e.g., microarray technology) that are capable of producing massive amounts of genetic data has highlighted the need for new pattern recognition techniques that can mine and discover biologically meaningful knowledge in large data sets. Many researchers have begun an endeavor in this direction to devise such data-mining techniques. As such, there is a need for survey articles that periodically review and summarize the work that has been done in the area. This article presents one such survey. The first portion of the paper is meant to provide the basic biology (mostly for non-biologists) that is required in such a project. This part is only meant to be a starting point for those experts in the technical fields who wish to embark on this new area of bioinformatics. The second portion of the paper is a survey of various data-mining techniques that have been used in mining microarray data for biological knowledge and information (such as sequence information). This survey is not meant to be treated as complete in any form, since the area is currently one of the most active, and the body of research is very large. Furthermore, the applications of the techniques mentioned here are not meant to be taken as the most significant applications of the techniques, but simply as examples among many.

Algorithms↗