Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Temporal abstraction in intelligent clinical data analysis: a survey.

OBJECTIVE: Intelligent clinical data analysis systems require precise qualitative descriptions of data to enable effective and context sensitive interpretation to take place. Temporal abstraction (TA) provides the means to achieve such descriptions, which can then be used as input to a reasoning engine where they are evaluated against a knowledge base to arrive at possible clinical hypotheses. This paper surveys previous research into the development of intelligent clinical data analysis systems that incorporate TA mechanisms and presents research synergies and trends across the research reviewed, especially those associated with the multi-dimensional nature of real-time patient data streams. The motivation for this survey is case study based research into the development of an intelligent real-time, high-frequency patient monitoring system to provide detection of temporal patterns within multiple patient data streams. RESULTS: The survey was based on factors that are of importance to broaden research into temporal abstraction and on characteristics we believe will assume an increasing level of importance for future clinical IDA systems. These factors were: aspects of the data that is abstracted such as source domain and sample frequency, complexity available within abstracted patterns, dimensionality of the TA and data environment and the knowledge and reasoning underpinning TA processes. CONCLUSION: It is evident from the review that for intelligent clinical data analysis systems to progress into the future where clinical environments are becoming increasingly data-intensive, the ability for managing multi-dimensional aspects of data at high observation and sample frequencies must be provided. Also, the detection of complex patterns within patient data requires higher levels of TA than are presently available. The conflicting matters of computational tractability and temporal reasoning within a real-time environment present a non-trivial problem for investigation in regard to these matters. Finally, to be able to fully exploit the value of learning new knowledge from stored clinical data through data mining and enable its application to data abstraction, the fusion of data mining and TA processes becomes a necessity.

Artificial Intelligence↗

Fuzzy OLAP association rules mining-based modular reinforcement learning approach for multiagent systems.

Multiagent systems and data mining have recently attracted considerable attention in the field of computing. Reinforcement learning is the most commonly used learning process for multiagent systems. However, it still has some drawbacks, including modeling other learning agents present in the domain as part of the state of the environment, and some states are experienced much less than others, or some state-action pairs are never visited during the learning phase. Further, before completing the learning process, an agent cannot exhibit a certain behavior in some states that may be experienced sufficiently. In this study, we propose a novel multiagent learning approach to handle these problems. Our approach is based on utilizing the mining process for modular cooperative learning systems. It incorporates fuzziness and online analytical processing (OLAP) based mining to effectively process the information reported by agents. First, we describe a fuzzy data cube OLAP architecture which facilitates effective storage and processing of the state information reported by agents. This way, the action of the other agent, not even in the visual environment. of the agent under consideration, can simply be predicted by extracting online association rules, a well-known data mining technique, from the constructed data cube. Second, we present a new action selection model, which is also based on association rules mining. Finally, we generalize not sufficiently experienced states, by mining multilevel association rules from the proposed fuzzy data cube. Experimental results obtained on two different versions of a well-known pursuit domain show the robustness and effectiveness of the proposed fuzzy OLAP mining based modular learning approach. Finally, we tested the scalability of the approach presented in this paper and compared it with our previous work on modular-fuzzy Q-learning and ordinary Q-learning.

Algorithms↗

[Automatic diagnosis of malignant degree of brain glioma based on Bayesian network].

Bayesian network connects graph theory with statistics, being an important research direction in data mining. Compared with other approaches used for data mining, Bayesian network can combine prior knowledge with observed data. Besides that, it can handle incomplete data sets. This paper applies Bayesian network to predict the malignant degree of brain glioma. Totally 280 cases are collected, and some of them contain missing values. Preprocessing is taken to make them applicable to the algorithms. Unlike MLP network, both Bayesian network and decision tree use attribute-value pairs to represent diagnostic knowledge derived from treated cases. These could improve both the understandability and applicability of their results. Results of all these algorithms can achieve accuracy rate over 80%, which satisfies the requirement of neuroradiologists.

Bayes Theorem↗

Multivariate image analysis in biomedicine.

In recent years, multivariate imaging techniques are developed and applied in biomedical research in an increasing degree. In research projects and in clinical studies as well m-dimensional multivariate images (MVI) are recorded and stored to databases for a subsequent analysis. The complexity of the m-dimensional data and the growing number of high throughput applications call for new strategies for the application of image processing and data mining to support the direct interactive analysis by human experts. This article provides an overview of proposed approaches for MVI analysis in biomedicine. After summarizing the biomedical MVI techniques the two level framework for MVI analysis is illustrated. Following this framework, the state-of-the-art solutions from the fields of image processing and data mining are reviewed and discussed. Motivations for MVI data mining in biology and medicine are characterized, followed by an overview of graphical and auditory approaches for interactive data exploration. The paper concludes with summarizing open problems in MVI analysis and remarks upon the future development of biomedical MVI analysis.

Algorithms↗

Mining protein data from two-dimensional gels: tools for systematic post-planned analyses.

There is a considerable need to develop comprehensive, systematic mechanisms to analyze the vast number of proteins that orchestrate various cellular functions and to identify proteins associated with disease or that are affected by pharmacological agents. Two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) continues to be relied upon to analyze protein constituents of cells and tissues. We have developed a Laboratory Information Processing System (LIPS) as a computer-based tool for capturing quantitative and qualitative changes in thousands of proteins detected in 2-D gels of various types. Protein databases have been developed to serve as a repository for data processing of the basic and derived data and of findings derived from different studies. There have been remarkable advances both in database technology as well as in the computer hardware that have benefited our effort at mining protein data from 2-D gels. We here review our current efforts aimed at improving the performance and features of our 2-D related protein databases, with particular emphasis on the tools we utilize for database mining via a systematic analysis of information known as post-planned analysis.

Acrylic Resins↗

Special report: mining admissions data. Benchmarking hospital admission rates by ages.

In this second part of our two-part series on mining hospital admissions data, Ronald Lagoe, executive director of the Hospital Executive Council, a four-hospital system in Syracuse, NY shows how admissions data readily available from various state agencies can be analyzed for patterns by age groups. As he points out below, this information will be particularly valuable as Medicare managed care takes over. For a discussion of data sources, methodology and analysis, see part one, beginning on p. 81 of the June Healthcare Benchmarks.

Adolescent↗

Datamining methodology for LC-MALDI-MS based peptide profiling.

This report will provide a brief overview of the application of data mining in proteomic peptide profiling used for medical biomarker research. Mass spectrometry based profiling of peptides and proteins is frequently used to distinguish disease from non-disease groups and to monitor and predict drug effects. It has the promising potential to enter clinical laboratories as a general purpose diagnostic tool. Data mining methodologies support biomedical science to manage the vast data sets obtained from these instrumentations. Here we will review the typical workflow of peptide profiling, together with typical data mining methodology. Mass spectrometric experiments in peptidomics raise numerous questions in the fields of signal processing, statistics, experimental design and discriminant analysis.

Animals↗

Web multimedia information retrieval using improved Bayesian algorithm.

The main thrust of this paper is application of a novel data mining approach on the log of user's feedback to improve web multimedia information retrieval performance. A user space model was constructed based on data mining, and then integrated into the original information space model to improve the accuracy of the new information space model. It can remove clutter and irrelevant text information and help to eliminate mismatch between the page author's expression and the user's understanding and expectation. User space model was also utilized to discover the relationship between high-level and low-level features for assigning weight. The authors proposed improved Bayesian algorithm for data mining. Experiment proved that the authors' proposed algorithm was efficient.

Algorithms↗

Highly scalable and robust rule learner: performance evaluation and comparison.

Business intelligence and bioinformatics applications increasingly require the mining of datasets consisting of millions of data points, or crafting real-time enterprise-level decision support systems for large corporations and drug companies. In all cases, there needs to be an underlying data mining system, and this mining system must be highly scalable. To this end, we describe a new rule learner called DataSqueezer. The learner belongs to the family of inductive supervised rule extraction algorithms. DataSqueezer is a simple, greedy, rule builder that generates a set of production rules from labeled input data. In spite of its relative simplicity, DataSqueezer is a very effective learner. The rules generated by the algorithm are compact, comprehensible, and have accuracy comparable to rules generated by other state-of-the-art rule extraction algorithms. The main advantages of DataSqueezer are very high efficiency, and missing data resistance. DataSqueezer exhibits log-linear asymptotic complexity with the number of training examples, and it is faster than other state-of-the-art rule learners. The learner is also robust to large quantities of missing data, as verified by extensive experimental comparison with the other learners. DataSqueezer is thus well suited to modern data mining and business intelligence tasks, which commonly involve huge datasets with a large fraction of missing data.

Algorithms↗

[Knowledge discovery in database and its application in clinical diagnosis].

Nowadays the tremendous amount of data has far exceeded our human ability for comprehension, and this has been particularly true for the medical database. However, traditional statistical techniques are no longer adequate for analyzing this vast collection of data. Knowledge discovery in database and data mining play an important role in analyzing data and uncovering important data patterns. This paper briefly presents the concepts of knowledge discovery in database and data mining, then describes the rough set theory, and gives some examples based on rough set.

Artificial Intelligence↗

Characteristic substructures and properties in chemical carcinogens studied by the cascade model.

MOTIVATION: Chemical carcinogenicity is an important subject in health and environmental sciences, and a reliable method is expected to identify characteristic factors for carcinogenicity. The predictive toxicology challenge (PTC) 2000-2001 has provided the opportunity for various data mining methods to evaluate their performance. The cascade model, a data mining method developed by the author, has the capability to mine for local correlations in data sets with a large number of attributes. The current paper explores the effectiveness of the method on the problem of chemical carcinogenicity. RESULTS: Rodent carcinogenicity of 417 compounds examined by the National Toxicology Program (NTP) was used as the training set. The analysis by the cascade model, for example, could obtain a rule 'Highly flexible molecules are carcinogenic, if they have no hydrogen bond acceptors in halogenated alkanes and alkenes'. Resulting rules are applied to predict the activity of 185 compounds examined by the FDA. The ROC analysis performed by the PTC organizers has shown that the current method has excellent predictive power for the female rat data. AVAILABILITY: The binary program of DISCAS 2.1 and samples of input data sets on Windows PC are available at http://www.clab.kwansei.ac.jp/mining/discas/discas.html upon request from the author. SUPPLEMENTARY INFORMATION: Summary of prediction results and cross validations is accessible via http://www.clab.kwansei.ac.jp/~okada/BIJ/BIJsupple.htm. Used rules and the prediction results for each molecule are also provided.

Algorithms↗

Contrast media and nephropathy: findings from systematic analysis and Food and Drug Administration reports of adverse effects.

CONTEXT: Recent studies suggest differences in the incidence of contrast-induced nephropathy (CIN) among contrast media (CM). OBJECTIVE: To determine whether there are significant differences among low-osmolality CM (LOCM) in the incidence of contrast-induced nephropathy (CIN), we reviewed published studies of CIN in renally impaired patients and conducted statistical data mining using databases of adverse events maintained by the U.S. Food and Drug Administration (FDA). DATA SOURCES: A systematic literature search was performed for prospective, controlled, English language studies published in peer-reviewed journals that reported CIN rates in renally impaired patients after a specific LOCM. Databases searched were EMBASE, MEDLINE, Biosis Previews, Derwent Drug File, Pascal, and SciScearch Cited Ref Sci. For the FDA analysis, we used the SRS and AERS databases. DATA SELECTION: Twenty-two studies reporting data in 3112 patients with renal impairment met the inclusion criteria. Most studies reported on the use of a pharmacologic intervention to prevent CIN. From the FDA databases, we evaluated 18 adverse event terms associated with renal injury or dysfunction after CM use. DATA EXTRACTION: Data from 22 studies were entered into a database. A meta-regression analysis using a mixed effect model was performed. CM effect was adjusted by the following covariates: baseline patient characteristics (mean age, gender distribution) and risk factors (prevalence of diabetes mellitus, degree of renal impairment, CM volume), and the use of prophylactic drug treatments. Multiple disproportionality analyses (adjusted odds ratio, adjusted empirical Bayesian estimate, or Bayesian logistic regression) were performed on the FDA databases to estimate associations between 4 CM and 18 AE terms related to CIN. DATA SYNTHESIS: Systematic analysis of clinical trials suggest the highest incidence of CIN occurs in patients receiving iohexol and the lowest incidence in patients receiving iopamidol, even when corrected for other CIN risk factors. Statistical data mining of FDA data also showed the highest association of CIN for iohexol and the lowest for iopamidol. CONCLUSIONS: The risk of CIN was higher in patients receiving iohexol compared with patients receiving iopamidol. No significant differences were found comparing iohexol to other LOCMs, including iodixanol.

Bayes Theorem↗

Neural networks in astronomy.

In the last decade, the use of neural networks (NN) and of other soft computing methods has begun to spread also in the astronomical community which, due to the required accuracy of the measurements, is usually reluctant to use automatic tools to perform even the most common tasks of data reduction and data mining. The federation of heterogeneous large astronomical databases which is foreseen in the framework of the astrophysical virtual observatory and national virtual observatory projects, is, however, posing unprecedented data mining and visualization problems which will find a rather natural and user friendly answer in artificial intelligence tools based on NNs, fuzzy sets or genetic algorithms. This review is aimed to both astronomers (who often have little knowledge of the methodological background) and computer scientists (who often know little about potentially interesting applications), and therefore will be structured as follows: after giving a short introduction to the subject, we shall summarize the methodological background and focus our attention on some of the most interesting fields of application, namely: object extraction and classification, time series analysis, noise identification, and data mining. Most of the original work described in the paper has been performed in the framework of the AstroNeural collaboration (Napoli-Salerno).

Astronomy↗

Genome-wide isolation of resistance gene analogs in maize (Zea mays L.).

Conserved domains or motifs shared by most known resistance (R) genes have been extensively exploited to identify unknown R-gene analogs (RGAs). In an attempt to isolate all potential RGAs from the maize genome, we adopted the following three methods: modified amplified fragment length polymorphism (AFLP), modified rapid amplification of cDNA ends (RACE), and data mining. The first two methods involved PCR-based isolations of RGAs with degenerate primers designed based on the conserved NBS domain; while the third method involved mining of RGAs from the maize EST database using full-length R-gene sequences. A total of 23 and 12 RGAs were obtained from the modified AFLP and RACE methods, respectively; while, as many as 109 unigenes and 77 singletons with high homology to known R-genes were recovered via data-mining. Moreover, R-gene-like ESTs (or RGAs) identified from the data-mining method could cover all RACE-derived RGAs and nearly half AFLP-derived RGAs. Totally, the three methods resulted in 199 non-redundant RGAs. Of them, at least 186 were derived from putative expressed R-genes. RGA-tagged markers were developed for 55 unique RGAs, including 16 STS and 39 CAPS markers.

Base Sequence↗

Applying Association Rule Discovery Algorithm to Multipoint Linkage Analysis.

Knowledge discovery in large databases (KDD) is being performed in several application domains, for example, the analysis of sales data, and is expected to be applied to other domains. We propose a KDD approach to multipoint linkage analysis, which is a way of ordering loci on a chromosome. Strict multipoint linkage analysis based on maximum likelihood estimation is a computationally tough problem. So far various kinds of approximate methods have been implemented. Our method based on the discovery of association between genetic recombinations is so different from others that it is useful to recheck the result of them. In this paper, we describe how to apply the framework of association rule discovery to linkage analysis, and also discuss that filtering input data and interpretation of discovered rules after data mining are practically important as well as data mining process itself.

Journal Article↗

Predictive model for survival at the conclusion of a damage control laparotomy.

BACKGROUND: We employed modern statistical and data mining methods to model survival based on preoperative and intraoperative parameters for patients undergoing damage control surgery. METHODS: One hundred seventy-four parameters were collected from 68 damage control patients in prehospital, emergency center, operating room, and intensive care unit (ICU) settings. Data were analyzed with logistic regression and data mining. Outcomes were survival and death after the initial operation. RESULTS: Overall mortality was 66.2%. Logistic regression identified pH at initial ICU admission (odds ratio: 4.4) and worst partial thromboplastin time from hospital admission to ICU admission (odds ratio: 9.4) as significant. Data mining selected the same factors, and generated a simple algorithm for patient classification. Model accuracy was 83%. CONCLUSION: Inability to correct pH at the conclusion of initial damage-control laparotomy and the worst PTT can be predictive of death. These factors may be useful to identify patients with a high risk of mortality.

Critical Illness↗

The bioinformatics resource for oral pathogens.

Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.

Bacteria↗

Mining SAGE data allows large-scale, sensitive screening of antisense transcript expression.

As a growing number of complementary transcripts, susceptible to exert various regulatory functions, are being found in eukaryotes, high throughput analytical methods are needed to investigate their expression in multiple biological samples. Serial Analysis of Gene Expression (SAGE), based on the enumeration of directionally reliable short cDNA sequences (tags), is capable of revealing antisense transcripts. We initially detected them by observing tags that mapped on to the reverse complement of known mRNAs. The presence of such tags in individual SAGE libraries suggested that SAGE datasets contain latent information on antisense transcripts. We raised a collection of virtual tags for mining these data. Tag pairs were assembled by searching for complementarities between 24-nt long sequences centered on the potential SAGE-anchoring sites of well-annotated human expressed sequences. An analysis of their presence in a large collection of published SAGE libraries revealed transcripts expressed at high levels from both strands of two adjacent, oppositely oriented, transcription units. In other cases, the respective transcripts of such cis-oriented genes displayed a mutually exclusive expression pattern or were co-expressed in a small number of libraries. Other tag pairs revealed overlapping transcripts of trans-encoded unique genes. Finally, we isolated a group of tags shared by multiple transcripts. Most of them mapped on to retroelements, essentially represented in humans by Alu sequences inserted in opposite orientations in the 3'UTR of otherwise different mRNAs. Registering these tags in separate files makes possible computational searches focused on unique sense-antisense pairs. The method developed in the present work shows that SAGE datasets constitute a major resource of rapidly investigating with high sensitivity the expression of antisense transcripts, so that a single tag may be detected in one library when screening a large number of biological samples.

Computational Biology↗