Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Surgeon and type of anesthesia predict variability in surgical procedure times.

BACKGROUND: Variability in surgical procedure times increases the cost of healthcare delivery by increasing both the underutilization and overutilization of expensive surgical resources. To reduce variability in surgical procedure times, we must identify and study its sources. METHODS: Our data set consisted of all surgeries performed over a 7-yr period at a large teaching hospital, resulting in 46,322 surgical cases. To study factors associated with variability in surgical procedure times, data mining techniques were used to segment and focus the data so that the analyses would be both technically and intellectually feasible. The data were subdivided into 40 representative segments of manageable size and variability based on headers adopted from the common procedural terminology classification. Each data segment was then analyzed using a main-effects linear model to identify and quantify specific sources of variability in surgical procedure times. RESULTS: The single most important source of variability in surgical procedure times was surgeon effect. Type of anesthesia, age, gender, and American Society of Anesthesiologists risk class were additional sources of variability. Intrinsic case-specific variability, unexplained by any of the preceding factors, was found to be highest for shorter surgeries relative to longer procedures. Variability in procedure times among surgeons was a multiplicative function (proportionate to time) of surgical time and total procedure time, such that as procedure times increased, variability in surgeons' surgical time increased proportionately. CONCLUSIONS: Surgeon-specific variability should be considered when building scheduling heuristics for longer surgeries. Results concerning variability in surgical procedure times due to factors such as type of anesthesia, age, gender, and American Society of Anesthesiologists risk class may be extrapolated to scheduling in other institutions, although specifics on individual surgeons may not. This research identifies factors associated with variability in surgical procedure times, knowledge of which may ultimately be used to improve surgical scheduling and operating room utilization.

Adolescent↗

Virtual instrumentation and real-time executive dashboards. Solutions for health care systems.

Successful organizations have the ability to measure and act on key indicators and events in real time. By leveraging the power of virtual instrumentation and open architecture standards, multidimensional executive dashboards can empower health care organizations to make better and faster data-driven decisions. This article will highlight how user-defined virtual instruments and dashboards can connect to hospital information systems (e.g., admissions/discharge/transfer systems, patient monitoring networks) and use statistical process control to "visualize" information and make timely, data-driven decisions. The case studies described will illustrate enterprisewide solutions for: bed management and census control, operational management, data mining and business intelligence applications, and clinical applications (physiological data acquisition and wound measurement and analysis).

Bed Occupancy↗

Adverse cardiac effects associated with clozapine.

OBJECTIVE: To review the published literature on serious adverse cardiac events associated with the atypical antipsychotic agent, clozapine, and to make recommendations for cardiac assessment of candidates for clozapine treatment and for monitoring of cardiac status after treatment is initiated. DATA SOURCES: We searched the PubMed and MEDLINE databases for articles published from 1970 to 2004 that contain the keywords "clozapine and myocarditis," "clozapine and cardiomyopathy," "clozapine and cardiotoxicity," "clozapine and sudden death" or "clozapine and mortality." We also manually searched the bibliographies of these articles for related sources. STUDY SELECTION: We reviewed the 30 case reports, case series, laboratory and clinical trials, data mining studies, and previous reviews identified by this search. DATA SYNTHESIS: Recent evidence suggests that clozapine is associated with a low (0.015% to 0.188%) risk of potentially fatal myocarditis or cardiomyopathy. The drug is not known to be independently associated with pathologic prolongation of the QTc interval, but it may contribute to pathologic QTc prolongation in patients with other risk factors for this condition. CONCLUSIONS: The low risk of a serious adverse cardiac event should be outweighed by a reduction in suicide risk for most patients taking clozapine. We provide recommendations for assessing and monitoring cardiac status in patients prior to and after initiation of treatment with clozapine.

Antipsychotic Agents↗

PlasmoDB: exploring genomics and post-genomics data of the malaria parasite, Plasmodium falciparum.

The recent completion of the genome sequence of Plasmodium falciparum 3D7 provides the foundation for genome-wide analysis of the parasite. In addition to DNA and gene sequence data, postgenomic methods including microarray-based transcript profiling and high-throughput proteomics are now accessible to Plasmodium researchers. The Plasmodium Genome database ( ) was developed to provide rapid and convenient access to the terabytes of genomic-scale data now being generated around the world. All data are available in a relational framework, permitting convenient downloading, browsing, and analysis. Combinatorial use of data analysis tools enables powerful data mining queries, such as combining gene and protein expression data to monitor changes through various life-cycle stages. Functional predictions can be used to explore potential targets for antimalarial drug development. This report outlines the use of PlasmoDB to examine redox-active functions in Plasmodium.

Animals↗

Prediction of protein structural class with Rough Sets.

BACKGROUND: A new method for the prediction of protein structural classes is constructed based on Rough Sets algorithm, which is a rule-based data mining method. Amino acid compositions and 8 physicochemical properties data are used as conditional attributes for the construction of decision system. After reducing the decision system, decision rules are generated, which can be used to classify new objects. RESULTS: In this study, self-consistency and jackknife tests on the datasets constructed by G.P. Zhou (Journal of Protein Chemistry, 1998, 17: 729-738) are used to verify the performance of this method, and are compared with some of prior works. The results showed that the rough sets approach is very promising and may play a complementary role to the existing powerful approaches, such as the component-coupled, neural network, SVM, and LogitBoost approaches. CONCLUSION: The results with high success rates indicate that the rough sets approach as proposed in this paper might hold a high potential to become a useful tool in bioinformatics.

Algorithms↗

Whole genome association mapping by incompatibilities and local perfect phylogenies.

BACKGROUND: With current technology, vast amounts of data can be cheaply and efficiently produced in association studies, and to prevent data analysis to become the bottleneck of studies, fast and efficient analysis methods that scale to such data set sizes must be developed. RESULTS: We present a fast method for accurate localisation of disease causing variants in high density case-control association mapping experiments with large numbers of cases and controls. The method searches for significant clustering of case chromosomes in the "perfect" phylogenetic tree defined by the largest region around each marker that is compatible with a single phylogenetic tree. This perfect phylogenetic tree is treated as a decision tree for determining disease status, and scored by its accuracy as a decision tree. The rationale for this is that the perfect phylogeny near a disease affecting mutation should provide more information about the affected/unaffected classification than random trees. If regions of compatibility contain few markers, due to e.g. large marker spacing, the algorithm can allow the inclusion of incompatibility markers in order to enlarge the regions prior to estimating their phylogeny. Haplotype data and phased genotype data can be analysed. The power and efficiency of the method is investigated on 1) simulated genotype data under different models of disease determination 2) artificial data sets created from the HapMap ressource, and 3) data sets used for testing of other methods in order to compare with these. Our method has the same accuracy as single marker association (SMA) in the simplest case of a single disease causing mutation and a constant recombination rate. However, when it comes to more complex scenarios of mutation heterogeneity and more complex haplotype structure such as found in the HapMap data our method outperforms SMA as well as other fast, data mining approaches such as HapMiner and Haplotype Pattern Mining (HPM) despite being significantly faster. For unphased genotype data, an initial step of estimating the phase only slightly decreases the power of the method. The method was also found to accurately localise the known susceptibility variants in an empirical data set--the DeltaF508 mutation for cystic fibrosis--where the susceptibility variant is already known--and to find significant signals for association between the CYP2D6 gene and poor drug metabolism, although for this dataset the highest association score is about 60 kb from the CYP2D6 gene. CONCLUSION: Our method has been implemented in the Blossoc (BLOck aSSOCiation) software. Using Blossoc, genome wide chip-based surveys of 3 million SNPs in 1000 cases and 1000 controls can be analysed in less than two CPU hours.

Chromosome Mapping↗

Integrated analysis of gene expression by Association Rules Discovery.

BACKGROUND: Microarray technology is generating huge amounts of data about the expression level of thousands of genes, or even whole genomes, across different experimental conditions. To extract biological knowledge, and to fully understand such datasets, it is essential to include external biological information about genes and gene products to the analysis of expression data. However, most of the current approaches to analyze microarray datasets are mainly focused on the analysis of experimental data, and external biological information is incorporated as a posterior process. RESULTS: In this study we present a method for the integrative analysis of microarray data based on the Association Rules Discovery data mining technique. The approach integrates gene annotations and expression data to discover intrinsic associations among both data sources based on co-occurrence patterns. We applied the proposed methodology to the analysis of gene expression datasets in which genes were annotated with metabolic pathways, transcriptional regulators and Gene Ontology categories. Automatically extracted associations revealed significant relationships among these gene attributes and expression patterns, where many of them are clearly supported by recently reported work. CONCLUSION: The integration of external biological information and gene expression data can provide insights about the biological processes associated to gene expression programs. In this paper we show that the proposed methodology is able to integrate multiple gene annotations and expression data in the same analytic framework and extract meaningful associations among heterogeneous sources of data. An implementation of the method is included in the Engene software package.

Algorithms↗

The Human Genome Project and the role of genetics in health care.

The Human Genome Project, the mapping of our 100,000 genes and the sequencing of all of our DNA, will have major impact on biomedical research and the therapeutic and preventive health care. The tracing of genetic diseases to their molecular causes is rapidly expanding diagnostic and preventive options, while the increased insights into molecular pathways open tremendous perspectives for pharmacological and genetic therapies. The design of animal model systems for the functional study of disease and development of bioinformatics and biostatistics to improve our pattern recognition abilities are greatly accelerating progress. However, the optimal value from the current explosion of 'data mining' possibilities will only be gained when the basic data are made and kept publicly accessible, at the same time preventing the jeopardisation of the protection of intellectual property, arising from downstream inventions. This is one of the goals of HUGO, the international Human Genome Organisation, established 9 years ago to assist coordinating data acquisition and exchange and societal implementation of the genome project. Additional points of major importance in this historic endeavour are the safeguarding of a worldwide balance in the contribution and benefits to countries and population, the prevention of stigmatisation and discrimination of individuals and groups and the maintenance of respect for the priceless diversity of our world's cultures and traditions.

Delivery of Health Care↗

Classification algorithms applied to narrative reports.

Narrative text reports represent a significant source of clinical data. However, the information stored in these reports is inaccessible to many automated decision support systems. Data mining techniques can assist in extracting information from narrative data. Multiple classification methods, such as rule generation, decision trees, Bayesian classifiers, and information retrieval were used to classify a set of 200 chest X-ray reports according to 6 clinical conditions indicated. A general-purpose natural language processor was used to convert the narrative text into a coded form that could be used by the classification algorithms. Significant differences in performance were found between algorithms. The best performing algorithm applied to the processor output was significantly better than information retrieval applied to raw text. Predictor variables from the coded processor output were limited to avoid overfitting. Methods that limited by domain knowledge performed significantly better than those that limited by conditional probabilities of the variables in the training set. Algorithms were also shown to be dependent on training set size.

Algorithms↗

MET system: a new approach to m-health in emergency triage.

The MET (Mobile Emergency Triage) system is an m-health application that supports emergency triage of various types of acute pain at the point of care. The system is designed for use in the Emergency Department (ED) of a hospital and to aid physicians in disposition decisions. Given patient's condition, MET recommends a triage by consulting decision rules stored in the system's knowledge base. The rules have been created using a data mining method (based on rough set methodology) applied to data collected during a retrospective chart study and verified by the clinicians. MET is designed following the extended client-server architecture, suited for weak-connectivity conditions, where stable connection between clients and a server cannot be provided. The MET server interacts with the hospital's patient information system in order to retrieve information about patients admitted to the ED. It also stores current patients' demographic and clinical data to be exchanged with mobile clients. The MET mobile client, running on a Personal Digital Assistant (PDA), is used for collecting clinical data and supporting triage decisions. The support function runs solely on the client side, thus it can be invoked anytime and anywhere, even if there is no communication link with the server (e.g., there is no wireless network available in the ED). Due to implementation on PDAs and working in weak-connectivity conditions, the MET system is very well suited for use in the ED and fits seamlessly into the regular clinical workflow without introducing any hindrances or disruptions that are often reported when using stationary (i.e., working on desktop computers) clinical systems. The system facilitates patient-centered service and timely, high quality patient management. It provides recommendations using a limited amount of clinical data, normally available at the point of care. Furthermore, it provides a possibility for the structured evaluation of this data by an attending physician.

Acute Disease↗

Mining time-dependent patient outcomes from hospital patient records.

We describe REMIND, a data mining framework that accurately infers missing clinical information by reasoning over the entire patient record. Hospitals collect computerized patient records (CPR's) in structured (database tables) and unstructured (free text) formats. Structured clinical data in the CPR's is often poorly recorded, and information may be missing about key outcomes and processes. For instance, for a population of 344 colon cancer patients, important clinical outcomes, such as disease state and its evolution, are stored only as unstructured data (doctors' dictations) in the CPR. Raw evidence (extracted directly from the CPR) is not a good predictor of disease state. Yet by combining this evidence in a principled fashion (using methods from uncertain and temporal reasoning), REMIND accurately infers disease state sequences for recurrence, a complex time-varying outcome, for these patients. These outcomes can now be added back into the CPR in structured form.

Colonic Neoplasms↗

Mining association rules with improved semantics in medical databases.

The discovery of new knowledge by mining medical databases is crucial in order to make an effective use of stored data, enhancing patient management tasks. One of the main objectives of data mining methods is to provide a clear and understandable description of patterns held in data. We introduce a new approach to find association rules among quantitative values in relational databases. The semantics of such rules are improved by introducing imprecise terms in both the antecedent and the consequent, as these terms are the most commonly used in human conversation and reasoning. The terms are modeled by means of fuzzy sets defined in the appropriate domains. However, the mining task is performed on the precise data. These "fuzzy association rules" are more informative than rules relating precise values. We also introduce a new measure of accuracy, based on Shortliffe and Buchanan's certainty factors [Shortliffe E, Buchanan B. Math Biosci 1975;23:351-79]. Also, the semantics of the usual measure of usefulness of an association rule, called support are discussed and some new criteria are introduced. Our new measures have been shown to be more understandable and appropriate than ordinary ones. Several experiments on large medical databases show that our new approach can provide useful knowledge with better semantics in this field.

Databases, Factual↗

PubMatrix: a tool for multiplex literature mining.

BACKGROUND: Molecular experiments using multiplex strategies such as cDNA microarrays or proteomic approaches generate large datasets requiring biological interpretation. Text based data mining tools have recently been developed to query large biological datasets of this type of data. PubMatrix is a web-based tool that allows simple text based mining of the NCBI literature search service PubMed using any two lists of keywords terms, resulting in a frequency matrix of term co-occurrence. RESULTS: For example, a simple term selection procedure allows automatic pair-wise comparisons of approximately 1-100 search terms versus approximately 1-10 modifier terms, resulting in up to 1,000 pair wise comparisons. The matrix table of pair-wise comparisons can then be surveyed, queried individually, and archived. Lists of keywords can include any terms currently capable of being searched in PubMed. In the context of cDNA microarray studies, this may be used for the annotation of gene lists from clusters of genes that are expressed coordinately. An associated PubMatrix public archive provides previous searches using common useful lists of keyword terms. CONCLUSIONS: In this way, lists of terms, such as gene names, or functional assignments can be assigned genetic, biological, or clinical relevance in a rapid flexible systematic fashion. http://pubmatrix.grc.nia.nih.gov/

Cell Line, Tumor↗

Using standard positions and image fusion to create proteome maps from collections of two-dimensional gel electrophoresis images.

Databases for two-dimensional protein gels pose new challenges in extracting meaningful information from large numbers of experiments. In order to create expression profiles, positions of corresponding protein spots across all gel images have to be established. In larger gel sets errors may accumulate rapidly during this spot matching process, effectively limiting the number of samples available for data mining. Here we present a novel approach for organizing spot data based on the concept of a standard position for a protein species. Standard positions are meaningful average positions that are determined using all occurrences of a protein species. They can be extended to spots that are not annotated via interpolation. The standard position of a spot can serve as a unifying index across all gels in a database, thus allowing creation and analysis of expression profiles that span the whole collection. The standard position gives a much more accurate estimation of a spot's position on a gel than can be obtained using theoretical isoelectric point and molecular weight. Positional indexing is a complement to a priori identifications (e.g. by mass spectrometry or Edman degradation). Moreover it can be used in advance to select spots that are worth identifying because they show relevant expression profiles. Furthermore, we show how to combine all spots that occur on any of the gels into one synthetic but nevertheless realistic-looking image. This composite image is produced such that all spots have their standard positions. It can serve as a proteome reference map for an organism. As an application, we have computed a reference map from 23 gel images of Bacillus subtilis, using an enhanced prerelease version of the gel analysis software Delta2D (DECODON, Greifswald, Germany).

Bacterial Proteins↗

In silico approaches to microarray-based disease classification and gene function discovery.

The automated analysis of transcriptional profiling data promises a wealth of information that may be used to develop a more complete understanding of gene function and interactions. Moreover, it may improve the effectiveness of complex diagnostic tasks. This article discusses important data mining and management techniques to analyse genome-wide expression data. It reviews some of the major discovery goals, methods and applications in a number of biomedical domains. Finally, this paper highlights key problems that need to be approached by a new generation of computational solutions.

Algorithms↗

A comparison of oligonucleotide and cDNA-based microarray systems.

Large-scale public data mining will become more common as public release of microarray data sets becomes a corequisite for publication. Therefore, there is an urgent need to clarify whether data from different microarray platforms are comparable. To assess the compatibility of microarray data, results were compared from the two main types of high-throughput microarray expression technologies, namely, an oligonucleotide-based and a cDNA-based platform, using RNA obtained from complex tissue (human colonic mucosa) of five individuals. From 715 sequence-verified genes represented on both platforms, 64% of the genes matched in "present" or "absent" calls made by both platforms. Calls were influenced by spurious signals caused by Alu repeats in cDNA clones, clone annotation errors, or matched probes that were designed to different regions of the gene; however, these factors could not completely account for the level of call discordance observed. Expression levels in sequence-verified, platform-overlapping genes were not related, as demonstrated by weakly positive rank order correlation. This study demonstrates that there is only moderate overlap in the results from the two array systems. This fact should be carefully considered when performing large-scale analyses on data originating from different microarray platforms.

Aged↗

Can the US minimum data set be used for predicting admissions to acute care facilities?

This paper is intended to give an overview of Knowledge Discovery in Large Datasets (KDD) and data mining applications in healthcare particularly as related to the Minimum Data Set, a resident assessment tool which is used in US long-term care facilities. The US Health Care Finance Administration, which mandates the use of this tool, has accumulated massive warehouses of MDS data. The pressure in healthcare to increase efficiency and effectiveness while improving patient outcomes requires that we find new ways to harness these vast resources. The intent of this preliminary study design paper is to discuss the development of an approach which utilizes the MDS, in conjunction with KDD and classification algorithms, in an attempt to predict admission from a long-term care facility to an acute care facility. The use of acute care services by long term care residents is a negative outcome, potentially avoidable, and expensive. The value of the MDS warehouse can be realized by the use of the stored data in ways that can improve patient outcomes and avoid the use of expensive acute care services. This study, when completed, will test whether the MDS warehouse can be used to describe patient outcomes and possibly be of predictive value.

Algorithms↗

[Differences in ethnicity and emergency department visits in the Negev].

The population of the Negev consists mainly of Jews and Bedouin, who have very different life styles. Patients of both ethnic groups use our emergency department exclusively, providing a unique opportunity to study comparative patient habits. In gathering and processing the information we used Data Mining technology, which allows search for unique patterns in large data bases. We examined demographic data on some 64,000 emergency department visits during 1997-8, mostly medical and surgical cases, but not trauma cases. Many more were by Bedouin than Jews, and between the ages of 25 and 44, more by women than men. There were changes in trends in comparison with an arrival survey conducted some 11 years before.

Adult↗