Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Knowledge discovery: Detecting elderly patients with impaired mobility.

Immobility is an important health concern for the elderly patients and healthcare providers who care for the elderly. The purpose of this study was to test a knowledge discovery method to detect elderly patients with impaired mobility in a large clinical dataset. The research method applied an exploratory design and a data mining classification method (cost sensitive Decision Tree J48 from WEKA) to classify patients. Important factors were identified by the Feature Selection method. The Decision Tree algorithm classified patients in the dataset with 65% sensitivity and 72% specificity for a reduced model. The results were evaluated by 10-fold cross validation. Examples of decision rules were also extracted. The study can be applied to classify different health problems in different populations and serves as a foundation for the development of healthcare decision support systems.

Aged↗

Markup of temporal information in electronic health records.

Temporal information plays a critical role in the understanding of clinical narrative (i.e., free text). We developed a representation for marking up temporal information in a narrative, consisting of five elements: 1) reference point, 2) direction, 3) number, 4) time unit, and 5) pattern. We identified 254 temporal expressions from 50 discharge summaries and represented them using our scheme. The overall inter-rater reliability among raters applying the representation model was 75 percent agreement. The model can contribute to temporal reasoning in computer systems for decision support, data mining, and process and outcomes analyses by providing structured temporal information.

Hospitals, Religious↗

CytoAccess, a relational laboratory information management system for a clinical cytogenetics laboratory.

We developed a CytoAccess laboratory management system based on the widely used Microsoft Access software to facilitate data processing, result reporting, and quality management for a full-service cytogenetics laboratory. The CytoAccess system consists of four functional modules. The data entry module is for logging in patient information. The result entry module is used to generate chromosome, fluorescent in situ hybridization (FISH), and array comparative genomic hybridization (aCGH) reports. The administrative module enables periodic monitoring of quality control and quality improvement (QA/QI) parameters and produces billing forms. The maintenance module allows users to update clinical demographics, report templates, code tables, and to refresh data links. We have integrated into this system over 15,000 chromosome and FISH results from prenatal, postnatal, and cancer cases for the past six years. This system is cost-effective, user-friendly, flexible in updating, and potentially adaptable for data mining.

Journal Article↗

[Symptom evaluation as an efficient diagnostic tool for upper abdominal diseases].

INTRODUCTION: GERD and dyspepsia are common conditions, affecting approximately 25-40% of the general population. In the absence of alarm symptoms, the current recommended policy in young dyspeptic patients is a "test and treat" strategy for H. pylori. On the other hand, in GERD patients, a therapeutic trial with proton pump inhibitors (PPI) is the treatment of choice. AIM: This study aimed to create short and simple clinical algorithms, for the diagnosis and treatment of patients with upper gastrointestinal complaints. METHODS: Data mining models and algorithms (neural networks, decision trees and logistic regression) were used to create a diagnostic symptom questionnaire to classify the patients into one of two diagnostic groups: GERD vs. non-GERD. The questionnaire was validated against endoscopic and clinical diagnoses of 132 patients in a cross-sectional study, and was designed to yield a "GERD score". The clinical and economical benefits of the new algorithm were evaluated in primary care clinics in Israel. RESULTS: The symptoms chosen for the diagnostic questionnaire were heartburn, acid regurgitation, sour oral taste, aggravation of symptoms after heavy meals, relief of symptoms by antacids, and nocturnal reflux. The use of the algorithm by primary care physicians improved clinical outcomes and reduced health resources consumption. CONCLUSIONS: This new short algorithm was found to be useful and easy to apply in clinical practice. The questionnaire can be also implemented in computerized medical systems.

Algorithms↗

Coronary angiography and its complications. The search for risk factors.

Previous efforts to define risk factors associated with coronary angiography have generally focused on only one or two potential candidates. These elements are often found by exploring data sets for relationships not generally specified in advance. We attempted to determine whether 14 factors were associated with increased risk in 995 patients undergoing angiocardiography in five medical centers. The study found a consistently increased risk for women, for patients having lengthier procedures, and for patients studied at one of the five centers. No evidence of increased risk was found for 11 other factors, including two related to operator experience. Data exploration ("data mining") was also undertaken, using sequential multiple logistic regression analyses. A strong risk of complications, not previously observed or hypothesized in advance, was noted for vascular complications in women studied by the brachial technique.

Angiocardiography↗

Aeromedical evacuation in Bosnia.

The objective of this study was to assess the value of aeromedical evacuation when compared to road ambulance transportation in predominantly trauma patients in a rural area. Uniquely, trauma was the most common presenting condition (75%), distances to secondary care facilities were long and road routes were poor with a risk of being mined. Data were collected of all British aeromedical flights in Multi-National Division Southwest, Bosnia-Herzegovina, over a six-month period, and benefit to the patient was assessed by a panel of experts when compared to calculated road ambulance evacuation. Sixty-nine patients were evacuated by air on 57 flights and transported to a secondary care facility for further management. The panel of experts found that only 15 of the 69 patients (22%) had benefited from aeromedical evacuation. This study again shows the low benefit to the patient from indiscriminate use of aeromedical evacuation, despite the air ambulance being operated in apparently ideal conditions of a high percentage of trauma, a rural setting and poor road communications. Crew safety and the high costs further highlight the need to devise a system that can screen out unnecessary flights and identify those patients who would benefit most.

Adult↗

Projecting corporate health plan utilization and charges from annual ICD-9-CM diagnostic rates: a value-added opportunity for pathologists.

OBJECTIVE: To develop an allocation method for corporate health plan resources and expenditures based on annual International Classification of Disease, 9th revision, Clinical Modification (ICD-9-CM) diagnostic rate stratification as a surrogate for disease incidence. DESIGN: A data-mining process was applied to a self-insured corporate health plan database. Annual membership rates of Current Procedural Terminology (CPT) procedure utilization and charges between 1990 and 1994 for a cohort of 7216 continuously employed plan members were stratified according to the annual rates of the 19 major ICD-9-CM diagnostic classifications. The stratified annual CPT utilization and charge rates were analyzed by correlation analysis and one-tailed t test. RESULTS: Laboratory and pathology procedure utilization and charge rates were highly correlated with specific rankings of ICD-9-CM diagnostic classifications. The health plan diagnostic rate, laboratory utilization rate, and all charge rates increased significantly during the 5 study years. CONCLUSION: Although all procedure utilization and charge rates in this health plan increased each year, their proportionality consistently was maintained among diagnostically related groups of patients. By restraining global expenditures, managed health plans conflict with historical utilization and charge patterns. Treating ICD-9-CM diagnostic groups as disease management services within a managed care plan allows procedures and expenses to be allocated according to medical necessity in the context of total membership benefits. For pathologists, who recently were mandated by the Health Care Financing Administration and the Office of Inspector General of the Department of Health and Human Services to become stewards of ICD-9-CM coding, this is a unique opportunity to lead an initiative to perfect managed care. The process will require permanent patient numbers, computerized longitudinal patient records, and standardized coded medical terminology.

Adult↗

CATCH/IT: a data warehouse to support comprehensive assessment for tracking community health.

A systematic methodology, Comprehensive Assessment for Tracking Community Health (CATCH), for analyzing the health status of communities has been under development at the University of South Florida since the early 1990s. CATCH draws 226 health status indicators from multiple data sources and uses an innovative comparative framework and weighted evaluation criteria to produce a rank-ordered list of community health problems. CATCH has been applied successfully in many Florida counties; focusing attention on high priority health issues and measuring the impact of health expenditures on community health status outcomes. Previously performed manually, we are using information technology (IT) to automate the CATCH methodology with a full-scale data warehouse, user-friendly forms and reports, and extended analysis and data mining capabilities. The automated system, CATCH/IT, will reduce the time to prepare community health status reports from months to days. In this paper, we present the current status of the project, along with the principal research and development issues and future directions of the project.

Community Health Planning↗

Scaling a data retrieval and mining application to the enterprise-wide level.

Most medical institutions have had difficulty in adopting practices that use stored clinical and administrative data effectively. This stems in part from the lack of available tools to easily and accurately retrieve datasets of interest. In this work, we describe the development of a data retrieval and mining application, Goldminer, which allows authorized personnel at our institution to query clinical and demographic data stores through a graphical, non-programmer interface. It builds upon DXtractor, our previously described tool that retrieves data from a smaller, more specialized dataset. We discuss the difficulties encountered in scaling this application to the enterprise-wide level, and our solutions.

Databases as Topic↗

New data analysis and mining approaches identify unique proteome and transcriptome markers of susceptibility to autoimmune diabetes.

Non-obese diabetic (NOD) mice spontaneously develop autoimmunity to the insulin producing beta cells leading to insulin-dependent diabetes. In this study we developed and used new data analysis and mining approaches on combined proteome and transcriptome (molecular phenotype) data to define pathways affected by abnormalities in peripheral leukocytes of young NOD female mice. Cells were collected before mice show signs of autoimmunity (age, 2-4 weeks). We extracted both protein and RNA from NOD and C57BL/6 control mice to conduct both proteome analysis by two-dimensional gel electrophoresis and transcriptome analysis on Affymetrix expression arrays. We developed a new approach to analyze the two-dimensional gel proteome data that included two-way analysis of variance, cluster analysis, and principal component analysis. Lists of differentially expressed proteins and transcripts were subjected to pathway analysis using a commercial service. From the list of 24 proteins differentially expressed between strains we identified two highly significant and interconnected networks centered around oncogenes (Myc and Mycn) and apoptosis-related genes (Bcl2 and Casp3). The 273 genes with significant strain differences in RNA expression levels created six interconnected networks with a significant over-representation of genes related to cancer, cell cycle, and cell death. They contained many of the same genes found in the proteome networks (including Myc and Mycn). The combination of the eight, highly significant networks created one large network of 272 genes of which 82 had differential expression between strains either at the protein or the RNA level. We conclude that new proteome data analysis strategies and combined information from proteome and transcriptome can enhance the insights gained from either type of data alone. The overall systems biology of prediabetic NOD mice points toward abnormalities in regulation of the opposing processes of cell renewal and cell death even before there are any clear signatures of immune system activation.

Analysis of Variance↗

Managing and mining protein crystallization data.

The crystallization of macromolecules remains a major bottleneck in structural biology. The routine screening of more than one thousand crystallization conditions and subsequent optimization by fine screening presents a challenge to conventional laboratory notebook keeping. In addition, the development of high-throughput robotic crystallization and imaging systems presents a pressing need for low-cost laboratory information management system (LIMS). Here we describe CLIMS2, a crystallization LIMS that features a simple, user-friendly graphical interface, allowing the storage, management, retrieval and mining of crystallization data. The CLIMS2 executable and documentation is freely available at http://clims.med.monash.edu.au.

Computer Graphics↗

Mining nematode genome data for novel drug targets.

Expressed sequence tag projects have currently produced over 400 000 partial gene sequences from more than 30 nematode species and the full genomic sequences of selected nematodes are being determined. In addition, functional analyses in the model nematode Caenorhabditis elegans have addressed the role of almost all genes predicted by the genome sequence. This recent explosion in the amount of available nematode DNA sequences, coupled with new gene function data, provides an unprecedented opportunity to identify pre-validated drug targets through efficient mining of nematode genomic databases. This article describes the various information sources available and strategies that can expedite this process.

Animals↗

Mining microarray expression data by literature profiling.

BACKGROUND: The rapidly expanding fields of genomics and proteomics have prompted the development of computational methods for managing, analyzing and visualizing expression data derived from microarray screening. Nevertheless, the lack of efficient techniques for assessing the biological implications of gene-expression data remains an important obstacle in exploiting this information. RESULTS: To address this need, we have developed a mining technique based on the analysis of literature profiles generated by extracting the frequencies of certain terms from thousands of abstracts stored in the Medline literature database. Terms are then filtered on the basis of both repetitive occurrence and co-occurrence among multiple gene entries. Finally, clustering analysis is performed on the retained frequency values, shaping a coherent picture of the functional relationship among large and heterogeneous lists of genes. Such data treatment also provides information on the nature and pertinence of the associations that were formed. CONCLUSIONS: The analysis of patterns of term occurrence in abstracts constitutes a means of exploring the biological significance of large and heterogeneous lists of genes. This approach should contribute to optimizing the exploitation of microarray technologies by providing investigators with an interface between complex expression data and large literature resources.

Cluster Analysis↗

Mining gene expression data for positive and negative co-regulated gene clusters.

MOTIVATION: Analysis of gene expression data can provide insights into the positive and negative co-regulation of genes. However, existing methods such as association rule mining are computationally expensive and the quality and quantities of the rules are sensitive to the support and confidence values. In this paper, we introduce the concept of positive and negative co-regulated gene cluster (PNCGC) that more accurately reflects the co-regulation of genes, and propose an efficient algorithm to extract PNCGCs. RESULTS: We experimented with the Yeast dataset and compared our resulting PNCGCs with the association rules generated by the Apriori mining algorithm. Our results show that our PNCGCs identify some missing co-regulations of association rules, and our algorithm greatly reduces the large number of rules involving uncorrelated genes generated by the Apriori scheme. AVAILABILITY: The software is available upon request.

Algorithms↗

Experimental Peptide Identification Repository (EPIR): an integrated peptide-centric platform for validation and mining of tandem mass spectrometry data.

LC MS/MS has become an established technology in proteomic studies, and with the maturation of the technology the bottleneck has shifted from data generation to data validation and mining. To address this bottleneck we developed Experimental Peptide Identification Repository (EPIR), which is an integrated software platform for storage, validation, and mining of LC MS/MS-derived peptide evidence. EPIR is a cumulative data repository where precursor ions are linked to peptide assignments and protein associations returned by a search engine (e.g. Mascot, Sequest, or PepSea). Any number of datasets can be parsed into EPIR and subsequently validated and mined using a set of software modules that overlay the database. These include a peptide validation module, a protein grouping module, a generic module for extracting quantitative data, a comparative module, and additional modules for extracting statistical information. In the present study, the utility of EPIR and associated software tools is demonstrated on LC MS/MS data derived from a set of model proteins and complex protein mixtures derived from MCF-7 breast cancer cells. Emphasis is placed on the key strengths of EPIR, including the ability to validate and mine multiple combined datasets, and presentation of protein-level evidence in concise, nonredundant protein groups that are based on shared peptide evidence.

Breast Neoplasms↗

EST pipeline system: detailed and automated EST data processing and mining.

Expressed sequence tags (ESTs) are widely used in gene survey research these years. The EST Pipeline System, software developed by Hangzhou Genomics Institute (HGI), can automatically analyze different scalar EST sequences by suitable methods. All the analysis reports, including those of vector masking, sequence assembly, gene annotation, Gene Ontology classification, and some other analyses, can be browsed and searched as well as downloaded in the Excel format from the web interface, saving research efforts from routine data processing for biological rules embedded in the data.

Automation↗

Mining of biological data I: identifying discriminating features via mean hypothesis testing.

Large volumes of data are routinely collected during bioprocess operations and, more recently, in basic biological research using genomics-based technologies. While these data often lack sufficient detail to be used for mechanism identification, it is possible that the underlying mechanisms affecting cell phenotype or process outcome are reflected as specific patterns in the overall or temporal sensor logs. This raises the possibility of identifying outcome-specific fingerprints that can be used for process or phenotype classification and the identification of discriminating characteristics, such as specific genes or process variables. The aim of this work is to provide a systematic approach to identifying and modeling patterns in historical records and using this information for process classification. This approach differs from others in that emphasis is placed on analyzing the data structure first and thereby extracting potentially relevant features prior to model creation. The initial step in this overall approach is to first identify the discriminating features of the relevant measurements and time windows, which can then be subsequently used to discriminate among different classes of process behavior. This is achieved via a mean hypothesis testing algorithm. Next, the homogeneity of the multivariate data in each class is explored via a novel cluster analysis technique called PC1 Time Series Clustering to ensure that the data subsets used accurately reflect the variability displayed in the historical records. This will be the topic of the second paper in this series. We present here the method for identifying discriminating features in data via mean hypothesis testing along with results from the analysis of case studies from industrial fermentations

Algorithms↗