Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Bayesian learning for cardiac SPECT image interpretation.

In this paper, we describe a system for automating the diagnosis of myocardial perfusion from single-photon emission computerized tomography (SPECT) images of male and female hearts. Initially we had several thousand of SPECT images, other clinical data and physician-interpreter's descriptions of the images. The images were divided into segments based on the Yale system. Each segment was described by the physician as showing one of the following conditions: normal perfusion, reversible perfusion defect, partially reversible perfusion defect, fixed perfusion defect, defect showing reverse redistribution, equivocal defect or artifact. The physician's diagnosis of overall left ventricular (LV) perfusion, based on the above descriptions, categorizes a study as showing one or more of eight possible conditions: normal, ischemia, infarct and ischemia, infarct, reverse redistribution, equivocal, artifact or LV dysfunction. Because of the complexity of the task, we decided to use the knowledge discovery approach, consisting of these steps: problem understanding, data understanding, data preparation, data mining, evaluating the discovered knowledge and its implementation. After going through the data preparation step, in which we constructed normal gender-specific models of the LV and image registration, we ended up with 728 patients for whom we had both SPECT images and corresponding diagnoses. Another major contribution of the paper is the data mining step, in which we used several new Bayesian learning classification methods. The approach we have taken, namely the six-step knowledge discovery process has proven to be very successful in this complex data mining task and as such the process can be extended to other medical data mining projects.

Bayes Theorem↗

[Predictive role of diagnostic information in treatment efficacy of rheumatoid arthritis based on neural network model analysis].

OBJECTIVE: To analyze the indications of the therapies for rheumatoid arthritis (RA) with neural network model analysis. METHODS: Three hundred and ninety-seven patients were included in the clinical trial from 9 clinical centers. They were randomly divided into Western medicine (WM) treated group, 194 cases; and traditional Chinese herbal medicine (CM) treated group, 203 cases. A complete physical examination and 18 common clinical manifestations were prepared before the randomization and after the treatment. The WM therapy included voltaren extended action tablet, methotrexate and sulfasalazine. The CM therapy included Glucosidorum Tripterygii Totorum Tablet and syndrome differentiation treatment. The American College of Rheumatology 20 (ACR20) was taken as efficacy evaluation. All data were analyzed on SAS 8.2 statistical package. The relationships between each variable and efficacy were analyzed, and the variables with P<0.2 were included for the data mining analysis with neural network model. All data were classified into training set (75%) and verification set (25%) for further verification on the data-mining model. RESULTS: Eighteen variables in CM and 24 variables in WM were included in the data-mining model. In CM, morning stiffness, swollen joint number, peripheral immunoglobulin M (IgM) level, tenderness joint number, tenderness, rheumatoid factor (RF), C-reactive protein (CRP) and joint pain were positively related to the efficacy, and disease duration and more urination at night negatively related to the efficacy. In WM, erythrocyte sedimentation rate (ESR), weak waist, white fur in tongue, joint pain, joint stiffness and swollen joint were positively related to the efficacy, and yellow fur in tongue, red tongue, white blood negatively related to the efficacy. In the analysis with the neural network model in the patients of verification set, the predictive response rates of 20% patients would be 100% and 90% in the treatment with CM and WM, respectively. CONCLUSION: Neural network model analysis, based on the full clinical trial data with collection of both traditional Chinese medicine and modern medicine diagnostic information, shows a good predictive role for the information in the efficacy in rheumatoid arthritis.

Adult↗

Use of screening algorithms and computer systems to efficiently signal higher-than-expected combinations of drugs and events in the US FDA's spontaneous reports database.

Since 1998, the US Food and Drug Administration (FDA) has been exploring new automated and rapid Bayesian data mining techniques. These techniques have been used to systematically screen the FDA's huge MedWatch database of voluntary reports of adverse drug events for possible events of concern. The data mining method currently being used is the Multi-Item Gamma Poisson Shrinker (MGPS) program that replaced the Gamma Poisson Shrinker (GPS) program we originally used with the legacy database. The MGPS algorithm, the technical aspects of which are summarised in this paper, computes signal scores for pairs, and for higher-order (e.g. triplet, quadruplet) combinations of drugs and events that are significantly more frequent than their pair-wise associations would predict. MGPS generates consistent, redundant, and replicable signals while minimising random patterns. Signals are generated without using external exposure data, adverse event background information, or medical information on adverse drug reactions. The MGPS interface streamlines multiple input-output processes that previously had been manually integrated. The system, however, cannot distinguish between already-known associations and new associations, so the reviewers must filter these events. In addition to detecting possible serious single-drug adverse event problems, MGPS is currently being evaluated to detect possible synergistic interactions between drugs (drug interactions) and adverse events (syndromes), and to detect differences among subgroups defined by gender and by age, such as paediatrics and geriatrics. In the current data, only 3.4% of all 1.2 million drug-event pairs ever reported (with frequencies > or = 1) generate signals [lower 95% confidence interval limit of the adjusted ratios of the observed counts over expected (O/E) counts (denoted EB05) of > or = 2]. The total frequency count that contributed to signals comprised 23% (2.4 million) of the total number, 10.4 million of drug-event pairs reported, greatly facilitating a more focused follow-up and evaluation. The algorithm provides an objective, systematic view of the data alerting reviewers to critically important, new safety signals. The study of signals detected by current methods, signals stored in the Center for Drug Evaluation and Research's Monitoring Adverse Reports Tracking System, and the signals regarding cerivastatin, a cholesterol-lowering drug voluntarily withdrawn from the market in August 2001, exemplify the potential of data mining to improve early signal detection. The operating characteristics of data mining in detecting early safety signals, exemplified by studying a drug recently well characterised by large clinical trials confirms our experience that the signals generated by data mining have high enough specificity to deserve further investigation. The application of these tools may ultimately improve usage recommendations.

Adverse Drug Reaction Reporting Systems↗

Pharmacogenomic responses of rat liver to methylprednisolone: an approach to mining a rich microarray time series.

A data set was generated to examine global changes in gene expression in rat liver over time in response to a single bolus dose of methylprednisolone. Four control animals and 43 drug-treated animals were humanely killed at 16 different time points following drug administration. Total RNA preparations from the livers of these animals were hybridized to 47 individual Affymetrix RU34A gene chips, generating data for 8799 different probe sets for each chip. Data mining techniques that are applicable to gene array time series data sets in order to identify drug-regulated changes in gene expression were applied to this data set. A series of 4 sequentially applied filters were developed that were designed to eliminate probe sets that were not expressed in the tissue, were not regulated by the drug treatment, or did not meet defined quality control standards. These filters eliminated 7287 probe sets of the 8799 total (82%) from further consideration. Application of judiciously chosen filters is an effective tool for data mining of time series data sets. The remaining data can then be further analyzed by clustering and mathematical modeling techniques.

Adrenalectomy↗

Molecular pathology and future developments.

There has already been a 'molecular' revolution in pathology. Demonstrating transcription of specific single genes or small gene sets and their protein products by in situ hybridisation and immunocytochemistry is routine in diagnostic and experimental pathology. A perhaps-greater revolution is imminent with the application of more recently established and emergent technologies in pathology. These include new approaches to polymerase chain reaction (PCR); simultaneous studies of multiple genes and their expression using oligonucleotide and cDNA arrays; serial analysis of gene expression (SAGE); expressed sequence tag (EST) sequencing, subtractive cloning and differential display; high-throughput sequencing; comparative genomic hybridization, multiplex fluorescence in situ hybridisation (FISH) (spectral karyotyping); reverse chromosome painting; knockout and transgenic organisms; laser microdissection and micro-machining; and new methods in bio-informatics, 'data mining' and data visualisation. Molecular methods will profoundly change diagnosis, prognosis and treatment targeting in oncology and elucidate fundamental mechanisms of neoplastic transformation. Individual susceptibility to specific diseases will become assessable and screening will be refined. The new molecular biology will be most fruitful in partnership with classical approaches to pathology: the expectation that molecular methods alone will answer all pathological questions is unrealistic. A further challenge for the biomedical community in the 'genome era' will be to ensure that the benefits of these sophisticated technologies are enjoyed globally.

Biopsy↗

Defining and maintaining a high quality screening collection: the GSK experience.

Understanding the quality of a screening collection is the first step to improving it and, as a result, the quality of the screening process. This article outlines how this issue was approached at GlaxoSmithKline and some of the hurdles that needed to be overcome to achieve success. The article focuses specifically on the necessary software and hardware infrastructure needed, and at some of the extra benefits of such a project in terms of data mining and data modelling.

Chromatography, High Pressure Liquid↗

HOX Pro DB: the functional genomics of hox ensembles.

The HOX Pro database contains information about the organization, function and evolution of gene ensembles, notably the homeobox-containing genes. It is now clear that a subset of genes containing the homeobox motif play key roles in the orchestration of genes which control embryonic patterning, morphogenesis, cell differentiation and malignant transformation. The HOX Pro contains a broad spectrum of information including images, diagrams and animations. Currently this amounts to approximately 700 HTML pages together with 400 images which contain information on 200 groups of genes and 90 promoters, in turn linked to maps of 13 HOX clusters and nine genetic networks. There are about 700 sequences of individual hox-genes of animals classified in approximately 200 homologous or paralogous groups. Graphical representation of HOX clusters and Hox-based networks is accomplished by means of flow and 3D diagrams, JavaScript animations and Java applets. The HOX Pro now includes sections presenting data mining and data simulation issues. The DB is located at http://www.iephb.nw.ru/hoxpro.

Animals↗

CoreGenes: a computational tool for identifying and cataloging "core" genes in a set of small genomes.

BACKGROUND: Improvements in DNA sequencing technology and methodology have led to the rapid expansion of databases comprising DNA sequence, gene and genome data. Lower operational costs and heightened interest resulting from initial intriguing novel discoveries from genomics are also contributing to the accumulation of these data sets. A major challenge is to analyze and to mine data from these databases, especially whole genomes. There is a need for computational tools that look globally at genomes for data mining. RESULTS: CoreGenes is a global JAVA-based interactive data mining tool that identifies and catalogs a "core" set of genes from two to five small whole genomes simultaneously. CoreGenes performs hierarchical and iterative BLASTP analyses using one genome as a reference and another as a query. Subsequent query genomes are compared against each newly generated "consensus." These iterations lead to a matrix comprising related genes from this set of genomes, e. g., viruses, mitochondria and chloroplasts. Currently the software is limited to small genomes on the order of 330 kilobases or less. CONCLUSION: A computational tool CoreGenes has been developed to analyze small whole genomes globally. BLAST score-related and putatively essential "core" gene data are displayed as a table with links to GenBank for further data on the genes of interest. This web resource is available at http://pumpkins.ib3.gmu.edu:8080/CoreGenes or http://www.bif.atcc.org/CoreGenes.

Algorithms↗

Mining biological data using self-organizing map.

This paper presents a novel method of mining biological data using a self-organizing map (SOM). After partitioning a set of protein sequences using SOM, conventional homology alignment is applied to each cluster to determine the conserved local motif (biological pattern) for the cluster. These local motifs are then regarded as rules for prediction and classification. In the application to the prediction of HIV protease cleavage sites in proteins, we found that the rules derived from this method are much more robust than those derived from the decision tree method.

Algorithms↗

Predicting gene function in Saccharomyces cerevisiae.

MOTIVATION: S.cerevisiae is one of the most important model organisms, and has has been the focus of over a century of study. In spite of these efforts, 40% of its open reading frames (ORFs) remain classified as having unknown function (MIPS: Munich Information Center for Protein Sequences). We wished to make predictions for the function of these ORFs using data mining, as we have previously successfully done for the genomes of M.tuberculosis and E.coli. Applying this approach to the larger and eukaryotic S.cerevisiae genome involves modifying the machine learning and data mining algorithms, as this is a larger organism with more data available, and a more challenging functional classification. RESULTS: Novel extensions to the machine learning and data mining algorithms have been devised in order to deal with the challenges. Accurate rules have been learned and predictions have been made for many of the ORFs whose function is currently unknown. The rules are informative, agree with known biology and allow for scientific discovery. AVAILABILITY: All predictions are freely available from http://www.genepredictions.org, all datasets used in this study are freely available from http://www.aber.ac.uk/compsci/Research/bio/dss/yeastdataand software for relational data mining is available from http://www.aber.ac.uk/compsci/Research/bio/dss/polyfarm.

Chromosome Mapping↗

Multiclass cancer classification using gene expression profiling and probabilistic neural networks.

Gene expression profiling by microarray technology has been successfully applied to classification and diagnostic prediction of cancers. Various machine learning and data mining methods are currently used for classifying gene expression data. However, these methods have not been developed to address the specific requirements of gene microarray analysis. First, microarray data is characterized by a high-dimensional feature space often exceeding the sample space dimensionality by a factor of 100 or more. In addition, microarray data exhibit a high degree of noise. Most of the discussed methods do not adequately address the problem of dimensionality and noise. Furthermore, although machine learning and data mining methods are based on statistics, most such techniques do not address the biologist's requirement for sound mathematical confidence measures. Finally, most machine learning and data mining classification methods fail to incorporate misclassification costs, i.e. they are indifferent to the costs associated with false positive and false negative classifications. In this paper, we present a probabilistic neural network (PNN) model that addresses all these issues. The PNN model provides sound statistical confidences for its decisions, and it is able to model asymmetrical misclassification costs. Furthermore, we demonstrate the performance of the PNN for multiclass gene expression data sets. Here, we compare the performance of the PNN with two machine learning methods, a decision tree and a neural network. To assess and evaluate the performance of the classifiers, we use a lift-based scoring system that allows a fair comparison of different models. The PNN clearly outperformed the other models. The results demonstrate the successful application of the PNN model for multiclass cancer classification.

Artificial Intelligence↗

Alkahest NuclearBLAST : a user-friendly BLAST management and analysis system.

BACKGROUND: Sequencing of EST and BAC end datasets is no longer limited to large research groups. Drops in per-base pricing have made high throughput sequencing accessible to individual investigators. However, there are few options available which provide a free and user-friendly solution to the BLAST result storage and data mining needs of biologists. RESULTS: Here we describe NuclearBLAST, a batch BLAST analysis, storage and management system designed for the biologist. It is a wrapper for NCBI BLAST which provides a user-friendly web interface which includes a request wizard and the ability to view and mine the results. All BLAST results are stored in a MySQL database which allows for more advanced data-mining through supplied command-line utilities or direct database access. NuclearBLAST can be installed on a single machine or clustered amongst a number of machines to improve analysis throughput. NuclearBLAST provides a platform which eases data-mining of multiple BLAST results. With the supplied scripts, the program can export data into a spreadsheet-friendly format, automatically assign Gene Ontology terms to sequences and provide bi-directional best hits between two datasets. Users with SQL experience can use the database to ask even more complex questions and extract any subset of data they require. CONCLUSION: This tool provides a user-friendly interface for requesting, viewing and mining of BLAST results which makes the management and data-mining of large sets of BLAST analyses tractable to biologists.

Algorithms↗

Outlier mining in high throughput screening experiments.

A data mining procedure for the rapid scoring of high-throughput screening (HTS) compounds is presented. The method is particularly useful for monitoring the quality of HTS data and tracking outliers in automated pharmaceutical or agrochemical screening, thus providing more complete and thorough structure-activity relationship (SAR) information. The method is based on the utilization of the assumed relationship between the structure of the screened compounds and the biological activity on a given screen expressed on a binary scale. By means of a data mining method, a SAR description of the data is developed that assigns probabilities of being a hit to each compound of the screen. Then, an inconsistency score expressing the degree of deviation between the adequacy of the SAR description and the actual biological activity is computed. The inconsistency score enables the identification of potential outliers that can be primed for validation experiments. The approach is particularly useful for detecting false-negative outliers and for identifying SAR-compliant hit/nonhit borderline compounds, both of which are classes of compounds that can contribute substantially to the development and understanding of robust SARs. In a first implementation of the method, one- and two-dimensional descriptors are used for encoding molecular structure information and logistic regression for calculating hits/nonhits probability scores. The approach was validated on three data sets, the first one from a publicly available screening data set and the second and third from in-house HTS screening campaigns. Because of its simplicity, robustness, and accuracy, the procedure is suitable for automation.

Algorithms↗

Endotoxin-like reactions with intravenous gentamicin: results from pharmacovigilance tools under investigation.

OBJECTIVE: To apply two data mining algorithms (DMAs) to Food and Drug Administration (FDA) Adverse Event Reporting System (AERS) reports that involved endotoxin-like reactions with intravenous gentamicin to determine whether a signal of disproportionate reporting of these events would have been generated concurrently with surveillance based on clinical observation. DESIGN: Multi-item gamma-Poisson shrinker (MGPS) and proportional reporting ratios (PRRs) were used. Data used for data mining consisted of an extract of the FDA AERS database. Previously published details of clusters of endotoxin-like reactions to intravenous gentamicin were used to select adverse events for data mining. RESULTS: The first signal of disproportionate reporting with any relevant event occurred in 1998, the year in which the outbreak was identified and evaluated by the Centers for Disease Control and Prevention and the FDA. In 1997, there were only 6 reports of rigors in the AERS; this jumped to 68 in 1998. In 1998, a signal was generated for endotoxic shock with PRRs but not with MGPS, based on one case. CONCLUSIONS: The two DMAs generated signals concurrently with the influx of reports. It would have been difficult for safety reviewers to ignore an increase in rigors by traditional methods of safety surveillance; therefore, DMAs might not have had a great deal to offer in this instance. If data mining were considered as a second-line defense to diligent clinical observations under similar circumstances, simple disproportionality methods such as PRRs might be more useful than DMAs such as MGPS when commonly cited thresholds are used.

Adverse Drug Reaction Reporting Systems↗

Assessment of potential biases in the application of MSHA respirable coal mine dust data to an epidemiologic study.

Systematic errors in exposure data will result in biased estimates of the exposure-response relationship derived from epidemiologic analyses. Thus, adjustment of exposure data to account for identified errors may provide for a more accurate assessment of effect. In preparing to apply respirable coal mine dust exposure data collected by the Mine Safety and Health Administration (MSHA) to a study of the pulmonary status of underground coal miners, an assessment of potential systematic errors was undertaken. Potential errors stemming from adjustment of controls during sampling, concentration-dependent sampling, truncation of sampling results, identified sampling equipment problems, and a disproportionate number of low concentration samples in mine operator-collected samples were identified and evaluated. Methods to account for these errors and adjust mean exposures by mine, occupation, and year are given.

Bias↗

Postmarketing surveillance of potentially fatal reactions to oncology drugs: potential utility of two signal-detection algorithms.

PURPOSE: Several data mining algorithms (DMAs) are being studied in hopes of enhancing screening of large post-marketing safety databases for signals of novel adverse events (AEs). The objective of this study was to apply two DMAs to the United States FDA Adverse Event Reporting System (AERS) database to see whether signals of potentially fatal AEs with cancer drugs might have been identified earlier than with traditional methods. METHODS: Screening algorithms used for analysis were the multi-item gamma Poisson shrinker (MGPS) and proportional reporting ratios (PRRs). Data mining was performed on data from the FDA AERS database. When a signal was identified, it was compared with that in the year in which the event was added to package insert and/or the year a "case series" was published. A recent publication summarizing the time of dissemination of information on potentially fatal AEs to cancer drugs provided the data set for analysis. RESULTS: The peer-reviewed published analysis contained 21 drugs and 26 drug-event combinations (DECs) that were considered sufficiently specific for data mining. Twenty-four of the DECs generated a signal of disproportionate reporting with PRRs (6 at 1 year and 16 from 2 years to 18 years prior to either a published "case series" or a package insert change) and 20 with MGPS (3 at 1 year and 11 from 2 years to 16 years prior to either a published "case series" or a package insert change). Two DECs did not signal with either DMA. CONCLUSION: At least one commonly cited DMA generated a signal of disproportionate reporting for 24 of 26 DECs for selected cancer drugs. For 16 DECs, one could conclude that a signal was generated well in advance (> or =2 years) of standard techniques in use with at least one DMA. DMAs might be useful in supplementing traditional surveillance strategies with oncology drugs and other drugs with similar features. (i.e., drugs that may be approved on an accelerated basis, are known to have serious toxicity, are administered to patients with substantial and complicated comorbid illness, are not available to the general medical community, and may have a high frequency of "off-label" use).

Adverse Drug Reaction Reporting Systems↗

[The working conditions and health status of miners in Donets Basin coal mines].

Data are reported on working conditions of coal miners considering the main physical (dust, noise, vibration, microclimate) and chemical environmental professional factors and their prognosis up to the year 2005. The authors analyze professional morbidity (pneumoconiosis, dust-induced bronchitis, vibration disease, cochlear neuritis etc.) and diseases with temporary loss of the working capacity invalidity and mortality of miners. The relation between working conditions and health status of miners were analyzed.

Absenteeism↗