Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Evaluation of mutual information and genetic programming for feature selection in QSAR.

Feature selection is a key step in Quantitative Structure Activity Relationship (QSAR) analysis. Chance correlations and multicollinearity are two major problems often encountered when attempting to find generalized QSAR models for use in drug design. Optimal QSAR models require an objective variable relevance analysis step for producing robust classifiers with low complexity and good predictive accuracy. Genetic algorithms coupled with information theoretic approaches such as mutual information have been used to find near-optimal solutions to such multicriteria optimization problems. In this paper, we describe a novel approach for analyzing QSAR data based on these methods. Our experiments with the Thrombin dataset, previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001 demonstrate the feasibility of this approach. It has been found that it is important to take into account the data distribution, the rule "interestingness", and the need to look at more invariant and monotonic measures of feature selection.

Algorithms↗

CGO: utilizing and integrating gene expression microarray data in clinical research and data management.

Clinical GeneOrganizer (CGO) is a novel windows-based archiving, organization and data mining software for the integration of gene expression profiling in clinical medicine. The program implements various user-friendly tools and extracts data for further statistical analysis. This software was written for Affymetrix GeneChip *.txt files, but can also be used for any other microarray-derived data. The MS-SQL server version acts as a data mart and links microarray data with clinical parameters of any other existing database and therefore represents a valuable tool for combining gene expression analysis and clinical disease characteristics.

Computational Biology↗

Gene expression informatics.

There are many methodologies for performing gene expression profiling on transcripts, and through their use scientists have been generating vast amounts of experimental data. Turning the raw experimental data into meaningful biological observation requires a number of processing steps; to remove noise, to identify the "true" expression value, normalize the data, compare it to reference data, and to extract patterns, or obtain insight into the underlying biology of the samples being measured. In this chapter we give a brief overview of how the raw data is processed, provide details on several data-mining methods, and discuss the future direction of expression informatics.

Animals↗

Mining associations between genetic markers, phenotypes, and covariates.

We used Haplotype Pattern Mining, HPM [Toivonen et al., Am J Hum Genet 67:133-45, 2000], for gene localization in Genetic Analysis Workshop (GAW) 12 isolate data. In HPM, association is analyzed by searching all trait-associated haplotype patterns. Data mining algorithms are utilized to make the search efficient. The strength of the haplotype-trait associations is measured by a linear model, into which a pre-seelected set of covariates is incorporated. Marker-wise patterns of association are used for predicting the disease gene location. Genome-wide scans of susceptibility genes for affection status as well as for the quantitative traits (Q1-Q5) were performed. First analyses were made with small sample sizes, 63-94 trios per trait, which is compared with a pilot study of a larger complex disease-mapping project. Subsequently, the analysis was repeated with approximately 600 cases and 600 controls per trait to give higher power to the analyses. With small sample sizes, only the susceptibility genes having the strongest effects on the traits could be localized. The larger sample size gave very good results: all susceptibility genes, except one, could be correctly localized. First experiments on candidate genes suggested that HPM is applicable even to fine mapping of mutations in DNA sequence.

Algorithms↗

C-QSAR: a database of 18,000 QSARs and associated biological and physical data.

The C-QSAR program is used to develop and search a database of over 18,000 equations that relate biological or physico-chemical properties of molecules to various molecular descriptors. The data used to derive the quantitative structure activity relationships (QSAR) are taken from various high quality journals. C-QSAR comprises two databases, one for structure-activity information biological systems (n = 9200) and the other for physical organic systems. Users can search the data in 20 different fields; for example by structure or substructure of the compounds involved, by the type of property correlated, by molecular properties, or by properties of the QSAR equation. Various ways in which information can be obtained is briefly discussed. Initially the database is often used for data mining, to search lead molecules, for substituent selection and "model mining" for lateral validation. The regression analysis is useful when the user wants to derive a new QSAR using his structures and activity data.

Chemical Phenomena↗

Characterizing the grape transcriptome. Analysis of expressed sequence tags from multiple Vitis species and development of a compendium of gene expression during berry development.

We report the analysis and annotation of 146,075 expressed sequence tags from Vitis species. The majority of these sequences were derived from different cultivars of Vitis vinifera, comprising an estimated 25,746 unique contig and singleton sequences that survey transcription in various tissues and developmental stages and during biotic and abiotic stress. Putatively homologous proteins were identified for over 17,752 of the transcripts, with 1,962 transcripts further subdivided into one or more Gene Ontology categories. A simple structured vocabulary, with modules for plant genotype, plant development, and stress, was developed to describe the relationship between individual expressed sequence tags and cDNA libraries; the resulting vocabulary provides query terms to facilitate data mining within the context of a relational database. As a measure of the extent to which characterized metabolic pathways were encompassed by the data set, we searched for homologs of the enzymes leading from glycolysis, through the oxidative/nonoxidative pentose phosphate pathway, and into the general phenylpropanoid pathway. Homologs were identified for 65 of these 77 enzymes, with 86% of enzymatic steps represented by paralogous genes. Differentially expressed transcripts were identified by means of a stringent believability index cutoff of > or =98.4%. Correlation analysis and two-dimensional hierarchical clustering grouped these transcripts according to similarity of expression. In the broadest analysis, 665 differentially expressed transcripts were identified across 29 cDNA libraries, representing a range of developmental and stress conditions. The groupings revealed expected associations between plant developmental stages and tissue types, with the notable exception of abiotic stress treatments. A more focused analysis of flower and berry development identified 87 differentially expressed transcripts and provides the basis for a compendium that relates gene expression and annotation to previously characterized aspects of berry development and physiology. Comparison with published results for select genes, as well as correlation analysis between independent data sets, suggests that the inferred in silico patterns of expression are likely to be an accurate representation of transcript abundance for the conditions surveyed. Thus, the combined data set reveals the in silico expression patterns for hundreds of genes in V. vinifera, the majority of which have not been previously studied within this species.

DNA, Complementary↗

"Saba"--application of knowledge and state-of-the-art technologies in the field of psychiatry for development of new diagnostics prevention and therapeutic tools for schizophrenia.

The aim of project is to build a European network, which will integrate the research capabilities of a group of research institutes and university departments to provide an infrastructure for the highest quality research in psychiatric disorders, particularly in schizophrenia and schizoaffective disturbances. This network will integrate original computer expert advisory system called "Saba" with modern brain imaging techniques and neurophysiological methods, which allows for the delineation of specific subtypes and particular episodes of mental disorders and their neural bases will be studied by state-of-the art (high tech) imaging techniques. This approach will lead to new investigatory, diagnostic and therapeutic techniques. Together the members of this network will comprise an unmatched critical mass of human and other resources aimed at fundamental and applied research into a group of disorders, which impose a huge burden on social and material capital. The relationships and mutual responsibilities between neuroscience and the society it serves will be addressed specifically. Top brain research is performed at several locations in Europe. In particular, in the area of linking classical psychiatric and psychological assessment methods and the newest brain imaging techniques in mental disorders, major progress can only be made when various research groups join their efforts. Large-scale studies using different databases are critically required, which demands standardization of the description of mental disorders and of the applied techniques and methods of analysis. Imaging techniques, including functional MRI (fMRI), Evoked Potentials (EPs), brain mapping, and the computer gathered information will be shared, standardized and further developed within the network. Developing new information technology tools for simulation, visualization and data-mining will be required to enable effective search for links between mental disorders and brain characteristics (function, structure) in very large scale data-sets acquired and stored in various research facilities.

Brain↗

Clinical information systems market - an insider's view.

Clinical information systems that provide electronic charting and documentation have been commercially available for over 15 years. These systems provide varying degrees of automation to flowsheets, forms, notes, worklists, care plans, and medication administration records. Although there are many benefits that an electronic system brings, such as accessibility, legibility, process adherence, and data mining, the market has been slow to adopt these systems. A variety of historical factors can explain the lack of widespread system implementations. Survey data of CEOs/CIOs from the Healthcare Information and Management Systems Society (HIMSS) shows promising data that clinically oriented applications will receive high prioritization in near term planning. Will this prioritization materialize in actual implementations? Market drivers appear to be in place to predict an increase in sales and implementations.

Computer Systems↗

Automatic discovery of cross-family sequence features associated with protein function.

BACKGROUND: Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterized protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. RESULTS: We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. CONCLUSION: We have developed a novel and useful approach for knowledge discovery in annotated sequence data. The technique is able to identify functionally important sequence features and does not require expert knowledge. By viewing protein function from a sequence perspective, the approach is also suitable for discovering unexpected links between biological processes, such as the recently discovered role of ubiquitination in transcription.

Algorithms↗

System-based proteomic analysis of the interferon response in human liver cells.

BACKGROUND: Interferons (IFNs) play a critical role in the host antiviral defense and are an essential component of current therapies against hepatitis C virus (HCV), a major cause of liver disease worldwide. To examine liver-specific responses to IFN and begin to elucidate the mechanisms of IFN inhibition of virus replication, we performed a global quantitative proteomic analysis in a human hepatoma cell line (Huh7) in the presence and absence of IFN treatment using the isotope-coded affinity tag (ICAT) method and tandem mass spectrometry (MS/MS). RESULTS: In three subcellular fractions from the Huh7 cells treated with IFN (400 IU/ml, 16 h) or mock-treated, we identified more than 1,364 proteins at a threshold that corresponds to less than 5% false-positive error rate. Among these, 54 were induced by IFN and 24 were repressed by more than two-fold, respectively. These IFN-regulated proteins represented multiple cellular functions including antiviral defense, immune response, cell metabolism, signal transduction, cell growth and cellular organization. To analyze this proteomics dataset, we utilized several systems-biology data-mining tools, including Gene Ontology via the GoMiner program and the Cytoscape bioinformatics platform. CONCLUSIONS: Integration of the quantitative proteomics with global protein interaction data using the Cytoscape platform led to the identification of several novel and liver-specific key regulatory components of the IFN response, which may be important in regulating the interplay between HCV, interferon and the host response to virus infection.

Cell Line, Tumor↗

Rationale and design of a large-scale trial using atrial natriuretic peptide (ANP) as an adjunct to percutaneous coronary intervention for ST-segment elevation acute myocardial infarction: Japan-Working groups of acute myocardial infarction for the reduction of Necrotic Damage by ANP (J-WIND-ANP).

BACKGROUND: The benefits of percutaneous coronary intervention (PCI) in acute myocardial infarction (AMI) are limited by reperfusion injury. In animal models, atrial natriuretic peptide (ANP) reduces infarct size, so the Japan-Working groups of acute myocardial Infarction for the reduction of Necrotic Damage by ANP (J-WIND-ANP) designed a prospective, randomized, multicenter study, to evaluate whether ANP as an adjunctive therapy for AMI reduces myocardial infarct size and improves regional wall motion. METHODS AND RESULTS: Twenty hospitals in Japan will participate in the J-WIND-ANP study. Patients with AMI who are candidates for PCI are randomly allocated to receive either intravenous ANP or placebo administration. The primary end-points are (1) estimated infarct size (Sigmacreatine kinase and troponin T) and (2) left ventricular function (left ventriculograms). Single nucleotide polymorphisms (SNPs) that may be associated with the function of ANP and susceptibility of AMI will be examined. Furthermore, a data mining method will be used to design the optimal combinational therapy for post-MI patients. CONCLUSIONS: J-WIND-ANP will provide important data on the effects of ANP as an adjunct to PCI for AMI and the SNPs information will open the field of tailor-made therapy. The optimal therapeutic drug combination will also be determined for post-MI patients.

Adult↗

Rough sets: a knowledge discovery technique for multifactorial medical outcomes.

Rough sets is a fairly new and promising technique for data mining and knowledge discovery from databases. This tutorial article presents the fundamentals of rough set theory in a nontechnical manner and outlines how the technique can be used to extract minimal if-then rules from tables of empirical data that either fully or approximately describe given example classifications. An example application for prediction of ambulation for patients with spinal cord injury is given. Because such rules are readily interpretable, they can be inspected to yield possible new insight into how various contributing factors interact and, thus, serve as hypothesis generators for further research. Additionally, the set of mined rules may function as a classifier of new, unseen cases.

Data Interpretation, Statistical↗

An environment for knowledge discovery in biology.

This paper describes a data mining environment for knowledge discovery in bioinformatics applications. The system has a generic kernel that implements the mining functions to be applied to input primary databases, with a warehouse architecture, of biomedical information. Both supervised and unsupervised classification can be implemented within the kernel and applied to data extracted from the primary database, with the results being suitably stored in a complex object database for knowledge discovery. The kernel also includes a specific high-performance library that allows designing and applying the mining functions in parallel machines. The experimental results obtained by the application of the kernel functions are reported.

Computational Biology↗

Recursive self-organizing network models.

Self-organizing models constitute valuable tools for data visualization, clustering, and data mining. Here, we focus on extensions of basic vector-based models by recursive computation in such a way that sequential and tree-structured data can be processed directly. The aim of this article is to give a unified review of important models recently proposed in literature, to investigate fundamental mathematical properties of these models, and to compare the approaches by experiments. We first review several models proposed in literature from a unifying perspective, thereby making use of an underlying general framework which also includes supervised recurrent and recursive models as special cases. We shortly discuss how the models can be related to different neuron lattices. Then, we investigate theoretical properties of the models in detail: we explicitly formalize how structures are internally stored in different context models and which similarity measures are induced by the recursive mapping onto the structures. We assess the representational capabilities of the models, and we shortly discuss the issues of topology preservation and noise tolerance. The models are compared in an experiment with time series data. Finally, we add an experiment for one context model for tree-structured data to demonstrate the capability to process complex structures.

Artifacts↗

Database development in toxicogenomics: issues and efforts.

The marriage of toxicology and genomics has created not only opportunities but also novel informatics challenges. As with the larger field of gene expression analysis, toxicogenomics faces the problems of probe annotation and data comparison across different array platforms. Toxicogenomics studies are generally built on standard toxicology studies generating biological end point data, and as such, one goal of toxicogenomics is to detect relationships between changes in gene expression and in those biological parameters. These challenges are best addressed through data collection into a well-designed toxicogenomics database. A successful publicly accessible toxicogenomics database will serve as a repository for data sharing and as a resource for analysis, data mining, and discussion. It will offer a vehicle for harmonizing nomenclature and analytical approaches and serve as a reference for regulatory organizations to evaluate toxicogenomics data submitted as part of registrations. Such a database would capture the experimental context of in vivo studies with great fidelity such that the dynamics of the dose response could be probed statistically with confidence. This review presents the collaborative efforts between the European Molecular Biology Laboratory-European Bioinformatics Institute ArrayExpress, the International Life Sciences Institute Health and Environmental Science Institute, and the National Institute of Environmental Health Sciences National Center for Toxigenomics Chemical Effects in Biological Systems knowledge base. The goal of this collaboration is to establish public infrastructure on an international scale and examine other developments aimed at establishing toxicogenomics databases. In this review we discuss several issues common to such databases: the requirement for identifying minimal descriptors to represent the experiment, the demand for standardizing data storage and exchange formats, the challenge of creating standardized nomenclature and ontologies to describe biological data, the technical problems involved in data upload, the necessity of defining parameters that assess and record data quality, and the development of standardized analytical approaches.

Animals↗

Contextual weighting for Support Vector Machines in literature mining: an application to gene versus protein name disambiguation.

BACKGROUND: The ability to distinguish between genes and proteins is essential for understanding biological text. Support Vector Machines (SVMs) have been proven to be very efficient in general data mining tasks. We explore their capability for the gene versus protein name disambiguation task. RESULTS: We incorporated into the conventional SVM a weighting scheme based on distances of context words from the word to be disambiguated. This weighting scheme increased the performance of SVMs by five percentage points giving performance better than 85% as measured by the area under ROC curve and outperformed the Weighted Additive Classifier, which also incorporates the weighting, and the Naive Bayes classifier. CONCLUSION: We show that the performance of SVMs can be improved by the proposed weighting scheme. Furthermore, our results suggest that in this study the increase of the classification performance due to the weighting is greater than that obtained by selecting the underlying classifier or the kernel part of the SVM.

Algorithms↗

A Bayesian neural network method for adverse drug reaction signal generation.

OBJECTIVE: The database of adverse drug reactions (ADRs) held by the Uppsala Monitoring Centre on behalf of the 47 countries of the World Health Organization (WHO) Collaborating Programme for International Drug Monitoring contains nearly two million reports. It is the largest database of this sort in the world, and about 35,000 new reports are added quarterly. The task of trying to find new drug-ADR signals has been carried out by an expert panel, but with such a large volume of material the task is daunting. We have developed a flexible, automated procedure to find new signals with known probability difference from the background data. METHOD: Data mining, using various computational approaches, has been applied in a variety of disciplines. A Bayesian confidence propagation neural network (BCPNN) has been developed which can manage large data sets, is robust in handling incomplete data, and may be used with complex variables. Using information theory, such a tool is ideal for finding drug-ADR combinations with other variables, which are highly associated compared to the generality of the stored data, or a section of the stored data. The method is transparent for easy checking and flexible for different kinds of search. RESULTS: Using the BCPNN, some time scan examples are given which show the power of the technique to find signals early (captopril-coughing) and to avoid false positives where a common drug and ADRs occur in the database (digoxin-acne; digoxin-rash). A routine application of the BCPNN to a quarterly update is also tested, showing that 1004 suspected drug-ADR combinations reached the 97.5% confidence level of difference from the generality. Of these, 307 were potentially serious ADRs, and of these 53 related to new drugs. Twelve of the latter were not recorded in the CD editions of The physician's Desk Reference or Martindale's Extra Pharmacopoea and did not appear in Reactions Weekly online. CONCLUSION: The results indicate that the BCPNN can be used in the detection of significant signals from the data set of the WHO Programme on International Drug Monitoring. The BCPNN will be an extremely useful adjunct to the expert assessment of very large numbers of spontaneously reported ADRs.

Adverse Drug Reaction Reporting Systems↗

Filtering erroneous protein annotation.

MOTIVATION: Automatically generated annotation on protein data of UniProt (Universal Protein Resource) is planned to be publicly available on the UniProt web pages in April 2004. It is expected that the data content of over 500,000 protein entries in the TrEMBL section will be enhanced by the output of an automated annotation pipeline. However, a part of the automatically added data will be erroneous, as are parts of the information coming from other sources. We present a post-processing system called Xanthippe that is based on a simple exclusion mechanism and a decision tree approach using the C4.5 data-mining algorithm. RESULTS: It is shown that Xanthippe detects and flags a large part of the annotation errors and considerably increases the reliability of both automatically generated data and annotation from other sources. As a cross-validation to Swiss-Prot shows, errors in protein descriptions, comments and keywords are successfully filtered out. Xanthippe is a contradictive application that can be combined seamlessly with predictive systems. It can be used either to improve the precision of automated annotation at a constant level of recall or increase the recall at a constant level of precision. AVAILABILITY: The application of the Xanthippe rules can be browsed at http://www.ebi.uniprot.org/

Algorithms↗