Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Harvesting chemical information from the Internet using a distributed approach: ChemXtreme.

The Internet is a comprehensive resource of chemical information which is at the same time largely unstructured. It provides a wealth of scientific information such as experimental data and requires a suitable automated data mining and analysis tool for its meaningful exploration. The Java based software presented here, ChemXtreme, is developed for harvesting chemical information from the Internet employing the Google API in combination with a distributed client/server text analysis architecture based on JavaRMI. It represents the first and until now the only toolkit for automated structured data retrieval from the Internet which is itself open source. ChemXtreme employs the "search the search engine" strategy, where the URLs returned from the search engine are analyzed further via textual pattern analysis. This process resembles the manual analysis of the hit list, where relevant data are captured and, by means of human intervention, are mined into a format suitable for further analysis. ChemXtreme on the other hand transforms chemical information automatically into a structured format suitable for storage in databases and further analysis and also provides links to the original information source. The query data retrieved from the search engine by the server is encoded, encrypted, and compressed and then sent to all the participating active clients in the network for parsing. Relevant information identified by the clients on the retrieved Web sites is sent back to the server, verified, and added to the database for data mining and further analysis. The distributed further analysis of URLs in a client/server architecture scales very favorably, thus producing only minimal overhead.

Chemistry↗

Analyzing microarray data using quantitative association rules.

MOTIVATION: We tackle the problem of finding regularities in microarray data. Various data mining tools, such as clustering, classification, Bayesian networks and association rules, have been applied so far to gain insight into gene-expression data. Association rule mining techniques used so far work on discretizations of the data and cannot account for cumulative effects. In this paper, we investigate the use of quantitative association rules that can operate directly on numeric data and represent cumulative effects of variables. Technically speaking, this type of quantitative association rules based on half-spaces can find non-axis-parallel regularities. RESULTS: We performed a variety of experiments testing the utility of quantitative association rules for microarray data. First of all, the results should be statistically significant and robust against fluctuations in the data. Next, the approach should be scalable in the number of variables, which is important for such high-dimensional data. Finally, the rules should make sense biologically and be sufficiently different from rules found in regular association rule mining working with discretizations. In all of these dimensions, the proposed approach performed satisfactorily. Therefore, quantitative association rules based on half-spaces should be considered as a tool for the analysis of microarray gene-expression data. AVAILABILITY: The code is available from the authors on request.

Algorithms↗

Mining protein sequences for motifs.

We use methods from Data Mining and Knowledge Discovery to design an algorithm for detecting motifs in protein sequences. The algorithm assumes that a motif is constituted by the presence of a "good" combination of residues in appropriate locations of the motif. The algorithm attempts to compile such good combinations into a "pattern dictionary" by processing an aligned training set of protein sequences. The dictionary is subsequently used to detect motifs in new protein sequences. Statistical significance of the detection results are ensured by statistically determining the various parameters of the algorithm. Based on this approach, we have implemented a program called GYM. The Helix-Turn-Helix motif was used as a model system on which to test our program. The program was also extended to detect Homeodomain motifs. The detection results for the two motifs compare favorably with existing programs. In addition, the GYM program provides a lot of useful information about a given protein sequence.

Algorithms↗

Microarray data analysis and mining.

DNA microarray is an innovative technology for obtaining information on gene function. Because it is a high-throughput method, computational tools are essential in data analysis and mining to extract the knowledge from experimental results. Filtering procedures and statistical approaches are frequently combined to identify differentially expressed genes. However, obtaining a list of differentially expressed genes is only the starting point because an important step is the integration of differential expression profiles in a biological context, which is a hot topic in data mining. In this chapter an integrated approach of filtering and statistical validation to select trustable differentially expressed genes is described together with a brief introduction on data mining focusing on the classification of co-regulated genes on the basis of their biological function.

Cluster Analysis↗

Mining HIV dynamics using independent component analysis.

MOTIVATION: We implement a data mining technique based on the method of Independent Component Analysis (ICA) to generate reliable independent data sets for different HIV therapies. We show that this technique takes advantage of the ICA power to eliminate the noise generated by artificial interaction of HIV system dynamics. Moreover, the incorporation of the actual laboratory data sets into the analysis phase offers a powerful advantage when compared with other mathematical procedures that consider the general behavior of HIV dynamics. RESULTS: The ICA algorithm has been used to generate different patterns of the HIV dynamics under different therapy conditions. The Kohonen Map has been used to eliminate redundant noise in each pattern to produce a reliable data set for the simulation phase. We show that under potent antiretroviral drugs, the value of the CD4+ cells in infected persons decreases gradually by about 11% every 100 days and the levels of the CD8+ cells increase gradually by about 2% every 100 days. AVAILABILITY: Executable code and data libraries are available by contacting the corresponding author. IMPLEMENTATION: Mathematica 4 has been used to simulate the suggested model. A Pentium III or higher platform is recommended.

Algorithms↗

Survey of microarray technologies suitable to elucidate transcriptional networks as exemplified by studying KRAB zinc finger gene families.

Current microarray systems are suitable to monitor genome-wide expression patterns, to detect single-nucleotide polymorphisms (SNP), to identify target genes of transcription factors and DNA-protein interaction sites thereof as well as to determine genomic sites that are modified by methylation of CpG islands. In this review, advantages and limitations of individual microarray technologies are presented as well as experiences from ongoing studies on KRAB zinc finger gene families are taken to exemplify how different microarray approaches are applicable to elucidate complex transcriptional networks of gene regulation. However, bioinformaticians should be aware that each microarray technology has limitations in its sensitivity and selectivity that has to be taken into account once data mining on comprehensive genome-wide microarray data is conducted. In many cases, microarray results are the initial step to identify target genes of interest and to study the molecular regulation of biological processes thereof followed and validated by complementary proteome, metabolome or toponome analysis. Thus, microarray technologies can be considered a reliable approach for determining gene functions that might be modulated by electromagnetic fields.

Gene Expression Profiling↗

Automated pharmacophore identification for large chemical data sets.

The identification of three-dimensional pharmacophores from large, heterogeneous data sets is still an unsolved problem. We developed a novel program, SCAMPI (statistical classification of activities of molecules for pharmacophore identification), for this purpose by combining a fast conformation search with recursive partitioning, a data-mining technique, which can easily handle large data sets. The pharmacophore identification process is designed to run recursively, and the conformation spaces are resampled under the constraints of the evolving pharmacophore model. This program is capable of deriving pharmacophores from a data set of 1000-2000 compounds, with thousands of conformations generated for each compound and in less than 1 day of computational time. For two test data sets, the identified pharmacophores are consistent with the known results from the literature.

Angiotensin-Converting Enzyme Inhibitors↗

The construction of meaning in computational integrative biology.

It is the contention of this paper that current data mining work in bioinformatics tends to emphasize data representation to the neglect of another essential aspect of biological systems, namely dynamics. This results in a divorce of both the enterprise and the teaching of bioinformatics from its central aim of meaning-construction. The paper argues that this neglect of dynamics is rooted in an information-processing view of cognitive psychology, and needs to be complemented by a more narrative perspective which emphasizes the explanation, rather than the mere description, of observed patterns in data. This increased emphasis on explanatory narrative in the form of dynamical modeling leads to both a deeper understanding of biological information and a more invigorating approach to the teaching of bioinformatics. The paper presents a cross-curricular teaching framework for a first-year undergraduate course in bioinformatic dynamical modeling which is based around the use of narrative plots.

Communication↗

PDA v.2: improving the exploration and estimation of nucleotide polymorphism in large datasets of heterogeneous DNA.

Pipeline Diversity Analysis (PDA) is an open-source, web-based tool that allows the exploration of polymorphism in large datasets of heterogeneous DNA sequences, and can be used to create secondary polymorphism databases for different taxonomic groups, such as the Drosophila Polymorphism Database (DPDB). A new version of the pipeline presented here, PDA v.2, incorporates substantial improvements, including new methods for data mining and grouping sequences, new criteria for data quality assessment and a better user interface. PDA is a powerful tool to obtain and synthesize existing empirical evidence on genetic diversity in any species or species group. PDA v.2 is available on the web at http://pda.uab.es/.

Algorithms↗

Automated variable weighting in k-means type clustering.

This paper proposes a k-means type clustering algorithm that can automatically calculate variable weights. A new step is introduced to the k-means clustering process to iteratively update variable weights based on the current partition of data and a formula for weight calculation is proposed. The convergency theorem of the new clustering process is given. The variable weights produced by the algorithm measure the importance of variables in clustering and can be used in variable selection in data mining applications where large and complex real data are often involved. Experimental results on both synthetic and real data have shown that the new algorithm outperformed the standard k-means type algorithms in recovering clusters in data.

Algorithms↗

Considerations for the future development of virtual technology as a rehabilitation tool.

BACKGROUND: Virtual environments (VE) are a powerful tool for various forms of rehabilitation. Coupling VE with high-speed networking [Tele-Immersion] that approaches speeds of 100 Gb/sec can greatly expand its influence in rehabilitation. Accordingly, these new networks will permit various peripherals attached to computers on this network to be connected and to act as fast as if connected to a local PC. This innovation may soon allow the development of previously unheard of networked rehabilitation systems. Rapid advances in this technology need to be coupled with an understanding of how human behavior is affected when immersed in the VE. METHODS: This paper will discuss various forms of VE that are currently available for rehabilitation. The characteristic of these new networks and examine how such networks might be used for extending the rehabilitation clinic to remote areas will be explained. In addition, we will present data from an immersive dynamic virtual environment united with motion of a posture platform to record biomechanical and physiological responses to combined visual, vestibular, and proprioceptive inputs. A 6 degree-of-freedom force plate provides measurements of moments exerted on the base of support. Kinematic data from the head, trunk, and lower limb was collected using 3-D video motion analysis. RESULTS: Our data suggest that when there is a confluence of meaningful inputs, neither vision, vestibular, or proprioceptive inputs are suppressed in healthy adults; the postural response is modulated by all existing sensory signals in a non-additive fashion. Individual perception of the sensory structure appears to be a significant component of the response to these protocols and underlies much of the observed response variability. CONCLUSION: The ability to provide new technology for rehabilitation services is emerging as an important option for clinicians and patients. The use of data mining software would help analyze the incoming data to provide both the patient and the therapist with evaluation of the current treatment and modifications needed for future therapies. Quantification of individual perceptual styles in the VE will support development of individualized treatment programs. The virtual environment can be a valuable tool for therapeutic interventions that require adaptation to complex, multimodal environments.

Journal Article↗

Recursive partitioning analysis of complex disease pharmacogenetic studies. I. Motivation and overview.

Identifying genetic variation predictive of important phenotypes, including disease susceptibility, drug efficacy, and adverse events, is a challenging task, and theory and computer science work is being carried out in an attempt to tackle this issue. For many important diseases, such as diabetes, schizophrenia, and depression, the etiology is complex; either the disease is a result of several multiple mechanisms or is caused by an interaction among multiple genes or gene-environment interactions, or both. There is a need for statistical methods to deal with the large, complex data sets that will be used to disentangle these diseases. Each putative genetic polymorphism can be tested for association sequentially. The most difficult problem, however, is the identification of combinations of polymorphisms or genetic markers with increased predictive characteristics. Data from clinical trials, where patients with a particular disease are treated with certain drugs, can be retrospectively assembled using a case-control design. Such data will typically include treatment assignment, demographics, medical history, and genotypes for a large number of genetic markers. The number of variables in such data is expected to be much larger than the number of subjects. This report focuses on some of the methods being employed to deal with this complex data and covers, in some detail, a data-mining method--recursive partitioning--to analyze such data. The methods are demonstrated using a complex simulated data set, as there are few available public data sets. This explication of recursive partitioning should provide researchers with a better idea of the current available analysis techniques, in order to allow them to plan their experiments more effectively.

Biomedical Research↗

Uncertainty and decisions in medical informatics.

This paper presents a tutorial introduction to the handling of uncertainty and decision-making in medical reasoning systems. It focuses on the central role of uncertainty in all of medicine and identifies the major themes that arise in research papers. It then reviews simple Bayesian formulations of the problem and pursues the generalization to the Bayesian network methods that are popular today. Decision making is presented from the decision analysis viewpoint, with brief mention of recently-developed methods. The paper concludes with review of more abstract characterization of uncertainty, and anticipates the growing importance of analytic and "data mining" techniques as growing amounts of clinical data become widely available.

Bayes Theorem↗

The cell-centered database: a database for multiscale structural and protein localization data from light and electron microscopy.

The creation of structured shared data repositories for molecular data in the form of web-accessible databases like GenBank has been a driving force behind the genomic revolution. These resources serve not only to organize and manage molecular data being created by researchers around the globe, but also provide the starting point for data mining operations to uncover interesting information present in the large amount of sequence and structural data. To realize the full impact of the genomic and proteomic efforts of the last decade, similar resources are needed for structural and biochemical complexity in biological systems beyond the molecular level, where proteins and macromolecular complexes are situated within their cellular and tissue environments. In this review, we discuss our efforts in the development of neuroinformatics resources for managing and mining cell level imaging data derived from light and electron microscopy. We describe the main features of our web-accessible database, the Cell Centered Database (CCDB; http://ncmir.ucsd.edu/CCDB/), designed for structural and protein localization information at scales ranging from large expanses of tissue to cellular microdomains with their associated macromolecular constituents. The CCDB was created to make 3D microscopic imaging data available to the scientific community and to serve as a resource for investigating structural and macromolecular complexity of cells and tissues, particularly in the rodent nervous system.

Brain↗

Mining frequent patterns for AMP-activated protein kinase regulation on skeletal muscle.

BACKGROUND: AMP-activated protein kinase (AMPK) has emerged as a significant signaling intermediary that regulates metabolisms in response to energy demand and supply. An investigation into the degree of activation and deactivation of AMPK subunits under exercise can provide valuable data for understanding AMPK. In particular, the effect of AMPK on muscle cellular energy status makes this protein a promising pharmacological target for disease treatment. As more AMPK regulation data are accumulated, data mining techniques can play an important role in identifying frequent patterns in the data. Association rule mining, which is commonly used in market basket analysis, can be applied to AMPK regulation. RESULTS: This paper proposes a framework that can identify the potential correlation, either between the state of isoforms of alpha, beta and gamma subunits of AMPK, or between stimulus factors and the state of isoforms. Our approach is to apply item constraints in the closed interpretation to the itemset generation so that a threshold is specified in terms of the amount of results, rather than a fixed threshold value for all itemsets of all sizes. The derived rules from experiments are roughly analyzed. It is found that most of the extracted association rules have biological meaning and some of them were previously unknown. They indicate direction for further research. CONCLUSION: Our findings indicate that AMPK has a great impact on most metabolic actions that are related to energy demand and supply. Those actions are adjusted via its subunit isoforms under specific physical training. Thus, there are strong co-relationships between AMPK subunit isoforms and exercises. Furthermore, the subunit isoforms are correlated with each other in some cases. The methods developed here could be used when predicting these essential relationships and enable an understanding of the functions and metabolic pathways regarding AMPK.

AMP-Activated Protein Kinases↗

Cytomics--new technologies: towards a human cytome project.

BACKGROUND: Molecular cell systems research (cytomics) aims at the understanding of the molecular architecture and functionality of cell systems (cytomes) by single-cell analysis in combination with exhaustive bioinformatic knowledge extraction. In this way, loss of information as a consequence of molecular averaging by cell or tissue homogenisation is avoided. PROGRESS: The cytomics concept has been significantly advanced by a multitude of current developments. Amongst them are confocal and laser scanning microscopy, multiphoton fluorescence excitation, spectral imaging, fluorescence resonance energy transfer (FRET), fast imaging in flow, optical stretching in flow, and miniaturised flow and image cytometry within laboratories on a chip or laser microdissection, as well as the use of bead arrays. In addition, biomolecular analysis techniques like tyramide signal amplification, single-cell polymerase chain reaction (PCR), and the labelling of biomolecules by quantum dots, magnetic nanobeads, or aptamers open new horizons of sensitivity and molecular specificity at the single-cell level. Data sieving or data mining of the vast amounts of collected multiparameter data for exhaustive multilevel bioinformatic knowledge extraction avoids the inadvertent loss of information from unknown molecular relations being inaccessible to an a priori hypothesis. CHALLENGE: It seems important to address the challenge of a human cytome project using hypothesis-driven molecular information collection from disease associated cell systems, supplemented by systematic and exhaustive knowledge extraction. This will allow the description of the molecular setup of normal and abnormal cell systems within a relational knowledge system, permitting the standardised discrimination of abnormal cell states in disease. As one of the consequences, individualised predictions of further disease course in patients (predictive medicine by cytomics) by characteristic discriminatory data patterns will permit individualised therapies, identification of new pharmaceutical targets, and establishment of a standardised framework of relevant molecular alterations in disease. This special issue of Cytometry, on new technologies in cytomics, focuses on prominent examples of this presently fast-moving scientific field, and represents one of the preconditions for the formulation of a human cytome project.

Cell Biology↗

The Human Genome Project and the future of diagnostics, treatment and prevention.

The Human Genome Project, the mapping of our 30,000-50,000 genes and the sequencing of all of our DNA, will have major impact on biomedical research and the whole of therapeutic and preventive health care. The tracing of genetic diseases to their molecular causes is rapidly expanding diagnostic and preventive options. The increased insights into molecular pathways, gained from high-throughput 'functional genomics', using DNA-chip and protein-chip approaches and specially designed animal model systems, will open great prospects for pharmacological and genetic therapies. Powerful bioinformatics and biostatistics will further improve our pattern recognition and accelerate progress. A rapidly expanding area of high expectations is that of 'pharmacogenomics': the design of more effective drugs with lower toxicity through tailoring of drug treatment to individual, genetically determined differences in drug metabolism. Not only will this decrease the cost of health care through reduction of adverse drug reactions, but a better stratification of populations will also provide more statistical power farther upstream in drug trials. However, the optimal benefits from the current explosion of 'data mining' will only be realized when the basic data are made and kept publicly accessible, while at the same time safeguarding the protection of intellectual property arising from downstream inventions. This is one of the goals of HUGO, the international Human Genome Organization, established 13 years ago to assist coordination of data acquisition and exchange and societal implementation of the genome project. Additional points of attention in this historic endeavour are the prevention of stigmatization and discrimination and the safeguarding of a worldwide balance in the contribution by--and benefits to--different populations, while respecting the diversity in cultures and traditions.

Ethics, Medical↗

Discovering motif pairs at interaction sites from protein sequences on a proteome-wide scale.

MOTIVATION: Protein-protein interaction, mediated by protein interaction sites, is intrinsic to many functional processes in the cell. In this paper, we propose a novel method to discover patterns in protein interaction sites. We observed from protein interaction networks that there exist a kind of significant substructures called interacting protein group pairs, which exhibit an all-versus-all interaction between the two protein-sets in such a pair. The full-interaction between the pair indicates a common interaction mechanism shared by the proteins in the pair, which can be referred as an interaction type. Motif pairs at the interaction sites of the protein group pairs can be used to represent such interaction type, with each motif derived from the sequences of a protein group by standard motif discovery algorithms. The systematic discovery of all pairs of interacting protein groups from large protein interaction networks is a computationally challenging problem. By a careful and sophisticated problem transformation, the problem is solved using efficient algorithms for mining frequent patterns, a problem extensively studied in data mining. RESULTS: We found 5349 pairs of interacting protein groups from a yeast interaction dataset. The expected value of sequence identity within the groups is only 7.48%, indicating non-homology within these protein groups. We derived 5343 motif pairs from these group pairs, represented in the form of blocks. Comparing our motifs with domains in the BLOCKS and PRINTS databases, we found that our blocks could be mapped to an average of 3.08 correlated blocks in these two databases. The mapped blocks occur 4221 out of total 6794 domains (protein groups) in these two databases. Comparing our motif pairs with iPfam consisting of 3045 interacting domain pairs derived from PDB, we found 47 matches occurring in 105 distinct PDB complexes. Comparing with another putative domain interaction database InterDom, we found 203 matches. AVAILABILITY: http://research.i2r.a-star.edu.sg/BindingMotifPairs/resources. SUPPLEMENTARY INFORMATION: http://research.i2r.a-star.edu.sg/BindingMotifPairs and Bioinformatics online.

Algorithms↗