Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Selected techniques for data mining in medicine.

Widespread use of medical information systems and explosive growth of medical databases require traditional manual data analysis to be coupled with methods for efficient computer-assisted analysis. This paper presents selected data mining techniques that can be applied in medicine, and in particular some machine learning techniques including the mechanisms that make them better suited for the analysis of medical databases (derivation of symbolic rules, use of background knowledge, sensitivity and specificity of induced descriptions). The importance of the interpretability of results of data analysis is discussed and illustrated on selected medical applications.

Adult↗

MineBlast: a literature presentation service supporting protein annotation by data mining of BLAST results.

MineBlast is a web service for literature search and presentation based on data-mining results received from UniProt. Users can submit a simple list of protein sequences via a web-based interface. MineBlast performs a BLASTP search in UniProt to identify names and synonyms based on homologous proteins and subsequently queries PubMed, using combined search terms inorder to find and present relevant literature.

Database Management Systems↗

Shotgun proteomics of cyanobacteria--applications of experimental and data-mining techniques.

Cyanobacteria are photosynthetic bacteria notable for their ability to produce hydrogen and a variety of interesting secondary metabolites. As a result of the growing number of completed cyanobacterial genome projects, the development of post-genomics analysis for this important group has been accelerating. DNA microarrays and classical two-dimensional gel electrophoresis (2DE) were the first technologies applied in such analyses. In many other systems, 'shotgun' proteomics employing multi-dimensional liquid chromatography and tandem mass spectrometry has proven to be a powerful tool. However, this approach has been relatively under-utilized in cyanobacteria. This study assesses progress in cyanobacterial shotgun proteomics to date, and adds a new perspective by developing a protocol for the shotgun proteomic analysis of the filamentous cyanobacterium Anabaena variabilis ATCC 29413, a model for N(2) fixation. Using approaches for enhanced protein extraction, 646 proteins were identified, which is more than double the previous results obtained using 2DE. Notably, the improved extraction method and shotgun approach resulted in a significantly higher representation of basic and hydrophobic proteins. The use of protein bioinformatics tools to further mine these shotgun data is illustrated through the application of PSORTb for localization, the grand average hydropathy (GRAVY) index for hydrophobicity, LipoP for lipoproteins and the exponentially modified protein abundance index (emPAI) for abundance. The results are compared with the most well-studied cyanobacterium, Synechocystis sp. PCC 6803. Some general issues in shotgun proteome identification and quantification are then addressed.

Computational Biology↗

XML-based visual data mining in medicine.

Medical databases in general are characterized by a high degree of complexity in terms of quantity of items, number of parameter values and data types (free text, categorical, numerical and other). Substantial domain knowledge is required for adequate formalization of medical entities. In this context we developed medical database plot (mdplot), a data mining tool to visualize both structure and quality of data in medical databases to identify items suitable for evaluation. Data models are provided in XML format. Missing data is identified to enable targeted efforts to improve data quality prior to analysis. Database items are classified as 1:1- related to the patient (i.e. variables are collected once per patient) and 1:n related. mdplot provides a list of all classes contained in a database, the number of records each and a condensed bar chart for semi-quantitative description of completeness according to four types of items: categorical, numerical, text and other. All items in a category are grouped from left to right, the height of each bar represents the proportion of non-missing values with respect to the total number of records in the class; thus the amount of content in a specific class is visualized. By selection of a specific class, a detailed description of it is provided including mean completeness in each item category as well as number of values per item. The new methodology was applied to a cardiological research database consisting of 619 items on 88 patients.

Atrial Fibrillation↗

[In silico data mining of the human programmed cell death 5 (PDCD5) sequences].

OBJECTIVE: To lay foundation for the functional studies of programmed cell death 5 (PDCD5) and develop new technical pathway for bioinformatics analysis of human functional genes. METHODS: Using PDCD5 as the target molecule, intensive bioinformatics analysis of the nucleic acid and protein sequences were conducted. Data mining and comprehensive analysis by sequence against database similarity searching, ortholog structure comparison, expression profile analysis and gene "neighbor" listing were performed. RESULTS: Two human putative pseudogenes on chromosomes 12 and 5, and one mouse putative pseudogene on chromosome 1 were identified. The methanobacterium thermoautotrophicum ortholog was classified as the same fold as ubiquitin and ribosomal protein S13. The C. elegans ortholog, ubiquitin and IAP (inhibitor of apoptosis proteins) belonged to the same expression profile cluster. This cluster was related to biosynthesis and protein synthesis. PDCD5 orthologs in various genomes were adjacent to various ribosomal proteins on the chromosome. CONCLUSION: The human genome contains at least two processed pseudogenes of PDCD5. Besides the relationship with cell apoptosis, PDCD5 is predicted to have functional relationship with ubiquitin and participate in the translation regulation.

Amino Acid Sequence↗

Data mining the NCI cancer cell line compound GI(50) values: identifying quinone subtypes effective against melanoma and leukemia cell classes.

Using data mining techniques, we have studied a subset (1400) of compounds from the large public National Cancer Institute (NCI) compounds data repository. We first carried out a functional class identity assignment for the 60 NCI cancer testing cell lines via hierarchical clustering of gene expression data. Comprised of nine clinical tissue types, the 60 cell lines were placed into six classes-melanoma, leukemia, renal, lung, and colorectal, and the sixth class was comprised of mixed tissue cell lines not found in any of the other five classes. We then carried out supervised machine learning, using the GI(50) values tested on a panel of 60 NCI cancer cell lines. For separate 3-class and 2-class problem clustering, we successfully carried out clear cell line class separation at high stringency, p < 0.01 (Bonferroni corrected t-statistic), using feature reduction clustering algorithms embedded in RadViz, an integrated high dimensional analytic and visualization tool. We started with the 1400 compound GI(50) values as input and selected only those compounds, or features, significant in carrying out the classification. With this approach, we identified two small sets of compounds that were most effective in carrying out complete class separation of the melanoma, non-melanoma classes and leukemia, non-leukemia classes. To validate these results, we showed that these two compound sets' GI(50) values were highly accurate classifiers using five standard analytical algorithms. One compound set was most effective against the melanoma class cell lines (14 compounds), and the other set was most effective against the leukemia class cell lines (30 compounds). The two compound classes were both significantly enriched in two different types of substituted p-quinones. The melanoma cell line class of 14 compounds was comprised of 11 compounds that were internal substituted p-quinones, and the leukemia cell line class of 30 compounds was comprised of 6 compounds that were external substituted p-quinones. Attempts to subclassify melanoma or leukemia cell lines based upon their clinical cancer subtype met with limited success. For example, using GI(50) values for the 30 compounds we identified as effective against all leukemia cell lines, we could subclassify acute lymphoblastic leukemia (ALL) origin cell lines from non-ALL leukemia origin cell lines without significant overlap from non-leukemia cell lines. Based upon clustering using GI(50) values for the 60 cancer cell lines laid out by the RadViz algorithm, these two compound subsets did not overlap with clusters containing any of the NCI's 92 compounds of known mechanism of action, a few of which are quinones. Given their structural patterns, the two p-quinone subtypes we identified would clearly be expected to possess different redox potentials/substrate specificities for enzymatic reduction in vivo. These two p-quinone subtypes represent valuable information that may be used in the elucidation of pharmacophores for the design of compounds to treat these two cancer tissue types in the clinic.

Algorithms↗

The microarray explorer tool for data mining of cDNA microarrays: application for the mammary gland.

The Microarray Explorer (MAExplorer) is a versatile Java-based data mining bioinformatic tool for analyzing quantitative cDNA expression profiles across multiple microarray platforms and DNA labeling systems. It may be run as either a stand-alone application or as a Web browser applet over the Internet. With this program it is possible to (i) analyze the expression of individual genes, (ii) analyze the expression of gene families and clusters, (iii) compare expression patterns and (iv) directly access other genomic databases for clones of interest. Data may be downloaded as required from a Web server or in the case of the stand-alone version, reside on the user's computer. Analyses are performed in real-time and may be viewed and directly manipulated in images, reports, scatter plots, histograms, expression profile plots and cluster analyses plots. A key feature is the clone data filter for constraining a working set of clones to those passing a variety of user-specified logical and statistical tests. Reports may be generated with hypertext Web access to UniGene, GenBank and other Internet databases for sets of clones found to be of interest. Users may save their explorations on the Web server or local computer and later recall or share them with other scientists in this groupware Web environment. The emphasis on direct manipulation of clones and sets of clones in graphics and tables provides a high level of interaction with the data, making it easier for investigators to test ideas when looking for patterns. We have used the MAExplorer to profile gene expression patterns of 1500 duplicated genes isolated from mouse mammary tissue. We have identified genes that are preferentially expressed during pregnancy and during lactation. One gene we identified, carbonic anhydrase III, is highly expressed in mammary tissue from virgin and pregnant mice and in gene knock-out mice with underdeveloped mammary epithelium. Other genes, which include those encoding milk proteins, are preferentially expressed during lactation.

Animals↗

Data mining for seeking accurate quantitative relationship between molecular structure and GC retention indices of alkanes by projection pursuit.

Primary data mining on alkanes for seeking accurate quantitative relationship between molecular structure and retention indices of gas chromatography is developed in this paper. Based on the results obtained from projection pursuit (PP), a new variable named class distance variable, which essentially describes the branching structure of the alkanes, is proposed. With the help of the new variable, both fitting and prediction accuracy of the regression model can be dramatically improved. The results obtained in this work show that the technique of PP developed in statistics is a quite promising tool for seeking accurate quantitative structure-activity relationship (QSAR) and/or quantitative structure-property relationship (QSPR) researches.

Journal Article↗

Data mining goes multidimensional.

The success of a healthcare organization depends on its ability to acquire, store, analyze and compare data across many parts of the enterprise, by many individuals. While relational databases have been around since the 1970s, their two-dimensional structure has limited--or made impossible--the kind of cross-dimensional trend analysis so necessary to healthcare today. Enter online analytical processing (OLAP), in which servers store data in multiple dimensions, opening a world of opportunity for data-mining across the enterprise. In this issue of HEALTHCARE INFORMATICS, we feature our first report from the National Software Testing Laboratories (NSTL) about technologies that will change the way healthcare does business. A division of The McGraw-Hill Companies, NSTL is an independent software and hardware testing lab offering services that include compatibility testing, bug testing, comparison testing, documentation evaluation and usability.

Computer User Training↗

Comparative genomics using data mining tools.

We have analysed the genomes of representatives of three kingdoms of life, namely, archaea, eubacteria and eukaryota using data mining tools based on compositional analyses of the protein sequences. The representatives chosen in this analysis were Methanococcus jannaschii, Haemophilus influenzae and Saccharomyces cerevisiae. We have identified the common and different features between the three genomes in the protein evolution patterns. M. jannaschii has been seen to have a greater number of proteins with more charged amino acids whereas S. cerevisiae has been observed to have a greater number of hydrophilic proteins. Despite the differences in intrinsic compositional characteristics between the proteins from the different genomes we have also identified certain common characteristics. We have carried out exploratory Principal Component Analysis of the multivariate data on the proteins of each organism in an effort to classify the proteins into clusters. Interestingly, we found that most of the proteins in each organism cluster closely together, but there are a few 'outliers'. We focus on the outliers for the functional investigations, which may aid in revealing any unique features of the biology of the respective organisms

Archaeal Proteins↗

Data mining and machine learning techniques for the identification of mutagenicity inducing substructures and structure activity relationships of noncongeneric compounds.

This paper explores the utility of data mining and machine learning algorithms for the induction of mutagenicity structure-activity relationships (SARs) from noncongeneric data sets. We compare (i) a newly developed algorithm (MOLFEA) for the generation of descriptors (molecular fragments) for noncongeneric compounds with traditional SAR approaches (molecular properties) and (ii) different machine learning algorithms for the induction of SARs from these descriptors. In addition we investigate the optimal parameter settings for these programs and give an exemplary interpretation of the derived models. The predictive accuracies of models using MOLFEA derived descriptors is approximately 10-15%age points higher than those using molecular properties alone. Using both types of descriptors together does not improve the derived models. From the applied machine learning techniques the rule learner PART and support vector machines gave the best results, although the differences between the learning algorithms are only marginal. We were able to achieve predictive accuracies up to 78% for 10-fold cross-validation. The resulting models are relatively easy to interpret and usable for predictive as well as for explanatory purposes.

Algorithms↗

Data mining for seeking an accurate quantitative relationship between molecular structure and GC retention indices of alkenes by projection pursuit.

Primary data mining on alkenes for seeking an accurate quantitative relationship between the molecular structure and retention indices of gas chromatography is developed in this paper. Based on the results obtained from projection pursuit, all alkenes investigated show an interesting classification. Thus, a new variable named class distance variable of alkenes, which essentially describes information about the branch, position of the double bonds, the number of double bonds, and so on for alkenes, is proposed. With the help of the new variable, both fitting and prediction accuracy of the regression model can be dramatically improved. The results obtained in this work show that the technique of projection pursuit developed in statistics is a quite promising tool for seeking an accurate quantitative structure-retention relationship (QSRR).

Journal Article↗

Developing an antituberculosis compounds database and data mining in the search of a motif responsible for the activity of a diverse class of antituberculosis agents.

A novel data mining procedure to look for new antitubercular agents and targets as well as to find a minimum common bioactive substructure (MCBS), has been reported here. The methodology extracts MCBS, both across the diverse chemical classes and within the particular chemical class, known to be present in the various marketed drugs alongside antimycobacterial compounds with known MICs. For this purpose a small in-house database of compounds has been created, for which MICs against Mycobacterium are known. The compounds have been collected from literature available on the synthetic compounds, having known MICs against Mycobacterium tuberculosis. An elaborate HQSAR (Hologram QSAR) study has been attempted to extract active fragment from a diverse class of compounds, in combination with the clustering technique to select a homogeneous group of compounds having good a profile toward the activity. The 2D pharmacophore (the 2D fragments extracted from HQSAR) has been validated searching the database. It has been found further that this validated 2D pharmacophore could be used for searching the orphan target in Mycobacterium effectively.

Antitubercular Agents↗

Data warehouse and data mining in a surgical clinic.

Hospitals and clinics have taken advantage of information systems to streamline many clinical and administrative processes. However, the potential of health care information technology as a source of data for clinical and administrative decision support has not been fully explored. In response to pressure for timely information, many hospitals are developing clinical data warehouses. This paper attempts to identify problem areas in the process of developing a data warehouse to support data mining in surgery. Based on the experience from a data warehouse in surgery several solutions are discussed.

Databases, Bibliographic↗

Data mining in bioinformatics using Weka.

UNLABELLED: The Weka machine learning workbench provides a general-purpose environment for automatic classification, regression, clustering and feature selection-common data mining problems in bioinformatics research. It contains an extensive collection of machine learning algorithms and data pre-processing methods complemented by graphical user interfaces for data exploration and the experimental comparison of different machine learning techniques on the same problem. Weka can process data given in the form of a single relational table. Its main objectives are to (a) assist users in extracting useful information from data and (b) enable them to easily identify a suitable algorithm for generating an accurate predictive model from it. AVAILABILITY: http://www.cs.waikato.ac.nz/ml/weka.

Algorithms↗

Data mining in medical time series.

This article proposes a modular, computer-based methodology to describe and compare medical problems using data mining methods. The methodology focuses on a mathematical formulation of typical classification problems, systematic extraction of interpretable features from time series, and an evaluation adapted to problem-specific preferences and limitations (computational power, interpretability, etc.). The approach is applied to instrumented gait analysis and to the individual design of myoelectric controllers for hand prostheses.

Algorithms↗

yMGV: a database for visualization and data mining of published genome-wide yeast expression data.

The yeast Microarray Global Viewer (yMGV) is an on-line database providing a synthetic view of the transcriptional expression profiles of Saccharomyces cerevisiae genes in most of the published expression datasets. yMGV displays a one-screen graphical representation of gene expression variations for each published genome-wide experiment, allowing quick retrieval of experimental conditions affecting expression of this gene. yMGV also provides tools to isolate groups of genes sharing similar transcription profiles in a defined subset of experiments. Additionally, yMGV furnishes a set of statistical tools for critical assessment of published data. We therefore believe that yMGV is an efficient tool that affords a quick and comprehensive overview of microarray data and generates new gene classifications. As of 20 March 2001 the yMGV database contains 6 000 000 measurements, representing genome-wide expression comparisons of 932 experiments from 39 microarray publications. The yMGV interface is available at http://transcriptome.ens.fr/ymgv/.

Computational Biology↗