Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Classification using partial least squares with penalized logistic regression.

MOTIVATION: One important aspect of data-mining of microarray data is to discover the molecular variation among cancers. In microarray studies, the number n of samples is relatively small compared to the number p of genes per sample (usually in thousands). It is known that standard statistical methods in classification are efficient (i.e. in the present case, yield successful classifiers) particularly when n is (far) larger than p. This naturally calls for the use of a dimension reduction procedure together with the classification one. RESULTS: In this paper, the question of classification in such a high-dimensional setting is addressed. We view the classification problem as a regression one with few observations and many predictor variables. We propose a new method combining partial least squares (PLS) and Ridge penalized logistic regression. We review the existing methods based on PLS and/or penalized likelihood techniques, outline their interest in some cases and theoretically explain their sometimes poor behavior. Our procedure is compared with these other classifiers. The predictive performance of the resulting classification rule is illustrated on three data sets: Leukemia, Colon and Prostate.

Algorithms↗

Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research.

SUMMARY: We present here Blast2GO (B2G), a research tool designed with the main purpose of enabling Gene Ontology (GO) based data mining on sequence data for which no GO annotation is yet available. B2G joints in one application GO annotation based on similarity searches with statistical analysis and highlighted visualization on directed acyclic graphs. This tool offers a suitable platform for functional genomics research in non-model species. B2G is an intuitive and interactive desktop application that allows monitoring and comprehension of the whole annotation and analysis process. AVAILABILITY: Blast2GO is freely available via Java Web Start at http://www.blast2go.de. SUPPLEMENTARY MATERIAL: http://www.blast2go.de -> Evaluation.

Algorithms↗

NASCArrays: a repository for microarray data generated by NASC's transcriptomics service.

NASC operates an Affymetrix 'GeneChip' (microarray) service for the Arabidopsis thaliana community. All data produced by the service are publicly available through our microarray data base 'NASCArrays' published at http://affymetrix. arabidopsis.info. The data are accessible through text searching and a series of data mining tools. All data are annotated with sample preparation details, and the original Affymetrix data are available for download. The database aims to be MIAME supportive and provide a coordinated resource for re searchers interested in the transcriptome of Arabidopsis. Using this database, data produced will be shared with other databases worldwide.

Arabidopsis↗

Detecting selective expression of genes and proteins.

Selective expression of a gene product (mRNA or protein) is a pattern in which the expression is markedly high, or markedly low, in one particular tissue compared with its level in other tissues or sources. We present a computational method for the identification of such patterns. The method combines assessments of the reliability of expression quantitation with a statistical test of expression distribution patterns. The method is applicable to small studies or to data mining of abundance data from expression databases, whether mRNA or protein. Though the method was developed originally for gene-expression analyses, the computational method is, in fact, rather general. It is well suited for the identification of exceptional values in many sorts of intensity data, even noisy data, for which assessments of confidences in the sources of the intensities are available. Moreover, the method is indifferent as to whether the intensities are experimentally or computationally derived. We show details of the general method and examples of computational results on gene abundance data.

Algorithms↗

Integrating data from biological experiments into metabolic networks with the DBE information system.

Modern 'omics'-technologies result in huge amounts of data about life processes. For analysis and data mining purposes this data has to be considered in the context of the underlying biological networks. This work presents an approach for integrating data from biological experiments into metabolic networks by mapping the data onto network elements and visualising the data enriched networks automatically. This methodology is implemented in DBE, an information system that supports the analysis and visualisation of experimental data in the context of metabolic networks. It consists of five parts: (1) the DBE-Database for consistent data storage, (2) the Excel-Importer application for the data import, (3) the DBE-Website as the interface for the system, (4) the DBE-Pictures application for the up- and download of binary (e. g. image) files, and (5) DBE-Gravisto, a network analysis and graph visualisation system. The usability of this approach is demonstrated in two examples.

Computational Biology↗

Use and misuse of p-values in designed and observational studies: guide for researchers and reviewers.

Analysis of scientific data involves many components, one of which is often statistical testing with the calculation of p-values. However, researchers too often pepper their papers with p-values in the absence of critical thinking about their results. In fact, statistical tests in their various forms address just one question: does an observed difference exceed that which might reasonably be expected solely as a result of sampling error and/or random allocation of experimental material? Such tests are best applied to the results of designed studies with reasonable control of experimental error and sampling error, as well as acquisition of a sufficient sample size. Nevertheless, attributing an observed difference to a specific treatment effect requires critical thinking on the part of the scientist. Observational studies involve data sets whose size is usually a matter of convenience with results that reflect a number of potentially confounding factors. In this situation, statistical testing is not appropriate and p-values may be misleading; other more modern statistical tools should be used instead, including graphic analysis, computer-intensive methods, regression trees, and other procedures broadly classified as bioinformatics, data mining, and exploratory data analysis. In this review, the utility of p-values calculated from designed experiments and observational studies are discussed, leading to the formation of a decision tree to aid researchers and reviewers in understanding both the benefits and limitations of statistical testing.

Clinical Trials as Topic↗

A classification of delivery patient groups using CART (Classification and Regression Trees) for an improvement of critical path.

Critical Path is able to offer high quality standardized medical treatment. In searching for the better Critical Path (CP) and improving CP, it is necessary to use it continuously. This paper indicates the improvement from two classes of delivery patient groups to five, using data mining method. The data used for analysis is medicine data based on the actual data of delivery patient groups. As a result, it became possible to offer higher quality medical treatment.

Critical Pathways↗

A collaborative international nursing informatics research project: predicting ARDS risk in critically ill patients.

An international nursing informatics research collaboration between Duke University Medical Center in Durham, North Carolina, USA and Allgemeine Krankenhaus Hospital in Vienna, Austria used data mining techniques called Knowledge Discovery in Databases (KDD) to explore the relationship between clinical data variables and adult respiratory distress syndrome (ARDS) in critically ill patients. Results of the study and logistics of international research collaboration will be presented at NI '97. The conceptual model, data mining methodology, and objectives for the collaboration are described here.

Austria↗

Web-based tools for mining the NCI databases for anticancer drug discovery.

In this paper, we describe the development of a set of integrated Web-based tools for mining the National Cancer Institute's (NCI) anticancer databases for anticancer drug discovery. For data mining, three different correlation algorithms were implemented, which included the commonly used Pearson's correlation algorithm available from the NCI's COMPARE program, the Spearman's and Kendall's correlation algorithms. In addition, we implemented the p-value test to evaluate the significance of the correlation results. These Web-based data mining tools allow robust analysis of the correlation between the in vitro anticancer activity of the drugs in the NCI anticancer database, the protein levels and mRNA levels of molecular targets (genes) in the NCI 60 human cancer cell lines for identification of potential lead compounds for a specific molecular target and for study of the molecular mechanism action of a drug. Examples were provided to identify PKC ligands using a lead compound and to identify potential ErbB-2 inhibitors using the mRNA levels of ErbB-2 in the NCI 60 tumor cell lines.

Antineoplastic Agents↗

Automated methods of predicting the function of biological sequences using GO and BLAST.

BACKGROUND: With the exponential increase in genomic sequence data there is a need to develop automated approaches to deducing the biological functions of novel sequences with high accuracy. Our aim is to demonstrate how accuracy benchmarking can be used in a decision-making process evaluating competing designs of biological function predictors. We utilise the Gene Ontology, GO, a directed acyclic graph of functional terms, to annotate sequences with functional information describing their biological context. Initially we examine the effect on accuracy scores of increasing the allowed distance between predicted and a test set of curator assigned terms. Next we evaluate several annotator methods using accuracy benchmarking. Given an unannotated sequence we use the Basic Local Alignment Search Tool, BLAST, to find similar sequences that have already been assigned GO terms by curators. A number of methods were developed that utilise terms associated with the best five matching sequences. These methods were compared against a benchmark method of simply using terms associated with the best BLAST-matched sequence (best BLAST approach). RESULTS: The precision and recall of estimates increases rapidly as the amount of distance permitted between a predicted term and a correct term assignment increases. Accuracy benchmarking allows a comparison of annotation methods. A covering graph approach performs poorly, except where the term assignment rate is high. A term distance concordance approach has a similar accuracy to the best BLAST approach, demonstrating lower precision but higher recall. However, a discriminant function method has higher precision and recall than the best BLAST approach and other methods shown here. CONCLUSION: Allowing term predictions to be counted correct if closely related to a correct term decreases the reliability of the accuracy score. As such we recommend using accuracy measures that require exact matching of predicted terms with curator assigned terms. Furthermore, we conclude that competing designs of BLAST-based GO term annotators can be effectively compared using an accuracy benchmarking approach. The most accurate annotation method was developed using data mining techniques. As such we recommend that designers of term annotators utilise accuracy benchmarking and data mining to ensure newly developed annotators are of high quality.

Benchmarking↗

Health system mines financial data to unearth clinical conclusions, improvements.

Making clinical conclusions from your financial data. Using artificial intelligence to provide cleaner and less biased data, Sentara Health System in Norfolk, VA, has changed the way care is delivered for 25 top inpatient and outpatient conditions and procedures--and saved $4 million in the last three years as a direct result. Here's a look at how Sentara is mining clinical information from financial data.

Algorithms↗

Microarray data normalization and transformation.

Underlying every microarray experiment is an experimental question that one would like to address. Finding a useful and satisfactory answer relies on careful experimental design and the use of a variety of data-mining tools to explore the relationships between genes or reveal patterns of expression. While other sections of this issue deal with these lofty issues, this review focuses on the much more mundane but indispensable tasks of 'normalizing' data from individual hybridizations to make meaningful comparisons of expression levels, and of 'transforming' them to select genes for further analysis and data mining.

Animals↗

Comparison of hospital charge prediction models for colorectal cancer patients: neural network vs. decision tree models.

Analysis and prediction of the care charges related to colorectal cancer in Korea are important for the allocation of medical resources and the establishment of medical policies because the incidence and the hospital charges for colorectal cancer are rapidly increasing. But the previous studies based on statistical analysis to predict the hospital charges for patients did not show satisfactory results. Recently, data mining emerges as a new technique to extract knowledge from the huge and diverse medical data. Thus, we built models using data mining techniques to predict hospital charge for the patients. A total of 1,022 admission records with 154 variables of 492 patients were used to build prediction models who had been treated from 1999 to 2002 in the Kyung Hee University Hospital. We built an artificial neural network (ANN) model and a classification and regression tree (CART) model, and compared their prediction accuracy. Linear correlation coefficients were high in both models and the mean absolute errors were similar. But ANN models showed a better linear correlation than CART model (0.813 vs. 0.713 for the hospital charge paid by insurance and 0.746 vs. 0.720 for the hospital charge paid by patients). We suggest that ANN model has a better performance to predict charges of colorectal cancer patients.

Algorithms↗

[Computer aided intracranial aneurysm embolization with GDC].

OBJECTIVE: To establish an expert system that automatically generates optimal GDC selection program for the embolization of intracranial aneurysm. METHODS: Twenty highly cost-effective cases of intracranial aneurysm embolized with GDC dense packing were collected. Each of them contains information including aneurysm's volume measured by three-dimension digital subtraction angiography (3D DSA), aneurysm's location, maximum transverse diameter, maximum length diameter, and GDC selection program. An expert system made up of a case base, a mathematical model simulating experts' experience (established with the help of data mining techniques combining multi-layer perceptron network with polyhedrons in high dimensional space), and data envelopment analysis (DEA), was implemented. RESULTS: When the user inputted four required parameters (volume, location, maximum transverse diameter, and maximum length diameter) into the expert system and clicked the "program design" button, candidate GDC selection program(s) would be presented in the result box. CONCLUSION: Case base, data mining techniques, and DEA can be used to establish the expert system that automatically generates optimal GDC selection program for the embolization of intracranial aneurysm. Its clinical value needs to be further evaluated.

Adult↗

Joint explorative analysis of neuroreceptor subsystems in the human brain: application to receptor-transporter correlation using PET data.

Positron emission tomography (PET) has proved to be a highly successful technique in the qualitative and quantitative exploration of the human brain's neurotransmitter-receptor systems. In recent years, the number of PET radioligands, targeted to different neuroreceptor systems of the human brain, has increased considerably. This development paves the way for a simultaneous analysis of different receptor systems and subsystems in the same individual. The detailed exploration of the versatility of neuroreceptor systems requires novel technical approaches, capable of operating on huge parametric image datasets. An initial step of such explorative data processing and analysis should be the development of novel exploratory data-mining tools to gain insight into the "structure" of complex multi-individual, multi-receptor data sets. For practical reasons, a possible and feasible starting point of multi-receptor research can be the analysis of the pre- and post-synaptic binding sites of the same neurotransmitter. In the present study, we propose an unsupervised, unbiased data-mining tool for this task and demonstrate its usefulness by using quantitative receptor maps, obtained with positron emission tomography, from five healthy subjects on (pre-synaptic) serotonin transporters (5-HTT or SERT) and (post-synaptic) 5-HT(1A) receptors. Major components of the proposed technique include the projection of the input receptor maps to a feature space, the quasi-clustering and classification of projected data (neighbourhood formation), trans-individual analysis of neighbourhood properties (trajectory analysis), and the back-projection of the results of trajectory analysis to normal space (creation of multi-receptor maps). The resulting multi-receptor maps suggest that complex relationships and tendencies in the relationship between pre- and post-synaptic transporter-receptor systems can be revealed and classified by using this method. As an example, we demonstrate the regional correlation of the serotonin transporter-receptor systems. These parameter-specific multi-receptor maps can usefully guide the researchers in their endeavour to formulate models of multi-receptor interactions and changes in the human brain.

Adult↗

In silico identification of breast cancer genes by combined multiple high throughput analyses.

Publicly available human genomic sequence data provide an unprecedented opportunity for researchers to decode the functionality of human genome. Such information is extremely valuable in cancer prevention diagnosis and treatment. Cancer Genome Anatomy Project (CGAP) and Gene Expression Omnibus (GEO) are two bioinformatic infrastructures for studying functional genomics. The goal of this study is to explore the feasibility of incorporating the Internet-available bioinformatic databases to discover human breast cancer-related genes. Several tools including the Gene Finder, Virtual Northern (vNorthern) and SAGE digital gene expression displayer (DGED) were used to analyze differential gene expression between benign and malignant breast tissues. A pilot study was performed using both EST and SAGE vNorthern to analyze the expression of a panel of known genes, including high abundance genes beta-actin and G3PDH, low abundance genes BRCA1 and p53, tissue-specific genes CEA and PSA and two breast cancer-related genes Her2/neu and MUC1. We found a high expression of beta-actin and G3PDH and a low expression of BRCA1 and p53 across different types of tissues as well as a tissue-specific expression of CEA in colon and PSA in prostate. A further analysis of 30 known breast cancer-related genes in breast cancer tissues by vNorthern demonstrated a high expression of oncogenes and low expression of tumor suppressor genes. An open-end analysis of two pools of breast cancer and benign breast tissue libraries by SAGE DGED produced 53 differentially expressed genes according to the screening criteria of a >five-fold difference and p<0.01. Further analysis by EST vNorthern and virtual microarray analysis reduced the candidate genes to six, with four down-regulated genes, ANXA1, CAV1, KRT5 and MMP7, and two up-regulated genes, ERBB2 and G1P3 in breast cancer. These findings were validated by a real-time RT-PCR analysis in eight paired human breast cancer tissue samples. We conclude that the combined multiple high throughput analyses is an effective data mining strategy in cancer gene identification. This approach may improve the usage of public available genomic data through strategic data mining of high throughput analysis.

Blotting, Northern↗

Scalable partitioning and exploration of chemical spaces using geometric hashing.

Virtual screening (VS) has become a preferred tool to augment high-throughput screening(1) and determine new leads in the drug discovery process. The core of a VS informatics pipeline includes several data mining algorithms that work on huge databases of chemical compounds containing millions of molecular structures and their associated data. Thus, scaling traditional applications such as classification, partitioning, and outlier detection for huge chemical data sets without a significant loss in accuracy is very important. In this paper, we introduce a data mining framework built on top of a recently developed fast approximate nearest-neighbor-finding algorithm(2) called locality-sensitive hashing (LSH) that can be used to mine huge chemical spaces in a scalable fashion using very modest computational resources. The core LSH algorithm hashes chemical descriptors so that points close to each other in the descriptor space are also close to each other in the hashed space. Using this data structure, one can perform approximate nearest-neighbor searches very quickly, in sublinear time. We validate the accuracy and performance of our framework on three real data sets of sizes ranging from 4337 to 249 071 molecules. Results indicate that the identification of nearest neighbors using the LSH algorithm is at least 2 orders of magnitude faster than the traditional k-nearest-neighbor method and is over 94% accurate for most query parameters. Furthermore, when viewed as a data-partitioning procedure, the LSH algorithm lends itself to easy parallelization of nearest-neighbor classification or regression. We also apply our framework to detect outlying (diverse) compounds in a given chemical space; this algorithm is extremely rapid in determining whether a compound is located in a sparse region of chemical space or not, and it is quite accurate when compared to results obtained using principal-component-analysis-based heuristics.

Journal Article↗

Securing electronic health records without impeding the flow of information.

OBJECTIVE: We present an integrated set of technologies, known as the Hippocratic Database, that enable healthcare enterprises to comply with privacy and security laws without impeding the legitimate management, sharing, and analysis of personal health information. APPROACH: The Hippocratic Database approach to securing electronic health records involves (1) active enforcement of fine-grained data disclosure policies using query modification techniques, (2) efficient auditing of past database access to verify compliance with policies and track security breaches, (3) data mining algorithms that preserve privacy by randomizing information at the individual level, (4) de-identification of personal health data using an optimal method of k-anonymization, and (5) information sharing across autonomous data sources using cryptographic protocols. CONCLUSIONS: Our research confirms that policies concerning the disclosure of electronic health records can be reliably and efficiently enforced and audited at the database level. We further demonstrate that advanced data mining and anonymization techniques can be employed to analyze aggregate health records without revealing individual patient identities. Finally, we show that web services and commutative encryption can be used to share sensitive information selectively among autonomous entities without compromising security or privacy.

Access to Information↗