Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Predictive models for protein crystallization.

Crystallization of proteins is a nontrivial task, and despite the substantial efforts in robotic automation, crystallization screening is still largely based on trial-and-error sampling of a limited subset of suitable reagents and experimental parameters. Funding of high throughput crystallography pilot projects through the NIH Protein Structure Initiative provides the opportunity to collect crystallization data in a comprehensive and statistically valid form. Data mining and machine learning algorithms thus have the potential to deliver predictive models for protein crystallization. However, the underlying complex physical reality of crystallization, combined with a generally ill-defined and sparsely populated sampling space, and inconsistent scoring and annotation make the development of predictive models non-trivial. We discuss the conceptual problems, and review strengths and limitations of current approaches towards crystallization prediction, emphasizing the importance of comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high throughput crystallography initiatives, exchange of information and standardization should be encouraged, aiming to effectively integrate data mining and machine learning efforts into a comprehensive predictive framework for protein crystallization. Similar experimental design and knowledge discovery strategies should be applied to valid analysis and prediction of protein expression, solubilization, and purification, as well as crystal handling and cryo-protection.

Bayes Theorem↗

ProteinChip technology reveals distinctive protein expression profiles in the urine of bladder cancer patients.

OBJECTIVE: Since accurate biomarkers for the early diagnosis or individual prognosis of the bladder carcinoma are still not available, we used the ProteinChip technology, to search for discriminating protein expressions associated with this cancer and its subtypes. METHODS: A training set consisting of 30 archival urine samples from bladder carcinoma patients and 30 urinary samples from healthy volunteers, was analyzed via ProteinChip technology and computer based data mining. Mass clusters of differentially expressed proteins were verified by a second set (test set) comprising 21 bladder carcinoma urine samples and 21 non-tumor urinary samples. Expression differences between carcinoma subtype sample groups of the initial training set were assessed by a trend test. RESULTS: Bladder carcinoma was segregated from control with a sensitivity and specificity of 80% and 90 to 97% in the trainings set, as well as 52 to 57% and 57 to 62% in the test set, respectively. Segregation of pooled tumor stages pT2-pT3 from stages pT1 and pTa was possible at the 53.3 kDa cluster of the CM10-chip array data derived rule base. CONCLUSION: ProteinChip technology together with adapted computer based data mining tools are useful for the rapid establishment of potential protein biomarkers.

Biomarkers, Tumor↗

Identification, expression, and evolutionary analyses of plant lipocalins.

Lipocalins are a group of proteins that have been characterized in bacteria, invertebrate, and vertebrate animals. However, very little is known about plant lipocalins. We have previously reported the cloning of the first true plant lipocalins. Here we report the identification and characterization of plant lipocalins and lipocalin-like proteins using an integrated approach of data mining, expression studies, cellular localization, and phylogenetic analyses. Plant lipocalins can be classified into two groups, temperature-induced lipocalins (TILs) and chloroplastic lipocalins (CHLs). In addition, violaxanthin de-epoxidases (VDEs) and zeaxanthin epoxidases (ZEPs) can be classified as lipocalin-like proteins. CHLs, VDEs, and ZEPs possess transit peptides that target them to the chloroplast. On the other hand, TILs do not show any targeting peptide, but localization studies revealed that the proteins are found at the plasma membrane. Expression analyses by quantitative real-time PCR showed that expression of the wheat (Triticum aestivum) lipocalins and lipocalin-like proteins is associated with abiotic stress response and is correlated with the plant's capacity to develop freezing tolerance. In support of this correlation, data mining revealed that lipocalins are present in the desiccation-tolerant red algae Porphyra yezoensis and the cryotolerant marine yeast Debaryomyces hansenii, suggesting a possible association with stress-tolerant organisms. Considering the plant lipocalin properties, tissue specificity, response to temperature stress, and their association with chloroplasts and plasma membranes of green leaves, we hypothesize a protective function of the photosynthetic system against temperature stress. Phylogenetic analyses suggest that TIL lipocalin members in higher plants were probably inherited from a bacterial gene present in a primitive unicellular eukaryote. On the other hand, CHLs, VDEs, and ZEPs may have evolved from a cyanobacterial ancestral gene after the formation of the cyanobacterial endosymbiont from which the chloroplast originated.

Amino Acid Sequence↗

Transcription-based prediction of response to IFNbeta using supervised computational methods.

Changes in cellular functions in response to drug therapy are mediated by specific transcriptional profiles resulting from the induction or repression in the activity of a number of genes, thereby modifying the preexisting gene activity pattern of the drug-targeted cell(s). Recombinant human interferon beta (rIFNbeta) is routinely used to control exacerbations in multiple sclerosis patients with only partial success, mainly because of adverse effects and a relatively large proportion of nonresponders. We applied advanced data-mining and predictive modeling tools to a longitudinal 70-gene expression dataset generated by kinetic reverse-transcription PCR from 52 multiple sclerosis patients treated with rIFNbeta to discover higher-order predictive patterns associated with treatment outcome and to define the molecular footprint that rIFNbeta engraves on peripheral blood mononuclear cells. We identified nine sets of gene triplets whose expression, when tested before the initiation of therapy, can predict the response to interferon beta with up to 86% accuracy. In addition, time-series analysis revealed potential key players involved in a good or poor response to interferon beta. Statistical testing of a random outcome class and tolerance to noise was carried out to establish the robustness of the predictive models. Large-scale kinetic reverse-transcription PCR, coupled with advanced data-mining efforts, can effectively reveal preexisting and drug-induced gene expression signatures associated with therapeutic effects.

Adolescent↗

Clustering of gene expression data: performance and similarity analysis.

BACKGROUND: DNA Microarray technology is an innovative methodology in experimental molecular biology, which has produced huge amounts of valuable data in the profile of gene expression. Many clustering algorithms have been proposed to analyze gene expression data, but little guidance is available to help choose among them. The evaluation of feasible and applicable clustering algorithms is becoming an important issue in today's bioinformatics research. RESULTS: In this paper we first experimentally study three major clustering algorithms: Hierarchical Clustering (HC), Self-Organizing Map (SOM), and Self Organizing Tree Algorithm (SOTA) using Yeast Saccharomyces cerevisiae gene expression data, and compare their performance. We then introduce Cluster Diff, a new data mining tool, to conduct the similarity analysis of clusters generated by different algorithms. The performance study shows that SOTA is more efficient than SOM while HC is the least efficient. The results of similarity analysis show that when given a target cluster, the Cluster Diff can efficiently determine the closest match from a set of clusters. Therefore, it is an effective approach for evaluating different clustering algorithms. CONCLUSION: HC methods allow a visual, convenient representation of genes. However, they are neither robust nor efficient. The SOM is more robust against noise. A disadvantage of SOM is that the number of clusters has to be fixed beforehand. The SOTA combines the advantages of both hierarchical and SOM clustering. It allows a visual representation of the clusters and their structure and is not sensitive to noises. The SOTA is also more flexible than the other two clustering methods. By using our data mining tool, Cluster Diff, it is possible to analyze the similarity of clusters generated by different algorithms and thereby enable comparisons of different clustering methods.

Algorithms↗

CLOE: identification of putative functional relationships among genes by comparison of expression profiles between two species.

BACKGROUND: Public repositories of microarray data contain an incredible amount of information that is potentially relevant to explore functional relationships among genes by meta-analysis of expression profiles. However, the widespread use of this resource by the scientific community is at the moment limited by the limited availability of effective tools of analysis. We here describe CLOE, a simple cDNA microarray data mining strategy based on meta-analysis of datasets from pairs of species. The method consists in ranking EST probes in the datasets of the two species according to the similarity of their expression profiles with that of two EST probes from orthologous genes, and extracting orthologous EST pairs from a given top interval of the ranked lists. The Gene Ontology annotation of the obtained candidate partners is then analyzed for keywords overrepresentation. RESULTS: We demonstrate the capabilities of the approach by testing its predictive power on three proteomically-defined mammalian protein complexes, in comparison with single and multiple species meta-analysis approaches. Our results show that CLOE can find candidate partners for a greater number of genes, if compared to multiple species co-expression analysis, but retains a comparable specificity even when applied to species as close as mouse and human. On the other hand, it is much more specific than single organisms co-expression analysis, strongly reducing the number of potential candidate partners for a given gene of interest. CONCLUSIONS: CLOE represents a simple and effective data mining approach that can be easily used for meta-analysis of cDNA microarray experiments characterized by very heterogeneous coverage. Importantly, it produces for genes of interest an average number of high confidence putative partners that is in the range of standard experimental validation techniques.

Animals↗

Integrative genomics based identification of potential human hepatocarcinogenesis-associated cell cycle regulators: RHAMM as an example.

DNA microarray has been widely used to examine gene expression profile of different human tumors. The information generated from microarray analysis usually represents the overall range of cancer-associated abnormality associated with gene regulation. In order to identify key regulatory genes involved in carcinogenesis of human cancer, hypothesis driven data mining of the microarray data plus experimental validation becomes a critical approach in the post-genome era. Here, we present an integrative genomic analysis of published microarray data and homolog gene database. Over 20,000 genes were examined to reveal 16 genes specific to vertebrates, cell cycle G2/M regulated, and overexpressed in human HCC. Using Affymetrix microarray analysis, we found that all 16 genes were up-regulated in human HCC. Among these 16 genes, we experimentally validated the up-regulation of receptor for hyaluronan-mediated motility (RHAMM) in different cell model systems. We first confirmed elevation of RHAMM in the G2/M phase of synchronized HeLa cells. We also found that RHAMM had an elevated level of expression in all the HCC samples we examined and it was induced during the G2/M phase of regenerating mouse hepatocytes after partial hepatectomy. Thus, the expression of RHAMM appears to be tightly regulated during mammalian cell cycle G2/M progression. The ectopic overexpression of RHAMM in 293T cells resulted in the accumulation of cells at G2/M phase. RHAMM-induced mitotic arrest of cells was predominantly in the prophase. Taken together, using an integrated functional genomic approach, we have uncovered a set of genes that may play specific roles in cell cycle progression and in HCC development. To elucidate the function of these genes in cell cycle regulation may shed light on the control mechanism of human HCC in the future.

Amino Acid Sequence↗

Mining parasite data using genetic programming.

Genetic programming is a technique that can be used to tackle the hugely demanding data-processing problems encountered in the natural sciences. Application of genetic programming to a problem using parasites as biological tags demonstrates its potential for developing explanatory models using data that are both complex and noisy.

Algorithms↗

Multiple intermolecular interaction modes of positively charged residues with adenine in ATP-binding proteins.

Adenosine 5'-triphosphate (ATP) plays an essential role in all forms of life. Molecular recognition of ATP in ATP-binding proteins is a subject of great importance for understanding enzymatic mechanisms and for drug design. We have carried out a large-scale data mining of the Protein Data Bank (PDB) to analyze molecular determinants for recognition of ATP, in particular, the adenine base, by ATP-binding proteins. A novel distribution pattern of charged residues around the adenine base was discovered: lysine residues tend to occupy the major groove N7 side of the adenine base, and the arginine residues situate preferentially above or below the adenine bases. Such an arrangement is advantageous because it facilitates multiple modes of intermolecular interactions, that is, cation-pi interactions and a hydrogen bond between lysine and adenine, and cation-pi and pi-pi stacking interactions between arginine and adenine. For the two representative Lys... Adenine and Arg... Adenine interactions, intermolecular interaction energies were subsequently analyzed by means of the supermolecular approach at the MP2 level with solvation free energy correction using the SM5.42R model of Cramer and Truhlar, which gave rise to significant interaction strengths.

Adenine↗

Feature selection and transduction for prediction of molecular bioactivity for drug design.

MOTIVATION: In drug discovery a key task is to identify characteristics that separate active (binding) compounds from inactive (non-binding) ones. An automated prediction system can help reduce resources necessary to carry out this task. RESULTS: Two methods for prediction of molecular bioactivity for drug design are introduced and shown to perform well in a data set previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001. The data is characterized by very few positive examples, a very large number of features (describing three-dimensional properties of the molecules) and rather different distributions between training and test data. Two techniques are introduced specifically to tackle these problems: a feature selection method for unbalanced data and a classifier which adapts to the distribution of the the unlabeled test data (a so-called transductive method). We show both techniques improve identification performance and in conjunction provide an improvement over using only one of the techniques. Our results suggest the importance of taking into account the characteristics in this data which may also be relevant in other problems of a similar type.

Algorithms↗

CryptoDB: a Cryptosporidium bioinformatics resource update.

The database, CryptoDB (http://CryptoDB.org), is a community bioinformatics resource for the AIDS-related apicomplexan-parasite, Cryptosporidium. CryptoDB integrates whole genome sequence and annotation with expressed sequence tag and genome survey sequence data and provides supplemental bioinformatics analyses and data-mining tools. A simple, yet comprehensive web interface is available for mining and visualizing the data. CryptoDB is allied with the databases PlasmoDB and ToxoDB via ApiDB, an NIH/NIAID-fundedBioinformatics Resource Center. Recent updates to CryptoDB include the deposition of annotated genome sequences for Cryptosporidium parvum and Cryptosporidium hominis, migration to a relational database (GUS), a new query and visualization interface and the introduction of Web services.

Animals↗

MetricMap: an embedding technique for processing distance-based queries in metric spaces.

In this paper, we present an embedding technique, called MetricMap, which is capable of estimating distances in a pseudometric space. Given a database of objects and a distance function for the objects, which is a pseudometric, we map the objects to vectors in a pseudo-Euclidean space with a reasonably low dimension while preserving the distance between two objects approximately. Such an embedding technique can be used as an approximate oracle to process a broad class of distance-based queries. It is also adaptable to data mining applications such as data clustering and classification. We present the theory underlying MetricMap and conduct experiments to compare MetricMap with other methods including MVP-tree and M-tree in processing the distance-based queries. Experimental results on both protein and RNA data show the good performance and the superiority of MetricMap over the other methods.

Algorithms↗

Mining medical data.

Explore the source record for details and available documents.

Data Interpretation, Statistical↗

Mining microarray data at NCBI's Gene Expression Omnibus (GEO)*.

The Gene Expression Omnibus (GEO) at the National Center for Biotechnology Information (NCBI) has emerged as the leading fully public repository for gene expression data. This chapter describes how to use Web-based interfaces, applications, and graphics to effectively explore, visualize, and interpret the hundreds of microarray studies and millions of gene expression patterns stored in GEO. Data can be examined from both experiment-centric and gene-centric perspectives using user-friendly tools that do not require specialized expertise in microarray analysis or time-consuming download of massive data sets. The GEO database is publicly accessible through the World Wide Web at http://www.ncbi.nlm.nih.gov/geo.

Algorithms↗

HTS quality control and data analysis: a process to maximize information from a high-throughput screen.

Changes in all aspects of HTS from compound management through to evaluation of hits and leads, strengthened by infrastructure improvements, in both automation and informatics, have made possible increased analysis and implementation of process and quality control throughout HTS. This paper focuses on the process of HTS with an emphasis on quality control, reducing the variability of all the processes that have an impact on the final result, and argue that by increasing the quality of the entire process that data mining of primary screening data is in fact possible and will reduce cycle times to medicinal chemistry.

Automation↗