Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Mining microarray data to identify transcription factors expressed in naïve resting but not activated T lymphocytes.

Transcriptional repressors controlling the expression of cytokine genes have been implicated in a variety of physiological and pathological phenomena. An unknown repressor that binds to the distal NFAT element of the interleukin-2 (IL-2) gene promoter in naive T-helper lymphocytes has been implicated in autoimmune phenomena and has emerged as a potentially important factor controlling the latency of HIV-1. The aim of this paper was the identification of this repressor. We resorted to public microarray databases looking for DNA-binding proteins that are present in naïve resting T cells but are downregulated when the cells are activated. A Bayesian data mining statistical analysis uncovered 25 candidate factors. Of the 25, NFAT4 and the oncogene ets-2 bind to the common motif AAGGAG found in the HIV-1 LTR and IL-2 probes. Ets-2 binding site contains the three G's that have been shown to be important for binding of the unknown factor; hence, we considered it the likeliest candidate. Electrophoretic mobility shift assays confirmed cross-reactivity between the unknown repressor and anti-ets-2 antibodies, and cotransfection experiments demonstrated the direct involvement of Ets-2 in silencing the IL-2 promoter. Designing experiments for transcription factor analysis using microarrays and Bayesian statistical methodologies provides a novel way toward elucidation of gene control networks.

CD4-Positive T-Lymphocytes↗

Mining combinatorial data in protein sequences and structures.

Combinatorial searches of structural and physical chemical properties involving the components of libraries of dipeptide, tripeptide and tetra peptide fragments were carried out in the Protein Data Bank and the SwissProt databases. The properties investigated are structural propensities, co localization of peptide fragments in protein sequences, interactions between peptide fragments in close structural proximity and the participation of physical chemical profiles in the distribution of structural motifs among peptide fragments. The results obtained for each combinatorial search in the study are classified according to the structural motifs alpha-helix, beta-sheet, reverse turn I and reverse turn II. The application of combinatorial data mined in protein databases to the design of new peptide libraries is discussed. The present findings have implications for the study of protein structure which are also discussed.

Algorithms↗

Mining genetic epidemiology data with Bayesian networks application to APOE gene variation and plasma lipid levels.

There is a critical need for data-mining methods that can identify SNPs that predict among individual variation in a phenotype of interest and reverse-engineer the biological network of relationships between SNPs, phenotypes, and other factors. This problem is both challenging and important in light of the large number of SNPs in many genes of interest and across the human genome. A potentially fruitful form of exploratory data analysis is the Bayesian or Belief network. A Bayesian or Belief network provides an analytic approach for identifying robust predictors of among-individual variation in a disease endpoints or risk factor levels. We have applied Belief networks to SNP variation in the human APOE gene and plasma apolipoprotein E levels from two samples: 702 African-Americans from Jackson, MS, and 854 non-Hispanic whites from Rochester, MN. Twenty variable sites in the APOE gene were genotyped in both samples. In Jackson, MS, SNPs 4036 and 4075 were identified to influence plasma apoE levels. In Rochester, MN, SNPs 3937 and 4075 were identified to influence plasma apoE levels. All three SNPs had been previously implicated in affecting measures of lipid and lipoprotein metabolism. Like all data-mining methods, Belief networks are meant to complement traditional hypothesis-driven methods of data analysis. These results document the utility of a Belief network approach for mining large scale genotype-phenotype association data.

Apolipoproteins E↗

Mining gene expression data by interpreting principal components.

BACKGROUND: There are many methods for analyzing microarray data that group together genes having similar patterns of expression over all conditions tested. However, in many instances the biologically important goal is to identify relatively small sets of genes that share coherent expression across only some conditions, rather than all or most conditions as required in traditional clustering; e.g. genes that are highly up-regulated and/or down-regulated similarly across only a subset of conditions. Equally important is the need to learn which conditions are the decisive ones in forming such gene sets of interest, and how they relate to diverse conditional covariates, such as disease diagnosis or prognosis. RESULTS: We present a method for automatically identifying such candidate sets of biologically relevant genes using a combination of principal components analysis and information theoretic metrics. To enable easy use of our methods, we have developed a data analysis package that facilitates visualization and subsequent data mining of the independent sources of significant variation present in gene microarray expression datasets (or in any other similarly structured high-dimensional dataset). We applied these tools to two public datasets, and highlight sets of genes most affected by specific subsets of conditions (e.g. tissues, treatments, samples, etc.). Statistically significant associations for highlighted gene sets were shown via global analysis for Gene Ontology term enrichment. Together with covariate associations, the tool provides a basis for building testable hypotheses about the biological or experimental causes of observed variation. CONCLUSION: We provide an unsupervised data mining technique for diverse microarray expression datasets that is distinct from major methods now in routine use. In test uses, this method, based on publicly available gene annotations, appears to identify numerous sets of biologically relevant genes. It has proven especially valuable in instances where there are many diverse conditions (10's to hundreds of different tissues or cell types), a situation in which many clustering and ordering algorithms become problematic. This approach also shows promise in other topic domains such as multi-spectral imaging datasets.

Algorithms↗

Metabolomics technology and bioinformatics.

Metabolomics is the global analysis of all or a large number of cellular metabolites. Like other functional genomics research, metabolomics generates large amounts of data. Handling, processing and analysis of this data is a clear challenge and requires specialized mathematical, statistical and bioinformatics tools. Metabolomics needs for bioinformatics span through data and information management, raw analytical data processing, metabolomics standards and ontology, statistical analysis and data mining, data integration and mathematical modelling of metabolic networks within a framework of systems biology. The major approaches in metabolomics, along with the modern analytical tools used for data generation, are reviewed in the context of these specific bioinformatics needs.

Animals↗

Mining genetic epidemiology data with Bayesian networks I: Bayesian networks and example application (plasma apoE levels).

MOTIVATION: The wealth of single nucleotide polymorphism (SNP) data within candidate genes and anticipated across the genome poses enormous analytical problems for studies of genotype-to-phenotype relationships, and modern data mining methods may be particularly well suited to meet the swelling challenges. In this paper, we introduce the method of Belief (Bayesian) networks to the domain of genotype-to-phenotype analyses and provide an example application. RESULTS: A Belief network is a graphical model of a probabilistic nature that represents a joint multivariate probability distribution and reflects conditional independences between variables. Given the data, optimal network topology can be estimated with the assistance of heuristic search algorithms and scoring criteria. Statistical significance of edge strengths can be evaluated using Bayesian methods and bootstrapping. As an example application, the method of Belief networks was applied to 20 SNPs in the apolipoprotein (apo) E gene and plasma apoE levels in a sample of 702 individuals from Jackson, MS. Plasma apoE level was the primary target variable. These analyses indicate that the edge between SNP 4075, coding for the well-known epsilon2 allele, and plasma apoE level was strong. Belief networks can effectively describe complex uncertain processes and can both learn from data and incorporate prior knowledge. AVAILABILITY: Various alternative and supplemental networks (not given in the text) as well as source code extensions, are available from the authors. SUPPLEMENTARY INFORMATION: http://bioinformatics.oxfordjournals.org.

Apolipoproteins E↗

Evaluation of statistical association measures for the automatic signal generation in pharmacovigilance.

Pharmacovigilance aims at detecting the adverse effects of marketed drugs. It is generally based on the spontaneous reporting of events thought to be the adverse effects of drugs. Spontaneous Reporting Systems (SRSs) supply huge databases that pharmacovigilance experts cannot exhaustively exploit without data mining tools. Data mining methods; i.e., statistical association measures in conjunction with signal generation criteria, have been proposed in the literature but there is no consensus regarding their applicability and efficiency, especially since such methods are difficult to evaluate on the basis of actual data. The objective of this paper is to evaluate association measures on simulated datasets obtained with SRS modeling. We compared association measures using the percentage of false positive signals among a given number of the most highly ranked drug-event combinations according to the values of the association measures. By considering 150 drugs and 100 adverse events, these percentages of false positives, among the 500 most highly ranked drug-event couples, vary from 1.1% to 53.4% (averages over 1000 simulated datasets). As the measures led to very different results, we could identify which measures appeared to be the most relevant for pharmacovigilance.

Adverse Drug Reaction Reporting Systems↗

Automatic classification and pattern discovery in high-throughput protein crystallization trials.

Conceptually, protein crystallization can be divided into two phases search and optimization. Robotic protein crystallization screening can speed up the search phase, and has a potential to increase process quality. Automated image classification helps to increase throughput and consistently generate objective results. Although the classification accuracy can always be improved, our image analysis system can classify images from 1,536-well plates with high classification accuracy (85%) and ROC score (0.87), as evaluated on 127 human-classified protein screens containing 5,600 crystal images and 189,472 non-crystal images. Data mining can integrate results from high-throughput screens with information about crystallizing conditions, intrinsic protein properties, and results from crystallization optimization. We apply association mining, a data mining approach that identifies frequently occurring patterns among variables and their values. This approach segregates proteins into groups based on how they react in a broad range of conditions, and clusters cocktails to reflect their potential to achieve crystallization. These results may lead to crystallization screen optimization, and reveal associations between protein properties and crystallization conditions. We also postulate that past experience may lead us to the identification of initial conditions favorable to crystallization for novel proteins.

Algorithms↗

Quantitative comparisons of in vitro assays for estrogenic activities.

Substances that may act as estrogens show a broad chemical structural diversity. To thoroughly address the question of possible adverse estrogenic effects, reliable methods are needed to detect and identify the chemicals of these diverse structural classes. We compared three assays--in vitro estrogen receptor competitive binding assays (ER binding assays), yeast-based reporter gene assays (yeast assays), and the MCF-7 cell proliferation assay (E-SCREEN assay)--to determine their quantitative agreement in identifying structurally diverse estrogens. We examined assay performance for relative sensitivity, detection of active/inactive chemicals, and estrogen/antiestrogen activities. In this examination, we combined individual data sets in a specific, quantitative data mining exercise. Data sets for at least 29 chemicals from five laboratories were analyzed pair-wise by X-Y plots. The ER binding assay was a good predictor for the other two assay results when the antiestrogens were excluded (r(2) is 0.78 for the yeast assays and 0.85 for the E-SCREEN assays). Additionally, the examination strongly suggests that biologic information that is not apparent from any of the individual assays can be discovered by quantitative pair-wise comparisons among assays. Antiestrogens are identified as outliers in the ER binding/yeast assay, while complete antagonists are identified in the ER binding and E-SCREEN assays. Furthermore, the presence of outliers may be explained by different mechanisms that induce an endocrine response, different impurities in different batches of chemicals, different species sensitivity, or limitations of the assay techniques. Although these assays involve different levels of biologic complexity, the major conclusion is that they generally provided consistent information in quantitatively determining estrogenic activity for the five data sets examined. The results should provide guidance for expanded data mining examinations and the selection of appropriate assays to screen estrogenic endocrine disruptors.

Binding, Competitive↗

Parallel and distributed methods for incremental frequent itemset mining.

Traditional methods for data mining typically make the assumption that the data is centralized, memory-resident, and static. This assumption is no longer tenable. Such methods waste computational and input/output (I/O) resources when data is dynamic, and they impose excessive communication overhead when data is distributed. Efficient implementation of incremental data mining methods is, thus, becoming crucial for ensuring system scalability and facilitating knowledge discovery when data is dynamic and distributed. In this paper, we address this issue in the context of the important task of frequent itemset mining. We first present an efficient algorithm which dynamically maintains the required information even in the presence of data updates without examining the entire dataset. We then show how to parallelize this incremental algorithm. We also propose a distributed asynchronous algorithm, which imposes minimal communication overhead for mining distributed dynamic datasets. Our distributed approach is capable of generating local models (in which each site has a summary of its own database) as well as the global model of frequent itemsets (in which all sites have a summary of the entire database). This ability permits our approach not only to generate frequent itemsets, but also to generate high-contrast frequent itemsets, which allows one to examine how the data is skewed over different sites.

Algorithms↗

Towards data warehousing and mining of protein unfolding simulation data.

OBJECTIVES: The prediction of protein structure and the precise understanding of protein folding and unfolding processes remains one of the greatest challenges in structural biology and bioinformatics. Computer simulations based on molecular dynamics (MD) are at the forefront of the effort to gain a deeper understanding of these complex processes. Currently, these MD simulations are usually on the order of tens of nanoseconds, generate a large amount of conformational data and are computationally expensive. More and more groups run such simulations and generate a myriad of data, which raises new challenges in managing and analyzing these data. Because the vast range of proteins researchers want to study and simulate, the computational effort needed to generate data, the large data volumes involved, and the different types of analyses scientists need to perform, it is desirable to provide a public repository allowing researchers to pool and share protein unfolding data. METHODS: To adequately organize, manage, and analyze the data generated by unfolding simulation studies, we designed a data warehouse system that is embedded in a grid environment to facilitate the seamless sharing of available computer resources and thus enable many groups to share complex molecular dynamics simulations on a more regular basis. RESULTS: To gain insight into the conformational fluctuations and stability of the monomeric forms of the amyloidogenic protein transthyretin (TTR), molecular dynamics unfolding simulations of the monomer of human TTR have been conducted. Trajectory data and meta-data of the wild-type (WT) protein and the highly amyloidogenic variant L55P-TTR represent the test case for the data warehouse. CONCLUSIONS: Web and grid services, especially pre-defined data mining services that can run on or 'near' the data repository of the data warehouse, are likely to play a pivotal role in the analysis of molecular dynamics unfolding data.

Computational Biology↗

Confidentiality issues for medical data miners.

The first task in any medical data mining effort is ensuring patient confidentiality. In the past, most data mining efforts ensured confidentiality by the dubious policy of withholding their raw data from colleagues and the public. A cursory review of medical informatics literature in the past decade reveals that much of what we have "learned" consists of assertions derived from confidential datasets unavailable for anyone's review. Without access to the original data, it is impossible to validate or improve upon a researcher's conclusions. Without access to research data, we are asked to accept findings as an act of faith, rather than as a scientific conclusion. This special issue of Artificial Intelligence in Medicine is devoted to medical data mining. The medical data miner has an obligation to conduct valid research in a way that protects human subjects. Today, data miners have the technical tools to merge large data collections and to distribute queries over disparate databases. In order to include patient-related data in shared databases, data miners will need methods to anonymize and deidentify data. This article reviews the human subject risks associated with medical data mining. This article also describes some of the innovative computational remedies that will permit researchers to conduct research AND share their data without risk to patient or institution.

Computer Security↗

Extraction of knowledge on protein-protein interaction by association rule discovery.

MOTIVATION: Protein-protein interactions are systematically examined using the yeast two-hybrid method. Consequently, a lot of protein-protein interaction data are currently being accumulated. Nevertheless, general information or knowledge on protein-protein interactions is poorly extracted from these data. Thus we have been trying to extract the knowledge from the protein-protein interaction data using data mining. RESULTS: A data mining method is proposed to discover association rules related to protein-protein interactions. To evaluate the detected rules by the method, a new scoring measure of the rules is introduced. The method allowed us to detect popular interaction rules such as "An SH3 domain binds to a proline-rich region." These results indicate that the method may detect novel knowledge on protein-protein interactions.

Algorithms↗

[Correlations between diagnostic information and therapeutic efficacy in rheumatoid arthritis analyzed with decision tree model].

OBJECTIVE: To explore the correlations between diagnostic information and therapeutic efficacy in rheumatoid arthritis (RA) with decision tree model analysis. METHODS: Three hundred and ninety seven patients came from 9 clinical centers were randomly divided into the Western medicine (WM) group (n=194) treated with non-steroidal anti-inflammatory drugs and slow-acting antirheumatic drug and the Chinese medicine (CM) group (n=203) with basic therapy and syndrome-differentiation dependant TCM treatment. TCM and WM diagnostic information were collected. The ACR 20 was used for efficacy evaluation and the information of patients before treatment was analyzed by SAS 8.2 statistical package. Through single-factor exploratory analysis, odds ratio of efficacy and variable was calculated taken P < 0.2 as the including criteria for data mining analysis with decision tree model. All data were classified into the training set (75%) and verifying set (25%) with efficacy as the variable for layering to make further verification of the data-mining analysis. RESULTS: Twenty variables were included in the CM group and 26 in the WM group in the data-mining model. In the former, 9 variables were positively correlated to the efficacy, including degree of arthralgia, tenderness and morning stiffness, number of swollen joint, and joint with tenderness, levels of IgM, rheumatoid factor (RF), C-reactive protein (CRP), and total assessment from doctor; and disease duration and degree of nocturnal polyuria were negatively correlated to that. While in the latter, 8 were positively correlated to the efficacy, including erythrocyte sedimentation rate (ESR), sour and weak waist and knees, white fur in tongue, joint ache and stiffness, swollen joint, and total assessment from doctor and patient, and red tongue with yellow fur and leucocyte count negatively correlated to it. Data mining with decision tree analysis revealed that different combinations of morning stiffness, slight red tongue, joint tenderness and nocturnal polyuria in the CM group, and those of white fur in tongue, CRP level, leucocyte count and morning stiffness in the WM group showed different efficacy, which were also verified in the randomly chosen verifying set. CONCLUSION: To analyze the correlations between diagnostic information and therapeutic efficacy with decision tree analysis is conformed to the theory of TCM in applying treatment according to syndrome differentiation individually, thus it would contribute to elevate the accuracy of therapy.

Adolescent↗

Combination of automated high throughput platforms, flow cytometry, and hierarchical clustering to detect cell state.

BACKGROUND: This study examined whether hierarchical clustering could be used to detect cell states induced by treatment combinations that were generated through automation and high-throughput (HT) technology. Data-mining techniques were used to analyze the large experimental data sets to determine whether nonlinear, non-obvious responses could be extracted from the data. METHODS: Unary, binary, and ternary combinations of pharmacological factors (examples of stimuli) were used to induce differentiation of HL-60 cells using a HT automated approach. Cell profiles were analyzed by incorporating hierarchical clustering methods on data collected by flow cytometry. Data-mining techniques were used to explore the combinatorial space for nonlinear, unexpected events. Additional small-scale, follow-up experiments were performed on cellular profiles of interest. RESULTS: Multiple, distinct cellular profiles were detected using hierarchical clustering of expressed cell-surface antigens. Data-mining of this large, complex data set retrieved cases of both factor dominance and cooperativity, as well as atypical cellular profiles. Follow-up experiments found that treatment combinations producing "atypical cell types" made those cells more susceptible to apoptosis. CONCLUSIONS Hierarchical clustering and other data-mining techniques were applied to analyze large data sets from HT flow cytometry. From each sample, the data set was filtered and used to define discrete, usable states that were then related back to their original formulations. Analysis of resultant cell populations induced by a multitude of treatments identified unexpected phenotypes and nonlinear response profiles.

Algorithms↗

[Analysis of factors affecting the results of assisted reproduction using a system for information mining of the SHLUK database].

OBJECTIVE: To retrospective explorating computer analysis of data about therapeutic cycles in assisted reproduction technology (ART) to confirm applicability of system for data mining SHLUK in partial analysis of fertilisation phase of therapeutic cycle. Relations between parameters of sperm count analysis and outcome of in vitro fertilisation were analyzed. DESIGN: Retrospective analysis. SETTING: 1st Depart of Obstet. and Gynaecol., Masaryk University, Brno; FEI, VSB, Ostrava. METHODS AND MATERIAL: Conditions of successful therapy in single phases of ART therapeutic cycles, were analysed using system SHLUK, which included a lot of methods for data mining. Analysis of relation between reasons and results in ART therapeutic cycles was done through method IMPL and method of group implication GRIMPL. Analysed file included data about 8516 therapeutic cycles ART in 4470 patients and data about 666 clinical pregnancies stored in electronical form in clinical data register. The model analysis of fertilisation tested relations between parameters of sperm analysis and outcome of in vitro fertilisation. Fertilisation rate (FR)--ratio of fertilized oocytes/obtained oocytes was evaluated as fertilisation stage outcome. RESULTS: Significantly higher FR--60.9% was in the group with sperm concentration before preparation 41-60 mil/ml. When sperm concentration before preparation was under 10 mil/ml--FR was significantly lower--42.2%. Motility of sperm before preparation under 10%--FR was significantly lower--45.3%. Motility of sperm before preparation 41-50%--FR was significantly higher--56.9%. Significantly higher FR--minimal 56.0% was in group of examinations with sperm after preparation was 41-90%, then FR was significantly higher--53.5%. In sperm survival test, where more than 30% of sperm survive 24 hours of cocultivation with oocytes FR was significantly higher--minimal 55.9%. CONCLUSION: Applicability of system for data mining SHLUK in the analysis of factors with influence on assisted reproduction outcome was proved. System for data mining SHLUK makes possible to define statistically significant relations between attributes of fertilisation stage of ART cycles and it is able to postulate basic hypothesis about existing reasons and results in therapeutic cycles of ART.

Data Interpretation, Statistical↗

Too much data, but little inter-changeability: a lesson learned from mining public data on tissue specificity of gene expression.

BACKGROUND: The tissue expression pattern of a gene often provides an important clue to its potential role in a biological process. A vast amount of gene expression data have been and are being accumulated in public repository through different technology platforms. However, exploitations of these rich data sources remain limited in part due to issues of technology standardization. Our objective is to test the data comparability between SAGE and microarray technologies, through examining the expression pattern of genes under normal physiological states across variety of tissues. RESULTS: There are 42-54% of genes showing significant correlations in tissue expression patterns between SAGE and GeneChip, with 30-40% of genes whose expression patterns are positively correlated and 10-15% of genes whose expression patterns are negatively correlated at a statistically significant level (p = 0.05). Our analysis suggests that the discrepancy on the expression patterns derived from technology platforms is not likely from the heterogeneity of tissues used in these technologies, or other spurious correlations resulting from microarray probe design, abundance of genes, or gene function. The discrepancy can be partially explained by errors in the original assignment of SAGE tags to genes due to the evolution of sequence databases. In addition, sequence analysis has indicated that many SAGE tags and Affymetrix array probe sets are mapped to different splice variants or different sequence regions although they represent the same gene, which also contributes to the observed discrepancies between SAGE and array expression data. CONCLUSION: To our knowledge, this is the first report attempting to mine gene expression patterns across tissues using public data from different technology platforms. Unlike previous similar studies that only demonstrated the discrepancies between the two gene expression platforms, we carried out in-depth analysis to further investigate the cause for such discrepancies. Our study shows that the exploitation of rich public expression resource requires extensive knowledge about the technologies, and experiment. Informatic methodologies for better interoperability among platforms still remain a gap. One of the areas that can be improved practically is the accurate sequence mapping of SAGE tags and array probes to full-length genes.

Journal Article↗

BISON: Bio-Interface for the Semi-global analysis Of Network patterns.

BACKGROUND: The large amount of genomics data that have accumulated over the past decade require extensive data mining. However, the global nature of data mining, which includes pattern mining, poses difficulties for users who want to study specific questions in a more local environment. This creates a need for techniques that allow a localized analysis of globally determined patterns. RESULTS: We developed a tool that determines and evaluates global patterns based on protein property and network information, while providing all the benefits of a perspective that is targeted at biologist users with specific goals and interests. Our tool uses our own data mining techniques, integrated into current visualization and navigation techniques. The functionality of the tool is discussed in the context of the transcriptional network of regulation in the enteric bacterium Escherichia coli. Two biological questions were asked: (i) Which functional categories of proteins (identified by hidden Markov models) are regulated by a regulator with a specific domain? (ii) Which regulators are involved in the regulation of proteins that contain a common hidden Markov model? Using these examples, we explain the gene-centered and pattern-centered analysis that the tool permits. CONCLUSION: In summary, we have a tool that can be used for a wide variety of applications in biology, medicine, or agriculture. The pattern mining engine is global in the way that patterns are determined across the entire network. The tool still permits a localized analysis for users who want to analyze a subportion of the total network. We have named the tool BISON (Bio-Interface for the Semi-global analysis Of Network patterns).

Journal Article↗