Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

New approaches towards integrated proteomic databases and depositories.

Since the publication of the human genome, two key points have emerged. First, it is still not certain which regions of the genome code for proteins. Second, the number of discrete protein-coding genes is far fewer than the number of different proteins. Proteomics has the potential to address some of these postgenomic issues if the obstacles that we face can be overcome in our efforts to combine proteomic and genomic data. There are many challenges associated with high-throughput and high-output proteomic technologies. Consequently, for proteomics to continue at its current growth rate, new approaches must be developed to ease data management and data mining. Initiatives have been launched to develop standard data formats for exchanging mass spectrometry proteomic data, including the Proteomics Standards Initiative formed by the Human Proteome Organization. Databases such as SwissProt and Uniprot are publicly available repositories for protein sequences annotated for function, subcellular location and known potential post-translational modifications. The availability of bioinformatics solutions is crucial for proteomics technologies to fulfil their promise of adding further definition to the functional output of the human genome. The aim of the Oxford Genome Anatomy Project is to provide a framework for integrating molecular, cellular, phenotypic and clinical information with experimental genetic and proteomics data. This perspective also discusses models to make the Oxford Genome Anatomy Project accessible and beneficial for academic and commercial research and development.

Databases, Protein↗

Survival and viral load in four groups of HIV-1 infected hemophiliacs compared by three-way data clustering.

We assigned a total of 131 hemophiliacs infected with HIV-1 into four clusters by applying a 3-way data analysis method. Sequentially acquired CD4+ and CD8+ cell counts obtained longitudinally over an observation period from 1986 to 1992 were analyzed. During the successive observation in this interval, a clustering of patients is not always coincident over all the times, because the cell counts vary with time. Therefore, the 3-way data clustering is to obtain the optimal result of the classification of patients through all the interval of observation. Examining patients' survival after that period, the cumulative mortality rate was highest among the 36 hemophiliacs in Cluster 1. Less mortality was found in Cluster 2, consisting of 49 hemophiliacs and none was reported in Clusters 3 and 4, which included 33 and 13 hemophiliacs, respectively. However, a significantly lower blood viral copy number was found in Cluster 3 than in Cluster 4. A total of six long-term non-progressors was found, five in Cluster 3 and one in Cluster 4, while none was found in Cluster 1 or 2. As demonstrated in this analysis, 3-way data clustering may represent a good data mining technique for handling various types of clinical data.

CD4 Lymphocyte Count↗

[Mining microarray gene expression data of metastatic colorectal cancer by literature profiling].

OBJECTIVE: To search new metastatic colorectal cancer-related genes. METHODS: Metastatic colorectal cancer microarray gene expression data was mined by literature profiling, which was based on the analysis of literature profiles generated by extracting the frequencies of certain terms from the abstracts in the Medline literature database. The terms are then filtered on the basis of both repetitive occurrence and co-occurrence among multiple gene entries. Clustering analysis was subsequently performed on the retained frequency values, shaping a coherent picture of the functional relationship among the heterogeneous genes identified in the long lists. The clustering result was analyzed against the result of document analysis and experiment. RESULTS AND CONCLUSION: Two new genes (TIAM1 and NM23H1) with potential relation to metastatic colorectal cancer were identified.

Colorectal Neoplasms↗

[Mining gene expression microarray data of nasopharyngeal carcinoma by literature profiling].

OBJECTIVE: To study abnormal signal pathway in nasopharyngeal carcinoma (NPC). METHOD: NPC gene expression microarray data was mined by analysis of literature profiles generated by extracting the frequencies of certain terms from the abstracts stored in the Medline literature database. The terms were then filtered on the basis of both repetitive occurrence and co-occurrence among multiple gene entries. Finally, clustering analysis was performed on the retained frequency values, shaping a coherent picture of the functional relationship among large and heterogeneous lists of genes. RESULT: Sixteen function groups were found among 112 abnormally expressed genes, including 4 groups indicative of Epstein-Barr virus (EBV) infection, 6 groups indicative of normal nasopharyngeal tissues that acquired essential capabilities to develop into tumor, 2 groups involved in energy metabolism, 1 group suggesting abnormal phosphorylation of proteins, 2 groups related to other diseases, and 1 group associated with muscle activities. The pathways of p53 and Rb, which were frequently abnormal during tumor progression, were not found in these groups. CONCLUSION: Initiation and progression of NPC may be caused by special signal transduction pathways.

Humans↗

Genome mining for human cancer genes: wherefore art thou?

In an initial data-mining effort, the draft human genome was searched to find paralogs of known tumor suppressor genes, and for gene arrangements, which are typical of oncogenes, in cancer cells. The results were disappointing, indicating that although knowledge of the human genome will undoubtedly be of great help, other approaches to identify new oncogenes are needed.

Databases, Factual↗

The different strategies for designing GPCR and kinase targeted libraries.

In recent years the trend in combinatorial library design has shifted to include target class focusing along with diversity and drug-likeness criteria. In this manuscript we review the computational tools available for target class library design and highlight the areas where they have proven useful in our work. The protein kinase family is used to illustrated structure-based target class focused library design, and the G-protein coupled receptor (GPCR) family is used to illustrate ligand-based target class focused library design. Most of the tools discussed are those designed for libraries targeted to a single protein and are simply applied "brute-force" to a large number of targets within the family. The tools that have proven to be the most useful in our work are those that can extract trends from the computational data such as docking and clustering or data mining large amounts of structure activity or high throughput screening data. Finally, areas where improvements are needed in the computational tools available for target class focusing are highlighted. These areas include tools to extract the relevant patterns from all available information for a family of targets, tools to efficiently apply models for all targets in the family rather than just a small subset, mining tools to extract the relevant information from the computational absorption, distribution, metabolism, excretion and toxicity (ADMET) and targeting data, and tools to allow interactive exploration of the virtual space around a library to facilitate the selection of the library that best suits the needs of the design team.

Animals↗

Mining gene expression databases for association rules.

MOTIVATION: Global gene expression profiling, both at the transcript level and at the protein level, can be a valuable tool in the understanding of genes, biological networks, and cellular states. As larger and larger gene expression data sets become available, data mining techniques can be applied to identify patterns of interest in the data. Association rules, used widely in the area of market basket analysis, can be applied to the analysis of expression data as well. Association rules can reveal biologically relevant associations between different genes or between environmental effects and gene expression. An association rule has the form LHS --> RHS, where LHS and RHS are disjoint sets of items, the RHS set being likely to occur whenever the LHS set occurs. Items in gene expression data can include genes that are highly expressed or repressed, as well as relevant facts describing the cellular environment of the genes (e.g. the diagnosis of a tumor sample from which a profile was obtained). RESULTS: We demonstrate an algorithm for efficiently mining association rules from gene expression data, using the data set from Hughes et al. (2000, Cell, 102, 109-126) of 300 expression profiles for yeast. Using the algorithm, we find numerous rules in the data. A cursory analysis of some of these rules reveals numerous associations between certain genes, many of which make sense biologically, others suggesting new hypotheses that may warrant further investigation. In a data set derived from the yeast data set, but with the expression values for each transcript randomly shifted with respect to the experiments, no rules were found, indicating that most all of the rules mined from the actual data set are not likely to have occurred by chance. AVAILABILITY: An implementation of the algorithm using Microsoft SQL Server with Access 2000 is available at http://dot.ped.med.umich.edu:2000/pub/assoc_rules/assoc_rules.zip. Our results from mining the yeast data set are available at http://dot.ped.med.umich.edu:2000/pub/assoc_rules/yeast_results.zip.

Algorithms↗

Human genetics in health care.

UNLABELLED: The Human Genome Project, the mapping of our 100,000 genes and the sequencing of all of our DNA, will have major impact on biomedical research and the whole of therapeutic and preventive health care. The tracing of genetic diseases to their molecular causes is rapidly expanding diagnostic and preventive options, while the increased insights into molecular pathways open tremendous perspectives for pharmacological and genetic therapies. The design of animal model systems for the functional study of disease and development and the use of bioinformatics and biostatistics to improve our pattern recognition abilities are greatly accelerating progress. However, the optimal value from the current explosion of 'data mining' possibilities will only be gained when the basic data are made and kept publicly accessible, while at the same time safeguarding the protection of intellectual property arising from downstream inventions. This is one of the goals of the international Human Genome Organisation, established 10 years ago to assist coordinating data acquisition and exchange and societal implementation of the genome project. Additional points of major attention in this historic endeavour are the safeguarding of a worldwide balance in the contribution and benefits to countries and populations, the prevention of stigmatisation and discrimination of individuals and groups and the maintenance of respect for the diversity of our world's cultures and traditions. CONCLUSION: The acquisition and use of genomic information for health care benefit should be seen in the light of a worldwide improvement without prejudice.

Ethics, Medical↗

Content-based image database system for epilepsy.

We have designed and implemented a human brain multi-modality database system with content-based image management, navigation and retrieval support for epilepsy. The system consists of several modules including a database backbone, brain structure identification and localization, segmentation, registration, visual feature extraction, clustering/classification and query modules. Our newly developed anatomical landmark localization and brain structure identification method facilitates navigation through an image data and extracts useful information for segmentation, registration and query modules. The database stores T1-, T2-weighted and FLAIR MRI and ictal/interictal SPECT modalities with associated clinical data. We confine the visual feature extractors within anatomical structures to support semantically rich content-based procedures. The proposed system serves as a research tool to evaluate a vast number of hypotheses regarding the condition such as resection of the hippocampus with a relatively small volume and high average signal intensity on FLAIR. Once the database is populated, using data mining tools, partially invisible correlations between different modalities of data, modeled in database schema, can be discovered. The design and implementation aspects of the proposed system are the main focus of this paper.

Brain↗

Bioinformatics in proteomics: application, terminology, and pitfalls.

Bioinformatics applies data mining, i.e., modern computer-based statistics, to biomedical data. It leverages on machine learning approaches, such as artificial neural networks, decision trees and clustering algorithms, and is ideally suited for handling huge data amounts. In this article, we review the analysis of mass spectrometry data in proteomics, starting with common pre-processing steps and using single decision trees and decision tree ensembles for classification. Special emphasis is put on the pitfall of overfitting, i.e., of generating too complex single decision trees. Finally, we discuss the pros and cons of the two different decision tree usages.

Computational Biology↗

Analysing six types of protein-protein interfaces.

Non-covalent residue side-chain interactions occur in many different types of proteins and facilitate many biological functions. Are these differences manifested in the sequence compositions and/or the residue-residue contact preferences of the interfaces? Previous studies analysed small data sets and gave contradictory answers. Here, we introduced a new data-mining method that yielded the largest high-resolution data set of interactions analysed. We introduced an information theory-based analysis method. On the basis of sequence features, we were able to differentiate six types of protein interfaces, each corresponding to a different functional or structural association between residues. Particularly, we found significant differences in amino acid composition and residue-residue preferences between interactions of residues within the same structural domain and between different domains, between permanent and transient interfaces, and between interactions associating homo-oligomers and hetero-oligomers. The differences between the six types were so substantial that, using amino acid composition alone, we could predict statistically to which of the six types of interfaces a pool of 1000 residues belongs at 63-100% accuracy. All interfaces differed significantly from the background of all residues in SWISS-PROT, from the group of surface residues, and from internal residues that were not involved in non-trivial interactions. Overall, our results suggest that the interface type could be predicted from sequence and that interface-type specific mean-field potentials may be adequate for certain applications.

Amino Acid Sequence↗

Analyzing tumor gene expression profiles.

A brief introduction to high throughput technologies for measuring and analyzing gene expression is given. Various supervised and unsupervised data mining methods for analyzing the produced high-dimensional data are discussed. The main emphasis is on supervised machine learning methods for classification and prediction of tumor gene expression profiles. Furthermore, methods to rank the genes according to their importance for the classification are explored. The approaches are illustrated by exploratory studies using two examples of retrospective clinical data from routine tests; diagnostic prediction of small round blue cell tumors (SRBCT) of childhood and determining the estrogen receptor (ER) status of sporadic breast cancer. The classification performance is gauged using blind tests. These studies demonstrate the feasibility of machine learning-based molecular cancer classification.

Adult↗

Physiogenomic resources for rat models of heart, lung and blood disorders.

Cardiovascular disorders are influenced by genetic and environmental factors. The TIGR rodent expression web-based resource (TREX) contains over 2,200 microarray hybridizations, involving over 800 animals from 18 different rat strains. These strains comprise genetically diverse parental animals and a panel of chromosomal substitution strains derived by introgressing individual chromosomes from normotensive Brown Norway (BN/NHsdMcwi) rats into the background of Dahl salt sensitive (SS/JrHsdMcwi) rats. The profiles document gene-expression changes in both genders, four tissues (heart, lung, liver, kidney) and two environmental conditions (normoxia, hypoxia). This translates into almost 400 high-quality direct comparisons (not including replicates) and over 100,000 pairwise comparisons. As each individual chromosomal substitution strain represents on average less than a 5% change from the parental genome, consomic strains provide a useful mechanism to dissect complex traits and identify causative genes. We performed a variety of data-mining manipulations on the profiles and used complementary physiological data from the PhysGen resource to demonstrate how TREX can be used by the cardiovascular community for hypothesis generation.

Animals↗

Charting gene regulatory networks: strategies, challenges and perspectives.

One of the foremost challenges in the post-genomic era will be to chart the gene regulatory networks of cells, including aspects such as genome annotation, identification of cis-regulatory elements and transcription factors, information on protein-DNA and protein-protein interactions, and data mining and integration. Some of these broad sets of data have already been assembled for building networks of gene regulation. Even though these datasets are still far from comprehensive, and the approach faces many important and difficult challenges, some strategies have begun to make connections between disparate regulatory events and to foster new hypotheses. In this article we review several different genomics and proteomics technologies, and present bioinformatics methods for exploring these data in order to make novel discoveries.

Animals↗

Identification of the maturation factor for dual oxidase. Evolution of an eukaryotic operon equivalent.

Dual oxidase 2 (DUOX2), an NADPH:O(2) oxidoreductase flavoprotein, is a component of the thyroid H(2)O(2) generator crucial for hormone synthesis at the apical membrane. Mutations in DUOX2 produce congenital hypothyroidism in humans. However, no functional DUOX-based NADPH oxidase has ever been reconstituted at the plasma membrane of transfected cells. It has been proposed that DUOX retention in the endoplasmatic reticulum (ER) of heterologous systems is due to the lack of an unidentified component required for functional maturation of the enzyme. By data mining of a massively parallel signature sequencing tissue expression data base, we identified an uncharacterized gene named DUOX maturation factor (DUOXA2) arranged head-to-head to and co-expressed with DUOX2. A paralog (DUOXA1) was similarly linked to DUOX1. The genomic rearrangement leading to linkage of ancient DUOX and DUOXA genes could be traced back before the divergence of echinoderms. We demonstrate that co-expression of DUOXA2, an ER-resident transmembrane protein, allows ER-to-Golgi transition, maturation, and translocation to the plasma membrane of functional DUOX2 in a heterologous system. The identification of DUOXA genes has important implications for studies of the molecular mechanisms controlling DUOX expression and the molecular genetics of congenital hypothyroidism.

Amino Acid Sequence↗

Data pre-processing in liquid chromatography-mass spectrometry-based proteomics.

MOTIVATION: In a liquid chromatography-mass spectrometry (LC-MS)-based expressional proteomics, multiple samples from different groups are analyzed in parallel. It is necessary to develop a data mining system to perform peak quantification, peak alignment and data quality assurance. RESULTS: We have developed an algorithm for spectrum deconvolution. A two-step alignment algorithm is proposed for recognizing peaks generated by the same peptide but detected in different samples. The quality of LC-MS data is evaluated using statistical tests and alignment quality tests. AVAILABILITY: Xalign software is available upon request from the author.

Algorithms↗

Enhancing instance-based classification with local density: a new algorithm for classifying unbalanced biomedical data.

MOTIVATION: Classification is an important data mining task in biomedicine. In particular, classification on biomedical data often claims the separation of pathological and healthy samples with highest discriminatory performance for diagnostic issues. Even more important than the overall accuracy is the balance of a classifier, particularly if datasets of unbalanced class size are examined. RESULTS: We present a novel instance-based classification technique which takes both information of different local density of data objects and local cluster structures into account. Our method, which adopts the basic ideas of density-based outlier detection, determines the local point density in the neighborhood of an object to be classified and of all clusters in the corresponding region. A data object is assigned to that class where it fits best into the local cluster structure. The experimental evaluation on biomedical data demonstrates that our approach outperforms most popular classification methods. AVAILABILITY: The algorithm LCF is available for testing under http://biomed.umit.at/upload/lcfx.zip.

Algorithms↗

BTW: a web server for Boltzmann time warping of gene expression time series.

UNLABELLED: Dynamic time warping (DTW) is a well-known quadratic time algorithm to determine the smallest distance and optimal alignment between two numerical sequences, possibly of different length. Originally developed for speech recognition, this method has been used in data mining, medicine and bioinformatics. For gene expression time series data, time warping distance is arguably a more flexible tool to determine genes having similar temporal expression, hence possibly related biological function, than either Euclidean distance or correlation coefficient--especially since time warping accommodates sequences of different length. The BTW web server allows a user to upload two tab-separated text files A,B of gene expression data, each possibly having a different number of time intervals of different durations. BTW then computes time warping distance between each gene of A with each gene of B, using a recently developed symmetric algorithm which additionally computes the Boltzmann partition function and outputs Boltzmann pair probabilities. The Boltzmann pair probabilities, not available with any other existent software, suggest possible biological significance of certain positions in an optimal time warping alignment. AVAILABILITY: http://bioinformatics.bc.edu/clotelab/BTW/.

Algorithms↗