Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

In silico analysis of 2085 clones from a normalized rat vestibular periphery 3' cDNA library.

The inserts from 2400 cDNA clones isolated from a normalized Rattus norvegicus vestibular periphery cDNA library were sequenced and characterized. The Wackym-Soares vestibular 3' cDNA library was constructed from the saccular and utricular maculae, the ampullae of all three semicircular canals and Scarpa's ganglia containing the somata of the primary afferent neurons, microdissected from 104 male and female rats. The inserts from 2400 randomly selected clones were sequenced from the 5' end. Each sequence was analyzed using the BLAST algorithm compared to the Genbank nonredundant, rat genome, mouse genome and human genome databases to search for high homology alignments. Of the initial 2400 clones, 315 (13%) were found to be of poor quality and did not yield useful information, and therefore were eliminated from the analysis. Of the remaining 2085 sequences, 918 (44%) were found to represent 758 unique genes having useful annotations that were identified in databases within the public domain or in the published literature; these sequences were designated as known characterized sequences. 1141 sequences (55%) aligned with 1011 unique sequences had no useful annotations and were designated as known but uncharacterized sequences. Of the remaining 26 sequences (1%), 24 aligned with rat genomic sequences, but none matched previously described rat expressed sequence tags or mRNAs. No significant alignment to the rat or human genomic sequences could be found for the remaining 2 sequences. Of the 2085 sequences analyzed, 86% were singletons. The known, characterized sequences were analyzed with the FatiGO online data-mining tool (http://fatigo.bioinfo.cnio.es/) to identify level 5 biological process gene ontology (GO) terms for each alignment and to group alignments with similar or identical GO terms. Numerous genes were identified that have not been previously shown to be expressed in the vestibular system. Further characterization of the novel cDNA sequences may lead to the identification of genes with vestibular-specific functions. Continued analysis of the rat vestibular periphery transcriptome should provide new insights into vestibular function and generate new hypotheses. Physiological studies are necessary to further elucidate the roles of the identified genes and novel sequences in vestibular function.

Afferent Pathways↗

Genome informatics: current status and future prospects.

This article reviews recent advances in genomics and informatics relevant to cardiovascular research. In particular, we review the status of (1) whole genome sequencing efforts in human, mouse, rat, zebrafish, and dog; (2) the development of data mining and analysis tools; (3) the launching of the National Heart, Lung, and Blood Institute Programs for Genomics Applications and Proteomics Initiative; (4) efforts to characterize the cardiac transcriptome and proteome; and (5) the current status of computational modeling of the cardiac myocyte. In each instance, we provide links to relevant sources of information on the World Wide Web and critical appraisals of the promises and the challenges of an expanding and diverse information landscape.

Animals↗

Transcriptional profiling in coronary artery disease: indications for novel markers of coronary collateralization.

BACKGROUND: The development of collateral circulation plays an important role in protecting tissues from ischemic damage, and its stimulation has emerged as one of principal approaches to therapeutic angiogenesis. Clinical observations have documented substantial differences in the extent of collateralization among patients with coronary artery disease (CAD), with some individuals demonstrating marked abundance and others showing nearly complete absence of these vessels. Recent studies have suggested that circulating monocytes play a major role in collateral growth. The present study was undertaken to determine transcriptional profiles of circulating monocytes in CAD patients with different extents of collateral growth. METHODS AND RESULTS: Monocyte transcriptomes from CAD patients with and without collateral vessels were obtained by use of high-throughput expression profiling. Using a newly developed redundancy-based data mining method, we have identified a set of molecular markers characteristic of a "noncollateralgenic" phenotype. Moreover, we show that these transcriptional abnormalities are independent of the severity of CAD or any other known clinical parameter thought to affect collateral development and correlated with protein expression levels in monocytes and plasma. CONCLUSIONS: Monocyte transcription profiling identifies sets of patients with extensive versus poorly developed collateral circulation. Thus, genetic factors may heavily influence coronary collateral vessel growth in CAD and affect prognosis and response to therapeutic interventions.

Aged↗

A Bayesian committee machine.

The Bayesian committee machine (BCM) is a novel approach to combining estimators that were trained on different data sets. Although the BCM can be applied to the combination of any kind of estimators, the main foci are gaussian process regression and related systems such as regularization networks and smoothing splines for which the degrees of freedom increase with the number of training data. Somewhat surprisingly, we find that the performance of the BCM improves if several test points are queried at the same time and is optimal if the number of test points is at least as large as the degrees of freedom of the estimator. The BCM also provides a new solution for on-line learning with potential applications to data mining. We apply the BCM to systems with fixed basis functions and discuss its relationship to gaussian process regression. Finally, we show how the ideas behind the BCM can be applied in a non-Bayesian setting to extend the input-dependent combination of estimators.

Algorithms↗

Accuracy-based learning classifier systems: models, analysis and applications to classification tasks.

Recently, Learning Classifier Systems (LCS) and particularly XCS have arisen as promising methods for classification tasks and data mining. This paper investigates two models of accuracy-based learning classifier systems on different types of classification problems. Departing from XCS, we analyze the evolution of a complete action map as a knowledge representation. We propose an alternative, UCS, which evolves a best action map more efficiently. We also investigate how the fitness pressure guides the search towards accurate classifiers. While XCS bases fitness on a reinforcement learning scheme, UCS defines fitness from a supervised learning scheme. We find significant differences in how the fitness pressure leads towards accuracy, and suggest the use of a supervised approach specially for multi-class problems and problems with unbalanced classes. We also investigate the complexity factors which arise in each type of accuracy-based LCS. We provide a model on the learning complexity of LCS which is based on the representative examples given to the system. The results and observations are also extended to a set of real world classification problems, where accuracy-based LCS are shown to perform competitively with respect to other learning algorithms. The work presents an extended analysis of accuracy-based LCS, gives insight into the understanding of the LCS dynamics, and suggests open issues for further improvement of LCS on classification tasks.

Algorithms↗

A comprehensive overview of the applications of artificial life.

We review the applications of artificial life (ALife), the creation of synthetic life on computers to study, simulate, and understand living systems. The definition and features of ALife are shown by application studies. ALife application fields treated include robot control, robot manufacturing, practical robots, computer graphics, natural phenomenon modeling, entertainment, games, music, economics, Internet, information processing, industrial design, simulation software, electronics, security, data mining, and telecommunications. In order to show the status of ALife application research, this review primarily features a survey of about 180 ALife application articles rather than a selected representation of a few articles. Evolutionary computation is the most popular method for designing such applications, but recently swarm intelligence, artificial immune network, and agent-based modeling have also produced results. Applications were initially restricted to the robotics and computer graphics, but presently, many different applications in engineering areas are of interest.

Artificial Intelligence↗

Using unsupervised learning with independent component analysis to identify patterns of glaucomatous visual field defects.

PURPOSE: Clustering by unsupervised learning with machine learning classifiers was shown to segment clusters of patterns in standard automated perimetry (SAP) for glaucoma in previous publications. In this study, unsupervised learning by independent component analysis decomposed SAP field patterns into axes, and the information represented by these axes was evaluated. METHODS: SAP fields were used that were obtained with the Humphrey Visual Field Analyzer (Carl Zeiss Meditec, Dublin, CA) from 189 normal eyes and 156 eyes with glaucomatous optic neuropathy (GON) determined by masked review with stereoscopic optic disc photographs. The variational Bayesian independent component analysis mixture model (vB-ICA-mm) partitioned the SAP fields into the most informative number of clusters. Simultaneously, the model learned an optimal number of maximally independent axes for each cluster. RESULTS: The most informative number of clusters in the SAP set was two. vB-ICA-mm placed 68.6% of the eyes with GON in a cluster labeled G and 98.4% of the eyes with normal optic discs in a cluster labeled N. Cluster G optimally contained six axes. Post hoc analysis of patterns generated at -1 SD and +2 SD from the cluster G mean on the six axes revealed defects similar to those identified by experts as indicative of glaucoma. SAP fields associated with an axis showed increasing severity, as they were located farther in the positive direction from the cluster G mean. CONCLUSIONS: vB-ICA-mm represented the SAP fields with patterns that were meaningful for glaucoma experts. This process also captured severity in the patterns uncovered. These findings should validate vB-ICA-mm as a data-mining technique for new and unfamiliar complex tests.

Artificial Intelligence↗

Developing a practical forecasting screener for domestic violence incidents.

In this article, the authors report on the development of a short screening tool that deputies in the Los Angeles Sheriff's Department could use in the field to help forecast domestic violence incidents in particular households. The data come from more than 500 households to which sheriff's deputies were dispatched in fall 2003. Information on potential predictors was collected at the scene. Outcomes were measured during a 3-month follow-up. Data were analyzed with modern data-mining procedures in which true forecasts were evaluated. A screening instrument was developed based on a small fraction of the information collected. Making the screening instrument more complicated did not improve forecasting skill. Taking the relative costs of false positives and false negatives into account, the instrument correctly forecasted future calls for service about 60% of the time. Future calls involving domestic violence misdemeanors and felonies were correctly forecast about 50% of the time. The 50% figure is important because such calls require a law enforcement response and yet are a relatively small fraction of all domestic violence calls for service.

Domestic Violence↗

A neural network applied to criminal psychological profiling: an Italian initiative.

The author presents a brief discussion of criminal profiling followed by an introduction to the Italian Neural Network for Psychological Criminal Profiling (NNPCP) project. This project, based on a so-called neural network and data mining, is an innovative technique being developed with the intention of extending criminal profiling to single serious crimes through the use of a computerized database.

Crime↗

Identification of gap junction blockers using automated fluorescence microscopy imaging.

Gap junctions coordinate electrical signals and facilitate metabolic synchronization between cells. In this study, the authors have developed a novel assay for the identification of gap junction blockers using fluorescence microscopy imaging-based high-content screening technology. In the assay, the communication between neighboring cells through gap junctions was measured by following the redistribution of a fluorescent marker. The movement of calcein dye from dye-loaded donor cells to dye-free acceptor cells through gap junctions overexpressed on cell surface membranes was monitored using automated fluorescence microscopy imaging in a high-throughput compatible format. The fluorescence imaging technology consisted of automated focusing, image acquisition, image processing, and data mining. The authors have successfully performed a high-throughput screening of a 486,000- compound program with this assay, and they were able to identify false positives without additional experiments. Selective and pharmacologically interesting compounds were identified for further optimization.

Animals↗

Virtual ligand screening against Escherichia coli dihydrofolate reductase: improving docking enrichment using physics-based methods.

Motivated by their participation in the McMaster Data-Mining and Docking Competition, the authors developed 2 new computational technologies and applied them to docking against Escherichia coli dihydrofolate reductase: a receptor preparation procedure that incorporates rotamer optimization of side chains and a physics-based rescoring procedure for estimating relative binding affinities of the protein-ligand complexes. Both methods use the same energy function, consisting of the all-atom OPLS-AA force field and a generalized Born solvent model, which treats the protein receptor and small-molecule ligands in a consistent manner. Thus, the energy function is similar to that used in more sophisticated approaches, such as free-energy perturbation and the molecular mechanics Poisson-Boltzmann/surface area, but sampling during the rescoring procedure is limited to simple energy minimization of the ligand. The use of a highly efficient minimization algorithm permitted the authors to apply this rescoring procedure to hundreds of thousands of protein-ligand complexes during the competition, using a modest Linux cluster. To test these methods, they used the 12 competitive inhibitors identified in the training set, plus methotrexate, as positive controls in enrichment studies with both the training and test sets, each containing 50,000 compounds. The key conclusion is that combining the receptor preparation and rescoring methods makes it possible to identify most of the positive controls within the top few tenths of a percent of the rank-ordered training and test set libraries.

Computational Biology↗

Evaluating the high-throughput screening computations.

The judges evaluated the submissions for the McMaster University High-Throughput Data-Mining and Docking Competition based on 3 criteria: identification of active compounds, percent enrichment, and overview of the competition. Using these metrics, 4 of the participating groups found meaningful enrichment, and 3 groups made perceptive comments about the general nature of the competition.

Computational Biology↗

Bioinformatics for cancer management in the post-genome era.

Human cancer is caused by multiple factors, such as genetic predisposition, chronic persistent inflammation, environmental factors, life style, and aging. Dysregulated proliferation, dysregulated adhesion, resistance to apoptosis, resistance to senescence, and resistance to anti-cancer drugs are features of cancer cells. Accumulation of multiple epigenetic changes and genetic alterations of cancer-associated genes during multi-stage carcinogenesis results in more malignant phenotypes. Post-genome science is characterized by omics data related to genome, transcriptome, proteome, metabolome, interactome, and epigenome as well as by high-throughput technology, such as whole-genome tiling oligonucleotide array, array CGH with 32,433 overlapping BAC clones, transcriptome microarray, mass spectrometry, tissue-based expression array, and cell-based transfection array. Benchtop oncology supplies Desktop oncology with large amounts of omics data produced by high-throughput technology. Desktop oncology establishes knowledge on cancer-related biomarkers, such as predisposition markers, diagnostic markers, prognostic markers, and therapeutic markers, by using bioinformatics and human intelligence of experts for data mining and text mining. Bedside oncology applies the knowledge established by Desktop oncology to determine therapeutics for cancer patients. Antibody drugs (Trastuzumab/Herceptin, Cetuximab/Erbitux, Bevacizumab/Avastin, et cetera), small molecule inhibitors for tyrosine kinases (Gefitinib/Iressa, Erlotinib/Tarceva, Imatinib/Gleevec, et cetera), conventional cytotoxic drugs, and anti-hormonal drugs are used for cancer chemotherapy. Biomarker monitoring contributes to therapeutic optional choice and drug dosage determination for cancer patients. Knowledge on biomarkers is feedforwarded from desktop to bedside in the translational research, and then biomarker monitoring is feedbacked from bedside to desktop in the reverse translational research. Desktop oncology is indispensable for cancer research in the post-genome era. Combination of genetic screening for cancer predisposition in the general population and precise selection of therapeutic options during cancer management could contribute to the realization of personalized prevention and to dramatically improve the prognosis of cancer patients in the future.

Antineoplastic Agents↗

TETRA: a web-service and a stand-alone program for the analysis and comparison of tetranucleotide usage patterns in DNA sequences.

BACKGROUND: In the emerging field of environmental genomics, direct cloning and sequencing of genomic fragments from complex microbial communities has proven to be a valuable source of new enzymes, expanding the knowledge of basic biological processes. The central problem of this so called metagenome-approach is that the cloned fragments often lack suitable phylogenetic marker genes, rendering the identification of clones that are likely to originate from the same genome difficult or impossible. In such cases, the analysis of intrinsic DNA-signatures like tetranucleotide frequencies can provide valuable hints on fragment affiliation. With this application in mind, the TETRA web-service and the TETRA stand-alone program have been developed, both of which automate the task of comparative tetranucleotide frequency analysis. AVAILABILITY: http://www.megx.net/tetra. RESULTS: TETRA provides a statistical analysis of tetranucleotide usage patterns in genomic fragments, either via a web-service or a stand-alone program. With respect to discriminatory power, such an analysis outperforms the assignment of genomic fragments based on the (G+C)-content, which is a widely-used sequence-based measure for assessing fragment relatedness. While the web-service is restricted to the calculation of correlation coefficients between tetranucleotide usage patterns of submitted DNA sequences, the stand-alone program generates a much more detailed output, comprising all raw data and graphical plots. The stand-alone program is controlled via a graphical user interface and can batch-process a multitude of sequences. Furthermore, it comes with pre-computed tetranucleotide usage patterns for 166 prokaryote chromosomes, providing a useful reference dataset and source for data-mining. CONCLUSIONS: Up to now, the analysis of skewed oligonucleotide distributions within DNA sequences is not a commonly used tool within metagenomics. With the TETRA web-service and stand-alone program, the method is now accessible in an easy to use manner for a broad audience. This will hopefully facilitate the interrelation of genomic fragments from metagenome libraries, ultimately leading to new insights into the genetic potentials of yet uncultured microorganisms.

Base Composition↗

Predicting binding sites of hydrolase-inhibitor complexes by combining several methods.

BACKGROUND: Protein-protein interactions play a critical role in protein function. Completion of many genomes is being followed rapidly by major efforts to identify interacting protein pairs experimentally in order to decipher the networks of interacting, coordinated-in-action proteins. Identification of protein-protein interaction sites and detection of specific amino acids that contribute to the specificity and the strength of protein interactions is an important problem with broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. RESULTS: In order to increase the power of predictive methods for protein-protein interaction sites, we have developed a consensus methodology for combining four different methods. These approaches include: data mining using Support Vector Machines, threading through protein structures, prediction of conserved residues on the protein surface by analysis of phylogenetic trees, and the Conservatism of Conservatism method of Mirny and Shakhnovich. Results obtained on a dataset of hydrolase-inhibitor complexes demonstrate that the combination of all four methods yield improved predictions over the individual methods. CONCLUSIONS: We developed a consensus method for predicting protein-protein interface residues by combining sequence and structure-based methods. The success of our consensus approach suggests that similar methodologies can be developed to improve prediction accuracies for other bioinformatic problems.

Algorithms↗

B.E.A.R. GeneInfo: a tool for identifying gene-related biomedical publications through user modifiable queries.

BACKGROUND: Once specific genes are identified through high throughput genomics technologies there is a need to sort the final gene list to a manageable size for validation studies. The triaging and sorting of genes often relies on the use of supplemental information related to gene structure, metabolic pathways, and chromosomal location. Yet in disease states where the genes may not have identifiable structural elements, poorly defined metabolic pathways, or limited chromosomal data, flexible systems for obtaining additional data are necessary. In these situations having a tool for searching the biomedical literature using the list of identified genes while simultaneously defining additional search terms would be useful. RESULTS: We have built a tool, BEAR GeneInfo, that allows flexible searches based on the investigators knowledge of the biological process, thus allowing for data mining that is specific to the scientist's strengths and interests. This tool allows a user to upload a series of GenBank accession numbers, Unigene Ids, Locuslink Ids, or gene names. BEAR GeneInfo takes these IDs and identifies the associated gene names, and uses the lists of gene names to query PubMed. The investigator can add additional modifying search terms to the query. The subsequent output provides a list of publications, along with the associated reference hyperlinks, for reviewing the identified articles for relevance and interest. An example of the use of this tool in the study of human prostate cancer cells treated with Selenium is presented. CONCLUSIONS: This tool can be used to further define a list of genes that have been identified through genomic or genetic studies. Through the use of targeted searches with additional search terms the investigator can limit the list to genes that match their specific research interests or needs. The tool is freely available on the web at http://prostategenomics.org1, and the authors will provide scripts and database components if requested mdatta@mcw.edu

Databases, Genetic↗

Building a protein name dictionary from full text: a machine learning term extraction approach.

BACKGROUND: The majority of information in the biological literature resides in full text articles, instead of abstracts. Yet, abstracts remain the focus of many publicly available literature data mining tools. Most literature mining tools rely on pre-existing lexicons of biological names, often extracted from curated gene or protein databases. This is a limitation, because such databases have low coverage of the many name variants which are used to refer to biological entities in the literature. RESULTS: We present an approach to recognize named entities in full text. The approach collects high frequency terms in an article, and uses support vector machines (SVM) to identify biological entity names. It is also computationally efficient and robust to noise commonly found in full text material. We use the method to create a protein name dictionary from a set of 80,528 full text articles. Only 8.3% of the names in this dictionary match SwissProt description lines. We assess the quality of the dictionary by studying its protein name recognition performance in full text. CONCLUSION: This dictionary term lookup method compares favourably to other published methods, supporting the significance of our direct extraction approach. The method is strong in recognizing name variants not found in SwissProt.

Abstracting and Indexing↗

Biclustering of gene expression data by Non-smooth Non-negative Matrix Factorization.

BACKGROUND: The extended use of microarray technologies has enabled the generation and accumulation of gene expression datasets that contain expression levels of thousands of genes across tens or hundreds of different experimental conditions. One of the major challenges in the analysis of such datasets is to discover local structures composed by sets of genes that show coherent expression patterns across subsets of experimental conditions. These patterns may provide clues about the main biological processes associated to different physiological states. RESULTS: In this work we present a methodology able to cluster genes and conditions highly related in sub-portions of the data. Our approach is based on a new data mining technique, Non-smooth Non-Negative Matrix Factorization (nsNMF), able to identify localized patterns in large datasets. We assessed the potential of this methodology analyzing several synthetic datasets as well as two large and heterogeneous sets of gene expression profiles. In all cases the method was able to identify localized features related to sets of genes that show consistent expression patterns across subsets of experimental conditions. The uncovered structures showed a clear biological meaning in terms of relationships among functional annotations of genes and the phenotypes or physiological states of the associated conditions. CONCLUSION: The proposed approach can be a useful tool to analyze large and heterogeneous gene expression datasets. The method is able to identify complex relationships among genes and conditions that are difficult to identify by standard clustering algorithms.

Algorithms↗