Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Fullerene data mining using bibliometrics and database tomography

Database tomography (DT) is a textual database analysis system consisting of two major components: (1) algorithms for extracting multiword phrase frequencies and phrase proximities (physical closeness of the multiword technical phrases) from any type of large textual database, to augment (2) interpretative capabilities of the expert human analyst. DT was used to derive technical intelligence from a fullerenes database derived from the Science Citation Index and the Engineering Compendex. Phrase frequency analysis by the technical domain experts provided the pervasive technical themes of the fullerenes database, and phrase proximity analysis provided the relationships among the pervasive technical themes. Bibliometric analysis of the fullerenes literature supplemented the DT results with author/journal/institution publication and citation data. Comparisons of fullerenes results with past analyses of similarly structured near-earth space, chemistry, hypersonic/supersonic flow, aircraft, and ship hydrodynamics databases are made. One important finding is that many of the normalized bibliometric distribution functions are extremely consistent across these diverse technical domains and could reasonably be expected to apply to broader chemical topics than fullerenes that span multiple structural classes. Finally, lessons learned about integrating the technical domain experts with the data mining tools are presented.

Journal Article↗

RNA expression profiles and data mining of sugarcane response to low temperature.

Tropical and subtropical plants are generally sensitive to cold and can show appreciable variation in their response to cold stress when exposed to low positive temperatures. Using nylon filter arrays, we analyzed the expression profile of 1,536 expressed sequence tags (ESTs) of sugarcane (Saccharum sp. cv SP80-3280) exposed to cold for 3 to 48 h. Thirty-four cold-inducible ESTs were identified, of which 20 were novel cold-responsive genes that had not previously been reported as being cold inducible, including cellulose synthase, ABI3-interacting protein 2, a negative transcription regulator, phosphate transporter, and others, as well as several unknown genes. In addition, 25 ESTs were identified as being down-regulated during cold exposure. Using a database of cold-regulated proteins reported for other plants, we searched for homologs in the sugarcane EST project database (SUCEST), which contains 263,000 ESTs. Thirty-three homologous putative cold-regulated proteins were identified in the SUCEST database. On the basis of the expression profiles of the cold-inducible genes and the data-mining results, we propose a molecular model for the sugarcane response to low temperature.

Cold Temperature↗

Data mining for regulatory elements in yeast genome.

We have examined methods and developed a general software tool for finding and analyzing combinations of transcription factor binding sites that occur relatively often in gene upstream regions (putative promoter regions) in the yeast genome. Such frequently occurring combinations may be essential parts of possible promoter classes. The regions upstream to all genes were first isolated from the yeast genome database MIPS using the information in the annotation files of the database. The ones that do not overlap with coding regions were chosen for further studies. Next, all occurrences of the yeast transcription factor binding sites, as given in the IMD database, were located in the genome and in the selected regions in particular. Finally, by using a general purpose data mining software in combination with our own software, which parametrizes the search, we can find the combinations of binding sites that occur in the upstream regions more frequently than would be expected on the basis of the frequency of individual sites. The procedure also finds so-called association rules present in such combinations. The developed tool is available for use through the WWW.

Binding Sites↗

Data-mining analyses of pharmacovigilance signals in relation to relevant comparison drugs.

OBJECTIVE: The aim of this paper is to demonstrate the usefulness of the Bayesian Confidence Propagation Neural Network (BCPNN) in the detection of drug-specific and drug-group effects in the database of adverse drug reactions of the World Health Organization Programme for International Drug Monitoring. METHODS: Examples of drug-adverse reaction combinations highlighted by the BCPNN as quantitative associations were selected. The anatomical therapeutic chemical (ATC) group to which the drug belonged was then identified, and the information component (IC) was calculated for this ATC group and the adverse drug reaction (ADR). The IC of the ATC group with the ADR was then compared with the IC of the drug-ADR by plotting the change in IC and its 95% confidence limit over time for both. RESULTS: The chosen examples show that the BCPNN data-mining approach can identify drug-specific as well as group effects. In the known examples that served as test cases, beta-blocking agents other than practolol are not associated with sclerosing peritonitis, but all angiotensin-converting enzyme inhibitors are associated with coughing, as are antihistamines with heart-rhythm disorders and antipsychotics with myocarditis. The recently identified association between antipsychotics and myocarditis remains even after consideration of concomitant medication. CONCLUSION: The BCPNN can be used to improve the ability of a signal detection system to highlight group and drug-specific effects.

Adverse Drug Reaction Reporting Systems↗

An expressed sequence tag (EST) data mining strategy succeeding in the discovery of new G-protein coupled receptors.

We have developed a comprehensive expressed sequence tag database search method and used it for the identification of new members of the G-protein coupled receptor superfamily. Our approach proved to be especially useful for the detection of expressed sequence tag sequences that do not encode conserved parts of a protein, making it an ideal tool for the identification of members of divergent protein families or of protein parts without conserved domain structures in the expressed sequence tag database. At least 14 of the expressed sequence tags found with this strategy are promising candidates for new putative G-protein coupled receptors. Here, we describe the sequence and expression analysis of five new members of this receptor superfamily, namely GPR84, GPR86, GPR87, GPR90 and GPR91. We also studied the genomic structure and chromosomal localization of the respective genes applying in silico methods. A cluster of six closely related G-protein coupled receptors was found on the human chromosome 3q24-3q25. It consists of four orphan receptors (GPR86, GPR87, GPR91, and H963), the purinergic receptor P2Y1, and the uridine 5'-diphosphoglucose receptor KIAA0001. It seems likely that these receptors evolved from a common ancestor and therefore might have related ligands. In conclusion, we describe a data mining procedure that proved to be useful for the identification and first characterization of new genes and is well applicable for other gene families.

Amino Acid Motifs↗

Discovery of a novel murine type C retrovirus by data mining.

Analysis of genomic and expression data allows both identification and characterization of novel retroviruses. We describe a recombinant type C murine retrovirus, similar to the Mus dunni endogenous retrovirus, with VL30-like long terminal repeats and murine leukemia virus-like coding sequences. This virus is present in multiple copies in the mouse genome and expressed in a range of mouse tissues.

Amino Acid Sequence↗

Data mining.

Explore the source record for details and available documents.

Biotechnology↗

The Merck Gene Index browser: an extensible data integration system for gene finding, gene characterization and EST data mining.

MOTIVATION: To make effective use of the vast amounts of expressed sequence tag (EST) sequence data generated by the Merck-sponsored EST project and other similar efforts, sequences must be organized into gene classes, and scientists must be able to 'mine' the gene class data in the context of related genomic data. RESULTS: This paper presents the Merck Gene Index browser, an easily extensible, World Wide Web-based system for mining the Merck Gene Index (MGI) and related genomic data. The MGI is a non-redundant set of clones and sequences, each representing a distinct gene, constructed from all high-quality 3' EST sequences generated by the Merck-sponsored EST project. The MGI browser integrates data from a variety of sources and storage formats, both local and remote, using an eclectic integration strategy, including a federation of relational databases, a local data warehouse and simple hypertext links. Data currently integrated include: LENS cDNA clone and EST data, dbEST protein and non-EST nucleic acid similarity data, WashU sequence chromatograms. Entrez sequence and Medline entries, and UniGene gene clusters. Flatfile sequence data are accessed using the Bioapps server, an internally developed client-server system that supports generic sequence analysis applications. Browser data are retrieved and formatted by means of the Bioinformatics Data Integration Toolkit (B-DIT), a new suite of Perl routines.

Abstracting and Indexing↗

A semiautomated approach to gene discovery through expressed sequence tag data mining: discovery of new human transporter genes.

Identification and functional characterization of the genes in the human genome remain a major challenge. A principal source of publicly available information used for this purpose is the National Center for Biotechnology Information database of expressed sequence tags (dbEST), which contains over 4 million human ESTs. To extract the information buried in this data more effectively, we have developed a semiautomated method to mine dbEST for uncharacterized human genes. Starting with a single protein input sequence, a family of related proteins from all species is compiled. This entire family is then used to mine the human EST database for new gene candidates. Evaluation of putative new gene candidates in the context of a family of characterized proteins provides a framework for inference of the structure and function of the new genes. When applied to a test data set of 28 families within the major facilitator superfamily (MFS) of membrane transporters, our protocol found 73 previously characterized human MFS genes and 43 new MFS gene candidates. Development of this approach provided insights into the problems and pitfalls of automated data mining using public databases.

Biological Transport, Active↗

Prospecting for gold in the data mine.

As a transaction based industry, health care is data rich. A new generation of business-supporting information technology is emerging that can transform such data into knowledge critical to sustain value in health care. In order to adopt successfully this new information technology, physicians and other health care leaders must come to understand the value of complete, accurate and consistent coding of clinical activities. The clinical laboratory has a pervasive role in health care. With its recent federally assigned responsibility to assure clinically relevant testing through ICD-9-CM and CPT coding, and its experience with computerized information systems, the clinical laboratory is in an ideal position to become a champion of the new information technology.

Clinical Laboratory Information Systems↗

Data mining of inputs: analysing magnitude and functional measures.

The problem of data encoding and feature selection for training back-propagation neural networks is well known. The basic principles are to avoid encrypting the underlying structure of the data, and to avoid using irrelevant inputs. This is not easy in the real world, where we often receive data which has been processed by at least one previous user. The data may contain too many instances of some class, and too few instances of other classes. Real data sets often include many irrelevant or redundant input fields. This paper examines the use of weight matrix analysis techniques and functional measures using two real (and hence noisy) data sets. The first part of this paper examines the use of the weight matrix of the trained neural network itself to determine which inputs are significant. A new technique is introduced and compared with two other techniques from the literature. We present our experience and results on some satellite data augmented by a terrain model. The task was to predict the forest supra-type based on the available information. A brute force technique eliminating randomly selected inputs was used to validate our approach. The second part of this paper examines the use of measures to determine the functional contribution of inputs to outputs. Inputs which include minor but unique information to the network are more significant than inputs with higher magnitude contribution but providing redundant information, which is also provided by another input. A comparison is made to sensitivity analysis, where the sensitivity of outputs to input perturbation is used as a measure of the significance of inputs. This paper presents a novel functional analysis of the weight matrix based on a technique developed for determining the behavioral significance of hidden neurons. This is compared with the application of the same technique to the training and test data. Finally, a novel aggregation technique is introduced.

Algorithms↗

The University of Minnesota Biocatalysis/Biodegradation Database: post-genomic data mining.

The University of Minnesota Biocatalysis/Biodegradation Database (UM-BBD, http://umbbd.ahc.umn.edu/) provides curated information on microbial catabolism and related biotransformations, primarily for environmental pollutants. Currently, it contains information on over 130 metabolic pathways, 800 reactions, 750 compounds and 500 enzymes. In the past two years, it has increased its breath to include more examples of microbial metabolism of metals and metalloids; and expanded the types of information it includes to contain microbial biotransformations of, and binding interactions with many chemical elements. It has also increased the ways in which this data can be accessed (mined). Structure-based searching was added, for exact matches, similarity, or substructures. Analysis of UM-BBD reactions has lead to a prototype, guided, pathway prediction system. Guided prediction means that the user is shown all possible biotransformations at each step and guides the process to its conclusion. Mining the UM-BBD's data provides a unique view into how the microbial world recycles organic functional groups. UM-BBD users are encouraged to comment on all aspects of the database, including the information it contains and the tools by which it can be mined. The database and prediction system develop under the direction of the scientific community.

Biodegradation, Environmental↗

Visualization of multiple influences on ocellar flight control in giant honeybees with the data-mining tool Viscovery SOMine.

Viscovery SOMine is a software tool for advanced analysis and monitoring of numerical data sets. It was developed for professional use in business, industry, and science and to support dependency analysis, deviation detection, unsupervised clustering, nonlinear regression, data association, pattern recognition, and animated monitoring. Based on the concept of self-organizing maps (SOMs), it employs a robust variant of unsupervised neural networks--namely, Kohonen's Batch-SOM, which is further enhanced with a new scaling technique for speeding up the learning process. This tool provides a powerful means by which to analyze complex data sets without prior statistical knowledge. The data representation contained in the trained SOM is systematically converted to be used in a spectrum of visualization techniques, such as evaluating dependencies between components, investigating geometric properties of the data distribution, searching for clusters, or monitoring new data. We have used this software tool to analyze and visualize multiple influences of the ocellar system on free-flight behavior in giant honeybees. Occlusion of ocelli will affect orienting reactivities in relation to flight target, level of disturbance, and position of the bee in the flight chamber; it will induce phototaxis and make orienting imprecise and dependent on motivational settings. Ocelli permit the adjustment of orienting strategies to environmental demands by enforcing abilities such as centering or flight kinetics and by providing independent control of posture and flight course.

Animals↗

Discovery of predictive models in an injury surveillance database: an application of data mining in clinical research.

A new, evolutionary computation-based approach to discovering prediction models in surveillance data was developed and evaluated. This approach was operationalized in EpiCS, a type of learning classifier system specially adapted to model clinical data. In applying EpiCS to a large, prospective injury surveillance database, EpiCS was found to create accurate predictive models quickly that were highly robust, being able to classify > 99% of cases early during training. After training, EpiCS classified novel data more accurately (p < 0.001) than either logistic regression or decision tree induction (C4.5), two traditional methods for discovering or building predictive models.

Artificial Intelligence↗