Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

An expressed sequence tag (EST) data mining strategy succeeding in the discovery of new G-protein coupled receptors.

We have developed a comprehensive expressed sequence tag database search method and used it for the identification of new members of the G-protein coupled receptor superfamily. Our approach proved to be especially useful for the detection of expressed sequence tag sequences that do not encode conserved parts of a protein, making it an ideal tool for the identification of members of divergent protein families or of protein parts without conserved domain structures in the expressed sequence tag database. At least 14 of the expressed sequence tags found with this strategy are promising candidates for new putative G-protein coupled receptors. Here, we describe the sequence and expression analysis of five new members of this receptor superfamily, namely GPR84, GPR86, GPR87, GPR90 and GPR91. We also studied the genomic structure and chromosomal localization of the respective genes applying in silico methods. A cluster of six closely related G-protein coupled receptors was found on the human chromosome 3q24-3q25. It consists of four orphan receptors (GPR86, GPR87, GPR91, and H963), the purinergic receptor P2Y1, and the uridine 5'-diphosphoglucose receptor KIAA0001. It seems likely that these receptors evolved from a common ancestor and therefore might have related ligands. In conclusion, we describe a data mining procedure that proved to be useful for the identification and first characterization of new genes and is well applicable for other gene families.

Amino Acid Motifs↗

Insufficient filling of vacuum tubes as a cause of microhemolysis and elevated serum lactate dehydrogenase levels. Use of a data-mining technique in evaluation of questionable laboratory test results.

Experienced physicians noted unexpectedly elevated concentrations of lactate dehydrogenase in some patient samples, but quality control specimens showed no bias. To evaluate this problem, we used a "latent reference individual extraction method", designed to obtain reference intervals from a laboratory database by excluding individuals who have abnormal results for basic analytes other than the analyte in question, in this case lactate dehydrogenase. The reference interval derived for the suspected year was 264-530 U/L, while that of the previous year was 248-495 U/L. The only change we found was the introduction of an order entry system, which requests precise sampling volumes rather than complete filling of vacuum tubes. The effect of vacuum persistence was tested using ten freshly drawn blood samples. Compared with complete filling, 1/5 filling resulted in average elevations of lactate dehydrogenase, aspartic aminotransferase, and potassium levels of 8.0%, 3.8%, and 3.4%, respectively (all p<0.01). Microhemolysis was confirmed using a urine stick method. The length of time before centrifugation determined the degree of hemolysis, while vacuum during centrifugation did not affect it. Microhemolysis is the probable cause of the suspected pseudo-elevation noted by the physicians. Data-mining methodology represents a valuable tool for monitoring long-term bias in laboratory results.

Adult↗

An application of a genetic algorithm in conjunction with other data mining methods for estimating outcome after hospitalization in cancer patients.

BACKGROUND: We investigated which factors predicted the risk of in-hospital mortality in a general population of cancer patients with non-terminal disease and whether employing the genetic algorithm technique would be useful in this regard. MATERIAL/METHODS: A total of 201 cancer patients, including all cases of in-hospital mortality over a 2-year period, as well as a control group of subjects discharged during the same period, all having an Eastern Cooperative Oncology Group (ECOG) performance status of of < or =3 at the time of admission, were retrospectively evaluated. Indicators of in-hospital mortality were determined by multivariate logistic regression, recursive partitioning analysis, neural network, and genetic algorithm (GA) techniques. The performance of the different techniques were compared by a number of measures, including receiver operating curve (ROC) analysis. RESULTS: All four analysis methods selected a combination of six explanatory variables to explain the risk of in-hospital mortality: lactate dehydrogenase (LDH), alanine transaminase (ALT), hemoglobin (Hb), white blood cell counts (Wbc), type of cancer, and reason for admission. Compared with the other 3 methods, GA selected the least number of explanatory variables, i.e. LDH and reason for admission, with similar fraction of cases explained (78.6%), and yielded a fitness score of 0.52. CONCLUSIONS: LDH is an important indicator of in-hospital mortality for hospitalized cancer patients not in terminal stage. GA reliably predicted in-hospital mortality and was shown to be as efficient as the other data mining techniques employed in this study. Its use in a clinical setting for prognostication in oncology appears promising.

Adolescent↗

Discovery of a novel murine type C retrovirus by data mining.

Analysis of genomic and expression data allows both identification and characterization of novel retroviruses. We describe a recombinant type C murine retrovirus, similar to the Mus dunni endogenous retrovirus, with VL30-like long terminal repeats and murine leukemia virus-like coding sequences. This virus is present in multiple copies in the mouse genome and expressed in a range of mouse tissues.

Amino Acid Sequence↗

Experimental screening of dihydrofolate reductase yields a "test set" of 50,000 small molecules for a computational data-mining and docking competition.

High-throughput screening (HTS) generates an abundance of data that are a valuable resource to be mined. Dockers and data miners can use "real-world" HTS data to test and further develop their tools. A screen of 50,000 diverse small molecules was carried out against Escherichia coli dihydrofolate reductase (DHFR) and compared with a previous screen of 50,000 compounds against the same target. Identical assays and conditions were maintained for both studies. Prior to the completion of the second screen, the original screening data were publicly released for use as a "training set", and computational chemists and data analysts were challenged to predict the activity of compounds in this second "test set". Upon completion, the primary screen of the test set generated no potent inhibitors of DHFR activity.

Computational Biology↗

Methods for mining HTS data.

Data mining is a fast-growing field that is finding application across a wide range of industries. HTS is a crucial part of the drug discovery process at most large pharmaceutical companies. Accurate analysis of HTS data is, therefore, vital to drug discovery. Given the large quantity of data generated during an HTS, and the importance of analyzing those data effectively, it is unsurprising that data-mining techniques are now increasingly applied to HTS data analysis. Taking a broad view of both the HTS process and the data-mining process, we review recent literature that describes the application of data-mining techniques to HTS data.

Animals↗

Data mining.

Explore the source record for details and available documents.

Biotechnology↗

The use of data mining to investigate a possible quality problem with ultrasensitive HIV viral load data at a large reference laboratory.

Suppression of HIV viral load to <50 copies/mL, the lower limit of detection for the ultrasensitive assays, has been shown to correlate with favorable clinical outcome. Patients periodically exhibited transient or sustained low-level viremia based on this test. In order to investigate a possible quality concern, we used our corporate data warehouse to examine the patterns in our data over time as well as across geographic regions.

Data Interpretation, Statistical↗

Predicting crystal structure by merging data mining with quantum mechanics.

Modern methods of quantum mechanics have proved to be effective tools to understand and even predict materials properties. An essential element of the materials design process, relevant to both new materials and the optimization of existing ones, is knowing which crystal structures will form in an alloy system. Crystal structure can only be predicted effectively with quantum mechanics if an algorithm to direct the search through the large space of possible structures is found. We present a new approach to the prediction of structure that rigorously mines correlations embodied within experimental data and uses them to direct quantum mechanical techniques efficiently towards the stable crystal structure of materials.

Journal Article↗

The Merck Gene Index browser: an extensible data integration system for gene finding, gene characterization and EST data mining.

MOTIVATION: To make effective use of the vast amounts of expressed sequence tag (EST) sequence data generated by the Merck-sponsored EST project and other similar efforts, sequences must be organized into gene classes, and scientists must be able to 'mine' the gene class data in the context of related genomic data. RESULTS: This paper presents the Merck Gene Index browser, an easily extensible, World Wide Web-based system for mining the Merck Gene Index (MGI) and related genomic data. The MGI is a non-redundant set of clones and sequences, each representing a distinct gene, constructed from all high-quality 3' EST sequences generated by the Merck-sponsored EST project. The MGI browser integrates data from a variety of sources and storage formats, both local and remote, using an eclectic integration strategy, including a federation of relational databases, a local data warehouse and simple hypertext links. Data currently integrated include: LENS cDNA clone and EST data, dbEST protein and non-EST nucleic acid similarity data, WashU sequence chromatograms. Entrez sequence and Medline entries, and UniGene gene clusters. Flatfile sequence data are accessed using the Bioapps server, an internally developed client-server system that supports generic sequence analysis applications. Browser data are retrieved and formatted by means of the Bioinformatics Data Integration Toolkit (B-DIT), a new suite of Perl routines.

Abstracting and Indexing↗

Support vector machines in HTS data mining: Type I MetAPs inhibition study.

This article reports a successful application of support vector machines (SVMs) in mining high-throughput screening (HTS) data of a type I methionine aminopeptidases (MetAPs) inhibition study. A library with 43,736 small organic molecules was used in the study, and 1355 compounds in the library with 40% or higher inhibition activity were considered as active. The data set was randomly split into a training set and a test set (3:1 ratio). The authors were able to rank compounds in the test set using their decision values predicted by SVM models that were built on the training set. They defined a novel score PT50, the percentage of the test set needed to be screened to recover 50% of the actives, to measure the performance of the models. With carefully selected parameters, SVM models increased the hit rates significantly, and 50% of the active compounds could be recovered by screening just 7% of the test set. The authors found that the size of the training set played a significant role in the performance of the models. A training set with 10,000 member compounds is likely the minimum size required to build a model with reasonable predictive power.

Algorithms↗

A semiautomated approach to gene discovery through expressed sequence tag data mining: discovery of new human transporter genes.

Identification and functional characterization of the genes in the human genome remain a major challenge. A principal source of publicly available information used for this purpose is the National Center for Biotechnology Information database of expressed sequence tags (dbEST), which contains over 4 million human ESTs. To extract the information buried in this data more effectively, we have developed a semiautomated method to mine dbEST for uncharacterized human genes. Starting with a single protein input sequence, a family of related proteins from all species is compiled. This entire family is then used to mine the human EST database for new gene candidates. Evaluation of putative new gene candidates in the context of a family of characterized proteins provides a framework for inference of the structure and function of the new genes. When applied to a test data set of 28 families within the major facilitator superfamily (MFS) of membrane transporters, our protocol found 73 previously characterized human MFS genes and 43 new MFS gene candidates. Development of this approach provided insights into the problems and pitfalls of automated data mining using public databases.

Biological Transport, Active↗

Prospecting for gold in the data mine.

As a transaction based industry, health care is data rich. A new generation of business-supporting information technology is emerging that can transform such data into knowledge critical to sustain value in health care. In order to adopt successfully this new information technology, physicians and other health care leaders must come to understand the value of complete, accurate and consistent coding of clinical activities. The clinical laboratory has a pervasive role in health care. With its recent federally assigned responsibility to assure clinically relevant testing through ICD-9-CM and CPT coding, and its experience with computerized information systems, the clinical laboratory is in an ideal position to become a champion of the new information technology.

Clinical Laboratory Information Systems↗

Diagnostic support for glaucoma using retinal images: a hybrid image analysis and data mining approach.

The availability of modern imaging techniques such as Confocal Scanning Laser Tomography (CSLT) for capturing high-quality optic nerve images offer the potential for developing automatic and objective methods for diagnosing glaucoma. We present a hybrid approach that features the analysis of CSLT images using moment methods to derive abstract image defining features. The features are then used to train classifers for automatically distinguishing CSLT images of normal and glaucoma patient. As a first, in this paper, we present investigations in feature subset selction methods for reducing the relatively large input space produced by the moment methods. We use neural networks and support vector machines to determine a sub-set of moments that offer high classification accuracy. We demonstratee the efficacy of our methods to discriminate between healthy and glaucomatous optic disks based on shape information automatically derived from optic disk topography and reflectance images.

Data Mining↗

Data mining of inputs: analysing magnitude and functional measures.

The problem of data encoding and feature selection for training back-propagation neural networks is well known. The basic principles are to avoid encrypting the underlying structure of the data, and to avoid using irrelevant inputs. This is not easy in the real world, where we often receive data which has been processed by at least one previous user. The data may contain too many instances of some class, and too few instances of other classes. Real data sets often include many irrelevant or redundant input fields. This paper examines the use of weight matrix analysis techniques and functional measures using two real (and hence noisy) data sets. The first part of this paper examines the use of the weight matrix of the trained neural network itself to determine which inputs are significant. A new technique is introduced and compared with two other techniques from the literature. We present our experience and results on some satellite data augmented by a terrain model. The task was to predict the forest supra-type based on the available information. A brute force technique eliminating randomly selected inputs was used to validate our approach. The second part of this paper examines the use of measures to determine the functional contribution of inputs to outputs. Inputs which include minor but unique information to the network are more significant than inputs with higher magnitude contribution but providing redundant information, which is also provided by another input. A comparison is made to sensitivity analysis, where the sensitivity of outputs to input perturbation is used as a measure of the significance of inputs. This paper presents a novel functional analysis of the weight matrix based on a technique developed for determining the behavioral significance of hidden neurons. This is compared with the application of the same technique to the training and test data. Finally, a novel aggregation technique is introduced.

Algorithms↗