Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Finding biological process modifications in cancer tissues by mining gene expression correlations.

BACKGROUND: Through the use of DNA microarrays it is now possible to obtain quantitative measurements of the expression of thousands of genes from a biological sample. This technology yields a global view of gene expression that can be used in several ways. Functional insight into expression profiles is routinely obtained by using Gene Ontology terms associated to the cellular genes. In this paper, we deal with functional data mining from expression profiles, proposing a novel approach that studies the correlations between genes and their relations to Gene Ontology (GO). By using this "functional correlations comparison" we explore all possible pairs of genes identifying the affected biological processes by analyzing in a pair-wise manner gene expression patterns and linking correlated pairs with Gene Ontology terms. RESULTS: We apply here this "functional correlations comparison" approach to identify the existing correlations in hepatocarcinoma (161 microarray experiments) and to reveal functional differences between normal liver and cancer tissues. The number of well-correlated pairs in each GO term highlights several differences in genetic interactions between cancer and normal tissues. We performed a bootstrap analysis in order to compute false detection rates (FDR) and confidence limits. CONCLUSION: Experimental results show the main advantage of the applied method: it both picks up general and specific GO terms (in particular it shows a fine resolution in the specific GO terms). The results obtained by this novel method are highly coherent with the ones proposed by other cancer biology studies. But additionally they highlight the most specific and interesting GO terms helping the biologist to focus his/her studies on the most relevant biological processes.

Carcinoma, Hepatocellular↗

Knowledge discovery from structured mammography reports using inductive logic programming.

The development of large mammography databases provides an opportunity for knowledge discovery and data mining techniques to recognize patterns not previously appreciated. Using a database from a breast imaging practice containing patient risk factors, imaging findings, and biopsy results, we tested whether inductive logic programming (ILP) could discover interesting hypotheses that could subsequently be tested and validated. The ILP algorithm discovered two hypotheses from the data that were 1) judged as interesting by a subspecialty trained mammographer and 2) validated by analysis of the data itself.

Algorithms↗

Flexible information storage in MUDR(II) EHR.

An important research task of the EuroMISE Centre is the applied research in the field of electronic health record (EHR) design including electronic medical guidelines and intelligent systems for data mining and decision support. The research in this field was inspired by several European projects. We have proposed a mathematical meta-description of a flexible information storage model based on the experience gathered in cooperation in those projects. In this model, we use two basic structures called a knowledge base and data files. We describe those two structures using the graph theory concepts. Furthermore, we use logical formulas to express conditions that should be valid. Additionally, we present a description of a global system architecture of a 3-tier EHR application with interfaces based on the latest technologies; predominately on Web Services, SOAP, XML, HTTP, CORBA, etc. According to our experience and test results gained from the MUDR EHR usage, we describe an open universal solution, which can be applied as the EHR kernel of hospital information systems. To realize this approach in a daily practice for health professionals we have started a co-operative project with clinical information systems developers. Within that project we are developing a new system for continual shared health care.

Biomedical Research↗

Analyzing and mining image databases.

Image mining is the application of computer-based techniques that extract and exploit information from large image sets to support human users in generating knowledge from these sources. This review focuses on biomedical applications, in particular automated imaging at the cellular level. An image database is an interactive software application that combines data management, image analysis and visual data mining. The main characteristic of such a system is a layer that represents objects within an image, and that represents a large spectrum of quantitative and semantic object features. The image analysis needs to be adapted to each particular experiment, so 'end-user programming' will be desirable to make the technology more widely applicable.

Humans↗

[Describing language of spectra and rough set].

It is the traditional way to analyze spectra by experiences in astronomical field. And until now there has never been a suitable theoretical frame to describe spectra, which is may be owing to small spectra datasets that astronomers can get by low-level instruments. With the high-speed development of telescopes, especially on behalf of LAMOST, a large telescope which can collect more than 20,000 spectra in an observing night, spectra datasets are becoming larger and larger very fast. Facing these voluminous datasets, the traditional spectra-processing way simply depending on experiences becomes unfit. In this paper, we develop a brand-new language--describing language of spectra (DLS) to describe spectra of celestial bodies by defining BE (Basic element). And based on DLS, we introduce the method of RSDA (Rough set and data analysis), which is a technique of data mining. By RSDA we extract some rules of stellar spectra, and this experiment can be regarded as an application of DLS.

Algorithms↗

Analysis of traffic injury severity: an application of non-parametric classification tree techniques.

Statistical regression models, such as logit or ordered probit/logit models, have been widely employed to analyze injury severity of traffic accidents. However, most regression models have their own model assumptions and pre-defined underlying relationships between dependent and independent variables. If these assumptions are violated, the model could lead to erroneous estimations of injury likelihood. The classification and regression tree (CART), one of the most widely applied data mining techniques, has been commonly employed in business administration, industry, and engineering. CART does not require any pre-defined underlying relationship between target (dependent) variable and predictors (independent variables) and has been shown to be a powerful tool, particularly for dealing with prediction and classification problems. This study uses the 2001 accident data for Taipei, Taiwan. A CART model was developed to establish the relationship between injury severity and driver/vehicle characteristics, highway/environmental variables and accident variables. The results indicate that the most important variable associated with crash severity is the vehicle type. Pedestrians, motorcycle and bicycle riders are identified to have higher risks of being injured than other types of vehicle drivers in traffic accidents.

Accidents, Traffic↗

Validation in pharmaceutical analysis. Part II: Central importance of precision to establish acceptance criteria and for verifying and improving the quality of analytical data.

Validation of analytical procedures is a vital aspect not just for regulatory purposes, but also for their efficient and reliable long-term application. In order to address the performance of the analytical procedure adequately, the analyst is responsible to identify the relevant parameters, to design the experimental validation studies accordingly and to define appropriate acceptance criteria. Establishing an acceptable analytical variability for the given application is of central importance as many other acceptance criteria can be derived from such a precision. Acceptable precision ranges for types of control tests and/or analytes can be obtained from validation, but also related activities such as transfer, control charts, or extracted from routine applications such as batch release or stability studies (data mining). Apart from compiling a database for general benchmarking, during such an information-building process, the reliability of the analytical variability of the specific procedure is more and more increased. This is important as a reliable target variability facilitates to detect or investigate atypical or out-of specification behaviour of analytical data in a routine application, thus improving the data quality and reliability. According to the life-cycle concept of validation, measures should be taken to maintain and control the validated status of analytical procedures during long-term routine application, such as monitoring relevant performance parameters (system suitability tests), control charts, etc. If the analytical system is demonstrated to be stable, i.e. under statistical control, a major variability contribution in LC originating from the standard preparation and analysis can be reduced. A concept of quantification by pre-determined calibration parameters instead of the classical approach of simultaneous calibration is described.

Chemistry, Pharmaceutical↗

Wrapping SRS with CORBA: from textual data to distributed objects.

MOTIVATION: Biological data come in very different shapes. Databanks are maintained and used by distinct organizations. Text is the de facto Standard exchange format. The SRS system can integrate heterogeneous textual databanks but it was lacking a way to structure the extracted data. RESULTS: This paper presents a CORBA interface to the SRS system which manages databanks in a flat file format. SRS Object Servers are CORBA wrappers for SRS. They allow client applications (visualisation tools, data mining tools, etc.) to access and query SRS servers remotely through an Object Request Broker (ORB). They provide loader objects that contain the information extracted from the databanks by SRS. Loader objects are not hard-coded but generated in a flexible way by using loader specifications which allow SRS administrators to package data coming from distinct databanks. AVAILABILITY: The prototype may be available for beta-testing. Please contact the SRS group (http://srs.ebi.ac.uk).

Computer Systems↗

Health care in the information society. A prognosis for the year 2013.

Our society is increasingly influenced by modern information and communication technology (ICT). Health care has profited greatly by this development. How could health care provision look in the near future, in 10 years, or more precisely, in the year 2013? What measures must be undertaken by political and self-governing health institutions, and by medical informatics research, to ensure an efficient, medically advanced and yet affordable future health care system? Three factors will greatly influence the further development of information processing in health care within the near future: the development of the population, medical advances, and advances in informatics. These factors have motivated us to set up 30 theses for health care provision in the year 2013. The theses cover areas of health care, such as its people, its information systems, and its ICT tools. Three major goals requiring achievement have been identified: patient-centered recording and use of medical data for cooperative care, process-integrated decision support through current medical knowledge, comprehensive use of patient data for research and health care reporting. In consequence, political institutions should provide a framework for networked, patient-centered health care. They are called on to regulate the storage and exchange of health care data and of appropriate information system architectures. Finally, the health care institutions themselves must emphasize professional information management more strongly. Relevant research topics in medical informatics are: comprehensive electronic patient records, modern health information system architectures, architectures for medical knowledge centers, specific data processing methods ('medical data mining'), and multi-functional, mobile ICT tools.

Adolescent↗

Exploration of multilocus effects in a highly polymorphic gene, the apolipoprotein (APOB) gene, in relation to plasma apoB levels.

A detailed exploration of all the polymorphisms in candidate genes is required to better characterize the relationship between gene variability and complex traits. We propose a novel strategy for investigating the association between a highly polymorphic gene and a phenotype, by combining a multilocus genotype analysis and an haplotype analysis. For the multilocus genotype analysis, a data mining tool--termed DICE (Detection of Informative Combined Effects)--was developed to identify the best subset of polymorphisms that are associated--individually or in combination--with the phenotype. For the haplotype analysis, we used our recently developed method of haplotype-phenotype association to determine the most informative and parsimonious haplotype model fitting the data. We illustrate this strategy by investigating the association between twelve polymorphisms of the APOB gene and plasma apoB levels in 1442 European subjects. After exploring all main effects and interactions between polymorphisms, DICE identified the N4311S polymorphism as the most informative polymorphism in relation to apoB levels. Haplotype analysis led to the same conclusion. Additionally, DICE identified the E4154K (EcoRI) and the T2488T (XbaI) polymorphisms as potentially interesting. This selection was not modified by inclusion of the common APOE polymorphism in the analysis.

Adult↗

Mapping knowledge domains: characterizing PNAS.

A review of data mining and analysis techniques that can be used for the mapping of knowledge domains is given. Literature mapping techniques can be based on authors, documents, journals, words, and/or indicators. Most mapping questions are related to research assessment or to the structure and dynamics of disciplines or networks. Several mapping techniques are demonstrated on a data set comprising 20 years of papers published in PNAS. Data from a variety of sources are merged to provide unique indicators of the domain bounded by PNAS. By using funding source information and citation counts, it is shown that, on an aggregate basis, papers funded jointly by the U.S. Public Health Service (which includes the National Institutes of Health) and non-U.S. government sources outperform papers funded by other sources, including by the U.S. Public Health Service alone. Grant data from the National Institute on Aging show that, on average, papers from large grants are cited more than those from small grants, with performance increasing with grant amount. A map of the highest performing papers over the 20-year period was generated by using citation analysis. Changes and trends in the subjects of highest impact within the PNAS domain are described. Interactions between topics over the most recent 5-year period are also detailed.

Documentation↗

In silico comparison of gene expression levels in ten human tumor types reveals candidate genes associated with carcinogenesis.

Most human cancers are characterized by genomic instability. Changes associated with such may result in altered expression of numerous genes. The sequence information available in the public databases can be used to identify transcripts differentially expressed in cancers. Determining cancer-related genes that are commonly deregulated in different tumor types may facilitate identification of targets for cancer diagnoses and therapeutic treatments. Using a data-mining tool named Digital Differential Display (DDD) from the UniGene database at the NCBI web site, gene expression levels of ten different tumor types and their counterpart normal tissues were analyzed. Unigenes which showed transcriptional regulation in more than five tumor types with > or =2-fold differences from normal tissues were identified. The expression data of selected Unigenes were subjected to clustering analysis. 127 commonly up-regulated genes and 92 commonly down-regulated genes were identified. Clustering analysis using these genes showed that most tumor types can be clustered into a separate branch from most normal tissues. Nineteen genes that have been shown to be involved in carcinogenesis by experimental evidence were also identified. Present computational analyses revealed 219 candidate cancer-related genes that are commonly deregulated in ten human tumor types which may contribute to the progress of carcinogenesis.

Gene Expression↗

A framework for scientific data modeling and automated software development.

MOTIVATION: The lack of standards for storage and exchange of data is a serious hindrance for the large-scale data deposition, data mining and program interoperability that is becoming increasingly important in bioinformatics. The problem lies not only in defining and maintaining the standards, but also in convincing scientists and application programmers with a wide variety of backgrounds and interests to adhere to them. RESULTS: We present a UML-based programming framework for the modeling of data and the automated production of software to manipulate that data. Our approach allows one to make an abstract description of the structure of the data used in a particular scientific field and then use it to generate fully functional computer code for data access and input/output routines for data storage, together with accompanying documentation. This code can be generated simultaneously for different programming languages from a single model, together with, for example for format descriptions and I/O libraries XML and various relational databases. The framework is entirely general and could be applied in any subject area. We have used this approach to generate a data exchange standard for structural biology and analysis software for macromolecular NMR spectroscopy. AVAILABILITY: The framework is available under the GPL license, the data exchange standard with generated subroutine libraries under the LGPL license. Both may be found at http://www.ccpn.ac.uk; http://sourceforge.net/projects/ccpn CONTACT: ccpn@mole.bio.cam.ac.uk.

Biopolymers↗

BIND--The Biomolecular Interaction Network Database.

The Biomolecular Interaction Network Database (BIND; http://binddb. org) is a database designed to store full descriptions of interactions, molecular complexes and pathways. Development of the BIND 2.0 data model has led to the incorporation of virtually all components of molecular mechanisms including interactions between any two molecules composed of proteins, nucleic acids and small molecules. Chemical reactions, photochemical activation and conformational changes can also be described. Everything from small molecule biochemistry to signal transduction is abstracted in such a way that graph theory methods may be applied for data mining. The database can be used to study networks of interactions, to map pathways across taxonomic branches and to generate information for kinetic simulations. BIND anticipates the coming large influx of interaction information from high-throughput proteomics efforts including detailed information about post-translational modifications from mass spectrometry. Version 2.0 of the BIND data model is discussed as well as implementation, content and the open nature of the BIND project. The BIND data specification is available as ASN.1 and XML DTD.

Binding, Competitive↗

The TIGR rice genome annotation resource: annotating the rice genome and creating resources for plant biologists.

Rice is not only a major food staple for the world's population but it also is a model species for a major group of flowering plants, the monocotyledonous plants. Draft genomic sequence of two subspecies of rice, Oryza sativa spp. japonica and indica ssp. are publicly available. To provide the community with a resource to data-mine the rice genome, we have constructed an annotation resource for rice (http://www.tigr.org/tdb/e2k1/osa1/). In this resource, we have annotated the rice genome for gene content, identified motifs/domains within the predicted genes, constructed a rice repeat database, identified related sequences in other plant species, and identified syntenic sequences between rice and maize. All of the data is available through web-based interfaces, FTP downloads, and a Distributed Annotation System.

Chromosomes, Artificial↗

Nearest neighbors by neighborhood counting.

Finding nearest neighbors is a general idea that underlies many artificial intelligence tasks, including machine learning, data mining, natural language understanding, and information retrieval. This idea is explicitly used in the k-nearest neighbors algorithm (kNN), a popular classification method. In this paper, this idea is adopted in the development of a general methodology, neighborhood counting, for devising similarity functions. We turn our focus from neighbors to neighborhoods, a region in the data space covering the data point in question. To measure the similarity between two data points, we consider all neighborhoods that cover both data points. We propose to use the number of such neighborhoods as a measure of similarity. Neighborhood can be defined for different types of data in different ways. Here, we consider one definition of neighborhood for multivariate data and derive a formula for such similarity, called neighborhood counting measure or NCM. NCM was tested experimentally in the framework of kNN. Experiments show that NCM is generally comparable to VDM and its variants, the state-of-the-art distance functions for multivariate data, and, at the same time, is consistently better for relatively large k values. Additionally, NCM consistently outperforms HEOM (a mixture of Euclidean and Hamming distances), the "standard" and most widely used distance function for multivariate data. NCM has a computational complexity in the same order as the standard Euclidean distance function and NCM is task independent and works for numerical and categorical data in a conceptually uniform way. The neighborhood counting methodology is proven sound for multivariate data experimentally. We hope it will work for other types of data.

Algorithms↗

The sumatriptan/naratriptan aggregated patient (SNAP) database: aggregation, validation and application.

Pooled data from multiple clinical trials can provide information for medical decision-making that typically cannot be derived from a single clinical trial. By increasing the sample size beyond that achievable in a single clinical trial, pooling individual-patient data from multiple trials provides additional statistical power to detect possible effects of study medication, confers the ability to detect rare outcomes, and facilitates evaluation of effects among subsets of patients. Data from pharmaceutical company-sponsored clinical trials lend themselves to data-pooling, meta-analysis, and data mining initiatives. Pharmaceutical company-sponsored clinical trials are arguably among the most rigorously designed and conducted of studies involving human subjects as a result of multidisciplinary collaboration involving clinical, academic and/or governmental investigators as well as the input and review of medical institutional bodies and regulatory authorities. This paper describes the aggregation, validation and initial analysis of data from the sumatriptan/naratriptan aggregate patient (SNAP) database, which to date comprises pooled individual-patient data from 128 clinical trials conducted from 1987 to 1998 with the migraine medications sumatriptan and naratriptan. With an extremely large sample size (>28000 migraineurs, >140000 treated migraine attacks), the SNAP database allows exploration of questions about migraine and the efficacy and safety of migraine medications that cannot be answered in single clinical trials enrolling smaller numbers of patients. Besides providing the adequate sample size to address specific questions, the SNAP database allows for subgroup analyses that are not possible in individual trial analyses due to small sample size. The SNAP database exemplifies how the wealth of data from pharmaceutical company-sponsored clinical trials can be re-used to continue to provide benefit.

Clinical Trials as Topic↗

THE CAENORHABDITIS ELEGANS GENOME: A Guide in The Post Genomics Age.

The completion of the entire genome sequence of the free-living nematode, Caenorhabditis elegans is a tremendous milestone in modern biology. Not only will scientists be poring over data mined from this resource, but techniques and methodologies developed along the way have changed the way we can approach biological questions. The completion of the C. elegans genomic sequence will be of particular importance to scientists working on parasitic nematodes. In many cases, these nematode species present intractable challenges to those interested in their biology and genetics. The data already compared from parasites to the C. elegans database reveals a wealth of opportunities for parasite biologists. It is likely that many of the same genes will be present in parasites and that these genes will have similar functions. Additional information regarding differences between free-living and parasitic species will provide insight into the evolution and nature of parasitism. Finally, genetic and genomic approaches to the study of parasitic nematodes now have a clearly marked path to follow.

Journal Article↗