Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

[Hierarchies according to the level of evidence of source data, before their integration in the synthesis, in the matter of therapeutic efficacy].

The discipline therapeutic information uses the concept of the level of evidence for source data concerning therapeutic efficacy, and its ordering before integration in syntheses. In this paper we will start by considering the problems raised by the definition of the level of evidence in terms of the dimensions it covers. We have differentiated three components: clinical pertinance of the question asked, the methods used to reply and the quality of the data collected. Second, we will examine the different criteria important for each of these three dimensions. There are many criteria possible which do not all have the same weight, and thus for any non-arbitrary tool developed to enable the level of evidence to be ordered, it is necessary to know the weight of the different criteria. Thirdly, we will present the techniques used for working with multicriteria situations in econometrics which represent a methodology we propose using to apply in our context. To do so we need to build a 'reference' base for the level of evidence using 'experts' opinions which will help us to examine the weights of the different criteria. This approach, in conjunction with some epistemological and sociological considerations, may contribute to a better understanding of the different dimensions of this concept.

Clinical Trials as Topic↗

Supervised enzyme network inference from the integration of genomic data and chemical information.

MOTIVATION: The metabolic network is an important biological network which relates enzyme proteins and chemical compounds. A large number of metabolic pathways remain unknown nowadays, and many enzymes are missing even in known metabolic pathways. There is, therefore, an incentive to develop methods to reconstruct the unknown parts of the metabolic network and to identify genes coding for missing enzymes. RESULTS: This paper presents new methods to infer enzyme networks from the integration of multiple genomic data and chemical information, in the framework of supervised graph inference. The originality of the methods is the introduction of chemical compatibility as a constraint for refining the network predicted by the network inference engine. The chemical compatibility between two enzymes is obtained automatically from the information encoded by their Enzyme Commission (EC) numbers. The proposed methods are tested and compared on their ability to infer the enzyme network of the yeast Saccharomyces cerevisiae from four datasets for enzymes with assigned EC numbers: gene expression data, protein localization data, phylogenetic profiles and chemical compatibility information. It is shown that the prediction accuracy of the network reconstruction consistently improves owing to the introduction of chemical constraints, the use of a supervised approach and the weighted integration of multiple datasets. Finally, we conduct a comprehensive prediction of a global enzyme network consisting of all enzyme candidate proteins of the yeast to obtain new biological findings. AVAILABILITY: Softwares are available upon request.

Algorithms↗

Improving report turnaround time: an integrated method using data from a radiology information system.

OBJECTIVE: In the face of a changing health care system and increased competition, radiology departments need to become more efficient. One measurement of efficiency is promptness in producing a final report. Many large radiology centers have radiology information systems (RIS) that track work flow, collecting tremendous amounts of data. Most, however, lack an appropriate analytic mechanism. We have developed an integrated system that allows continual monitoring of radiology work flow and thus of opportunities to apply interventions. This system can form an important component of the quality management process in the radiology department. MATERIALS AND METHODS: In developing the system, we identified seven key steps in the work-flow process. When left to chance, these steps occur out of sequence and large delays occur. A scheme was devised to improve the sequencing of the work flow by using the data collected from the RIS, sorted by radiology division and patient type. Biweekly, the appropriate data file is transferred to each division for analysis, via the department's computer network. A one-step process follows, using desktop Macintosh computers and a custom program written in Microsoft Excel. Extracted data are quickly converted into a tailored division summary, and a report is automatically generated. RESULTS: The result summary format is uniform throughout the department, allowing ease of review at divisional and departmental meetings. Problems can be immediately localized to a specific step in the work-flow process. Automation of much of the system allows continual, near-real-time review of work flow. Using this approach, we have seen a sustained reduction of average report turnaround time. CONCLUSION: This system allows continual monitoring of work flow. It is largely automated and lends itself well to inclusion in the quality management program of any radiology department.

Medical Records Systems, Computerized↗

MetaGeneAlyse: analysis of integrated transcriptional and metabolite data.

UNLABELLED: New techniques in sample preparation allow high throughput analysis of samples on the transcriptional as well as on the metabolic level. We present a service accessible via the web that allows the analysis of integrated data sets that combine gene-expression data and metabolic data. After uploading, data sets can be normalized, clustered by various methods and results can be graphically visualized. All calculations are carried out on a server, so even time- and memory-consuming analyses can be done independently of the performance of the client. AVAILABILITY: The service is accessible via web-interface at http://metagenealyse.mpimp-golm.mpg.de/

Algorithms↗

Integration of genomic data for pharmacology and toxicology using Internet resources.

Genome based technologies such as sequencing and gene expression profiling using microarrays are creating massive amounts of data. Results from these studies have provided unique insights into targets, biochemical pathways, and biological systems affected by drug or xenobiotic chemical treatments. Moreover, these genomic technologies offer the potential to identify biomarkers for pharmacological development or toxicological prediction. Nonetheless, microarray studies involving a single compound produce useful although limited data. To gain further power from these individual studies, the ability to combine datasets through integration schemes has become imperative. In the current study, we describe and analyze currently available Internet resources designed to address this problem. Many functionalities, such as ability to cross reference orthologous genes across species or to combine same technology platform data, are present in these resources. Nonetheless, these resources are limited in the number of technology platforms they can support. While the ability to integrate all currently existing gene expression datasets remains enigmatic, the current tools provide a partial solution that may still yield unique insights into the affects of exogenous molecules at the level of gene expression.

Databases, Genetic↗

Federating data with Information Integrator.

Information Integrator is an extension to IBM's relational database DB2, which uses data federation to provide benefits to molecular biology researchers through two unique capabilities: increased flexibility in combining data from disparate sources, and SQL access to non-SQL data, easing the task of automating data analysis.

Databases as Topic↗

Development of a biodynamics work environment (BWE) which integrates a biodynamics data bank, models, and analytical tools.

For more than thirty years the Armstrong Laboratory (AL) has studied the response of human volunteers and human surrogates to impact accelerations. Data from these studies form a national resource which should be made available to biodynamics model builders and safety researchers everywhere. Recently AL joined with the National Highway Transportation Safety Administration (NHTSA) to develop a multi-media Biodynamics Work Environment (BWE) which will facilitate access to both their respective data bases and provide a means for sharing that access within the scientific community. The BWE concept is sufficiently flexible to allow integration of other databases as well as biodynamics models and tools. The AL data, for example, consists of over 10,000 impact tests which are currently archived on optical media. Each test includes acceleration and motion time histories, subject anthropometry, documentation photos, and high speed video. Among the models to be included are the Articulated Total Body, Head Spine, Dynamic Response Index, Multi-axis Dynamic Response Criteria, and BURNSIM. Complementary analytical tools would involve the use of spreadsheets, statistics, modeling, data analysis, and math packages. This paper will discuss the development, current status, and future plans of the BWE.

Acceleration↗

Data management and preliminary data analysis in the pilot phase of the HUPO Plasma Proteome Project.

The pilot phase of the HUPO Plasma Proteome Project (PPP) is an international collaboration to catalog the protein composition of human blood plasma and serum by analyzing standardized aliquots of reference serum and plasma specimens using a variety of experimental techniques. Data management for this project included collection, integration, analysis, and dissemination of findings from participating organizations world-wide. Accomplishing this task required a communication and coordination infrastructure specific enough to support meaningful integration of results from all participants, but flexible enough to react to changing requirements and new insights gained during the course of the project and to allow participants with varying informatics capabilities to contribute. Challenges included integrating heterogeneous data, reducing redundant information to minimal identification sets, and data annotation. Our data integration workflow assembles a minimal and representative set of protein identifications, which account for the contributed data. It accommodates incomplete concordance of results from different laboratories, ambiguity and redundancy in contributed identifications, and redundancy in the protein sequence databases. Recommendations of the PPP for future large-scale proteomics endeavors are described.

Algorithms↗

Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks.

MOTIVATION: Clinical data, such as patient history, laboratory analysis, ultrasound parameters--which are the basis of day-to-day clinical decision support--are often underused to guide the clinical management of cancer in the presence of microarray data. We propose a strategy based on Bayesian networks to treat clinical and microarray data on an equal footing. The main advantage of this probabilistic model is that it allows to integrate these data sources in several ways and that it allows to investigate and understand the model structure and parameters. Furthermore using the concept of a Markov Blanket we can identify all the variables that shield off the class variable from the influence of the remaining network. Therefore Bayesian networks automatically perform feature selection by identifying the (in)dependency relationships with the class variable. RESULTS: We evaluated three methods for integrating clinical and microarray data: decision integration, partial integration and full integration and used them to classify publicly available data on breast cancer patients into a poor and a good prognosis group. The partial integration method is most promising and has an independent test set area under the ROC curve of 0.845. After choosing an operating point the classification performance is better than frequently used indices.

Bayes Theorem↗

How probability survey data can help integrate 305(b) and 303(d) monitoring and assessment of state waters.

Section 305(b) of the United States' Clean Water Act (CWA) requires states to assess the overall quality of waters in the states, while Section 303(d) requires states to develop a list of the specific waters in their state not attaining water quality standards (a.k.a impaired waters). An integrated, efficient and cost-effective process is needed to acquire and assess the data needed to meet both these mandates. A subset of presentations at the 2002 Environmental Monitoring and Assessment Program (EMAP) Symposium provided information on how probability data, tools and methods could be used by states and other entities to aid in development of their overall assessment of condition and list of impaired waters. Discussion identified some of the technical and institutional problems that hinder the use of EMAP methods and data in the analysis to identify impaired waters as well as development needs to overcome these problems.

Data Collection↗

Neuronal database integration: the Senselab EAV data model.

We discuss an approach towards integrating heterogeneous nervous system data using an augmented Entity-Attribute-Value (EAV) schema design. This approach, widely used in implementing electronic patient record systems (EPRSs), allows the physical schema of the database to be relatively immune to changes in domain knowledge. This is because new kinds of facts are added as data (or as metadata) rather than hard-coded as the names of newly created tables or columns. Because the domain knowledge is stored as metadata, a framework developed in one scientific domain can be ported to another with only modest revision. We describe our progress in creating a code framework that handles browsing and hyperlinking of the different kinds of data.

Databases, Factual↗