Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Genomic and proteomic databases: large-scale analysis and integration of data.

With the completion of the human genome draft sequencing and assembly, an upcoming targeted goal will be identifying the proteins and changes in protein expression, which are derived from the genome template. There is important information already available on many proteins that can be a valuable resource to scientists in cardiovascular medicine as well as other fields. This article will summarize many of the different databases, their content and growth, and the differences between growth of nucleic acid and protein databases. Linkage between and integration of different protein and nuclei acid databases, along with related annotation, will greatly improve the information content and knowledge base with regards to protein data.

Databases, Factual↗

Inter-platform comparability of microarrays in acute lymphoblastic leukemia.

BACKGROUND: Acute lymphoblastic leukemia (ALL) is the most common pediatric malignancy and has been the poster-child for improved therapeutics in cancer, with life time disease-free survival (LTDFS) rates improving from <10% in 1970 to >80% today. There are numerous known genetic prognostic variables in ALL, which include T cell ALL, the hyperdiploid karyotype and the translocations: t(12;21)[TEL-AML1], t(4;11)[MLL-AF4], t(9;22)[BCR-ABL], and t(1;19)[E2A-PBX]. ALL has been studied at the molecular level through expression profiling resulting in un-validated expression correlates of these prognostic indices. To date, the great wealth of expression data, which has been generated in disparate institutions, representing an extremely large cohort of samples has not been combined to validate any of these analyses. The majority of this data has been generated on the Affymetrix platform, potentially making data integration and validation on independent sample sets a possibility. Unfortunately, because the array platform has been evolving over the past several years the arrays themselves have different probe sets, making direct comparisons difficult. To test the comparability between different array platforms, we have accumulated all Affymetrix ALL array data that is available in the public domain, as well as two sets of cDNA array data. In addition, we have supplemented this data pool by profiling additional diagnostic pediatric ALL samples in our lab. Lists of genes that are differentially expressed in the six major subclasses of ALL have previously been reported in the literature as possible predictors of the subclass. RESULTS: We validated the predictability of these gene lists on all of the independent datasets accumulated from various labs and generated on various array platforms, by blindly distinguishing the prognostic genetic variables of ALL. Cross-generation array validation was used successfully with high sensitivity and high specificity of gene predictors for prognostic variables. We have also been able to validate the gene predictors with high accuracy using an independent dataset generated on cDNA arrays. CONCLUSION: Interarray comparisons such as this one will further enhance the ability to integrate data from several generations of microarray experiments and will help to break down barriers to the assimilation of existing datasets into a comprehensive data pool.

Bone Marrow↗

Model-based methodology for analyzing incomplete quality-of-life data and integrating them into the Q-TWiST framework.

BACKGROUND: The standard Q-TWiST approach defines a series of health states and weights each state's duration according to its quality of life (QOL) to calculate quality-adjusted lifetimes. However, a fixed weight may not adequately reflect time variations in QOL. METHODS: To account for measurements derived from irregular visits and informative missing data, the authors estimated the mean QOL profile using a mixed-effect growth curve model for the response, combined with a logistic regression model for the drop-out process. RESULTS: Using data from a clinical study of lymphoma patients, the authors demonstrated better readaptation to normal life for patients younger than 30. Sensitivity analyses and computer simulations demonstrated that modeling the drop-out probability as a function of the QOL measurements is necessary if conditioning by health state is not possible. CONCLUSION: Our model-based approach is useful to analyze studies with incomplete QOL data, especially when approximate QOL assessment by health state is not possible.

Adult↗

The importance of Java and CORBA in medicine.

One of the most powerful tools available for telemedicine is a multimedia medical record accessible over a wide area and simultaneously editable by multiple physicians. The ability to do this through an intuitive interface linking multiple distributed data repositories while maintaining full data integrity is a fundamental enabling technology in healthcare. We discuss the role of distributed object technology using Java and CORBA in providing this capability including an example of such a system (TeleMed) which can be accessed through the World Wide Web. Issues of security, scalability, data integrity, and usability are emphasized.

Computer Communication Networks↗

Integrating biological data through the genome.

Owing to the ongoing success of the genome sequencing and structural genomics projects, the increase in both sequence and structural data is rapid. The development of tools for the annotation of sequence and structural data has become more important in the hope of keeping up with this data explosion. Scientists in this field have addressed these issues over the last 10 years and there now exists a wealth of methods and approaches to help interpret these data. However, there is no current way in which these methods can be incorporated easily so that the resulting annotations can be viewed together. This review discusses the development of these annotation methods and introduces the BioSapiens Network of Excellence, which has been formed in order to integrate the methods which have been developed in Europe.

Animals↗

Integrative analysis of gene expression patterns predicts specific modulations of defined cell functions by estrogen and tamoxifen in MCF7 breast cancer cells.

To explore the mechanisms whereby estrogen and antiestrogen (tamoxifen (TAM)) can regulate breast cancer cell growth, we investigated gene expression changes in MCF7 cells treated with 17beta-estradiol (E2) and/or with 4-OH-TAM. The patterns of differential expression were determined by the ValiGen Gene IDentification (VGID) process, a subtractive hybridization approach combined with microarray validation screening. Their possible biologic consequences were evaluated by integrative data analysis. Over 1000 cDNA inserts were isolated and subsequently cloned, sequenced and analyzed against nucleotide and protein databases (NT/NR/EST) with BLAST software. We revealed that E2 induced differential expression of 279 known and 28 unknown sequences, whereas TAM affected the expression of 286 known and 14 unknown sequences. Integrative data analysis singled out a set of 32 differentially expressed genes apparently involved in broad cellular mechanisms. The presence of E2 modulated the expression patterns of 23 genes involved in anchors and junction remodeling; extracellular matrix (ECM) degradation; cell cycle progression, including G1/S check point and S-phase regulation; and synthesis of genotoxic metabolites. In tumor cells, these four mechanisms are associated with the acquisition of a motile and invasive phenotype. TAM partly reversed the E2-induced differential expression patterns and consequently restored most of the biologic functions deregulated by E2, except the mechanisms associated with cell cycle progression. Furthermore, we found that TAM affects the expression of nine additional genes associated with cytoskeletal remodeling, DNA repair, active estrogen receptor formation and growth factor synthesis, and mitogenic pathways. These modulatory effects of E2 and TAM upon the gene expression patterns identified here could explain some of the mechanisms associated with the acquisition of a more aggressive phenotype by breast cancer cells, such as E2-independent growth and TAM resistance.

Breast Neoplasms↗

An integrated genetic data environment (GDE)-based LINUX interface for analysis of HIV-1 and other microbial sequences.

MOTIVATION: Sequence databases encode a wealth of information needed to develop improved vaccination and treatment strategies for the control of HIV and other important pathogens. To facilitate effective utilization of these datasets, we developed a user-friendly GDE-based LINUX interface that reduces input/output file formatting. DESIGN AND RESULTS: GDE was adapted to the Linux operating system, bioinformatics tools were integrated with microbe-specific databases, and up-to-date GDE menus were developed for several clinically important viral, bacterial and parasitic genomes. Each microbial interface was designed for local access and contains Genbank, BLAST-formatted and phylogenetic databases. AVAILABILITY: GDE-Linux is available for research purposes by direct application to the corresponding author. Application-specific menus and support files can be downloaded from (http://www.bioafrica.net).

Database Management Systems↗

Integrating mutation data and structural analysis of the TP53 tumor-suppressor protein.

TP53 encodes p53, which is a nuclear phosphoprotein with cancer-inhibiting properties. In response to DNA damage, p53 is activated and mediates a set of antiproliferative responses including cell-cycle arrest and apoptosis. Mutations in the TP53 gene are associated with more than 50% of human cancers, and 90% of these affect p53-DNA interactions, resulting in a partial or complete loss of transactivation functions. These mutations affect the structural integrity and/or p53-DNA interactions, leading to the partial or complete loss of the protein's function. We report here the results of a systematic automated analysis of the effects of p53 mutations on the structure of the core domain of the protein. We found that 304 of the 882 (34.4%) distinct mutations reported in the core domain can be explained in structural terms by their predicted effects on protein folding or on protein-DNA contacts. The proportion of "explained" mutations increased to 55.6% when substitutions of evolutionary conserved amino acids were included. The automated method of structural analysis developed here may be applied to other frequently mutated gene mutations such as dystrophin, BRCA1, and G6PD.

Amino Acid Substitution↗

BIOZON: a hub of heterogeneous biological data.

Biological entities are strongly related and mutually dependent on each other. Therefore, there is a growing need to corroborate and integrate data from different resources and aspects of biological systems in order to analyze them effectively. Biozon is a unified biological database that integrates heterogeneous data types such as proteins, structures, domain families, protein-protein interactions and cellular pathways, and establishes the relationships between them. All data are integrated on to a single graph schema centered around the non-redundant set of biological objects that are shared by each source. This integration results in a highly connected graph structure that provides a more complete picture of the known context of a given object that cannot be determined from any one source. Currently, Biozon integrates roughly 2 million protein sequences, 42 million DNA or RNA sequences, 32,000 protein structures, 150,000 interactions and more from sources such as GenBank, UniProt, Protein Data Bank (PDB) and BIND. Biozon augments source data with locally derived data such as 5 billion pairwise protein alignments and 8 million structural alignments. The user may form complex cross-type queries on the graph structure, add similarity relations to form fuzzy queries and rank the results based on analysis of the edge structure similar to Google PageRank, online at Biozon.org.

Computer Graphics↗

dbZach: A MIAME-compliant toxicogenomic supportive relational database.

Quantitative risk assessment and the elucidation of mechanisms of toxicity requires computational infrastructure and innovative analysis approaches that systematically consider available data at all levels of biological organization. dbZach (http://dbzach.fst.msu.edu) is a modular relational database with associated data insertion, retrieval, and mining tools that manages traditional toxicology and complementary toxicogenomic data to facilitate comprehensive data integration, analysis, and sharing. It consists of four Core Subsystems (i.e., Clones, Genes, Sample Annotation, and Protocols), four Experimental Subsystems (i.e., Microarray, Affymetrix, Real-Time PCR, and Toxicology), and three Computational Subsystems (i.e., Gene Regulation, Pathways, Orthology) that comply with the Minimum Information About a Microarray Experiment (MIAME) standard. It is capable of including emerging technologies and other model systems, including ecologically relevant species. dbZach represents an enterprise toxicogenomic data management system which facilitates data integration and analysis, and reduces uncertainties in the continuum from initial exposure to toxicity while facilitating more comprehensive elucidations of mechanisms of toxicity and supporting mechanistically-based quantitative risk assessment.

Animals↗

Biologically based validation of PC electrophysiology data collection systems utilizing the Good Automated Laboratory Practices.

Since there was a scientific need to conduct electrophysiology measurements to detect possible ocular (electroretinography, ERG), central neurotoxic (quantitative electroencephalography, qEEG), and cardiac (electrocardiography, ECG) effects in animals used in certain regulatory studies, the acquisition of suitable automated PC software systems were required. This article describes the process by which these systems were validated to ensure that they met the scientific requirements, while also addressing the principles of Good Automated Laboratory Practices (GALP). After a thorough search of existing commercial packages, a plan was developed specific for each PC-based collection system selected for evaluation. The common elements of each plan included consideration of both scientific and GALP elements, such as necessary biological response variables, raw data acquisition and identification, acceptance criteria, security, protection, storage media, data integrity, audit requirements and standard operating procedures. The authors' approach to validation for each electrophysiology system was to determine scientific needs for accuracy, precision, and detection limit of biological effects concurrent with GALP requirements. The selected software systems were employed in separate scientific GLP studies conducted in dogs, rats, and mini-pigs to demonstrate the ability to detect cholinesterase effects due to multiple infusions of physostigmine, based on parallel measurement of cholinesterase biomarkers. Since the systems were designed for human usage, certain adaptations were necessary. A critical assumption to be tested was the ability of the system's algorithms to adequately capture and assimilate the data in an accurate fashion. Concomitantly, the related GALP needs, such as data integrity, security, CD-ROM archive, and personnel training requirements were evaluated, implemented, and defined to accommodate the application and process needs. The biological approach to validation of these PC-based electrophysiology systems met the necessary scientific acceptance criteria as well as compliance requirements in order to be used in regulatory studies.

Algorithms↗

Integrating safety data: the expert report.

Complete evaluation of the toxicity of a new chemical entity requires critical analysis of the pattern of positive and negative findings in all types of toxicity tests, pharmacokinetics and metabolism in the species examined, and correlation of the effects with information about its pharmacodynamic and pharmacological properties. The goal is to obtain sufficient understanding of the mechanisms underlying the therapeutic and toxic effects of the compound to permit a well-supported extrapolation from the test to the target species. The Expert Report system in the European Community is based on comprehensive 25-page reviews of information about a compound arranged under 3 headings: Chemistry and Pharmacy, Pharmacology and Toxicology, and Clinical Studies. Each section requires a searching review of the experimental work and the relevant literature, integration of the findings, and then careful correlation among these 3 main areas of knowledge to indicate the circumstances of safe and effective use of the drug.

Documentation↗

BioDataServer: a SQL-based service for the online integration of life science data.

Regarding molecular biology, we see an exponential growth of data and knowledge. Among others, this fact is reflected in more than 300 molecular databases which are readily available on the Internet. The usage of these data requires integration tools capable of complex information fusion processes. This paper will present a novel concept for user specific integration of life science data. Our approach is based on a mediator architecture in conjunction with freely adjustable data schemes. The implemented prototype is called BioDataServer and can be accessed on the Internet: http://integration.genophen.de. To realize a comfortable usage of the resulted data sets of the integration process, a SQL-based query language and a XML data format were developed and implemented.

Computational Biology↗

Security requirements for electronic patients records: the Norwegian view.

Information security, including secrecy, privacy and data integrity, quality and availability, are fundamental issues when electronic patient records are introduced and used. In this paper we outline the results of a project that KITH carried out on an assignment from the Norwegian Ministry of Health. We describe the general and detailed requirements set up in the project for electronic patient records, and to the hospitals that take such into use. We believe that these requirements are the minimum necessary to give a positive answer to the question of whether electronic patient records can meet all the security-related requirements and intentions given by current regulations. In our work we have focused on secrecy, privacy and data integrity, but also on elaborating requirements that allow a user friendly and suitable implementation of electronic patient record systems.

Computer Communication Networks↗

MAGIC Tool: integrated microarray data analysis.

SUMMARY: Several programs are now available for analyzing the large datasets arising from cDNA microarray experiments. Most programs are expensive commercial packages or require expensive third party software. Some are freely available to academic researchers, but are limited to one operating system. MicroArray Genome Imaging and Clustering Tool (MAGIC Tool) is an open source program that works on all major platforms, and takes users 'from tiff to gif'. Several unique features of MAGIC Tool are particularly useful for research and teaching. AVAILABILITY: http://www.bio.davidson.edu/MAGIC

Algorithms↗

Database and knowledge base integration--a data mapping method for Arden Syntax knowledge modules.

One of the most important categories of decision-support systems in medicine are data driven systems where the inference engine is linked to a database. It is, therefore, important to find methods that facilitate the implementation of database queries referred to in the knowledge modules. A method is described for linking clinical databases to a knowledge base with Arden Syntax modules. The method is based on a query meta-database including templates for SQL queries which is maintained by a database administrator. During knowledge module authoring the medical expert refers only to a code in the query meta-database; no knowledge is needed about the database model or the naming of attributes and relations. The method uses standard tools, such as C+2 and ODBC, which makes it possible to implement the method at many platforms and to link to different clinical databases in a standardized way.

Databases, Factual↗

Integration of omics data: how well does it work for bacteria?

In the current omics era, innovative high-throughput technologies allow measuring temporal and conditional changes at various cellular levels. Although individual analysis of each of these omics data undoubtedly results into interesting findings, it is only by integrating them that gaining a global insight into cellular behaviour can be aimed at. A systems approach thus is predicated on data integration. However, because of the complexity of biological systems and the specificities of the data-generating technologies (noisiness, heterogeneity, etc.), integrating omics data in an attempt to reconstruct signalling networks is not trivial. Developing its methodologies constitutes a major research challenge. Besides for their intrinsic value towards health care, environment and industry, prokaryotes are ideal model systems to further develop these methods because of their lower regulatory complexity compared with eukaryotes, and the ease with which they can be manipulated. Several successful examples outlined in this review already show the potential of the systems approach for both fundamental and industrial applications, which would be time-consuming or impossible to develop solely through traditional reductionist approaches.

Bacteria↗