Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans↗

FlyBase: genes and gene models.

FlyBase (http://flybase.org) is the primary repository of genetic and molecular data of the insect family Drosophilidae. For the most extensively studied species, Drosophila melanogaster, a wide range of data are presented in integrated formats. Data types include mutant phenotypes, molecular characterization of mutant alleles and aberrations, cytological maps, wild-type expression patterns, anatomical images, transgenic constructs and insertions, sequence-level gene models and molecular classification of gene product functions. There is a growing body of data for other Drosophila species; this is expected to increase dramatically over the next year, with the completion of draft-quality genomic sequences of an additional 11 Drosophila species.

Animals↗

Remote access to anatomical information: an integration between semantic knowledge and visual data.

A novel internet-based application is presented which provides access to anatomy knowledge through symbolic modality expressed by keywords taken from controlled or non-controlled terminology. The system is based on a database where anatomical concepts have been organized into a hierarchical framework. Along with term queries that allow retrieving concepts containing or exactly matching the used keyword, the system also provides semantic access to anatomical information. Queries can be setup, which retrieve concepts relying to a particular meaning and sharing a particular relationship. Moreover, the application has the capability to refine the search of the terms by querying the UMLS knowledge server. Anatomical image data have been integrated by using Visible Human Dataset. A set of these images has been indexed according to our anatomical classification and is used inside the application. The system has been implemented through Java client-server technology and works within standard Internet browsers.

Anatomy↗

"KARIBIN," an information resource for obtaining genomic information in a cytogenetic band.

KARIBIN () is a karyotypic region-based integrated information resource that provides a comprehensive view of the integrated mapping and sequencing data for the human genome. A cytogenetic band is linked to a genetic or physical location using fluorescence in situ hybridization (FISH) mapping data. The genetic, physical mapping data and the sequencing data are integrated using STS markers positioned on multiple maps. For each cytogenetic band, the user can obtain the most up-to-date information that includes genetic and physical maps, human transcript gene map, YAC and PAC/BAC clone coverage, disease gene phenotype, and high throughput genomic sequences from the major human genome sequencing centers. This information provides a framework for future experiments and may accelerate the process of disease gene hunting. It is envisioned that other cytogenetic-based information such as chromosome aberrations can be linked to this framework.

Chromosome Banding↗

Epidural motor cortex stimulation with functional imaging guidance.

Chronic epidural motor cortex stimulation (MCS) has been shown to have promise in the treatment of patients with refractory deafferentation pain. Precise placement of the electrode over the motor cortex region corresponding to the area of pain is essential for the success of this procedure. Whereas standard anatomical landmarks have been used in the past in conjunction with image guidance, the use of functional brain imaging can be beneficial in the precise surgical planning. The authors report the use of functional imaging-guided frameless stereotactic surgery for epidural MCS. Five patients underwent MCS in which functional imaging guidance was used. Prior to surgery, patients underwent magnetic resonance (MR) imaging with skin fiducial markers placed on standard anatomical reference prints, followed by magnetoencephalography (MEG) mapping of the sensory and motor cortices. In two patients, functional MR imaging was also performed using a motor task paradigm. The functional imaging data were integrated into a frameless stereotactic database by using a three-dimensional coregistration algorithm. Subsequently, a frameless stereotactic craniotomy was performed using the integrated anatomical and functional imaging data for surgical planning. Intraoperative somatosensory evoked potentials (SSEPs) and direct stimulation were used to confirm the target and final placement of the electrode. Direct stimulation and SSEPs performed intraoperatively confirmed the accuracy of the functional imaging data. Trial periods of stimulation successfully reduced pain in three of the five patients who then underwent permanent internal placement of the system. At a mean 6-month follow up, these patients reported an average reduction in pain of 55% on a visual analog scale. The integration of functional and anatomical imaging data allows for precise and efficient surgical planning and may reduce the time necessary for intraoperative physiological verification.

Humans↗

PRIME: a graphical interface for integrating genomic/proteomic databases.

Data mining, finding and integration of information about proteins of interest, is an essential component in modern biological and biomedical research. Even when focusing on a single organism and only on a small number of proteins, there are often dozens fo data sources containing relevant information. We are developing PRIME, a protein information environment, to serve as a virtual central database which integrates distributed heterogeneous information about proteins (linked by common identifier). PRIME has powerful capabilities to visualize all kinds of protein annotation in specialized views. These views can be displayed side by side at the same time and can be synchronized in order to show simultaneously different aspects of identical proteins. These features allow a quick and comprehensive overview of properties of single proteins or protein sets.

Computational Biology↗

Integrated access to metabolic and genomic data.

The EcoCyc system consists of a knowledge base (KB) that describes the genes and intermediary metabolism of Escherichia coli, and a graphical user interface (GUI) for accessing that knowledge. This paper addresses two problems: How can we create a GUI that provides integrated access to metabolic and genomic data? We describe the design and implementation of visual presentations that closely mimic those found in the biology literature, and that offer hypertext navigation among related entities, and multiple views of the same entity. We employ a frame knowledge representation system (FRS) called HyperTHEO to manage the EcoCyc knowledge base. Among the advantages of FRSs are an expressive data model for capturing the complexities of biological information, and schema-evolution capabilities that facilitate the constant schema changes that biological databases tend to undergo. HyperTHEO also includes rule-based inference facilities that are the foundation of expert systems, a constraint language for maintaining data integrity, and a declarative query language. A graphic KB editor and browser allow the EcoCyc developers to interactively inspect and modify this evolving KB.

Artificial Intelligence↗

Managing the measurement: a model of data support in an integrated delivery system.

There are many functions in an integrated delivery system charged with focused, clinical excellence and process improvement based on data. Information system complexity, enhanced technology and software, and increased demand for data have created the need for one healthcare organization to look at an alternative system for data support of these key functions. This article describes a clinical data department model that has been implemented successfully in an integrated delivery system. It describes the structure and function of this department and its relationship in supporting the areas where clinical data is needed.

Data Collection↗

A data-hiding technique with authentication, integration, and confidentiality for electronic patient records.

A data-hiding technique called the "bipolar multiple-number base" was developed to provide capabilities of authentication, integration, and confidentiality for an electronic patient record (EPR) transmitted among hospitals through the Internet. The proposed technique is capable of hiding those EPR related data such as diagnostic reports, electrocardiogram, and digital signatures from doctors or a hospital into a mark image. The mark image could be the mark of a hospital used to identify the origin of an EPR. Those digital signatures from doctors and a hospital could be applied for the EPR authentication. Thus, different types of medical data can be integrated into the same mark image. The confidentiality is ultimately achieved by decrypting the EPR related data and digital signatures with an exact copy of the original mark image. The experimental results validate the integrity and the invisibility of the hidden EPR related data. This newly developed technique allows all of the hidden data to be separated and restored perfectly by authorized users.

Algorithms↗

Achieving maximum ROI from corporate databases: exploiting your databases with integrated querying for better decision-making.

In order to increase the rate of drug discovery, pharmaceutical and biotechnology companies spend billions of dollars a year assembling research databases. Current trends still indicate a falling rate in the discovery of New Molecular Entities (NMEs). It is widely accepted that the data need to be integrated in order for it to add value. The degree to which this must be achieved is often misunderstood. The true goal of data integration must be to provide accessible knowledge. If knowledge cannot be gained from these data, then it will invalidate the business case for gathering it. Current data integration solutions focus on the initial task of integrating the actual data and to some extent, also address the need to allow users to access integrated information. Typically the search tools that are provided are either restrictive forms or free text based. While useful, neither of these solutions is suitable for providing full coverage of large numbers of integrated structured data sources. One solution to this accessibility problem is to present the integrated data in a collated manner that allows users to browse and explore it and also perform complex ad-hoc searches on it within a scientific context and without the need for advanced Information Technology (IT) skills. Additionally, the solution should be maintainable by 'in-house' administrators rather than requiring expensive consultancy. This paper examines the background to this problem, investigates the requirements for effective exploitation of corporate data and presents a novel effective solution.

Databases, Factual↗

Design of a Multi Dimensional Database for the Archimed DataWarehouse.

The Archimed data warehouse project started in 1993 at the Geneva University Hospital. It has progressively integrated seven data marts (or domains of activity) archiving medical data such as Admission/Discharge/Transfer (ADT) data, laboratory results, radiology exams, diagnoses, and procedure codes. The objective of the Archimed data warehouse is to facilitate the access to an integrated and coherent view of patient medical in order to support analytical activities such as medical statistics, clinical studies, retrieval of similar cases and data mining processes. This paper discusses three principal design aspects relative to the conception of the database of the data warehouse: 1) the granularity of the database, which refers to the level of detail or summarization of data, 2) the database model and architecture, describing how data will be presented to end users and how new data is integrated, 3) the life cycle of the database, in order to ensure long term scalability of the environment. Both, the organization of patient medical data using a standardized elementary fact representation and the use of the multi dimensional model have proved to be powerful design tools to integrate data coming from the multiple heterogeneous database systems part of the transactional Hospital Information System (HIS). Concurrently, the building of the data warehouse in an incremental way has helped to control the evolution of the data content. These three design aspects bring clarity and performance regarding data access. They also provide long term scalability to the system and resilience to further changes that may occur in source systems feeding the data warehouse.

Databases, Factual↗

The CAP cancer protocols--a case study of caCORE based data standards implementation to integrate with the Cancer Biomedical Informatics Grid.

BACKGROUND: The Cancer Biomedical Informatics Grid (caBIG) is a network of individuals and institutions, creating a world wide web of cancer research. An important aspect of this informatics effort is the development of consistent practices for data standards development, using a multi-tier approach that facilitates semantic interoperability of systems. The semantic tiers include (1) information models, (2) common data elements, and (3) controlled terminologies and ontologies. The College of American Pathologists (CAP) cancer protocols and checklists are an important reporting standard in pathology, for which no complete electronic data standard is currently available. METHODS: In this manuscript, we provide a case study of Cancer Common Ontologic Representation Environment (caCORE) data standard implementation of the CAP cancer protocols and checklists model--an existing and complex paper based standard. We illustrate the basic principles, goals and methodology for developing caBIG models. RESULTS: Using this example, we describe the process required to develop the model, the technologies and data standards on which the process and models are based, and the results of the modeling effort. We address difficulties we encountered and modifications to caCORE that will address these problems. In addition, we describe four ongoing development projects that will use the emerging CAP data standards to achieve integration of tissue banking and laboratory information systems. CONCLUSION: The CAP cancer checklists can be used as the basis for an electronic data standard in pathology using the caBIG semantic modeling methodology.

Clinical Protocols↗

An integrated medical record and data system for primary care. Part 8: the individual patient's medical record.

This is the last in a series of eight articles describing an integrated system of recording medical data as developed and used by the Family Medicine Program at the University of Rochester-Highland Hospital. Compatability of manual and automated systems has been described. The total system allows the practicing family physician to assess morbidity patterns within his/her practice more effectively, to record and monitor patient care, to perform audit, and to conduct research in primary care.

Humans↗

Integrated analysis of genetic data with R.

Genetic data are now widely available. There is, however, an apparent lack of concerted effort to produce software systems for statistical analysis of genetic data compared with other fields of statistics. It is often a tremendous task for end-users to tailor them for particular data, especially when genetic data are analysed in conjunction with a large number of covariates. Here, R (http://www.r-project.org), a free, flexible and platform-independent environment for statistical modelling and graphics is explored as an integrated system for genetic data analysis. An overview of some packages currently available for analysis of genetic data is given. This is followed by examples of package development and practical applications. With clear advantages in data management, graphics, statistical analysis, programming, internet capability and use of available codes, it is a feasible, although challenging, task to develop it into an integrated platform for genetic analysis; this will require the joint efforts of many researchers.

Algorithms↗

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals↗

Tools enabling the elucidation of molecular pathways active in human disease: application to Hepatitis C virus infection.

BACKGROUND: The extraction of biological knowledge from genome-scale data sets requires its analysis in the context of additional biological information. The importance of integrating experimental data sets with molecular interaction networks has been recognized and applied to the study of model organisms, but its systematic application to the study of human disease has lagged behind due to the lack of tools for performing such integration. RESULTS: We have developed techniques and software tools for simplifying and streamlining the process of integration of diverse experimental data types in molecular networks, as well as for the analysis of these networks. We applied these techniques to extract, from genomic expression data from Hepatitis C virus-infected liver tissue, potentially useful hypotheses related to the onset of this disease. Our integration of the expression data with large-scale molecular interaction networks and subsequent analyses identified molecular pathways that appear to be induced or repressed in the response to Hepatitis C viral infection. CONCLUSION: The methods and tools we have implemented allow for the efficient dynamic integration and analysis of diverse data in a major human disease system. This integrated data set in turn enabled simple analyses to yield hypotheses related to the response to Hepatitis C viral infection.

Computational Biology↗

Simultaneous quantitative evaluation of visual-evoked responses and background EEG activity in rat: normative data.

An integrated quantitative electroencephalography system (Phegra) for pharmacological and toxicological research in rat is described. Peak latencies and amplitudes of visual-evoked potentials, occurrence, duration, and linear excursions of photically evoked afterdischarges, "activity," "mobility," "complexity" of Hjorth, and absolute spectral powers of delta, theta, alpha, and beta frequency bands of background activity of visual cortex and frontal-visual leads were measured in freely moving rats. Counts of small and large movements were also registered. Data of baseline measurements performed in large amount of animals are presented. None of the parameters except the occurrence of photically evoked afterdischarge and the linear excursion of its averaged waveshape changed significantly in five measurements performed within six hours following the intraperitoneal and oral administration of two commonly used drug vehicles.

Animals↗