Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans↗

The evolution of an integrated timeline for oncology patient healthcare.

The introduction of computers in the medical environment has contributed to the proliferation of medical data, often making it difficult to consolidate information on a single patient. In patients with complex medical problems, such as oncology patients, the lack of data integration can negatively impact on patient care. This paper presents an infrastructure for the creation of an integrated multimedia timeline that automatically combines patient information from distributed hospital information sources, and creates a visual summary of pertinent events in a patient's medical history. In this prototype, we focus on oncology patients under treatment for advanced cancers.

Database Management Systems↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

Networking and data management for health care monitoring of mobile patients.

The problem of medical devices and data integration in health care is discussed and a proposal for remote monitoring of patients based on recent developments in networking and data management is presented. In particular the paper discusses the benefits of the integration of personal medical devices into a Medical Information System and how wireless sensor networks and open protocols could be employed as building blocks of a patient monitoring system.

Biomedical Technology↗

A strategy capitalizing on synergies: the Reporting Structure for Biological Investigation (RSBI) working group.

In this article we present the Reporting Structure for Biological Investigation (RSBI), a working group under the Microarray Gene Expression Data (MGED) Society umbrella. RSBI brings together several communities to tackle the challenges associated with integrating data and representing complex biological investigations, employing multiple OMICS technologies. Currently, RSBI includes environmental genomics, nutrigenomics and toxicogenomics communities, where independent activities are underway to develop databases and establish data communication standards within their respective domains. The RSBI working group has been conceived as a "single point of focus" for these communities, conforming to general accepted view that duplication and incompatibility should be avoided where possible. This endeavour has aimed to synergize insular solutions into one common terminology between biologically driven standardisation efforts and has also resulted in strong collaborations and shared understanding between those in the technological domain. Through extensive liaisons with many standards efforts, several threads have been woven with the hope that ultimately technology-centered standards and their specific extensions into biological domains of interest will not only stand alone, but will also be able to function together, as interchangeable modules.

Databases, Genetic↗

Precautions in topographic mapping and in evoked potential map reading.

First, we consider the main points that must be addressed when constructing topographic maps: types of projection, methods of interpolation, number and locations of recording electrodes, and color scales. Data integrity and precautions in map interpretation are then examined for the case of evoked potential data.

Brain↗

An allosteric model for transmembrane signaling in bacterial chemotaxis.

Bacteria are able to sense chemical gradients over a wide range of concentrations. However, calculations based on the known number of receptors do not predict such a range unless receptors interact with one another in a cooperative manner. A number of recent experiments support the notion that this remarkable sensitivity in chemotaxis is mediated by localized interactions or crosstalk between neighboring receptors. A number of simple, elegant models have proposed mechanisms for signal integration within receptor clusters. What is a lacking is a model, based on known molecular mechanisms and our accumulated knowledge of chemotaxis, that integrates data from multiple, heterogeneous sources. To address this question, we propose an allosteric mechanism for transmembrane signaling in bacterial chemotaxis based on the "trimer of dimers" model, where three receptor dimers form a stable complex with CheW and CheA. The mechanism is used to integrate a diverse set of experimental data in a consistent framework. The main predictions are: (1) trimers of receptor dimers form the building blocks for the signaling complexes; (2) receptor methylation increases the stability of the active state and retards the inhibition arising from ligand-bound receptors within the signaling complex; (3) trimer of dimer receptor complexes aggregate into clusters through their mutual interactions with CheA and CheW; (4) cooperativity arises from neighboring interaction within these clusters; and (5) cluster size is determined by the concentration of receptors, CheA, and CheW. The model is able to explain a number of seemingly contradictory experiments in a consistent manner and, in the process, explain how bacteria are able to sense chemical gradients over a wide range of concentrations by demonstrating how signals are integrated within the signaling complex.

Allosteric Regulation↗

Fast fully 3-D image reconstruction in PET using planograms.

We present a method of performing fast and accurate three-dimensional (3-D) backprojection using only Fourier transform operations for line-integral data acquired by planar detector arrays in positron emission tomography. This approach is a 3-D extension of the two-dimensional (2-D) linogram technique of Edholm. By using a special choice of parameters to index a line of response (LOR) for a pair of planar detectors, rather than the conventional parameters used to index a LOR for a circular tomograph, all the LORs passing through a point in the field of view (FOV) lie on a 2-D plane in the four-dimensional (4-D) data space. Thus, backprojection of all the LORs passing through a point in the FOV corresponds to integration of a 2-D plane through the 4-D "planogram." The key step is that the integration along a set of parallel 2-D planes through the planogram, that is, backprojection of a plane of points, can be replaced by a 2-D section through the origin of the 4-D Fourier transform of the data. Backprojection can be performed as a sequence of Fourier transform operations, for faster implementation. In addition, we derive the central-section theorem for planogram format data, and also derive a reconstruction filter for both backprojection-filtering and filtered-backprojection reconstruction algorithms. With software-based Fourier transform calculations we provide preliminary comparisons of planogram backprojection to standard 3-D backprojection and demonstrate a reduction in computation time by a factor of approximately 15.

Algorithms↗

Cognitive evaluation of decision making processes and assessment of information technology in medicine.

This paper describes cognitive methods for analyzing medical decision making and evaluating medical information systems. The overall approach focuses on understanding the processes involved in the decision making and reasoning of health care workers, both with and without the use of information technologies. The issue of developing appropriate evaluation tools, for use in the design and analysis of medical information systems is considered to be of great importance. However, conventional methods are limited in their ability to identify and characterize the effects of information technology on the cognitive processes involved in decision making and reasoning. In this paper a range of methods are described involving video recording for collecting data on the use of information systems. The techniques described allow for the collection of an integrated data set consisting of transcripts of health care workers as they 'think aloud' in interacting with a medical system, along with complete video records of user-computer interaction. In addition, the methods can be extended to allow for the collection of process data from video recording of systems in actual clinical and emergency situations. The use of a variety of approaches, borrowing from research in cognitive science, is discussed. The development and application of these evaluation methods within the Canadian Centres of Excellence network HEALNet is subsequently described. Finally, implications for the development and evaluation of medical information systems are considered.

Cognition↗

Physical mapping: integrating computational and molecular genetic data.

A crucial step beyond the identification of genetic linkage of a disease to a chromosomal region is the production of a physical map that will allow the identification of candidate genes. Although the process of physical map building has been facilitated by the flow of data released by the Human Genome Project, gathering all the information together requires significant effort. In a previous study, we reported linkage between Bipolar Affective Disorder and the chromosomal location 4p15.3--p16.1. In this review we use this example to describe how to collect publicly available sequence, DNA fingerprint, and genetic marker data and integrate these with empirical data to build a large scale high resolution physical map of a region. Methods used to identify new genetic markers and candidate genes within a circumscribed region are also presented.

Databases, Factual↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

Health care informatics: the key to successful disease management.

Health services integration and disease state management (DSM) require improved health care informatics systems. Accurate, comprehensive patient information and an integrated data infrastructure are needed for all stages of DSM, from development and implementation of programs to evaluation and continuous program improvement. The lack of an integrated information infrastructure is one of the leading obstacles to achieving a comprehensive electronic patient data system. This article examines initiatives underway to make the computer-based patient record a reality.

Attitude of Health Personnel↗

GenBank.

GenBank (R) is a comprehensive database that contains publicly available DNA sequences for more than 205 000 named organisms, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the Web-based BankIt or standalone Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in Europe and the DNA Data Bank of Japan ensures worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, go to the NCBI Homepage at www.ncbi.nlm.nih.gov.

Animals↗

GenBank.

GenBank (R) is a comprehensive database that contains publicly available nucleotide sequences for more than 240 000 named organisms, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects. Most submissions are made using the web-based BankIt or standalone Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the EMBL Data Library in Europe and the DNA Data Bank of Japan ensures worldwide coverage. GenBank is accessible through NCBI's retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. To access GenBank and its related retrieval and analysis services, begin at the NCBI Homepage (www.ncbi.nlm.nih.gov).

Animals↗

Unlocking the Full Potential of Spatial Omics in Plants: Practical Challenges, Solutions, and a Path Forward.

Spatial omics technologies are providing new opportunities for plant biology by enabling molecular profiling within structurally intact tissues, revealing spatially organised cell states, developmental gradients, and regulatory interactions. While spatial transcriptomics has driven early advances, the field is rapidly expanding toward integrated spatial multi-omics by combining single-cell and spatial transcriptomic, epigenomic, proteomic, and metabolomic data. These approaches offer new opportunities to study development, physiology, and plant biotic and abiotic interactions in spatially preserved cellular contexts. However, despite rapid adoption, the field remains constrained by plant-specific challenges when applying technologies largely developed for animal systems. Compared with animal systems, plant tissues pose additional challenges due to rigid cell walls, and diverse chemistries, complicating sample preparation, cell and subcellular segmentation, signal detection, and data integration. As a result, many studies rely on bespoke protocols and analysis pipelines that are often difficult to reproduce or generalise. Here, we provide a practical, solution-oriented synthesis of current bottlenecks across experimental and computational pipelines, highlight emerging strategies to overcome these limitations, and propose a roadmap for community-driven protocol sharing, benchmarking, and integration across spatial and multi-omics modalities. Addressing these challenges will be essential to establish spatial omics as a routine and scalable tool for plant biology.

Journal Article↗

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans↗

Initiating informatics and GIS support for a field investigation of Bioterrorism: The New Jersey anthrax experience.

BACKGROUND: The investigation of potential exposure to anthrax spores in a Trenton, New Jersey, mail-processing facility required rapid assessment of informatics needs and adaptation of existing informatics tools to new physical and information-processing environments. Because the affected building and its computers were closed down, data to list potentially exposed persons and map building floor plans were unavailable from the primary source. RESULTS: Controlling the effects of anthrax contamination required identification and follow-up of potentially exposed persons. Risk of exposure had to be estimated from the geographic relationship between work history and environmental sample sites within the contaminated facility. To assist in establishing geographic relationships, floor plan maps of the postal facility were constructed in ArcView Geographic Information System (GIS) software and linked to a database of personnel and visitors using Epi Info and Epi Map 2000. A repository for maintaining the latest versions of various documents was set up using Web page hyperlinks. CONCLUSIONS: During public health emergencies, such as bioterrorist attacks and disease epidemics, computerized information systems for data management, analysis, and communication may be needed within hours of beginning the investigation. Available sources of data and output requirements of the system may be changed frequently during the course of the investigation. Integrating data from a variety of sources may require entering or importing data from a variety of digital and paper formats. Spatial representation of data is particularly valuable for assessing environmental exposure. Written documents, guidelines, and memos important to the epidemic were frequently revised. In this investigation, a database was operational on the second day and the GIS component during the second week of the investigation.

Journal Article↗

[Use of satellites for public health purposes in tropical areas].

The epidemiological hallmark of the new millennium has been the emergence or recrudescence of transmissible diseases with high epidemic potential. Disease tracking is becoming an increasingly global task requiring implementation of more and more sophisticated control strategies and facilities for sustainable development. A promising initiative involves the use of satellite technology to monitor and forecast the spread of disease. The Health Early Warning System (HEWS) was designed based on successful application of satellite data in food programs as well as in other areas (e.g. weather, farming and fishing). The HEWS integrates data from communications, remote-sensing and positioning satellites. The purpose of this review is to present the main studies containing satellite data on public health in tropical areas. Satellite data has allowed development of more reactive epidemiological tracking networks better suited to increasing population mobility, correlation of environmental factors (vegetation index, rainfall and ocean surface color) with human, animal and insect factors in epidemiological studies and assessment of the role of such factors in the development or reappearance of disease. Satellite technology holds great promise for more efficient management of public health problems in tropical areas.

Cholera↗