Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A statistical method for chromatographic alignment of LC-MS data.

Integrated liquid-chromatography mass-spectrometry (LC-MS) is becoming a widely used approach for quantifying the protein composition of complex samples. The output of the LC-MS system measures the intensity of a peptide with a specific mass-charge ratio and retention time. In the last few years, this technology has been used to compare complex biological samples across multiple conditions. One challenge for comparative proteomic profiling with LC-MS is to match corresponding peptide features from different experiments. In this paper, we propose a new method--Peptide Element Alignment (PETAL) that uses raw spectrum data and detected peak to simultaneously align features from multiple LC-MS experiments. PETAL creates spectrum elements, each of which represents the mass spectrum of a single peptide in a single scan. Peptides detected in different LC-MS data are aligned if they can be represented by the same elements. By considering each peptide separately, PETAL enjoys greater flexibility than time warping methods. While most existing methods process multiple data sets by sequentially aligning each data set to an arbitrarily chosen template data set, PETAL treats all experiments symmetrically and can analyze all experiments simultaneously. We illustrate the performance of PETAL on example data sets.

Animals↗

The VA's use of DICOM to integrate image data seamlessly into the online patient record.

The US Department of Veterans Affairs (VA) is using the Digital Imaging and Communications in Medicine (DICOM) standard to integrate image data objects from multiple systems for use across the healthcare enterprise. DICOM uses a structured representation of image data and a communication mechanism that allows the VA to easily acquire radiology images and store them directly into the online patient record. Images can then be displayed on low-cost clinician's workstations throughout the medical center. High-resolution diagnostic quality multi-monitor VistA workstations with specialized viewing software can be used for reading radiology images. Various image and study specific items from the DICOM data object are essential for the correct display of images. The VA's DICOM capabilities are now used to interface seven different commercial Picture Archiving and Communication Systems (PACS) and over twenty different radiology image acquisition modalities.

Evaluation Studies as Topic↗

scMGCL: accurate and efficient integration representation of single-cell multi-omics data.

MOTIVATION: Single-cell multi-omics data integration is essential for understanding cellular states and disease mechanisms, yet integrating heterogeneous data modalities remains a challenge. We present scMGCL, a graph contrastive learning framework for robust integration of single-cell ATAC-seq and RNA-seq data. Our approach leverages self-supervised learning on cell-cell similarity graphs, in which each modality's graph structure serves as an augmentation for the other. This cross-modality contrastive paradigm enables the learning of biologically meaningful, shared representations while preserving modality-specific features. RESULTS: Benchmarking against state-of-the-art methods demonstrates that scMGCL outperforms others in cell-type clustering, label transfer accuracy, and preservation of marker-gene correlations. Additionally, scMGCL significantly improves computational efficiency, reducing runtime and memory usage. The method's effectiveness is further validated through extensive analyses of cell-type similarity and functional consistency, providing a powerful tool for multi-omics data exploration. AVAILABILITY AND IMPLEMENTATION: Code and datasets are released at https://github.com/zlCreator/scMGCL.

Single-Cell Analysis↗

An integrated data-warehouse-concept for clinical and biological information.

The development of medical research networks within the framework of translational research has fostered interest in the integration of clinical and biological research data in a common database. The building of one single database integrating clinical data and biological research data requires a concept which enables scientists to retrieve information and to connect known facts to new findings. Clinical parameters are collected by a Patient Data Management System and viewed in a database which also includes genomic data. This database is designed as an Entity Attribute Value model, which implicates the development of a data warehouse concept. For the realization of this project, various requirements have to be taken into account which has to be fulfilled sufficiently in order to align with international standards. Data security and protection of data privacy are most important parts of the data warehouse concept. It has to be clear how patient pseudonymization has to be carried out in order to be within the scope of data security law. To be able to evaluate the data stored in a database consisting of clinical data collected by a Patient Data Management System and genomic research data easily, a data warehouse concept based on an Entity Attribute Value datamodel has been developed.

Computer Security↗

caCORE: a common infrastructure for cancer informatics.

MOTIVATION: Sites with substantive bioinformatics operations are challenged to build data processing and delivery infrastructure that provides reliable access and enables data integration. Locally generated data must be processed and stored such that relationships to external data sources can be presented. Consistency and comparability across data sets requires annotation with controlled vocabularies and, further, metadata standards for data representation. Programmatic access to the processed data should be supported to ensure the maximum possible value is extracted. Confronted with these challenges at the National Cancer Institute Center for Bioinformatics, we decided to develop a robust infrastructure for data management and integration that supports advanced biomedical applications. RESULTS: We have developed an interconnected set of software and services called caCORE. Enterprise Vocabulary Services (EVS) provide controlled vocabulary, dictionary and thesaurus services. The Cancer Data Standards Repository (caDSR) provides a metadata registry for common data elements. Cancer Bioinformatics Infrastructure Objects (caBIO) implements an object-oriented model of the biomedical domain and provides Java, Simple Object Access Protocol and HTTP-XML application programming interfaces. caCORE has been used to develop scientific applications that bring together data from distinct genomic and clinical science sources. AVAILABILITY: caCORE downloads and web interfaces can be accessed from links on the caCORE web site (http://ncicb.nci.nih.gov/core). caBIO software is distributed under an open source license that permits unrestricted academic and commercial use. Vocabulary and metadata content in the EVS and caDSR, respectively, is similarly unrestricted, and is available through web applications and FTP downloads. SUPPLEMENTARY INFORMATION: http://ncicb.nci.nih.gov/core/publications contains links to the caBIO 1.0 class diagram and the caCORE 1.0 Technical Guide, which provide detailed information on the present caCORE architecture, data sources and APIs. Updated information appears on a regular basis on the caCORE web site (http://ncicb.nci.nih.gov/core).

Animals↗

The ASH HematOmics Program supports integrative analysis of genomic and clinical data in hematologic diseases.

The increasing availability of genomic and transcriptomic sequencing has uncovered diverse genomic alterations and distinct gene expression profiles driving hematologic diseases, yet a data integration and sharing platform dedicated to hematology remains lacking. We developed the American Society of Hematology (ASH) HematOmics Program (ASHOP; ashop.hematology.org), a resource for exploring somatic alterations and gene fusions, transcriptomic results, and clinical data from 5960 patients spanning B-cell precursor and T-cell acute lymphoblastic leukemia, acute myeloid leukemia, myelodysplastic syndromes, and chronic lymphocytic leukemia. Users can explore genomic alteration landscapes and comutation patterns via lollipop and matrix plots and analyze significantly altered genes in user-defined subcohorts. Transcriptomes can be explored through interactive uniform manifold approximation and projections, clustering, differential expression, and pathway enrichment. Genomic, transcriptomic features, and clinical outcomes can be correlated in a user-driven manner or combined to precisely define study cohorts. We illustrate the following 4 use cases of ASHOP: (1) stratification of DUX4-rearranged B-cell leukemias into Early/Multipotent and Committed subgroups with distinct outcomes, (2) characterization of HOXA/HOXB expression patterns in acute myeloid leukemias, (3) correlating mutational burden with mismatch repair deficiency and mutational signatures, and (4) investigation of TP53 alteration landscape. ASHOP is an open-access resource to inform genomic and transcriptomic data interpretation for hematologic malignancies and will expand to support additional diseases and data modalities from the ASH community.

Humans↗

Dyslexia: the possible benefit of multimodal integration of fMRI- and EEG-data.

Biological research about dyslexia has been conducted using various neuroimaging methods like functional Magnetic Resonance Imaging (fMRI) or Electroencephalography (EEG). Since language functions are characterized by both distributed network activities and speed of processing within milliseconds, high temporal as well as high spatial resolution of activation profiles are of interest: "where" can dyslexia specific activations be detected and "when" do language processes start to diverge between dyslexics and controls? Due to the network character of language processing, fMRI-constrained distributed source models based on EEG-data were computed for multimodal data integration. First single-case results show that this method could be a promising approach for the understanding of a repeatedly described experimental finding for dyslexia like that of an overactivation in inferior frontal language areas. Multimodal data analysis for the subjects presented here could probably demonstrate that inferior frontal overactivations are the consequence of a phonological deficit and could represent ongoing articulation processes used to solve phonologically challenging tasks.

Adolescent↗

ArrayXPath: mapping and visualizing microarray gene-expression data with integrated biological pathway resources using Scalable Vector Graphics.

Biological pathways can provide key information on the organization of biological systems. ArrayXPath (http://www.snubi.org/software/ArrayXPath/) is a web-based service for mapping and visualizing microarray gene-expression data for integrated biological pathway resources using Scalable Vector Graphics (SVG). By integrating major bio-databases and searching pathway resources, ArrayXPath automatically maps different types of identifiers from microarray probes and pathway elements. When one inputs gene-expression clusters, ArrayXPath produces a list of the best matching pathways for each cluster. We applied Fisher's exact test and the false discovery rate (FDR) to evaluate the statistical significance of the association between a cluster and a pathway while correcting the multiple-comparison problem. ArrayXPath produces Javascript-enabled SVGs for web-enabled interactive visualization of pathways integrated with gene-expression profiles.

Cluster Analysis↗

A profile for managing sensory integrative test data.

A concise method for compiling a data profile from a general sensory integrative test battery has been presented. Subtests from each test used were categorized according to the sensory integrative and motor functions being tested. These categories have been defined and include: tactile-kinesthetic perception, visual perception-figure ground, visual perception-constancy, ocular control, gross motor control, fine motor control, integration of function-two sides of the body, orientation in space, body awareness, and auditory discrimination. A method for converting the various scores into descriptive terminology is provided in which the test results are reported as above age expectancy, appropriate for age, somewhat deficient for age, and markedly deficient for age. The clinical implications of the technique are discussed.

Auditory Perception↗

The relation between observed and perceived family structures: an integrative model.

Two methods of family structure evaluation were compared, the Georgia Family Q-sort which provides an observational measure of family interaction and the FACES III which is a self-report measure of cohesion and adaptability of the real and ideal family as perceived. Both measures were used with family groups of various sizes and compositions. Integrative data association of family-relevant information was used to assess the interactional structure of 20 one-parent and 20 two-parent family groups. The Q-sort procedure provided integrative data in terms of the situations in which the family is observed. The self-report measure provided integrative data in terms of responders' self-relevant perceptions (social or personal). Some support was found for both procedures. Correlations between the Q-sort clusters and self-report clusters were considered. The integrated use of the two methods is preferred because the integration provides more information about family structure.

Adolescent↗

Integration of data obtained at fixed intervals.

This article discusses the design and implementation of a program well suited to integrating experimental or simulated data obtained at fixed intervals. The program uses Simpson's method and produces substantially better accuracy than trapezoidal rule integration at little extra computational cost. It accepts command line specification of integration parameters (step size and/or number) and source files. Multiple source files and integration parameters can be specified at runtime. Output can be displayed on the console or redirected to an ASCII file.

Computers↗

RNAcare: integrating clinical data with transcriptomic evidence using rheumatoid arthritis as a case study.

BACKGROUND: Gene expression analysis is a crucial tool for uncovering the biological mechanisms that underlie differences between patient subgroups, offering insights that can inform clinical decisions. However, despite its potential, gene expression analysis remains challenging for clinicians due to the specialised skills required to access, integrate, and analyse large datasets. Existing tools primarily focus on RNA-Seq data analysis, providing user-friendly interfaces but often falling short in several critical areas: they typically do not integrate clinical data, lack support for patient-specific analyses, and offer limited flexibility in exploring relationships between gene expression and clinical outcomes in disease cohorts. Users, including clinicians with a general knowledge of transcriptomics, however, who may have limited programming experience, are increasingly seeking tools that go beyond traditional analysis. To overcome these issues, computational tools must incorporate advanced techniques, such as machine learning, to better understand how gene expression correlates with patient symptoms of interest. RESULTS: Our RNAcare platform, addresses these limitations by offering an interactive and reproducible solution specifically designed for analysing transcriptomic data from patient samples in a clinical context. This enables researchers to directly integrate gene expression data with clinical features, perform exploratory data analysis, and identify patterns among patients with similar diseases. By enabling users to integrate transcriptomic and clinical data, and customise the target label, the platform facilitates the analysis of the relationships between gene expression and clinical symptoms like pain and fatigue. This allows users to generate hypotheses and illustrative visualisations/reports to support their research. As proof of concept, we use RNAcare to link inflammation-related genes to pain and fatigue in rheumatoid arthritis (RA) and detect signatures in the drug response group, confirming previous findings. CONCLUSION: We present a novel computational platform allowing the interpretation of clinical and transcriptomics data in real-time. The platform can be used for data generated by the user, such as the patient data presented here or using published datasets. The platform is available at https://rna-care.mvls.gla.ac.uk/ , and its source code is https://github.com/sii-scRNA-Seq/RNAcare/ .

Humans↗

Model selection for integrated recovery/recapture data.

Catchpole et al. (1998, Biometrics 54, 33-46) provide a novel scheme for integrating both recovery and recapture data analyses and derive sufficient statistics that facilitate likelihood computations. In this article, we demonstrate how their efficient likelihood expression can facilitate Bayesian analyses of these kinds of data and extend their methodology to provide a formal framework for model determination. We consider in detail the issue of model selection with respect to a set of recapture/recovery histories of shags (Phalacrocorax aristotelis) and determine, from the enormous range of biologically plausible models available, which best describe the data. By using reversible jump Markov chain Monte Carlo methodology, we demonstrate how this enormous model space can be efficiently and effectively explored without having to resort to performing an infeasibly large number of pairwise comparisons or some ad hoc stepwise procedure. We find that the model used by Catchpole et al. (1998) has essentially zero posterior probability and that, of the 477,144 possible models considered, over 60% of the posterior mass is placed on three neighboring models with biologically interesting interpretations.

Animals↗

Ecological data in Integrated Coastal Zone Management: case study of Posidonia oceanica meadows along the Corsican coastline (Mediterranean Sea).

Integrated Coastal Zone Management (ICZM) contributes towards maximizing the benefits provided by the coastal zone and minimizing conflicts and the harmful effects of activities upon each other. The coastal zone includes highly productive and biologically diverse ecosystems, but ecological data (including structure and processes) seem to be neglected. The purpose of this article is to present a case study of Posidonia oceanica meadows (seagrass beds) along the Corsican coastline (Mediterranean Sea) in order to exemplify the usefulness of ecological data to Integrated Costal Management programs. We will try to determine how the use of organisms could be enhanced. These investigations show the undoubted success of the Corsican Posidonia oceanica protection program, with a detailed description of the ICZM that precisely presents each component (e.g., mapping, assessment of water quality, implementation of a system to aid decision-making concerning the installation of new aquaculture units). This experience on the Corsican coasts could be used as an example in order to transfer to other locations in the Mediterranean Sea and/or to other target species.

Alismatales↗

3D RBI-EM reconstruction with spherically-symmetric basis function for SPECT rotating slat collimator.

A single photon emission computed tomography (SPECT) rotating slat collimator with strip detector acquires distance-weighted plane integral data, along with the attenuation factor and distance-dependent detector response. In order to image a 3D object, the slat collimator device has first to spin around its axis and then rotate around the object to produce 3D projection measurements. Compared to the slice-by-slice 2D reconstruction for the parallel-hole collimator and line integral data, a more complex 3D reconstruction is needed for the slat collimator and plane integral data. In this paper, we propose a 3D RBI-EM reconstruction algorithm with spherically-symmetric basis function, also called 'blobs', for the slat collimator. It has a closed and spherically symmetric analytical expression for the 3D Radon transform, which makes it easier to compute the plane integral than the voxel. It is completely localized in the spatial domain and nearly band-limited in the frequency domain. Its size and shape can be controlled by several parameters to have desired reconstructed image quality. A mathematical lesion phantom study has demonstrated that the blob reconstruction can achieve better contrast-noise trade-offs than the voxel reconstruction without greatly degrading the image resolution. A real lesion phantom study further confirmed this and showed that a slat collimator with CZT detector has better image quality than the conventional parallel-hole collimator with NaI detector. The improvement might be due to both the slat collimation and the better energy resolution of the CZT detector.

Algorithms↗

Successful integration requires data on disease distribution and utilization patterns.

Planning and development: Where should integrated networks locate or establish contracts with physician offices, hospitals, nursing homes or other facilities? A close look at demand and supply data is the only way to effectively determine your community's needs. Sources abound for such data, but there are a few things you need to be careful about when using national or regional information.

Catchment Area, Health↗

Integrated geophysical and chemical study of saline water intrusion.

Surface geophysical surveys provide an effective way to image the subsurface and the ground water zone without a large number of observation wells. DC resistivity sounding generally identifies the subsurface formations-the aquifer zone as well as the formations saturated with saline/brackish water. However, the method has serious ambiguities in distinguishing the geological formations of similar resistivities such as saline sand and saline clay, or water quality such as fresh or saline, in a low resistivity formation. In order to minimize the ambiguity and ascertain the efficacy of data integration techniques in ground water and saline contamination studies, a combined geophysical survey and periodic chemical analysis of ground water were carried out employing DC resistivity profiling, resistivity sounding, and shallow seismic refraction methods. By constraining resistivity interpretation with inputs from seismic refraction and chemical analysis, the data integration study proved to be a powerful method for identification of the subsurface formations, ground water zones, the subsurface saline/brackish water zones, and the probable mode and cause of saline water intrusion in an inland aquifer. A case study presented here illustrates these principles. Resistivity sounding alone had earlier failed to identify the different formations in the saline environment. Data integration and resistivity interpretation constrained by water quality analysis led to a new concept of minimum resistivity for ground water-bearing zones, which is the optimum value of resistivity of a subsurface formation in an area below which ground water contained in it is saline/brackish and unsuitable for drinking.

Data Collection↗