Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation↗

SurvGRN: a multi-feature fusion framework for bladder cancer survival prediction.

Bladder cancer survival outcomes exhibit significant heterogeneity, influenced by multifaceted factors. While digital pathology-based survival models leveraging artificial intelligence show promise, they often overlook complementary data sources. Conversely, imaging lacks cellular detail, and genomics/proteomics entail complexity and cost. To integrate multidimensional data for enhanced survival prediction, we propose SurvGRN, a multi-feature fusion framework. SurvGRN synergistically combines clinical variables, transcriptomics, and digital pathology slides using a gated residual network architecture. Pathological features are extracted via multiple instance learning, while clinical and transcriptomic data are processed as static inputs. These features are dynamically fused using a long short-term memory (LSTM) network for comprehensive survival risk assessment. Evaluated on 400 bladder cancer patients, SurvGRN significantly outperformed existing methods: improving the C-index by 12.6% over DeepMISL; 20.6% and 7.1% over graph-based models (DeepGraphConv and Patch-GCN); and 5.4% and 4.0% over attention-based approaches (Surformer and HVTSurv). Ablation studies confirmed the contributions of pathology features (extracted via ResNet-50 pre-trained on bladder tissue), clinical/transcriptomic data, and the LSTM fusion. SurvGRN also enabled significant stratification of patients into distinct risk cohorts. This work demonstrates that holistic integration of multi-source data through tailored fusion architectures substantially improves bladder cancer survival prediction.

bladder cancer↗

The integration of macromolecular diffraction data.

The objective of any modern data-processing program is to produce from a set of diffraction images a set of indices (hkls) with their associated intensities (and estimates of their uncertainties), together with an accurate estimate of the crystal unit-cell parameters. This procedure should not only be reliable, but should involve an absolute minimum of user intervention. The process can be conveniently divided into three stages. The first (autoindexing) determines the unit-cell parameters and the orientation of the crystal. The unit-cell parameters may indicate the likely Laue group of the crystal. The second step is to refine the initial estimate of the unit-cell parameters and also the crystal mosaicity using a procedure known as post-refinement. The third step is to integrate the images, which consists of predicting the positions of the Bragg reflections on each image and obtaining an estimate of the intensity of each reflection and its uncertainty. This is carried out while simultaneously refining various detector and crystal parameters. Basic features of the algorithms employed for each of these three separate steps are described, principally with reference to the program MOSFLM.

Algorithms↗

Unsupervised pattern recognition: an introduction to the whys and wherefores of clustering microarray data.

Clustering has become an integral part of microarray data analysis and interpretation. The algorithmic basis of clustering -- the application of unsupervised machine-learning techniques to identify the patterns inherent in a data set -- is well established. This review discusses the biological motivations for and applications of these techniques to integrating gene expression data with other biological information, such as functional annotation, promoter data and proteomic data.

Algorithms↗

Computational metabolomics at scale: from open data to insight.

Metabolomics data are currently generated at scale thanks to the evolution of technologies that have led to marked improvements in the number of metabolites detected, spanning all chemical classes. These data are increasingly submitted to public repositories for data reuse, integration, and interpretation. Despite the availability of public resources and associated computational tools, the field still lacks a widely adopted, consistent data and analytics infrastructure capable of transforming this wealth of information into scientific insight. Indeed, the metabolomics field is just now scratching the surface of being able to harness the power of new computational technologies. In this review, we summarize discussions from the "Dagstuhl-Seminar 24181 Computational Metabolomics: Towards Molecules, Models, and their Meaning" with a focus on public data availability, open data standards, data and knowledge integration, and education. Our goal is to raise awareness and adoption of the latest open science resources while highlighting key areas needing further development.

Metabolomics↗

DBMap: a space-conscious data visualization and knowledge discovery framework for biomedical data warehouse.

Advances in digital imaging modalities as well as other diagnosis and therapeutic techniques have generated a massive amount of diverse data for clinical research. The purpose of this study is to investigate and implement a new intuitive and space-conscious visualization framework, called DBMap, to facilitate efficient multidimensional data visualization and knowledge discovery against the large-scale data warehouses of integrated image and nonimage data. The DBMap framework is built upon the TreeMap concept. TreeMap is a space constrained graphical representation of large hierarchical data sets, mapped to a matrix of rectangles, whose size and color represent interested database fields. It allows the display of a large amount of numerical and categorical information in limited real estate of the computer screen with an intuitive user interface. DBMap has been implemented and integrated into a large brain research data warehouse to support neurologic and neuroradiologic research at the University of California, San Francisco Medical Center. For imaging specialists and clinical researchers, this novel DBMap framework facilitates another way to better explore and classify the hidden knowledge embedded in medical image data warehouses.

Algorithms↗

Data standards: a call to action.

Access to data is something that every molecular biologist takes for granted nowadays, but data alone is of little use unless it is made available in a useable form through the development and global uptake of data standards. The challenge of standards development has been taken up by grass-roots movements working within several different branches of the biomedical research community. Many of these initiatives are proving extremely successful; for example, the Gene Ontology, which provides a controlled vocabulary for describing the properties of gene products, the Microarray Gene Expression Data Society's standards for describing microarray experiments, and the emerging standards developed by the Proteomics Standards Initiative are gaining broad acceptance. Standards development now faces its greatest ever challenge--the integration of diverse data types to fulfill the goals of systems biology. Now is the time for the communities that are developing these standards, the funding bodies that have invested so heavily in high-throughput data generation, and the publishers of biomedical research papers to cooperate fully to make the goals of integrated data analysis a reality.

Animals↗

[Development of psychogenic disorders: an integrative model].

On the basis of clinical experience and empirically based data an integrative model of how psychogenic disorders develop is described in this article. The development-psychological steps of maturation from the uterine period to adolescence are examined with regard to the respective basic conflict to be derived from the step, and the disorder forms neurotization, structural disorder, and traumatisation are differentiated. Especially the process character of the respective development from the basic conflicts over the different coping strategies up to the symptom outbreak is emphasized.

Adaptation, Psychological↗

MR image reconstruction algorithms for sparse k-space data: a Java-based integration.

We have worked on multi-dimensional magnetic resonance imaging (MRI) data acquisition and related image reconstruction methods that aim at reducing the MRI scan time. To achieve this scan-time reduction we have combined the approach of 'increasing the speed' of k-space acquisition with that of 'deliberately omitting' acquisition of k-space trajectories (sparse sampling). Today we have a whole range of (sparse) sampling distributions and related reconstruction methods. In the context of a European Union Training and Mobility of Researchers project we have decided to integrate all methods into one coordinating software system. This system meets the requirements that it is highly structured in an object-oriented manner using the Unified Modeling Language and the Java programming environment, that it uses the client-server approach, that it allows multi-client communication sessions with facilities for sharing data and that it is a true distributed computing system with guaranteed reliability using core activities of the Java Jini package.

Algorithms↗

CRC Tissue Core Management System (TCMS): integration of basic science and clinical data for translational research.

The Chronic Lymphocytic Leukemia (CLL) Research Consortium (CRC) consists of 9 geographically distributed sites conducting a program of research including both basic science and clinical components. The CRC TCMS was designed to capture and integrate basic science and clinical data sets. The system utilizes multiple data modeling methodologies and web-application platforms, and was designed with the high level objectives of providing an extensible, generalizable model for integrating data as required to conduct translational research.

Biomedical Research↗

TongueTwister: an integrated program for analyzing lickometer data.

The analysis of lickometer data is often rendered prohibitively tedious by the large volume of data generated by the typical experiment. TongueTwister is an integrated program for the rapid and automatic analysis, presentation, and summary of long- and medium-access data collected by lickometers or of brief-access data collected by multi-bottle lickometers such as the DiLog Instruments MS80. The program was written in C+2 for Macintosh computers, and analyzes data collected by MS-DOS PCs. It takes advantage of the Macintosh user interface to provide quick and convenient output from all the files of a single experimental session, and to export the data to third-party statistical software or other documents. It can batch-process data files by automatically opening and analyzing all the files in a directory; thus, the user can employ directories as a simple database for organizing experimental groups. When a lickometer data file is opened, a textual summary, a raster plot of the lick pattern, the cumulative licks, the lick rate, a histogram of inter-lick intervals, and a breakdown of the session by fractions are automatically calculated and displayed. When an MS80 brief-access file is opened, the lick pattern for each tube presentation and a textual summary of the mean values derived for each tube are automatically displayed. If a directory of files is opened, the mean values derived across all the individual files are calculated and graphed. Analysis parameters can be tailored to the investigator's liking. Tables or graphs can be saved to disk, or copied and pasted into other Macintosh programs for additional analysis. The program may also be used for general-purpose analysis of periodic event records.

Animals↗

Integration of gel-based proteome data with pProRep.

UNLABELLED: pProRep is a web application integrating electrophoretic and mass spectral data from proteome analyses into a relational database. The graphical web-interface allows users to upload, analyse and share experimental proteome data. It offers researchers the possibility to query all previously analysed datasets and can visualize selected features, such as the presence of a certain set of ions in a peptide mass spectrum, on the level of the two-dimensional gel. AVAILABILITY: The pProRep package and instructions for its use can be downloaded from http://www.ptools.ua.ac.be/pProRep. The application requires a web server that runs PHP 5 (http://www.php.net) and MySQL. Some (non-essential) extensions need additional freely available libraries: details are described in the installation instructions.

Computational Biology↗

Telemedicine: the new must for surgery.

The vision of telesurgery comprises a multitude of new communicative elements influencing the way surgeons will treat their patients in the future. The first prerequisite for effective telecommunication is to digitize surgical data. Many medical imaging modalities provide primarily digital data sets, and digital image communication is already entering clinical practice under the labels of teleradiology and telepathology. However, for any surgical purpose, images must refer to tissues. Three-dimensional image reconstruction is warranted, and if such data shall be useful during surgery, different image sources must be combined into some virtual, multiparametric body model and matched to an intraoperatively distorted organ contour. A multitude of detail problems arise, beginning with image standards, data interfaces, data transport, image fusion, registering, contour matching, and, once the data are integrated, all the aspects of surgery-suitable data display and interaction. We refer here to several demonstration projects illustrating such a complex surgical data set and its interactive telecommunication. In all instances, telecommunication was to enable a concentration of distributed medical intelligence at the site where the patient was treated. With further technological development, such telesurgical applications will have a growing influence on patient management and surgical decision making. In the very near future, computer-aided navigation and robotic assistance, based on the same surgical data sets, will be available to all fields of surgery. How decisive the role these methods will play for specific procedures or diseases needs to be determined.

Computer Simulation↗

Integrating health-related data from various sources: combining surveys, records and routine data.

The paper discusses the necessity of combining data from various sources in order to enhance their usefulness for a variety of applications. As future data use can hardly be forseen in advance, major data sets in health services should fulfill several formal requirements in order to make them suitable for future linkage. These formal requirements are that there be references to defined populations, to specific persons, to defined time periods, to specific places or regions. It would be necessary for terms, definition and classification schemes to agree between data sets which are to be linked and be in wide use. Three facets of data linkage are discussed specifically namely linking data at one level of aggregation, linking different data components, and combining data sets from different sources at several levels of aggregation. Three examples are provided, describing linkages of data from various sources for epidemiological studies and a study in health services research. They show that at this point in descriptive epidemiological studies linkage on the basis of regions is of great importance. This implies that it would be desirable for large scale data collection activities in health services to provide for a uniform representation of the geographic areas. Such uniformity would greatly enhance the linkage potential of data sets and thus their usefulness for small area and regional analyses.

Adult↗

Integrating behavioral and neural data in a model of zebrafish network interaction.

The spinal neural networks of larval zebrafish (Danio rerio) generate a variety of movements such as escape, struggling, and swimming. Various mechanisms at the neural and network levels have been proposed to account for switches between these behaviors. However, there are currently no detailed demonstrations of such mechanisms. This makes determining which mechanisms are plausible extremely difficult. In this paper, we propose a detailed biologically plausible model of the interactions between the swimming and escape networks in the larval zebrafish, while taking into account anatomical and physiological evidence. We show that the results of our neural model generate the expected behavior when used to control a hydrodynamic model of carangiform locomotion. As a result, the model presented here is a clear demonstration of a plausible mechanism by which these distinct behaviors can be controlled. Interestingly, the networks are anatomically overlapping, despite clear differences in behavioral function and physiology.

Animals↗

Trihalomethanes in drinking water and cancer: risk assessment and integrated evaluation of available data, in animals and humans.

In our study, we attempted to jointly consider THM concentration data collected from drinking waters and carcinogenic risk assessment derived from mathematical models commonly used in this field (multi-stage models for laboratory animal experimentation data, and 'unit risk' derived from the relative risk in the case of epidemiological data). In order to estimate the risks related to joint exposure to different THMs, in this study the risk additivity hypothesis is taken into account. Based on animal data for the various tumors, carcinogenic risk estimates for different THM combinations vary from 2.7 x 10(-7) to 4.6 x 10(-6) per micrograms/l in relation to different carcinogenic substances published in the literature or specifically calculated in this study. The carcinogenic risk parameters derived from experimental studies and from epidemiological data were substantially consistent. Our study uses also as an example some data on concentration levels of THMs for drinking water supplies in Sardinia. The area mean THM concentration values for each supply varied, for ground waters, from 8.1 to 13.6 micrograms/l and, for surface waters, from 52.8 to 168 micrograms/l. For the 1976-1989 period, bladder cancer standardized mortality rates in the water distribution system areas where the THMs were measured indicate values similar, but generally lower, than the national ones, except in the province of Cagliari where the values were not significantly different. The risk estimates derived from animal studies are of the same order of magnitude as the epidemiological data in literature.

Animals↗

Genomic biomarkers for cancer assessment: implementation challenges for laboratory practice.

Genomic biomarkers are an emerging class of laboratory tests, which present special implementation challenges for clinical laboratory services, compared to conventional laboratory tests. These challenges, which include analytical, bioinformatics, bioethical, interpretation and commercialization issues, represent real obstacles to widespread implementation of these tests. Technical challenges include the capacity to detect and identify many different kinds of markers for different diseases in a short time period, capacity to identify simultaneously gene rearrangements, amplification, inhibition, deletions and replications. Bioinformatics challenges include rapid analysis of genomic data, as well as the cross reference to other genomic data, and to other laboratory tests. Bioethical issues relate to consent to retain and use genetic data, which may be obtained inadvertently during analysis for genomic markers. Interpretation challenges include observations that the particular genomic markers may not be independent variables, as other undetected genomic alterations could invalidate or alter genomic marker interpretation. Further, as early experience with predictive genetic markers for cancer has shown, proprietary commercial interests may conflict with public health values of identifying genomic markers in subject populations. Based on our 10 years of experience with genomic biomarkers, important implementation strategies for genomic markers include development of:Standard high throughput analyzers capable of detecting any alteration of any genomic variant at any time. Bioinformatics analysis online, coupled to stored patient data. Laboratory service framework that preserves confidentiality but integrates genomic data with other laboratory tests. Laboratory service framework, which links consents, genomic analysis, reports to both specimen and data repositories. Overall, the laboratory service challenges for genomic markers are to manage very large analytical sets and very large data sets in finite time with responsible interpretation, all within finite funding. To meet these challenges, implementation strategies beyond the one disease, one diagnosis, one genomic marker concept must begin now.

Biomarkers, Tumor↗