Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Proposed new data file standard for flow cytometry, version FCS 3.0.

In 1984, the first flow cytometry data file format was proposed as Flow Cytometry Standard 1.0 (FCS1.0). FCS 1.0 provided a uniform file format allowing data acquired on one computer to be correctly read and interpreted on other computers running a variety of operating systems. That standard was modified in 1990 and adopted by the Society of Analytical Cytology as FCS 2.0. Here, we report on an update of the FCS 2.0 standard which we propose to designate FCS 3.0. We have retained the basic four segment structure of earlier versions (HEADER, TEXT, DATA and ANALYSIS) in order to maintain analysis software compatibility, where possible. The changes described in this proposal include a method to collect files larger than 100 megabytes (not possible in earlier versions of the standard), the inclusion of international characters in the TEXT portions of the file, a method of verifying data integrity using a 16-bit cyclic redundancy check, and increased keyword support for cluster analysis and time acquisition. This report summarizes the work of the ISAC Data File Standards Committee. The complete and detailed FCS 3.0 standard is available through the ISAC office [Sherwood Group, 60 Revere Drive, Ste 500, Northbrook, IL 60062, phone: (847) 480-9080 ext. 231, fax: (847) 480-9282, E-mail: isac@sherwood-group.com] or through the internet at the ISAC WWW site, http://nucleus.immunol.washington.edu/ISAC.ht ml.

Database Management Systems↗

Threshold integration of bi-amplitude signals.

This study examined the pattern of intensity integration at threshold. The stimuli studied are unique in that they have a compound peak-to-peak amplitude envelope. This waveform was partitioned into two segments, each having a different peak-to-peak magnitude (bi-amplitude). All signals in the bi-amplitude series were the same duration (100 ms). Therefore, threshold differences between these signals are due solely to the integration of intensity in the amplitude dimension. A prediction of the pattern of thresholds, based on the diverted-input hypothesis, suggested that little or no integration would occur when the amplitude difference between segments is greater than a specific magnitude. Our results indicate that there are similarities in the integration process found with variable duration signals and with bi-amplitude signals. We conclude that previous estimates of the minimum intensity level based on temporal integration data underestimates the intensity levels that can contribute to threshold. Our results suggest that there are no apparent constraints on the intensity levels that can be integrated near threshold. The auditory system integrates distributed stimulus intensity in both the time and amplitude dimensions. Temporal integration in the auditory system can be viewed as a signal process, where an enhanced internal representation is given low-level stimuli.

Acoustic Stimulation↗

A framework for medical visual information exchange on the web.

The web has become such an extensive health information repository in the world that it is increasingly difficult to search for relevant medical information. Most medical information available on the web is not peer reviewed, and is retrieved imprecisely by current web search mechanisms (i.e. based on keywords). This paper presents the MedISeek metadata model that allows one to describe medical visual information (i.e. medical images) of different modalities, including their properties, components, relationships and authorship. The model uses the web architecture and supports the international classification of diseases and related health problems (i.e. ICD-10). An RDF schema (Resource Description Framework (RDF), http://www.w3.org/RDF/.) derived from this metadata model is integrated to each medical image, and specifies the semantics of each property in the image. Thus, relevant information can be extracted directly from the images, and data integrity is better preserved in the web. A prototype, presented here, has been built to validate the metadata model, and the mechanism for medical visual information exchange on the web. Our preliminary experimental results indicate that authorized users of our system have been able to describe, store and retrieve medical images and their associated diagnostic information.

Computer Systems↗

Metabolomics technology and bioinformatics.

Metabolomics is the global analysis of all or a large number of cellular metabolites. Like other functional genomics research, metabolomics generates large amounts of data. Handling, processing and analysis of this data is a clear challenge and requires specialized mathematical, statistical and bioinformatics tools. Metabolomics needs for bioinformatics span through data and information management, raw analytical data processing, metabolomics standards and ontology, statistical analysis and data mining, data integration and mathematical modelling of metabolic networks within a framework of systems biology. The major approaches in metabolomics, along with the modern analytical tools used for data generation, are reviewed in the context of these specific bioinformatics needs.

Animals↗

Surface reflections of cardiac excitation and the assessment of infarct volume in dogs. A comparison of methods.

Ventricular depolarization was analyzed in intact dogs by simultaneously recording body surface potential maps, McFee axial vectorcardiograms, and a 5 X 4 lead precordial grid of QRS complexes. The purpose of this study was to compare the effectiveness of subtraction approaches, using the simultaneously acquired data. The totally closed chest approach avoided the problem of volume conductor alteration by thoracotomy. Infarct volume was calculated morphologically from measurements of serial ventricular sections. The maximal correlation with anatomic infarct size using the precordial QRS grid approach was 0.51, using cumulative difference data between 1 and 38 msec when the postinfarction grid was substracted from the preinfarction grid. A correlation coefficient of 0.80 was achieved using the numerically integrated data between 1 and 31 msec from the vectorcardiogram, and the body surface potential map achieved a correlation coefficient above 0.88 when the electrical difference of msec 16 was used. These data suggest that estimates of infarct size from selected surface reflections of the activation process are feasible if some sort of preinfarction control data are available. Caution must be exercised to avoid inclusion of electrical effects late in the activation process which contain contamination by highly variable alterations in the excitation sequence due to delayed conduction or alteration in conduction pathway in or near the infarct zone.

Action Potentials↗

Quality improvement for stroke management at the Cleveland Clinic Health System.

BACKGROUND: The Cleveland Clinic Health System established a stroke quality improvement (QI) initiative across its nine hospitals. IMPLEMENTING THE STROKE QI INITIATIVE: A stroke QI team took a three-pronged approach to QI: professional education, public education, and hospital process improvements. Its activities and subsequent data analysis needs were divided into four cycles (1999-2003). All data were provided to the stroke QI team and then to the Medical Operations Council to review results, consider data integrity issues, and plan dissemination. The dissemination of performance results permitted broad organizational responses to facilitate improvement. Improvement activities included professional education, public awareness, process improvement, focused data collection with routine feedback, protocol refinement, and coordination of clinical personnel within and between hospitals. RESULTS: The frequency of brain hemorrhagic complications decreased by more than half, from 13.4% to 6.4%; the rate of intravenous tissue plasminogen activator use increased from 1.5% to 3.9% of all stroke patients; and protocol deviations were reduced from 33% to 17%. DISCUSSION: The keys to this initiative's success were the health system's leadership's support, physicians' engagement via multidisciplinary project committees at the health system and hospital levels, and flexibility in implementing locally tailored process interventions.

Humans↗

Privacy-Preserving Linkage of Distributed Biological, Clinical, and Imaging Data Supporting Artificial Intelligence in Pediatric Oncology.

BACKGROUND: Cancer remains the leading cause of disease-related mortality in children over the age of one in Europe, with over 35,000 new pediatric cases and more than 6,000 deaths annually. Due to the rarity of pediatric cancers, clinical trial protocols often substitute for formal treatment guidelines, resulting in many children being enrolled in multiple trials, with biological samples and genomic data stored in various biobanks. Data collection in pediatric oncology is challenging, with sparse data acquired over extended periods, underscoring the need for optimal utilization of all available information through linked, privacy-preserving datasets. METHODS: Here, we report the development of a distributed, privacy-preserving data infrastructure for the PRIMAGE project, a European initiative aimed at supporting artificial intelligence (AI)-driven image analysis for pediatric cancer prognostics. The infrastructure leverages the European Patient Identity (EUPID) Services for Privacy-Preserving Record Linkage, enabling pseudonymized data integration across clinical, biological, and imaging sources. The system incorporates EUPID's hashing and phonetic matching protocols to pseudonymize patient identifiers and link distributed datasets, facilitating secondary data use in compliance with the General Data Protection Regulation. RESULTS: Data from over 700 neuroblastoma patients from European trials and hospitals were linked and uploaded to the PRIMAGE platform, where AI models predict clinical outcomes. CONCLUSION: This infrastructure successfully facilitated AI model development, advancing pediatric oncology research, and offering a scalable framework for future European health data initiatives, such as the European Health Data Space.

Journal Article↗

Bridging genotype, phenotype, and clinical insight: the role of multi-omics in cardiovascular disease.

INTRODUCTION: It is increasingly evident that the multifactorial nature of cardiovascular disease requires the combination of different omics approaches for improving our mechanistic understanding, identifying novel drug targets, and developing accurate diagnostic, predictive, and prognostic biomarker panels. AREAS COVERED: We review the current state and the potential of multi-omics in cardiovascular disease, with a specific focus on plasma-, spatial-, and single-cell approaches. We discuss lipidomics as a genotype‑to‑phenotype bridge, the utility of remote longitudinal monitoring via microsampling/dried blood spots, and emerging clinical‑trial integrations of multi-omics approaches. We outline critical gaps in standardization and how to overcome these, pre‑analytical challenges and constraints that are often neglected, and data‑integration methods spanning from canonical correlation analysis to modern machine learning approaches. EXPERT OPINION: Multi‑omics can shape cardiovascular care by identifying drug targets in diseased tissue and by yielding small, usable biomarker panels.

Humans↗

Crafting new methods of systems integration.

A handful of pioneering organizations are using intranets in their systems integration efforts. They have concluded that intranets, when used in combination with interface engines, can "virtually integrate" data.

Computer Communication Networks↗

Estimation of systemic toxicity of acrylamide by integration of in vitro toxicity data with kinetic simulations.

Neurodegenerative properties of acrylamide were studied in vitro by exposure of differentiated SH-SY5Y human neuroblastoma cells for 72 h. The number of neurites per cell and the total cellular protein content were determined every 24 h throughout the exposure and the subsequent 96-h recovery period. Using kinetic data on the metabolism of acrylamide in rat, a biokinetic model was constructed in which the in vitro toxicity data were integrated. Using this model, we estimated the acute and subchronic toxicity of acrylamide for the rat in vivo. These estimations were compared to experimentally derived lowest observed effect doses (LOEDs) for daily intraperitoneal exposure (1, 10, 30, and 90 days) to acrylamide. The estimated LOEDs differed maximally twofold from the experimental LOEDs, and the nonlinear response to acrylamide exposure over time was simulated correctly. It is concluded that the integration of the present in vitro toxicity data with kinetic data gives adequate estimates of acute and subchronic neurotoxicity resulting from acrylamide exposure.

Acrylamide↗

Physician decision-making--evaluation of data used in a computerized ICU.

New instrumentation, techniques and computers have made such large amounts of information rapidly available to ICU clinicians that there is now a danger of information overload. To help with this problem at LDS Hospital, a computerized system was implemented in the Shock-Trauma ICU. This ICU is almost totally computerized with each patient's physiologic, laboratory, drug, demographic, fluid input/output and nutritional data integrated into the patient's computer record. In the ICU, physician decision-making takes place in two situations: during rounds and on-site. For this study, data usage in decision-making was evaluated in both of these environments. The items of data used in decision-making were tabulated into six categories: bedside monitor, laboratory, drugs, input/output and IV, blood gas laboratory, observations and other. Comparisons were made between the portion of the computerized database occupied by a category and its use in decision-making. Combined laboratory data (clinical, microbiology and blood gas) made up 38 to 41% of total patient data reviewed and occupied 16.3% of the database. Observations made up 21-22% of the data reviewed and occupied 6.8% of the database. Drugs, input/output and IV data usage ranged from 13% to 23%, but occupied 36% of the database. Bedside monitor data usage was 12.5% to 22% and occupied 32.5% of the database. The 'other' category, used 2.5% to 5% of the time, made up 8.4% of the database. These results indicate that patient data collection and storage must be evaluated and optimized. This evaluation, along with implementation of the computerized ICU Rounds Report developed for optimal data presentation, will help physicians to evaluate patient status and should facilitate effective decisions.

Diagnosis, Computer-Assisted↗

WebGestalt: an integrated system for exploring gene sets in various biological contexts.

High-throughput technologies have led to the rapid generation of large-scale datasets about genes and gene products. These technologies have also shifted our research focus from 'single genes' to 'gene sets'. We have developed a web-based integrated data mining system, WebGestalt (http://genereg.ornl.gov/webgestalt/), to help biologists in exploring large sets of genes. WebGestalt is composed of four modules: gene set management, information retrieval, organization/visualization, and statistics. The management module uploads, saves, retrieves and deletes gene sets, as well as performs Boolean operations to generate the unions, intersections or differences between different gene sets. The information retrieval module currently retrieves information for up to 20 attributes for all genes in a gene set. The organization/visualization module organizes and visualizes gene sets in various biological contexts, including Gene Ontology, tissue expression pattern, chromosome distribution, metabolic and signaling pathways, protein domain information and publications. The statistics module recommends and performs statistical tests to suggest biological areas that are important to a gene set and warrant further investigation. In order to demonstrate the use of WebGestalt, we have generated 48 gene sets with genes over-represented in various human tissue types. Exploration of all the 48 gene sets using WebGestalt is available for the public at http://genereg.ornl.gov/webgestalt/wg_enrich.php.

Computer Graphics↗

Auditing drug metabolism protocols, data, and reports.

A critical component of the drug discovery process is to assess the safety and metabolic disposition of a drug. This can be accomplished by evaluating drug concentration levels as part of safety assessment studies (toxicokinetics) and identifying the drug's presence in and throughout the body by conducting ADME (adsorption, distribution, metabolism, and excretion) studies. From a regulatory perspective, the major difference between these two types of evaluations is that toxicokinetic studies are under the purview of the Good Laboratory Practice (GLP) regulations (21 CFR 58) and ADME studies are not. International debate continues to resolve around the need to consider the applicability of the GLPs to ADME studies. While it is recognized that the current regulatory intent between ADME and toxicokinetic studies differs, this inspecting/auditing approach treats them generally the same. It assures management and external reviewers that data integrity measures are in place for all drug metabolism studies conducted by the facility.

Animals↗

High-density rat radiation hybrid maps containing over 24,000 SSLPs, genes, and ESTs provide a direct link to the rat genome sequence.

The laboratory rat is a major model organism for systems biology. To complement the cornucopia of physiological and pharmacological data generated in the rat, a large genomic toolset has been developed, culminating in the release of the rat draft genome sequence. The rat draft sequence used a variety of assembly packages, as well as data from the Radiation Hybrid (RH) map of the rat as part of their validation. As part of the Rat Genome Project, we have been building a high-density RH map to facilitate data integration from multiple maps and now to help validate the genome assembly. By incorporating vectors from our lab and several other labs, we have doubled the number of simple sequence length polymorphisms (SSLPs), genes, expressed sequence tags (ESTs), and sequence-tagged sites (STSs) compared to any other genome-wide rat map, a total of 24,437 elements. During the process, we also identified a novel approach for integrating the RH placement results from multiple maps. This new integrated RH map contains approximately 10 RH-mapped elements per Mb on the genome assembly, enabling the RH maps to serve as a scaffold for a variety of data visualization tools.

Animals↗

Nutritional surveillance.

The concept of nutritional surveillance is derived from disease surveillance, and means "to watch over nutrition, in order to make decisions that lead to improvements in nutrition in populations". Three distinct objectives have been defined for surveillance systems, primarily in relation to problems of malnutrition in developing countries: to aid long-term planning in health and development; to provide input for programme management and evaluation; and to give timely warning of the need for intervention to prevent critical deteriorations in food consumption. Decisions affecting nutrition are made at various administrative levels, and the uses of different types of nutritional surveillance information can be related to national policies, development programmes, public health and nutrition programmes, and timely warning and intervention programmes. The information should answer specific questions, for example concerning the nutritional status and trends of particular population groups.Defining the uses and users of the information is the first essential step in designing a system; this is illustrated with reference to agricultural and rural development planning, the health sector, and nutrition and social welfare programmes. The most usual data outputs are nutritional outcome indicators (e.g., prevalence of malnutrition among preschool children), disaggregated by descriptive or classifying variables, of which the commonest is simply administrative area. Often, additional "status" indicators, such as quality of housing or water supply, are presented at the same time. On the other hand, timely warning requires earlier indicators of the possibility of nutritional deterioration, and agricultural indicators are often the most appropriate.DATA COME FROM TWO MAIN TYPES OF SOURCE: administrative (e.g., clinics and schools) and household sample surveys. Each source has its own advantages and disadvantages: for example, administrative data often already exist, and can be disaggregated to village level, but are of unknown representativeness and often cannot be linked with other variables of interest; sample surveys provide integrated data of more or less known representativeness, but sample sizes usually do not allow disaggregation to, for example, specific villages. A combination of these sources, with a capability for ad hoc surveys (formal or informal) is often the best solution. Finally, much depends on adequate facilities for data analysis, even though simple, comprehensible data outputs are what is required. Intersectoral cooperation is needed to provide realistic options for the decision-making process.

Agriculture↗

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans↗

Development and implementation of a multi-centre information system for paediatric and infant critical care.

BACKGROUND: With no UK collective information system, a need existed to establish an integrated information system for public and private sector hospitals providing paediatric and infant critical care services. A lack of information in the past made it difficult for those procuring, providing and monitoring services to make informed, evidence-based decisions using reliable integrated data. OBJECTIVES: To develop and implement a collective multi-purpose information system for paediatric and infant critical care that was easily adaptable to any UK infant or paediatric critical care setting. Information outputs had to fulfil policy requirements and meet the needs of stakeholders. METHOD: Two minimum datasets, corresponding data definitions, survey forms and a user database were developed through a process of consultation by utilising an information partnership. Design, content, development and implementation issues were identified, discussed and resolved through a co-ordinated collaborative process. RESULTS: Data collection was implemented in all London and Brighton National Health Service (NHS) general and cardio-thoracic paediatric intensive care (PIC) units, several private PIC units and one NHS tertiary referral neonatal unit (NNU) 24 months from project start. CONCLUSIONS: The development of universal integrated information systems for defined settings of care is achievable within reasonable timeframes; however, successful development and implementation requires working within an information partnership to maximise co-ordination, co-operation and collaboration. Those collecting and using data must be identified and involved in all aspects of development from project start. Financial and manpower resources must be well planned. Datasets should be as small as possible in order to make the collection of complete and valid data realistically achievable. When considering service-based information needs, considerable thought should be given to a multi-purpose; multi-use approach based on the most refined minimum dataset possible.

Child↗

Evidence-based treatment of stuttering: III. Evidence-based practice in a clinical setting.

UNLABELLED: At the heart of evidence-based practice in stuttering treatment are four issues: (1) the collection of data to inform treatment; (2) the long standing concern with maintenance of treatment gains; (3) the need to demonstrate accountability to clients, payers and our profession as service providers; and (4) the desire to advance theoretical knowledge. This article addresses the first three of these issues from a practical point of view, illustrating how data collection for stuttering treatment outcome research in a clinical setting is intimately blended with that required for clinical purposes and providing an example of a process of evaluating data for clinical and research purposes. EDUCATIONAL OBJECTIVES: The reader will learn about and be able to (1) differentiate between treatment outcome and treatment efficacy research, (2) describe models for integrating data collection for treatment outcome and clinical purposes, and (3) utilize guidelines for treatment efficacy that are applicable to outcome research to evaluate data for use in treatment outcome studies and to design outcome studies.

Data Collection↗