Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

The BioMediator system as a data integration tool to answer diverse biologic queries.

We present the BioMediator (www.biomediator.org) system and the process of executing queries on it. The system was designed as a tool for posing queries across semantically and syntactically heterogeneous data particularly in the biological arena. We use examples from researchers at the University of Washington, and the University of Missouri-Columbia, to discuss the BioMediator system architecture, query execution, modifications to the system to support the queries, and summarize our findings and our future directions. Finally, we discuss the system's flexibility and generalized approach and give examples of how the system can be extended for a variety of objectives.

Computational Biology↗

Data integrity in a general practice computer system (CLINICS).

The accuracy of computer held medical information may be of critical importance in patient care, therefore it is important not only to know the error rate in the stored data but also to know the effectiveness of error checking and detection programmes. This paper reports on the errors which were detected in the University of Southampton Primary Medical Care computer system (CLINICS) by checking the consistency between stored data and incoming data. Seven per cent of incoming data had important errors of kinds not normally detected by many medical record systems. The majority were traced either to the registration of new patients or to the doctors failing to pay adequate attention to detail in their record keeping (or to their legibility). They have been subsequently corrected, and it is calculated that the stored data contains less than 1% errors. We suggest ways of improving this; and conclude that certain items are essential to general practice information systems.

Computers↗

VisANT: data-integrating visual framework for biological networks and modules.

VisANT is a web-based software framework for visualizing and analyzing many types of networks of biological interactions and associations. Networks are a useful computational tool for representing many types of biological data, such as biomolecular interactions, cellular pathways and functional modules. Given user-defined sets of interactions or groupings between genes or proteins, VisANT provides: (i) a visual interface for combining and annotating network data, (ii) supporting function and annotation data for different genomes from the Gene Ontology and KEGG databases and (iii) the statistical and analytical tools needed for extracting topological properties of the user-defined networks. Users can customize, modify, save and share network views with other users, and import basic network data representations from their own data sources, and from standard exchange formats such as PSI-MI and BioPAX. The software framework we employ also supports the development of more sophisticated visualization and analysis functions through its open API for Java-based plug-ins. VisANT is distributed freely via the web at http://visant.bu.edu and can also be downloaded for individual use.

Computer Graphics↗

[Integrated data processing systems for a functional diagnostic service].

The paper discusses the problem of automation of the functional diagnostic service system on the basis of up-to-date informational technologies, gives recommendations on the minimum list of diagnostic procedures used for various service levels in medical institutions, outlines requirements for hard- and softwares needed for designing this system of integrated processing, provides specific recommendations on the structure of automatic complexes for the first two levels of functional diagnostic service.

Diagnosis, Computer-Assisted↗

[Viewpoints on preparations for integrating data processing systems in routine microbiology diagnosis].

The enormously risen and further increasing numbers of examinations and tests in microbiological diagnostics within the last years need new methods for treatment. One possibility to meet the higher requirements for information of the clinic without loss in quality at constant staff is the integration of the microcomputer technique into the laboratory as direct "tool". Demands for a qualitatively high empirical antimicrobial chemotherapy, chemotherapy according to antibiotic susceptibility tests, indicated use of antimicrobial drugs and control measures of infectious processes in general are met only by means of a fast information processing. The microcomputer technique in the laboratory provides also the chance to automate still manually performed tests and comprises according to algorithm the strict observation of the diagnostic process and its control. The application of the microcomputer technique on the one hand means for the technical assistant the omission of much manually performed work, on the other hand enables work of higher quality and supports decisions in the diagnostic process. Mathematical and statistical calculations are no longer connected with great losses of activity. The actual need for information of the clinician is met in time in different ways.

Bacteriological Techniques↗

Network security and data integrity in academia: an assessment and a proposal for large-scale archiving.

A direct impediment to the optimal use of online databases is the increasing prevalence, severity, and toll of computer and network security incidents. Funding agencies should set up working groups that can provide essential services such as universal backup, archival storage, and mirroring of community resources, consistent with the key goal of security in academia: to preserve data and results for posterity.

Archives↗

Experimental-neuromodeling framework for understanding auditory object processing: integrating data across multiple scales.

In this article, we review a combined experimental-neuromodeling framework for understanding brain function with a specific application to auditory object processing. Within this framework, a model is constructed using the best available experimental data and is used to make predictions. The predictions are verified by conducting specific or directed experiments and the resulting data are matched with the simulated data. The model is refined or tested on new data and generates new predictions. The predictions in turn lead to better-focused experiments. The auditory object processing model was constructed using available neurophysiological and neuroanatomical data from mammalian studies of auditory object processing in the cortex. Auditory objects are brief sounds such as syllables, words, melodic fragments, etc. The model can simultaneously simulate neuronal activity at a columnar level and neuroimaging activity at a systems level while processing frequency-modulated tones in a delayed-match-to-sample task. The simulated neuroimaging activity was quantitatively matched with neuroimaging data obtained from experiments; both the simulations and the experiments used similar tasks, sounds, and other experimental parameters. We then used the model to investigate the neural bases of the auditory continuity illusion, a type of perceptual grouping phenomenon, without changing any of its parameters. Perceptual grouping enables the auditory system to integrate brief, disparate sounds into cohesive perceptual units. The neural mechanisms underlying auditory continuity illusion have not been studied extensively with conventional neuroimaging or electrophysiological techniques. Our modeling results agree with behavioral studies in humans and an electrophysiological study in cats. The results predict a particular set of bottom-up cortical processing mechanisms that implement perceptual grouping, and also attest to the robustness of our model.

Acoustic Stimulation↗

Detection threshold microstructure and its effect on temporal integration data.

Auditory detection thresholds were measured at several preselected frequencies using an adaptive 2IFC procedure. Results confirm the existence of shifts in detection threshold ranging from 2 to 14 dB with quite small changes in signal frequency. There does not appear to be a uniform pattern associated with the microstructure of the detection threshold curve. An additional experiment was performed to determine the effect of signal duration on detection threshold microstructure. Results indicate that the temporal integration function is considerably steeper for more sensitive frequencies (3.7 dB/doubling of duration), than for less sensitive frequencies (1.7 dB/doubling). This probably is not due to differences in processing as much as it is to the effect of the energy spread associated with decreasing signal duration.

Acoustic Stimulation↗

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software↗

SeqUIaSCOPE: multi-omics data integration platform for single-patient clinical oncology pathway exploration.

SUMMARY: SeqUIaSCOPE is an open-source platform designed for routine clinical oncology diagnostics through case-centric integration and visualization of genomic variants, fusion events, and expression profiles. The platform combines molecular-level validation via embedded genome browsing with systems-level interpretation through dynamic pathway visualization, enabling geneticists to assess how alterations converge across biological networks. Flexible reporting with customizable templates accommodates diverse institutional requirements, while secure cluster-based or local deployment ensures compliance with data protection policies, making advanced multi-omics diagnostics accessible to academic and clinical institutions. AVAILABILITY AND IMPLEMENTATION: SeqUIaSCOPE is freely available on GitHub at https://github.com/BioIT-CEITEC/sequiascope under the MIT license and archived at Zenodo (https://zenodo.org/records/21338445). Due to the sensitive nature of patient data, the repository provides simulated datasets that mimic the structure of real clinical data for testing and exploration. Documentation and a live demo accompany these datasets, allowing users to explore the application without any prior setup. The repository also includes a Helm chart for Kubernetes deployment and Docker containers for local deployment, ensuring compatibility across Linux, macOS, and Windows. No user registration is required, and all data remains on local or institutional infrastructure.

Humans↗

Challenges of target/compound data integration from disease to chemistry: a case study of dihydrofolate reductase inhibitors.

Despite the improvements in informatics associated with initiatives in the structure-based design and genomics fields, no straight-forward links are available between a given disease class and drug chemistry. This involves effective linking of disease to protein targets, and then mapping these targets to drug chemistry. In practice, protein-ligand structural analyses and high-throughput screening experiments generate the links between targets implicated in disease and chemical leads. Additionally, large volumes of relevant data are also being produced by high-throughput X-ray crystallography and in-silico docking initiatives. Each of these efforts takes a distinctly different approach to how data is managed and mined, resulting in difficulties in sharing data across each area. This review discusses the diverse approaches taken to data management in these areas, and the challenges associated with the construction of a data warehouse that meets all of the needs of each data type. Using the current work available for dihydrofolate reductase inhibitors, we demonstrate the challenges and opportunities associated with data mining from disease to drug chemistry.

Animals↗

Functional Genomics meets neurodegenerative disorders. Part II: application and data integration.

The transcriptomic and proteomic techniques presented in part I (Functional Genomics meets neurodegenerative disorders. Part I: transcriptomic and proteomic technology) of this back-to-back review have been applied to a range of neurodegenerative disorders, including Huntington's disease (HD), Prion diseases (PrD), Creutzfeldt-Jakob disease, amyotrophic lateral sclerosis (ALS), Alzheimer's disease (AD), frontotemporal dementia (FTD) and Parkinson's disease (PD). Samples have been derived either from human brain and cerebrospinal fluid, tissue culture cells or brains and spinal cord of experimental animal models. With the availability of huge data sets it will firstly be a major challenge to extract meaningful information and secondly, not to obtain contradicting results when data are collected in parallel from the same source of biological specimen using different techniques. Reliability of the data highly depends on proper normalization and validation both of which are discussed together with an outlook on developments that can be anticipated in the future and are expected to fuel the field. The new insight undoubtedly will lead to a redefinition and subdivision of disease entities based on biochemical criteria rather than the clinical presentation. This will have important implications for treatment strategies.

Animals↗

E-neuroscience: challenges and triumphs in integrating distributed data from molecules to brains.

Imaging, from magnetic resonance imaging (MRI) to localization of specific macromolecules by microscopies, has been one of the driving forces behind neuroinformatics efforts of the past decade. Many web-accessible resources have been created, ranging from simple data collections to highly structured databases. Although many challenges remain in adapting neuroscience to the new electronic forum envisioned by neuroinformatics proponents, these efforts have succeeded in formalizing the requirements for effective data sharing and data integration across multiple sources. In this perspective, we discuss the importance of spatial systems and ontologies for proper modeling of neuroscience data and their use in a large-scale data integration effort, the Biomedical Informatics Research Network (BIRN).

Animals↗

Reporting, appraising, and integrating data on genotype prevalence and gene-disease associations.

The recent completion of the first draft of the human genome sequence and advances in technologies for genomic analysis are generating tremendous opportunities for epidemiologic studies to evaluate the role of genetic variants in human disease. Many methodological issues apply to the investigation of variation in the frequency of allelic variants of human genes, of the possibility that these influence disease risk, and of assessment of the magnitude of the associated risk. Based on a Human Genome Epidemiology workshop, a checklist for reporting and appraising studies of genotype prevalence and studies of gene-disease associations was developed. This focuses on selection of study subjects, analytic validity of genotyping, population stratification, and statistical issues. Use of the checklist should facilitate the integration of evidence from these studies. The relation between the checklist and grading schemes that have been proposed for the evaluation of observational studies is discussed. Although the limitations of grading schemes are recognized, a robust approach is proposed. Other issues in the synthesis of evidence that are particularly relevant to studies of genotype prevalence and gene-disease association are discussed, notably identification of studies, publication bias, criteria for causal inference, and the appropriateness of quantitative synthesis.

Case-Control Studies↗

Prioritization of causal genes from genome-wide association studies by Bayesian data integration across loci.

MOTIVATION: Genome-wide association studies (GWAS) have identified genetic variants, usually single-nucleotide polymorphisms (SNPs), associated with human traits, including disease and disease risk. These variants (or causal variants in linkage disequilibrium with them) usually affect the regulation or function of a nearby gene. A GWAS locus can span many genes, however, and prioritizing which gene or genes in a locus are most likely to be causal remains a challenge. Better prioritization and prediction of causal genes could reveal disease mechanisms and suggest interventions. RESULTS: We describe a new Bayesian method, termed SigNet for significance networks, that combines information both within and across loci to identify the most likely causal gene at each locus. The SigNet method builds on existing methods that focus on individual loci with evidence from gene distance and expression quantitative trait loci (eQTL) by sharing information across loci using protein-protein and gene regulatory interaction network data. In an application to cardiac electrophysiology with 226 GWAS loci, only 46 (20%) have within-locus evidence from Mendelian genes, protein-coding changes, or colocalization with eQTL signals. At the remaining 180 loci lacking functional information, SigNet selects 56 genes other than the minimum distance gene, equal to 31% of the information-poor loci and 25% of the GWAS loci overall. Assessment by pathway enrichment demonstrates improved performance by SigNet. Review of individual loci shows literature evidence for genes selected by SigNet, including PMP22 as a novel causal gene candidate.

Genome-Wide Association Study↗

An integrated data analysis approach to characterize genes highly expressed in hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is one of the major causes of cancer deaths worldwide. New diagnostic and therapeutic options are needed for more effective and early detection and treatment of this malignancy. We identified 703 genes that are highly expressed in HCC using DNA microarrays, and further characterized them in order to uncover novel tumor markers, oncogenes, and therapeutic targets for HCC. Using Gene Ontology annotations, genes with functions related to cell proliferation and cell cycle, chromatin, repair, and transcription were found to be significantly enriched in this list of highly expressed genes. We also identified a set of genes that encode secreted (e.g. GPC3, LCN2, and DKK1) or membrane-bound proteins (e.g. GPC3, IGSF1, and PSK-1), which may be attractive candidates for the diagnosis of HCC. A significant enrichment of genes highly expressed in HCC was found on chromosomes 1q, 6p, 8q, and 20q, and we also identified chromosomal clusters of genes highly expressed in HCC. The microarray analyses were validated by RT-PCR and PCR. This approach of integrating other biological information with gene expression in the analysis helps select aberrantly expressed genes in HCC that may be further studied for their diagnostic or therapeutic utility.

Biomarkers, Tumor↗

Discovery of functional genes for systemic acquired resistance in Arabidopsis thaliana through integrated data mining.

Various data mining techniques combined with sequence motif information in the promoter region of genes were applied to discover functional genes that are involved in the defense mechanism of systemic acquired resistance (SAR) in Arabidopsis thaliana. A series of K-Means clustering with difference-in-shape as distance measure was initially applied. A stability measure was used to validate this clustering process. A decision tree algorithm with the discover-and-mask technique was used to identify a group of most informative genes. Appearance and abundance of various transcription factor binding sites in the promoter region of the genes were studied. Through the combination of these techniques, we were able to identify 24 candidate genes involved in the SAR defense mechanism. The candidate genes fell into 2 highly resolved categories, each category showing significantly unique profiles of regulatory elements in their promoter regions. This study demonstrates the strength of such integration methods and suggests a broader application of this approach.

Algorithms↗

Integrating data for sustainable development: introducing the distribution of resources framework.

A complex, information-rich world requires frameworks that organize data to reveal succinct views and interrelationships. The goal of sustainable development requires a particularly large range of data covering many disciplines--economic, environmental, social and institutional. This paper introduces the Distribution of Resources (DOR) framework. DOR explicitly quantifies the changes in our natural and produced resources and traces the use of resources through production and consumption efficiency to distribution as welfare. The DOR framework accommodates many of the recognized concepts and principles of sustainable development resulting in a comprehensive policy-focused framework with a more extensive range of desirable features than other frameworks developed to date.

Data Interpretation, Statistical↗