Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The network approach: a means for the collection of integrated data following standardized protocols.

The understanding of the epidemiology of a vector borne disease, involving various vector and host species in a defined area requires a multidisciplinary approach. It is essential that specialists obtain data relevant to common objectives. In the case of African Trypanosomiasis this means that the observations being made on the health status of the human and domestic--and wild animal populations are made over the same period of time as those on the tsetse population to which these hosts are exposed. Standardization of methodology is a pre-condition for reliable comparison of observations. The creation of a network of research situations is one possibility for the fulfillment of this pre-condition; while at the same time it is a suitable means for the collection of integrated data. The African Trypanotolerant Livestock Network, created in 1982 is one example of such a network. Examples of conclusions which could be drawn after analysis of data collected over a two-year period through this Network are presented.

Animals↗

Mutational data integration in gene-oriented files of the Hermansky-Pudlak Syndrome database.

Hermansky-Pudlak Syndrome (HPS) is a genetically heterogeneous disorder characterized by oculocutaneous albinism and prolonged bleeding due to abnormal vesicle trafficking to lysosomes and related organelles such as melanosomes and platelet dense granules. This HPS database (HPSD; http://liweilab.genetics.ac.cn/HPSD/) provides integrated, annotatory, and curative data that is distributed in a variety of public databases or predicted by bioinformatics servers for the recently cloned human and mouse HPS genes, as well as for the genes responsible for HPSrelated syndromes, such as ChediakHigashi Syndrome (CHS), Griscelli syndrome (GS), oculocutaneous albinism (OCA), Usher syndrome type 1B (USH1B), and ocular albinism (OA). The HPSD is designed by using a unique GeneOriented File (GOF) format. Seven blocks (genomic, transcript, protein, function, mutation, phenotype, and reference) are carefully annotated in each userfriendly GOF entry. The HPSD emphasizes paired human and mouse GOF entries. The genes included in this database (currently 58 in total) are arbitrarily divided into four categories: 1) Human and Mouse HPS, 2) Mouse HPS Only, 3) Putative Mouse or Human HPS, and 4) HPS Related Syndromes. All the mutations in these genes are integrated in the GOFs. We expect that these very informative and peerreviewed GOFs will be shortcuts to utilize the webbased information for the emerging interdisciplinary studies of HPS.

Animals↗

Inferring gene expression networks via static and dynamic data integration.

This paper presents a novel approach for the extraction of gene regulatory networks from DNA microarray data. The approach is characterized by the integration of data coming from static and dynamic experiments, exploiting also prior knowledge on the biological process under analysis. A starting network topology is built by analyzing gene expression data measured during knockout experiments. The analysis of time series expression profiles allows to derive the complete network structure and to learn a model of the gene expression dynamics: to this aim a genetic algorithm search coupled with a regression model of the gene interactions is exploited. The method has been applied to the reconstruction of a network of genes involved into the Saccharomyces Cerevisiae cell cycle. The proposed approach was able to reconstruct known relationships among genes and to provide meaningful biological results.

Artificial Intelligence↗

Global inequities in hepatitis B and C genomic surveillance revealed through an interactive data integration dashboard.

OBJECTIVES: To assess global disparities in hepatitis B virus (HBV) and hepatitis C virus (HCV) genomic surveillance and to develop an integrated platform that links genomic data with epidemiological burden. STUDY DESIGN: Retrospective observational analysis. METHODS: We reviewed existing viral genomic repositories to identify structural and analytical limitations. Subsequently, we integrated 10 996 HBV and 3533 HCV whole-genome sequences (WGS) from public databases with Global Burden of Disease (GBD) estimates to quantify inequities in genomic surveillance across countries and genotypes. Using these data, we developed the open-access Hepatitis Dashboard, incorporating >14 000 sequences from 141 countries with GBD metrics to evaluate representativeness and sequencing coverage relative to disease burden. RESULTS: Marked inequities in hepatitis genomic surveillance were identified. Despite increasing HBV- and HCV-associated mortality, virus sequence availability remains geographically and genotypically skewed-dominated by China and the United States, with substantial underrepresentation of HBV genotype E and HCV genotypes 5 and 8. Many high-endemic countries in Africa and the Western Pacific remain severely undersampled. We detected circulating antiviral drug-resistance mutations and developed a burden-adjusted sequencing coverage metric, revealing that several high-burden countries, including China, Nigeria and India, are among the least represented in global genomic datasets. Projections to 2030 indicate that neither HBV nor HCV are currently on track to meet WHO elimination targets. CONCLUSIONS: The Hepatitis Dashboard provides an integrated, continuously updated resource that links genomic and epidemiological data to quantify and visualise global surveillance gaps. This analysis highlights a critical disconnect between sequencing efforts and public health needs, which may limit the effectiveness of surveillance-informed strategies to support progress toward WHO 2030 elimination goals. By enabling burden-adjusted prioritisation and longitudinal tracking of genomic coverage, the platform supports evidence-based sampling strategies, equitable resource allocation, and monitoring of global progress toward hepatitis elimination.

Humans↗

Putting data integration into practice: using biomedical terminologies to add structure to existing data sources.

A major purpose of biomedical terminologies is to provide uniform concept representation, allowing for improved methods of analysis of biomedical information. While this goal is being realized in bioinformatics, with the emergence of the Gene Ontology as a standard, there is still no real standard for the representation of clinical concepts. As discoveries in biology and clinical medicine move from parallel to intersecting paths, standardized representation will become more important. A large portion of significant data, however, is mainly represented as free text, upon which conducting computer-based inferencing is nearly impossible. In order to test our hypothesis that existing biomedical terminologies, specifically the UMLS Metathesaurus and SNOMED CT, could be used as templates to implement semantic and logical relationships over free text data that is important both clinically and biologically, we chose to analyze OMIM (Online Mendelian Inheritance in Man). After finding OMIM entries' conceptual equivalents in each respective terminology, we extracted the semantic relationships that were present and evaluated a subset of them for semantic, logical, and biological legitimacy. Our study reveals the possibility of putting the knowledge present in biomedical terminologies to its intended use, with potentially clinically significant consequences.

Databases, Genetic↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the worldwide Protein Data Bank (wwPDB) and to work towards the integration of various bioinformatics data resources. One of the major obstacles to the improved integration of structural databases such as MSD and sequence databases like UniProt is the absence of up to date and well-maintained mapping between corresponding entries. We have worked closely with the UniProt group at the EBI to clean up the taxonomy and sequence cross-reference information in the MSD and UniProt databases. This information is vital for the reliable integration of the sequence family databases such as Pfam and Interpro with the structure-oriented databases of SCOP and CATH. This information has been made available to the eFamily group (http://www.efamily.org.uk/) and now forms the basis of the regular interchange of information between the member databases (MSD, UniProt, Pfam, Interpro, SCOP and CATH). This exchange of annotation information has enriched the structural information in the MSD database with annotation from wider sequence-oriented resources. This work was carried out under the 'Structure Integration with Function, Taxonomy and Sequences (SIFTS)' initiative (http://www.ebi.ac.uk/msd-srv/docs/sifts) in the MSD group.

Amino Acid Sequence↗

Research data integrity: a result of an integrated information system.

The toxicologic problems of today frequently require long-term, multidisciplinary experimentation involving large numbers of animals. In order to provide the extensive safety evaluation necessary to produce data that can be reasonably extrapolated to humans, automated research support systems have transcended the position of useful tools and have become an integral part of the total design of experimental protocols. For an automated information system to fully represent the reality of the experiment, it must be able to assure integrity, as well as provide for the storage, calculation, and retrieval of data values of the quality and quantity necessary for fulfilling protocol requirements. Guarantees against error and loss of data, in addition to flexibility and easy access, must be an inherent part of the system if the acceptance and condifence of the investigator are to be obtained. This paper discusses the criteria, philosophies, and benefits of integrated data systems that ensure integrity of toxicologic research support.

Computers↗

Data integration: blue skies ahead for network management.

Managed care organizations with an eye on the future realize that they must make the transition from cost management to quality management. This process relies heavily on the integrity and immediacy of the data that are collected, collated, and interpreted. Physicians and other medical providers will now play key roles in this process because, to date, MCOs have proved more adept at collecting reams of data than at employing data to effectively improve quality and reduce costs.

Computer Communication Networks↗

Describing the longitudinal course of major depression using Markov models: data integration across three national surveys.

BACKGROUND: Most epidemiological studies of major depression report period prevalence estimates. These are of limited utility in characterizing the longitudinal epidemiology of this condition. Markov models provide a methodological framework for increasing the utility of epidemiological data. Markov models relating incidence and recovery to major depression prevalence have been described in a series of prior papers. In this paper, the models are extended to describe the longitudinal course of the disorder. METHODS: Data from three national surveys conducted by the Canadian national statistical agency (Statistics Canada) were used in this analysis. These data were integrated using a Markov model. Incidence, recurrence and recovery were represented as weekly transition probabilities. Model parameters were calibrated to the survey estimates. RESULTS: The population was divided into three categories: low, moderate and high recurrence groups. The size of each category was approximated using lifetime data from a study using the WHO Mental Health Composite International Diagnostic Interview (WMH-CIDI). Consistent with previous work, transition probabilities reflecting recovery were high in the initial weeks of the episodes, and declined by a fixed proportion with each passing week. CONCLUSION: Markov models provide a framework for integrating psychiatric epidemiological data. Previous studies have illustrated the utility of Markov models for decomposing prevalence into its various determinants: incidence, recovery and mortality. This study extends the Markov approach by distinguishing several recurrence categories.

Journal Article↗

E-MSD: an integrated data resource for bioinformatics.

The Macromolecular Structure Database (MSD) group (http://www.ebi.ac.uk/msd/) continues to enhance the quality and consistency of macromolecular structure data in the Protein Data Bank (PDB) and to work towards the integration of various bioinformatics data resources. We have implemented a simple form-based interface that allows users to query the MSD directly. The MSD 'atlas pages' show all of the information in the MSD for a particular PDB entry. The group has designed new search interfaces aimed at specific areas of interest, such as the environment of ligands and the secondary structures of proteins. We have also implemented a novel search interface that begins to integrate separate MSD search services in a single graphical tool. We have worked closely with collaborators to build a new visualization tool that can present both structure and sequence data in a unified interface, and this data viewer is now used throughout the MSD services for the visualization and presentation of search results. Examples showcasing the functionality and power of these tools are available from tutorial webpages (http://www. ebi.ac.uk/msd-srv/docs/roadshow_tutorial/).

Algorithms↗

Integrating data to facilitate clinical research: a case study.

The integration of routine clinical administrative activities into ongoing rigorous clinical research poses challenges for both clinicians and researchers. This case study describes the development of a responsive database system used to facilitate comprehensive longitudinal research into the outcomes of patients waiting for hip and knee replacement surgery in a large public teaching hospital. The initial research procedure was paper-based, with manual patient matching and data entry. This process was time-consuming and associated with substantial risk of error and omissions, necessitating the design of a better system. An integrated database system was designed to receive daily electronic updates of the orthopaedic waiting-list and scheduled clinic and surgery dates. Using readily available software (Microsoft Access), new patients were identified through specifying inclusion and exclusion criteria which allowed rapid and complete recruitment at time of entry to the waiting-list. The integrated system specified the appropriate timing of multiple follow-up assessments, provided prompt information on recruitment for reporting purposes and integrated multiple linked research projects within one database. Seamless exporting of data to statistical programs for analysis was also enabled. This simple integrated approach facilitated efficient execution of a longitudinal study from recruitment to statistical analysis while maximising confidentiality and minimising resources required. This case study describes the development and design of a simple system which could be easily adapted for database management in hospital or clinic-based settings according to local requirements.

Biomedical Research↗

Interpretable data integration for single-cell and spatial multi-omics.

Integrating single-cell or spatial transcriptomic and epigenomic data enables scrutinizing the transcriptional regulatory mechanisms controlling cell fate. Current integration methods usually align multi-omics data into a shared latent space but fail to reveal the underlying connections between genes and regulatory elements. The correlation- or regression-based regulatory inference methods cannot dissect different transcriptional regulation codes for cells under different spatial and temporal states. To address both problems, we develop a feature-guided optimal transport (FGOT) method, which simultaneously uncovers cellular heterogeneity and their associated transcriptional regulatory links. FGOT also provides post hoc interpretability for existing integration methods. FGOT is applicable for paired/unpaired single-cell multi-omics data and paired spatial multi-omics data. Benchmarking and validating via histone modification data or three-dimensional (3D) genomics data show good robustness and accuracy in integration and inference of regulatory links. The method allows systematic screening of cell-state and spatial-location-specific regulatory elements in diseases at the single-cell level. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis↗

Improving recombinant protein productivity in CHO cells via multi-omics data integration.

Chinese hamster ovary (CHO) cells represent the dominant host system for the production of recombinant therapeutic proteins. In recent decades, extensive research has focused on process/media optimization and cell line engineering to improve both the productivity and quality of biopharmaceutical proteins produced in CHO cells. Nevertheless, the inherent complexity of biological pathways and the heterogeneous cellular responses to different environmental conditions have posed substantial challenges to traditional methodologies. Recent advances in omics technologies have enabled comprehensive characterization of CHO cell physiology, providing multidimensional molecular and phenotypic insights that facilitate the enhancement of recombinant protein production. This review first summarizes the methodologies and advances in CHO omics research, including genomics, transcriptomics, proteomics, metabolomics, and epigenomics. It then examines contemporary approaches to integrate and analyze multi-omics data in CHO cells. The review further elucidates how these multi-omics datasets can be strategically applied across various developmental stages, including cell line selection, genetic engineering, expression vector design, and bioprocess optimization. Finally, we explore the transformative potential of integrating multi-omics with artificial intelligence and discuss promising future research directions in CHO cell studies. These emerging paradigms offer novel opportunities for data-driven cell engineering and bioprocess optimization in CHO-based biomanufacturing.

Bioprocessing↗

Nomenclature-based data retrieval without prior annotation: facilitating biomedical data integration with fast doublet matching.

Assigning nomenclature codes to biomedical data is an arduous, expensive and error-prone task. Data records are coded to to provide a common representation of contained concepts, allowing facile retrieval of records via a standard terminology. In the medical field, cancer registrars, nurses, pathologists, and private clinicians all understand the importance of annotating medical records with vocabularies that codify the names of diseases, procedures, billing categories, etc. Molecular biologists need codified medical records so that they can discover or validate relationships between experimental data and clinical data. This paper introduces a new approach to retrieving data records without prior coding. The approach achieves the same result as a search over pre-coded records. It retrieves all records that contain any terms that are synonymous with a user's query-term. A recently described fast algorithm (the doublet method) permits quick iterative searches over every synonym for any term from any nomenclature occurring in a dataset of any size. As a demonstration, a 105+ Megabyte corpus of Pubmed abstracts was searched for medical terms. Query terms were matched against either of two vocabularies and expanded as an array of equivalent search items. A single search term may have over one hundred nomenclature synonyms, all of which were searched against the full database. Iterative searches of a list of concept-equivalent terms involves many more operations than a single search over pre-annotated concept codes. Nonetheless, the doublet method achieved fast query response times (0.05 seconds using Snomed and 5 seconds using the Developmental Lineage Classification of Neoplasms, on a computer with a 2.89 GHz processor). Pre-annotated datasets lose their value when the chosen vocabulary is replaced by a different vocabulary or by a different version of the same vocabulary. The doublet method can employ any version of any vocabulary with no pre-annotation. In many instances, the enormous effort and expense associated with data annotation can be eliminated by on-the-fly doublet matching. The algorithm for nomenclature-based database searches using the doublet method is described. Perl scripts for implementing the algorithm and testing execution speed are provided as open source documents available from the Association for Pathology Informatics (www.pathologyinformatics.org/informatics_r.htm).

Abstracting and Indexing↗

Approaches to integrating data within enterprise healthcare information systems.

The benefits of an Enterprise Healthcare Information System are related in large part to the degree with which its data and the processes it supports are integrated. There are several technical approaches to achieve integration. The strategic decision to put a group or several groups of applications within a single product to improve integration depends on the degree to which best of breed solutions or an integrated whole is needed. It also depends on a number of factors specific to each organization. It is important to understand the challenge of interfaces before choosing the best solution.

Hospital Information Systems↗

GABAagent: a system for integrating data on GABA receptors.

MOTIVATION: Scientific data pertaining to GABA receptors, which are of medical importance, are widely scattered throughout numerous heterogeneous Internet resources. This situation has made the integrated acquisition of such data difficult and substantially time consuming even for researchers who are Internet aficionados. Thus, there exists a genuine need for the development of Internet applications, such as GABAagent, which provide efficient and timely access to concise and integrated information. RESULTS: We report here the establishment of a novel server (GABAagent) which has been written in Perl script, and which is freely accessible through the Internet. GABAagent is designed to assist researchers in retrieving focused and integrated information related to GABA receptors from various public domain databases. GABAagent relies on server-side flat-file databases that have been created through data mining from Internet sources such as the PubMed, DDBJ, SWISS-PROT and TrEMBL, in addition to the many World Wide Web (Web) sites which are accessible through Excite (E-Web). These warehouse databases are regularly updated and contain among other things, information concerning: (i) GABA receptor publications, (ii) DNA and protein sequences and (iii) the contents of related E-Web sites along with their addresses. Our system also provides hard links to the above-mentioned Web sites and E-Web sites; the feature which adds to it the character of virtual federation type of database. The current version of GABAagent provides two user-friendly services. The first is a search engine possessing intelligent query reformulation support (GABAengine), the second an elaborate email alert service was designed into the system (GABAalert). The GABAengine allows the user to search server-side databases exclusively for GABA receptor-related queries. Whereas, GABAalert allows the user, by means of subscription, to receive immediate and/or monthly updates automatically. AVAILABILITY: GABAagent is freely accessible at the following Web address http://www.ust.hk/gaba.

Databases, Factual↗

The Adult Mouse Anatomical Dictionary: a tool for annotating and integrating data.

We have developed an ontology to provide standardized nomenclature for anatomical terms in the postnatal mouse. The Adult Mouse Anatomical Dictionary is structured as a directed acyclic graph, and is organized hierarchically both spatially and functionally. The ontology will be used to annotate and integrate different types of data pertinent to anatomy, such as gene expression patterns and phenotype information, which will contribute to an integrated description of biological phenomena in the mouse.

Animals↗