Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Genomic structural equation modeling elucidates the shared genetic architecture of allergic disorders.

BACKGROUND: The intricate shared genetic architecture underlying allergic disorders-including allergic asthma, atopic dermatitis, contact dermatitis, allergic rhinitis, allergic conjunctivitis, allergic urticaria, anaphylaxis, and eosinophilic esophagitis-remains incompletely characterized. METHODS: Our study employed genomic structural equation modeling (Genomic SEM) to define the common factor representing the shared genetic architecture of allergic disorders. Coupled with diverse post-GWAS analytical methods, we aimed to discover susceptible loci and investigate genetic associations with external traits. Furthermore, we explored enriched genetic pathways, cellular layers, and genomic elements, and investigated putative plasma protein biomarkers. Polygenic risk score (PRS) analyses, leveraging our integrated GWAS data, were conducted to assess chromosomal-level risk associations for allergic disorders. RESULTS: A well-fitted genomic SEM integrated GWAS data, revealing the shared genetic architecture of allergic disorders. We identified a total of 2038 genome-wide significant SNP loci (p&#x2009;<&#x2009;5e-8), including 31 previously unreported loci. Fine-mapping of variants and gene sets pinpointed 2 causal variants and 31 candidate susceptible genes. Genetic correlation analyses further illuminated the shared genetic architecture underlying multiple traits, notably psychiatric disorders. Preliminary findings identified four putative causal plasma protein biomarkers. CONCLUSION: Notably, this study presents the first comprehensive genetic characterization of allergic disorders through a GWAS analysis of an unmeasured composite phenotype, providing novel insights into shared etiological pathways across these conditions.

Humans↗

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans↗

Integrating in vitro ADMET data through generic physiologically based pharmacokinetic models.

Early estimation of kinetics in man currently relies on extrapolation from experimental data generated in animals. Recent results from the application of a generic physiologically based model, Cloe PK) (Cyprotex), which is parameterised for human and rat physiology, to the estimation of plasma pharmacokinetics, are summarised in this paper. A comparison with predictive methods that involve scaling from in vivo animal data can also be made from recently published data. On average, the divergence of the predicted plasma concentrations from the observed data was 0.47 log units. For the external test set, > 70% of the predicted values of the AUC were within threefold of the observed values. Furthermore, the model was found to match or exceed the performance of three published interspecies scaling methods for estimating clearance, all of which showed a distinct bias towards overprediction. It is concluded that Cloe PK, as a means of integrating readily determined in vitro and/or in silico data, is a powerful, cost-effective tool for estimating exposure and kinetics in drug discovery and risk assessment that should, if widely adopted, lead to major reductions in the need for animal experimentation.

Data Interpretation, Statistical↗

Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes.

BACKGROUND: Due to the high cost and low reproducibility of many microarray experiments, it is not surprising to find a limited number of patient samples in each study, and very few common identified marker genes among different studies involving patients with the same disease. Therefore, it is of great interest and challenge to merge data sets from multiple studies to increase the sample size, which may in turn increase the power of statistical inferences. In this study, we combined two lung cancer studies using microarray GeneChip, employed two gene shaving methods and a two-step survival test to identify genes with expression patterns that can distinguish diseased from normal samples, and to indicate patient survival, respectively. RESULTS: In addition to common data transformation and normalization procedures, we applied a distribution transformation method to integrate the two data sets. Gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) were then applied separately to the joint data set for cancer gene selection. The two methods discovered 13 and 10 marker genes (5 in common), respectively, with expression patterns differentiating diseased from normal samples. Among these marker genes, 8 and 7 were found to be cancer-related in other published reports. Furthermore, based on these marker genes, the classifiers we built from one data set predicted the other data set with more than 98% accuracy. Using the univariate Cox proportional hazard regression model, the expression patterns of 36 genes were found to be significantly correlated with patient survival (p < 0.05). Twenty-six of these 36 genes were reported as survival-related genes from the literature, including 7 known tumor-suppressor genes and 9 oncogenes. Additional principal component regression analysis further reduced the gene list from 36 to 16. CONCLUSION: This study provided a valuable method of integrating microarray data sets with different origins, and new methods of selecting a minimum number of marker genes to aid in cancer diagnosis. After careful data integration, the classification method developed from one data set can be applied to the other with high prediction accuracy.

Adenocarcinoma↗

Field evaluation of a method for estimating gaseous fluxes from area sources using open-path Fourier transform infrared.

This paper describes results from the first field experiment designed to evaluate a new approach for quantifying gaseous fugitive emissions of area air pollution sources. The approach combines path-integrated concentration data acquired with any path-integrated optical remote sensing (PI-ORS) technique and computed tomography (CT) technique. In this study, an open-path Fourier transform infrared (OP-FTIR) instrument sampled path-integrated concentrations along five radial beam paths in a vertical plane downwind from the source. A meteorological station collected measurements of wind direction and wind speed. Nitrous oxide (N2O) was released from a controlled area source simulator. The innovative CT technique, which applies the smooth basis function minimization method to the beam data in conjunction with measured wind data, was used to estimate the total flux from the simulated area source. The new approach estimates consistently underestimated the true emission rates in unstable atmospheric conditions and agreed with the true emission rate in neutral atmospheric conditions. This approach is applicable to many types of industrial areas or volume sources, given the use of an adequate PI-ORS system.

Air Pollution↗

Integration is required for productive infection of monocyte-derived macrophages by human immunodeficiency virus type 1.

Certain human immunodeficiency virus type 1 (HIV-1) isolates are able to productively infect nondividing cells of the monocyte/macrophage lineage. We have used a molecular genetic approach to construct two different HIV-1 integrase mutants that were studied in the context of an infectious, macrophage-tropic HIV-1 molecular clone. One mutant, HIV-1 delta D(35)E, containing a 37-residue deletion within the central, catalytic domain of integrase, was noninfectious in both peripheral blood mononuclear cells and monocyte-derived macrophages. The HIV-1 delta D(35)E mutant, however, exhibited defects in the assembly and/or release of progeny virions in transient transfection assays, as well as defects in entry and/or viral DNA synthesis during the early stages of monocyte-derived macrophage infection. The second mutant, HIV-1D116N/8, containing a single Asp-to-Asn substitution at the invariant Asp-116 residue of integrase, was also noninfectious in both peripheral blood mononuclear cells and monocyte-derived macrophages but, in contrast to HIV-1 delta D(35)E, was indistinguishable from wild-type virus in reverse transcriptase production. PCR analysis indicated that HIV-1D116N/8 entered monocyte-derived macrophages efficiently and reverse transcribed its RNA but was unable to complete its replication cycle because of a presumed block to integration. These data are consistent with the hypothesis that integration is an obligate step in productive HIV-1 infection of activated peripheral blood mononuclear cells and primary human macrophage cultures.

Amino Acid Sequence↗

Unifying heterogeneous distributed clinical data in a relational database.

Access to clinical data which are distributed among multiple satellite information systems is crucial to delivering better care and reducing costs in many hospitals and medical centers. An integrated view of these data is needed to reduce the effort of users requiring data from multiple systems. We have addressed the issue of distributed data integration while developing both production and research decision-support applications. We describe an ideal integration solution, obstacles to realizing this solution, and our integration requirements and architecture. Our focus is a description of our specific schema and data integration techniques. We conclude with an analysis of our approach.

Computer Communication Networks↗

Syndromic surveillance using automated collection of computerized discharge diagnoses.

The Syndromic Surveillance Information Collection (SSIC) system aims to facilitate early detection of bioterrorism attacks (with such agents as anthrax, brucellosis, plague, Q fever, tularemia, smallpox, viral encephalitides, hemorrhagic fever, botulism toxins, staphylococcal enterotoxin B, etc.) and early detection of naturally occurring disease outbreaks, including large foodborne disease outbreaks, emerging infections, and pandemic influenza. This is accomplished using automated data collection of visit-level discharge diagnoses from heterogeneous clinical information systems, integrating those data into a common XML (Extensible Markup Language) form, and monitoring the results to detect unusual patterns of illness in the population. The system, operational since January 2001, collects, integrates, and displays data from three emergency department and urgent care (ED/UC) departments and nine primary care clinics by automatically mining data from the information systems of those facilities. With continued development, this system will constitute the foundation of a population-based surveillance system that will facilitate targeted investigation of clinical syndromes under surveillance and allow early detection of unusual clusters of illness compatible with bioterrorism or disease outbreaks.

Bioterrorism↗

MagnaportheDB: a federated solution for integrating physical and genetic map data with BAC end derived sequences for the rice blast fungus Magnaporthe grisea.

We have created a federated database for genome studies of Magnaporthe grisea, the causal agent of rice blast disease, by integrating end sequence data from BAC clones, genetic marker data and BAC contig assembly data. A library of 9216 BAC clones providing >25-fold coverage of the entire genome was end sequenced and fingerprinted by HindIII digestion. The Image/FPC software package was then used to generate an assembly of 188 contigs covering >95% of the genome. The database contains the results of this assembly integrated with hybridization data of genetic markers to the BAC library. AceDB was used for the core database engine and a MySQL relational database, populated with numerical representations of BAC clones within FPC contigs, was used to create appropriately scaled images. The database is being used to facilitate sequencing efforts. The database also allows researchers mapping known genes or other sequences of interest, rapid and easy access to the fundamental organization of the M.grisea genome. This database, MagnaportheDB, can be accessed on the web at http://www.cals.ncsu.edu/fungal_genomics/mgdatabase/int.htm.

Base Sequence↗

The art of assessment in psychology: ethics, expertise, and validity.

Psychological assessment is a hybrid, both art and science. The empirical foundations of testing are indispensable in providing reliable and valid data. At the level of the integrated assessment, however, science gives way to art. Standards of reliability and validity account for the individual instrument; they do not account for the integration of data into a comprehensive assessment. This article examines the current climate of psychological assessment, selectively reviewing the literature of the past decade. Ethics, expertise, and validity are the components under discussion. Psychologists can and do take precautions to ensure that the "art" of their work holds as much merit as the science.

Ethics, Medical↗

An SSR-based genetic linkage map for perennial ryegrass ( Lolium perenne L.).

A simple sequence repeat (SSR)-based linkage map has been constructed for perennial ryegrass ( Lolium perenne L.) using a one-way pseudo-testcross reference population. A total of 309 unique perennial ryegrass SSR (LPSSR) primer pairs showing efficient amplification were evaluated for genetic polymorphism, with 31% detecting segregating alleles. Ninety-three loci have been assigned to positions on seven linkage groups. The majority of the mapped loci are derived from cloned sequences containing (CA)(n)-type dinucleotide SSR arrays. A small number (7%) of primer pairs amplified fragments that mapped to more than one locus. The SSR locus data has been integrated with selected data for RFLP, AFLP and other loci mapped in the same population to produce a composite map containing 258 loci. The SSR loci cover 54% of the genetic map and show significant clustering around putative centromeric regions. BLASTN and BLASTX analysis of the sequences flanking mapped SSRs indicated that a majority (84%) are derived from non-genic sequences, with a small proportion corresponding to either known repetitive DNA sequence families or predicted genes. The mapped LPSSR loci provide the basis for linkage group assignment across multiple mapping populations.

Journal Article↗

Direct adverse effects of Sendai virus DI particles on virus budding and on M protein fate and stability.

Upon infections of BHK cells with a mixture of Sendai standard and defective interfering (DI) viruses (mixed virus infection), viral budding was found to be restricted by factors ranging from 5 to more than 20. The reduced viral budding correlated with a high intracellular M protein turnover. M appeared to be degraded shortly after its synthesis, and seemed not to be able to self-associate in a stable way under the plasma membrane as it did in St virus-infected cells. These data, added to the previous findings that infection with DI particles allowed infected cell survival and favored the cell-surface turnover of the hemagglutinin-neuraminidase protein, led to the hypothesis that DI genomes directly act by preventing the stable formation inside the cells of a viral structure composed of M/HN/nucleocapsids. When involved in this structure M would be protected from degradation and HN would be stably anchored in the plasma membrane. Formation of this structure would be necessary for viral budding and would be damaging for the cells. Comparison with results published by other authors shows that such a model is consistent with other data. It can integrate, as well, data obtained in the analysis of mutant viruses involved in persistence.

Defective Viruses↗

GSD: a genetic screen database.

The systematic assignment of gene function to a sequenced genome is one of the outstanding challenges in the post-genomic era. Large-scale systematic mutagenesis screens are important tools for reaching this goal. Here we describe GSD, a software package that allows storage and integration of data from genetic screens. GSD was initially developed for a large-scale F3 mutagenesis screen for developmental mutants of medaka (Oryzias latipes). The version presented here supports a wide range of different screens (mutagenesis, RNAi, morpholinos, transgenesis and others) using different organisms. Data are stored in a relational database and can be made accessible through web interfaces. Researchers can enter data describing their screened embryos: They can track statistics, submit images and describe the resulting phenotypes using a phenotype classification ontology. We developed a fish phenotype classification ontology of medaka and zebrafish for this software package and made it available to the public. In addition, a list of genetic lines resulting from each screen can be generated. These lines (mutant alleles, transgenic lines) can be described and categorized in the same ways as the screened individuals. Raw data from the screen can be integrated to describe these lines. A query module that searches this list can be used to publish the screen results on the Internet. A test version is available at and the software can be downloaded from this site.

Animals↗

Experience with airborne detection of radioactive pollution (ENMOS, IRIS).

This paper discusses the advantages of airborne monitoring of radioactive pollution and shows example maps indicating manmade pollution from different sources. The sensitivity of airborne radioactive detection is discussed. Comparisons of airborne and different ground measurements are presented. New instrumentation for airborne or ground moving vehicles is briefly described. Airborne footprinting provides rapid, well-defined spatial images of natural and manmade radioactive contamination. Data acquisition integrated with GPS navigation provides consistent data and guarantees proper data location. Real-time airborne measurements are re-calculated, with the use of special algorithms, into absolute units for individual radioactive nuclei contamination of the ground together with dose calculation. Raw records and calculated data are provided after enhanced post-flight processing. Dose rates and detection of different radioactive elements are presented. (ENMOS is a product of Picodas Group Inc. and IRIS is the product of Pico Envirotec Inc.)

Air Pollutants, Radioactive↗

Audio-visual integration in schizophrenia.

Integration of information provided simultaneously by audition and vision was studied in a group of 18 schizophrenic patients. They were compared to a control group, consisting of 12 normal adults of comparable age and education. By administering two tasks, each focusing on one aspect of audio-visual integration, the study could differentiate between a spatial integration deficit and a speech-based integration deficit. Experiment 1 studied audio-visual interactions in the spatial localisation of sounds. Experiment 2 investigated integration of auditory and visual speech. The schizophrenic group performed as the control group on the sound localisation task, but in the audio-visual speech task, there was an impairment in lipreading as well as a smaller impact of lipreading on auditory speech information. Combined with findings about functional and neuro-anatomical specificity of intersensory integration, the data suggest that there is an integration deficit in the schizophrenic group that is related to the processing of phonetic information.

Adult↗

A managed care perspective.

This paper highlights some of the problems associated with lipid therapy in the primary and secondary prevention of cardiovascular disorders and to make some potentially useful suggestions in the context of managed care. For managed care organizations, financial and logistical issues create obstacles to the provision of primary prevention of cardiovascular disease. These current obstacles necessitate the generation of external forces, perhaps regulatory or standards agencies, that may help increase accountability in managed care organizations for midterm and distant outcomes. In contrast, the provision of secondary prevention by managed care organizations has fewer limitations. One of the major challenges in secondary prevention, however, is the low rate of physician compliance with national treatment guidelines and standards. Among possible explanations for this observation are limitations in health data collection and integration. Improvements in data management are vital to the achievement of treatment goal optimization in secondary prevention.

Journal Article↗

ChimerDB--a knowledgebase for fusion sequences.

Chromosome translocation and gene fusion are frequent events in the human genome and are often the cause of many types of tumor. ChimerDB is the database of fusion sequences encompassing bioinformatics analysis of mRNA and expressed sequence tag (EST) sequences in the GenBank, manual collection of literature data and integration with other known database such as OMIM. Our bioinformatics analysis identifies the fusion transcripts that have non-overlapping alignments at multiple genomic loci. Fusion events at exon-exon borders are selected to filter out the cloning artifacts in cDNA library preparation. The result is classified into two groups--genuine chromosome translocation and fusion between neighboring genes owing to intergenic splicing. We also integrated manually collected literature and OMIM data for chromosome translocation as an aid to assess the validity of each fusion event. The database is available at http://genome.ewha.ac.kr/ChimerDB/ for human, mouse and rat genomes.

Animals↗

A rigorous method for multigenic families' functional annotation: the peptidyl arginine deiminase (PADs) proteins family example.

BACKGROUND: large scale and reliable proteins' functional annotation is a major challenge in modern biology. Phylogenetic analyses have been shown to be important for such tasks. However, up to now, phylogenetic annotation did not take into account expression data (i.e. ESTs, Microarrays, SAGE, ...). Therefore, integrating such data, like ESTs in phylogenetic annotation could be a major advance in post genomic analyses. We developed an approach enabling the combination of expression data and phylogenetic analysis. To illustrate our method, we used an example protein family, the peptidyl arginine deiminases (PADs), probably implied in Rheumatoid Arthritis. RESULTS: the analysis was performed as follows: we built a phylogeny of PAD proteins from the NCBI's NR protein database. We completed the phylogenetic reconstruction of PADs using an enlarged sequence database containing translations of ESTs contigs. We then extracted all corresponding expression data contained in EST database This analysis allowed us 1/To extend the spectrum of homologs-containing species and to improve the reconstruction of genes' evolutionary history. 2/To deduce an accurate gene expression pattern for each member of this protein family. 3/To show a correlation between paralogous sequences' evolution rate and pattern of tissular expression. CONCLUSION: coupling phylogenetic reconstruction and expression data is a promising way of analysis that could be applied to all multigenic families to investigate the relationship between molecular and transcriptional evolution and to improve functional annotation.

Animals↗