Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

Integrating genotypic and expression data in a segregating mouse population to identify 5-lipoxygenase as a susceptibility gene for obesity and bone traits.

Forward genetic approaches to identify genes involved in complex traits such as common human diseases have met with limited success. Fine mapping of linkage regions and validation of positional candidates are time-consuming and not always successful. Here we detail a hybrid procedure to map loci involved in complex traits that leverages the strengths of forward and reverse genetic approaches. By integrating genotypic and expression data in a segregating mouse population, we show how clusters of expression quantitative trait loci linking to regions of the genome accurately reflect the underlying perturbation to the transcriptional network induced by DNA variations in genes that control the complex traits. By matching patterns of gene expression in a segregating population with expression responses induced by single-gene perturbation experiments, we show how genes controlling clusters of expression and clinical quantitative trait loci can be mapped directly. We demonstrate the utility of this approach by identifying 5-lipoxygenase as underlying previously identified quantitative trait loci in an F(2) cross between strains C57BL/6J and DBA/2J and showing that it has pleiotropic effects on body fat, lipid levels and bone density.

Animals↗

Clinical evaluation of the effects of signal integrity and saturation on data availability and accuracy of Masimo SE and Nellcor N-395 oximeters in children.

UNLABELLED: Pulse oximetry manufacturers have introduced technologies that claim improved detection of hypoxemic events. Because improvements in signal processing and data rejection algorithms may differentially affect data reporting, we compared the data reporting and signal heuristic performance and agreement among the Nellcor N-395, Masimo SET, and GE Solar 8000 oximeters under a spectrum of conditions of signal integrity and arterial oxygen saturations. A blinded side-by-side comparison of technologies was performed in 27 patients, and data were analyzed for time of data availability, measures of agreement and signal heuristics, and warnings stratified by signal integrity and SpO(2). The Solar 8000 had less total data dropout than either of the new technologies. Masimo's LoSIQ (signal quality) heuristic rejected less data than Nellcor's MOT/PS (motion/pulse search) flag. When no signal heuristic was displayed, there was little difference in precision and bias between the two newer technologies; however, agreement between devices deteriorated in the presence of SIQ, MOT, or hypoxemia. Both newer devices flagged questionable data, but their use of different rejection algorithms resulted in different probabilities of presenting data. Therefore, with poor SIQ or during hypoxemia, the Nellcor N-395 and Masimo oximeters are not clinically equivalent to each other or to the older Solar 8000 oximeter. IMPLICATIONS: We compared new pulse oximeters from Nellcor and Masimo and found that, with good signal conditions, both new devices performed similarly to older technology. Overall, Masimo reported less data as questionable than Nellcor. With poor signal conditions or during hypoxemia, the new devices are not clinically equivalent to each other or to the older technology.

Algorithms↗

The semantic metadatabase (SEMEDA): ontology based integration of federated molecular biological data sources.

A system for "intelligent" semantic integration and querying of federated databases is being implemented by using three main components: A component which enables SQL access to integrated databases by database federation (MARGBench), an ontology based semantic metadatabase (SEMEDA) and an ontology based query interface (SEMEDA-query). In this publication we explain and demonstrate the principles, architecture and the use of SEMEDA. Since SEMEDA is implemented as 3 tiered web application database providers can enter all relevant semantic and technical information about their databases by themselves via a web browser. SEMEDA' s collaborative ontology editing feature is not restricted to database integration, and might also be useful for ongoing ontology developments, such as the "Gene Ontology" [2]. SEMEDA can be found at http://www-bm.cs.uni-magdeburg.de/semeda/. We explain how this ontologically structured information can be used for semantic database integration. In addition, requirements to ontologies for molecular biological database integration are discussed and relevant existing ontologies are evaluated. We further discuss how ontologies and structured knowledge sources can be used in SEMEDA and whether they can be merged supplemented or updated to meet the requirements for semantic database integration.

Databases, Genetic↗

MEANtools integrates multi-omics data to identify metabolites and predict biosynthetic pathways.

During evolution, plants have developed the ability to produce a vast array of specialized metabolites, which play crucial roles in helping plants adapt to different environmental niches. However, their biosynthetic pathways remain largely elusive. In the past decades, increasing numbers of plant biosynthetic pathways have been elucidated based on approaches utilizing genomics, transcriptomics, and metabolomics. These efforts, however, are limited by the fact that they typically adopt a target-based approach, requiring prior knowledge. Here, we present MEANtools, a systematic and unsupervised computational integrative omics workflow to predict candidate metabolic pathways de novo by leveraging knowledge of general reaction rules and metabolic structures stored in public databases. In our approach, possible connections between metabolites and transcripts that show correlated abundance across samples are identified using reaction rules linked to the transcript-encoded enzyme families. MEANtools thus assesses whether these reactions can connect transcript-correlated mass features within a candidate metabolic pathway. We validate MEANtools using a paired transcriptomic-metabolomic dataset recently generated to reconstruct the falcarindiol biosynthetic pathway in tomato. MEANtools correctly anticipated five out of seven steps of the characterized pathway and also identified other candidate pathways involved in specialized metabolism, which demonstrates its potential for hypothesis generation. Altogether, MEANtools represents a significant advancement to integrate multi-omics data for the elucidation of biochemical pathways in plants and beyond.

Metabolomics↗

Fatigue associated with congestive heart failure: use of Levine's Conservation Model.

This study aimed to refine and extend the findings of an original study which focused on the description of fatigue associated with congestive heart failure. A descriptive approach based on Levine's Conservation Model provided both quantitative and qualitative data. Qualitative data addressed personal integrity and quantitative data measured energy conservation, structural and social integrity. Patients described fatigue as being tired and exhausted and containing both physical and emotional components. Fatigue occurred as a result of stress, physical activity and disease. Patient-identified interventions included rest, distraction, medicine, and physical and spiritual activities. Age, pH and oxygen saturation were significantly related to fatigue. The findings are examined using the concept of adaptation as defined by Levine. Implications for nursing are discussed within the framework of the Conservation Model with emphasis on a holistic approach to patient care.

Adaptation, Physiological↗

Exploiting big biology: integrating large-scale biological data for function inference.

The amount of data produced by molecular biologists is growing at an exponential rate. Some of the fastest growing sets of data are measurements of gene expression, comparable in quantity only to gene sequences and the vast biological literature. Both gene expression data and sequence data offer hints as to the functions of thousands of newly discovered genes, but neither give complete answers. Therefore, much effort is being focused on integrating these large data sets and combining them with all available functional data to draw inferences about the functions of uncharacterised genes. This review discusses the most pertinent functional data for genome-wide functional inference and describes several methods by which these disparate data types are being integrated.

Data Collection↗

Data management solutions for protein therapeutic research and development.

Protein therapeutics, including monoclonal antibodies, are a growing focus of drug discovery research organizations. High-throughput screening of large libraries of protein variants is therefore becoming increasingly important in R&D. As a result, there is a need to link large numbers of variant protein sequences with chemical and biological assay data. This integration will allow more efficient data mining and facilitate decision-making regarding hit identification, lead optimization and drug development. In this paper, we present an implementation in which a widely used small-molecule high-throughput screening data management system has been adapted to meet the unique needs of protein drug discovery and development.

Antibodies, Monoclonal↗

Mining ChIP-chip data for transcription factor and cofactor binding sites.

MOTIVATION: Identification of single motifs and motif pairs that can be used to predict transcription factor localization in ChIP-chip data, and gene expression in tissue-specific microarray data. RESULTS: We describe methodology to identify de novo individual and interacting pairs of binding site motifs from ChIP-chip data, using an algorithm that integrates localization data directly into the motif discovery process. We combine matrix-enumeration based motif discovery with multivariate regression to evaluate candidate motifs and identify motif interactions. When applied to the HNF localization data in liver and pancreatic islets, our methods produce motifs that are either novel or improved known motifs. All motif pairs identified to predict localization are further evaluated according to how well they predict expression in liver and islets and according to how conserved are the relative positions of their occurrences. We find that interaction models of HNF1 and CDP motifs provide excellent prediction of both HNF1 localization and gene expression in liver. Our results demonstrate that ChIP-chip data can be used to identify interacting binding site motifs. AVAILABILITY: Motif discovery programs and analysis tools are available on request from the authors.

Algorithms↗

Genome-wide analysis of glucose-6-phosphate dehydrogenases in Arabidopsis.

In green tissues of plants under illumination, photosynthesis is the primary source of reduced nicotinamide adenine dinucleotide phosphate (NADPH), which is utilized in reductive reactions such as carbon fixation and nitrogen assimilation. In non-photosynthetic tissues or under non-photosynthetic conditions, the oxidative pentose phosphate pathway contributes to basic metabolism as one of the major sources of NADPH. The first and committed reaction is catalyzed by glucose-6-phosphate dehydrogenase (G6PDH). We characterized the six members of the G6PDH gene family in Arabidopsis. Transit peptide analysis predicted two cytosolic and four plastidic isoforms. Five of the six genes encode active G6PDHs. The recombinant isoforms showed differences in substrate requirements and sensitivities to feedback inhibition. Plastidic isoforms were redox sensitive. One cytosolic isoform was insensitive to redox changes, while the other was inactivated by oxidation. The respective genes had distinct expression patterns that did not correlate with the activity of the proteins, implying a regulatory mechanism beyond the control of mRNA abundance. Two cytosolic and one plastidic isoform were detected in vivo using zymograms, and the respective genes were identified using T-DNA insertion lines. The activity of a plastidic isoform was detected in all tissues including photosynthetic tissues despite its sensitivity to reduction observed in vitro. Genomic data, gene expression, and in vivo enzyme activity data were integrated with in vitro biochemical data to propose in vivo roles for individual G6PDH isoforms in Arabidopsis.

Amino Acid Sequence↗

Lost human capital from early-onset chronic depression.

OBJECTIVE: Chronic depression starts at an early age for many individuals and could affect their accumulation of "human capital" (i.e., education, higher amounts of which can broaden occupational choice and increase earnings potential). The authors examined the impact, by gender, of early- (before age 22) versus late-onset major depressive disorder on educational attainment. They also determined whether the efficacy and sustainability of antidepressant treatments and psychosocial outcomes vary by age at onset and quantified the impact of early- versus late-onset, as well as never-occurring, major depressive disorder on expected lifetime earnings. METHOD: The authors used logistic and multivariate regression methods to analyze data from a three-phase, multicenter, double-blind, randomized trial that compared sertraline and imipramine treatment of 531 patients with chronic depression aged 30 years and older. These data were integrated with U.S. Census Bureau data on 1995 earnings by age, educational attainment, and gender. RESULTS: Early-onset major depressive disorder adversely affected the educational attainment of women but not of men. No significant difference in treatment responsiveness by age at onset was observed after 12 weeks of acute treatment or, for subjects rated as having responded, after 76 weeks of maintenance treatment. A randomly selected 21-year-old woman with early-onset major depressive disorder in 1995 could expect future annual earnings that were 12%-18% lower than those of a randomly selected 21-year-old woman whose onset of major depressive disorder occurred after age 21 or not at all. CONCLUSIONS: Early-onset major depressive disorder causes substantial human capital loss, particularly for women. Detection and effective treatment of early-onset major depressive disorder may have substantial economic benefits.

Adult↗

Surveillance system of infectious diseases in Japan.

The surveillance system of infectious disease in Japan started in 1981 and has been providing useful epidemiological information on 27 communicable diseases. The system consists of medical institutions (fixed monitoring stations), institutions of hygienic sciences, health centers, local governments and the ministry of health and welfare. There are two types of information about infectious diseases. One is clinical reports of incidence cases from medical institutions, and the other is laboratory information about etiologic agents. Between health centers, local governments and the department of statistics and information in the ministry of health and welfare, information is transmitted through the on-line network. Collected information is analyzed and submitted by both local and central committees of analysis. From the epidemiological point of view, quality control of the data and integration of other sources of data would be the next goal of the system.

Databases, Factual↗

The doctor-patient relationship in the practice of medicine.

The patient-doctor relationship is based on the principles of interaction, collecting data and integration of both interaction and data into an overall diagnosis/therapy. Patients with functional abdominal disorders are seen as representatives of today's general patients and a study of their management in present medical practice is reported, as revealed through literature. The literature reveals an almost complete neglect of intractional and intergrational principles. This holds true even for psychosomatically oriented literature, which offers some crude clinical guidelines at best. Thus the primary physician gets little support from psychosomatic medicine in understanding the full meaning of the doctor-patient relationship. The clinical implications of the relationship are demonstrated through a short case history and implications for future training are described which are based on the primary physician's actual working experiences.

Adaptation, Psychological↗

Clinical genomics data standards for pharmacogenetics and pharmacogenomics.

This special report concerns a talk on data standards given at a workshop entitled 'An International Perspective on Pharmacogenetics: The Intersections between Innovation, Regulation and Health Delivery', which was held by the Organization for Economic Co-operation and Development (OECD) on October 17-19, 2005, in Rome, Italy. The worlds of healthcare and life sciences (HCLS) are extremely fragmented in terms of their underlying information technology, making it difficult to semantically exchange information between disparate entities. While we have reached the point where functional interoperability is ubiquitous, we are still far from achieving true semantic interoperability where a receiving system can use incoming data as though it was created internally. The critical enablers of semantic interoperability are information standards dedicated to HCLS data, spanning all the way from biological research data to clinical research and clinical trials, and finally to healthcare clinical data. The challenge lies in integrating various data standards based on predetermined goals, thereby improving the quality of care provided to patients.

Clinical Medicine↗

Interactive data collection: benefits of integrating new media into pediatric research.

Despite the prevalence of children's computerized games for recreational and educational purposes, the use of interactive technology to obtain pediatric research data remains underexplored. This article describes the development of laptop interactive data collection (IDC) software for a children's health intervention study. The IDC integrates computer technology, children's developmental needs, and quantitative research methods that are engaging for school-age children as well as reliable and efficient for the pediatric health researcher. Using this methodology, researchers can address common problems such as maintaining a child's attention throughout an assessment session while potentially increasing their response rate and reducing missing data rates. The IDC also promises to produce more reliable data by eliminating the need for manual double entry of data and reducing much of the time and costs associated with data cleaning and management. Development and design considerations and recommendations for further use are discussed.

Child↗

A regression-based K nearest neighbor algorithm for gene function prediction from heterogeneous data.

BACKGROUND: As a variety of functional genomic and proteomic techniques become available, there is an increasing need for functional analysis methodologies that integrate heterogeneous data sources. METHODS: In this paper, we address this issue by proposing a general framework for gene function prediction based on the k-nearest-neighbor (KNN) algorithm. The choice of KNN is motivated by its simplicity, flexibility to incorporate different data types and adaptability to irregular feature spaces. A weakness of traditional KNN methods, especially when handling heterogeneous data, is that performance is subject to the often ad hoc choice of similarity metric. To address this weakness, we apply regression methods to infer a similarity metric as a weighted combination of a set of base similarity measures, which helps to locate the neighbors that are most likely to be in the same class as the target gene. We also suggest a novel voting scheme to generate confidence scores that estimate the accuracy of predictions. The method gracefully extends to multi-way classification problems. RESULTS: We apply this technique to gene function prediction according to three well-known Escherichia coli classification schemes suggested by biologists, using information derived from microarray and genome sequencing data. We demonstrate that our algorithm dramatically outperforms the naive KNN methods and is competitive with support vector machine (SVM) algorithms for integrating heterogenous data. We also show that by combining different data sources, prediction accuracy can improve significantly CONCLUSION: Our extension of KNN with automatic feature weighting, multi-class prediction, and probabilistic inference, enhance prediction accuracy significantly while remaining efficient, intuitive and flexible. This general framework can also be applied to similar classification problems involving heterogeneous datasets.

Algorithms↗

The bioinformatics resource for oral pathogens.

Complete genomic sequences of several oral pathogens have been deciphered and multiple sources of independently annotated data are available for the same genomes. Different gene identification schemes and functional annotation methods used in these databases present a challenge for cross-referencing and the efficient use of the data. The Bioinformatics Resource for Oral Pathogens (BROP) aims to integrate bioinformatics data from multiple sources for easy comparison, analysis and data-mining through specially designed software interfaces. Currently, databases and tools provided by BROP include: (i) a graphical genome viewer (Genome Viewer) that allows side-by-side visual comparison of independently annotated datasets for the same genome; (ii) a pipeline of automatic data-mining algorithms to keep the genome annotation always up-to-date; (iii) comparative genomic tools such as Genome-wide ORF Alignment (GOAL); and (iv) the Oral Pathogen Microarray Database. BROP can also handle unfinished genomic sequences and provides secure yet flexible control over data access. The concept of providing an integrated source of genomic data, as well as the data-mining model used in BROP can be applied to other organisms. BROP can be publicly accessed at http://www.brop.org.

Bacteria↗

Caveat doctor: how to analyze claims-based report cards.

"Report cards" based on claims (billing) data are being widely used to evaluate the quality of care given by providers, even though they often lack sufficient clinical detail to render definitive judgments. Furthermore, their accuracy, especially for outpatient care, is quite variable. Nevertheless, claims data will continue to be used until better clinical information becomes widely available. To determine the suitability of automated claims data for measuring clinical performance, careful attention should be paid to the integrity of the data. Providers profiled by claims-based report cards should ask four questions about the source, robustness, management, and analysis of the data: 1. What are the key characteristics of the data set used to construct the profile? These include the insurer's name, coverage type, time period, geographic area, and number of patients, claims lines, and providers. 2. What clinical conditions and events are being measured and how well? In short, are the patients' conditions and their clinical encounters reasonably well characterized? 3. Is the information about the patients and providers accurate and up to date? 4. Once the insurer receives the medical claim, are data elements deleted or altered in ways that might affect their accuracy and completeness? Ensuring data integrity is not sufficient; the analysis of the data must be scrutinized. Potential pitfalls in analyzing claims data arise in choosing clinically meaningful measures, recognizing important differences in patients and their providers, and making fair comparisons against appropriate benchmarks. Monitoring patient care outcomes is no longer voluntary. By routinely constructing and augmenting profiles using outpatient claims data, provider groups become proactive rather than reactive in evaluating their patients' care.

Benchmarking↗