Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Customized dual data entry for computerized data analysis.

A major responsibility of any Quality Assurance Unit (QUA) is to ensure data integrity. Errors made during data entry can lead to many problems in the study review process and decrease the quality, accuracy, and overall efficiency of data management. One technique that can reduce the number of data entry errors in computer data sets is the use of a dual entry data system. Currently available software allows creation of customized data entry screens that either closely resemble or duplicate the data collection forms used during studies. Two data entry operators enter data into two independent data sets. The use of an on-screen display that resembles the data collection form reduces the potential for keypunch errors. The two data sets can then be electronically compared. The comparison reports differences between the two data sets. When differences exist, the correct values can be determined by reference to the original data sheets and the two data files can then be corrected. Theoretically, the only key punch errors that will exist after making these corrections are when the two independent entry operators make the same exact data entry error. Typically, the time required for two people to enter data is minimal compared to the time required to manually identify and correct data entry discrepancies. With error-free data entry, we have found that electronic data quality, accuracy, and audit efficiency are improved at every subsequent step of data management, analysis, quality assurance auditing, and report generation.

Information Systems↗

[General principles for safety evaluation of pharmaceuticals in man based on integration of data from various sources].

Safety evaluation of pharmaceuticals consists of two processes; firstly, to grade the adverse effects of a test material on individuals based on scientific evidence, and, secondly, to judge whether the adverse effects occurring under the dose condition of the material capable of exhibiting its efficacy in patients remains within the acceptable safe range. Accordingly, the basic criteria for safety evaluation of pharmaceuticals can be, simply, said to know how high the dose-response curve of adverse effects lies above that of efficacy. On the other hand, the judgement concerning how much difference is necessary between both dose-response curves with regard to the safety would require a careful consideration on a case by case basis taking into account various information on the risk/benefit balance of the drug such as 1) medical usefulness and social needs of the drug and 2) the presumed severity of adverse effects in man.

Data Collection↗

Race/ethnicity, substance abuse, and mental illness among suicide victims in 13 US states: 2004 data from the National Violent Death Reporting System.

OBJECTIVE: To calculate the prevalence of substance abuse and mental illness among suicide victims of different racial/ethnic groups and to identify race/ethnicity trends in mental health and substance abuse that may be used to improve suicide prevention. METHODS: Data are from the National Violent Death Reporting System (NVDRS), a state-based data integration system that, for 2004, includes data from 13 US states. The NVDRS integrates medical examiner, toxicology, death certificate, and law enforcement data. RESULTS: Within participating states, for data year 2004, 6865 suicide incidents in which race/ethnicity are known were identified. This included 5797 (84.4%) non-Hispanic whites, 501 (7.3%) non-Hispanic blacks, 257 (3.7%) Hispanics, and 310 (4.5%) persons from other racial/ethnic groups. At the time of the suicide event, non-Hispanic blacks had lower blood alcohol contents than other groups. Non-Hispanic whites had less cocaine but more antidepressants and opiates. There were no differences in the levels of amphetamines or marijuana by race/ethnicity. Hispanics were less likely to have been diagnosed with a mental illness or to have received treatment, although family reports of depression were comparable to non-Hispanic whites and other racial/ethnic groups. Non-Hispanic whites were more likely to be diagnosed with depression or bipolar disorder and non-Hispanic blacks with schizophrenia. Comorbid substance abuse and mental health problems were more likely among non-Hispanic whites and non-Hispanic blacks, while Hispanics were more likely to have a substance abuse problem without comorbid mental health problems. CONCLUSION: The results support earlier research documenting differences in race/ethnicity, substance abuse, and mental health problems as they relate to completed suicide. The data suggest that suicide prevention efforts must address not only substance abuse and mental health problems in general, but the unique personal, family, and social characteristics of different racial/ethnic groups.

Adolescent↗

Future of toxicology--predictive toxicology: An expanded view of "chemical toxicity".

A chemistry approach to predictive toxicology relies on structure-activity relationship (SAR) modeling to predict biological activity from chemical structure. Such approaches have proven capabilities when applied to well-defined toxicity end points or regions of chemical space. These approaches are less well-suited, however, to the challenges of global toxicity prediction, i.e., to predicting the potential toxicity of structurally diverse chemicals across a wide range of end points of regulatory and pharmaceutical concern. New approaches that have the potential to significantly improve capabilities in predictive toxicology are elaborating the "activity" portion of the SAR paradigm. Recent advances in two areas of endeavor are particularly promising. Toxicity data informatics relies on standardized data schema, developed for particular areas of toxicological study, to facilitate data integration and enable relational exploration and mining of data across both historical and new areas of toxicological investigation. Bioassay profiling refers to large-scale high-throughput screening approaches that use chemicals as probes to broadly characterize biological response space, extending the concept of chemical "properties" to the biological activity domain. The effective capture and representation of legacy and new toxicity data into mineable form and the large-scale generation of new bioassay data in relation to chemical toxicity, both employing chemical structure information to inform and integrate diverse biological data, are opening exciting new horizons in predictive toxicology.

Animals↗

Integrating syndromic surveillance data across multiple locations: effects on outbreak detection performance.

Syndromic surveillance systems are being deployed widely to monitor for signals of covert bioterrorist attacks. Regional systems are being established through the integration of local surveillance data across multiple facilities. We studied how different methods of data integration affect outbreak detection performance. We used a simulation relying on a semi-synthetic dataset, introducing simulated outbreaks of different sizes into historical visit data from two hospitals. In one simulation, we introduced the synthetic outbreak evenly into both hospital datasets (aggregate model). In the second, the outbreak was introduced into only one or the other of the hospital datasets (local model). We found that the aggregate model had a higher sensitivity for detecting outbreaks that were evenly distributed between the hospitals. However, for outbreaks that were localized to one facility, maintaining individual models for each location proved to be better. Given the complementary benefits offered by both approaches, the results suggest building a hybrid system that includes both individual models for each location, and an aggregate model that combines all the data. We also discuss options for multi-level signal integration hierarchies.

Bioterrorism↗

Gene Aging Nexus: a web database and data mining platform for microarray data on aging.

The recent development of microarray technology provided unprecedented opportunities to understand the genetic basis of aging. So far, many microarray studies have addressed aging-related expression patterns in multiple organisms and under different conditions. The number of relevant studies continues to increase rapidly. However, efficient exploitation of these vast data is frustrated by the lack of an integrated data mining platform or other unifying bioinformatic resource to enable convenient cross-laboratory searches of array signals. To facilitate the integrative analysis of microarray data on aging, we developed a web database and analysis platform 'Gene Aging Nexus' (GAN) that is freely accessible to the research community to query/analyze/visualize cross-platform and cross-species microarray data on aging. By providing the possibility of integrative microarray analysis, GAN should be useful in building the systems-biology understanding of aging. GAN is accessible at http://gan.usc.edu.

Aging↗

BiologicalNetworks: visualization and analysis tool for systems biology.

Systems level investigation of genomic scale information requires the development of truly integrated databases dealing with heterogeneous data, which can be queried for simple properties of genes or other database objects as well as for complex network level properties, for the analysis and modelling of complex biological processes. Towards that goal, we recently constructed PathSys, a data integration platform for systems biology, which provides dynamic integration over a diverse set of databases [Baitaluk et al. (2006) BMC Bioinformatics 7, 55]. Here we describe a server, BiologicalNetworks, which provides visualization, analysis services and an information management framework over PathSys. The server allows easy retrieval, construction and visualization of complex biological networks, including genome-scale integrated networks of protein-protein, protein-DNA and genetic interactions. Most importantly, BiologicalNetworks addresses the need for systematic presentation and analysis of high-throughput expression data by mapping and analysis of expression profiles of genes or proteins simultaneously on to regulatory, metabolic and cellular networks. BiologicalNetworks Server is available at http://brak.sdsc.edu/pub/BiologicalNetworks.

Computer Graphics↗

Linking experimental results, biological networks and sequence analysis methods using Ontologies and Generalised Data Structures.

The structure of a closely integrated data warehouse is described that is designed to link different types and varying numbers of biological networks, sequence analysis methods and experimental results such as those coming from microarrays. The data schema is inspired by a combination of graph based methods and generalised data structures and makes use of ontologies and meta-data. The core idea is to consider and store biological networks as graphs, and to use generalised data structures (GDS) for the storage of further relevant information. This is possible because many biological networks can be stored as graphs: protein interactions, signal transduction networks, metabolic pathways, gene regulatory networks etc. Nodes in biological graphs represent entities such as promoters, proteins, genes and transcripts whereas the edges of such graphs specify how the nodes are related. The semantics of the nodes and edges are defined using ontologies of node and relation types. Besides generic attributes that most biological entities possess (name, attribute description), further information is stored using generalised data structures. By directly linking to underlying sequences (exons, introns, promoters, amino acid sequences) in a systematic way, close interoperability to sequence analysis methods can be achieved. This approach allows us to store, query and update a wide variety of biological information in a way that is semantically compact without requiring changes at the database schema level when new kinds of biological information is added. We describe how this datawarehouse is being implemented by extending the text-mining framework ONDEX to link, support and complement different bioinformatics applications and research activities such as microarray analysis, sequence analysis and modelling/simulation of biological systems. The system is developed under the GPL license and can be downloaded from http://sourceforge.net/projects/ondex/

Algorithms↗

Model-driven user interfaces for bioinformatics data resources: regenerating the wheel as an alternative to reinventing it.

BACKGROUND: The proliferation of data repositories in bioinformatics has resulted in the development of numerous interfaces that allow scientists to browse, search and analyse the data that they contain. Interfaces typically support repository access by means of web pages, but other means are also used, such as desktop applications and command line tools. Interfaces often duplicate functionality amongst each other, and this implies that associated development activities are repeated in different laboratories. Interfaces developed by public laboratories are often created with limited developer resources. In such environments, reducing the time spent on creating user interfaces allows for a better deployment of resources for specialised tasks, such as data integration or analysis. Laboratories maintaining data resources are challenged to reconcile requirements for software that is reliable, functional and flexible with limitations on software development resources. RESULTS: This paper proposes a model-driven approach for the partial generation of user interfaces for searching and browsing bioinformatics data repositories. Inspired by the Model Driven Architecture (MDA) of the Object Management Group (OMG), we have developed a system that generates interfaces designed for use with bioinformatics resources. This approach helps laboratory domain experts decrease the amount of time they have to spend dealing with the repetitive aspects of user interface development. As a result, the amount of time they can spend on gathering requirements and helping develop specialised features increases. The resulting system is known as Pierre, and has been validated through its application to use cases in the life sciences, including the PEDRoDB proteomics database and the e-Fungi data warehouse. CONCLUSION: MDAs focus on generating software from models that describe aspects of service capabilities, and can be applied to support rapid development of repository interfaces in bioinformatics. The Pierre MDA is capable of supporting common database access requirements with a variety of auto-generated interfaces and across a variety of repositories. With Pierre, four kinds of interfaces are generated: web, stand-alone application, text-menu, and command line. The kinds of repositories with which Pierre interfaces have been used are relational, XML and object databases.

Computational Biology↗

Electronic data collection options for practice-based research networks.

PURPOSE: We wanted to describe the potential benefits and problems associated with selected electronic methods of collecting data within practice-based research networks (PBRNs). METHODS: We considered a literature review, discussions with PBRN researchers, industry information, and personal experience. This article presents examples of selected PBRNs' use of electronic data collection. RESULTS: Collecting research data in the geographically dispersed PBRN environment requires considerable coordination to ensure completeness, accuracy, and timely transmission of the data, as well as a limited burden on the participants. Electronic data collection, particularly at the point of care, offers some potential solutions. Electronic systems allow use of transparent decision algorithms and improved data entry and data integrity. These systems may improve data transfer to the central office as well as tracking systems for monitoring study progress. PBRNs have available to them a wide variety of electronic data collection options, including notebook computers, tablet PCs, personal digital assistants (PDAs), and browser-based systems that operate independent of or over the Internet. Tablet PCs appear particularly advantageous for direct patient data collection in an office environment. PDAs work well for collecting defined data elements at the point of care. Internet-based systems work well for data collection that can be completed after the patient visit, as most primary care offices do not support Internet connectivity in examination rooms. CONCLUSIONS: When planning to collect data electronically, it is important to match the electronic data collection method to the study design. Focusing an inappropriate electronic data collection method onto users can interfere with accurate data gathering and may also anger PBRN members.

Biomedical Research↗

Munich information center for protein sequences plant genome resources: a framework for integrative and comparative analyses 1(W).

With several plant genomes sequenced, the power of comparative genome analysis can now be applied. However, genome-scale cross-species analyses are limited by the effort for data integration. To develop an integrated cross-species plant genome resource, we maintain comprehensive databases for model plant genomes, including Arabidopsis (Arabidopsis thaliana), maize (Zea mays), Medicago truncatula, and rice (Oryza sativa). Integration of data and resources is emphasized, both in house as well as with external partners and databases. Manual curation and state-of-the-art bioinformatic analysis are combined to achieve quality data. Easy access to the data is provided through Web interfaces and visualization tools, bulk downloads, and Web services for application-level access. This allows a consistent view of the model plant genomes for comparative and evolutionary studies, the transfer of knowledge between species, and the integration with functional genomics data.

Computational Biology↗

Functional annotation and network reconstruction through cross-platform integration of microarray data.

The rapid accumulation of microarray data translates into a need for methods to effectively integrate data generated with different platforms. Here we introduce an approach, 2(nd)-order expression analysis, that addresses this challenge by first extracting expression patterns as meta-information from each data set (1(st)-order expression analysis) and then analyzing them across multiple data sets. Using yeast as a model system, we demonstrate two distinct advantages of our approach: we can identify genes of the same function yet without coexpression patterns and we can elucidate the cooperativities between transcription factors for regulatory network reconstruction by overcoming a key obstacle, namely the quantification of activities of transcription factors. Experiments reported in the literature and performed in our lab support a significant number of our predictions.

Algorithms↗

[Methodology for the elaboration of an Organizational Perceptions Index].

This article presents the methodological basis for elaborating an Organizational Perceptions Index. The tool allows one to grasp the perceptions of organization members and solve some problems identified in the process. It is assumed that perceptions are the result of individual and collective subjectivities, the latter referring to culturally-bounded, socially-constructed representations. The Organizational Perceptions Index expresses the positive or negative perceptions of organization members and integrates data related to four organizational dimensions: infra-structure, management, environment, and culture. The methodology integrates research data in an evaluation tool, with the aim of orienting interventions and monitoring changes over time.

Attitude of Health Personnel↗

Understanding the yeast proteome: a bioinformatics perspective.

Rapid development of genomic and proteomic methodologies has provided a wealth of data for deciphering the biomolecular circuitry of a living cell. The main areas of computational research of proteomes outlined in this review are: understanding the system, its features and parameters to help plan the experiments; data integration, to help produce more reliable data sets; visualization and other forms of data representation to simplify interpretation; modeling of the functional regulation; and systems biology. With false-positive rates reaching 50% even in the more reliable data sets, handling the experimental error remains one of the most challenging tasks. Integrative approaches, incorporating results of various genome- and proteome-wide experiments, allow for minimizing the error and bring with them significant predictive power.

Computational Biology↗

Identification of contrastive and comparable school neighborhoods for childhood obesity and physical activity research.

UNLABELLED: The neighborhood social and physical environments are considered significant factors contributing to children's inactive lifestyles, poor eating habits, and high levels of childhood obesity. Understanding of neighborhood environmental profiles is needed to facilitate community-based research and the development and implementation of community prevention and intervention programs. We sought to identify contrastive and comparable districts for childhood obesity and physical activity research studies. We have applied GIS technology to manipulate multiple data sources to generate objective and quantitative measures of school neighborhood-level characteristics for school-based studies. GIS technology integrated data from multiple sources (land use, traffic, crime, and census tract) and available social and built environment indicators theorized to be associated with childhood obesity and physical activity. We used network analysis and geoprocessing tools within a GIS environment to integrate these data and to generate objective social and physical environment measures for school districts. We applied hierarchical cluster analysis to categorize school district groups according to their neighborhood characteristics. We tested the utility of the area characterizations by using them to select comparable and contrastive schools for two specific studies. RESULTS: We generated school neighborhood-level social and built environment indicators for all 412 Chicago public elementary school districts. The combination of GIS and cluster analysis allowed us to identify eight school neighborhoods that were contrastive and comparable on parameters of interest (land use and safety) for a childhood obesity and physical activity study. CONCLUSION: The combination of GIS and cluster analysis makes it possible to objectively characterize urban neighborhoods and to select comparable and/or contrasting neighborhoods for community-based health studies.

Adolescent↗

An integrative genomic approach to uncover molecular mechanisms of prokaryotic traits.

With mounting availability of genomic and phenotypic databases, data integration and mining become increasingly challenging. While efforts have been put forward to analyze prokaryotic phenotypes, current computational technologies either lack high throughput capacity for genomic scale analysis, or are limited in their capability to integrate and mine data across different scales of biology. Consequently, simultaneous analysis of associations among genomes, phenotypes, and gene functions is prohibited. Here, we developed a high throughput computational approach, and demonstrated for the first time the feasibility of integrating large quantities of prokaryotic phenotypes along with genomic datasets for mining across multiple scales of biology (protein domains, pathways, molecular functions, and cellular processes). Applying this method over 59 fully sequenced prokaryotic species, we identified genetic basis and molecular mechanisms underlying the phenotypes in bacteria. We identified 3,711 significant correlations between 1,499 distinct Pfam and 63 phenotypes, with 2,650 correlations and 1,061 anti-correlations. Manual evaluation of a random sample of these significant correlations showed a minimal precision of 30% (95% confidence interval: 20%-42%; n = 50). We stratified the most significant 478 predictions and subjected 100 to manual evaluation, of which 60 were corroborated in the literature. We furthermore unveiled 10 significant correlations between phenotypes and KEGG pathways, eight of which were corroborated in the evaluation, and 309 significant correlations between phenotypes and 166 GO concepts evaluated using a random sample (minimal precision = 72%; 95% confidence interval: 60%-80%; n = 50). Additionally, we conducted a novel large-scale phenomic visualization analysis to provide insight into the modular nature of common molecular mechanisms spanning multiple biological scales and reused by related phenotypes (metaphenotypes). We propose that this method elucidates which classes of molecular mechanisms are associated with phenotypes or metaphenotypes and holds promise in facilitating a computable systems biology approach to genomic and biomedical research.

Algorithms↗

Defining, measuring, and predicting impulsive aggression: a heuristic model.

Aggression research does not lack data--it lacks a model for integrating data. One of the problems confronting aggression researchers is the extensive body of multidisciplinary data that is difficult to synthesize to generate new directions in research. This paper proposes one solution that starts by asking "what is the minimal number of categories of concepts and measurements which are necessary to describe a person?". The answer is four categories of concepts: biological; cognitive; behavioral; environmental (physical and social). One way of many for integrating these four categories of concepts is a proposed discipline neutral heuristic model that is used herein to compare two different research approaches to the study of impulsive aggression. This comparison identifies clearly the differences in the two approaches with regard to different emphases among the four categories of constructs for each program. Using the model an example of common ground between the two approaches is sought as a basis for extending aggression research. The main conclusion of one of the research programs was that central nervous arousal is related to impulsive aggression. This program demonstrated that phenytoin will reduce impulsive aggressive acts and has an effect on CNS arousal. The other research program on impulsive aggression has been at the forefront in demonstrating the well established inverse relationship between serotonin levels and aggression. The comparison resulted in the suggestion that both serotonin and phenytoin may relate to a common neurochemical substrate which interacts in part to control CNS arousal, especially at the cortical level. The proposed heuristic model made obvious the need to use synthesizing concepts (e.g. information processing or language) which can interrelate multidisciplinary concepts and data from different research programs within the four categories of constructs when comparing interdisciplinary research.

Aggression↗

Biodiversity informatics.

Biodiversity informatics is an emerging field that applies information management tools to the management and analysis of species-occurrence, taxonomic character, and image data. A wide and growing range of tools is available for both curators and researchers. The development and implementation of formal data exchange standards and query protocols have made it possible to integrate data holdings from collections around the world. The current technological environment is summarized; protocols, standards, and tools for data management, sharing, and integration are reviewed; and methods and tools for analyzing species-occurrence and character data are examined. Direct access to primary data and imagery has the power to transform the means by which taxonomy is practiced and its results disseminated to the general community.

Animals↗