Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Prediction of in vivo drug clearance from in vitro data. II: potential inter-ethnic differences.

Potential differences in drug clearance between Japanese and Caucasians were investigated by integrating data on demography, liver size, the abundance of the major cytochromes P450 and in vitro metabolic parameters. Eleven drugs (alprazolam, caffeine, chlorzoxazone, cyclosporine, midazolam, omeprazole, sildenafil, tolbutamide, triazolam, S-warfarin and zolpidem) fulfilled the entry criteria of the study (i.e. the necessary in vitro metabolism data were available and clearance values had been reported both in Caucasians and Japanese). Values of relevant biological variables were obtained from the literature, and clearance predictions were made using the Simcyp Population-Based ADME Simulator. The ratios of observed oral clearance (CLp.o.) values in Caucasians compared with Japanese ranged from 0.6 to 2.8 (integrating data from 82 sources). The CLp.o. values for alprazolam, caffeine and zolpidem were not statistically different between Caucasian and Japanese (p>0.05), whereas those for chorzoxazone, cyclosporine, omeprazole, tolbutamide and triazolam were higher in Caucasians (p<0.05), and those for midazolam, sildenafil and S-warfarin were higher in Japanese (p<0.05). CLp.o. values, predicted from in vitro data, were within 3-fold of observed in vivo values for seven of the 11 drugs in Japanese. Values for the predicted ratios ranged from 1.6 to 4.9. The predicted ratios were not significantly different from observed ratios for cyclosporine, omeprazole, tolbutamide and triazolam. Only partial success in predicting ethnic differences in clearance indicates the need for larger and more reliable databases on relevant variables. With such information, in silico predictions might be used with more confidence to decrease the need for repeating pharmacokinetic studies in different ethnic groups.

Asian People↗

PQL: a declarative query language over dynamic biological schemata.

We introduce the PQL query language (PQL) used in the GeneSeek genetic data integration project. PQL incorporates many features of query languages for semi-structured data. To this we add the ability to express metadata constraints like intended semantics and database curation approach. These constraints guide the dynamic generation of potential query plans. This allows a single query to remain relevant even in the presence of source and mediated schemas that are continually evolving, as is often the case in data integration.

Computational Biology↗

Management information systems--green light for better info.

An EIS gathers financial and non-financial data from a variety of sources, both internal and external to an organisation, and presents them accessibly and understandably. It should facilitate the presentation of data so that it is timely and relevant to senior managers' needs. Generally, there are three categories of EIS: Front-end tools, which enhance the presentation of output from existing systems. For example, they can take the output from a general ledger system and enhance its appearance but do not change the content. Internal consolidation tools, which take data from a number of internal sources (such as the general ledger), store it in a central database, and present it to users. This is often done in report book format to which all EIS style functions may be applied. Integrators of information, which integrate data from internal and external sources. Much of the data is non-financial and a major emphasis is on the users' ability to communicate with each other. The key functions that may be included within an EIS are: Drill down. The facility to explore increasingly detailed levels of data. Trends and variances against pre-set targets, such as financial budgets. Graphics and tabular reporting. Data integrity checking. Analysis of the data, modelling and the production of forecasts using time series analysis techniques. Exception reporting through the use of some form of alert. Incorporation of text into the output.

Diagnosis-Related Groups↗

Scriptable access to the Caenorhabditis elegans genome sequence and other ACEDB databases.

Much of the world's genomic data are available to the community through networked databases that are accessed via Web interfaces. Although this paradigm provides browse-level access and has greatly facilitated linking between databases, it does not provide any convenient mechanism for programmatically fetching and integrating data from diverse databases. We have created a library and an application programming interface (API) named AcePerl that provides simple, direct access to ACEDB databases from the Perl programming language. With this library, programmers and computer-savvy biologists can write software to pose complex queries on local and remote ACEDB databases, retrieve the data, integrate the results, and move data objects from one database to another. In addition, a set of Web scripts running on top of AcePerl provides Web-based browsing of any local or remote ACEDB database. AcePerl and the AceBrowser Web browser run on Unix systems and are available under a license that allows for unrestricted use and redistribution. Both packages can be downloaded from URL. A Microsoft Windows port of AcePerl is in the planning stages.

Animals↗

Congruence of tissue expression profiles from Gene Expression Atlas, SAGEmap and TissueInfo databases.

BACKGROUND: Extracting biological knowledge from large amounts of gene expression information deposited in public databases is a major challenge of the postgenomic era. Additional insights may be derived by data integration and cross-platform comparisons of expression profiles. However, database meta-analysis is complicated by differences in experimental technologies, data post-processing, database formats, and inconsistent gene and sample annotation. RESULTS: We have analysed expression profiles from three public databases: Gene Expression Atlas, SAGEmap and TissueInfo. These are repositories of oligonucleotide microarray, Serial Analysis of Gene Expression and Expressed Sequence Tag human gene expression data respectively. We devised a method, Preferential Expression Measure, to identify genes that are significantly over- or under-expressed in any given tissue. We examined intra- and inter-database consistency of Preferential Expression Measures. There was good correlation between replicate experiments of oligonucleotide microarray data, but there was less coherence in expression profiles as measured by Serial Analysis of Gene Expression and Expressed Sequence Tag counts. We investigated inter-database correlations for six tissue categories, for which data were present in the three databases. Significant positive correlations were found for brain, prostate and vascular endothelium but not for ovary, kidney, and pancreas. CONCLUSION: We show that data from Gene Expression Atlas, SAGEmap and TissueInfo can be integrated using the UniGene gene index, and that expression profiles correlate relatively well when large numbers of tags are available or when tissue cellular composition is simple. Finally, in the case of brain, we demonstrate that when PEM values show good correlation, predictions of tissue-specific expression based on integrated data are very accurate.

Brain↗

[Internet communication between family physicians and the university hospital].

The potential of electronic communication in medicine is assessed based on an analysis of a pilot project pertaining to internet based communication among referring and hospital physicians. Advantages of electronic data exchange in medicine pertain to speed and capacity for data transfer, availability of data and data integration, ultimately enabling consistent medical case management. Quality requirements of electronic communication of medical data are related to safety, availability, data integration, potential for case management and system qualities. Medical efficiency can be increased by use of electronic communication only if complex functions beyond the substitution of conventional mail by e-mail are implemented and an exhaustive use of the technology can be achieved.

Case Management↗

Integrating genomic data to predict transcription factor binding.

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

Algorithms↗

Innovations in neonatal case management: an integrated, data-driven approach.

The intent of this article is to provide one company's perspective on the challenging and complex care management of the high-risk neonate. The strategies presented herein should enable and encourage case managers to implement an integrated management process for the frail neonatal population.

Case Management↗

Can we integrate bioinformatics data on the Internet?

The NETTAB (Network Tools and Applications in Biology) 2001 Workshop entitled 'CORBA and XML: towards a bioinformatics-integrated network environment' was held at the Advanced Biotechnology Centre, Genoa, Italy, 17-18 May 2001.

Computational Biology↗

Protecting participants in family medicine research: a consensus statement on improving research integrity and participants' safety in educational research, community-based participatory research, and practice network research.

Recent events that include the deaths of research subjects and the falsification of data have drawn greater scrutiny on assuring research data integrity and protecting participants. Several organizations have created guidelines to help guide researchers working in the area of clinical trials and ensure that their research is safe and valid. However, family medicine researchers often engage in research that differs from a typical clinical trial. Investigators working in the areas of educational research, community-based participatory research, and practice-based network research would benefit from similar recommendations to guide their own research. With funding from the US Office of Research Integrity and the Association of American Medical Colleges, we convened a panel to review issues of data integrity and participant protection in educational research, community-based participatory research, and research conducted by practice-based networks. The panel generated 11 recommendations for researchers working in these areas. Three key recommendations include the need for (1) all educational research to undergo review and approval by an institutional review board (IRB), (2) community-based participatory research to be approved not just by an IRB but also by appropriate community representatives, and (3) practice-based researchers to undertake only valid and meaningful studies that can be reviewed by a central IRB, rather than separate IRBs for each participating practice.

Biomedical Research↗

Collection and analysis of intake data from the integrated survey.

Intake data from the combined CSFII/NHANES survey will be used for many different purposes, each with specific data requirements and appropriate analytic methods. For monitoring and surveillance, the availability of Dietary Reference Intakes will allow estimates of the prevalence of inadequate intakes and the prevalence of intakes with a risk of adverse effects. The accuracy of the nutrient intake estimates will be enhanced by the 5-pass dietary recall methodology, availability of quantified dietary supplement intake data and expanded food and supplement composition data. Food-level dietary monitoring will be improved by using new databases to calculate servings of food groups from the Food Guide Pyramid and intakes of food commodities. Another major strength of the survey is the ability to relate intake data to health measures for individuals. Inferences will continue to be limited by a lack of usual intake for each individual, but the attenuation will be less with 2 d of data than with only 1 d, as in the past. Better data collection and analysis will also lead to more informed nutrition policies and programs. Innovative methods of analyzing the data should be investigated to minimize the effects of underreporting, provide better estimates of usual intake at both the group and individual levels and accurately combine nutrient intakes from foods and supplements. Future modifications to the intake collection methods might be considered to allow larger sample sizes for certain subgroups, more detailed information on supplement use, an expanded food frequency questionnaire, a different number of recall days and incorporation of diet and health knowledge questions.

Child↗

Integration of data driven decision support into the HELIOS environment.

The development of large-scale, clinically accepted decision support systems (DSS) calls for powerful and commonly available methods and tools for knowledge acquisition, system realisation, and knowledge base maintenance. The paper addresses problems associated with the integration of knowledge-based systems within the clinical setting with special reference to (i) data driven decision support, (ii) the Arden Syntax as a knowledge representation format and, (iii) the HELIOS software engineering environment. Architecture of a DSS based on Arden Syntax and its integration in the HELIOS environment are presented. Realisation of the DSS is discussed in relation to client-server architecture and object-oriented databases, which are essential concepts of the HELIOS environment. Sharability and reusability of the knowledge, together with commonality of used software tools are also discussed.

Database Management Systems↗

iProClass: an integrated database of protein family, function and structure information.

The iProClass database provides comprehensive, value-added descriptions of proteins and serves as a framework for data integration in a distributed networking environment. The protein information in iProClass includes family relationships as well as structural and functional classifications and features. The current version consists of about 830 000 non-redundant PIR-PSD, SWISS-PROT, and TrEMBL proteins organized with more than 36 000 PIR superfamilies, 145 000 families, 4000 domains, 1300 motifs and 550 000 FASTA similarity clusters. It provides rich links to over 50 database of protein sequences, families, functions and pathways, protein-protein interactions, post-translational modifications, protein expressions, structures and structural classifications, genes and genomes, ontologies, literature and taxonomy. Protein and superfamily summary reports present extensive annotation information and include membership statistics and graphical display of domains and motifs. iProClass employs an open and modular architecture for interoperability and scalability. It is implemented in the Oracle object-relational database system and is updated biweekly. The database is freely accessible from the web site at http://pir.georgetown.edu/iproclass/ and searchable by sequence or text string. The data integration in iProClass supports exploration of protein relationships. Such knowledge is fundamental to the understanding of protein evolution, structure and function and crucial to functional genomic and proteomic research.

Amino Acid Motifs↗

Informational aspects of telepathology in routine surgical pathology.

Application of computer and telecommunication technology calls serious challenges in routine diagnostic pathology. Complete data integration, fast access patients' data to usage of diagnosis thesaurus labeled with standardized codes and free text supplements, complex inquiry of the data contents, data exchange via teleconsultation and multilevel data protection are required functions of an integrated information system. Increasing requirement for teleconsultation transferring a large amount of multimedia data among different pathology information systems raises new questions in telepathology. Creation of complex telematic systems in pathology requires efficient methods of software engineering and implementation. Information technology of object-oriented modeling, usage of client server architecture and relational database management systems enables more compatible systems in field of telepathology. The aim of this paper is to present a practical example how to unify text based database, image archive and teleconsultation in a frame of an integrated telematic system and to discuss the main conceptual questions of information technology of telepathology.

Humans↗

Ontological integration of data models for cell signaling pathways by defining a factor of causality called 'signal'.

Databases have collected masses of information concerning cell signaling pathways that includes information on pathways, molecular interactions as well as molecular complexes. However we have no general data model to represent comprehensive properties of cell signaling pathways, so that this type of information has been represented by two different data models that we call 'binary relation' and 'state transition'. The disagreement between the existing models derives from lack of consensus about a factor of causality in reactions in cell signaling pathways, which is often called 'signal'. We developed an ontology named CSNO (Cell Signaling Networks Ontology) based on device ontology. As device ontology is a research product of knowledge engineering, CSNO is the first application of it to biological knowledge. CSNO defines the factor of causality called 'signal', offers an integrative viewpoint for the two different data models, explicates intrinsic distinctions between signaling and metabolic pathways, and eliminates ambiguity from representation of complex molecules.

Cell Physiological Phenomena↗

Health care fraud and abuse data collection program: technical revisions to Healthcare Integrity and Protection Data Bank data collection activities. Final rule.

The rule finalizes technical changes to the Healthcare Integrity and Protection Data Bank (HIPDB) data collection reporting requirements by clarifying the types of personal numeric identifiers that may be reported to the data bank in connection with adverse actions. The rule clarifies that in lieu of a Social Security Number (SSN), an individual taxpayer identification number (ITIN) may be reported to the data bank when, in those limited situations, an individual does not have an SSN.

Credentialing↗

Using forest health monitoring data to integrate above and below ground carbon information.

The national Forest Health Monitoring (FHM) program conducted a remeasurement study in 1999 to evaluate the usefulness and feasibility of collecting data needed for investigating carbon budgets in forests. This study indicated that FHM data are adequate for detecting a 20% change over 10 years (2% change per year) in percent total carbon and carbon content (MgC/ha) when sampling by horizon, with greater than 80% probability that a change in carbon content will be determined when a change has truly occurred (P < or = 0.33). The data were also useful in producing estimates of forest floor and soil carbon stocks by depth that were somewhat lower than literature values used for comparison. The scale at which the data were collected lends itself to producing standing stock estimates needed for carbon budget development and carbon cycle modeling. The availability of site-specific forest mensuration data enables the exploration of above ground and below ground linkages.

Biomass↗