Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data integration”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Atlas - a data warehouse for integrative bioinformatics.

BACKGROUND: We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for bioinformatics research and development. DESCRIPTION: The Atlas system is based on relational data models that we developed for each of the source data types. Data stored within these relational models are managed through Structured Query Language (SQL) calls that are implemented in a set of Application Programming Interfaces (APIs). The APIs include three languages: C++, Java, and Perl. The methods in these API libraries are used to construct a set of loader applications, which parse and load the source datasets into the Atlas database, and a set of toolbox applications which facilitate data retrieval. Atlas stores and integrates local instances of GenBank, RefSeq, UniProt, Human Protein Reference Database (HPRD), Biomolecular Interaction Network Database (BIND), Database of Interacting Proteins (DIP), Molecular Interactions Database (MINT), IntAct, NCBI Taxonomy, Gene Ontology (GO), Online Mendelian Inheritance in Man (OMIM), LocusLink, Entrez Gene and HomoloGene. The retrieval APIs and toolbox applications are critical components that offer end-users flexible, easy, integrated access to this data. We present use cases that use Atlas to integrate these sources for genome annotation, inference of molecular interactions across species, and gene-disease associations. CONCLUSION: The Atlas biological data warehouse serves as data infrastructure for bioinformatics research and development. It forms the backbone of the research activities in our laboratory and facilitates the integration of disparate, heterogeneous biological sources of data enabling new scientific inferences. Atlas achieves integration of diverse data sets at two levels. First, Atlas stores data of similar types using common data models, enforcing the relationships between data types. Second, integration is achieved through a combination of APIs, ontology, and tools. The Atlas software is freely available under the GNU General Public License at: http://bioinformatics.ubc.ca/atlas/

Computational Biology↗

Architectural design and tools to support the transparent access to hospital information systems, radiology information systems, and picture archiving and communication systems.

The fragmentation of the electronic patient record among hospital information systems (HIS), radiology information systems (RIS), and picture archiving and communication systems (PACS) makes the viewing of the complete medical patient record inconvenient. The purpose of this report is to describe the system architecture, development tools, and implementation issues related to providing transparent access to HIS, RIS, and PACS information. A client-mediator-server architecture was implemented to facilitate the gathering and visualization of electronic medical records from these independent heterogeneous information systems. The architecture features intelligent data access agents, run-time determination of data access strategies, and an active patient cache. The development and management of the agents were facilitated by data integration CASE (computer-assisted software engineering) tools. HIS, RIS, and PACS data access and translation agents were successfully developed. All pathology, radiology, medical, laboratory, admissions, and radiology reports for a patient are available for review from a single integrated workstation interface. A data caching system provides fast access to active patient data. New network architectures are evolving that support the integration of heterogeneous software subsystems. Commercial tools are available to assist in the integration procedure.

Computer Systems↗

ALFRED: An allele frequency database for anthropology.

The deluge of data from the human genome project (HGP) presents new opportunities for molecular anthropologists to study human variation through the promise of vast numbers of new polymorphisms (e.g., single nucleotide polymorphisms or SNPs). Collecting the resulting data into a single, easily accessible resource will be important to facilitate this research. We created a prototype Web-accessible database named ALFRED (ALelle FREquency Database, http://alfred.med.yale.edu/alfred/) to store and make publicly available allele frequency data on diverse polymorphic sites for many populations. In constructing this database, we considered many different concerns relating to the types of information needed for anthropology, population genetics, molecular genetics, and statistics, as well as issues of data integrity and ease of access to data. We also developed links to other Web-based databases as well as procedures for others to make links to the data in ALFRED. Here we present an overview of the issues considered and provisional solutions, as well as an example of data already available. It is our hope that this database will be useful for research and teaching in a wide range of fields, and that colleagues from various fields will contribute to making ALFRED an important resource for many studies as yet unforeseen.

Anthropology, Physical↗

ProteomeWeb: a web-based interface for the display and interrogation of proteomes.

The analysis of proteomes, i.e., the proteins expressed by biological organisms under a given set of conditions at a given time, requires separating complex protein mixtures into discrete protein components, measuring their relative abundances, and identifying the individual protein components. Many types of data are generated during the course of proteome analysis, including graphic images of the protein profiles, flat files containing numeric data, spreadsheets for assimilating numeric data, and relational database tables for integrating data from multiple experiments. As part of a project to describe the proteomes of microbes of interest to the U.S. Department of Energy, a World-Wide Web-based interface has been developed for the display of protein profiles generated by two-dimensional gel electrophoresis. The web interface is capable of obtaining protein identifications on the fly, interrogating the quantitative data in the context of available genome sequence information, and relating the proteome data to existing metabolic pathway databases. Analysis of protein expression profiles is expedited, providing the capability to efficiently determine the gene locations for proteins modulated in abundance in response to different growth conditions and to locate the positions of the proteins within specific metabolic pathways. The proteome of the archaeon Methanococcus jannaschii, a microbe for which the complete genome sequence is available, is used to demonstrate the capabilities of this evolving web interface (http://proteomeweb.anl.gov).

Amino Acid Sequence↗

Partial-filling micellar electrokinetic chromatography and non-aqueous capillary electrophoresis for the analysis of selected agrochemicals.

Selected agrochemicals (s-triazines and phenoxy acids) have been investigated with partial-filling micellar electrokinetic chromatography (PFMEKC) and non-aqueous capillary electrophoresis (NACE). Because these two techniques are compatible for coupling of capillary electrophoresis with mass spectrometry, different conditions affecting the separation efficiency (reproducibility, method linearity) were systematically tested, and the results were compared with those from classical MEKC. The conditions tested included buffer molarity, pH, the concentrations of the organic modifier and surfactant, the applied voltage, the injection time of the sample, and the length of the partial-filling plug. The respective limits of detection (LOD) using UV-detection were determined. Reduction of the electrophoretic raw data using the mobility scale transformation (micro-scale) improved qualitative comparison of the electropherograms and the reproducibility of quantitative data (integrated peak area) thus extending this data treatment from CZE to other endoosmotic flow-driven CE-techniques such as PFMEKC and NACE.

Journal Article↗

Implementation of a continuous quality improvement program for percutaneous coronary intervention and cardiac surgery at a large community hospital.

BACKGROUND: Continuous quality improvement (CQI) is widely used in other industries and has been promoted as a method for quality control in medicine. The national databases developed by the American College of Cardiology and the Society of Thoracic Surgeons have greatly facilitated data collection for CQI. Hospitals can encounter barriers to CQI, however, which include creating the proper organizational infrastructure and engaging physicians and hospital administrators in the process. These barriers are particularly evident in large community hospitals. METHODS: We describe the organizational infrastructure for CQI, including committee structure, methods of repeated data collection and feedback, and maintenance of data integrity and confidentiality. We report demographic data and clinical outcomes for patients undergoing percutaneous coronary intervention and coronary artery bypass surgery before and after implementation of our CQI program. RESULTS: Since 1995, we have maintained a CQI process driven by repeated collection of valid, confidential, operator-specific data. We have observed sustained physician and administration participation and buy-in. During the follow-up period, patient complexity increased and observed outcomes improved, although the improvement was clearly multifactorial. CONCLUSIONS: We describe the organization of a CQI program at a large complex community hospital. Our CQI program was successfully implemented, has been sustained, and is associated in observed improvement in patient outcomes. The program described here may be a useful model for other similar hospitals that are attempting to create a program to address quality improvement.

Aged↗

Pathway information for systems biology.

Pathway information is vital for successful quantitative modeling of biological systems. The almost 170 online pathway databases vary widely in coverage and representation of biological processes, making their use extremely difficult. Future pathway information systems for querying, visualization and analysis must support standard exchange formats to successfully integrate data on a large scale. Such integrated systems will greatly facilitate the constructive cycle of computational model building and experimental verification that lies at the heart of systems biology.

Animals↗

GenBank.

The GenBank(R) sequence database (http://www.ncbi.nlm.nih.gov/) incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from individual laboratories and from large-scale sequencing projects. Most submitters use the BankIt (WWW) or Sequin programs to send their sequence data. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez , which integrates data from the major DNA and protein sequence databases along with taxonomy, genome and protein structure information. MEDLINE(R) abstracts from published articles describing the sequences are also included as an additional source of biological annotation. Sequence similarity searching is offered through the BLAST series of database search programs. In addition to FTP, e-mail and server/client versions of Entrez and BLAST, NCBI offers a wide range of World Wide Web retrieval and analysis services of interest to biologists.

Animals↗

GenBank.

The GenBank (Registered Trademark symbol) sequence database incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from individual laboratories and from large-scale sequencing projects. Most submitters use the BankIt (Web) or Sequin programs to format and send sequence data. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome and protein structure information. MEDLINE (Registered Trademark symbol) s from published articles describing the sequences are included as an additional source of biological annotation through the PubMed search system. Sequence similarity searching is offered through the BLAST series of database search programs. In addition to FTP, Email, and server/client versions of Entrez and BLAST, NCBI offers a wide range of World Wide Web retrieval and analysis services based on GenBank data. The GenBank database and related resources are freely accessible via the URL: http://www.ncbi.nlm.nih.gov

Amino Acid Sequence↗

GenBank.

The GenBank((R))sequence database incorporates publicly available DNA sequences of >55 000 different organisms, primarily through direct submission of sequence data from individual laboratories and large-scale sequencing projects. Most submissions are made using the BankIt (Web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping and protein structure information, plus the biomedical literature via PubMed. Sequence similarity searching is provided by the BLAST family of programs. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. NCBI also offers a wide range of WWW retrieval and analysis services based on GenBank data. The GenBank database and related resources are freely accessible via the NCBI home page at http://www.ncbi.nlm.nih.gov

Animals↗

GenBank.

The GenBank sequence database incorporates publicly available DNA sequences of more than 105 000 different organisms, primarily through direct submission of sequence data from individual laboratories and large-scale sequencing projects. Most submissions are made using the BankIt (web) or Sequin programs and accession numbers are assigned by GenBank staff upon receipt. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive worldwide coverage. GenBank data is accessible through NCBI's integrated retrieval system, Entrez, which integrates data from the major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and the biomedical literature via PubMed. Sequence similarity searching is provided by the BLAST family of programs. Complete bimonthly releases and daily updates of the GenBank database are available by FTP. NCBI also offers a wide range of World Wide Web retrieval and analysis services based on GenBank data. The GenBank database and related resources are freely accessible via the NCBI home page at http://www.ncbi.nlm.nih.gov.

Animals↗

Identification of neonatal hearing impairment: experimental protocol and database management.

OBJECTIVE: The purposes of this article are to describe the overall protocol for the Identification of Neonatal Hearing Impairment (INHI) project and to describe the management of the data collected as part of this project. A well-defined protocol and database management techniques were needed to ensure that data were 1) collected accurately and in the same way across sites; 2) maintained in a database that could be used to provide feedback to individual sites regarding enrollment and the extent to which the protocol was complete on individual subjects; and 3) available to answer project questions. This article describes techniques that were used to meet these needs. DESIGN: This study was a prospective, randomized study that was designed to evaluate auditory brain stem responses, transient evoked otoacoustic emissions, and distortion product otoacoustic emissions as hearing-screening tools, and to relate neonatal test findings to hearing status, defined by visual reinforcement audiometry at 8 to 12 mo of age. Measures of middle-ear function also were obtained at some sites as part of the neonatal test battery. In addition, other clinical and demographic data were gathered to determine the extent to which factors, other than auditory status, influenced test behavior. Three groups were evaluated: neonatal intensive care unit (NICU) infants (those who spent 3 or more days in a NICU), well babies with risk factors for hearing loss, and well babies without risk factors. Six centers participated in the trial. The testers for the project included audiologists, technicians, audiology graduate students, and medical research staff. The same computerized neonatal test program was applied at each center. This program generated the neonatal test database automatically. Clinical and demographic data were collected by means of concise data collection forms and were entered into a database at each site. After the neonatal test, subjects from the NICU and at-risk well babies were evaluated with visual reinforcement audiometry starting at 8 to 12 mo of age. All data were electronically transmitted to the core site where they were merged into one overall database. This database was exercised to provide feedback and to identify discrepancies throughout the course of the study. In its final form, it served as the database on which all analyses were performed. RESULTS AND CONCLUSION: The protocol was a departure from typical hearing screening procedures in terms of 1) its regimented application of three screening measures; 2) the detailed information that was obtained regarding subject clinical and demographic factors; and 3) its application of the same procedures across six centers having diverse geographic location and subject demographics. A learning curve for successfully executing the study protocols was observed. Throughout the study, monthly reports were generated to monitor subject enrollment, check for data completeness, and to perform data integrity checks. In combination with monthly data reports and checks that occurred throughout the progression of the study, miscellaneous data audits were performed to check accuracy of neonatal testing programs and to cross-check information entered in the clinical and demographic database. The data management techniques used in this project helped to ensure the quality of the data collection process and also allowed for detailed analyses once data were collected. This was particularly important because it enabled us to evaluate not only the performance of individual measures as screening tools, but also permitted an evaluation of the influence of other variables on screening test results.

Acoustic Stimulation↗

An application of statistical matching with the survey of income and education and the 1976 Health Interview Survey.

This article outlines an alternative procedure to household surveys for obtaining individual observation-level data. The procedure, called statistical matching, integrates data on an individual observation from one source with data on a different observation identified as the "best matching" or "most similar" record from a second source. The best match is determined by objective statistical criteria. Also reported is a significant application of the procedure between the Survey of Income and Education and the 1976 National Health Interview Survey. The success of merging these two large, nationally representative data files shows statistical matching as a viable method of creating databases for health services research.

Adolescent↗

An automated clinic management system for a family planning network.

The medical information, financial, and logistic aspects of a comprehensive computer-based Appointment, Registration, Information System, and Evaluation (ARISE) are analyzed for the management of a family planning program serving 30,000 patients annually. An overview of the existing computer system network is presented with descriptions of the interactive master patient index, the batch appointment process, the management statistics package, and Department of Health, Education, and Welfare (HEW) reporting. Emphasis is placed on the financial management control system which includes 1) procedures for third-party submission of claims for payment, in particular Titles IVA, XX, and XIX (Social Security Act), together with discussion of related administrative requirements; 2) technics of auditing data integrity including systematic sampling of collected data; and 3) the process of billing and receipts collection. Methodology and implementation aspects of ARISE may have wide applicability to other family planning and similarly structured clinical programs.

Computers↗

Selection of biomarkers by a multivariate statistical processing of composite metabonomic data sets using multiple factor analysis.

We introduce a statistical approach for integrating data from several analytical platforms. We illustrate this approach using (1)H-(13)C Heteronuclear Multiple Bond Connectivity nuclear magnetic resonance spectroscopy ((1)H-(13)C HMBC NMR) and Pyrolysis Metastable Atom Bombardment Time-of-Flight mass spectrometry (Py-MAB-TOF-MS) to perform metabolic fingerprinting on cattle treated with anabolic steroids. Multiple factor analysis (MFA) integrates complementary aspects from NMR and MS data into a unique metabolic signature describing the biomarkers related to the dose-response. This work also indicates that, from a practical point of view, metabonomics and other "-omics" biotechnologies can benefit significantly from a generalized multi-platform integrative approach using multiple factor analysis.

Animals↗

An XML message broker framework for exchange and integration of microarray data.

MOTIVATION: Microarrays are an important research tool for the advancement of basic biological sciences. However this technology has yet to be integrated with clinical decision making. We have implemented an information framework based on the Microarray Gene Expression Markup Language (MAGE-ML) specification. We are using this framework to develop a test-bed integrated database application to identify genomic and imaging markers for diagnosis of breast cancer. RESULTS: We developed extensible software architecture for retrieving data from different microarray databases using MAGE-ML and for combining microarray data with breast cancer image analysis and clinical data for correlation studies. The framework we developed will provide the necessary data integration to move microarray research from basic biological sciences to clinical applications. AVAILABILITY: Open source software will be available from SourceForge (http://sourceforge.net/projects/microsoap/).

Database Management Systems↗

An integrated analysis of thirteen trials summarizing the long-term safety of alefacept in psoriasis patients who have received up to nine courses of therapy.

BACKGROUND: Information on longer-term safety and tolerability is needed to confidently prescribe alefacept therapy for chronic plaque psoriasis beyond 1 or 2 courses. OBJECTIVE: The aim of this work was to further examine the safety profile of alefacept by integrating data from clinical trials involving patients with chronic plaque psoriasis who received up to 9 courses of therapy over a 5-year period. METHODS: Data from 13 clinical trials conducted in patients with plaque psoriasis were integrated because they had similar inclusion/exclusion criteria and assessments. Patients who enrolled in the analyzed trials were aged > or =15 years with chronic plaque psoriasis for > or =12 months that involved > or =10% of body surface area, and CD4+ T lymphocyte counts above the lower limit of normal (>404 cells/microL). The incidences of adverse events (AEs), serious AEs, discontinuations for AEs, infections, serious infections, malignancies, and anti-alefacept antibodies were summarized for each course of alefacept. The incidence of infections was stratified according to CD4+ T lymphocyte counts (<250 cells/microL vs > or =250 cells/microL). RESULTS: Data from 13 clinical trials of alefacept were integrated and summarized (multicenter, randomized, double-blind studies, n = 6; multicenter, open-label studies, n = 5; other, n = 2). The analyzed population (n = 1869) included 1291 (69.1%) men and 578 (30.9%) women, between the ages of 15 and 84 years (mean, 44.8 years), of whom 1648 (88.2%) were white. Weights ranged from 40 kg to 206 kg (mean, 90.0 kg). A total of 1369 of these patients had been included in a previous analysis. Among the most commonly reported AEs in each treatment course were headache (0%-14.2%), nasopharyngitis (7.7%-25.0%), influenza (0%-8.1%), upper respiratory tract infection (0%-12.5%), and pruritus (0%-7.5%). The rates of discontinuations due to AEs (0%-4.8%), serious AEs (0%-4.8%), serious infections (0%-0.9%), or malignancies (0%-4.8%) did not appear to increase with repeated exposure. Fewer than 1 % of patients in each course developed a serious infection. No opportunistic infections or infection-related deaths were reported. The incidence of infections appeared to be unrelated to CD4+ T lymphocyte counts. Fewer than 2.5% of patients tested positive for anti-alefacept antibodies during any course of therapy. CONCLUSIONS: This integrated analysis of data from 13 trials with 1869 patients supports the safety and tolerability of alefacept for longer-term treatment of psoriasis.

Adolescent↗

The role of organizational infrastructure in implementation of hospitals' quality improvement.

Quality improvement (QI) is an organized approach to planning and implementing continuous improvement in performance. Although QI holds promise for improving quality of care and patient safety, hospitals that adopt QI often struggle with its implementation. This article examines the role of organizational infrastructure in implementation of quality improvement practices and structures in hospitals. The authors focus specifically on four elements of hospital support and infrastructure for QI-integrated data systems, financial support for QI, clinical integration, and information system capability. These macrolevel factors provide consistent, ongoing support for the QI efforts of clinical teams engaging in direct patient care, thus promoting institutionalization of QI. Results from the multivariate analysis of 1997 survey data on 2350 hospitals provide strong support for the hypotheses. Results signal that organizations intent upon improving quality must attend to the context in which QI efforts are practiced, and that such efforts are unlikely to be effective unless appropriate support systems are in place to ensure full implementation.

Data Collection↗