Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reference database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases. Recent developments of the database include: cross-references to additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT. TrEMBL consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except the CDS already included in SWISS-PROT. The URLs for SWISS-PROT on the WWW are: http://www.expasy.ch/sprot and http://www. ebi.ac.uk/sprot

Amino Acid Sequence↗

Design of a Multi Dimensional Database for the Archimed DataWarehouse.

The Archimed data warehouse project started in 1993 at the Geneva University Hospital. It has progressively integrated seven data marts (or domains of activity) archiving medical data such as Admission/Discharge/Transfer (ADT) data, laboratory results, radiology exams, diagnoses, and procedure codes. The objective of the Archimed data warehouse is to facilitate the access to an integrated and coherent view of patient medical in order to support analytical activities such as medical statistics, clinical studies, retrieval of similar cases and data mining processes. This paper discusses three principal design aspects relative to the conception of the database of the data warehouse: 1) the granularity of the database, which refers to the level of detail or summarization of data, 2) the database model and architecture, describing how data will be presented to end users and how new data is integrated, 3) the life cycle of the database, in order to ensure long term scalability of the environment. Both, the organization of patient medical data using a standardized elementary fact representation and the use of the multi dimensional model have proved to be powerful design tools to integrate data coming from the multiple heterogeneous database systems part of the transactional Hospital Information System (HIS). Concurrently, the building of the data warehouse in an incremental way has helped to control the evolution of the data content. These three design aspects bring clarity and performance regarding data access. They also provide long term scalability to the system and resilience to further changes that may occur in source systems feeding the data warehouse.

Databases, Factual↗

A summarization approach for Affymetrix GeneChip data using a reference training set from a large, biologically diverse database.

BACKGROUND: Many of the most popular pre-processing methods for Affymetrix expression arrays, such as RMA, gcRMA, and PLIER, simultaneously analyze data across a set of predetermined arrays to improve precision of the final measures of expression. One problem associated with these algorithms is that expression measurements for a particular sample are highly dependent on the set of samples used for normalization and results obtained by normalization with a different set may not be comparable. A related problem is that an organization producing and/or storing large amounts of data in a sequential fashion will need to either re-run the pre-processing algorithm every time an array is added or store them in batches that are pre-processed together. Furthermore, pre-processing of large numbers of arrays requires loading all the feature-level data into memory which is a difficult task even with modern computers. We utilize a scheme that produces all the information necessary for pre-processing using a very large training set that can be used for summarization of samples outside of the training set. All subsequent pre-processing tasks can be done on an individual array basis. We demonstrate the utility of this approach by defining a new version of the Robust Multi-chip Averaging (RMA) algorithm which we refer to as refRMA. RESULTS: We assess performance based on multiple sets of samples processed over HG U133A Affymetrix GeneChip arrays. We show that the refRMA workflow, when used in conjunction with a large, biologically diverse training set, results in the same general characteristics as that of RMA in its classic form when comparing overall data structure, sample-to-sample correlation, and variation. Further, we demonstrate that the refRMA workflow and reference set can be robustly applied to naïve organ types and to benchmark data where its performance indicates respectable results. CONCLUSION: Our results indicate that a biologically diverse reference database can be used to train a model for estimating probe set intensities of exclusive test sets, while retaining the overall characteristics of the base algorithm. Although the results we present are specific for RMA, similar versions of other multi-array normalization and summarization schemes can be developed.

Algorithms↗

Health technology assessment in social care: a case study of randomized controlled trial retrieval.

OBJECTIVES: The aim of this study was to evaluate the success of search strategies in retrieving key documents for a technology assessment report (TAR) on a social care topic. METHODS: This study measured the differential yield of relevant studies from various information sources and evaluated strategies in different databases, with particular reference to capturing randomized controlled trials (RCTs) as a study design. RESULTS: A combination of four major databases would have found all thirty-two key references. One database alone would have found 78 percent, with another two each locating 59 percent. Sixteen percent of the trials were unique references. In non-health care databases, more sensitive search strategies would have resulted in a higher yield of relevant studies, in part due to inconsistent indexing and in part to attempts to restrict searches to RCTs. Although additional terms could be used to increase the sensitivity of the original strategies, this raises the question of trading off time against exhaustiveness, given the greater number of irrelevant references likely to be retrieved. CONCLUSIONS: A successful search for evidence on this social care topic would be possible using a combination of MEDLINE, EMBASE, the Cochrane Library and PsyclNFO, supplemented by only limited use of supplementary databases. In areas such as social care where evidence-based research is not yet well established, attempts to replicate searches based on study design do not seem to be advisable, although this may be an area for future research.

Databases, Bibliographic↗

FragMatch--a program for the analysis of DNA fragment data.

FragMatch is a user-friendly Java-supported program that automates the identification of taxa present in mixed samples by comparing community DNA fragment data against a database of reference patterns for known species. The program has a user-friendly Windows interface and was primarily designed for the analysis of fragment data derived from terminal restriction fragment length polymorphism analysis of ectomycorrhizal fungal communities, but may be adapted for other applications such as microsatellite analyses. The program uses a simple algorithm to check for the presence of reference fragments within sample files that can be directly imported, and the results appear in a clear summary table that also details the parameters that were used for the analysis. This program is significantly more flexible than earlier programs designed for matching RFLP patterns as it allows default or user-defined parameters to be used in the analysis and has an unlimited database size in terms of both the number of reference species/individuals and the number of diagnostic fragments per database entry. Although the program has been developed with mycorrhizal fungi in mind, it can be used to analyse any DNA fragment data regardless of biological origin. FragMatch, along with a full description and users guide, is freely available to download from the Aberdeen Mycorrhiza Group web page (http://www.aberdeenmycorrhizas.com).

Algorithms↗

Allergen databases.

Allergies represent a significant medical and industrial problem. Molecular and clinical data on allergens are growing exponentially and in this article we have reviewed nine specialized allergen databases and identified data sources related to protein allergens contained in general purpose molecular databases. An analysis of allergens contained in public databases indicates a high level of redundancy of entries and a relatively low coverage of allergens by individual databases. From this analysis we identify current database needs for allergy research and, in particular, highlight the need for a centralized reference allergen database.

Allergens↗

IDEAS internal contamination database: a compilation of published internal contamination cases. A tool for the internal dosimetry community.

In the scope of the IDEAS project to develop General Guidelines for the Assessment of Internal Dose from Monitoring data, two databases were compiled. The IDEAS Bibliography database contains references dealing with problems related to cases of internal contamination. The IDEAS Internal Contamination Database now contains more than 200 cases of internal contamination. In the near future, the IDEAS Internal Contamination database will be made available to the internal dosimetry community. The database has several potential applications, including: training, testing biokinetic models, testing software for calculating intakes and doses from bioassay data, comparison of data from a new accidental intake with that from previous exposures to similar materials. The database is by no means complete, and this presentation is also an appeal for internal contamination cases to extend and update it.

Biological Assay↗

NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.

The National Center for Biotechnology Information (NCBI) Reference Sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) provides a non-redundant collection of sequences representing genomic data, transcripts and proteins. Although the goal is to provide a comprehensive dataset representing the complete sequence information for any given species, the database pragmatically includes sequence data that are currently publicly available in the archival databases. The database incorporates data from over 2400 organisms and includes over one million proteins representing significant taxonomic diversity spanning prokaryotes, eukaryotes and viruses. Nucleotide and protein sequences are explicitly linked, and the sequences are linked to other resources including the NCBI Map Viewer and Gene. Sequences are annotated to include coding regions, conserved domains, variation, references, names, database cross-references, and other features using a combined approach of collaboration and other input from the scientific community, automated annotation, propagation from GenBank and curation by NCBI staff.

Animals↗

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

An efficient disk based data structure for rapid searching of quantitative two-dimensional gel databases.

Fast access of two-dimensional (2-D) gel quantitative databases is important for rapid searching for protein differences between sets of 2-D gels from an experiment. The GELLAB-II system organizes corresponding spots from the gels in the database into reference or "Rspot" sets. These Rspot numeric names index fixed regions in the paged composite gel database file. This is adequate for an existing database, but has several problems. (i) Building the initial database requires guessing how much disk space to pre-allocate for each corresponding spot (i.e. spots from different gels). If it ever runs out of pre-allocated space during this process, it must expand the size of each corresponding set of spots copying the old database data into the new in-place on the disk. (ii) When adding new gels or editing the database, if a new spot is created, the system may also go into this expansion mode. The time spent and wasted disk space can be appreciable--depending on the size of the database (order of 100 gel database). (iii) Because each set of corresponding spots is the same size, we waste space in most spot sets since they do not require the additional space a few spot sets require which contain additional fragmented spots. We present a new low-level disk object-based structure and algorithm, paged indexed buckets (PIB), which optimizes disk space usage while having similar retrieval speed to the original method.

Algorithms↗

Hospital charges for a community inpatient palliative care program.

Defining financial parameters of palliative care (PC) is important for providing sustainable programming. In our study, we evaluated hospital length of stay (LOS) and charges for the first 164 inpatient PC consultations performed by the Advanced Illness Assistance (AIA) team at Blount Memorial Hospital (BMH). These AIA patients had a median LOS of 11 days (range, 3-114 days), mean total charges per patient of 65,795 dollars, and mean daily charges of 3,809 dollars. Higher mean daily charges (p = 2.74 E-08, chi-square) were associated with patients who received consultation because of nonphysical symptom reasons. Patients were followed in PC consultation (AIA follow-up days) for a median of five days (range, 1-48), and had mean daily charges of 3,117 dollars. These mean daily charges were 414 dollars less than the charges for the five days prior to PC consultation (pre-AIA days) (p = 0.04, t-test). There was a significant decrease in laboratory and imaging charges during AIA follow-up (p = 0.04, t-test). The study included a reference group of patients whose information was obtained retrospectively from the BMH Atlas (MediQual, Marlborough, MA) database. These reference group patients were hospitalized at BMH during the same time, but they were not seen by the AIA team. The reference group was matched by Diagnosis Related Group (DRG), Admission Severity Grade (ASG), and disposition to the AIA patients. The Atlas patients had a shorter median LOS of six days (range, 1-105 days), and significantly greater mean daily charges of 4,105 dollars (p = 0.006, t-test) compared with AIA patients. Mean daily charges decreased for Atlas patients, as their day of discharge approached (p < 0.001). Estimates of potential charge savings were calculated in two ways: 1) by evaluating the effect of decreasing the LOS of Atlas patients with long LOS (more than seven days) to the level of AIA patients with long LOS, and 2) by comparing the actual mean patient charges during AIA follow-up with using the pre-AIA mean daily charges during the AIA follow-up period and correcting for the effect of decreasing charges that occurred as discharge approached. The estimated savings achieved by decreasing long LOS were more than 100,000 dollars per year, and estimated savings achieved using AIA follow-up charges were more than 1,801,930 dollars per year.

Adult↗

Online Mendelian Inheritance in Man (OMIM).

Online Mendelian Inheritance In Man (OMIM) is a public database of bibliographic information about human genes and genetic disorders. Begun by Dr. Victor McKusick as the authoritative reference Mendelian Inheritance in Man, it is now distributed electronically by the National Center for Biotechnology Information (NCBI). Material in OMIM is derived from the biomedical literature and is written by Dr. McKusick and his colleagues at Johns Hopkins University and elsewhere. Each OMIM entry has a full text summary of a genetic phenotype and/or gene and has copious links to other genetic resources such as DNA and protein sequence, PubMed references, mutation databases, approved gene nomenclature, and more. In addition, NCBI's neighboring feature allows users to identify related articles from PubMed selected on the basis of key words in the OMIM entry. Through its many features, OMIM is increasingly becoming a major gateway for clinicians, students, and basic researchers to the ever-growing literature and resources of human genetics.

Alleles↗

DBTSS: DataBase of human Transcriptional Start Sites and full-length cDNAs.

Although the information of cDNAs is indispensable for analyzing gene function, most of the cDNA sequences stored in current databases are imperfect in the sense that they lack the precise information of 5' end termini. To overcome this difficulty, we have developed the oligo-capping method to obtain full-length cDNAs, the information of which has been partly deposited in public databases. In this study, we further constructed human cDNA libraries enriched in clones containing the cap structure to systematically explore the 5' end structure of expressed genes. Of approximately 217 402 5' end sequences obtained, 111 382 have been matched to cDNA sequences of known genes (7889 genes) and are presented in our new database, DataBase of Transcriptional Start Sites (DBTSS; http://elmo.ims.u-tokyo.ac.jp/dbtss/). Sequence comparison between our entries and those of a reference sequence database, RefSeq, revealed that 4683 (34%) of RefSeq sequences should be extended towards the 5' ends. We also mapped each sequence on the human draft genome sequence to identify its transcriptional start site, which provides us with more detailed information on distribution patterns of transcriptional start sites and adjacent regulatory regions.

5' Flanking Region↗

Gap analysis of pediatric reference intervals for risk biomarkers of cardiovascular disease and the metabolic syndrome.

The childhood obesity epidemic has begun to compromise the health of the pediatric population by promoting premature development of atherosclerosis and the metabolic syndrome (MS), both of which significantly increase the risk of cardiovascular disease (CVD) early in life. As a result, recently, there has been increased recognition of the need to assess and closely monitor children and adolescents for risk factors of CVD and components of the MS. Serum/Plasma biomarkers including total cholesterol, triglycerides, HDL-C, LDL-C, insulin and C-peptide have been used for this purpose for many years. Recently, emerging biomarkers such as apolipoprotein AI, apolipoprotein B, leptin, adiponectin, free fatty acids, and ghrelin have been proposed as tools that provide valuable complementary information to that obtained from traditional biomarkers, if not more powerful predictions of risk. In order for biomarkers to be clinically useful in accurately diagnosing and treating disorders, age-specific reference intervals that account for differences in gender, pubertal stage, and ethnic origin are a necessity. Unfortunately, to date, many critical gaps exist in the reference interval database of most of the biomarkers that have been identified. This review contains a comprehensive gap analysis of the reference intervals for emerging and traditional risk biomarkers of CVD and the MS and discusses the clinical significance and analytical considerations of each biomarker.

Biomarkers↗

Multicentre study of fetal cardiac time intervals using magnetocardiography.

OBJECTIVE: A database with reference values of the durations of the various waveforms in a magnetocardiogram of fetuses in uncomplicated pregnancies is assessed. This database will be of help to discriminate between pathologic and healthy fetuses. A fetal magnetocardiogram is a recording of the magnetic field in a location near the maternal abdomen and reflects the electric activity within the fetal heart. It is a non-invasive method, which can be used with nearly 100% reliability from the 20th week of gestation onward. DESIGN: Durations of the waveforms were assembled from averaged magnetocardiograms and statistically processed. SETTING: Fetal magnetocardiograms were measured with different magnetocardiographs. All measurements were carried out in magnetically shielded rooms. SAMPLE: Fetal magnetocardiograms were obtained for 582 healthy patients. METHOD: The durations of the waveforms were extracted from fetal magnetocardiograms measured at the cooperating centres. The variables collected included the duration of the P-wave, the PR interval, the PQ interval, the QRS complex, the QT interval and the T-wave and QTc value. The results were compared with values extracted from electrocardiograms of fetuses measured via electrodes attached to the maternal abdomen, from electrocardiograms measured during labour using a scalp electrode, and from electrocardiograms recorded in newborns, that were found in the literature. MAIN OUTCOME MEASURES: Values of the durations are given as a function of gestational age including the regression line as well as the bounds marking the 90%, 95% and 98% prediction interval. RESULTS: The durations of the P-wave, the PR interval, the QRS complex, the QT interval and QTc value increase linearly with gestational age. The durations of the PQ interval and the T-wave are independent of fetal age. CONCLUSION: The values found agree with those found in the literature. The scatter of the data is wide due to the variation in normal physiology, the measuring system and signal processing and the subjectivity of the researcher. However, the system can define normal ranges and may be used in diagnosis.

Analysis of Variance↗

Intravenous immunoglobulin for preventing infection in preterm and/or low-birth-weight infants.

BACKGROUND: Nosocomial infections continue to be a significant cause of morbidity and mortality among preterm and/or low birth weight infants. Maternal transport of immunoglobulins to the fetus mainly occurs after 32 weeks gestation and endogenous synthesis does not begin until several months after birth. Administration of intravenous immunoglobulin provides IgG that can bind to cell surface receptors, provide opsonic activity, activate complement, promote antibody dependent cytotoxicity, and improve neutrophilic chemoluminescence. Intravenous immunoglobulin thus has the potential of preventing or altering the course of nosocomial infections. OBJECTIVES: To assess the effectiveness/safety of intravenous immunoglobulin (IVIG) administration (compared to placebo or no intervention) to preterm (< 37 weeks gestational age at birth) and/or low birth weight (LBW) (< 2500 g BW) infants in preventing nosocomial infections. SEARCH STRATEGY: Medline, Embase, Cochrane Library and Reference Update Databases were searched in November 1997 using keywords: immunoglobulin and infant-newborn and random allocation or controlled trial or randomized controlled trial (RCT). The reference lists of identified RCTs, personal files and Science Citation Index were searched. No language restrictions were applied. SELECTION CRITERIA: The criteria used to select studies for inclusion in this overview were: 1) DESIGN: RCTs in which administration of IVIG was compared to a control group that received a placebo or no intervention. 2) POPULATION: preterm (< 37 weeks gestational age) and/or LBW (<2500 g) infants. 3) INTERVENTION: IVIG for the prevention of bacterial/fungal infection during initial hospital stay (8 days or longer). (Studies that were primarily designed to assess the effect of IVIG on humoral immune markers were excluded as were studies in which the follow-up period was one week or less). 4) At least one of the following outcomes was reported: sepsis, any serious infection, death from all causes, death from infection, length of hospital stay, intraventricular haemorrhage (IVH), necrotizing enterocolitis (NEC), bronchopulmonary dysplasia (BPD). DATA COLLECTION AND ANALYSIS: Two reviewers independently abstracted information for each outcome reported in each study, and one researcher (AO) checked for any discrepancies and pooled the results. Relative risk (RR) and Risk Difference (RD) with 95% confidence intervals (CI) using the fixed effects model are reported. When a statistically significant RD was found the number needed to treat (NNT) was also calculated with 95% CIs. The results include all accepted studies in which the outcome of interest was reported. When statistically significant heterogeneity was found for an outcome, secondary (sensitivity) analyses were performed including only studies of the highest quality. MAIN RESULTS: Fifteen studies met inclusion criteria. These included 5,054 preterm and/or LBW infants and reported on at least one of the outcomes of interest for this systematic review. When all studies were combined there was a statistically significant reduction in sepsis, one or more episodes [RR 0.83 (95% CI 0.72, 0.97); RD -0.028 (95% CI -0.006, -0.051); NNT 36 (95% CI 20, 167)]. There was significant between-study heterogeneity. When, in a sensitivity analysis, the high quality studies were combined, the results remained significant [RR 0.78 (95% CI 0.62, 0.98); RD -0.031(95% CI -0.003, -0.059); NNT 32 (95% CI 17, 333]. For this analysis there was no statistically significant between-study heterogeneity. A statistically significant reduction was also found for any serious infection, one or more episodes, when all studies were combined [RR 0.85 (95% CI 0.75, 0. 95); RD -0.032 (95% CI -0.010, -0.054,); NNT 31 (95% CI 19, 100). There was statistically significant between-study heterogeneity. When, in a sensitivity analysis, the high quality studies were combined the results remained statistically significant [RR 0.80 (95% CI

Cross Infection↗

Antiepileptic drug use and birth rate in patients with epilepsy--a population-based cohort study in Finland.

BACKGROUND: Antiepileptic medication use affects reproductive endocrine function, but its impact on fertility is not well known. METHODS: All epilepsy patients, who were approved as being eligible for reimbursement for antiepileptic drug (AED) costs from the Social Insurance Institution (SII) of Finland for the first time 1985-94, were identified from the SII database. A reference cohort without epilepsy was identified from the Finnish Population Register Centre. Information on AED purchases 1996-2000 was obtained from the SII database through computerized record linkage with the unique personal identification number assigned to all residents of Finland. The three AEDs included were carbamazepine, oxcarbazepine (OXC) and valproate. RESULTS: Birth rate was lower in both men and women with epilepsy on AEDs than in the reference cohort without epilepsy. However, compared with patients not using AED during the study period, the birth rate was lowered only among men on OXC [rate ratio (RR) = 0.52, 95% confidence interval (CI) = 0.32, 0.84]. CONCLUSIONS: The birth rate was lower in both women and men on any of the three AEDs compared with the reference cohort without epilepsy. However, a statistically significant difference between treated and untreated patients was only seen in men on OXC. It is unclear to what extent the differences found in this study are due to social or biological factors.

Adolescent↗

BioPD: a web-based information center for bioactive peptides.

Bioactive peptide database (BioPD) is a web-based knowledge base that contains more than 1100 protein sequences from human, mouse and rat, which are putative or are known to be bioactive peptides. In addition to peptide sequences and the annotation, the database also contains gene sequences with annotation, protein interaction and disease data related to the peptides. Each entry has as many references as possible to support the information represented. BioPD consists of six parts: PROTEIN, GENE, DISEASE, LINKS, INTERACTION, and REFERENCE. The database is searchable through keyword, gene and protein name, receptor name, etc. The links to PDB, InterPro, Pfam, OMIM, etc. are provided in each entry. Thus BioPD is formed as an information center for the bioactive peptide and serves as a gateway for exploration of bioactive peptides. The database can be accessed at http://biopd.bjmu.edu.cn.

Animals↗