Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Population of the HLA ligand database.

We have established an HLA ligand database to provide scientists and clinicians with access to Major Histocompatibility Complex (MHC) class I and II motif and ligand data. The HLA Ligand Database is available on the world wide web at http://hlaligand.ouhsc.edu and contains ligands that have been published in peer-reviewed journals. HLA peptide datasets prove useful in several areas: ligands are important as targets for various immune responses while algorithms built upon ligand datasets allow identification of new peptides without time-consuming experimental procedures. A review of the HLA class I ligands in the database identifies strengths and deficiencies in the database and, therefore, the utility of the dataset for identifying new peptides. For instance, 212 HLA-A phenotypes exist of which 23 have a motif determined and 43 have peptides characterized. In terms of number of ligands, HLA-A*0201 has 258 characterized ligands, A*1101 has 25 peptides, while the remaining two-thirds of the HLA-A phenotypes have less than 10 associated peptide sequences. Characterization of ligands and motifs remains roughly the same at the HLA-B locus while the peptides of the HLA-C locus tend to be less characterized. These data show that 74% of HLA class I molecules do not have ligands represented in the database and thus algorithms based on the dataset could not predict ligands for a majority of the US population. Building upon this dataset and knowledge of HLA allelic frequencies, it is possible to plan a systematic expansion of the HLA class I ligand database to better identify ligands useful throughout the population.

Databases, Protein↗

Overview of large database analysis in renal transplantation.

The discipline of renal transplantation has been fortunate in having one of the largest and most complete databases of any field of medical inquiry. Many analyses have been performed utilizing these databases and a wide array of views regarding the role and limitations of these analyses exists. In this manuscript, we hope to present the merits and limitations of large database analysis in renal transplantation. In addition, we will attempt to go over the major databases' structures and the critical issues of coding, verification and clinical awareness that help maintain the integrity of this type of analysis. Other fundamental issues that will be covered include the relationship of single center and randomized prospective studies with database analyses, statistical considerations, and hidden selection bias inherent in registry analysis. We hope to present both a guide to investigators who are first embarking on large-scale database analysis and also to present a fair view of the utility and limitations of a potentially very useful research tool in renal transplantation.

Databases, Factual↗

Validity of self-reported energy intake in lean and obese young women, using two nutrient databases, compared with total energy expenditure assessed by doubly labeled water.

OBJECTIVE: To compare self-reported total energy intake (TEI) estimated using two databases with total energy expenditure (TEE) measured by doubly labeled water in physically active lean and sedentary obese young women, and to compare reporting accuracy between the two subject groups. DESIGN: A cross-sectional study in which dietary intakes of women trained in diet-recording procedures were analyzed using the Minnesota Nutrition Data System (NDS; versions 2.4/6A/21, 2.6/6A/23 and 2.6/8.A/23) and Nutritionist III (N3; version 7.0) software. Reporting accuracy was determined by comparison of average TEI assessed by an 8 day estimated diet record with average TEE for the same period. RESULTS: Reported TEI differed from TEE for both groups irrespective of nutrient database (P<0.01). Measured TEE was 11.10+/-2.54 and 11.96+/-1.21 MJ for lean and obese subjects, respectively. Reported TEI, using either database, did not differ between groups. For lean women, TEI calculated by NDS was 7.66+/-1.73 MJ and by N3 was 8.44+/-1.59 MJ. Corresponding TEI for obese women were 7.46+/-2.17 MJ from NDS and 7.34+/-2.27 MJ from N3. Lean women under-reported by 23% (N3) and 30% (NDS), and obese women under-reported by 39% (N3) and 38% (NDS). Regardless of database, lean women reported higher carbohydrate intakes, and obese women reported higher total fat and individual fatty acid intakes. Higher energy intakes from mono- and polyunsaturated fatty acids were estimated by NDS than by N3 in both groups of women (P< or =0.05). CONCLUSIONS: Both physically active lean and sedentary obese women under-reported TEI regardless of database, although the magnitude of under-reporting may be influenced by the database for the lean women. SPONSORSHIP: USDA Hatch Project award (ARZT-136528-H-23-111) to LB Houtkooper and WH Howell.

Adolescent↗

Use of the UK General Practice Research Database for pharmacoepidemiology.

The last decade has seen a surge in the use of computerized health care data for pharmacoepidemiology. Of all European databases, the General Practice Research Database (GPRD) in the UK, has been the most widely used for pharmacoepidemiological research. Since 1994, this database has belonged to the UK Department of Health, and is maintained by the Office of National Statistics (ONS). Currently, around 1500 general practitioners with a population coverage in excess of 3 million, systematically provide their computerized medical data anonymously to ONS. Validation studies of the GPRD have documented the recording of medical data into general practitioners' computers to be near to complete. The GPRD collects truly population-based data, has a size that makes it possible to follow-up large cohorts of users of specific drugs, and includes both outpatient and inpatient clinical information. The access to original medical records is excellent. Desirable improvements to the GPRD would be additional computerized information on certain variables and linkage to other health care databases. Most published studies to date have been in the area of drug safety. The General Practice Research Database has proved that valuable data can be collected in a general practice setting. The full potential of this rich computerized database has yet to come. This experience should serve to encourage others to develop similar population-based data in other countries.

Databases, Factual↗

A scientific relational database combined with a report generator for endoscopy in networks: EndoNet.

BACKGROUND AND STUDY AIMS: The flexibility required in academic endoscopy units is not provided by the available database systems. In a project involving substantial cooperation between endoscopists and computer scientists, we have developed an adaptable database, combined with a report generator embedded in the hospital's intranet. PATIENTS AND METHODS: Six workstations in different areas of the hospital were clustered with a UNIX operating system to implement multi-user capability and access control. A relational database was used to design an application appropriate to the specific needs of the endoscopy unit in a teaching hospital engaged in scientific research. Both the terminology used in standardized endoscopy nomenclature and a free text block facility were included. A graphical user interface was developed to assemble pertinent data, generate the reports, and supervise the database. RESULTS: A total of 4936 examinations including 2988 patients were entered consecutively during continuous routine operation of the system. Complete report generation required five minutes (median; range 1-9 minutes). Both structured items and free text were used in all the reports. Querying of the database was possible, concerning matters such as the need for repeated endoscopic therapy in acute gastrointestinal bleeding (4%), the search for Helicobacter pylori in appropriate patients (64%), the rate of accidental pancreatic duct visualization in endoscopic retrograde cholangiography (24%), and links between examinations and active trials (2%). Indicating improved report quality, the number and the diameter of esophageal varices in patients with varices were more frequently reported with the new report system than with previous typed reports (P<0.001). An anonymous questionnaire revealed that the readability of the computer-generated reports was better than that of the previous typewritten reports (P=0.01). CONCLUSIONS: This report describes the creation of a database application and a report generator meeting the needs of scientific and routine use, and the successful application of this system in an academic endoscopy unit.

Computer Communication Networks↗

A nutrigenomics database--integrated repository for publications and associated microarray data in nutrigenomics research.

In the current situation where microarray data in the field of nutritional genomics (nutrigenomics) are accumulating rapidly, there is imminent need for an efficient data infrastructure to support research workflow. We have established a web-based, integrated database of the publications and microarray expression data in the field of nutrigenomics. The registered data include links to external databases such as PubMed of the National Center for Biotechnology Information and public microarray databases that contain Minimum Information About a Microarray Experiment-compliant microarray expression data. Using this database, all data sets created will be effectively utilized and shared with other researchers. This database is built on an open-source database system and is freely accessible via the World Wide Web (http://a-yo5.ch.a.u-tokyo.ac.jp/index.phtml).

Databases, Genetic↗

Estimation of exposure to food packaging materials. 1: Development of a food-packaging database.

A food-packaging database was developed to provide qualitative information on the types of packaging materials used for foods. Packaging information was collected from a sample of 594 children aged 5-12 years as part of a national children's food survey carried out in Ireland during 2003-04. All the food packaging collected during the survey was forwarded to the coordinating centre for further analysis and entry into the Irish Food Packaging Database. The database was created in Microsoft Access and stored information on: the brand of the food, the packaging type, the unit weight, the contact layer, the European Union food type (i.e. aqueous, acidic, alcoholic or fatty) and other relevant parameters. Of the 5551 different brand foods consumed by children in the food survey, packaging information was collected on 3441 (62%). As some brand foods had different unit weights and packaging formats, there was duplication of some brand foods in the database to account for this fact. Therefore, there were 3672 packaging entries in the database. Of these, plastics were the most common packaging contact layer (n = 2874, 78.3%). Multimaterial multilayers with a plastic contact layer accounted for 459 (12.5%) entries. Polyethylene was the most frequently used contact layer (n = 941), with polypropylene a close second (n = 809). This database is unique in Europe for the quality and amount of food packaging information it contains and could be used to develop packaging use factors for a more refined exposure assessment to food packaging materials in the European Union.

Beverages↗

Collection and preparation of molecular databases for virtual screening.

Drug discovery and development research is undergoing a paradigm shift from a linear and sequential nature of the various steps involved in the drug discovery process of the past to the more parallel approach of the present, due to a lack of sufficient correlation between activities estimated by in vitro and in vivo assays. This is attributed to the non-drug-likeness of the lead molecules, which has often been detected at advanced drug development stages. Thus a striking aspect of this paradigm shift has been early/parallel in silico prioritization of drug-like molecular databases (also database pre-processing), in addition to prioritizing compounds with high affinity and selectivity for a protein target. In view of this, a drug-like database useful for virtual screening has been created by prioritizing molecules from 36 catalog suppliers, using our recently derived binary QSAR based drug-likeness model as a filter. The performance of this model was assessed by a comparative evaluation with respect to commonly used filters implemented by the ZINC database. Since the model was derived considering all the limitations that have plagued the existing rules and models, it performs better than the existing filters and thus the molecules prioritized by this filter represent a better subset of drug-like compounds. The application of this model on exhaustive subsets of 4,972,123 molecules, many of which have passed the ZINC database filters for drug-likeness, led to a further prioritization of 2,920,551 drug-like molecules. This database may have a great potential for in silico virtual screening for discovering molecules, which may survive the later stages of the drug development research.

Computational Biology↗

Using the NTP database to assess the value of rodent carcinogenicity studies for determining human cancer risk.

The large database of carcinogenicity results generated by the National Toxicology Program (NTP) provides a unique opportunity to critically evaluate important scientific issues such as (1) the frequency of positive outcomes, (2) the interspecies correlation in carcinogenic response between rats and mice, (3) the correlation between body weight and tumor incidence, (4) estimates of the false-positive and false-negative rates, and (5) the frequency of decreasing tumor incidences. Such database evaluations enable us to better understand the value and limitations of rodent carcinogenicity studies for determining human cancer risk. However, as the NTP database becomes increasingly accessible to the general scientific community, there is also increased opportunity for misuse of the database. This article reexamines and updates previous database evaluations, presents four scientific principles that should be employed by anyone attempting to use this database, and illustrates how failure to apply these principles can lead to misleading results.

Animals↗

Access to databases in complementary medicine.

Access to medical databases is a keystone for obtaining up-to-date and complete information for physicians. In the last few years, the rapid growth of the World Wide Web has given rise to an information revolution, enabling health care providers to gain access (often free) to an expanding volume of information that was previously inaccessible. Search engines and online databases assist the search for health information. In this article we examine the biomedical databases of primary interest in the field of alternative and complementary medicine, dividing them into Web accessible and nonaccessible databases and emphasizing the freely available ones. A further classification is major biomedical bibliographic databases specific to complementary medicine, and dedicated therapy or modality-specific databases.

Complementary Therapies↗

A database for cell signaling networks.

We developed a data and knowledge base for cellular signal transduction in human cells, to make this rapidly growing information available. The database includes all the biological properties of cellular signal transduction, including biological reactions that transfer cellular signals and molecular attributes characterized by sequences, structures, and functions. Since the database is based on the object-oriented technique, highly flexible methods of data definition and modification are necessary to handle this diverse and complex biological information. The database includes attractive graphical representations of signaling cascades and the three-dimensional structure of molecules. The database is a novel application of ACEDB, which was the database originally developed to store the C. elegans genome. The database can be accessed through the Internet at http://geo.nihs.go.jp/csndb.html.

Cells↗

Use of an automated database to evaluate markers for early detection of pregnancy.

The objective of this study was to develop and validate algorithms to detect pregnancies from the time of first clinical recognition by using Kaiser Permanente automated databases from Portland, Oregon. In 1993--1994, the authors evaluated these databases retrospectively to identify markers indicative of initial clinical detection of pregnancy and pregnancy outcomes. Pregnancy markers were found for 99% of the women for whom pregnancy outcomes were included in the automated databases, and pregnancy outcomes were identified for 77% of the women for whom there were pregnancy markers. The earliest marker most predictive of a pregnancy outcome was a positive human chorionic gonadotropin test; least predictive was an obstetric outpatient visit. Medical record review indicated that in a sample of women with pregnancy markers in the database, an estimated 6% of pregnancy outcomes (primarily early fetal deaths and elective terminations) were lost. Pregnancies were first captured in automated databases 6--8 weeks after the last menstrual period, and a combination of a positive human chorionic gonadotropin test and an outpatient obstetric visit was the most sensitive and specific early marker of pregnancy. When combined with automated pharmacy records, these databases may be valuable tools for evaluating prescription drug effects on all major outcomes of clinically recognized pregnancies.

Abortion, Legal↗

Post-processing of BLAST results using databases of clustered sequences.

MOTIVATION: When evaluating the results of a sequence similarity search, there are many situations where it can be useful to determine whether sequences appearing in the results share some distinguishing characteristic. Such dependencies between database entries are often not readily identifiable, but can yield important new insights into the biological function of a gene or protein. RESULTS: We have developed a program called CBLAST that sorts the results of a BLAST sequence similarity search according to sequence membership in user-defined 'clusters' of sequences. To demonstrate the utility of this application, we have constructed two cluster databases. The first describes clusters of nucleotide sequences representing the same gene, as documented in the UNIGENE database, and the second describes clusters of protein sequences which are members of the protein families documented in the PROSITE database. Cluster databases and the CBLAST post-processor provide an efficient mechanism for identifying and exploring relationships and dependencies between new sequences and database entries.

Algorithms↗

A set-theoretic approach to database searching and clustering.

MOTIVATION: In this paper, we introduce an iterative method of database searching and apply it to design a database clustering algorithm applicable to an entire protein database. The clustering procedure relies on the quality of the database searching routine and further improves its results based on a set-theoretic analysis of a highly redundant yet efficient to generate cluster system. RESULTS: Overall, we achieve unambiguous assignment of 80% of SWISS-PROT sequences to non-overlapping sequence clusters in an entirely automatic fashion. Our results are compared to an expert-generated clustering for validation. The database searching method is fast and the clustering technique does not require time-consuming all-against-all comparison. This allows for fast clustering of large amounts of sequences. AVAILABILITY: The resulting clustering for the PIR1 (Release 51) and SWISS-PROT (Release 34) databases is available over the Internet from http://www.dkfz-heidelberg.de/tbi/services/modest/b rowsesysters.pl. CONTACT: a.krause@dkfz-heidelberg.de; m.vingron@dkfz-heidelberg.de

Algorithms↗

Improved database searches for orthologous sequences by conditioning on outgroup sequences.

MOTIVATION: Searches of biological sequence databases are usually focussed on distinguishing significant from random matches. However, the increasing abundance of related sequences on databases present a second challenge: to distinguish the evolutionarily most closely related sequences (often orthologues) from more distantly related homologues. This is particularly important when searching a database of partial sequences, where short orthologous sequences from a non-conserved region will score much more poorly than non-orthologous (outgroup) sequences from a conserved region. RESULTS: Such inferences are shown to be improved by conditioning the search results on the scores of an outgroup sequence. The log-odds score for each target sequence identified on the database has the log-odds score of the outgroup sequence subtracted from it. A test group of Caenorhabditis elegans kinase sequences and their identified C.elegans outgroups were searched against a test database of human Expressed Sequence Tag (EST) sequences, where the sets of true target sequences were known in advance. The outgroup conditioned method was shown to identify 58% more true positives ahead of the first false positive, compared to the straightforward search without an outgroup. A test dataset of 151 proteins drawn from the C.elegans genome, where the putative 'outgroup' was assigned automatically, similarly found 50% more true positives using outgroup conditioning. Thus, outgroup conditioning provides a means to improve the results of database searching with little increase in the search computation time.

Algorithms↗

Rapid 3D protein structure database searching using information retrieval techniques.

MOTIVATION: As the sizes of three-dimensional (3D) protein structure databases are growing rapidly nowadays, exhaustive database searching, in which a 3D query structure is compared to each and every structure in the database, becomes inefficient. We propose a rapid 3D protein structure retrieval system named 'ProtDex2', in which we adopt the techniques used in information retrieval systems in order to perform rapid database searching without having access to every 3D structure in the database. The retrieval process is based on the inverted-file index constructed on the feature vectors of the relationships between the secondary structure elements (SSEs) of all the 3D protein structures in the database. ProtDex2 is a significant improvement, both in terms of speed and accuracy, upon its predecessor system, ProtDex. RESULTS: The experimental results show that ProtDex2 is very much faster than two well-known protein structure comparison methods, DALI and CE, yet not sacrificing on the accuracy of the comparison. When comparing with a similar SSE-based method, namely TopScan, ProtDex2 is much faster with comparable degree of accuracy. AVAILABILITY: The software is available at: http://xena1.ddns.comp.nus.edu.sg/~genesis/PD2.htm

Algorithms↗

Efficient selection of unique and popular oligos for large EST databases.

MOTIVATION: Expressed sequence tag (EST) databases have grown exponentially in recent years and now represent the largest collection of genetic sequences. An important application of these databases is that they contain information useful for the design of gene-specific oligonucleotides (or simply, oligos) that can be used in PCR primer design, microarray experiments and genomic library screening. RESULTS: In this paper, we study two complementary problems concerning the selection of short oligos, e.g. 20-50 bases, from a large database of tens of thousands of ESTs: (i) selection of oligos each of which appears (exactly) in one unigene but does not appear (exactly or approximately) in any other unigene and (ii) selection of oligos that appear (exactly or approximately) in many unigenes. The first problem is called the unique oligo problem and has applications in PCR primer and microarray probe designs, and library screening for gene-rich clones. The second is called the popular oligo problem and is also useful in screening genomic libraries. We present an efficient algorithm to identify all unique oligos in the unigenes and an efficient heuristic algorithm to enumerate the most popular oligos. By taking into account the distribution of the frequencies of the words in the unigene database, the algorithms have been engineered carefully to achieve remarkable running times on regular PCs. Each of the algorithms takes only a couple of hours (on a 1.2 GHz CPU, 1 GB RAM machine) to run on a dataset 28 Mb of barley unigenes from the HarvEST database. We present simulation results on the synthetic data and a preliminary analysis of the barley unigene database. AVAILABILITY: Available on request from the authors.

Algorithms↗

TRbase: a database relating tandem repeats to disease genes for the human genome.

MOTIVATION: Tandem repeats are associated with disease genes, play an important role in evolution and are important in genomic organization and function. Although much research has been done on short perfect patterns of repeats, there has been less focus on imperfect repeats. Thus, there is an acute need for a tandem repeats database that provides reliable and up to date information on both perfect and imperfect tandem repeats in the human genome and relates these to disease genes. RESULTS: This paper presents a web-accessible relational tandem repeats database that relates tandem repeats to gene locations and disease genes of the human genome. In contrast to other available databases, this database identifies both perfect and imperfect repeats of 1-2000 bp unit lengths. The utility of this database has been illustrated by analysing these repeats for their distribution and frequencies across chromosomes and genomic locations and between protein-coding and non-coding regions. The applicability of this database to identify diseases associated with previously uncharacterized tandem repeats is demonstrated.

Chromosome Mapping↗