Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

MHCBN: a comprehensive database of MHC binding and non-binding peptides.

MHCBN is a comprehensive database of Major Histocompatibility Complex (MHC) binding and non-binding peptides compiled from published literature and existing databases. The latest version of the database has 19 777 entries including 17 129 MHC binders and 2648 MHC non-binders for more than 400 MHC molecules. The database has sequence and structure data of (a) source proteins of peptides and (b) MHC molecules. MHCBN has a number of web tools that include: (i) mapping of peptide on query sequence; (ii) search on any field; (iii) creation of data sets; and (iv) online data submission. The database also provides hypertext links to major databases like SWISS-PROT, PDB, IMGT/HLA-DB, GenBank and PUBMED.

Amino Acid Sequence↗

BioQuery: an object framework for building queries to biomedical databases.

SUMMARY: BioQuery is an application that helps scientists automate database searches. Users can build and store queries to public biomedical databases, and receive periodic updates on the results of those queries when new data is available. The application is implemented on a portable object framework that can provide database-searching capability to other applications. This framework is easily extensible, allowing users to develop plug-ins that provide access to new databases. BioQuery thus provides end-users with a complete database searching interface and updating service, and gives developers a toolkit to provide database-searching capability to their applications. AVAILABILITY: Free to all users: http://www.bioquery.org.

Biomedical Research↗

DAtA: database of Arabidopsis thaliana annotation.

The Database of Arabidopsis thaliana Annotation (D At A) was created to enable easy access to and analysis of all the Arabidopsis genome project annotation. The database was constructed using the completed A.thaliana genomic sequence data currently in GenBank. An automated annotation process was used to predict coding sequences for GenBank records that do not include annotation. D At A also contains protein motifs and protein similarities derived from searches of the proteins in D At A with motif databases and the non-redundant protein database. The database is routinely updated to include new GenBank submissions for Arabidopsis genomic sequences and new Blast and protein motif search results. A web interface to D At A allows coding sequences to be searched by name, comment, blast similarity or motif field. In addition, browse options present lists of either all the protein names or identified motifs present in the sequenced A.thaliana genome. The database can be accessed at http://baggage. stanford.edu/group/arabprotein/

Arabidopsis↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

Characteristics of the U.S. EPA's Office of Pesticide Programs' toxicity information databases.

The United States Environmental Protection Agency's Office of Pesticide Programs (OPP) requires that data from toxicity testing be submitted to the OPP to support the registration of pesticide chemicals. Once the toxicity data are submitted, they are entered into various toxicity databases. The studies are listed in an archival database to catalog and allow retrieval of the study for review. Reviews of toxicity studies are then placed into a separate database that can be retrieved to support a regulatory position. Toxicity information for health effects other than cancer and gene mutations from chronic exposure is reviewed through a reference dose (RfD) approach, and these decisions and supporting data are entered into an RfD database. Carcinogenicity data are reviewed by a peer review process, and these decisions are entered into a newly developed database to show the regulatory decision with supporting data. The mutagenicity data are reviewed and acceptable data are entered into the Genetic Activity Profile system to catalog and display the submitted information. These databases contain the information used for hazard evaluations as part of the OPP review of pesticide chemicals.

Animals↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

Measuring use patterns of online journals and databases.

PURPOSE: This research sought to determine use of online biomedical journals and databases and to assess current user characteristics associated with the use of online resources in an academic health sciences center. SETTING: The Library of the Health Sciences-Peoria is a regional site of the University of Illinois at Chicago (UIC) Library with 350 print journals, more than 4,000 online journals, and multiple online databases. METHODOLOGY: A survey was designed to assess online journal use, print journal use, database use, computer literacy levels, and other library user characteristics. A survey was sent through campus mail to all (471) UIC Peoria faculty, residents, and students. RESULTS: Forty-one percent (188) of the surveys were returned. Ninety-eight percent of the students, faculty, and residents reported having convenient access to a computer connected to the Internet. While 53% of the users indicated they searched MEDLINE at least once a week, other databases showed much lower usage. Overall, 71% of respondents indicated a preference for online over print journals when possible. CONCLUSIONS: Users prefer online resources to print, and many choose to access these online resources remotely. Convenience and full-text availability appear to play roles in selecting online resources. The findings of this study suggest that databases without links to full text and online journal collections without links from bibliographic databases will have lower use. These findings have implications for collection development, promotion of library resources, and end-user training.

Computer Literacy↗

Non-sequence databases for biological activity and physicochemical properties.

A biological activity database and a physicochemical property database are described. They are intended to complement the protein sequence database of PIR-International. The Biological Activity Database and the Physicochemical Property Database contain information regarding the biological activity and the physicochemical properties of proteins, respectively. In addition they also provide information about wild-type molecules with which information concerning variant molecules may be compared. Data on artificial variant molecules are stored in the Artificial Variant Database which is described separately.

Amino Acid Sequence↗

A database model for studies of cocaine-dependent pregnant women and their families.

The database management functions for the Mothers Project are arranged into administrative and analytic task groups, and separate systems are devised for each. The task groups can be distinguished not only by differences in data structure but also by interface requirements. The administrative database system uses a relational database technology, whereas the analytic database system employs more traditional flat-file methods. Although the database management systems are complex, they are based on standard database practices, used in widely available software packages, and run on inexpensive desktop computing equipment.

Cocaine↗

Creation and maintenance of Helix, a Web based database of medical genetics laboratories, to serve the needs of the genetics community.

Helix (healthlinks.washington.edu/helix) is a web accessible database that serves as the main U.S. directory of laboratories offering genetic testing. The database was designed to address the previously unmet need for a centralized, continuously updated source of information about clinical and research genetic testing to keep pace with the rapid rate of gene discovery resulting from the Human Genome Project. The Helix project began in 1992 at the University of Washington and Children's Hospital and Regional Medical Center. It has evolved from a single user stand alone relational database to a fully Web enabled database queried and maintained via the web and linked to other web accessible genomic databases. As of February, 1998 it lists more than 500 diseases and 290 laboratories, with over 5,200 registered users making approximately 250 queries/day (90% via the Internet). We describe the iterative design, implementation, population and assessment of the database over a six year period.

Database Management Systems↗

Human gene mutation database-a biomedical information and research resource.

Although 20 years have elapsed since the first single basepair substitution underlying an inherited disease in humans was characterised at the DNA level, the initiative has only recently been taken to establish central database resources for pathological genetic variants. Disease-associated gene lesions are currently collected and publicised by the Human Gene Mutation Database (HGMD) in Cardiff, locus-specific mutation databases, and to some extent also by the Genome Database (GDB) and Online Mendelian Inheritance in Man (OMIM). To date, HGMD represents the only comprehensive and publicly available database of gene lesions underlying human inherited disease. By July 1999, HGMD contained over 18,000 different mutations from some 900 human genes, the majority being single basepair substitutions. In addition to its potential as an information resource for clinicians and genetic counsellors, HGMD has allowed molecular geneticists to address a variety of biological questions through meta-analysis of the collated data. HGMD also promises to assist research workers in optimising mutation search strategies for a given gene. A questionnaire sent out to, and answered by, the editors of 20 key journals revealed that human genetics journals are increasingly reluctant to publish mutation reports. Electronic data submission and publication facilities are therefore urgently required. The World Wide Web (WWW) provides an excellent medium within which to combine the centralised management of basic mutation data, including rigorous quality control, with the possibility of publishing additional mutation-related information. In response to these needs, HGMD has both instituted a collaboration with Springer-Verlag GmbH, Heidelberg, to potentiate free online submission and electronic publication of human gene mutation data and developed links with the curators of locus-specific mutation databases.

Databases, Factual↗

A dynamic two-dimensional polyacrylamide gel electrophoresis database: the mycobacterial proteome via Internet.

Proteome analysis by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry, in combination with protein chemical methods, is a powerful approach for the analysis of the protein composition of complex biological samples. Data organization is imperative for efficient handling of the vast amount of information generated. Thus we have constructed a 2-D PAGE database to store and compare protein patterns of cell-associated and culture-supernatant proteins of different mycobacterial strains. In accordance with the guidelines for federated 2-DE databases, we developed a program that generates a dynamic 2-D PAGE database for the World-Wide-Web to organise and publish, via the internet, our results from proteome analysis of different Mycobacterium tuberculosis as well as Mycobacterium bovis BCG strains. The uniform resource locator for the database is http://www.mpiib-berlin.mpg.de/2D-PAGE and can be read with a Java compatible browser. The interactive hypertext markup language documents displayed are generated dynamically in each individual session from a rational data file, a 2-D gel image file and a map file describing the protein spots as polygons. The program consists of common gateway interface scripts written in PERL, minimizing the administrative workload of the database. Furthermore, the database facilitates not only interactive use, but also worldwide active participation of other scientific groups with their own data, requiring only minimal computer hardware and knowledge of information technology.

Bacterial Proteins↗

The structure and dipole moment of globular proteins in solution and crystalline states: use of NMR and X-ray databases for the numerical calculation of dipole moment.

The large dipole moment of globular proteins has been well known because of the detailed studies using dielectric relaxation and electro-optical methods. The search for the origin of these dipolemoments, however, must be based on the detailed knowledge on protein structure with atomic resolutions. At present, we have two sources of information on the structure of protein molecules: (1) x-ray databases obtained in crystalline state; (2) NMR databases obtained in solution state. While x-ray databases consist of only one model, NMR databases, because of the fluctuation of the protein folding in solution, consist of a number of models, thus enabling the computation of dipole moment repeated for all these models. The aim of this work, using these databases, is the detailed investigation on the interdependence between the structure and dipole moment of protein molecules. The dipole moment of protein molecules has roughly two components: one dipole moment is due to surface charges and the other, core dipole moment, is due to polar groups such as N--H and C==O bonds. The computation of surface charge dipole moment consists of two steps: (A) calculation of the pK shifts of charged groups for electrostatic interactions and (B) calculation of the dipole moment using the pK corrected for electrostatic shifts. The dipole moments of several proteins were computed using both NMR and x-ray databases. The dipole moments of these two sets of calculations are, with a few exceptions, in good agreement with one another and also with measured dipole moments.

Animals↗

Use of mass spectrometric molecular weight information to identify proteins in sequence databases.

During the last decade new ionization techniques have made it possible to measure the molecular weight of many intact proteins by mass spectrometry, and they have made it much easier to obtain a mass spectrometric peptide map of a protein. At the same time advances in protein and DNA sequencing technology are resulting in an exponential increase in the number of sequences deposited in databases. Here we investigate the possibility to use mass spectrometric data to identify proteins in databases. Searching a database by total molecular weight is found to be an easy and sometimes sufficient approach. For more specificity and for error tolerance in both the mass spectrometric data and the database information we search by partial mass spectrometric peptide map of the protein. In general, just four to six proteolytic peptides measured with a mass accuracy between 0.1 and 0.01% allow a useful search of databases such as the Protein Identification Resource (PIR). As the size of DNA and protein sequence databases grows, protein identification by partial mass spectrometric peptide maps should become increasingly powerful and may become a general method to identify and characterize proteins.

Amino Acid Sequence↗

Microcomputer-assisted filing system of cardiac catheterization records using a relational database management system.

To efficiently store and retrieve cardiac catheterization records, we have developed a computer-assisted database, which comprises a 16-bit microcomputer with dual floppy disk drives, a 20 MB random-access memory, hard disk drive, and a line printer. All programmings were accomplished using a relational database management system (R:base 5000, Microrim, Inc.). Data inquiry procedures could be performed with direct operational commands of the system as well as with preprogrammed command files, and final results of searches were printed out with a line printer. The major advantages of the present system described in this report include: (1) the relatively easy and rapid creation of the database, (2) ease of modification of the database structures even after the system design is finished, (3) operational commands in combination with conditional operator(s) are flexible and powerful enough to allow the end user to retrieve data based on various kinds of criteria, (4) a high-level programming language provided by the R:base automates a series of database procedures with relative ease, (5) relational capabilities of the database management system can enhance the possibility of reconstruction of a new data file from a single or several preexisting data files, and (6) the system can be realized at reasonable cost.

Cardiac Catheterization↗

The gene-protein database of Escherichia coli: edition 4.

The gene-protein database of Escherichia coli has as its core an index that links each of the protein spots from a two-dimensional polyacrylamide gel to the gene that encodes the protein. Additional information about each protein and its gene is generated from two-dimensional gel analysis or collated from the literature to form the database. Earlier editions of the database have provided periodic updates of information. The current edition does this, but also introduces a new reference gel image produced by an electrophoresis system recently adopted in this laboratory. The new gel system was chosen because it offers an improved opportunity for other investigations to produce close replicas of the reference gel pattern, thereby allowing easier access to the information of the database and encouraging independent contribution to the database. The new gel format also is larger and hence more compatible with computer assisted image analysis, which has become essential for a project of this magnitude. This edition continues the use of the former reference gel images, but adds a reference image of an equilibrium gel of E. coli strain W3110 produced by the new standardized gel system. At this time, 55% of the protein spots annotated on the previous equilibrium reference gel for this organism have been located on the new reference image, and these identifications are included in the tables of the database.

Bacterial Proteins↗

Mouse liver protein database: a catalog of proteins detected by two-dimensional gel electrophoresis.

Alterations in the abundance or structure of mouse liver proteins are being studied using two-dimensional gel electrophoresis (2-DE) to build a database of protein changes correlating with exposure to ionizing radiation or toxic chemicals. Thus far, studies have included the analysis of proteins from the offspring of exposed parents or from the exposed individuals themselves. In order to characterize and identify proteins found altered by such exposures, sex- and strain-related differences in protein patterns have been analyzed, and the subcellular locations of a large portion of the mapped proteins have been determined. As part of these studies, data are collected and stored using a variety of computer hardware and software tools that allow the accumulation of information on the origin of samples, gel identification, experiment description, and protein similarities and differences. This accumulation of information constitutes the mouse liver protein database. Relational database software is used to tie the different facets of the database together so that the results of a variety of experiments can be compared and interrelated. The database optimizes the information obtained from 2-DE gel sets by allowing use of the data for many purposes, including monitoring of gel resolution to ensure the collection of high quality data and correlation of protein effects induced by different agents. This first edition of the Argonne National Laboratory mouse liver protein database lays the foundation for future work and communication that should elucidate the significance of observed protein effects as possible markers of exposure to toxic agents.

Animals↗

HSC-2DPAGE and the two-dimensional gel electrophoresis database of dog heart proteins.

A two-dimensional gel electrophoresis database of dog (Canis familiaris) proteins is presented. The database contains 1212 protein spots which have been characterised in terms of their pI and Mr. This database has been integrated into the HSC-2DPAGE database which is accessible on the Internet via the World Wide Web with the uniform resource location (URL): (http://www.harefield.nthames.nhs.uk/nhli/ protein/index.html). Identifications for 80 of the protein spots have been obtained by visual cross-matching with the human heart protein database in HSC-2DPAGE (42 spots), N-terminal microsequence analysis (25 spots) and peptide mass fingerprinting (20 spots). This database is being used in studies of alterations in protein expression in models of heart failure and heart disease.

Amino Acid Sequence↗