Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

Improvements to GALA and dbERGE II: databases featuring genomic sequence alignment, annotation and experimental results.

We describe improvements to two databases that give access to information on genomic sequence similarities, functional elements in DNA and experimental results that demonstrate those functions. GALA, the database of Genome ALignments and Annotations, is now a set of interlinked relational databases for five vertebrate species, human, chimpanzee, mouse, rat and chicken. For each species, GALA records pairwise and multiple sequence alignments, scores derived from those alignments that reflect the likelihood of being under purifying selection or being a regulatory element, and extensive annotations such as genes, gene expression patterns and transcription factor binding sites. The user interface supports simple and complex queries, including operations such as subtraction and intersections as well as clustering and finding elements in proximity to features. dbERGE II, the database of Experimental Results on Gene Expression, contains experimental data from a variety of functional assays. Both databases are now run on the DB2 database management system. Improved hardware and tuning has reduced response times and increased querying capacity, while simplified query interfaces will help direct new users through the querying process. Links are available at http://www.bx.psu.edu/.

Animals↗

A relational database application in support of integrated neuroscience research.

The development of relational databases has significantly improved the performance of storage, search, and retrieval functions and has made it possible for applications that perform real-time data acquisition and analysis to interact with these types of databases. The purpose of this research was to develop a user interface for interaction between a data acquisition and analysis application and a relational database using the Oracle9i system. The overall system was designed to have an indexing capability that threads into the data acquisition and analysis programs. Tables were designed and relations within the database for indexing the files and information contained within the files were established. The system provides retrieval capabilities over a broad range of media, including analog, event, and video data types. The system's ability to interact with a data capturing program at the time of the experiment to create both multimedia files as well as the meta-data entries in the relational database avoids manual entries in the database and ensures data integrity and completeness for further interaction with the data by analysis applications.

Database Management Systems↗

BioBuilder as a database development and functional annotation platform for proteins.

BACKGROUND: The explosion in biological information creates the need for databases that are easy to develop, easy to maintain and can be easily manipulated by annotators who are most likely to be biologists. However, deployment of scalable and extensible databases is not an easy task and generally requires substantial expertise in database development. RESULTS: BioBuilder is a Zope-based software tool that was developed to facilitate intuitive creation of protein databases. Protein data can be entered and annotated through web forms along with the flexibility to add customized annotation features to protein entries. A built-in review system permits a global team of scientists to coordinate their annotation efforts. We have already used BioBuilder to develop Human Protein Reference Database http://www.hprd.org, a comprehensive annotated repository of the human proteome. The data can be exported in the extensible markup language (XML) format, which is rapidly becoming as the standard format for data exchange. CONCLUSIONS: As the proteomic data for several organisms begins to accumulate, BioBuilder will prove to be an invaluable platform for functional annotation and development of customizable protein centric databases. BioBuilder is open source and is available under the terms of LGPL.

Computational Biology↗

Specialized microbial databases for inductive exploration of microbial genome sequences.

BACKGROUND: The enormous amount of genome sequence data asks for user-oriented databases to manage sequences and annotations. Queries must include search tools permitting function identification through exploration of related objects. METHODS: The GenoList package for collecting and mining microbial genome databases has been rewritten using MySQL as the database management system. Functions that were not available in MySQL, such as nested subquery, have been implemented. RESULTS: Inductive reasoning in the study of genomes starts from "islands of knowledge", centered around genes with some known background. With this concept of "neighborhood" in mind, a modified version of the GenoList structure has been used for organizing sequence data from prokaryotic genomes of particular interest in China. GenoChore http://bioinfo.hku.hk/genochore.html, a set of 17 specialized end-user-oriented microbial databases (including one instance of Microsporidia, Encephalitozoon cuniculi, a member of Eukarya) has been made publicly available. These databases allow the user to browse genome sequence and annotation data using standard queries. In addition they provide a weekly update of searches against the world-wide protein sequences data libraries, allowing one to monitor annotation updates on genes of interest. Finally, they allow users to search for patterns in DNA or protein sequences, taking into account a clustering of genes into formal operons, as well as providing extra facilities to query sequences using predefined sequence patterns. CONCLUSION: This growing set of specialized microbial databases organize data created by the first Chinese bacterial genome programs (ThermaList, Thermoanaerobacter tencongensis, LeptoList, with two different genomes of Leptospira interrogans and SepiList, Staphylococcus epidermidis) associated to related organisms for comparison.

Algorithms↗

Characteristics of the U.S. EPA's Office of Pesticide Programs' toxicity information databases.

The United States Environmental Protection Agency's Office of Pesticide Programs (OPP) requires that data from toxicity testing be submitted to the OPP to support the registration of pesticide chemicals. Once the toxicity data are submitted, they are entered into various toxicity databases. The studies are listed in an archival database to catalog and allow retrieval of the study for review. Reviews of toxicity studies are then placed into a separate database that can be retrieved to support a regulatory position. Toxicity information for health effects other than cancer and gene mutations from chronic exposure is reviewed through a reference dose (RfD) approach, and these decisions and supporting data are entered into an RfD database. Carcinogenicity data are reviewed by a peer review process, and these decisions are entered into a newly developed database to show the regulatory decision with supporting data. The mutagenicity data are reviewed and acceptable data are entered into the Genetic Activity Profile system to catalog and display the submitted information. These databases contain the information used for hazard evaluations as part of the OPP review of pesticide chemicals.

Animals↗

HCVDB: hepatitis C virus sequences database.

UNLABELLED: To date, more than 30 000 hepatitis C virus (HCV) sequences have been deposited in the generalist databases DNA Data Bank of Japan (DDBJ), EMBL Nucleotide Sequence Database (EMBL) and GenBank. The main difficulties with HCV sequences in these databases are their retrieval, annotation and analyses. To help HCV researchers face the increasing needs of HCV sequence analyses, we developed a specialised database of computer-annotated HCV sequences, called HCVDB. HCVDB is re-built every month from an up-to-date EMBL database by an automated process. HCVDB provides key data about the HCV sequences (e.g. genotype, genomic region, protein names and functions, known 3-dimensional structures) and ensures consistency of the annotations, which enables reliable keyword queries. The database is highly integrated with sequence and structure analysis tools and the SRS (LION bioscience) keywords query system. Thus, any user can extract subsets of sequences matching particular criteria or enter their own sequences and analyse them with various bioinformatics programs available on the same server. AVAILABILITY: HCVDB is available from http://hepatitis.ibcp.fr.

Amino Acid Sequence↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

Measuring use patterns of online journals and databases.

PURPOSE: This research sought to determine use of online biomedical journals and databases and to assess current user characteristics associated with the use of online resources in an academic health sciences center. SETTING: The Library of the Health Sciences-Peoria is a regional site of the University of Illinois at Chicago (UIC) Library with 350 print journals, more than 4,000 online journals, and multiple online databases. METHODOLOGY: A survey was designed to assess online journal use, print journal use, database use, computer literacy levels, and other library user characteristics. A survey was sent through campus mail to all (471) UIC Peoria faculty, residents, and students. RESULTS: Forty-one percent (188) of the surveys were returned. Ninety-eight percent of the students, faculty, and residents reported having convenient access to a computer connected to the Internet. While 53% of the users indicated they searched MEDLINE at least once a week, other databases showed much lower usage. Overall, 71% of respondents indicated a preference for online over print journals when possible. CONCLUSIONS: Users prefer online resources to print, and many choose to access these online resources remotely. Convenience and full-text availability appear to play roles in selecting online resources. The findings of this study suggest that databases without links to full text and online journal collections without links from bibliographic databases will have lower use. These findings have implications for collection development, promotion of library resources, and end-user training.

Computer Literacy↗

Non-sequence databases for biological activity and physicochemical properties.

A biological activity database and a physicochemical property database are described. They are intended to complement the protein sequence database of PIR-International. The Biological Activity Database and the Physicochemical Property Database contain information regarding the biological activity and the physicochemical properties of proteins, respectively. In addition they also provide information about wild-type molecules with which information concerning variant molecules may be compared. Data on artificial variant molecules are stored in the Artificial Variant Database which is described separately.

Amino Acid Sequence↗

A database model for studies of cocaine-dependent pregnant women and their families.

The database management functions for the Mothers Project are arranged into administrative and analytic task groups, and separate systems are devised for each. The task groups can be distinguished not only by differences in data structure but also by interface requirements. The administrative database system uses a relational database technology, whereas the analytic database system employs more traditional flat-file methods. Although the database management systems are complex, they are based on standard database practices, used in widely available software packages, and run on inexpensive desktop computing equipment.

Cocaine↗

Creation and maintenance of Helix, a Web based database of medical genetics laboratories, to serve the needs of the genetics community.

Helix (healthlinks.washington.edu/helix) is a web accessible database that serves as the main U.S. directory of laboratories offering genetic testing. The database was designed to address the previously unmet need for a centralized, continuously updated source of information about clinical and research genetic testing to keep pace with the rapid rate of gene discovery resulting from the Human Genome Project. The Helix project began in 1992 at the University of Washington and Children's Hospital and Regional Medical Center. It has evolved from a single user stand alone relational database to a fully Web enabled database queried and maintained via the web and linked to other web accessible genomic databases. As of February, 1998 it lists more than 500 diseases and 290 laboratories, with over 5,200 registered users making approximately 250 queries/day (90% via the Internet). We describe the iterative design, implementation, population and assessment of the database over a six year period.

Database Management Systems↗

A Taxonomic Search Engine: federating taxonomic databases using web services.

BACKGROUND: The taxonomic name of an organism is a key link between different databases that store information on that organism. However, in the absence of a single, comprehensive database of organism names, individual databases lack an easy means of checking the correctness of a name. Furthermore, the same organism may have more than one name, and the same name may apply to more than one organism. RESULTS: The Taxonomic Search Engine (TSE) is a web application written in PHP that queries multiple taxonomic databases (ITIS, Index Fungorum, IPNI, NCBI, and uBIO) and summarises the results in a consistent format. It supports "drill-down" queries to retrieve a specific record. The TSE can optionally suggest alternative spellings the user can try. It also acts as a Life Science Identifier (LSID) authority for the source taxonomic databases, providing globally unique identifiers (and associated metadata) for each name. CONCLUSION: The Taxonomic Search Engine is available at http://darwin.zoology.gla.ac.uk/~rpage/portal/ and provides a simple demonstration of the potential of the federated approach to providing access to taxonomic names.

Classification↗

Human gene mutation database-a biomedical information and research resource.

Although 20 years have elapsed since the first single basepair substitution underlying an inherited disease in humans was characterised at the DNA level, the initiative has only recently been taken to establish central database resources for pathological genetic variants. Disease-associated gene lesions are currently collected and publicised by the Human Gene Mutation Database (HGMD) in Cardiff, locus-specific mutation databases, and to some extent also by the Genome Database (GDB) and Online Mendelian Inheritance in Man (OMIM). To date, HGMD represents the only comprehensive and publicly available database of gene lesions underlying human inherited disease. By July 1999, HGMD contained over 18,000 different mutations from some 900 human genes, the majority being single basepair substitutions. In addition to its potential as an information resource for clinicians and genetic counsellors, HGMD has allowed molecular geneticists to address a variety of biological questions through meta-analysis of the collated data. HGMD also promises to assist research workers in optimising mutation search strategies for a given gene. A questionnaire sent out to, and answered by, the editors of 20 key journals revealed that human genetics journals are increasingly reluctant to publish mutation reports. Electronic data submission and publication facilities are therefore urgently required. The World Wide Web (WWW) provides an excellent medium within which to combine the centralised management of basic mutation data, including rigorous quality control, with the possibility of publishing additional mutation-related information. In response to these needs, HGMD has both instituted a collaboration with Springer-Verlag GmbH, Heidelberg, to potentiate free online submission and electronic publication of human gene mutation data and developed links with the curators of locus-specific mutation databases.

Databases, Factual↗

A dynamic two-dimensional polyacrylamide gel electrophoresis database: the mycobacterial proteome via Internet.

Proteome analysis by two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) and mass spectrometry, in combination with protein chemical methods, is a powerful approach for the analysis of the protein composition of complex biological samples. Data organization is imperative for efficient handling of the vast amount of information generated. Thus we have constructed a 2-D PAGE database to store and compare protein patterns of cell-associated and culture-supernatant proteins of different mycobacterial strains. In accordance with the guidelines for federated 2-DE databases, we developed a program that generates a dynamic 2-D PAGE database for the World-Wide-Web to organise and publish, via the internet, our results from proteome analysis of different Mycobacterium tuberculosis as well as Mycobacterium bovis BCG strains. The uniform resource locator for the database is http://www.mpiib-berlin.mpg.de/2D-PAGE and can be read with a Java compatible browser. The interactive hypertext markup language documents displayed are generated dynamically in each individual session from a rational data file, a 2-D gel image file and a map file describing the protein spots as polygons. The program consists of common gateway interface scripts written in PERL, minimizing the administrative workload of the database. Furthermore, the database facilitates not only interactive use, but also worldwide active participation of other scientific groups with their own data, requiring only minimal computer hardware and knowledge of information technology.

Bacterial Proteins↗

The structure and dipole moment of globular proteins in solution and crystalline states: use of NMR and X-ray databases for the numerical calculation of dipole moment.

The large dipole moment of globular proteins has been well known because of the detailed studies using dielectric relaxation and electro-optical methods. The search for the origin of these dipolemoments, however, must be based on the detailed knowledge on protein structure with atomic resolutions. At present, we have two sources of information on the structure of protein molecules: (1) x-ray databases obtained in crystalline state; (2) NMR databases obtained in solution state. While x-ray databases consist of only one model, NMR databases, because of the fluctuation of the protein folding in solution, consist of a number of models, thus enabling the computation of dipole moment repeated for all these models. The aim of this work, using these databases, is the detailed investigation on the interdependence between the structure and dipole moment of protein molecules. The dipole moment of protein molecules has roughly two components: one dipole moment is due to surface charges and the other, core dipole moment, is due to polar groups such as N--H and C==O bonds. The computation of surface charge dipole moment consists of two steps: (A) calculation of the pK shifts of charged groups for electrostatic interactions and (B) calculation of the dipole moment using the pK corrected for electrostatic shifts. The dipole moments of several proteins were computed using both NMR and x-ray databases. The dipole moments of these two sets of calculations are, with a few exceptions, in good agreement with one another and also with measured dipole moments.

Animals↗

Use of mass spectrometric molecular weight information to identify proteins in sequence databases.

During the last decade new ionization techniques have made it possible to measure the molecular weight of many intact proteins by mass spectrometry, and they have made it much easier to obtain a mass spectrometric peptide map of a protein. At the same time advances in protein and DNA sequencing technology are resulting in an exponential increase in the number of sequences deposited in databases. Here we investigate the possibility to use mass spectrometric data to identify proteins in databases. Searching a database by total molecular weight is found to be an easy and sometimes sufficient approach. For more specificity and for error tolerance in both the mass spectrometric data and the database information we search by partial mass spectrometric peptide map of the protein. In general, just four to six proteolytic peptides measured with a mass accuracy between 0.1 and 0.01% allow a useful search of databases such as the Protein Identification Resource (PIR). As the size of DNA and protein sequence databases grows, protein identification by partial mass spectrometric peptide maps should become increasingly powerful and may become a general method to identify and characterize proteins.

Amino Acid Sequence↗

Microcomputer-assisted filing system of cardiac catheterization records using a relational database management system.

To efficiently store and retrieve cardiac catheterization records, we have developed a computer-assisted database, which comprises a 16-bit microcomputer with dual floppy disk drives, a 20 MB random-access memory, hard disk drive, and a line printer. All programmings were accomplished using a relational database management system (R:base 5000, Microrim, Inc.). Data inquiry procedures could be performed with direct operational commands of the system as well as with preprogrammed command files, and final results of searches were printed out with a line printer. The major advantages of the present system described in this report include: (1) the relatively easy and rapid creation of the database, (2) ease of modification of the database structures even after the system design is finished, (3) operational commands in combination with conditional operator(s) are flexible and powerful enough to allow the end user to retrieve data based on various kinds of criteria, (4) a high-level programming language provided by the R:base automates a series of database procedures with relative ease, (5) relational capabilities of the database management system can enhance the possibility of reconstruction of a new data file from a single or several preexisting data files, and (6) the system can be realized at reasonable cost.

Cardiac Catheterization↗