Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

PATMAT: a searching and extraction program for sequence, pattern and block queries and databases.

A program has been developed that provides molecular biologists with multiple tools for searching databases, yet uses a very simple interface. PATMAT can use protein or (translated) DNA sequences, patterns or blocks of aligned proteins as queries of databases consisting of amino acid or nucleotide sequences, patterns or blocks. The ability to search databases of blocks by 'on-the-fly' conversion to scoring matrices provides a new tool for detection and evaluation of distant relationships. PATMAT uses a pull-down, menu-driven interface to carry out its multiple searching, extraction and viewing functions. Each query or database type is recognized, reported, and the appropriate search carried out, with matches and alignments reported in windows as they occur. Any of the high scoring matches can be exported to a file, viewed and recalled as a query using only a few keystrokes or mouse selections. Searches of multiple database files are carried out by user selection within a window. PATMAT runs under DOS; the searching engine also runs under UNIX.

Amino Acid Sequence↗

Searching for amino acid sequence motifs among enzymes: the Enzyme-Reaction Database.

Recently we have constructed a database--the Enzyme-Reaction Database--which links a chemical structure to amino acid sequences of enzymes that recognize the chemical structure as their ligand. The total number of enzymes registered in the database is 1103 with 6668 NBRF-PIR entry codes and 1756 chemical compounds. The chemical structures and chemical names for 842 compounds are registered in the Chemical-Structure Database on the MACCS system. For each enzyme, the sequences were divided into clusters, and multiply aligned in each cluster to extract a conserved sequence. A total of 158,781 five-residue-long fragments were constructed from 433 conserved sequences and compared among different clusters of different enzymes. One of these motifs shared by different enzymes was S-G-G-L-D. The motif was conserved in both argininosuccinate synthase (EC 6.3.4.5) and asparagine synthase (glutamine-hydrolysing) (EC 6.3.5.4). This result showed that the database was useful for the analysis of the relationship between chemical structures and amino acid sequence motifs.

Algorithms↗

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

A comprehensive and non-redundant database of protein domain movements.

MOTIVATION: The current DynDom database of protein domain motions is a user-created database that suffers from selectivity and redundancy. The aim of the analysis presented here was to overcome both these limitations and to produce both a comprehensive and a non-redundant description of domain movements from structures stored in the current protein data bank. RESULTS: A multi-step procedure is applied that starts with grouping proteins in the structural databank into families based on sequence similarity. Multiple sequence alignment, conformational clustering and a dimensional clustering method based on the Gram-Schmidt algorithm are applied to members of each family to remove dynamic redundancy in their domain movements. Representative domain movements are described in terms of domains, hinge axes and hinge-bending residues using the DynDom program. The results show that within an average family of 11.5 members, there are on average only 1.31 different domain movements indicating a high redundancy in the movements these structures represent. This verifies earlier findings that domain movements are usually highly controlled. Despite the removal of this considerable redundancy, the process has resulted in double the number of domain movements stored in the user-created database. The data are organized in a relational database with a web-interface. AVAILABILITY: The database can be browsed and searched at http://www.cmp.uea.ac.uk/dyndom CONTACT: sjh@cmp.uea.ac.uk.

Algorithms↗

A comprehensive approach for establishment of the platform to analyze functions of KIAA proteins II: public release of inaugural version of InGaP database containing gene/protein expression profiles for 127 mouse KIAA genes/proteins.

The inaugural version of the InGaP database (Integrative Gene and Protein expression database; http://www.kazusa.or.jp/ingap/index.html) is a comprehensive database of gene/protein expression profiles of 127 mKIAA genes/proteins related to hypothetical ones obtained in our ongoing cDNA project. Information about each gene/protein consists of cDNA microarray analysis, subcellular localization of the ectopically expressed gene, and experimental data using anti-mKIAA antibody such as Western blotting and immunohistochemical analyses. KIAA cDNAs and their mouse counterparts, mKIAA cDNAs, were mainly isolated from cDNA libraries derived from brain tissues, thus we expect our database to contribute to the field of neuroscience. In fact, cDNA microarray analysis revealed that nearly half of our gene collection is predominantly expressed in brain tissues. Immunohistochemical analysis of the mouse brain provides functional insight into the specific area and/or cell type of the brain. This database will be a resource for the neuroscience community by seamlessly integrating the genomic and proteomic information about the mouse KIAA genes/proteins.

Animals↗

Development of a food database of nitrosamines, heterocyclic amines, and polycyclic aromatic hydrocarbons.

Some nitrosocompounds that are formed during food preservation, as well as polycyclic aromatic hydrocarbons (PAH) and heterocyclic amines (HA) formed during cooking, may have carcinogenic activity. An accurate assessment of dietary intake of such compounds is difficult, mainly because they are not naturally present in foods, and they are not included in standard food composition tables. Our objective was to develop a food composition database of nitrates, nitrites, nitrosamines, HA, and PAH. We conducted a literature search on the food content of these compounds using the Medline and EMBASE databases. We gathered the following information: 1) Food information: name, cooking methods, preservation methods, cooking doneness, temperature, and time; 2) compound information: type, quantity, value type, analytic method, and sampling methods; and 3) publication information: year, author, and country. We developed a table that includes 207 food items with information concerning the concentration of nitrites, nitrates, and nitrosamines, 297 food items with information about HA concentration and 313 food items with information about PAH. The database is based on 139 references from 23 different countries. It is arranged according to compounds and food groups to facilitate its practical use. The potential limitations are due to the quality of the information we could obtain through Medline and EMBASE databases. This database will allow investigators to quantify dietary exposure to several potential carcinogens, and to analyze their relation to the risk of cancer.

Carcinogens↗

Development of a glycemic index database for food frequency questionnaires used in epidemiologic studies.

Consumption of foods with a high glycemic index (GI) or glycemic load (GL) is hypothesized to contribute to insulin resistance, which is associated with increased risk of diabetes mellitus, obesity, cardiovascular disease, and some cancers. However, dietary assessment of GI and GL is difficult because values are not included in standard food composition databases. Our objective was to develop a database of GI and GL values that could be integrated into an existing dietary database used for the analysis of FFQ. Food GI values were obtained from published human experimental studies or imputed from foods with a similar carbohydrate and fiber content. We then applied the values to the Women's Health Initiative (WHI) FFQ database and tested the output in a random sample of previously completed WHI FFQs. Of the 122 FFQ line items (disaggregated into 350 foods), 83% had sufficient carbohydrate (>5 g/serving) for receipt of GI and GL values. The foods on the FFQ food list with the highest GL were fried breads, potatoes, pastries, pasta, and soft drinks. The fiber content of foods had very little influence on calculated GI or GL estimates. The augmentation of this FFQ database with GI and GL values will enable etiologic investigations of GI and GL with numerous disease outcomes in the WHI and other epidemiologic studies that utilize this FFQ.

Databases, Factual↗

The mammalian gene mutation database.

The Mammalian Gene Mutation Database (MGMD) is a comprehensive collection of published mutation data from the open literature on mammalian cell-based gene model mutation detection systems. The database currently contains approximately 30000 comprehensively described mutant spectra records and it is maintained and up- dated on a daily basis. The major objectives of the MGMD were (i) to provide an Internet-accessible database (http://lisntweb.swan.ac. uk/cmgt/index.htm) for chemically induced and spontaneous mutation types and spectra in selected genes; (ii) to standardize the reporting of mutations within different genes where ambiguity exists in the literature; and (iii) to provide interactive and user-friendly access to the information. A multi-option search facility has been included that allows the user to search the database for parameters such as mutagen, gene or cell type of interest. The structure of the database permits easy retrieval of specific mutation data for further analysis. Thus, the MGMD should become a useful and necessary reference source and provides an analysis tool for genetic toxicologists.

Animals↗

Corruption of genomic databases with anomalous sequence.

We describe evidence that DNA sequences from vectors used for cloning and sequencing have been incorporated accidentally into eukaryotic entries in the GenBank database. These incorporations were not restricted to one type of vector or to a single mechanism. Many minor instances may have been the result of simple editing errors, but some entries contained large blocks of vector sequence that had been incorporated by contamination or other accidents during cloning. Some cases involved unusual rearrangements and areas of vector distant from the normal insertion sites. Matches to vector were found in 0.23% of 20,000 sequences analyzed in GenBank Release 63. Although the possibility of anomalous sequence incorporation has been recognized since the inception of GenBank and should be easy to avoid, recent evidence suggests that this problem is increasing more quickly than the database itself. The presence of anomalous sequence may have serious consequences for the interpretation and use of database entries, and will have an impact on issues of database management. The incorporated vector fragments described here may also be useful for a crude estimate of the fidelity of sequence information in the database. In alignments with well-defined ends, the matching sequences showed 96.8% identity to vector; when poorer matches with arbitrary limits were included, the aggregate identity to vector sequence was 94.8%.

Base Sequence↗

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence↗

Histone Sequence Database: a compilation of highly-conserved nucleoprotein sequences.

By searching the current protein sequence databases using sequences from human and chicken histones H1/H5, H2A, H2B, H3 and H4, a database of aligned histone protein sequences with statistically significant sequence similarity to the search sequence was constructed. In addition, a nucleotide sequence database of the corresponding coding regions for these proteins has been assembled. The region of each of the core histones containing the histone fold motif is identified in the protein alignments. The database contains >1300 protein and nucleotide sequences. All sequences and alignments in this database are available through the World Wide Web at http://www.ncbi.nlm.nih.gov/Baxevani/HISTO NES.

Amino Acid Sequence↗

The PAH mutation analysis consortium database: update 1996.

A website (http://www.mcgill.ca/pahdb ) is maintained by the curators for a Consortium (88 investigators, 28 countries) and all other users; it serves a relational database for human locus-specific genetic variation in a defined DNA sequence (GenBank U49897); (100 kb on human chromosome 12q24.1, gene symbol PAH). The intragenic nucleotide variation is both rare (Q< 0.01), extensive (>320 different mutations) and phenotype modifying, causing hyperphenylalaninemia by impairing phenylalanine hydroxylase function (see OMIM 261600), as well as polymorphic and neutral, the latter providing informative locus-specific haplotypes (>1200 different mutation/haplotype associations). The PAH database contains both offline core components (mutations, population associations and data source information) and several accessory online components: (i) relative frequencies of mutations by populations/regions (expanding file); (ii) data on genotype- phenotype correlations both in vitro and in vivo (new file); (iii) polymorphic haplotype structures (new file); (iv) intron sequence data (new file for design of primers); (v) description of mouse homologues (new file for mutations and phenotypes); (vi) the predicted PAH gene mutability profile (improved graphic); (vii) a clinical field for patient use (new interface with database). The website home page has been revised and a counter is recording >15 visits per day. Linkages to other mutation databases and an alliance of mutation database curators (new) are expanding. The primary 'electronic publication' reports now vastly exceed print reports. PAHdb serves as a prototype for obtaining, storing and distributing records of human genetic variation.

Animals↗

The GDB Human Genome Database Anno 1997.

The value of the Genome Database (GDB) for the human genome research community has been greatly increased since the release of version 6. 0 last year. Thanks to the introduction of significant technical improvements, GDB has seen dramatic growth in the type and volume of information stored in the database. This article summarizes the types of data that are now available in the Genome Database, demonstrates how the database is interconnected with other biomedical resources on the World Wide Web, discusses how researchers can contribute new or updated information to the database, and describes our current efforts as well as planned improvements for the future.

Base Sequence↗

The Organelle Genome Database Project (GOBASE).

The taxonomically broad organelle genome database (GOBASE) organizes and integrates diverse data related to organelles (mitochondria and chloroplasts). The current version of GOBASE focuses on the mitochondrial subset of data and contains molecular sequences, RNA secondary structures and genetic maps, as well as taxonomic information for all eukaryotic species represented. The database has been designed so that complex biological queries, especially ones posed in a comparative genomics context, are supported. GOBASE has been implemented as a relational database with a web-based user interface (http://megasun.bch.umontreal.ca/gobase/gobas e.html ). Custom software tools have been written in house to assist in the population of the database, data validation, nomenclature standardization and front-end design. The database is fully operational and publicly accessible via the World Wide Web, allowing interactive browsing, sophisticated searching and easy downloading of data.

Amino Acid Sequence↗

The Genome Sequence DataBase (GSDB): improving data quality and data access.

In 1997 the primary focus of the Genome Sequence DataBase (GSDB; www. ncgr.org/gsdb ) located at the National Center for Genome Resources was to improve data quality and accessibility. Efforts to increase the quality of data within the database included two major projects; one to identify and remove all vector contamination from sequences in the database and one to create premier sequence sets (including both alignments and discontiguous sequences). Data accessibility was improved during the course of the last year in several ways. First, a graphical database sequence viewer was made available to researchers. Second, an update process was implemented for the web-based query tool, Maestro. Third, a web-based tool, Excerpt, was developed to retrieve selected regions of any sequence in the database. And lastly, a GSDB flatfile that contains annotation unique to GSDB (e.g., sequence analysis and alignment data) was developed. Additionally, the GSDB web site provides a tool for the detection of matrix attachment regions (MARs), which can be used to identify regions of high coding potential. The ultimate goal of this work is to make GSDB a more useful resource for genomic comparison studies and gene level studies by improving data quality and by providing data access capabilities that are consistent with the needs of both types of studies.

Base Sequence↗

Superior performance in protein homology detection with the Blocks Database servers.

The Blocks Database World Wide Web (http://www.blocks.fhcrc.org ) and Email (blocks@blocks.fhcrc.org) servers provide tools for the detection and analysis of protein homology based on alignment blocks representing conserved regions of proteins. During the past year, searching has been augmented by supplementation of the Blocks Database with blocks from the Prints Database, for a total of 4754 blocks from 1163 families. Blocks from both the Blocks and Prints Databases and blocks that are constructed from sequences submitted to Block Maker can be used for blocks-versus-blocks searching of these databases with LAMA, and for viewing logos and bootstrap trees. Sensitive searches of up-to-date protein sequence databanks are carried out via direct links to the MAST server using position-specific scoring matrices and to the BLAST and PSI-BLAST servers using consensus-embedded sequence queries. Utilizing the trypsin family to evaluate performance, we illustrate the superiority of blocks-based tools over expert pairwise searching or Hidden Markov Models.

Amino Acid Sequence↗

MIPS: a database for protein sequences and complete genomes.

The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled.

Amino Acid Sequence↗

Databases on transcriptional regulation: TRANSFAC, TRRD and COMPEL.

TRANSFAC, TRRD (Transcription Regulatory Region Database) and COMPEL are databases which store information about transcriptional regulation in eukaryotic cells. The three databases provide distinct views on the components involved in transcription: transcription factors and their binding sites and binding profiles (TRANSFAC), the regulatory hierarchy of whole genes (TRRD), and the structural and functional properties of composite elements (COMPEL). The quantitative and qualitative changes of all three databases and connected programs are described. The databases are accessible via WWW:http://transfac.gbf.de/TRANSFAC orhttp://www.bionet.nsc.ru/TRRD

Animals↗