Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

The FSSP database: fold classification based on structure-structure alignment of proteins.

The FSSP database presents a continuously updated classification of 3-D protein folds based on an all-against-all comparison of structures currently in the Protein Data Bank (PDB) [Bernstein et al. (1977) J. Mol. Biol., 112, 535- 542]. The database currently contains an extended structural family for each of 600 representative protein chains which have <25% mutual sequence identity. The results of the exhaustive pairwise structure comparisons are reported in the form of a fold tree generated by hierarchical clustering and as a series of structurally representative sets of folds at varying levels of uniqueness. For each query structure from the representative set, there is a database entry containing structure-structure alignments with its structural neighbours in the representative set and its sequence homologs in the PDB. All alignments are based purely on the 3-D co-ordinates of the proteins and are derived by an automatic structure comparison program (Dali). The FSSP database is accessible electronically on the World Wide Web and by anonymous ftp.

Amino Acid Sequence↗

Codon usage tabulated from the international DNA sequence databases.

Codon usage in 87 602 genes has been calculated using the nucleotide sequence data obtained from the GenBank Genetic Sequence Data Bank (Release 90.0; September 1995). The database is called the CUTG Database; the complete form of the database can be obtained by anonymous ftp from DDBJ and a part of the database, which lists the frequency of codon use in each organism, is made searchable through our World Wide Web server.

Base Sequence↗

GRBase, a database linking information on proteins involved in gene regulation.

The Gene Regulation Database (GRBase) is a compendium of information on the structure and function of proteins involved in the control of gene expression in eukaryotes. These proteins include transcription factors, proteins involved in signal transduction, and receptors. GRBase is now accessible via the World Wide Web (http://www.access.digex.net/regulate). A key feature of this database is the linking of each entry to data in other databases. The database is also available by anonymous ftp (URL ftp://ftp.trevigen.com/pub/Tfactors/) in both text and Filemaker pro formats.

Computer Communication Networks↗

O-GLYCBASE: a revised database of O-glycosylated proteins.

O-GLYCBASE is a comprehensive database of information on glycoproteins and their O-linked glycosylation sites. Entries are compiled and revised from the SWISS-PROT and PIR databases as well as directly from recently published reports. Nineteen percent of the entries extracted from the databases needed revision with respect to O-linked glycosylation. Entries include information about species, sequence, glycosylation site and glycan type, and are fully referenced. Sequence logos displaying the acceptor specificity for the GaINAc transferase are shown. A neural network method for prediction of mucin type O-glycosylation sites in mammalian glycoproteins exclusively from the primary sequence is made available by E-mail or WWW. The O-GLYCBASE database is also available electronically through our WWW server or by anonymous FTP.

Amino Acid Sequence↗

NRSub: a non-redundant database for Bacillus subtilis.

In the context of the international project aimed at sequencing the whole genome of Bacillus subtilis we have developed a non-redundant, fully annotated database of sequences from this organism. Starting from the B.subtilis sequences available in the EMBL, GenBank and DDBJ collections we have removed all encountered duplications and then added extra annotations to the sequences (e.g. accession numbers for the genes, locations on the genetic map, codon usage, etc.) We have also added cross-references to the EMBL, MEDLINE, SWISS-PROT and ENZYME data banks. The present system results from merging of the NRSub and SubtiList databases and the sequence contigs used in the two systems are identical. NRSub is distributed as a flatfile in EMBL format (which is supported by most sequence analysis software packages) and as an ACNUC database, while SubtiList is distributed as a relational database under 4th Dimension. It is possible to access the data through two dedicated World Wide Web servers located in France and Japan.

Bacillus subtilis↗

LISTA, LISTA-HOP and LISTA-HON: a comprehensive compilation of protein encoding sequences and its associated homology databases from the yeast Saccharomyces.

We continued our effort to make a comprehensive database (LISTA) for the yeast Saccharomyces cerevisiae. As in previous editions the genetic names are consistently associated to each sequence with a known and confirmed ORF. If necessary, synonyms are given in the case of allelic duplicated sequences. Although the first publication of a sequence gives-according to our rules-the genetic name of a gene, in some instances more commonly used names are given to avoid nomenclature problems and the use of ancient designations which are no longer used. In these cases the old designation is given as synonym. Thus sequences can be found either by the name or by synonyms given in LISTA. Each entry contains the genetic name, the mnemonic from the EMBL data bank, the codon bias, reference of the publication of the sequence, Chromosomal location as far as known, SWISSPROT and EMBL accession numbers. New entries will also contain the name from the systematic sequencing efforts. Since the release of LISTA4.1 we update the database continuously. To obtain more information on the included sequences, each entry has been screened against non-redundant nucleotide and protein data bank collections resulting in LISTA-HON and LISTA-HOP. This release includes reports from full Smith and Watermann peptide-level searches against a non-redundant protein sequence database. The LISTA data base can be linked to the associated data sets or to nucleotide and protein banks by the Sequence Retrieval System (SRS). The database is available by FTP and on World Wide Web.

Amino Acid Sequence↗

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence↗

The Factor VIII Mutation Database on the World Wide Web: the haemophilia A mutation, search, test and resource site. HAMSTeRS update (version 3.0).

The HAMSTeRS WWW site was set up in 1996 in order to facilitate easy access to, and aid understanding of, the causes of haemophilia A at the molecular level; previously, the first and second text editions of the database have been published in Nucleic Acids Research. This report describes the facilities originally available at the site and the recent additions which we have made to increase its usefulness to clinicians, the molecular genetics community and structural biologists interested in factor VIII. The database (version 3.0) has been completely updated with easy submission of point mutations, deletions and insertions via e-mail of custom-designed forms. The searching of point mutations in the database has been made simpler and more robust, with a concomitantly expanded real-time bioinformatic analysis of the database. A methods section devoted to mutation detection has been added, highlighting issues such as choice of technique and PCR primer sequences. Finally, a FVIII structure section gives access to 3D VRML (Virtual Reality Modelling Language) files for any user-definable residue in a FVIII A domain homology model based on the crystal structure of human caeruloplasmin, together with secondary structural data and a sound+video animation of the model. It is intended that the general availability of this model will assist both in interpretation of causative mutations and selection of candidate residues forin vitromutagenesis. The HAMSTeRS URL is http://europium.mrc.rpms.ac.uk.

Computer Communication Networks↗

The androgen receptor gene mutations database.

The current version of the androgen receptor (AR) gene mutations database is described. The total number of reported mutations has risen from 212 to 272. We have expanded the database: (i) by adding a large amount of new data on somatic mutations in prostatic cancer tissue; (ii) by defining a new constitutional phenotype, mild androgen insensitivity (MAI); (iii) by placing additional relevant information on an internet site (http://www.mcgill.ca/androgendb/ ). The database has allowed us to examine the contribution of CpG sites to the multiplicity of reports of the same mutation in different families. The database is also available from EMBL (ftp.ebi.ac.uk/pub/databases/androgen) or as a Macintosh Filemaker Pro or Word file (MC33@musica,mcgill.ca)

Computer Communication Networks↗

IMGT, the international ImMunoGeneTics database.

IMGT, the international ImMunoGeneTics database, is an integrated database specializing in immunoglobulins, T-cell receptors (TcR) and major histocompatibility complex (MHC) of all vertebrate species, initiated and co-ordinated by Marie-Paule Lefranc, CNRS, Montpellier II University, Montpellier, France (lefranc@ligm.crbm.cnrs-mop.fr). IMGT includes two databases: LIGM-DB (for immunoglobulins and TcR) and MHC/HLA-DB. IMGT comprises expertly annotated sequences and alignment tables. LIGM-DB contains more than 19 000 immunoglobulin and TcR sequences from 78 species. MHC/HLA-DB contains class I and class II human leukocyte antigen alignment tables. An IMGT tool, DNAPLOT, developed for immunoglobulins, TcR and MHC sequence alignments, is also available. IMGT works in close collaboration with the EMBL database. IMGT goals are to establish a common data access to all immunogenetics data, including sequences, oligonucleotide primers, gene maps and other genetic data of immunoglobulins, TcR and MHC molecules, and to provide a graphical user-friendly data access. IMGT will have important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas), therapeutical approaches (antibody engineering), genome diversity and genome evolution studies. IMGT can be accessed at http://imgt.cnusc.fr:8104 and http://www.ebi.ac.uk/IMGT

Amino Acid Sequence↗

Novel developments with the PRINTS protein fingerprint database.

The PRINTS database of protein family 'fingerprints' is a diagnostic resource that complements the PROSITE dictionary of sites and patterns. Unlike regular expressions, fingerprints exploit groups of conserved motifs within sequence alignments to build characteristic signatures of family membership. Thus fingerprints inherently offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 600 fingerprints have been constructed and stored in PRINTS, representing a 50% increase in the size of the database in the last year. The current version, 13.0, encodes approximately 3000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via UCL's Bioinformatics World Wide Web (WWW) server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser / . We describe here progress with the database, its Web interface, and a recent exciting development: the integration of a novel colour alignment editor (http://www.biochem.ucl.ac.uk/bsm/dbbrowser++ +/CINEMA ), which allows visualisation and interactive manipulation of PRINTS alignments over the Internet.

Amino Acid Sequence↗

TRANSFAC, TRRD and COMPEL: towards a federated database system on transcriptional regulation.

Three databases that provide data on transcriptional regulation are described. TRANSFAC is a database on transcription factors and their DNA binding sites. TRRD (Transcription Regulatory Region Database) collects information about complete regulatory regions, their regulation properties and architecture. COMPEL comprises specific information on composite regulatory elements. Here, we describe the present status of these databases and the first steps towards their federation.

Animals↗

O-GLYCBASE version 2.0: a revised database of O-glycosylated proteins.

O-GLYCBASE is an updated database of information on glycoproteins and their O-linked glycosylation sites. Entries are compiled and revised from the literature, and from the SWISS-PROT database. Entries include information about species, sequence, glycosylation sites and glycan type. O-GLYCBASE is now fully cross-referenced to the SWISS-PROT, PIR, PROSITE, PDB, EMBL, HSSP, LISTA and MIM databases. Compared with version 1.0 the number of entries have increased by 34%. Revision of the O-glycan assignment was performed on 20% of the entries. Sequence logos displaying the acceptor specificity patterns for the GalNAc, mannose and GlcNAc transferases are shown. The O-GLYCBASE database is available through WWW or by anonymous FTP.

Algorithms↗

The NRSub database: update 1997.

In the context of the international project aiming at sequencing the whole genome of Bacillus subtilis we have developed NRSub, a non-redundant database of sequences from this organism. Starting from the B.subtilis sequences available in the repository collections we have removed all encountered duplications, then we have added extra annotations to the sequences (e.g. accession numbers for the genes, locations on the genetic map, codon usage index). We have also added cross-references with EMBL/GenBank/DDBJ, MEDLINE, SWISS-PROT and ENZYME databases. NRSub is distributed through anonymous FTP as a text file in EMBL format and as an ACNUC database. It is also possible to access the database through two dedicated World Wide Web servers located in France (http://acnuc.univ-lyon1.fr/nrsub/nrsub.++ +html ) and in Japan (http://ddbjs4h.genes.nig.ac.jp/ ).

Academies and Institutes↗

GIF-DB, a WWW database on gene interactions involved in Drosophila melanogaster development.

GIF-DB (Gene Interactions in the Fly Database) is a new WWW database (http://www-biol.univ-mrs.fr/ approximately lgpd/GIFTS_home_page. html ) describing gene molecular interactions involved in the process of embryonic pattern formation in the flyDrosophila melanogaster. The detailed information is distributed in specific lines arranged into an EMBL- (or SWISS-PROT-) like format. GIF-DB achieves a high level of integration with other databases such as FlyBase, EMBL and SWISS-PROT through numerous hyperlinks. The original concept of interaction databases examplified by GIF-DB could be extended to other biological subjects and organisms so as to study gene regulatory networks in an evolutionary perspective.

Animals↗

HuGeMap: a distributed and integrated Human Genome Map database.

The HuGeMap database stores the major genetic and physical maps of the human genome. It is also interconnected with the gene radiation hybrid mapping database RHdb. HuGeMap is accessible through a Web server for interactive browsing at URL http://www.infobiogen. fr/services/Hugemap , as well as through a CORBA server for effective programming. HuGeMap is intended as an attempt to build open, interconnected databases, that is databases that distribute their objects worldwide in compliance with a recognized standard of distribution. Maps can be displayed and compared with a java applet (http://babbage.infobiogen.fr:15000/Mappet/Show. html ) that queries the HuGeMap ORB server as well as the RHdb ORB server at the EBI.

Chromosome Mapping↗

The Androgen Receptor Gene Mutations Database.

The current version of the androgen receptor (AR) gene mutations database is described. The total number of reported mutations has risen from 272 to 309 in the past year. We have expanded the database: (i) by giving each entry an accession number; (ii) by adding information on the length of polymorphic polyglutamine (polyGln) and polyglycine (polyGly) tracts in exon 1; (iii) by adding information on large gene deletions; (iv) by providing a direct link with a completely searchable database (courtesy EMBL-European Bioinformatics Institute). The addition of the exon 1 polymorphisms is discussed in light of their possible relevance as markers for predisposition to prostate or breast cancer. The database is also available on the internet (http://www.mcgill. ca/androgendb/ ), from EMBL-European Bioinformatics Institute (ftp. ebi.ac.uk/pub/databases/androgen ), or as a Macintosh FilemakerPro or Word file (MC33@musica.mcgill.ca).

Computer Communication Networks↗

The Human PAX6 Mutation Database.

The Human PAX6 Mutation Database contains details of 94 mutations of the PAX6 gene. A Microsoft Access program is used by the Curator to store, update and search the database entries. Mutations can be entered directly by the Curator, or imported from submissions made via the World Wide Web. The PAX6 Mutation Database web page at URL http://www.hgu.mrc.ac.uk/Softdata/PAX6/ provides information about PAX6, as well as a fill-in form through which new mutations can be submitted to the Curator. A search facility allows remote users to query the database. A plain text format file of the data can be downloaded via the World Wide Web. The Curation program contains prior knowledge of the genetic code and of the PAX6 gene including cDNA sequence, location of intron/exon boundaries, and protein domains, so that the minimum of information need be provided by the submitter or Curator.

Computer Communication Networks↗