Search PubMed⌕ Search

PubMed · 15153304

Biological database design and implementation.

Abstract

We present our experience of building biological databases. Such databases have most aspects in common with other complex databases in other fields. We do not believe that biological data are that different from complex data in other fields. Our experience has led us to emphasise simplicity and conservative technology choices when building these databases. This is a short paper of advice that we hope is useful to people designing their own biological database.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ewan Birney, Michele Clamp. 2004. Biological database design and implementation.. https://doi.org/10.1093/bib%2F5.1.31

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A comprehensive classification system for lipids.

Lipids are produced, transported, and recognized by the concerted actions of numerous enzymes, binding proteins, and receptors. A comprehensive analysis of lipid molecules, "lipidomics," in the context of genomics and proteomics is crucial to understanding cellular physiology and pathology; consequently, lipid biology has become a major research target of the postgenomic revolution and systems biology. To facilitate international communication about lipids, a comprehensive classification of lipids with a common platform that is compatible with informatics requirements has been developed to deal with the massive amounts of data that will be generated by our lipid community. As an initial step in this development, we divide lipids into eight categories (fatty acyls, glycerolipids, glycerophospholipids, sphingolipids, sterol lipids, prenol lipids, saccharolipids, and polyketides) containing distinct classes and subclasses of molecules, devise a common manner of representing the chemical structures of individual lipids and their derivatives, and provide a 12 digit identifier for each unique lipid molecule. The lipid classification scheme is chemically based and driven by the distinct hydrophobic and hydrophilic elements that compose the lipid. This structured vocabulary will facilitate the systematization of lipid biology and enable the cataloging of lipids and their properties in a way that is compatible with other macromolecular databases.

Database Management Systems↗

MPSS: an integrated database system for surveying a set of proteins.

SUMMARY: We design and implement an integrated database system called 'multi-protein survey system' (MPSS), which provides a platform to retrieve information about many proteins at a time. This system integrates several important and widely used databases including SwissProt, TrEMBL, PDB and InterPro, plus useful references such as GO and KEGG to other databases. Users may submit a group of protein IDs, entry names, SwissProt/TrEMBL accession numbers or GenBank GIs through MPSS' web interface, and obtain protein annotation information from public databases and pre-computed molecular properties speedily. MPSS can also supply comprehensive information about query proteins, including 3D structures, domains, pathway, gene ontology and visual presentation of mapping to the GO tree and KEGG pathway, to provide an up-to-date view of available knowledge with regard to the structures and molecular functions of proteins under study. AVAILABILITY: MPSS is freely accessible at http://www.scbit.org/mpss/

Database Management Systems↗