Search PubMed⌕ Search

PubMed · 10089484

Macromolecular structure databases: past progress and future challenges.

Abstract

Databases containing macromolecular structure data provide a crystallographer with important tools for use in solving, refining and understanding the functional significance of their protein structures. Given this importance, this paper briefly summarizes past progress by outlining the features of the significant number of relevant databases developed to date. One recent database, PDB+, containing all current and obsolete structures deposited with the Protein Data Bank (PDB) is discussed in more detail. PDB+ has been used to analyze the self-consistency of the current (1 January 1998) corpus of over 7000 structures. A summary of those findings is presented (a full discussion will appear elsewhere) in the form of global and temporal trends within the data. These trends indicate that challenges exist if crystallographers are to provide the community with complete and consistent structural results in the future. It is argued that better information management practices are required to meet these challenges.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

H Weissig, I N Shindyalov, P E Bourne. 1998-11-01. Macromolecular structure databases: past progress and future challenges.. https://doi.org/10.1107/s0907444998009846

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Retention index database for identification of general green leaf volatiles in plants by coupled capillary gas chromatography-mass spectrometry.

A series of ubiquitously occurring saturated and monounsaturated six-carbon aldehydes, alcohols and esters thereof is summarised as 'green leaf volatiles' (GLVs). The present study gives a comprehensive data collection of retention indices of 35 GLVs on commonly used non-polar DB-5, mid-polar DB-1701, and polar DB-Wax stationary phases. Seventeen commercially not available compounds were synthesised. Thus, the present study allows reliable identification of most known GLV in natural plant volatile samples. Applications revealed the presence of several seldom reported GLVs in headspace samples of mechanically damaged plant leaves of Carpinus betulus and Fagus sylvatica.

Database Management Systems↗

Toward general methods of targeted library design: topomer shape similarity searching with diverse structures as queries.

A promising strategy for selecting synthetic targets is similarity-based searching of very large "virtual libraries", which comprise all structures accessible by linking two or three commercially available building blocks with combinatorial syntheses. To assess the general applicability of this strategy, leading structures taken from each of 34 recent medicinal chemistry publications were used as queries to search a virtual library containing 2.6 x 10(13) products from seven reactions, using a topomer shape similarity metric. Eighty-five percent of these searches succeeded, by yielding, with a search radius no greater than 120 topomer shape units, either at least 400 hits or hits from at least six sublibraries. From these 34 sets of search results, 122 representative structures were selected, illustrating potential "lead hops", or otherwise novel structures. Overall shape similarity to the query structure was confirmed for up to 95% of these representative structures, according to FLEXS, an algorithmically distinct program. Experimentally, there were 28 structures among those reported in the 34 query publications that were identified within the virtual library. Among these, the frequency of high activity was 87% for the 16 structures whose similarity to their query was 90 topomer units or less, compared to a frequency of 50% for the other 12 structures.

Database Management Systems↗

NCBI's LocusLink and RefSeq.

The NCBI has introduced two new web resources-LocusLink and RefSeq-that facilitate retrieval of gene-based information and provide reference sequence standards. These resources are designed to provide a non-redundant view of current knowledge about human genes, transcripts and proteins. Additional information about these resources is available on the LocusLink web site at http://www.ncbi.nlm.nih.gov/LocusLink/

Database Management Systems↗