Search PubMed⌕ Search

Biomedical subjects

Shawn M Douglas

Publications and source records attributed to Shawn M Douglas.

6 recordsLinked to original sources

Network security and data integrity in academia: an assessment and a proposal for large-scale archiving.

A direct impediment to the optimal use of online databases is the increasing prevalence, severity, and toll of computer and network security incidents. Funding agencies should set up working groups that can provide essential services such as universal backup, archival storage, and mirroring of community resources, consistent with the key goal of security in academia: to preserve data and results for posterity.

Archives↗

PubNet: a flexible system for visualizing literature derived networks.

We have developed PubNet, a web-based tool that extracts several types of relationships returned by PubMed queries and maps them into networks, allowing for graphical visualization, textual navigation, and topological analysis. PubNet supports the creation of complex networks derived from the contents of individual citations, such as genes, proteins, Protein Data Bank (PDB) IDs, Medical Subject Headings (MeSH) terms, and authors. This feature allows one to, for example, examine a literature derived network of genes based on functional similarity.

Databases, Bibliographic↗

Robotic cloning and Protein Production Platform of the Northeast Structural Genomics Consortium.

In this chapter we describe the core Protein Production Platform of the Northeast Structural Genomics Consortium (NESG) and outline the strategies used for producing high-quality protein samples using Escherichia coli host vectors. The platform is centered on 6X-His affinity-tagged protein constructs, allowing for a similar purification procedure for most targets, and the implementation of high-throughput parallel methods. In most cases, these affinity-purified proteins are sufficiently homogeneous that a single subsequent gel filtration chromatography step is adequate to produce protein preparations that are greater than 98% pure. Using this platform, over 1000 different proteins have been cloned, expressed, and purified in tens of milligram quantities over the last 36-month period (see Summary Statistics for All Targets, ). Our experience using a hierarchical multiplex expression and purification strategy, also described in this chapter, has allowed us to achieve success in producing not only protein samples but also many three-dimensional structures. As of December 2004, the NESG Consortium has deposited over 145 new protein structures to the Protein Data Bank (PDB); about two-thirds of these protein samples were produced by the NESG Protein Production Facility described here. The methods described here have proven effective in producing quality samples of both eukaryotic and prokaryotic proteins. These improved robotic and?or parallel cloning, expression, protein production, and biophysical screening technologies will be of broad value to the structural biology, functional proteomics, and structural genomics communities.

Chromatography, Gel↗

Mining the structural genomics pipeline: identification of protein properties that affect high-throughput experimental analysis.

Structural genomics projects represent major undertakings that will change our understanding of proteins. They generate unique datasets that, for the first time, present a standardized view of proteins in terms of their physical and chemical properties. By analyzing these datasets here, we are able to discover correlations between a protein's characteristics and its progress through each stage of the structural genomics pipeline, from cloning, expression, purification, and ultimately to structural determination. First, we use tree-based analyses (decision trees and random forest algorithms) to discover the most significant protein features that influence a protein's amenability to high-throughput experimentation. Based on this, we identify potential bottlenecks in various stages of the structural genomics process through specialized "pipeline schematics". We find that the properties of a protein that are most significant are: (i.) whether it is conserved across many organisms; (ii). the percentage composition of charged residues; (iii). the occurrence of hydrophobic patches; (iv). the number of binding partners it has; and (v). its length. Conversely, a number of other properties that might have been thought to be important, such as nuclear localization signals, are not significant. Thus, using our tree-based analyses, we are able to identify combinations of features that best differentiate the small group of proteins for which a structure has been determined from all the currently selected targets. This information may prove useful in optimizing high-throughput experimentation. Further information is available from http://mining.nesg.org/.

Algorithms↗

SPINE 2: a system for collaborative structural proteomics within a federated database framework.

We present version 2 of the SPINE system for structural proteomics. SPINE is available over the web at http://nesg.org. It serves as the central hub for the Northeast Structural Genomics Consortium, allowing collaborative structural proteomics to be carried out in a distributed fashion. The core of SPINE is a laboratory information management system (LIMS) for key bits of information related to the progress of the consortium in cloning, expressing and purifying proteins and then solving their structures by NMR or X-ray crystallography. Originally, SPINE focused on tracking constructs, but, in its current form, it is able to track target sample tubes and store detailed sample histories. The core database comprises a set of standard relational tables and a data dictionary that form an initial ontology for proteomic properties and provide a framework for large-scale data mining. Moreover, SPINE sits at the center of a federation of interoperable information resources. These can be divided into (i) local resources closely coupled with SPINE that enable it to handle less standardized information (e.g. integrated mailing and publication lists), (ii) other information resources in the NESG consortium that are inter-linked with SPINE (e.g. crystallization LIMS local to particular laboratories) and (iii) international archival resources that SPINE links to and passes on information to (e.g. TargetDB at the PDB).

Cooperative Behavior↗