Search PubMed⌕ Search

Biomedical subjects

A D Michie

Publications and source records attributed to A D Michie.

8 recordsLinked to original sources

Classifying a protein in the CATH database of domain structures.

The CATH database of protein domain structures classifies structures according to their (C)lass, (A)rchitecture, (T)opology or fold and (H)omologous family (http://www.biochem.ucl.ac.uk/bsm/cath). Although the protocol used is mostly automatic, manual inspection is used to check assignments at some critical stages, such as the detection of very distantly related homologues and anologues and the assignment of novel architectures. Described in this article is a recently established facility to search the database with the coordinates of a newly determined structure. The CATH server first locates domain boundaries and then uses automatic sequence and structure comparison methods to assign this new structure to one or more of the domain families within CATH. Diagnostic reports are generated, together with multiple structural alignments for close relatives. The Server can be accessed over the World Wide Web (WWW) and mirror sites are planned to improve access.

Amino Acid Sequence↗

CINEMA--a novel colour INteractive editor for multiple alignments.

CINEMA is a new editor for manipulating and generating multiple sequence alignments. The program provides both an interface to existing databases of alignments on the Internet and a tool for constructing and modifying alignments locally. It is written in Java, so executable code will run on most major desktop platforms without modification. The implementation is highly flexible, so the applet can be easily customised with additional functions; and the object classes are reusable, promoting rapid development of program extensions. Formerly, such extended functionality might have been provided via browser plug-ins, which have to be downloaded and installed on every client before loading data. Now, for the first time, an applet is available that allows interactive client-side processing of an alignment, which can then be stored or processed automatically on the server. The program is embedded in a comprehensive help file and is accessible both as a stand-alone tool on UCL's Bioinformatics Server; http:/(/)www.biochem.ucl.ac.uk/bsm/dbbrowser+ ++/CINEMA2.02/, and as an integral part of the PRINTS protein fingerprint database. Exploitation of such novel technologies revolutionises the way users may interact with public databases in the future: bioinformatics centres need not simply provide data, but are now able to offer the means by which information is visualised and manipulated, without the requirement for users to install software.

Color Perception↗

CATH--a hierarchic classification of protein domain structures.

BACKGROUND: Protein evolution gives rise to families of structurally related proteins, within which sequence identities can be extremely low. As a result, structure-based classifications can be effective at identifying unanticipated relationships in known structures and in optimal cases function can also be assigned. The ever increasing number of known protein structures is too large to classify all proteins manually, therefore, automatic methods are needed for fast evaluation of protein structures. RESULTS: We present a semi-automatic procedure for deriving a novel hierarchical classification of protein domain structures (CATH). The four main levels of our classification are protein class (C), architecture (A), topology (T) and homologous superfamily (H). Class is the simplest level, and it essentially describes the secondary structure composition of each domain. In contrast, architecture summarises the shape revealed by the orientations of the secondary structure units, such as barrels and sandwiches. At the topology level, sequential connectivity is considered, such that members of the same architecture might have quite different topologies. When structures belonging to the same T-level have suitably high similarities combined with similar functions, the proteins are assumed to be evolutionarily related and put into the same homologous superfamily. CONCLUSIONS: Analysis of the structural families generated by CATH reveals the prominent features of protein structure space. We find that nearly a third of the homologous superfamilies (H-levels) belong to ten major T-levels, which we call superfolds, and furthermore that nearly two-thirds of these H-levels cluster into nine simple architectures. A database of well-characterised protein structure families, such as CATH, will facilitate the assignment of structure-function/evolution relationships to both known and newly determined protein structures.

Databases, Factual↗

Novel developments with the PRINTS protein fingerprint database.

The PRINTS database of protein family 'fingerprints' is a diagnostic resource that complements the PROSITE dictionary of sites and patterns. Unlike regular expressions, fingerprints exploit groups of conserved motifs within sequence alignments to build characteristic signatures of family membership. Thus fingerprints inherently offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 600 fingerprints have been constructed and stored in PRINTS, representing a 50% increase in the size of the database in the last year. The current version, 13.0, encodes approximately 3000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via UCL's Bioinformatics World Wide Web (WWW) server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser / . We describe here progress with the database, its Web interface, and a recent exciting development: the integration of a novel colour alignment editor (http://www.biochem.ucl.ac.uk/bsm/dbbrowser++ +/CINEMA ), which allows visualisation and interactive manipulation of PRINTS alignments over the Internet.

Amino Acid Sequence↗

Analysis of domain structural class using an automated class assignment protocol.

The extent to which the contemporary dataset of protein structures can be segregated into four structural "classes" as originally defined by Levitt & Chothia in 1976 is examined and a simple method presented for the assignment of protein domains into these classes. Assignments are based on known three-dimensional structures, and for successful assignment it was found that helix/sheet content, contacts between secondary structures and their sequential order had to be used. The procedure attempts to maximise the automatic separation into classes for a dataset of 197 manually classified, non-homologous domains. It was found that approximately 90% of the structures were classified automatically; the remainder were borderline and were left for manual inspection. The method was then applied to a test set of 43 protein domains with similar results. The data support the concept of distinct classes of protein structure, although a few intermediate structures are found, demonstrating that it is possible to define relatively simple parameters complying with commonly accepted nomenclature that automatically define 90% of protein domains with essentially 100% accuracy. However, re-examination of the data also suggested that the previously separate alpha/beta and alpha + beta classes show considerable overlap and are more naturally represented as a single alpha beta class. This large alpha beta class can then be most easily subdivided by consideration of whether the sheets are mainly parallel, antiparallel or mixed. The correlation between structural class and function is discussed, together with the conservation of class within a sequence superfamily. This represents the first step in an automated phenetic description of protein structure complementing the usual phylogenetic approach to protein structure classification.

Algorithms↗

Structural similarity between the pleckstrin homology domain and verotoxin: the problem of measuring and evaluating structural similarity.

An unexpected structural similarity is described between the pleckstrin homology (PH) domain and verotoxin. This similarity has escaped detection primarily due to the differences in topology that exist between the two proteins. By comparing this result with two previously reported similarities for the PH domain, one with the lipocalins and another with the FK506 binding protein, we discuss the problems of measuring and assessing structural similarities.

Amino Acid Sequence↗