Search PubMed⌕ Search

Biomedical subjects

T K Attwood

Publications and source records attributed to T K Attwood.

At least 19 recordsLinked to original sources

METIS: multiple extraction techniques for informative sentences.

SUMMARY: METIS is a web-based integrated annotation tool. From single query sequences, the PRECIS component allows users to generate structured protein family reports from sets of related Swiss-Prot entries. These reports may then be augmented with pertinent sentences extracted from online biomedical literature via support vector machine and rule-based sentence classification systems. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/metis/

Algorithms↗

CADRE: the Central Aspergillus Data REpository.

CADRE is a public resource for housing and analysing genomic data extracted from species of Aspergillus. It arose to enable maintenance of the complete annotated genomic sequence of Aspergillus fumigatus and to provide tools for searching, analysing and visualizing features of fungal genomes. By implementing CADRE using Ensembl, a framework is in place for storing and comparing several genomes: the resource will thus expand by including other Aspergillus genomes (such as Aspergillus nidulans) as they become available. CADRE is accessible at http://www.cadre. man.ac.uk.

Aspergillus↗

PRECIS: protein reports engineered from concise information in SWISS-PROT.

MOTIVATION: There have been several endeavours to address the problem of annotating sequence data computationally, but the task is non-trivial and few tools have emerged that gather useful information on a given sequence, or set of sequences, in a simple and convenient manner. As more genome projects bear fruit, the mass of uncharacterized sequence data accumulating in public repositories grows ever larger. There is thus a pressing need for tools to support the process of automatic analysis and annotation of newly determined sequences. With this in mind, we have developed PRECIS, which automatically creates protein reports from sets of SWISS-PROT entries, collating results into structured reports, detailing known biological and medical information, literature and database cross-references, and relevant keywords.

Abstracting and Indexing↗

PRINTS and its automatic supplement, prePRINTS.

The PRINTS database houses a collection of protein fingerprints. These may be used to assign uncharacterised sequences to known families and hence to infer tentative functions. The September 2002 release (version 36.0) includes 1800 fingerprints, encoding approximately 11 000 motifs, covering a range of globular and membrane proteins, modular polypeptides and so on. In addition to its continued steady growth, we report here the development of an automatic supplement, prePRINTS, designed to increase the coverage of the resource and reduce some of the manual burdens inherent in its maintenance. The databases are accessible for interrogation and searching at http://www.bioinf.man.ac.uk/dbbrowser/PRINTS/.

Amino Acid Motifs↗

Phylogenomic analysis and evolution of the potassium channel gene family.

Potassium channels govern the permeability of cells to potassium ions, thereby controlling the membrane potential. In metazoa, potassium channels are encoded by a large, diverse gene family. Previous analyses of this gene family have focused on its diversity in mammals. Here we have pursued a more comprehensive study in Caenorhabditis elegans, Drosophila melanogaster, and mammalian genomes. The investigation revealed 164 potassium channel encoding genes in C. elegans, D. melanogaster, and mammals, classified into seven conserved families, which we applied to phylogenetic analysis. The trees are discussed in relation to the assignment of orthologous relationships between genes and vertebrate genome duplication.

Animals↗

PRINTS and PRINTS-S shed light on protein ancestry.

The PRINTS database houses a collection of protein fingerprints. These may be used to make family and tentative functional assignments for uncharacterised sequences. The September 2001 release (version 32.0) includes 1600 fingerprints, encoding approximately 10 000 motifs, covering a range of globular and membrane proteins, modular polypeptides and so on. In addition to its continued steady growth, we report here its use as a source of annotation in the InterPro resource, and the use of its relational cousin, PRINTS-S, to model relationships between families, including those beyond the reach of conventional sequence analysis approaches. The database is accessible for BLAST, fingerprint and text searches at http://www.bioinf.man.ac.uk/dbbrowser/PRINTS/.

Amino Acid Motifs↗

Progress in bioinformatics and the importance of being earnest.

In silico biology has gathered momentum as, worldwide, scientists have united in a common quest to sequence, store and analyse complete genomes. This year, a pivotal achievement of this cooperative endeavour was realised in the release of a public draft of the human genome, and with it the promises to improve our understanding of diverse aspects of biology and to yield a healthier future with safe personalized medicines. Key to these goals will be the need to elucidate and characterise the genes and gene products encoded not just in the human genome, but in many genomes. These tasks are underpinned by the concepts and processes of genome and gene/protein evolution, regulation of gene expression, mechanisms of protein folding, the manifestation of protein function, and so on, all of which must be understood in the context of complex, dynamic biological systems. Our use of computers to model such concepts and systems must be placed in the context of the current limits of our understanding of them:- it is important to recognise, for example, that we don't have a common understanding either of what constitutes a gene or a protein function; we can't invariably say that a particular sequence or fold has arisen via divergent or convergent evolution; and we don't fully understand the rules of protein folding. Accepting what we can't do in silico is essential in appreciating what we can do. Without this understanding, it is easy to be misled, as notions of what particular computational approaches can achieve are sometimes rather optimistic. There are valuable lessons to be learned here from the field of Artificial Intelligence, principal among which is the realisation that capturing and representing complex knowledge is time consuming, expensive and hard. Thus, we argue here that if bioinformatics is to tackle biological complexity in earnest, it would be wise to absorb the experience distilled from decades of artificial intelligence research, and to approach the road ahead with caution, rigour and pragmatism.

Artificial Intelligence↗

CINEMA-MX: a modular multiple alignment editor.

UNLABELLED: Analyzing and visualizing multiple sequence alignments is a common task in many areas of molecular biology and bioinformatics. Many tools exist for this purpose, but are not easily customizable for specific in-house uses. Here we report the development of an editor, CINEMA-MX, that addresses these issues. CINEMA-MX is highly modular and configurable, and we present examples to illustrate its extensibility. AVAILABILITY: The program and full source code, which are available from http://www.bioinf.man.ac.uk/dbbrowser/cinema-mx, are being released under a combination of the LGPL and GPL, for Unix or Windows platforms.

Computer Graphics↗

Deriving structural and functional insights from a ligand-based hierarchical classification of G protein-coupled receptors.

G protein-coupled receptors (GPCRs) constitute the largest known family of cell-surface receptors. With hundreds of members populating the rhodopsin-like GPCR superfamily and many more awaiting discovery in the human genome, they are of interest to the pharmaceutical industry because of the opportunities they afford for yielding potentially lucrative drug targets. Typical sequence analysis strategies for identifying novel GPCRs tend to involve similarity searches using standard primary database search tools. This will reveal the most similar sequence, generally without offering any insight into its family or superfamily relationships. Conversely, searches of most 'pattern' or family databases are likely to identify the superfamily, but not the closest matching subtype. Here we describe a diagnostic resource that allows identification of GPCRs in a hierarchical fashion, based principally upon their ligand preference. This resource forms part of the PRINTS database, which now houses approximately 250 GPCR-specific fingerprints (http://www.bioinf.man.ac.uk/dbbrowser/gpcrPRINTS/). This collection of fingerprints is able to provide more sensitive diagnostic opportunities than have been realized by related approaches and is currently the only diagnostic tool for assigning GPCR subtypes. Mapping such fingerprints on to three-dimensional GPCR models offers powerful insights into the structural and functional determinants of subtype specificity.

Algorithms↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

Which craft is best in bioinformatics?

'Silicon-based' biology has gathered momentum as the world-wide sequencing projects have made possible the investigation and comparative analysis of complete genomes. Central to the quest to elucidate and characterise the genes and gene products encoded within genomes are pivotal concepts concerning the processes of evolution, the mechanisms of protein folding, and, crucially, the manifestation of protein function. Our use of computers to model such concepts is limited by, and must be placed in the context of, the current limits of our understanding of these biological processes. It is important to recognise that we do not have a common understanding of what constitutes a gene; we cannot invariably say that a particular sequence or fold has arisen via divergence or convergence; we do not fully understand the rules of protein folding, so we cannot predict protein structure; and we cannot invariably diagnose protein function, given knowledge only of its sequence or structure in isolation. Accepting what we cannot do with computers plays an essential role in forming an appreciation of what we can do. Without this understanding, it is easy to be misled, as spurious arguments are often used to promote over-enthusiastic notions of what particular programs can achieve. There are valuable lessons to be learned here from the field of artificial intelligence, principal among which is the realisation that capturing and representing complex knowledge is time consuming, expensive and hard. If bioinformatics is to tackle biological complexity meaningfully, the road ahead must therefore be paved with caution, rigour and pragmatism.

Amino Acid Sequence↗

A compendium of specific motifs for diagnosing GPCR subtypes.

Analysis of G-protein-coupled receptor (GPCR) subtypes has attracted considerable interest because some drugs that act on GPCRs cause therapeutic problems as a result of their failure to differentiate between subtypes. In this article, an extensive compendium of diagnostic 'fingerprints' for GPCR subtypes and their families will be described. These fingerprints offer new opportunities to investigate correlations between specific sequence motifs and ligand binding or G-protein coupling, and are likely to prove valuable both in seeking novel receptors in genome data and in the characterization of orphan receptors.

Computational Biology↗

EASY--an Expert Analysis SYstem for interpreting database search outputs.

With the ever-increasing need to handle large volumes of sequence data efficiently and reliably, we have developed the EASY system for performing combined protein sequence and pattern database searches. EASY runs searches simultaneously and distils results into a concise 1-line diagnosis. By bringing together results of several different analyses, EASY provides a rapid means of evaluating biological significance, minimising the risk of inferring false relationships, for example from relying exclusively on top BLAST hits. The program has been tested using a variety of protein families and was instrumental in resolving family assignments in a major update of the PRINTS database.

Computational Biology↗

PRINTS-S: the database formerly known as PRINTS.

The PRINTS database houses a collection of protein family fingerprints. These are groups of motifs that together are diagnostically more potent than single motifs by virtue of the biological context afforded by matching motif neighbours. Around 1200 fingerprints have now been created and stored in the database. The September 1999 release (version 24.0) encodes approximately 7200 motifs, covering a range of globular and membrane proteins, modular polypeptides and so on. In addition to its continued steady growth, we report here several major changes to the resource, including the design of an automated strategy for database maintenance, and implementation of an object-relational schema for more efficient data management. The database is accessible for BLAST, fingerprint and text searches at http://www.bioinf.man.ac. uk/dbbrowser/PRINTS/

Database Management Systems↗

The quest to deduce protein function from sequence: the role of pattern databases.

In the wake of the numerous now-fruitful genome projects, we have witnessed a 'tsunami' of sequence data and with it the birth of the field of bioinformatics. Bioinformatics involves the application of information technology to the management and analysis of biological data. For many of us, this means that databases and their search tools have become an essential part of the research environment. However, the rate of sequence generation and the haphazard proliferation of databases have made it difficult to keep pace with developments, even for the cognoscenti. Moreover, increasing amounts of sequence information do not necessarily equate with an increase in knowledge, and in the panic to automate the route from raw data to biological insight, we may be generating and propagating innumerable errors in our precious databases. In the genome era upon us, researchers want rapid, easy-to-use, reliable tools for functional characterisation of newly determined sequences. For the pharmaceutical industry in particular, the Pandora's box of bioinformatics harbours an information-rich nugget, ripe with potential drug targets and possible new avenues for the development of therapeutic agents. This review outlines the current status of the major pattern databases now used routinely in the analysis of protein sequences. The review is divided into three main sections. In the first, commonly used terms are defined and the methods behind the databases are briefly described; in the second, the structure and content of the principal pattern databases are discussed; and in the final part, several alignment databases, which are frequently confused with pattern databases, are mentioned. For the new-comer, the array of resources, the range of methods behind them and the different tools required to search them can be confusing. The review therefore also briefly mentions a current international endeavour to integrate the diverse databases, which effort should facilitate sequence analysis in the future. This is particularly important for target-discovery programmes, where the challenge is to rationalise the enormous numbers of potential targets generated by sequence database searches. This problem may be addressed, at least in part, by reducing search outputs to the more focused and manageable subsets suggested by searches of integrated groups of family-specific pattern databases.

Amino Acid Motifs↗