Search PubMed⌕ Search

Biomedical subjects

H Hermjakob

Publications and source records attributed to H Hermjakob.

12 recordsLinked to original sources

IntAct--open source resource for molecular interaction data.

IntAct is an open source database and software suite for modeling, storing and analyzing molecular interaction data. The data available in the database originates entirely from published literature and is manually annotated by expert biologists to a high level of detail, including experimental methods, conditions and interacting domains. The database features over 126,000 binary interactions extracted from over 2100 scientific publications and makes extensive use of controlled vocabularies. The web site provides tools allowing users to search, visualize and download data from the repository. IntAct supports and encourages local installations as well as direct data submission and curation collaborations. IntAct source code and data are freely available from http://www.ebi.ac.uk/intact.

DNA↗

Disrupted in Schizophrenia 1 Interactome: evidence for the close connectivity of risk genes and a potential synaptic basis for schizophrenia.

Disrupted in Schizophrenia 1 (DISC1) is a schizophrenia risk gene associated with cognitive deficits in both schizophrenics and the normal ageing population. In this study, we have generated a network of protein-protein interactions (PPIs) around DISC1. This has been achieved by utilising iterative yeast-two hybrid (Y2H) screens, combined with detailed pathway and functional analysis. This so-called 'DISC1 interactome' contains many novel PPIs and provides a molecular framework to explore the function of DISC1. The network implicates DISC1 in processes of cytoskeletal stability and organisation, intracellular transport and cell-cycle/division. In particular, DISC1 looks to have a PPI profile consistent with that of an essential synaptic protein, which fits well with the underlying molecular pathology observed at the synaptic level and the cognitive deficits seen behaviourally in schizophrenics. Utilising a similar approach with dysbindin (DTNBP1), a second schizophrenia risk gene, we show that dysbindin and DISC1 share common PPIs suggesting they may affect common biological processes and that the function of schizophrenia risk genes may converge.

Biological Transport↗

Integr8: enhanced inter-operability of European molecular biology databases.

OBJECTIVES: The increasing production of molecular biology data in the post-genomic era, and the proliferation of databases that store it, require the development of an integrative layer in database services to facilitate the synthesis of related information. The solution of this problem is made more difficult by the absence of universal identifiers for biological entities, and the breadth and variety of available data. METHODS: Integr8 was modelled using UML (Universal Modelling Language). Integr8 is being implemented as an n-tier system using a modern object-oriented programming language (Java). An object-relational mapping tool, OJB, is being used to specify the interface between the upper layers and an underlying relational database. RESULTS: The European Bioinformatics Institute is launching the Integr8 project. Integr8 will be an automatically populated database in which we will maintain stable identifiers for biological entities, describe their relationships with each other (in accordance with the central dogma of biology), and store equivalences between identified entities in the source databases. Only core data will be stored in Integr8, with web links to the source databases providing further information. CONCLUSIONS: Integr8 will provide the integrative layer of the next generation of bioinformatics services from the EBI. Web-based interfaces will be developed to offer gene-centric views of the integrated data, presenting (where known) the links between genome, proteome and phenotype.

Computational Biology↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

The role SWISS-PROT and TrEMBL play in the genome research environment.

SWISS-PROT, a curated protein sequence data bank, contains not only sequence data but also annotation relevant to a particular sequence. The annotation added to each entry is done by a team of biologists and comes, primarily, from articles in journals reporting the actual sequencing and sometimes characterisation. Review articles and collaboration with external experts also play a role along with the use of secondary databases like PROSITE and Pfam in addition to a variety of feature prediction methods. Annotation added by these methods is checked for relevance and likelihood to a particular sequence. The onset of genome sequencing has led to a dramatic increase in sequence data to be included in SWISS-PROT. This has led to the production of TrEMBL (Translation of the EMBL database). TrEMBL consists of entries in a SWISS-PROT format that are derived from the translation of all coding sequences in the EMBL nucleotide sequence database, that are not in SWISS-PROT. Unlike SWISS-PROT entries those in TrEMBL are awaiting manual annotation. However, rather than just representing basic sequence and source information, steps have been taken to add features and annotation automatically. In taking these steps it is hoped that TrEMBL entries are enhanced with some indication as to what a protein is, could or may be.

Amino Acid Sequence↗

VARSPLIC: alternatively-spliced protein sequences derived from SWISS-PROT and TrEMBL.

UNLABELLED: The program varsplic.pl uses information present in the SWISS-PROT and TrEMBL databases to create new records for alternatively spliced isoforms. These new records can be used in similarity searches. AVAILABILITY: The program is available at ftp://ftp.ebi.ac.uk/pub/software/swissprot/, together with regularly updated output files. CONTACT: pkersey@ebi.ac.uk

Alternative Splicing↗

InterPro--an integrated documentation resource for protein families, domains and functional sites.

MOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL).

Computational Biology↗

On the frequency of protein glycosylation, as deduced from analysis of the SWISS-PROT database.

The SWISS-PROT protein sequence data bank contains at present nearly 75,000 entries, almost two thirds of which include the potential N-glycosylation consensus sequence, or sequon, NXS/T (where X can be any amino acid but proline) and thus may be glycoproteins. The number of proteins filed as glycoproteins is however considerably smaller, 7942, of which 749 have been characterized with respect to the total number of their carbohydrate units and sites of attachment of the latter to the protein, as well as the nature of the carbohydrate-peptide linking group. Of these well characterized glycoproteins, about 90% carry either N-linked carbohydrate units alone or both N- and O-linked ones, attached at 1297 N-glycosylation sites (1.9 per glycoprotein molecule) and the rest are O-glycosylated only. Since the total number of sequons in the well characterized glycoproteins is 1968, their rate of occupancy is 2/3. Assuming that the same number of N-linked units and rate of sequon occupancy occur in all sequon containing proteins and that the proportion of solely O-glycosylated proteins (ca. 10%) will also be the same as among the well characterized ones, we conclude that the majority of sequon containing proteins will be found to be glycosylated and that more than half of all proteins are glycoproteins.

Animals↗

Swissknife - 'lazy parsing' of SWISS-PROT entries.

UNLABELLED: We present Swissknife, a set of Perl modules which provides a fast and reliable object-oriented interface to parsing and modifying files in SWISS-PROT format. AVAILABILITY: The Swissknife modules are available at ftp://ftp.ebi.ac. uk/pub/software/swissprot/. CONTACT: hhe@ebi.ac.uk

Databases, Factual↗

Databases on transcriptional regulation: TRANSFAC, TRRD and COMPEL.

TRANSFAC, TRRD (Transcription Regulatory Region Database) and COMPEL are databases which store information about transcriptional regulation in eukaryotic cells. The three databases provide distinct views on the components involved in transcription: transcription factors and their binding sites and binding profiles (TRANSFAC), the regulatory hierarchy of whole genes (TRRD), and the structural and functional properties of composite elements (COMPEL). The quantitative and qualitative changes of all three databases and connected programs are described. The databases are accessible via WWW:http://transfac.gbf.de/TRANSFAC orhttp://www.bionet.nsc.ru/TRRD

Animals↗

RIFLE: rapid identification of microorganisms by fragment length evaluation.

Biological macromolecules represent a valuable source of information for the identification and phylogenetic classification of microorganisms. One of the most commonly used macromolecules for this task is the 16S rDNA. The WWW-based RIFLE system presented here supports large-scale identification tasks by comparing 16S rDNA restriction patterns to a database of restriction patterns derived from sequence databases. Computing efficiency and robustness against experimental errors are gained by employing a new distance measure for restriction patterns, the fragment length distance. Results from the application of the system to the identification of uncultured microorganisms associated with the seagrass halophila stipulacea show the reliability of the method.

Algorithms↗