Search PubMedSearch

Biomedical subjects

P Bork

Publications and source records attributed to P Bork.

At least 19 recordsLinked to original sources

SMART: identification and annotation of domains from signalling and extracellular protein sequences.

SMART is a simple modular architecture research tool and database that provides domain identification and annotation on the WWW (http://coot.embl-heidelberg.de/SMART). The tool compares query sequences with its databases of domain sequences and multiple alignments whilst concurrently identifying compositionally biased regions such as signal peptide, transmembrane and coiled coil segments. Annotated and unannotated regions of the sequence can be used as queries in searches of sequence databases. The SMART alignment collection represents more than 250 signalling and extracellular domains. Each alignment is curated to assign appropriate domain boundaries and to ensure its quality. In addition, each domain is annotated extensively with respect to cellular localisation, species distribution, functional class, tertiary structure and functionally important residues.

Amino Acid Sequence

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins

Evolution of new protein function: recombinational enhancer Fis originated by horizontal gene transfer from the transcriptional regulator NtrC.

New protein function is thought to evolve mostly by gene duplication and divergence. Here we present phylogenetic evidence that the multifunctional protein Fis of the gamma proteobacterial species derived from the COOH-terminal domain of an ancestral alpha proteobacterial NtrC transcriptional regulatory protein. All of the known enterobacterial fis genes are preceded by an open reading frame, named yhdG, that is highly similar to nifR3, a gene that forms an operon with ntrC in several alpha proteobacterial species. Thus, we propose that yhdG and fis were acquired by a lineage ancestral to the gamma proteobacteria in a single horizontal gene transfer event, and later diverged to their present functions.

Amino Acid Sequence

Homology-based fold predictions for Mycoplasma genitalium proteins.

Homology search techniques based on the iterative PSI-BLAST method in combination with various filters for low sequence complexity are applied to assign folds to all Mycoplasma genitalium proteins. The resulting procedure (implemented as a web server) is able to predict at least one domain in 37% of these proteins automatically, with an estimated accuracy higher than 98%. Taking structural features such as coiled coil or transmembrane regions aside, folds can be assigned to more than half of the globular proteins in a bacterium just by iterative sequence comparison.

Bacterial Proteins

Conformational stability studies of the pleckstrin DEP domain: definition of the domain boundaries.

Pleckstrin is the major substrate of protein kinase C in platelets. It contains at its N- and C-termini two pleckstrin homology (PH) domains which have been proposed to mediate protein-protein and protein-lipid interactions. A new module, called DEP, has recently been identified by sequence analysis in the central region of pleckstrin. In order to study this module, several recombinant polypeptides corresponding to the DEP module and N- and C-termini extended forms have been expressed. Using circular dichroism (CD) and nuclear magnetic resonance (NMR) techniques, the domain boundaries have been determined that yield a soluble and folded pleckstrin DEP domain. This comprises 93 amino acids with an alpha/beta fold in agreement with secondary structure predictions. Stability studies indicate that the regions surrounding the DEP domain do not contribute to its stability suggesting that the phosphorylation sites at S113, T114 and S117 are in an unstructured region. Identification of the regions of pleckstrin that are folded shall facilitate determination of its structure and function.

Amino Acid Sequence

Measuring genome evolution.

The determination of complete genome sequences provides us with an opportunity to describe and analyze evolution at the comprehensive level of genomes. Here we compare nine genomes with respect to their protein coding genes at two levels: (i) we compare genomes as "bags of genes" and measure the fraction of orthologs shared between genomes and (ii) we quantify correlations between genes with respect to their relative positions in genomes. Distances between the genomes are related to their divergence times, measured as the number of amino acid substitutions per site in a set of 34 orthologous genes that are shared among all the genomes compared. We establish a hierarchy of rates at which genomes have changed during evolution. Protein sequence identity is the most conserved, followed by the complement of genes within the genome. Next is the degree of conservation of the order of genes, whereas gene regulation appears to evolve at the highest rate. Finally, we show that some genomes are more highly organized than others: they show a higher degree of the clustering of genes that have orthologs in other genomes.

Animals

SMART, a simple modular architecture research tool: identification of signaling domains.

Accurate multiple alignments of 86 domains that occur in signaling proteins have been constructed and used to provide a Web-based tool (SMART: simple modular architecture research tool) that allows rapid identification and annotation of signaling domain sequences. The majority of signaling proteins are multidomain in character with a considerable variety of domain combinations known. Comparison with established databases showed that 25% of our domain set could not be deduced from SwissProt and 41% could not be annotated by Pfam. SMART is able to determine the modular architectures of single sequences or genomes; application to the entire yeast genome revealed that at least 6.7% of its genes contain one or more signaling domains, approximately 350 greater than previously annotated. The process of constructing SMART predicted (i) novel domain homologues in unexpected locations such as band 4.1-homologous domains in focal adhesion kinases; (ii) previously unknown domain families, including a citron-homology domain; (iii) putative functions of domain families after identification of additional family members, for example, a ubiquitin-binding role for ubiquitin-associated domains (UBA); (iv) cellular roles for proteins, such predicted DEATH domains in netrin receptors further implicating these molecules in axonal guidance; (v) signaling domains in known disease genes such as SPRY domains in both marenostrin/pyrin and Midline 1; (vi) domains in unexpected phylogenetic contexts such as diacylglycerol kinase homologues in yeast and bacteria; and (vii) likely protein misclassifications exemplified by a predicted pleckstrin homology domain in a Candida albicans protein, previously described as an integrin.

Amino Acid Sequence

Differential genome analysis applied to the species-specific features of Helicobacter pylori.

We introduce a simple and rapid strategy to identify genes that are responsible for species-specific phenotypes. The genome of a species that has a specific phenotype is compared with at least one, closely related, species that lacks this phenotype. Homologous genes that are shared among the species compared are identified and discarded from the list of candidates for species-specific genes. The process is automated and rapidly yields a small subset of the genome that likely contains genes responsible for the species-specific features. Functions are assigned to the genes, and dubious annotations are filtered out. Information is extracted not only from the presence of genes, but also from their absence with respect to known phenotypes. We have applied the technique to identify a set of species-specific genes in Helicobacter pylori by comparing it with its closest relatives for which complete genome sequences are available, Haemophilus influenzae and Escherichia coli. Of the genes of this set for which functional features can be obtained, a large fraction (63%, 123 proteins) is (potentially) involved in H. pylori's interaction with its host. We hypothesize that a family of outer membrane proteins is critical for the ability of H. pylori to colonize host cells in highly acidic environments.

Amino Acid Sequence

Merging extracellular domains: fold prediction for laminin G-like and amino-terminal thrombospondin-like modules based on homology to pentraxins.

Using a new method for construction and database searches of sequence consensus strings, we have identified a new superfamily of protein modules comprising laminin G, thrombospondin N and the pentraxin families. The conserved patterns correspond mainly to hydrophobic core residues located in central beta strands of the known three-dimensional structures of two pentraxins, the human C-reactive protein and the serum amyloid P-component. Thus, we predict a similar jellyroll fold for all members of this superfamily. In addition, the conservation of two exposed aspartate residues in the majority of superfamily members suggests hitherto unrecognised functional sites.

Amino Acid Sequence

Conservation of gene order: a fingerprint of proteins that physically interact.

A systematic comparison of nine bacterial and archaeal genomes reveals a low level of gene-order (and operon architecture) conservation. Nevertheless, a number of gene pairs are conserved. The proteins encoded by conserved gene pairs appear to interact physically. This observation can therefore be used to predict functions of, and interactions between, prokaryotic gene products.

Archaeal Proteins

Predicting functions from protein sequences--where are the bottlenecks?

The exponential growth of sequence data does not necessarily lead to an increase in knowledge about the functions of genes and their products. Prediction of function using comparative sequence analysis is extremely powerful but, if not performed appropriately, may also lead to the creation and propagation of assignment errors. While current homology detection methods can cope with the data flow, the identification, verification and annotation of functional features need to be drastically improved.

Amino Acid Sequence

Systematic genomic screening and analysis of mRNA in untranslated regions and mRNA precursors: combining experimental and computational approaches.

MOTIVATION: The untranslated regions (UTRs) of mRNA upstream (5'UTR) and downstream (3'UTR) of the open reading frame, as well as the mRNA precursor, carry important regulatory sequences. To reveal unidentified regulatory signals, we combine information from experiments with computational approaches. Depending on available knowledge, three different strategies are employed. RESULTS: Searching with a consensus template, new RNAs with regulatory RNA elements can be identified in genomic screens. By this approach, we identify new candidate regulatory motifs resembling iron-responsive elements in the 5'UTRs of HemA, FepB and FrdB mRNA from Escherichia coli. If an RNA element is not yet defined, it may be analyzed by combining results from SELEX (selective enrichment of ligands by exponential amplification) and a search of databases from RNA or genomic sequences. A cleavage stimulating factor (CstF) binding element 3 of the polyadenylation site in the mRNA precursor serves as a test example. Alternatively, the regulatory RNA element may be found by studying different RNA foldings and their correlation with simple experimental tests. We delineate a novel instability element in the 3'UTR of the estrogen receptor mRNA in this way. AVAILABILITY: Strategy, methods and programs are available on request from T.Dandekar. CONTACT: dandekar@embl-heidelberg.de

3' Untranslated Regions

Towards detection of orthologues in sequence databases.

MOTIVATION: Numerous homologous sequences from diverse species can be retrieved from databases using programs such as BLAST. However, due to multigene families, evolutionary relationship often cannot be easily determined and proper functional assignment becomes difficult. Thus, discrimination between orthologues and paralogues within BLAST output lists of homologous sequences becomes more and more important. RESULT: We therefore developed a method that attempts to construct a reconciled tree from a gene tree of selected sequences and its corresponding phylogenetic tree of the species involved (species tree). An interface on the Web is developed to enable users to analyse the BLAST result. BLAST outputs are parsed and, for the selected sequences, multiple alignments are constructed either globally or for local regions. Bootstrapped trees are returned and compared with the expected species tree. In cases of discrepancies, gene duplications are assumed and a reconciled tree is computed. The reconciled tree shows probable orthologues and paralogues as predicted.

Computational Biology

Frame: detection of genomic sequencing errors.

MOTIVATION: The underlying error rate for genomic sequencing sometimes results in the introduction of artificial frameshifts and in-frame stop codons into putative protein encoding genes. Severe errors are then introduced into the inferred transcripts through mis-translation or premature termination. RESULTS: We describe a system for screening segments of DNA for frameshift and in-frame stop errors in coding regions. The method is based on homology matching using blastx to compare all six reading frames of the query nucleotide sequence against selected protein sequence databases. Fragments of protein matching neighbouring regions of the query DNA are united and extended laterally to define candidate open reading frames, within which, frameshifts and stops are identified. Suitable targets include prokaryotic or other intron-free genomic sequence and complementary DNAs. As an example of its use, we report here two frameshifted ORFs that deviate from the original TIGR sequence annotations for the recently released Helicobacter pylori genome. AVAILABILITY: The tool is accessible via the URL http://www.sander.ebi.ac.uk/frame/. CONTACT: brown@ebi.ac.uk.

Amino Acid Sequence

Gene families: the taxonomy of protein paralogs and chimeras.

Ancient duplications and rearrangements of protein-coding segments have resulted in complex gene family relationships. Duplications can be tandem or dispersed and can involve entire coding regions or modules that correspond to folded protein domains. As a result, gene products may acquire new specificities, altered recognition properties, or modified functions. Extreme proliferation of some families within an organism, perhaps at the expense of other families, may correspond to functional innovations during evolution. The underlying processes are still at work, and the large fraction of human and other genomes consisting of transposable elements may be a manifestation of the evolutionary benefits of genomic flexibility.

Amino Acid Sequence