Search PubMed⌕ Search

Biomedical subjects

Michael J E Sternberg

Publications and source records attributed to Michael J E Sternberg.

At least 19 recordsLinked to original sources

The proteome: structure, function and evolution.

This paper reports two studies to model the inter-relationships between protein sequence, structure and function. First, an automated pipeline to provide a structural annotation of proteomes in the major genomes is described. The results are stored in a database at Imperial College, London (3D-GENOMICS) that can be accessed at www.sbg.bio.ic.ac.uk. Analysis of the assignments to structural superfamilies provides evolutionary insights. 3D-GENOMICS is being integrated with related proteome annotation data at University College London and the European Bioinformatics Institute in a project known as e-protein (http://www.e-protein.org/). The second topic is motivated by the developments in structural genomics projects in which the structure of a protein is determined prior to knowledge of its function. We have developed a new approach PHUNCTIONER that uses the gene ontology (GO) classification to supervise the extraction of the sequence signal responsible for protein function from a structure-based sequence alignment. Using GO we can obtain profiles for a range of specificities described in the ontology. In the region of low sequence similarity (around 15%), our method is more accurate than assignment from the closest structural homologue. The method is also able to identify the specific residues associated with the function of the protein family.

Computational Biology↗

Prediction of viable circular permutants using a graph theoretic approach.

MOTIVATION: In recent years graph-theoretic descriptions have been applied to aid the analysis of a number of complex biological systems. However, such an approach has only just begun to be applied to examine protein structures and the network of interactions between residues with promising results. Here we examine whether a graph measure known as closeness is capable of predicting regions where a protein can be split to form a viable circular permutant. Circular permutants are a powerful experimental tool to probe folding mechanisms and more recently have been used to design split enzyme reporter proteins. RESULTS: We test our method on an extensive set of experiments carried out on dihydrofolate reductase in which circular permutants were constructed for every amino acid position in the sequence, together with partial data from studies on other proteins. Results show that closeness is capable of correctly identifying significantly more residues which are suitable for circular permutation than solvent accessibility. This has potential implications for the design of successful split enzymes having particular importance for the development of protein-protein interaction screening methods and offers new perspectives on protein folding. More generally, the method illustrates the success with which graph-theoretic measures encapsulate the variety of long and short range interactions between residues during the folding process.

Binding Sites↗

Capturing expert knowledge with argumentation: a case study in bioinformatics.

MOTIVATION: The output of a bioinformatic tool such as BLAST must usually be interpreted by an expert before reliable conclusions can be drawn. This may be based upon the expert's experience, additional data and statistical analysis. Often the process is laborious, goes unrecorded and may be biased. Argumentation is an established technique for reasoning about situations where absolute truth or precise probability is impossible to determine. RESULTS: We demonstrate the application of argumentation to 3D-PSSM, a protein structure prediction tool. The expert's interpretation of results is represented as an argumentation framework. Given a 3D-PSSM result, an automated procedure constructs arguments for and against the conclusion that the result is a good predictor of protein structure. In addition to capturing the unique expertise of the author of 3D-PSSM for distribution to users, an improvement in recall of 5-10 percentage points is achieved. This technique can be applied to a wide range of bioinformatic tools. AVAILABILITY: Example public server and benchmarking data are available at http://www.sbg.bio.ic.ac.uk/~brj03/argumentation/paper/. Source code available on request.

Amino Acid Sequence↗

Assessing protein co-evolution in the context of the tree of life assists in the prediction of the interactome.

The identification of the whole set of protein interactions taking place in an organism is one of the main tasks in genomics, proteomics and systems biology. One of the computational techniques used by many investigators for studying and predicting protein interactions is the comparison of evolutionary histories (phylogenetic trees), under the hypothesis that interacting proteins would be subject to a similar evolutionary pressure resulting in a similar topology of the corresponding trees. Here, we present a new approach to predict protein interactions from phylogenetic trees, which incorporates information on the overall evolutionary histories of the species (i.e. the canonical "tree of life") in order to correct by the expected background similarity due to the underlying speciation events. We test the new approach in the largest set of annotated interacting proteins for Escherichia coli. This assessment of co-evolution in the context of the tree of life leads to a highly significant improvement (P(N) by sign test approximately 10E-6) in predicting interaction partners with respect to the previous technique, which does not incorporate information on the overall speciation tree. For half of the proteins we found a real interactor among the 6.4% top scores, compared with the 16.5% by the previous method. We applied the new method to the whole E.coli proteome and propose functions for some hypothetical proteins based on their predicted interactors. The new approach allows us also to detect non-canonical evolutionary events, in particular horizontal gene transfers. We also show that taking into account these non-canonical evolutionary events when assessing the similarity between evolutionary trees improves the performance of the method predicting interactions.

Algorithms↗

Prediction of the conformation and geometry of loops in globular proteins: testing ArchDB, a structural classification of loops.

In protein structure prediction, a central problem is defining the structure of a loop connecting 2 secondary structures. This problem frequently occurs in homology modeling, fold recognition, and in several strategies in ab initio structure prediction. In our previous work, we developed a classification database of structural motifs, ArchDB. The database contains 12,665 clustered loops in 451 structural classes with information about phi-psi angles in the loops and 1492 structural subclasses with the relative locations of the bracing secondary structures. Here we evaluate the extent to which sequence information in the loop database can be used to predict loop structure. Two sequence profiles were used, a HMM profile and a PSSM derived from PSI-BLAST. A jack-knife test was made removing homologous loops using SCOP superfamily definition and predicting afterwards against recalculated profiles that only take into account the sequence information. Two scenarios were considered: (1) prediction of structural class with application in comparative modeling and (2) prediction of structural subclass with application in fold recognition and ab initio. For the first scenario, structural class prediction was made directly over loops with X-ray secondary structure assignment, and if we consider the top 20 classes out of 451 possible classes, the best accuracy of prediction is 78.5%. In the second scenario, structural subclass prediction was made over loops using PSI-PRED (Jones, J Mol Biol 1999;292:195-202) secondary structure prediction to define loop boundaries, and if we take into account the top 20 subclasses out of 1492, the best accuracy is 46.7%. Accuracy of loop prediction was also evaluated by means of RMSD calculations.

Models, Molecular↗

Isolation of a small molecule inhibitor of DNA base excision repair.

The base excision repair (BER) pathway is essential for the removal of DNA bases damaged by alkylation or oxidation. A key step in BER is the processing of an apurinic/apyrimidinic (AP) site intermediate by an AP endonuclease. The major AP endonuclease in human cells (APE1, also termed HAP1 and Ref-1) accounts for >95% of the total AP endonuclease activity, and is essential for the protection of cells against the toxic effects of several classes of DNA damaging agents. Moreover, APE1 overexpression has been linked to radio- and chemo-resistance in human tumors. Using a newly developed high-throughput screen, several chemical inhibitors of APE1 have been isolated. Amongst these, CRT0044876 was identified as a potent and selective APE1 inhibitor. CRT0044876 inhibits the AP endonuclease, 3'-phosphodiesterase and 3'-phosphatase activities of APE1 at low micromolar concentrations, and is a specific inhibitor of the exonuclease III family of enzymes to which APE1 belongs. At non-cytotoxic concentrations, CRT0044876 potentiates the cytotoxicity of several DNA base-targeting compounds. This enhancement of cytotoxicity is associated with an accumulation of unrepaired AP sites. In silico modeling studies suggest that CRT0044876 binds to the active site of APE1. These studies provide both a novel reagent for probing APE1 function in human cells, and a rational basis for the development of APE1-targeting drugs for antitumor therapy.

Antineoplastic Agents↗

Protein-protein docking using 3D-Dock in rounds 3, 4, and 5 of CAPRI.

In rounds 3-5 of CAPRI, the community-wide experiment on the comparative evaluation of protein-protein docking for structure prediction, we applied the 3D-Dock software package to predict the atomic structures of nine biophysical interactions. This approach starts with an initial grid-based shape complementarity search. The product of this is a large number of potential interacting conformations that are subsequently ranked by interface residue propensities and interaction energies. Refinement through detailed energetics and optimization of side-chain positions using a rotamer library is also performed. For rounds 3, 4, and 5 of the CAPRI evaluation, where possible, we clustered functional residues on the surfaces of the monomers as an indication of binding sites, using sequence based evolutionary conservations. In certain targets this provided a very useful tool for identifying the areas of interaction. During round 5, we also applied the techniques of side-chain trimming and geometrical clustering described in the literature. Of the nine target complexes in rounds 3-5, we predicted conformations that contained at least some correct contact residues for seven of these systems. For two of the targets, we submitted predictions that were considered as medium-quality. These were a nidogen-laminin complex for target 8 (T08) and a serine-threonine phosphatase bound to a targeting subunit (T14). For a further three target systems, we produced models that were rated as acceptable predictions.

Algorithms↗

The relationship between the flexibility of proteins and their conformational states on forming protein-protein complexes with an application to protein-protein docking.

We investigate the extent to which the conformational fluctuations of proteins in solution reflect the conformational changes that they undergo when they form binary protein-protein complexes. To do this, we study a set of 41 proteins that form such complexes and whose three-dimensional structures are known, both bound in the complex and unbound. We carry out molecular dynamics simulations of each protein, starting from the unbound structure, and analyze the resulting conformational fluctuations in trajectories of 5 ns in length, comparing with the structure in the complex. It is found that fluctuations take some parts of the molecules into regions of conformational space close to the bound state (or give information about it), but at no point in the simulation does each protein as whole sample the complete bound state. Subsequent use of conformations from a clustered MD ensemble in rigid-body docking is nevertheless partially successful when compared to docking the unbound conformations, as long as the unbound conformations are themselves included with the MD conformations and the whole globally rescored. For one key example where sub-domain motion is present, a ribonuclease inhibitor, principal components analysis of the MD was applied and was also able to produce conformations for docking that gave enhanced results compared to the unbound. The most significant finding is that core interface residues show a tendency to be less mobile (by size of fluctuation or entropy) than the rest of the surface even when the other binding partner is absent, and conversely the peripheral interface residues are more mobile. This surprising result, consistent across up to 40 of the 41 proteins, suggests different roles for these regions in protein recognition and binding, and suggests ways that docking algorithms could be improved by treating these regions differently in the docking process.

Cluster Analysis↗

Automated prediction of protein function and detection of functional sites from structure.

Current structural genomics projects are yielding structures for proteins whose functions are unknown. Accordingly, there is a pressing requirement for computational methods for function prediction. Here we present PHUNCTIONER, an automatic method for structure-based function prediction using automatically extracted functional sites (residues associated to functions). The method relates proteins with the same function through structural alignments and extracts 3D profiles of conserved residues. Functional features to train the method are extracted from the Gene Ontology (GO) database. The method extracts these features from the entire GO hierarchy and hence is applicable across the whole range of function specificity. 3D profiles associated with 121 GO annotations were extracted. We tested the power of the method both for the prediction of function and for the extraction of functional sites. The success of function prediction by our method was compared with the standard homology-based method. In the zone of low sequence similarity (approximately 15%), our method assigns the correct GO annotation in 90% of the protein structures considered, approximately 20% higher than inheritance of function from the closest homologue.

Amino Acid Sequence↗

Clustering of protein domains in the human genome.

We present a systematic study of the clustering of genes within the human genome based on homology inferred from both sequence and structural similarity. The 3D-Genomics automated proteome annotation pipeline () was utilised to infer homology for each protein domain in the genome, for the 26 superfamilies most highly represented in the Structural Classification Of Proteins (SCOP) database. This approach enabled us to identify homologues that could not be detected by sequence-based methods alone. For each superfamily, we investigated the distribution, both within and among chromosomes, of genes encoding at least one domain within the superfamily. The results indicate a diversity of clustering behaviours: some superfamilies showed no evidence of any clustering, and others displayed significant clustering either within or among chromosomes, or both. Removal of tandem repeats reduced the levels of clustering observed, but some superfamilies still displayed highly significant clustering. Thus, our study suggests that either the process of gene duplication, or the evolution of the resulting clusters, differs between structural superfamilies.

Cadherins↗

Analysis of phenetic trees based on metabolic capabilites across the three domains of life.

Here, we used data of complete genomes to study comparatively the metabolism of different species. We built phenetic trees based on the enzymatic functions present in different parts of metabolism. Seven broad metabolic classes, comprising a total of 69 metabolic pathways, were comparatively analyzed for 27 fully sequenced organisms of the domains Eukarya, Bacteria and Archaea. Phylogenetic profiles based on the presence/absence of enzymatic functions for each metabolic class were determined and distance matrices for all the organisms were then derived from the profiles. Unrooted phenetic trees based upon the matrices revealed the distribution of the organisms according to their metabolic capabilities, reflecting the ecological pressures and adaptations that those species underwent during their evolution. We found that organisms that are closely related in phylogenetic terms could be distantly related metabolically and that the opposite is also true. For example, obligate bacterial pathogens were usually grouped together in our metabolic trees, demonstrating that obligate pathogens share common metabolic features regardless of their diverse phylogenetic origins. The branching order of proteobacteria often did not match their classical phylogenetic classification and Gram-positive bacteria showed diverse metabolic affinities. Archaea were found to be metabolically as distant from free-living bacteria as from eukaryotes, and sometimes were placed close to the metabolically highly specialized group of obligate bacterial pathogens. Metabolic trees represent an integrative approach for the comparison of the evolution of the metabolism and its correlation with the evolution of the genome, helping to find new relationships in the tree of life.

Animals↗

ArchDB: automated protein loop classification as a tool for structural genomics.

The annotation of protein function has become a crucial problem with the advent of sequence and structural genomics initiatives. A large body of evidence suggests that protein structural information is frequently encoded in local sequences, and that folds are mainly made up of a number of simple local units of super-secondary structural motifs, consisting of a few secondary structures and their connecting loops. Moreover, protein loops play an important role in protein function. Here we present ArchDB, a classification database of structural motifs, consisting of one loop plus its bracing secondary structures. ArchDB currently contains 12,665 super-secondary elements classified into 1496 motif subclasses. The database provides an easy way to retrieve functional information from protein structures sharing a common motif, to search motifs found in a given SCOP family, superfamily or fold, or to search by keywords on proteins with classified loops. The ArchDB database of loops is located at http://sbi.imim.es/archdb.

Amino Acid Motifs↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

The automatic discovery of structural principles describing protein fold space.

The study of protein structure has been driven largely by the careful inspection of experimental data by human experts. However, the rapid determination of protein structures from structural-genomics projects will make it increasingly difficult to analyse (and determine the principles responsible for) the distribution of proteins in fold space by inspection alone. Here, we demonstrate a machine-learning strategy that automatically determines the structural principles describing 45 folds. The rules learnt were shown to be both statistically significant and meaningful to protein experts. With the increasing emphasis on high-throughput experimental initiatives, machine-learning and other automated methods of analysis will become increasingly important for many biological problems.

Algorithms↗

CAPRI: a Critical Assessment of PRedicted Interactions.

CAPRI is a communitywide experiment to assess the capacity of protein-docking methods to predict protein-protein interactions. Nineteen groups participated in rounds 1 and 2 of CAPRI and submitted blind structure predictions for seven protein-protein complexes based on the known structure of the component proteins. The predictions were compared to the unpublished X-ray structures of the complexes. We describe here the motivations for launching CAPRI, the rules that we applied to select targets and run the experiment, and some conclusions that can already be drawn. The results stress the need for new scoring functions and for methods handling the conformation changes that were observed in some of the target systems. CAPRI has already been a powerful drive for the community of computational biologists who development docking algorithms. We hope that this issue of Proteins will also be of interest to the community of structural biologists, which we call upon to provide new targets for future rounds of CAPRI, and to all molecular biologists who view protein-protein recognition as an essential process.

Algorithms↗

Evaluation of the 3D-Dock protein docking suite in rounds 1 and 2 of the CAPRI blind trial.

The 3D-Dock suite of programs has been used to make predictions for the seven targets in rounds 1 and 2 of the CAPRI method evaluation exercise. Some correct contacts were obtained in at least one prediction for four of seven targets. Target 06 was predicted very well, with an RMSD of the ligand after superimposition of the receptor of only 0.77 A. We investigate the performance of the various stages of the method, with the aim of finding where improvements need to be made, and in particular whether the manual interventions that were made were essential, and whether results of the level of accuracy obtained for target 06 may be expected with confidence.

Algorithms↗

Evolution of enzymes in metabolism: a network perspective.

Several models have been proposed to explain the origin and evolution of enzymes in metabolic pathways. Initially, the retro-evolution model proposed that, as enzymes at the end of pathways depleted their substrates in the primordial soup, there was a pressure for earlier enzymes in pathways to be created, using the later ones as initial template, in order to replenish the pools of depleted metabolites. Later, the recruitment model proposed that initial templates from other pathways could be used as long as those enzymes were similar in chemistry or substrate specificity. These two models have dominated recent studies of enzyme evolution. These studies are constrained by either the small scale of the study or the artificial restrictions imposed by pathway definitions. Here, a network approach is used to study enzyme evolution in fully sequenced genomes, thus removing both constraints. We find that homologous pairs of enzymes are roughly twice as likely to have evolved from enzymes that are less than three steps away from each other in the reaction network than pairs of non-homologous enzymes. These results, together with the conservation of the type of chemical reaction catalyzed by evolutionarily related enzymes, suggest that functional blocks of similar chemistry have evolved within metabolic networks. One possible explanation for these observations is that this local evolution phenomenon is likely to cause less global physiological disruptions in metabolism than evolution of enzymes from other enzymes that are distant from them in the metabolic network.

Databases, Protein↗