Search PubMedSearch

Biomedical subjects

M J Sternberg

Publications and source records attributed to M J Sternberg.

At least 19 recordsLinked to original sources

Use of pair potentials across protein interfaces in screening predicted docked complexes.

Empirical residue-residue pair potentials are used to screen possible complexes for protein-protein dockings. A correct docking is defined as a complex with not more than 2.5 A root-mean-square distance from the known experimental structure. The complexes were generated by "ftdock" (Gabb et al. J Mol Biol 1997;272:106-120) that ranks using shape complementarity. The complexes studied were 5 enzyme-inhibitors and 2 antibody-antigens, starting from the unbound crystallographic coordinates, with a further 2 antibody-antigens where the antibody was from the bound crystallographic complex. The pair potential functions tested were derived both from observed intramolecular pairings in a database of nonhomologous protein domains, and from observed intermolecular pairings across the interfaces in sets of nonhomologous heterodimers and homodimers. Out of various alternate strategies, we found the optimal method used a mole-fraction calculated random model from the intramolecular pairings. For all the systems, a correct docking was placed within the top 12% of the pair potential score ranked complexes. A combined strategy was developed that incorporated "multidock," a side-chain refinement algorithm (Jackson et al. J Mol Biol 1998;276:265-285). This placed a correct docking within the top 5 complexes for enzyme-inhibitor systems, and within the top 40 complexes for antibody-antigen systems.

Algorithms

Progress in protein structure prediction: assessment of CASP3.

The third comparative assessment of techniques of protein structure prediction (CASP3) was held during 1998. This is a blind trial in which structures are predicted prior to having knowledge of the coordinates, which are then revealed to enable the assessment. Three sections at the meeting evaluated different methodologies - comparative modelling, fold recognition and ab initio methods. For some, but not all of the target coordinates, high quality models were submitted in each of these sections. There have been improvements in prediction techniques since CASP2 in 1996, most notably for ab initio methods.

Algorithms

An analysis of conformational changes on protein-protein association: implications for predictive docking.

Conformational changes on complex formation have been measured for 39 pairs of structures of complexed proteins and unbound equivalents, averaged over interface and non-interface regions and for individual residues. We evaluate their significance by comparison with the differences seen in 12 pairs of independently solved structures of identical proteins, and find that just over half have some substantial overall movement. Movements involve main chains as well as side chains, and large changes in the interface are closely involved with complex formation, while those of exposed non-interface residues are caused by flexibility and disorder. Interface movements in enzymes are similar in extent to those of inhibitors. All eight of the complexes (six enzyme-inhibitor and two antibody-antigen) that have structures of both components in an unbound form available show some significant interface movement. However, predictive docking is successful even when some of the largest changes occur. We note however that the situation may be different in systems other than the enzyme-inhibitors which dominate this study. Thus the general model is induced fit but, because there is only limited conformational change in many systems, recognition can be treated as lock and key to a first approximation.

Algorithms

Crystal structure at 1.95 A resolution of the breast tumour-specific antibody SM3 complexed with its peptide epitope reveals novel hypervariable loop recognition.

The anti-breast tumour antibody SM3 has a high selectivity in reacting specifically with carcinoma-associated mucin. SM3 recognises the core repeating motif (Pro-Asp-Thr-Arg-Pro) of aberrantly glycosylated epithelial mucin MUC1, and has potential as a therapeutic and diagnostic tool. Here we report the crystal structure of the Fab fragment of SM3 in complex with a 13-residue MUC1 peptide antigen (Thr1P-Ser2P-Ala3P-Pro4P-Asp5P-Thr6P -Arg7P-Pro8P-Ala9P-Pro10P-Gly11P- Ser12P-Thr13P). The SM3-MUC1 peptide structure was solved by molecular replacement, and the current model is refined at 1.95 A resolution with an R-factor of 21.3% and R-free 28.3%. The MUC1 peptide is bound both by non-polar interactions and hydrogen bonds in an elongated groove in the antibody-combining site through interactions with Complimentarity Determining Regions (CDRs), three of the light chain (L1, L2, L3) and two of the heavy chain (H1 and H3). The conformation of the peptide is mainly extended with no discernable standard secondary structure. There is a single non-proline cis-peptide bond in H3 (Val95H-Gly96H-Gln97H-Phe98H-Ala101H-Ty r102H) between Gly96H and Gln97H, which appears to play a role in SM3-peptide antigen interactions, and represents the first such example within an antibody hypervariable loop. The SM3-MUC1 peptide structure has implications for rational therapeutic and diagnostic antibody engineering.

Amino Acid Sequence

Conformational analysis of the first observed non-proline cis-peptide bond occurring within the complementarity determining region (CDR) of an antibody.

An analysis has been performed on the first example of a non-proline cis- peptide bond found within a complementarity determining region (CDR) of an antibody. The bond is located in CDR 3 of the heavy chain (H3) and makes substantial interactions to a peptide from a breast tumour-associated antigen. The antibody-peptide complex is compared, both in H3 length (six residues) and peptide conformation, to a number of other such complexes in the Brookhaven Data Bank (PDB). There is only one other H3 loop of the same length. Analysis of loop searches of the PDB, taken over the H3 framework of SM3, suggest that there is a limited repertoire of conformations for loops of length 6 compared to loops of length 5 and 7. It is argued that the cis-peptide bond is present because of the limited number of loop conformations of length 6, plus, the requirement of the H3 loop to contact the bound peptide. Modelling suggests that an all-trans-peptide loop conformation can replace the H3 loop and this raises the question of whether there is a trans- to cis-peptide bond isomerization upon peptide binding.

Antibodies

Modelling repressor proteins docking to DNA.

The docking of repressor proteins to DNA starting from the unbound protein and model-built DNA coordinates is modeled computationally. The approach was evaluated on eight repressor/DNA complexes that employed different modes for protein/ DNA recognition. The global search is based on a protein-protein docking algorithm that evaluates shape and electrostatic complementarity, which was modified to consider the importance of electrostatic features in DNA-protein recognition. Complexes were then ranked by an empirical score for the observed amino acid /nucleotide pairings (i.e., protein-DNA pair potentials) derived from a database of 20 protein/ DNA complexes. A good prediction had at least 65% of the correct contacts modeled. This approach was able to identify a good solution at rank four or better for three out of the eight complexes. Predicted complexes were filtered by a distance constraint based on experimental data defining the DNA footprint. This improved coverage to four out of eight complexes having a good model at rank four or better. The additional use of amino acid mutagenesis and phylogenetic data defining residues on the repressor resulted in between 2 and 27 models that would have to be examined to find a good solution for seven of the eight test systems. This study shows that starting with unbound coordinates one can predict three-dimensional models for protein/DNA complexes that do not involve gross conformational changes on association.

Algorithms

Structure of an XRCC1 BRCT domain: a new protein-protein interaction module.

The BRCT domain (BRCA1 C-terminus), first identified in the breast cancer suppressor protein BRCA1, is an evolutionarily conserved protein-protein interaction region of approximately 95 amino acids found in a large number of proteins involved in DNA repair, recombination and cell cycle control. Here we describe the first three-dimensional structure and fold of a BRCT domain determined by X-ray crystallography at 3.2 A resolution. The structure has been obtained from the C-terminal region of the human DNA repair protein XRCC1, and comprises a four-stranded parallel beta-sheet surrounded by three alpha-helices, which form an autonomously folded domain. The compact XRCC1 structure explains the observed sequence homology between different BRCT motifs and provides a framework for modelling other BRCT domains. Furthermore, the established structure of an XRCC1 BRCT homodimer suggests potential protein-protein interaction sites for the complementary BRCT domain in DNA ligase III, since these two domains form a stable heterodimeric complex. Based on the XRCC1 BRCT structure, we have constructed a model for the C-terminal BRCT domain of BRCA1, which frequently is mutated in familial breast and ovarian cancer. The model allows insights into the effects of such mutations on the fold of the BRCT domain.

Amino Acid Sequence

HAD, a data bank of heavy-atom binding sites in protein crystals: a resource for use in multiple isomorphous replacement and anomalous scattering.

Information on the preparation and characterization of heavy-atom derivatives of protein crystals has been collected, either from the literature or directly from protein crystallographers, and assembled in the form of a heavy-atom data bank (HAD). The data bank contains coordinate data for the heavy-atom positions in a form that is compatible with the crystallographic data in the Brookhaven Protein Data Bank, together with a wealth of information on the crystallization conditions, the nature of the heavy-atom reagent and references to relevant publications. Some statistical information derived from the data bank, such as the most popular heavy-atom derivatives, is also included. The information can be directly accessed and should be useful to protein crystallographers seeking to improve their success in preparing heavy-atom derivatives for the methods of isomorphous replacement and anomalous dispersion. The World Wide Web address of HAD is http://www.icnet.uk/bmm/had.

Binding Sites

Supersites within superfolds. Binding site similarity in the absence of homology.

A method is presented to assess the significance of binding site similarities within superimposed protein three-dimensional (3D) structures and applied to all similar structures in the Protein Data Bank. For similarities between 3D structures lacking significant sequence similarity, the important distinction was made between remote homology (an ancient common ancestor) and analogy (likely convergence to a folding motif) according to the structural classification of proteins (SCOP) database. Supersites were defined as structural locations on groups of analogous proteins (i.e. superfolds) showing a statistically significant tendency to bind substrates despite little evidence of a common ancestor for the proteins considered. We identify three potentially new superfolds containing supersites: ferredoxin-like folds, four-helical bundles and double-stranded beta helices. In addition, the method quantifies binding site similarities within homologous proteins and previously identified supersites such as that found in the beta/alpha (TIM) barrels. For the nine superfolds, the accuracy of predictions of binding site locations is assessed. Implications for protein evolution, and the prediction of protein function either through fold recognition or tertiary structure comparison, are discussed.

Animals

Automated classification of antibody complementarity determining region 3 of the heavy chain (H3) loops into canonical forms and its application to protein structure prediction.

A computer-based algorithm was used to cluster the loops forming the complementarity determining region (CDR) 3 of the heavy chain (H3) into canonical classes. Previous analyses of the three-dimensional structures of CDR loops (also known as the hypervariable regions) within antibody immunoglobulin variable domains have shown that for five of the six CDRs there are only a few main-chain conformations (known as canonical forms) that show clear relationships between sequence and structure. However, the larger variation in length and conformation of loops within H3 has limited the classification of these loops into canonical forms. The clustering procedure presented here is based on aligning the Ramachandran-coded main-chain conformation of the residues using a dynamic algorithm that allows the insertion of gaps to obtain an optimum alignment. A total of 41 H3 loops out of 62 non-identical loops, extracted from the Brookhaven Protein Data Bank, have been automatically grouped into 22 clusters. Inspection of the clusters for consensus sequences or intra-loop interactions or invariant conformation led to the proposal of 13 canonical forms representing 31 loops. These canonical forms include a consideration of the geometry of both the take-off region adjacent to the bracing beta-strands and the remaining loop apex. Subsequently a new set of 15 H3 loops not included in the initial analysis was considered. The clustering procedure was repeated and nine of these 15 loops could be assigned to original clusters, including seven to canonical forms. A sequence profile was generated for each canonical form from the original set of loops and matched against the sequences of the new H3 loops. For five out of the seven new H3 loops that were in a canonical form, the correct form was identified at first rank by this predictive scheme.

Amino Acid Sequence

Rapid refinement of protein interfaces incorporating solvation: application to the docking problem.

A computationally tractable strategy has been developed to refine protein-protein interfaces that models the effects of side-chain conformational change, solvation and limited rigid-body movement of the subunits. The proteins are described at the atomic level by a multiple copy representation of side-chains modelled according to a rotamer library on a fixed peptide backbone. The surrounding solvent environment is described by "soft" sphere Langevin dipoles for water that interact with the protein via electrostatic, van der Waals and field-dependent hydrophobic terms. Energy refinement is based on a two-step process in which (1) a probability-based conformational matrix of the protein side-chains is refined iteratively by a mean field method. A side-chain interacts with the protein backbone and the probability-weighted average of the surrounding protein side-chains and solvent molecules. The resultant protein conformations then undergo (2) rigid-body energy minimization to relax the protein interface. Steps (1) and (2) are repeated until convergence of the interaction energy. The influence of refinement on side-chain conformation starting from unbound conformations found improvement in the RMSD of side-chains in the interface of protease-inhibitor complexes, and shows that the method leads to an improvement in interface geometry. In terms of discriminating between docked structures, the refinement was applied to two classes of protein-protein complex: five protease-protein inhibitor and four antibody-antigen complexes. A large number of putative docked complexes have already been generated for the test systems using our rigid-body docking program, FTDOCK. They include geometries that closely resemble the crystal complex, and therefore act as a test for the refinement procedure. In the protease-inhibitors, geometries that resemble the crystal complex are ranked in the top four solutions for four out of five systems when solvation is included in the energy function, against a background of between 26 and 364 complexes in the data set. The results for the antibody-antigen complexes are not as encouraging, with only two of the four systems showing discrimination. It would appear that these results reflect the somewhat different binding mechanism dominant in the two types of protein-protein complex. Binding in the protease-inhibitors appears to be "lock and key" in nature. The fixed backbone and mobile side-chain representation provide a good model for binding. Movements in the backbone geometry of antigens on binding represent an "induced-fit" and provides more of a challenge for the model. Given the limitations of the conformational sampling, the ability of the energy function to discriminate between native and non-native states is encouraging. Development of the approach to include greater conformational sampling could lead to a more general solution to the protein docking problem.

Animals

Predictive docking of protein-protein and protein-DNA complexes.

Recent developments in algorithms to predict the docking of two proteins have considered both the initial rigid-body global search and subsequent screening and refinement. The result of two blind trials of protein docking are encouraging--for complexes that are not too large and do not undergo sizeable conformational change upon association, the algorithms are now able to suggest reasonably accurate models.

Algorithms

Recognition of analogous and homologous protein folds--assessment of prediction success and associated alignment accuracy using empirical substitution matrices.

Fold recognition methods aim to use the information in the known protein structures (the targets) to identify that the sequence of a protein of unknown structure (the probe) will adopt a known fold. This paper highlights that the structural similarities sought by these methods can be divided into two types: remote homologues and analogues. Homologues are the result of divergent evolution and often share a common function. We define remote homologues as those that are not easily detectable by sequence comparison methods alone. Analogues do not have a common ancestor and generally do not have a common function. Several sets of empirical matrices for residue substitution, secondary structure conservation and residue accessibility conservation have previously been derived from aligned pairs of remote homologues and analogues (Russell et al., J. Mol. Biol., 1997, 269, 423-439). Here a method for fold recognition, FOLDFIT, is introduced that uses these matrices to match the sequences, secondary structures and residue accessibilities of the probe and target. The approach is evaluated on distinct datasets of analogous and remotely homologous folds. The accuracy of FOLDFIT with the different matrices on the two datasets is contrasted to results from another fold recognition method (THREADER) and to searches using mutation matrices in the absence of any structural information. FOLDFIT identifies at top rank 12 out of 18 remotely homologous folds and five out of nine analogous folds. The average alignment accuracies for residue and secondary structure equivalencing are much higher for homologous folds (residue approximately 42%, secondary structure approximately 78%) than for analogues folds (approximately 12%, approximately 47%). Sequence searches alone can be successful for several homologues in the testing sets but nearly always fail for the analogues. These results suggest that the recognition of analogous and remotely homologous folds should be assessed separately. This study has implications for the development and comparative evaluation of fold recognition algorithms.

Evolution, Molecular

Misleading local sequence alignments: implications for comparative protein modelling.

Although it is well known that significant sequence similarity between proteins is reflected at the structural level, it is commonly assumed that any misaligned regions, as judged by the correct structure based alignment, are those where the local sequence identity is lower than the global. Recent studies have shown that this is not always the case and there can exist short stretches of high local identity which is not reflected in the structure based alignment. An analysis is presented of 290 pairs of homologous proteins with a view to quantifying the occurrence of these misleading local sequence alignments (MLSAs). It is found that such MLSAs are likely if the global sequence identity is less than 40% and can occur even when it is greater than 60%. The results have implications for automated homology modelling and also for the inference of function made by comparison.

Algorithms

A computational system for modelling flexible protein-protein and protein-DNA docking.

A computational system is described that predicts the structure of protein/protein and protein/DNA complexes starting from unbound coordinate sets. The approach is (i) a global search with rigid-body docking for complexes with shape complementarity and favourable electrostatics; (ii) use of distance constraints from experimental (or predicted) knowledge of critical residues; (iii) use of pair potential to screen docked complexes and (iv) refinement and further screening by protein-side chain optimisation and interfacial energy minimisation. The system has been applied to model ten protein/protein and eight protein-repressor/DNA (steps i to iii only) complexes. In general a few complexes, one of which is close to the true structure, can be generated.

Algorithms

Modelling protein docking using shape complementarity, electrostatics and biochemical information.

A protein docking study was performed for two classes of biomolecular complexes: six enzyme/inhibitor and four antibody/antigen. Biomolecular complexes for which crystal structures of both the complexed and uncomplexed proteins are available were used for eight of the ten test systems. Our docking experiments consist of a global search of translational and rotational space followed by refinement of the best predictions. Potential complexes are scored on the basis of shape complementarity and favourable electrostatic interactions using Fourier correlation theory. Since proteins undergo conformational changes upon binding, the scoring function must be sufficiently soft to dock unbound structures successfully. Some degree of surface overlap is tolerated to account for side-chain flexibility. Similarly for electrostatics, the interaction of the dispersed point charges of one protein with the Coulombic field of the other is measured rather than precise atomic interactions. We tested our docking protocol using the native rather than the complexed forms of the proteins to address the more scientifically interesting problem of predictive docking. In all but one of our test cases, correctly docked geometries (interface Calpha RMS deviation </=2 A from the experimental structure) are found during a global search of translational and rotational space in a list that was always less than 250 complexes and often less than 30. Varying degrees of biochemical information are still necessary to remove most of the incorrectly docked complexes.

Algorithms

Recognition of analogous and homologous protein folds: analysis of sequence and structure conservation.

An analysis was performed on 335 pairs of structurally aligned proteins derived from the structural classification of proteins (SCOP http://scop.mrc-lmb.cam.ac.uk/scop/) database. These similarities were divided into analogues, defined as proteins with similar three-dimensional structures (same SCOP fold classification) but generally with different functions and little evidence of a common ancestor (different SCOP superfamily classification). Homologues were defined as pairs of similar structures likely to be the result of evolutionary divergence (same superfamily) and were divided into remote, medium and close sub-divisions based on the percentage sequence identity. Particular attention was paid to the differences between analogues and remote homologues, since both types of similarities are generally undetectable by sequence comparison and their detection is the aim of fold recognition methods. Distributions of sequence identities and substitution matrices suggest a higher degree of sequence similarity in remote homologues than in analogues. Matrices for remote homologues show similarity to existing mutation matrices, providing some validity for their use in previously described fold recognition methods. In contrast, matrices derived from analogous proteins show little conservation of amino acid properties beyond broad conservation of hydrophobic or polar character. Secondary structure and accessibility were more conserved on average in remote homologues than in analogues, though there was no apparent difference in the root-mean-square deviation between these two types of similarities. Alignments of remote homologues and analogues show a similar number of gaps, openings (one or more sequential gaps) and inserted/deleted secondary structure elements, and both generally contain more gaps/openings/deleted secondary structure elements than medium and close homologues. These results suggest that gap parameters for fold recognition should be more lenient than those used in sequence comparison. Parameters were derived from the analogue and remote homologue datasets for potential used in fold recognition methods. Implications for protein fold recognition and evolution are discussed.

Computer Simulation

Chemical synthesis, structural modeling, and biological activity of the epidermal growth factor-like domain of human cripto.

Cripto, also known as human teratocarcinoma-derived growth factor 1 (TDGF-1), contains a 40 amino acid region with some similarity to the epidermal growth factor (EGF) domain. However, sequence homology is largely restricted to the classical cysteine/glycine motif with only limited similarities in other regions. Significant differences to human EGF include the absence of all seven residues between the two N-terminal half-cystines and a five-residue shorter loop between the third and fourth half-cystines. We examine the hypothesis that, in spite of these differences, cripto can adopt the characteristic EGF-like 1-3, 2-4, 5-6 disulfide bond pattern. A comparative structural model of the growth factor cripto was constructed on the basis of its similarity to EGF, transforming growth factor alpha (TGF-alpha), and the EGF-like domain of human clotting factor IX. The predicted disulfide bridges and disulfide-bridged loops were analyzed and appear viable in the modeled structure. Moreover, to ascertain the importance of disulfide arrangement for cripto bioactivity, two 47-residue peptides were synthesized and then refolded using either a simple oxidative or a controlled sequential refolding protocol. The cripto peptides were tested for their ability to stimulate MAP-kinase activity, for inhibition of beta-casein induction, and for Shc phosphorylation in MDA-MB 453 human mammary carcinoma cells and HC-11 mouse mammary epithelial cells. Data suggest that cripto does adopt the 1-3, 2-4, 5-6 disulfide pattern and thus forms the classical EGF-like fold in spite of the significant deletions within the folding domain. The predicted structure of cripto shows some of the characteristics of both the ErbB1- and ErbB3/ErbB4-binding growth factors.

Amino Acid Sequence