Search PubMed⌕ Search

Biomedical subjects

R B Russell

Publications and source records attributed to R B Russell.

At least 37 records · Page 2Linked to original sources

Detection of protein three-dimensional side-chain patterns: new examples of convergent evolution.

Detection of recurring three-dimensional side-chain patterns is a potential means of inferring protein function. This paper presents a new method for detecting such patterns and discusses various implications. The method allows detection of side-chain patterns without any prior knowledge of function, requiring only protein structure data and associated multiple sequence alignments. A recursive, depth-first search algorithm finds all possible groups of identical amino acids common to two protein structures independent of sequence order. The search is highly constrained by distance constraints, and by ignoring amino acids unlikely to be involved in protein function. A weighted root-mean-square deviation (RMSD) between equivalenced groups of amino acids is used as a measure of similarity. The statistical significance of any RMSD is assigned by reference to a distribution fitted to simulated data. Searches with the Ser/His/Asp catalytic triad, a His/His porphyrin binding pattern, and the zinc-finger Cys/Cys/His/His pattern are performed to test the method on known examples. An all-against-all comparison of representatives from the structural classification of proteins (SCOP) is performed, revealing several new examples of evolutionary convergence to common patterns of side-chains within different tertiary folds and in different orders along the sequence. These include a di-zinc binding Asp/Asp/His/His/Ser pattern common to alkaline phosphatase/bacterial aminopeptidase, and an Asp/Glu/His/His/Asn/Asn pattern common to the active sites of DNase I and endocellulase E1. Implications for protein evolution, function prediction and the rational design of functional regulators are discussed.

Acetylglucosaminidase↗

Protein fold irregularities that hinder sequence analysis.

The detection of homologous protein sequences frequently provides useful predictions of function and structure. Methods for homology searching have continued to improve, such that very distant evolutionary relationships can now be detected. Little attention has been paid, however, to the problems of detecting homology when domains are inserted or permuted. Here we review recent occurrences of these phenomena and discuss methods that permit their detection.

Amino Acid Sequence↗

Recognition of analogous and homologous protein folds--assessment of prediction success and associated alignment accuracy using empirical substitution matrices.

Fold recognition methods aim to use the information in the known protein structures (the targets) to identify that the sequence of a protein of unknown structure (the probe) will adopt a known fold. This paper highlights that the structural similarities sought by these methods can be divided into two types: remote homologues and analogues. Homologues are the result of divergent evolution and often share a common function. We define remote homologues as those that are not easily detectable by sequence comparison methods alone. Analogues do not have a common ancestor and generally do not have a common function. Several sets of empirical matrices for residue substitution, secondary structure conservation and residue accessibility conservation have previously been derived from aligned pairs of remote homologues and analogues (Russell et al., J. Mol. Biol., 1997, 269, 423-439). Here a method for fold recognition, FOLDFIT, is introduced that uses these matrices to match the sequences, secondary structures and residue accessibilities of the probe and target. The approach is evaluated on distinct datasets of analogous and remotely homologous folds. The accuracy of FOLDFIT with the different matrices on the two datasets is contrasted to results from another fold recognition method (THREADER) and to searches using mutation matrices in the absence of any structural information. FOLDFIT identifies at top rank 12 out of 18 remotely homologous folds and five out of nine analogous folds. The average alignment accuracies for residue and secondary structure equivalencing are much higher for homologous folds (residue approximately 42%, secondary structure approximately 78%) than for analogues folds (approximately 12%, approximately 47%). Sequence searches alone can be successful for several homologues in the testing sets but nearly always fail for the analogues. These results suggest that the recognition of analogous and remotely homologous folds should be assessed separately. This study has implications for the development and comparative evaluation of fold recognition algorithms.

Evolution, Molecular↗

Misleading local sequence alignments: implications for comparative protein modelling.

Although it is well known that significant sequence similarity between proteins is reflected at the structural level, it is commonly assumed that any misaligned regions, as judged by the correct structure based alignment, are those where the local sequence identity is lower than the global. Recent studies have shown that this is not always the case and there can exist short stretches of high local identity which is not reflected in the structure based alignment. An analysis is presented of 290 pairs of homologous proteins with a view to quantifying the occurrence of these misleading local sequence alignments (MLSAs). It is found that such MLSAs are likely if the global sequence identity is less than 40% and can occur even when it is greater than 60%. The results have implications for automated homology modelling and also for the inference of function made by comparison.

Algorithms↗

Recognition of analogous and homologous protein folds: analysis of sequence and structure conservation.

An analysis was performed on 335 pairs of structurally aligned proteins derived from the structural classification of proteins (SCOP http://scop.mrc-lmb.cam.ac.uk/scop/) database. These similarities were divided into analogues, defined as proteins with similar three-dimensional structures (same SCOP fold classification) but generally with different functions and little evidence of a common ancestor (different SCOP superfamily classification). Homologues were defined as pairs of similar structures likely to be the result of evolutionary divergence (same superfamily) and were divided into remote, medium and close sub-divisions based on the percentage sequence identity. Particular attention was paid to the differences between analogues and remote homologues, since both types of similarities are generally undetectable by sequence comparison and their detection is the aim of fold recognition methods. Distributions of sequence identities and substitution matrices suggest a higher degree of sequence similarity in remote homologues than in analogues. Matrices for remote homologues show similarity to existing mutation matrices, providing some validity for their use in previously described fold recognition methods. In contrast, matrices derived from analogous proteins show little conservation of amino acid properties beyond broad conservation of hydrophobic or polar character. Secondary structure and accessibility were more conserved on average in remote homologues than in analogues, though there was no apparent difference in the root-mean-square deviation between these two types of similarities. Alignments of remote homologues and analogues show a similar number of gaps, openings (one or more sequential gaps) and inserted/deleted secondary structure elements, and both generally contain more gaps/openings/deleted secondary structure elements than medium and close homologues. These results suggest that gap parameters for fold recognition should be more lenient than those used in sequence comparison. Parameters were derived from the analogue and remote homologue datasets for potential used in fold recognition methods. Implications for protein fold recognition and evolution are discussed.

Computer Simulation↗

Two new examples of protein structural similarities within the structure-function twilight zone.

Two three-dimensional protein structural similarities accompanied by slight similarities in function are reported to highlight the present difficulties in discerning the relationship between structure and function. The similarity between the structures of uteroglobin and the CAP domain of haloalkane dehalogenase reveals a common four helix hydrophobic ligand binding motif. The similarity between the beta-barrel domains of glutaminyl tRNA synthetase and domain 1 of glutamine synthetase is accompanied by some similarity in the location of nucleotide binding sites and suggests a possible ancient domain common to these two proteins involved in the synthesis or binding of glutaminyl moieties. The problems raised by these two similarities in the structure-function 'twilight zone' are discussed.

Amino Acid Sequence↗

Protein fold recognition by mapping predicted secondary structures.

A strategy is presented for protein fold recognition from secondary structure assignments (alpha-helix and beta-strand). The method can detect similarities between protein folds in the absence of sequence similarity. Secondary structure mapping first identifies all possible matches (maps) between a query string of secondary structures and the secondary structures of protein domains of known three-dimensional structure. The maps are then passed through a series of structural filters to remove those that do not obey simple rules of protein structure. The surviving maps are ranked by scores from the alignment of predicted and experimental accessibilities. Searches made with secondary structure assignments for a test set of 11 fold-families put the correct sequence-dissimilar fold in the first rank 8/11 times. With cross-validated predictions of secondary structure this drops to 4/11 which compares favourably with the widely used THREADER program (1/11). The structural class is correctly predicted 10/11 times by the method in contrast to 5/11 for THREADER. The new technique obtains comparable accuracy in the alignment of amino acid residues and secondary structure elements. Searches are also performed with published secondary structure predictions for the von-Willebrand factor type A domain, the proteasome 20 S alpha subunit and the phosphotyrosine interaction domain. These searches demonstrate how the method can find the correct fold for a protein from a carefully constructed secondary structure prediction, multiple sequence alignment and distant restraints. Scans with experimentally determined secondary structures and accessibility, recognise the correct fold with high alignment accuracy (86% on secondary structures). This suggests that the accuracy of mapping will improve alongside any improvements in the prediction of secondary structure or accessibility. Application to NMR structure determination is also discussed.

Algorithms↗

Structure prediction. How good are we?

Recent successes show that, in certain circumstances, protein secondary structures can be predicted with high accuracy. How far are we from being able to predict the complete structure of a protein from its sequence?

Amino Acid Sequence↗

Towards an intelligent system for the automatic assignment of domains in globular proteins.

The automatic identification of protein domains from coordinates is the first step in the classification of protein folds and hence is required for databases to guide structure prediction. Most algorithms encode a single concept based and sometimes do not yield assignments that are consistent with the generally accepted perception. Our development of an automatic approach to identify reliably domains from protein coordinates is described. The algorithm is benchmarked against a manual identification of the domains in 284 representative protein chains. The first step is the domain assignment by distance (DAD) algorithm that considers the density of inter-residue contacts represented in a contact matrix. The algorithm yields 85% agreement with the manual assignment. The paper then considers how the reliability of these assignments could be evaluated. Finally the use of structural comparisons using the STAMP algorithm to validate domain assignment is reported on a test case.

Algorithms↗

Structural features can be unconserved in proteins with similar folds. An analysis of side-chain to side-chain contacts secondary structure and accessibility.

Side-chain to side-chain contacts, accessibility, secondary structure and RMS deviation were compared within 607 pairs of proteins having similar three-dimensional (3D) structures. Three types of protein 3D structural similarities were defined: type A having sequence and usually functional similarity; type B having functional, but no sequence similarity; and type C having only 3D structural similarity. Within proteins having little or no sequence similarity (types B and C), structural features frequently had a degree of conservation comparable to dissimilar 3D structures. Despite similar protein folds, as few as 30% of residues within similar protein 3D structures can form a common core. RMS deviations on core C alpha atoms can be as high as 3.2 A. Similar protein structures can have secondary structure identities as low as 41%, which is equivalent to that expected by chance. By defining three categories of amino acid accessibility (buried, half buried and exposed), some similar protein 3D structures have as few as 30% of positions in the same category, making them indistinguishable from pairs of dissimilar protein structures. Similar structures can also have as few as 12% of common side-chain to side-chain contacts, and virtually no similar energetically favourable side-chain to side-chain interactions. Complementary changes are defined as structurally equivalent pairs of interacting residues in two structures with energetically favourable but different side-chain interactions. For many proteins with similar three-dimensional structures, the proportion of complementary changes is near to that expected by chance, suggesting that many similar structures have fundamentally different stabilising interactions. All of the results suggest that proteins having similar 3D structures can have little in common apart from a scaffold of core secondary structures. This has profound implications for methods of protein fold detection, since many of the properties assumed to be conserved across similar protein 3D structures (e.g. accessibility, side-chain to side-chain contacts, etc.) are often unconserved within weakly similar (i.e. type B and C) protein 3D structures. Little difference was found between type B and C similarities suggesting that the structure of similar proteins can evolve beyond recognition even when function is conserved. Our findings suggest that it is more general features of protein structure, such as the requirements for burial of hydrophobic residues and exposure of polar residues, rather than specific residue-residue interactions that determine how well a particular sequence adopts a particular fold.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Domain insertion.

Explore the source record for details and available documents.

Models, Molecular↗

The limits of protein secondary structure prediction accuracy from multiple sequence alignment.

The expected best residue-by-residue accuracies for secondary structure prediction from multiple protein sequence alignment have been determined by an analysis of known protein structural families. The results show substantial variation is possible among homologous protein structures, and that 100% agreement is unlikely between a consensus prediction and one member of a protein structural family. The study provides the range of agreement to be expected between a perfect secondary structure prediction from a multiple alignment and each protein within the alignment. The results of this study overcome the difficulties inherent in the use of residue-by-residue accuracy for assessing the quality of consensus secondary structure predictions. The accuracies of recent consensus predictions for the annexins, SH2 domains and SH3 domains fall within the expected range for a perfect prediction.

Amino Acid Sequence↗

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence↗

Multiple protein sequence alignment from tertiary structure comparison: assignment of global and residue confidence levels.

An algorithm is presented for the accurate and rapid generation of multiple protein sequence alignments from tertiary structure comparisons. A preliminary multiple sequence alignment is performed using sequence information, which then determines an initial superposition of the structures. A structure comparison algorithm is applied to all pairs of proteins in the superimposed set and a similarity tree calculated. Multiple sequence alignments are then generated by following the tree from the branches to the root. At each branchpoint of the tree, a structure-based sequence alignment and coordinate transformations are output, with the multiple alignment of all structures output at the root. The algorithm encoded in STAMP (STructural Alignment of Multiple Proteins) is shown to give alignments in good agreement with published structural accounts within the dehydrogenase fold domains, globins, and serine proteinases. In order to reduce the need for visual verification, two similarity indices are introduced to determine the quality of each generated structural alignment. Sc quantifies the global structural similarity between pairs or groups of proteins, whereas Pij' provides a normalized measure of the confidence in the alignment of each residue. STAMP alignments have the quality of each alignment characterized by Sc and Pij' values and thus provide a reproducible resource for studies of residue conservation within structural motifs.

Algorithms↗