Computer speed and sequence comparison.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to G J Barton.
Explore the source record for details and available documents.
Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.
Tryptases are trypsin-like serine proteinases found in the granules of mast cells. Although they show 40% sequence identity with trypsin and contain only 20 or 21 additional residues, tryptases display several unusual features. Unlike trypsin, the tryptases only make limited cleavages in a few proteins and are not inhibited by natural trypsin inhibitors, they form tetramers, bind heparin, and their activity on synthetic substrates is progressively inhibited as the concentration of salt increases above 0.2 M. Unique sequence features of seven tryptases were identified by comparison to other serine proteinases. The three-dimensional structures of the tryptases were then predicted by molecular modeling based on the crystal structure of bovine trypsin. The models show two large insertions to lie on either side of the active-site cleft, suggesting an explanation for the limited activity of tryptases on protein substrates and the lack of inhibition by natural inhibitors. A group of conserved Trp residues and a unique proline-rich region make two surface hydrophobic patches that may account for the formation of tetramers and/or inhibition with increasing salt. Although they contain no consensus heparin-binding sequence, the tryptases have 10-13 more His residues than trypsin, and these are positioned on the surface of the model. In addition, clustering of Arg and Lys residues may also contribute to heparin binding. Putative Asn-linked glycosylation sites are found on the opposite side of the model from the active site. The model provides structural explanations for some to the unusual characteristics of the tryptases and a rational basis for future experiments, such as site-directed mutagenesis.
Explore the source record for details and available documents.
An algorithm is presented for the accurate and rapid generation of multiple protein sequence alignments from tertiary structure comparisons. A preliminary multiple sequence alignment is performed using sequence information, which then determines an initial superposition of the structures. A structure comparison algorithm is applied to all pairs of proteins in the superimposed set and a similarity tree calculated. Multiple sequence alignments are then generated by following the tree from the branches to the root. At each branchpoint of the tree, a structure-based sequence alignment and coordinate transformations are output, with the multiple alignment of all structures output at the root. The algorithm encoded in STAMP (STructural Alignment of Multiple Proteins) is shown to give alignments in good agreement with published structural accounts within the dehydrogenase fold domains, globins, and serine proteinases. In order to reduce the need for visual verification, two similarity indices are introduced to determine the quality of each generated structural alignment. Sc quantifies the global structural similarity between pairs or groups of proteins, whereas Pij' provides a normalized measure of the confidence in the alignment of each residue. STAMP alignments have the quality of each alignment characterized by Sc and Pij' values and thus provide a reproducible resource for studies of residue conservation within structural motifs.
The annexins are a widespread family of calcium-dependent membrane-binding proteins. No common function has been identified for the family and, until recently, no crystallographic data existed for an annexin. In this paper we draw together 22 available annexin sequences consisting of 88 similar repeat units, and apply the techniques of multiple sequence alignment, pattern matching, secondary structure prediction and conservation analysis to the characterisation of the molecules. The analysis clearly shows that the repeats cluster into four distinct families and that greatest variation occurs within the repeat 3 units. Multiple alignment of the 88 repeats shows amino acids with conserved physicochemical properties at 22 positions, with only Gly at position 23 being absolutely conserved in all repeats. Secondary structure prediction techniques identify five conserved helices in each repeat unit and patterns of conserved hydrophobic amino acids are consistent with one face of a helix packing against the protein core in predicted helices a, c, d, e. Helix b is generally hydrophobic in all repeats, but contains a striking pattern of repeat-specific residue conservation at position 31, with Arg in repeats 4 and Glu in repeats 2, but unconserved amino acids in repeats 1 and 3. This suggests repeats 2 and 4 may interact via a buried saltbridge. The loop between predicted helices a and b of repeat 3 shows features distinct from the equivalent loop in repeats 1, 2 and 4, suggesting an important structural and/or functional role for this region. No compelling evidence emerges from this study for uteroglobin and the annexins sharing similar tertiary structures, or for uteroglobin representing a derivative of a primordial one-repeat structure that underwent duplication to give the present day annexins. The analyses performed in this paper are re-evaluated in the Appendix, in the light of the recently published X-ray structure for human annexin V. The structure confirms most of the predictions and shows the power of techniques for the determination of tertiary structural information from the amino acid sequences of an aligned protein family.
The programme pscan has been developed to distribute protein databank scans over a network of computers that share a common file system. pscan may be used in conjunction with most conventional sequence comparison programmes with few modifications. In test runs using the Smith-Waterman dynamic programming algorithm, the time required to scan a 6858 sequence databank using a query sequence 740 residues long was reduced from approximately 50 min for a single processor, to approximately 11 minutes for five processors. Accordingly, pscan provides a low-cost, portable alternative to dedicated parallel processing computers.
Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.
The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.
Explore the source record for details and available documents.
Methods for analyzing the amino-acid sequence of a protein for the purposes of predicting its three-dimensional structure were systematically analyzed using knowledge engineering techniques. The resulting entities (data) and relations (processing methods and constraints) have been represented within a generalized dependency network consisting of 29 nodes and over 100 links. It is argued that such a representation meets the requirements of knowledge-based systems in molecular biology. This network is used as the architecture for a prototype knowledge-based system that simulates logically the processes used in protein structure prediction. Although developed specifically for applications in protein structure prediction, the network architecture provides a strategy for tackling the general problem of orchestrating and integrating the diverse sources of knowledge that are characteristic of many areas of science.
Explore the source record for details and available documents.