Search PubMedSearch

Biomedical subjects

J M Thornton

Publications and source records attributed to J M Thornton.

At least 19 recordsLinked to original sources

Assigning genomic sequences to CATH.

We report the latest release (version 1.6) of the CATH protein domains database (http://www.biochem.ucl. ac.uk/bsm/cath ). This is a hierarchical classification of 18 577 domains into evolutionary families and structural groupings. We have identified 1028 homo-logous superfamilies in which the proteins have both structural, and sequence or functional similarity. These can be further clustered into 672 fold groups and 35 distinct architectures. Recent developments of the database include the generation of 3D templates for recognising structural relatives in each fold group, which has led to significant improvements in the speed and accuracy of updating the database and also means that less manual validation is required. We also report the establishment of the CATH-PFDB (Protein Family Database), which associates 1D sequences with the 3D homologous superfamilies. Sequences showing identifiable homology to entries in CATH have been extracted from GenBank using PSI-BLAST. A CATH-PSIBLAST server has been established, which allows you to scan a new sequence against the database. The CATH Dictionary of Homologous Superfamilies (DHS), which contains validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies, has been updated to include annotations associated with sequence relatives identified in GenBank. The DHS is a powerful tool for considering the variation of functional properties within a given CATH superfamily and in deciding what functional properties may be reliably inherited by a newly identified relative.

Amino Acid Sequence

Protein domain interfaces: characterization and comparison with oligomeric protein interfaces.

The physical and chemical properties of domain-domain interactions have been analysed in two-domain proteins selected from the protein classification, CATH. The two-domain structures were divided into those derived from (i) monomeric proteins, or (ii) oligomeric or complexed proteins. The size, polarity, hydrogen bonding and packing of the intra-chain domain interface were calculated for both sets of two-domain structures. The results were compared with inter-chain interface parameters from permanent and non-obligate protein-protein complexes. In general, the intra-chain domain and inter-chain interfaces were remarkably similar. Many of the intra-chain interface properties are intermediate between those calculated for permanent and non-obligate inter-chain complexes. Residue interface propensities were also found to be very similar, with hydrophobic residues playing a major role, together with positively charged arginine residues. In addition, the residue composition of the domain interfaces were found to be more comparable with domain surfaces than domain cores. The implications of these results for domain swapping and protein folding are discussed.

Amino Acids

Analysis and prediction of carbohydrate binding sites.

An analysis of the characteristic properties of sugar binding sites was performed on a set of 19 sugar binding proteins. For each site six parameters were evaluated: solvation potential, residue propensity, hydrophobicity, planarity, protrusion and relative accessible surface area. Three of the parameters were found to distinguish the observed sugar binding sites from the other surface patches. These parameters were then used to calculate the probability for a surface patch to be a carbohydrate binding site. The prediction was optimized on a set of 19 non-homologous carbohydrate binding structures and a test prediction was carried out on a set of 40 protein-carbohydrate complexes. The overall accuracy of prediction achieved was 65%. Results were in general better for carbohydrate-binding enzymes than for the lectins, with a rate of success of 87%.

Algorithms

The CATH Dictionary of Homologous Superfamilies (DHS): a consensus approach for identifying distant structural homologues.

A consensus approach has been developed for identifying distant structural homologues. This is based on the CATH Dictionary of Homologous Superfamilies (DHS), a database of validated multiple structural alignments annotated with consensus functional information for evolutionary protein superfamilies (URL: http://www. biochem.ucl.ac.uk/bsm/dhs). Multiple structural alignments have been generated for 362 well-populated superfamilies in the CATH structural domain database and annotated with secondary structure, physicochemical properties, functional sequence patterns and protein-ligand interaction data. Consensus functional information for each superfamily includes descriptions and keywords extracted from SWISS-PROT and the ENZYME database. The Dictionary provides a powerful resource to validate, examine and visualize key structural and functional features of each homologous superfamily. The value of the DHS, for assessing functional variability and identifying distant evolutionary relationships, is illustrated using the pyridoxal-5'-phosphate (PLP) binding aspartate aminotransferase superfamily. The DHS also provides a tool for examining sequence-structure relationships for proteins within each fold group.

Amino Acid Sequence

Protein folds, functions and evolution.

The evolution of proteins and their functions is reviewed from a structural perspective in the light of the current database. Protein domain families segregate unequally between the three major classes, the 32 different architectures and almost 700 folds observed to date. We find that the number of new topologies is still increasing, although 25 new structures are now determined for each new topology. The corresponding analysis and classification of function is only just beginning, fuelled by the genome data. The structural data revealed unexpected conservations and divergence of function both within and between families. The next five years will see the compilation of a definitive dictionary of protein families and their related functions, based on structural data which reveals relationships hidden at the sequence level. Such information will provide the foundation to build a better understanding of the molecular basis of biological complexity and hopefully to facilitate rational molecular design.

Animals

Correlation of observed fold frequency with the occurrence of local structural motifs.

It is well known that some protein folds (superfolds) occur very frequently. We show that compared to other folds, most superfold structures have a higher proportion of their alpha-helical or beta-strand residues in one of three basic units of supersecondary structure (alpha-hairpin, beta-hairpin or betaalphabeta-unit). Furthermore, by taking into consideration two more complex motifs, the four-stranded Greek-key (beta4) and the betaalpha-Greek key (betaalphabetabeta), we demonstrate that the remaining superfold structures contain many of these higher order units of three-dimensional packing. The implications of these results for folding are discussed.

Models, Molecular

Protein-DNA interactions: A structural analysis.

A detailed analysis of the DNA-binding sites of 26 proteins is presented using data from the Nucleic Acid Database (NDB) and the Protein Data Bank (PDB). Chemical and physical properties of the protein-DNA interface, such as polarity, size, shape, and packing, were analysed. The DNA-binding sites shared common features, comprising many discontinuous sequence segments forming hydrophilic surfaces capable of direct and water-mediated hydrogen bonds. These interface sites were compared to those of protein-protein binding sites, revealing them to be more polar, with many more intermolecular hydrogen bonds and buried water molecules than the protein-protein interface sites. By looking at the number and positioning of protein residue-DNA base interactions in a series of interaction footprints, three modes of DNA binding were identified (single-headed, double-headed and enveloping). Six of the eight enzymes in the data set bound in the enveloping mode, with the protein presenting a large interface area effectively wrapped around the DNA.A comparison of structural parameters of the DNA revealed that some values for the bound DNA (including twist, slide and roll) were intermediate of those observed for the unbound B-DNA and A-DNA. The distortion of bound DNA was evaluated by calculating a root-mean-square deviation on fitting to a canonical B-DNA structure. Major distortions were commonly caused by specific kinks in the DNA sequence, some resulting in the overall bending of the helix. The helix bending affected the dimensions of the grooves in the DNA, allowing the binding of protein elements that would otherwise be unable to make contact. From this structural analysis a preliminary set of rules that govern the bending of the DNA in protein-DNA complexes, are proposed.

Binding Sites

Three-dimensional structure analysis of PROSITE patterns.

Pattern matches for each of the sequence patterns in PROSITE, a database of sequence patterns, were searched in all protein sequences in the Brookhaven Protein Data Bank (PDB). The three-dimensional structures of the pattern matches for the 20 patterns with the largest numbers of hits were analysed. We found that the true positives have a common three-dimensional structure for each pattern; the structures of false positives, found for six of the 20 patterns, were clearly different from those of the true positives. The results suggest that the true pattern matches each have a characteristic common three-dimensional structure, which could be used to create a template to define a three-dimensional functional pattern.

Adenosine Triphosphate

Small subunit rRNA gene sequences of Aeromonas salmonicida subsp. smithia and Haemophilus piscium reveal pronounced similarities with A. salmonicida subsp. salmonicida.

The small subunit ribosomal RNA (SSU rRNA) encoding genes from reference strains of Aeromonas salmonicida subsp. smithia and Haemophilus piscium were amplified by polymerase chain reaction and cloned into Escherichia coli cells. Almost the entire SSU rRNA gene sequence (1505 nucleotides) from both organisms was determined. These DNA sequences were compared with those previously described from A. salmonicida subsp. salmonicida, subsp. achromogenes and subsp. masoucida. This genetic analysis revealed that A. salmonicida subsp. smithia and H. piscium showed 99.4 and 99.6% SSU rRNA gene sequence identity, respectively, with A. salmonicida subsp. salmonicida.

Aeromonas

The CATH Database provides insights into protein structure/function relationships.

We report the latest release (version 1.4) of the CATH protein domains database (http://www.biochem.ucl.ac.uk/bsm/cath). This is a hierarchical classification of 13 359 protein domain structures into evolutionary families and structural groupings. We currently identify 827 homologous families in which the proteins have both structual similarity and sequence and/or functional similarity. These can be further clustered into 593 fold groups and 32 distinct architectures. Using our structural classification and associated data on protein functions, stored in the database (EC identifiers, SWISS-PROT keywords and information from the Enzyme database and literature) we have been able to analyse the correlation between the 3D structure and function. More than 96% of folds in the PDB are associated with a single homologous family. However, within the superfolds, three or more different functions are observed. Considering enzyme functions, more than 95% of clearly homologous families exhibit either single or closely related functions, as demonstrated by the EC identifiers of their relatives. Our analysis supports the view that determining structures, for example as part of a 'structural genomics' initiative, will make a major contribution to interpreting genome data.

Algorithms

COVOL: an interactive program for evaluating second virial coefficients from the triaxial shape or dimensions of rigid macromolecules.

An interactive program is described for calculating the second virial coefficient contribution to the thermodynamic nonideality of solutions of rigid macromolecules based on their triaxial dimensions. The FORTRAN-77 program, available in precompiled form for the PC, is based on theory for the covolume of triaxial ellipsoid particles [Rallison, J. M., and S.E Harding. (1985). J. Colloid Interface Sci. 103:284-289]. This covolume has the potential to provide a magnitude for the second virial coefficient of macromolecules bearing no net charge. Allowance for a charge-charge contribution is made via an expression based on Debye-Hückel theory and uniform distribution of the net charge over the surface of a sphere with dimensions governed by the Stokes radius of the macromolecule. Ovalbumin, ribonuclease A, and hemoglobin are used as model systems to illustrate application of the COVOL routine.

Animals

Polymerase chain reaction (PCR)-based typing analysis of atypical isolates of the fish pathogen Aeromonas salmonicida.

Two hundred and five isolates of atypical Aeromonas salmonicida, recovered from a wide range of hosts and countries were characterized by polymerase chain reaction (PCR) targeting four genes. The chosen genes were those encoding the extracellular A-layer protein (AP), the serine protease (Sprot), the glycerophospholipid:cholestrol acetyltransferase protein (GCAT), and the 16S rRNA (16S rDNA). All the atypical A. salmonicida isolates could be assigned to 4 PCR groups. Group 1 comprised 45 strains which tested positive for PCR amplification, using the 16S rDNA, GCAT2, Sprot2, and AP primer-sets. Group 2 comprised 88 strains with produced PCR products using the 16S rDNA, GCAT2 and AP primer-sets. Group 3 comprised 21 strains which produced PCR products using 16S rDNA, GCAT2 and Sprot2 primer-sets, and group 4 comprised 51 strains which produced PCR products using the 16S rDNA and GCAT2 primer-sets only. A. salmonicida subsp. salmonicida isolates tested, belonged to group 1. The PCR primer-sets separated A. salmonicida from other reference strains of Aeromonas species and related bacteria with the exception of Aeromonas hydrophila. The results indicated that PCR typing is a useful framework for characterization of the increasing number of isolations of atypical A. salmonicida.

Acyltransferases

From protein structure to function.

Several databases of protein structural families now exist-organised according to both evolutionary relationships and common folding arrangements. Although these lag behind sequence databases in size, the prospect of structural genomics initiatives means that they may soon include representatives of many of the sequence families. To some extent, functional information can be derived from structural similarity. For some structural families, their function is highly conserved, whereas, for others, it can only be inherited or derived on the basis of additional information (e.g. sequence patterns, common residue clusters and characteristic surface properties).

Computational Biology

Evolution of protein function, from a structural perspective.

The recent growth in structural data, and ensuing analyses, have revealed the structural and functional versatility of protein families. With respect to enzymes, local active-site mutations, variations in surface loops and recruitment of additional domains accommodate the diverse substrate specificities and catalytic activities observed within several superfamilies. Conversely, some functions have more than one structural solution, having evolved independently several times during evolution. Combined with the existence of multi-functional genes, which have arisen by gene recruitment, these phenomena must be considered in the process of genome annotation.

Animals

DOMPLOT: a program to generate schematic diagrams of the structural domain organization within proteins, annotated by ligand contacts.

A program is described for automatically generating schematic linear representations of protein chains in terms of their structural domains. The program requires the co-ordinates of the chain, the domain assignment, PROSITE information and a file listing all intermolecular interactions in the protein structure. The output is a PostScript file in which each protein is represented by a set of linked boxes, each box corresponding to all or part of a structural domain. PROSITE motifs and residues involved in ligand interactions are highlighted. The diagrams allow immediate visualization of the domain arrangement within a protein chain, and by providing information on sequence motifs, and metal ion, ligand and DNA binding at the domain level, the program facilitates detection of remote evolutionary relationships between proteins.

Adenosine Diphosphate

Protein side-chain conformation: a systematic variation of chi 1 mean values with resolution - a consequence of multiple rotameric states?

A systematic variation with resolution of the mean values of the gauche-, trans and gauche+ chi1 rotamers in protein structures determined by X-ray crystallography has been observed. Further analysis revealed that these correlations differ considerably between residue types, being highly significant for some residue types (e.g. Ser, Thr, Leu, Lys) and absent for others (e.g. aromatics). For the individual residue types which exhibited the trend most strongly, these changes were accompanied by corresponding systematic variations in the percentage relative populations in the three energy wells. Examination of a uniformly sized subset of monomers showed that this effect, while attenuated, was still present, and was thus not entirely a consequence of the change in size and surface area which also correlates with resolution. An analysis of B values in the disfavoured high-energy barrier region between the rotameric wells showed a pronounced tendency towards larger than average values. As a plausible hypothesis, it is suggested here that these observations can be accounted for by the presence of multiple rotameric states. The averaged electron density produced by dual occupancy at low resolution giving an averaged conformation is resolved at high resolution into its individual components.

Asparagine

Barrel structures in proteins: automatic identification and classification including a sequence analysis of TIM barrels.

Automated methods for identifying and characterizing regular beta-barrels from coordinate data have been developed to analyze and classify various kinds of barrel structures based on geometric parameters such as the barrel strand number (n) and shear number (S). In total, we find 1,316 barrels in the January 1998 release of Protein Data Bank. Of 1,316 barrels, 1,277 barrels had an even shear number, corresponding to 50 nonhomologous families. The (beta alpha)8 triose phosphate isomerase (TIM) barrel (n = 8, S = 8) fold has the largest number of apparently nonhomologous entries, 16, although the trypsin like antiparallel (n = 6, S = 8) barrels (representing only three families) are the most common with 527 barrels. Of all the protein families that exhibit barrel structures, 68% are found to be various kinds of enzymes, the remainder being binding proteins and transport membrane proteins. In addition, the layers of side chains, which form the cores of barrels with S = n and S = 2n, are also analyzed. More sophisticated methods were developed for detecting TIM barrels specifically, including consideration of the amino acid propensities for the side chains that form the layers. We found that the residues on the outside of the eight stranded parallel beta-barrel, buried by the alpha-helices, are much more hydrophobic than the residues inside the barrel.

Hydrogen Bonding

Factors limiting the performance of prediction-based fold recognition methods.

In the past few years, a new generation of fold recognition methods has been developed, in which the classical sequence information is combined with information obtained from secondary structure and, sometimes, accessibility predictions. The results are promising, indicating that this approach may compete with potential-based methods (Rost B et al., 1997, J Mol Biol 270:471-480). Here we present a systematic study of the different factors contributing to the performance of these methods, in particular when applied to the problem of fold recognition of remote homologues. Our results indicate that secondary structure and accessibility prediction methods have reached an accuracy level where they are not the major factor limiting the accuracy of fold recognition. The pattern degeneracy problem is confirmed as the major source of error of these methods. On the basis of these results, we study three different options to overcome these limitations: normalization schemes, mapping of the coil state into the different zones of the Ramachandran plot, and post-threading graphical analysis.

Algorithms