Fibronectin type III domains in yeast detected by a hidden Markov model.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to C Chothia.
Explore the source record for details and available documents.
The L1 cell adhesion molecule has six domains homologous to members of the immunoglobulin superfamily and five homologous to fibronectin type III domains. We determined the outline structure of the L1 domains by showing that they have, at the key sites that determine conformation, residues similar to those in proteins of known structure. The outline structure describes the relative positions of residues, the major secondary structures and residue solvent accessibility. We use the outline structure to investigate the likely effects of 22 mutations that cause neurological diseases. The mutations are not randomly distributed but cluster in a few regions of the structure. They can be divided into those that act mainly by changing conformation or denaturing their domain and those that alter its surface properties.
We have determined the packing efficiency at the protein-water interface by calculating the volumes of atoms on the protein surface and nearby water molecules in 22 crystal structures. We find that an atom on the protein surface occupies, on average, a volume approximately 7% larger than an atom of equivalent chemical type in the protein core. In these calculations, larger volumes result from voids between atoms and thus imply a looser or less efficient packing. We further find that the volumes of individual atoms are not related to their chemical type but rather to their structural location. More exposed atoms have larger volumes. Moreover, the packing around atoms in locally concave, grooved regions of protein surfaces is looser than that around atoms in locally convex, ridge regions. This as a direct manifestation of surface curvature-dependent hydration. The net volume increase for atoms on the protein surface is compensated by volume decreases in water molecules near the surface. These waters occupy volumes smaller than those in the bulk solvent by up to 20%; the precise amount of this decrease is directly related to the extent of contact with the protein.
We report a prediction that two prokaryotic proteins contain immunoglobulin superfamily domains. Immunoglobulin-like folds have been identified previously in prokaryotic proteins, but these share no recognizable sequence similarity with eukaryotic immunoglobulin superfamily (IgSF) folds, and may be the result of the physics and chemistry of proteins favoring certain common folds. In contrast, the prokaryotic proteins identified have sequences whose match to the immunoglobulin superfamily can be detected by hidden Markov modeling, BLASTP matches, key residue analysis, and secondary structure predictions. We propose that these prokaryotic immunoglobulin-like domains are almost certain to be related by divergence from a common ancestor to eukaryotic immunoglobulin superfamily domains.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
In humans, the gene for the V kappa domain is produced by the recombination of one of 40 functional V kappa segments and one of five functional J kappa segments. We have analysed the sequences of these germline segments and of 736 rearranged V kappa genes to determine the repertoire of main chain conformations, or canonical structures, they encode. Over 96% of the sequences correspond to one of four canonical structures for the first antigen binding loop (L1) and one canonical structure for the second antigen binding loop (L2). Junctional diversity produces some variation in the length of the third antigen binding loop (L3) and in the identity of residues at the V kappa-J kappa join. However, this is limited and 70% of the rearranged sequences correspond to one of three known canonical structures for the L3 region. Furthermore, we show that the canonical structures selected during the primary response are conserved during affinity maturation: the key residues that determine the conformations of the antigen binding loops are unmutated or undergo conservative mutation. The implications of these results for immune recognition are discussed.
To facilitate understanding of, and access to, the information available for protein structures, we have constructed the Structural Classification of Proteins (scop) database. This database provides a detailed and comprehensive description of the structural and evolutionary relationships of the proteins of known structure. It also provides for each entry links to co-ordinates, images of the structure, interactive viewers, sequence data and literature references. Two search facilities are available. The homology search permits users to enter a sequence and obtain a list of any structures to which it has significant levels of sequence similarity. The key word search finds, for a word entered by the user, matches from both the text of the scop database and the headers of Brookhaven Protein Databank structure files. The database is freely accessible on World Wide Web (WWW) with an entry point to URL http: parallel scop.mrc-lmb.cam.ac.uk magnitude of scop.
Fibroblast growth factor receptors (FGFRs) have three extracellular domains that belong to the immunoglobulin superfamily. We have determined the outline structures for these domains on the basis of their homology to the I set molecule telokin. The outline structures describe the relative positions of residues in each domain; their major secondary structures, and the extent to which residues are accessible to the solvent. They also provide the basis of a coherent description of the change in recognition properties that occur when the IIIb and IIIc exons are switched and of the effects of mutations in FGFRs that cause genetic diseases.
BACKGROUND: Protein volumes change very little on folding at low pressure, but at high pressure the unfolded state is more compact. So far, the molecular origins of this behaviour have not been explained: it is the opposite of that expected from the model of the hydrophobic effect based on the transfer of non-polar solutes from water to organic solvent. RESULTS: We redetermined the mean volumes occupied by residues in the interior of proteins. The new residue volumes are smaller than those given by previous calculations which were based on much more limited data. They show that the packing density in protein interiors is exceptionally high. Comparison of the volumes that residues occupy in proteins with those they occupy in solution shows that aliphatic groups have smaller volumes in protein interiors than in solution, while peptide and charged groups have larger volumes. The cancellation of these volume changes is the reason that the net change on folding is very small. CONCLUSIONS: The exceptionally high density of the protein interior shown here implies that packing forces play a more important role in protein stability than has been believed hitherto.
We survey all the known instances of domain movements in proteins for which there is crystallographic evidence for the movement. We explain these domain movements in terms of the repertoire of low-energy conformation changes that are known to occur in proteins. We first describe the basic elements of this repertoire, hinge and shear motions, and then show how the elements of the repertoire can be combined to produce domain movements. We emphasize that the elements used in particular proteins are determined mainly by the structure of the interfaces between the domains.
On the basis of similarities in sequence and structure, the protein domains that form the immunoglobulin superfamily have been divided into three sets: one with variable-like domains, the V set, and two with different variants of the constant-like domains, the C1 and C2 sets. Examination of a muscle member of the immunoglobulin superfamily, telokin, shows that its structure is closely related to those of the variable domains found in antibodies, CD2, CD4 and CD8. However, it also contains structural features that, previously, have only been found in constant domains. Telokin represents a new structural set in the superfamily which we call the I set. Using the structures of telokin, and variable domains from antibodies, CD4 and CD8, we constructed a profile that describes the sequence characteristics of the structural core common to those proteins. This sequence profile makes a good match to the sequences of many of the immunoglobulin superfamily domains that form the cell adhesion molecules and surface receptors. This match implies that these domains also have structures that belong to the I set.
The major feature of many proteins is a large beta-sheet that twists and coils to form a closed structure in which the first strand is hydrogen bonded to the last: the beta-sheet barrel. McLachlan classified barrels in terms of two integral parameters: the number of strands in the beta-sheet, n, and the "shear number", S, a measure of the stagger of the strands in the beta-sheet. He showed that the mean radius of a barrel and the extent to which strands are tilted relative to its axis are determined by the values of n and S. Here we show that the (n, S) values determine all the other general structural features of regular beta-sheet barrels, in particular, optimal values of the twist and coiling angles that produce the closed beta-sheet, the hyperboloidal shape and the arrangement of residues in the barrel interior. Consideration of the residue arrangements in the interiors of different potential barrel structures, and of side-chain volumes, suggest that barrels, in which the interiors are close packed by the residues in beta-sheets with good geometries, have structures that correspond to one of only ten different combinations of n and S. In the accompanying paper, we demonstrate, by an analysis of all observed protein structures that contain beta-sheet barrels and for which atomic co-ordinates are available, the validity of these theoretical results.
In the accompanying paper we derived a set of principles that, we argue, govern the structure of beta-sheet barrels. Barrel structures are classified in terms of two integral parameters: the number of strands in the beta-sheet, n, and a measure of the stagger in the beta-sheet, S. We derived a set of equations that show how the (n, S) values of a barrel structure determine the arrangement of its strands; its general shape; the twist and coiling of the beta-sheet, and the arrangement of residues in the barrel interior. This work suggested that there are ten different combinations of n and S that form barrels with good beta-sheet geometries and interiors close packed by beta-sheet residues. In this paper we demonstrate the validity of these principles. We analyse in detail the observed structures of 39 different beta-sheet barrels. These structures include representatives of all the different barrel structures currently known and for which atomic co-ordinates are available. We show that the observed arrangement of the strands, and the extent of the twist and coiling of the beta-sheets, are very close to those calculated from the (n, S) values for the barrel. Of the 39 structures, 34 have one of the ten (n, S) values that we expect to form barrels with good beta-sheet geometries and interiors close packed by beta-sheet residues. The other five have one of two (n, S) values that give good beta-sheet geometries but radii so large the beta-sheet residues leave cavities at the centre of the barrels. In at least four of these cavities have a functional role.
We have determined the variations in volume that occur during evolution in the buried core of three different families of proteins. The variation of the whole core is very small (approximately 2.5%) compared to the variation at individual sites (approximately 13%). However, by comparing our results to those expected from random sequences with no correlations between sites, we show that the small variation observed may simply be a manifestation of the statistical "law of large numbers" and not reflect any compensating changes in, or global constraints upon, protein sequences. We have also analysed in detail the volume variations at individual sites, both in the core and on the surface, and compared these variations with those expected from random sequences. Individual sites on the surface have nearly the same variation as random sequences (24% versus 28% variation). However, individual sites in the core have about half the variation of random sequences (13% versus 30%). Roughly, half of these core sites strongly conserve their volume (0 to 10% variation); one quarter have moderate variation (10 to 20%); and the remaining quarter vary randomly (20 to 40%). Our results have clear implications for the relationship between protein sequence and structure. For our analysis, we have developed a new and simple method for weighting protein sequences to correct for unequal representation, which we describe in an Appendix.
Analysis of human and mouse immunoglobulins has shown that five of six hypervariable regions that form the antigen binding site have a small repertoire of main chain conformations (canonical structures). Cartilaginous fishes are the most distantly related species to humans known to have an immune system, their evolutionary lines having diverged 450 million years ago. An analysis of VH and V kappa sequences from these fishes shows that all the main chain structures in their L1, L2, H1 and H2 hypervariable regions, and one of those in the L3 region, are the same as those most commonly found in human and mouse. This implies that the canonical structures occurring most commonly in hypervariable regions arose very early in the stages of the evolution of the immune system.
The evolution of development involves the development of new proteins. Estimates based on the initial results of the genome projects, and on the data banks of protein sequences and structures, suggest that the large majority of proteins come from no more than one thousand families. Members of a family are descended from a common ancestor. Protein families evolve by gene duplication and mutation. Mutations change the conformation of the peripheral regions of proteins; i.e. the regions that are involved, at least in part, in their function. If mutations proceed until only 20% of the residues in related proteins are identical, it is common for the conformational changes to affect half the structure. Most of the proteins involved in the interactions of cells, and in their assembly to form multicellular organisms, are mosaic proteins. These are large and have a modular structure, in that they are built of sets of homologous domains that are drawn from a relatively small number of protein families. Patthy's model for the evolution of mosaic proteins describes how they arose through the insertion of introns into genes, gene duplications and intronic recombination. The rates of progress in the genome sequencing projects, and in protein structure analyses, means that in a few years we will have a fairly complete outline description of the molecules responsible for the structure and function of organisms at several different levels of developmental complexity. This should make a major contribution to our understanding of the evolution of development.
Explore the source record for details and available documents.