Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “structure”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Structurally diverse copper(II)-carboxylato complexes: neutral and ionic mononuclear structures and a novel binuclear structure.

The copper complexes with the commercial auxin herbicides MCPA, 2,4-D, and 2,4,5-T in the presence of a nitrogen donor heterocyclic ligand, phen or bipyam, were prepared and characterized. The available evidence supports a dimeric structure for the 2,4-D complex in the presence of bipyam while phen leads to monomeric forms. The EPR spectrum of Cu2(2,4-D)4(bipyam)2 at 4 K in the solid state exhibits an axial signal which corresponds to almost isolated S = 1/2 magnetic ions. Magnetic data for the dimer show a weak antiferromagnetic interaction between the two metal ions with J = -08 cm-1. The crystal structures of tetrakis[(2,4-dichlorophenoxy)acetato]bis(2,2'-bipyridylamine)dicopper(II), 1, bis(1,10-phenanthroline)[(2,4,5-trichlorophenoxy)acetato]copper(II) chloride, 2, and aqua(1,10-phenanthroline)bis[((2-methyl-4-chlorophenoxyacetato]copper(II), 3, were determined and refined by least-squares methods using three-dimensional MoK alpha data. 1 crystallizes in space group P1, in a cell of dimensions a = 10813(1) A, b = 12138(1) A, c = 11909(1) A, alpha = 86448(3) degrees, beta = 80127(3) degrees, and gamma = 63982(3) degrees, and V = 13837(2) A3, with Z = 1 2 crystallizes in space group I2/a, in a cell of dimensions a = 29958(9) A, b = 11342(3) A, c = 21196(7) A, beta = 10794(1) degrees, and V = 68522(4) A3, with Z = 8 3 crystallizes in space group P1, in a cell of dimensions a = 87419(8) A, b = 12512(1) A, c = 14598(1) A, alpha = 110737(1) degrees, beta = 95742(2) degrees, gamma = 103286(2) degrees, V = 14241(2) A3, with Z = 2.

Journal Article↗

The high-resolution X-ray crystal structure of the complex formed between subtilisin Carlsberg and eglin c, an elastase inhibitor from the leech Hirudo medicinalis. Structural analysis, subtilisin structure and interface geometry.

Triclinic crystals of the complex formed by eglin with subtilisin Carlsberg were analyzed by X-ray diffraction. The crystal and molecular structure of this complex was determined with data that extended to 0.12-nm resolution by a combination of Patterson search methods and isomorphous replacement techniques. Its structure was refined to a crystallographic R value of 0.178 (1.0-0.12 nm) using an energy-restraint least-squares procedure. The complete subtilisin molecule could be traced without ambiguity in the refined electron density. The eglin component, from which an amino-terminal segment is cleaved off, is only defined from Lys8I (i.e. the lysine residue 8 of the inhibitor) onwards. Per unit cell, 436 fixed solvent molecules and 2 calcium ions were located. In spite of 84 amino acid replacements and one deletion, subtilisin Carlsberg exhibits a very similar polypeptide fold to subtilisin BPN'. The root-mean-square deviations of all alpha-carbon atoms (excluding those at the deletion site) from models of subtilisin BPN' [Alden, R. A., Birktoft, J. J., Kraut, J., Robertus, J. D. & Wright, C. S. (1971) Biochem. Biophys. Res. Commun. 45, 337-344] and subtilisin Novo [Drenth, J., Hol, W. G. J., Jansonius, J. N. & Kockoek, R. (1972) Eur. J. Biochem. 25, 177-181] are 0.077 nm and 0.103 nm. Most of these deviations result from global shifts rather than changes of the local geometry. The single-residue deletion at position 56 affects only the surrounding conformation. Two sites of high electron density and close distances to surrounding oxygen ligands have been found in the Carlsberg enzyme which are probably occupied by calcium ions. Eglin consists of a twisted four-stranded beta-sheet flanked by an alpha-helix and by an exposed proteinase binding loop on opposite sides. Around the reactive site, Leu45I-Asp46I, this loop is mainly stabilized by electrostatic/hydrogen bond interactions with the side chains of two arginine residues which project from the hydrophobic core [Bode, W., Papamokos, E., Musil, D., Seemüller, W. & Fritz, H. (1986) EMBO J. 5, 813-818]. The reactive site loop conformation resembles that found in other 'small' proteinase inhibitors. The scissile peptide bond is not cleaved but its carbonyl group is slightly distorted from planar geometry. Most of the intermolecular contacts are contributed by the nine residues of the reactive-site loop Gly40I-Arg48I.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Segment 8 encodes a structural protein of infectious salmon anaemia virus (ISAV); the co-linear transcript from Segment 7 probably encodes a non-structural or minor structural protein.

In this study we present the cloning, expression and partial identification of Genomic Segment 7 of infectious salmon anaemia virus (ISAV). The nucleotide sequence corresponding to Segment 7 was isolated from a bacteriophage lambda cDNA library and contained 2 overlapping open reading frames (ORFs) of 903 and 522 bases respectively. It also contained an ISAV-specific conserved nucleotide motif in the mRNA 5' region. The co-linear transcript representing the large ORF undergoes a splicing event that removes a 526 nucleotide intron to form a mRNA corresponding to the smaller reading frame. Thus, ISAV Genomic Segment 7 has a similar coding strategy as influenza A virus Segments 7 and 8. The largest ORF of Segment 7 and the first ORF of Segment 8 was expressed in E. coli as fusion proteins and rabbit antiserum was raised against the recombinant protein from Segment 8. Immunoblot studies using this antiserum and a serum against purified virus, show that Segment 8 encodes one of the major structural proteins of the virus whereas the co-linear ORF of Segment 7 probably encodes a non- or minor structural protein

Amino Acid Sequence↗

[Rule of antibody structure: the primary structure of a human monoclonal IgA1-immunoglobulin (myeloma protein Tro), V. The arrangement of the tryptic peptides and a discussion of the complete primary structure of the H-chain (author's transl)].

This communication deals with the sequence work done with tryptic and chymotryptic peptides and some cyanogen bromide splitting products. With these peptides, and if necessary with their splitting peptides, the whole primary structure of the alpha1-H-chain of myeloma protein Tro is established. The position of the amides is determined by electrophoresis and digestion with aminopeptidase M. The alpha1-chain Tro comprises 475 amino acid residues. Because of its specific exchanges and deletions the variable part of alpha1-chain Tro belongs to subgroup III of variable parts of H-chains. The switch from the variable to the constant part occurs at position 119/120 and is analogous to other chains which have been sequenced up to now. The large number of cysteine residues in the alpha-chain which may influence the tertiary structure, especially in the hinge and the subsequent CH2-region, is noteworthy. Furthermore, myeloma protein Tro is compared with the other alpha1-chain Bur[5] sequenced in the meantime, and protein But[6], which is an IgA2 molecule of the allotype A2m(2).

Amino Acid Sequence↗

[Primary structure of the elongation factor G from Escherichia coli. IX. Structure of peptides generated by cyanogen bromide cleavage of the G-factor isolated on thiol-activated sepharose and of the products of the G-factor cleavage at Asp-Pro bonds. Complete primary structure].

The amino acid sequence of cysteine- and cystine-containing peptides resulting from cleavage of the G-factor by cyanogen bromide has been determined. For structure analysis cyanogen bromide peptides were further degradated using trypsin, chymotrypsin, thermolysin, staphylococcal glutamic protease, or limited acid hydrolysis. The products of the G-factor cleavage at Asp-Pro bonds were also studied. The obtained data together with those published earlier permitted to establish the complete primary structure of the elongation factor G. The polypeptide chain consists of 701 amino acid residues and has molecular mass of 77321,46.

Amino Acid Sequence↗

The primary structure of human serum transferrin. The structures of seven cyanogen bromide fragments and the assembly of the complete structure.

The amino acid sequences of seven cyanogen bromide fragments of human serum transferrin have been determined, and the primary structure of transferrin established by determining the order of these and three additional fragments (Sutton, M. R., MacGillivray, R. T. A., and Brew, K. (1975) Eur. J. Biochem. 51, 43-48) in the polypeptide chain. The order of the fragments was deduced from peptides that overlap methionyl residues which were obtained by thermolysin digestion of performic acid-oxidized transferrin or by partial peptic hydrolysis of unmodified transferrin, together with other evidence. The polypeptide chain of transferrin contains 679 amino acid residues, which together with the two N-linked oligosaccharide chains gives a calculated molecular weight of 79,570. Transferrin consists of two homologous domains (residues 1-336, 337-679), each associated with a single Fe-binding site, with both sites of glycosylation in the carboxyl-terminal domain at positions 413 and 611. Consideration of the primary structure in relation to previously published results provides information concerning the evolutionary development of transferrins and related proteins, and the locations of metal-binding residues in the transferrin molecule.

Amino Acid Sequence↗

[The rule of antibody structure. The primary structure of a monoclonal IgG1 immunoglobulin (myeloma protein Nie). III. The chymotryptic peptides of the H-chain, alignment of the tryptic peptides and discussion of the complete structure].

In this final paper the complete primary structure of the H-chain of immunoglobulin Nie (IgG1, Gm1+, 17+) is established by overlapping tryptic fragments with chymotryptic peptides. The preceding papers dealt with the purification of the protein, the characterization of the light and heavy chains, the purification and characterization of the cyanogen bromide cleavage products, the location of the disulfide bonds, the isolation of the tryptic peptides and their sequence determination. The gamma1-chain Nie comprises 448 amino acid residues. When the protein is compared with other H-chains, the switch from the variable to the constant part occurs at position 119/120. Based on the amino acid sequence of the variable part, protein Nie belongs to subgroup III of the H-chains. It was the first protein of this subgroup to be sequenced. In the meantime several other proteins are known which have been assigned to the same subgroup on the basis of linked amino acid exchanges in comparison to members of other subgroups. This confirms the evolutionary origin of antibody variability and hence the genetically fixed antibody specific. Furthermore protein Nie is the first completely determined chain with the genetic factors Gm1+, 17+. These factors are inherited codominantly and are localized on the constant part of the gamma1-chain. By comparison with protein Eu, which is Gml-, 4+ and therefore an allele of Nie, these serologically defined factors are correlated with Eu. Besides the amino acid exchanges caused by the Gm-factors we elucidated a series of differences to the constant part of the protein Eu. These differences include 6 amide postions and the sequence from residues 387 to 391. Using the structure of IgG1 Nie as an example some rules for the evolution of immunoglobulin sequences have been described. In particular the "elongation-rule" and the "Disulfide-rule" are discussed. While chain-elongation of the H-chains can simply be explained by repeated gene duplications of a basic unit containing ca 110 amino acids, the location of disulfide bonds is determined partly by gene duplication, which implies multiplication of evolutionary "old" cystein residues and partly by the relatively recent acquisition of "new" cystein in appropriate sites. Most evident is the origin of the "hindge-region" by partial gene duplication on the C-terminal residues of the first homology region.

Amino Acid Sequence↗

Structural features can be unconserved in proteins with similar folds. An analysis of side-chain to side-chain contacts secondary structure and accessibility.

Side-chain to side-chain contacts, accessibility, secondary structure and RMS deviation were compared within 607 pairs of proteins having similar three-dimensional (3D) structures. Three types of protein 3D structural similarities were defined: type A having sequence and usually functional similarity; type B having functional, but no sequence similarity; and type C having only 3D structural similarity. Within proteins having little or no sequence similarity (types B and C), structural features frequently had a degree of conservation comparable to dissimilar 3D structures. Despite similar protein folds, as few as 30% of residues within similar protein 3D structures can form a common core. RMS deviations on core C alpha atoms can be as high as 3.2 A. Similar protein structures can have secondary structure identities as low as 41%, which is equivalent to that expected by chance. By defining three categories of amino acid accessibility (buried, half buried and exposed), some similar protein 3D structures have as few as 30% of positions in the same category, making them indistinguishable from pairs of dissimilar protein structures. Similar structures can also have as few as 12% of common side-chain to side-chain contacts, and virtually no similar energetically favourable side-chain to side-chain interactions. Complementary changes are defined as structurally equivalent pairs of interacting residues in two structures with energetically favourable but different side-chain interactions. For many proteins with similar three-dimensional structures, the proportion of complementary changes is near to that expected by chance, suggesting that many similar structures have fundamentally different stabilising interactions. All of the results suggest that proteins having similar 3D structures can have little in common apart from a scaffold of core secondary structures. This has profound implications for methods of protein fold detection, since many of the properties assumed to be conserved across similar protein 3D structures (e.g. accessibility, side-chain to side-chain contacts, etc.) are often unconserved within weakly similar (i.e. type B and C) protein 3D structures. Little difference was found between type B and C similarities suggesting that the structure of similar proteins can evolve beyond recognition even when function is conserved. Our findings suggest that it is more general features of protein structure, such as the requirements for burial of hydrophobic residues and exposure of polar residues, rather than specific residue-residue interactions that determine how well a particular sequence adopts a particular fold.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Local structure prediction with local structure-based sequence profiles.

MOTIVATION: A large body of experimental and theoretical evidence suggests that local structural determinants are frequently encoded in short segments of protein sequence. Although the local structural information, once recognized, is particularly useful in protein structural and functional analyses, it remains a difficult problem to identify embedded local structural codes based solely on sequence information. RESULTS: In this paper, we describe a local structure prediction method aiming at predicting the backbone structures of nine-residue sequence segments. Two elements are the keys for this local structure prediction procedure. The first key element is the LSBSP1 database, which contains a large number of non-redundant local structure-based sequence profiles for nine-residue structure segments. The second key element is the consensus approach, which identifies a consensus structure from a set of hit structures. The local structure prediction procedure starts by matching a query sequence segment of nine consecutive amino acid residues to all the sequence profiles in the local structure-based sequence profile database (LSBSP1). The consensus structure, which is at the center of the largest structural cluster of the hit structures, is predicted to be the native state structure adopted by the query sequence segment. This local structure prediction method is assessed with a large set of random test protein structures that have not been used in constructing the LSBSP1 database. The benchmark results indicate that the prediction capacities of the novel local structure prediction procedure exceed the prediction capacities of the local backbone structure prediction methods based on the I-sites library by a significant margin. AVAILABILITY: All the computational and assessment procedures have been implemented in the integrated computational system PrISM.1 (Protein Informatics System for Modeling). The system and associated databases for LINUX systems can be downloaded from the website: http://www.columbia.edu/~ay1/.

Amino Acid Sequence↗

Local structure-based sequence profile database for local and global protein structure predictions.

MOTIVATION: A large body of evidence suggests that protein structural information is frequently encoded in local sequences-sequence-structure relationships derived from local structure/sequence analyses could significantly enhance the capacities of protein structure prediction methods. In this paper, the prediction capacity of a database (LSBSP2) that organizes local sequence-structure relationships encoded in local structures with two consecutive secondary structure elements is tested with two computational procedures for protein structure prediction. The goal is twofold: to test the folding hypothesis that local structures are determined by local sequences, and to enhance our capacity in predicting protein structures from their amino acid sequences. RESULTS: The LSBSP2 database contains a large set of sequence profiles derived from exhaustive pair-wise structural alignments for local structures with two consecutive secondary structure elements. One computational procedure makes use of the PSI-BLAST alignment program to predict local structures for testing sequence fragments by matching the testing sequence fragments onto the sequence profiles in the LSBSP2 database. The results show that 54% of the test sequence fragments were predicted with local structures that match closely with their native local structures. The other computational procedure is a filter system that is capable of removing false positives as possible from a set of PSI-BLAST hits. An assessment with a large set of non-redundant protein structures shows that the PSI-BLAST + filter system improves the prediction specificity by up to two-fold over the prediction specificity of the PSI-BLAST program for distantly related protein pairs. Tests with the two computational procedures above demonstrate that local sequence-structure relationships can indeed enhance our capacity in protein structure prediction. The results also indicate that local sequences encoded with strong local structure propensities play an important role in determining the native state folding topology.

Algorithms↗

Refinement of the three-dimensional solution structure of barley serine proteinase inhibitor 2 and comparison with the structures in crystals.

The three-dimensional structure of barley serine proteinase inhibitor, CI-2, has been determined using nuclear magnetic resonance spectroscopy. The present structure determination is a refinement of the structure previously determined by us, using in the present case stereo-specific assignments, and a virtually complete set of assignments of the two-dimensional nuclear Overhauser spectrum. The structure determination is based on the identification of more than 1300 nuclear Overhauser effects, of which 961 were used in the structure calculation as distance restraints, and on 94 dihedral angle restraints, of which 31 are for chi 1 angles in defined chiral centers. These have been used to calculate a series of 20 three-dimensional structures using a combination of distance geometry, simulated annealing and restrained molecular dynamics. Each of the 20 structures was in agreement within less than 0.5 A of each of the distance restraints and with all dihedral angle restraints. When compared to the geometric average structure of the 20 refined structures the root-mean-square differences for the backbone atoms were 0.8 (+/- 0.2) A and for all atoms were 1.6 (+/- 0.2) A. By comparison, the values obtained for the structures determined previously were 1.4 (+/- 0.2) A and 2.1 (+/- 0.1) A, respectively. The structures were also compared to the structure determined in the crystalline state by X-ray diffraction showing root-mean-square differences of 1.6 (+/- 0.2) A and 2.8 (+/- 0.2) A for the backbone and all atoms, respectively. Common features of the solution structure and the two crystal structures are the four-stranded beta-structure, composed of a pair of parallel strands, and three pairs of antiparallel beta-strands flanked on one side by a 12-residue alpha-helix and on the other side by a loop containing the serine proteinase binding site. The new analysis of the structure has revealed an additional pair of antiparallel beta-strands, consisting of residues 65 to 67 and 81 to 83, that was not seen in either of the crystal structures or the previous solution structure. Identification of this was based on nuclear magnetic resonance evidence for the hydrogen bond (67HN to 81CO) not reported previously. Also the presence of a bifurcated hydrogen bond involving Phe69 CO and HN atoms of Ala77 and Gln78 was observed in solution but not in crystals. Minor differences between the two structures were observed in the phi-angles of residues Met59 and Glu60 in the inhibitory site.

Crystallography↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

Dissociable processes for learning the surface structure and abstract structure of sensorimotor sequences.

A sensorimotor sequence may contain information structure at several different levels. In this study, we investigated the hypothesis that two dissociable processes are required for the learning of surface structure and abstract structure, respectively, of sensorimotor sequences. Surface structure is the simple serial order of the sequence elements, whereas abstract structure is defined by relationships between repeating sequence elements. Thus, sequences ABCBAC and DEFEDF have different surface structures but share a common abstract structure, 123213, and are therefore isomorphic. Our simulations of sequence learning performance in serial reaction time (SRT) tasks demonstrated that (1) an existing model of the primate fronto-striatal system is capable of learning surface structure but fails to learn abstract structure, which requires an additional capability, (2) surface and abstract structure can be learned independently by these independent processes, and (3) only abstract structure transfers to isomorphic sequences. We tested these predictions in human subjects. For a sequence with predictable surface and abstract structure, subjects in either explicit or implicit conditions learn the surface structure, but only explicit subjects learn and transfer the abstract structure. For sequences with only abstract structure, learning and transfer of this structure occurs only in the explicit group. These results are parallel to those from the simulations and support our dissociable process hypothesis. Based on the synthesis of the current simulation and empirical results with our previous neuropsychological findings, we propose a neuro-physiological basis for these dissociable processes: Surface structure can be learned by processes that operate under implicit conditions and rely on the fronto-striatal system, whereas learning abstract structure requires a more explicit activation of dissociable processes that rely on a distributed network that includes the left anterior cortex.

Computer Simulation↗

Prediction of secondary structural content of proteins from their amino acid composition alone. II. The paradox with secondary structural class.

The success rates reported for secondary structural class prediction with different methods are contradictory. On one side, the problem of recognizing the secondary structural class of a protein knowing only its amino acid composition appears completely solved by simply applying jury decision with an elliptically scaled distance function. Chou and coworkers repeatedly (see Crit. Rev. Biochem. Mol. Biol. 30:275-349, 1995) published prediction accuracies near 100%. On the other hand, traditional secondary structure prediction techniques achieve success rates of about 70% for the secondary structural state per residue and about 75% for structural class only with extensive input information (full sequence of the query protein, its amino acid composition and length, multiple alignments with homologous sequences). In this article, we resolve the paradox and consider (1) the question of the secondary structural class definition, (2) the role of the representativity of the test set of protein tertiary structure for the current state of the Protein Data Bank (PDB); and (3) we estimate the real impact of amino acid composition on secondary structural class. We formulate three objective criteria for a reasonable definition of secondary structural classes and show that only the criterion of Nakashima et al. (J. Biochem. 99:153-162, 1986) complies with all of them. Only this definition matches the distribution of secondary structural content in representative PDB subsets, whereas other criteria leave many proteins (up to 65% of all PDB entries) simply unassigned. We review critically specialized secondary-structural class prediction methods, especially those of Chou and coworkers, which claim almost 100% accuracy using only amino acid composition, and resolve the paradox that these prediction accuracies are better than those from secondary structure predictions from multiple alignments. We show (i) that these techniques rely on a preselection of test sets which removes irregular proteins and other proteins without any class assignment (about 35% of all PDB entries); and (ii) that even for preselected representative test sets, the success rate drops to 60% and lower for a 4-type classification (alpha, beta, alpha + beta, alpha/beta). The prediction accuracies fall to about 50% if the secondary structural class definition of Nakashima et al. is applied and only few irregular proteins are preselected and removed from automatically generated, representative subsets of the PDB. We have applied two new vector decomposition methods for secondary structural content prediction from amino acid composition alone, with and without consideration of amino acid compositional coupling in the learning set of tertiary structures respectively, to the problem of class prediction and achieve about 60% correct assignment among four classes (alpha, beta, mixed, irregular) as well as single sequence-based secondary structure prediction methods like GORIII and COMBI. Our results demonstrate that 60% correctness is the upper limit for a 4-type class prediction from amino acid composition alone for an unknown query protein and that consideration of compositional coupling does not improve the prediction success. The prediction program SSCP offering secondary structural class assignment for query compositions and sequences has been made available as a World Wide Web and E-mail service.

Amino Acids↗

The solution structure of eglin c based on measurements of many NOEs and coupling constants and its comparison with X-ray structures.

A high-precision solution structure of the elastase inhibitor eglin c was determined by NMR and distance geometry calculations. A large set of 947 nuclear Overhauser (NOE) distance constraints was identified, 417 of which were quantified from two-dimensional NOE spectra at short mixing times. In addition, a large number of homonuclear 1H-1H and heteronuclear 1H-15N vicinal coupling constants were used, and constraints on 42 chi 1 and 38 phi angles were obtained. Structure calculations were carried out using the distance geometry program DG-II. These calculations had a high convergence rate, in that 66 out of 75 calculations converged with maximum residual NOE violations ranging from 0.17 A to 0.47 A. The spread of the structures was characterized with average root mean square deviations ( ) between the structures and a mean structure. To calculate the unbiased toward any single structure, a new procedure was used for structure alignment. A canonical structure was calculated from the mean distances, and all structures were aligned relative to that. Furthermore, an angular order parameter S was defined and used to characterize the spread of structures in torsion angle space. To obtain an accurate estimate of the precision of the structure, the number of calculations was increased until the and the angular order parameters stabilized. This was achieved after approximately 40 calculations. The structure consists of a well-defined core whose backbone deviates from the canonical structure ca. 0.4 A, a disordered N-terminal heptapeptide whose backbone deviates by 0.8-12 A, and a proteinase-binding loop whose backbone deviates up to 3.0 A. Analysis of the angular order parameters and inspection of the structures indicates that a hinge-bending motion of the binding loop may occur in solution. Secondary structures were analyzed by comparison of dihedral angle patterns. The high precision of the structure allows one to identify subtle differences with four crystal structures of eglin c determined in complexes with proteinases.

Amino Acid Sequence↗

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence↗

[Analysis of structural motifs of proteins using sets of codes, describing local structures].

An amino acid sequence pattern conserved among a family of proteins is called motif. It is usually related to the specific function of the family. On the other hand, functions of proteins are achieved by their 3D structures. Specific local structures, called structural motifs, are considered related to their functions. However, searching for common structural motifs in different proteins is much more difficult than for common sequence motifs. We are attempting in this study to convert the information about the structural motifs into a set of one-dimensional digital strings, i.e., a set of codes, to compare them more easily by computer and to investigate their relationship to functions more quantitatively. By applying the Delaunay tessellation to a 3D structure of a protein, we can assign each local structure to a unique code that is defined so as to reflect its structural feature. Since a structural motif is defined as a set of the local structures in this paper, the structural motif is represented by a set of the codes. In order to examine the ability of the set of the codes to distinguish differences among the sets of local structures with a given PROSITE pattern that contain both true and false positives, we clustered them by introducing a similarity measure among the set of the codes. The obtained clustering shows a good agreement with other results by direct structural comparison methods such as a superposition method. The structural motifs in homologous proteins are also properly clustered according to their sources. These results suggest that the structural motifs can be well characterized by these sets of the codes, and that the method can be utilized in comparing structural motifs and relating them with function.

Amino Acid Motifs↗

Solution structure and dynamics of linked cell attachment modules of mouse fibronectin containing the RGD and synergy regions: comparison with the human fibronectin crystal structure.

We report the three-dimensional solution structure of the mouse fibronectin cell attachment domain consisting of the linked ninth and tenth type III modules, mFnFn3(9,10). Because the tenth module contains the RGD cell attachment sequence while the ninth contains the synergy region, mFnFn3(9,10) has the cell attachment activity of intact fibronectin. Essentially complete signal assignments and approximately 1800 distance and angle restraints were derived from multidimensional heteronuclear NMR spectra. These restraints were used with a hybrid distance geometry/simulated annealing protocol to generate an ensemble of 20 NMR structures having no distance or angle violations greater than 0.3 A or 3 degrees. Although the beta-sheet core domains of the individual modules are well-ordered structures, having backbone atom rmsd values from the mean structure of 0.51(+/-0.12) and 0.40(+/-0.07) A, respectively, the rmsd of the core atom coordinates increases to 3.63(+/-1.41) A when the core domains of both modules are used to align the coordinates. The latter result is a consequence of the fact that the relative orientation of the two modules is not highly constrained by the NMR restraints. Hence, while structures of the beta-sheet core domains of the NMR structures are very similar to the core domains of the crystal structure of hFnFn3(9,10), the ensemble of NMR structures suggests that the two modules form a less extended and more flexible structure than the fully extended rod-like crystal structure. The radius of gyration, Rg, of mFnFn3(9,10) derived from small-angle neutron scattering measurements, 20.5(+/-0.5) A, agrees with the average Rg calculated for the NMR structures, 20.4 A, and is ca 1 A less than the value of Rg calculated for the X-ray structure. The values of the rotational anisotropy, D ||/D perpendicular, derived from an analysis of 15N relaxation data, range from 1.7 to 2.1, and are significantly less than the anisotropy of 2.67 predicted by hydrodynamic modeling of the crystal coordinates. In contrast, hydrodynamic modeling of the NMR coordinates yields anisotropies in the range of 1.9 to 2.7 (average 2.4(+/-0.2)), with NMR structures bent by more than 20 degrees relative the crystal structure having calculated anisotropies in best agreement with experiment. In addition, the relaxation parameters indicate that several loops in mFnFn3(9,10), including the RGD loop, are flexible on the nanosecond to picosecond time-scale. Taken together, our results suggest that, in solution, the limited set of interactions between the mFnFn3(9,10) modules position the RGD and synergy regions to interact specifically with cell surface integrins, and at the same time permit sufficient flexibility that allows mFnFn3(9,10) to adjust for some variation in integrin structure or environment.

Amino Acid Sequence↗