Search PubMed⌕ Search

Biomedical subjects

A Godzik

Publications and source records attributed to A Godzik.

At least 55 records · Page 3Linked to original sources

Similarities and differences between nonhomologous proteins with similar folds: evaluation of threading strategies.

BACKGROUND: There are many pairs and groups of proteins with similar folds and interaction patterns, but whose sequence similarity is below the threshold of easily recognizable sequence homology. The existence of multiple sequence solutions for a given fold has inspired fold prediction methods in which structural information from one protein is used to estimate the energy of another, putatively similar, structure. RESULTS: A set of 68 pairs of proteins with similar folds and sequence identity in the 8-30% range is identified from the literature. for each pair, the energy of one protein, calculated using knowledge-based statistical potentials, is compared to the estimated energy, calculated with the same potentials but using the structural information (burial status and interaction pattern) of another protein with the same fold. Different energy estimates, corresponding to approximations used in various fold recognition algorithms, are calculated and compared to each other, as well as to the correct energy. It is shown that the local energy terms, based on burial and secondary structure preferences, can be reliably estimated with an accuracy close to 70%. At the same time, the two-body nonlocal energy loses over 60% of its value due to the repacking of the structure. Further approximations, such as the 'frozen approximation', can bring it to an essentially random value. CONCLUSIONS: Local energy terms could be used safely to improve fold recognition algorithms. To utilize pair interaction information, specially designed pair potentials and/or a self-consistent description of pair interactions is necessary.

Algorithms↗

Secondary structure prediction using segment similarity.

We present a secondary structure prediction method based on finding similarities between sequence segments from the target sequence and segments contained in the database of proteins with known structures. The similarity definition is optimized using a genetic algorithm and is based on a 21 x 40 similarity matrix, comparing a target sequence with the sequence and burial status of the proteins from the database. The three-state secondary structure prediction accuracy reaches 72.4% on a non homologous (maximum sequence identity <25%) data set derived from PDB and is reproduced on two independent testing sets, including the set of CASP2 prediction targets and a group of newly solved PDB structures. The prediction method was developed with simplicity and open architecture in mind, allowing for an easy extension to other types of predictions and to the analysis of the contributions to the local structure formation. For instance, the design of the prediction procedure allows us to trace back segments of the database that contributed to the prediction. It can be shown that those segments came from various structural classes and that even complete exclusion of related folds from the database does not result in a significant decrease in prediction accuracy.

Databases, Factual↗

Sequence-structure specificity--how does an inverse folding approach work?

The inverse folding approach is a powerful tool in protein structure prediction when the native state of a sequence adopts one of the known protein folds. This is because some proteins show strong sequence-structure specificity in inverse folding experiments that allow gaps and insertions in the sequence-structure alignment. In those cases when structures similar to their native folds are included in the structure database, the z-scores (which measure the sequence-structure specificity) of these folds are well separated from those of other alternative structures. In this paper, we seek to understand the origin of this sequence-structure specificity and to identify how the specificity arises on passing from a short peptide chain to the entire protein sequence. To accomplish this objective, a simplified version of inverse folding, gapless inverse folding, is performed using sequence fragments of different sizes from 53 proteins. The results indicate that usually a significant portion of the entire protein sequence is necessary to show sequence-structure specificity, but there are regions in the sequence that begin to show this specificity at relatively short fragment size (15-20 residues). An island picture, in which the regions in the sequence that recognize their own native structure grow from some seed fragments, is observed as the fragment size increases. Usually, more similar structures to the native states are found in the top-scoring structural fragments in these high-specificity regions.

Amino Acid Sequence↗

A method for the prediction of surface "U"-turns and transglobular connections in small proteins.

A simple method for predicting the location of surface loops/turns that change the overall direction of the chain that is, "U" turns, and assigning the dominant secondary structure of the intervening transglobular blocks in small, single-domain globular proteins has been developed. Since the emphasis of the method is on the prediction of the major topological elements that comprise the global structure of the protein rather than on a detailed local secondary structure description, this approach is complementary to standard secondary structure prediction schemes. Consequently, it may be useful in the early stages of tertiary structure prediction when establishment of the structural class and possible folding topologies is of interest. Application to a set of small proteins of known structure indicates a high level of accuracy. The prediction of the approximate location of the surface turns/loops that are responsible for the change in overall chain direction is correct in more than 95% of the cases. The accuracy for the dominant secondary structure assignment for the linear blocks between such surface turns/loops is in the range of 82%.

Algorithms↗

Multiple model approach--dealing with alignment ambiguities in protein modeling.

Sequence alignments for distantly homologous proteins are often ambiguous, which creates a weak link in structure prediction by homology. We address this problem by using several plausible alignments in a modeling procedure, obtaining many models of the target. All are subsequently evaluated by a threading algorithm. It is shown that this approach can identify best alignments and produce reasonable models, whose quality is now limited only by the extent of the structural similarity between the known and predicted protein. Using a similar approach structure prediction for the oxidized dimer of S100A1 protein, for which the structure is not known, is presented.

Amino Acid Sequence↗

Structural diversity in a family of homologous proteins.

An interesting example of a structurally diverse group of sequentially homologous proteins is analyzed at the level of molecular interactions. In this family, the EF-hand calcium-binding proteins, there are examples of at least three distinct mutual positions of the N and C-terminal domains, despite significant sequence homology between all members of this family. Why does a particular protein choose one arrangement over another? To answer this question, detailed models of all proteins in their native structures as well as all alternative sequence/structure combinations are built by comparative modeling. By studying and comparing interactions stabilizing native structures and destabilizing alternative conformations, it is possible to gain insight into how such conformational diversity is achieved. It is shown that some mechanisms used to achieve it are: correlated mutations on the surface of two units and the presence of additional domains/chain fragments stabilizing desired topologies. The implications of these findings, both for structure predictions for other members of this family as well as the general problem of quaternary structure formation, are discussed.

Amino Acid Sequence↗

The structural alignment between two proteins: is there a unique answer?

Structurally similar but sequentially unrelated proteins have been discovered and rediscovered by many researchers, using a variety of structure comparison tools. For several pairs of such proteins, existing structural alignments obtained from the literature, as well as alignments prepared using several different similarity criteria, are compared with each other. It is shown that, in general, they differ from each other, with differences increasing with diminishing sequence similarity. Differences are particularly strong between alignments optimizing global similarity measures, such as RMS deviation between C alpha atoms, and alignments focusing on more local features, such as packing or interaction pattern similarity. Simply speaking, by putting emphasis on different aspects of structure, different structural alignments show the unquestionable similarity in a different way. With differences between various alignments extending to a point where they can differ at all positions, analysis of structural similarities leads to contradictory results reported by groups using different alignment techniques. The problem of uniqueness and stability of structural alignments is further studied with the help of visualization of the suboptimal alignments. It is shown that alignments are often degenerate and whole families of alignments can be generated with almost the same score as the "optimal alignment." However, for some similarity criteria, specially those based on side-chain positions, rather than C alpha positions, alignments in some areas of the protein are unique. This opens the question of how and if the structural alignments can be used as "standards of truth" for protein comparison.

Amino Acid Sequence↗

An algorithm for prediction of structural elements in small proteins.

A method for predicting the location of surface loops/turns and assigning the intervening secondary structure of the transglobular linkers in small, single domain globular proteins has been developed. Application to a set of 10 proteins of known structure indicates a high level of accuracy. The secondary structure assignment in the center of transglobular connections is correct in more than 85% of the cases. A similar error rate is found for loops. Since more global information about the fold is provided, it is complementary to standard secondary structure prediction approaches. Consequently, it may be useful in early stages of tertiary structure prediction when establishment of the structural class and possible folding topologies is of interest.

Algorithms↗

Are proteins ideal mixtures of amino acids? Analysis of energy parameter sets.

Various existing derivations of the effective potentials of mean force for the two-body interactions between amino acid side chains in proteins are reviewed and compared to each other. The differences between different parameter sets can be traced to the reference state used to define the zero of energy. Depending on the reference state, the transfer free energy or other pseudo-one-body contributions can be present to various extents in two-body parameter sets. It is, however, possible to compare various derivations directly by concentrating on the "excess" energy-a term that describes the difference between a real protein and an ideal solution of amino acids. Furthermore, the number of protein structures available for analysis allows one to check the consistency of the derivation and the errors by comparing parameters derived from various subsets of the whole database. It is shown that pair interaction preferences are very consistent throughout the database. Independently derived parameter sets have correlation coefficients on the order of 0.8, with the mean difference between equivalent entries of 0.1 kT. Also, the low-quality (low resolution, little or no refinement) structures show similar regularities. There are, however, large differences between interaction parameters derived on the basis of crystallographic structures and structures obtained by the NMR refinement. The origin of the latter difference is not yet understood.

Amino Acid Sequence↗

In search of the ideal protein sequence.

The inverse of a folding problem is to find the ideal sequence that folds into a particular protein structure. This problem has been addressed using the topology fingerprint-based threading algorithm, capable of calculating a score (energy) of an arbitrary sequence-structure pair. At first, the search is conducted by unconstrained minimization of the energy in sequence space. It is shown that using energy as the only design criterion leads to spurious solutions with incorrect amino acid composition. The problem lies in the general features of the protein energy surface as a function of both structure and sequence. The proposed solution is to design the sequence by maximizing the difference between its energy in the desired structure and in other known protein structures. Depending on the size of the database of structures 'to avoid', sequences bearing significant similarity to the native sequence of the target protein are obtained using this procedure.

Algorithms↗

Flexible algorithm for direct multiple alignment of protein structures and sequences.

The recently described equivalence between the alignment of two proteins and a conformation of a lattice chain on a two-dimensional square lattice is extended to multiple alignments. The search for the optimal multiple alignment between several proteins, which is equivalent to finding the energy minimum in the conformational space of a multi-dimensional lattice chain, is studied by the Monte Carlo approach. This method, while not deterministic, and for two-dimensional problems slower than dynamic programming, can accept arbitrary scoring functions, including non-local ones, and its speed decreases slowly with increasing number of dimensions. For the local scoring functions, the MC algorithm can also reproduce known exact solutions for the direct multiple alignments. As illustrated by examples, both for structure- and sequence-based alignments, direct multi-dimensional alignments are able to capture weak similarities between divergent families much better than ones built from pairwise alignments by a hierarchical approach.

Algorithms↗

A method for predicting protein structure from sequence.

BACKGROUND: The ability to predict the native conformation of a globular protein from its amino-acid sequence is an important unsolved problem of molecular biology. We have previously reported a method in which reduced representations of proteins are folded on a lattice by Monte Carlo simulation, using statistically-derived potentials. When applied to sequences designed to fold into four-helix bundles, this method generated predicted conformations closely resembling the real ones. RESULTS: We now report a hierarchical approach to protein-structure prediction, in which two cycles of the above-mentioned lattice method (the second on a finer lattice) are followed by a full-atom molecular dynamics simulation. The end product of the simulations is thus a full-atom representation of the predicted structure. The application of this procedure to the 60 residue, B domain of staphylococcal protein A predicts a three-helix bundle with a backbone root mean square (rms) deviation of 2.25-3 A from the experimentally determined structure. Further application to a designed, 120 residue monomeric protein, mROP, based on the dimeric ROP protein of Escherichia coli, predicts a left turning, four-helix bundle native state. Although the ultimate assessment of the quality of this prediction awaits the experimental determination of the mROP structure, a comparison of this structure with the set of equivalent residues in the ROP dime- crystal structure indicates that they have a rms deviation of approximately 3.6-4.2 A. CONCLUSION: Thus, for a set of helical proteins that have simple native topologies, the native folds of the proteins can be predicted with reasonable accuracy from their sequences alone. Our approach suggest a direction for future work addressing the protein-folding problem.

Journal Article↗

De novo and inverse folding predictions of protein structure and dynamics.

In the last two years, the use of simplified models has facilitated major progress in the globular protein folding problem, viz., the prediction of the three-dimensional (3D) structure of a globular protein from its amino acid sequence. A number of groups have addressed the inverse folding problem where one examines the compatibility of a given sequence with a given (and already determined) structure. A comparison of extant inverse protein-folding algorithms is presented, and methodologies for identifying sequences likely to adopt identical folding topologies, even when they lack sequence homology, are described. Extension to produce structural templates or fingerprints from idealized structures is discussed, and for eight-membered beta-barrel proteins, it is shown that idealized fingerprints constructed from simple topology diagrams can correctly identify sequences having the appropriate topology. Furthermore, this inverse folding algorithm is generalized to predict elements of supersecondary structure including beta-hairpins, helical hairpins and alpha/beta/alpha fragments. Then, we describe a very high coordination number lattice model that can predict the 3D structure of a number of globular proteins de novo; i.e. using just the amino acid sequence. Applications to sequences designed by DeGrado and co-workers [Biophys. J., 61 (1992) A265] predict folding intermediates, native states and relative stabilities in accord with experiment. The methodology has also been applied to the four-helix bundle designed by Richardson and co-workers [Science, 249 (1990) 884] and a redesigned monomeric version of a naturally occurring four-helix dimer, rop. Based on comparison to the rop dimer, the simulations predict conformations with rms values of 3-4 A from native. Furthermore, the de novo algorithms can assess the stability of the folds predicted from the inverse algorithm, while the inverse folding algorithms can assess the quality of the de novo models. Thus, the synergism of the de novo and inverse folding algorithm approaches provides a set of complementary tools that will facilitate further progress on the protein-folding problem.

Algorithms↗

Regularities in interaction patterns of globular proteins.

The description of protein structure in the language of side chain contact maps is shown to offer many advantages over more traditional approaches. Because it focuses on side chain interactions, it aids in the discovery, study and classification of similarities between interactions defining particular protein folds and offers new insights into the rules of protein structure. For example, there is a small number of characteristic patterns of interactions between protein supersecondary structural fragments, which can be seen in various non-related proteins. Furthermore, the overlap of the side chain contact maps of two proteins provides a new measure of protein structure similarity. As shown in several examples, alignments based on contact map overlaps are a powerful alternative to other structure-based alignments.

Computer Simulation↗

Sequence-structure matching in globular proteins: application to supersecondary and tertiary structure determination.

A methodology designed to address the inverse globular protein-folding problem (the identification of which sequences are compatible with a given three-dimensional structure) is described. By using a library of protein finger-prints, defined by the side chain interaction pattern, it is possible to match each structure to its own sequence in an exhaustive data base search. It is shown that this is a permissive requirement for the validation of the methodology. To pass the more rigorous test of identifying proteins that are not close sequence homologs, but that have similar structure, the method has been extended to include insertions and deletions in the sequence, which is compared to the fingerprint. This allows for the identification of sequences having little or no sequence homology to the fingerprint. Examples include plastocyanin/azurin/pseudoazurin, the globin family, different families of proteases and cytochromes, including cytochromes c' and b-562, actinidin/papain, and lysozyme/alpha-lactalbumin. Turning to supersecondary structure prediction, we find that alpha/beta/alpha fragments possess sufficient specificity to identify their own and related sequences. By threading a beta-hairpin through a sequence, it is possible to predict the location of such hairpins and turns with remarkable fidelity. Thus, the method greatly extends existing techniques for the prediction of both global structural homology and local supersecondary structure.

Amino Acid Sequence↗

Topology fingerprint approach to the inverse protein folding problem.

We describe the most general solution to date of the problem of matching globular protein sequences to the appropriate three-dimensional structures. The screening template, against which sequences are tested, is provided by a protein "structural fingerprint" library based on the contact map and the buried/exposed pattern of residues. Then, a lattice Monte Carlo algorithm validates or dismisses the stability of the proposed fold. Examples of known structural similarities between proteins having weakly or unrelated sequences such as the globins and phycocyanins, the eight-member alpha/beta fold of triose phosphate isomerase and even a close structural equivalence between azurin and immunoglobulins are found.

Algorithms↗