Search PubMed⌕ Search

Biomedical subjects

A Kolinski

Publications and source records attributed to A Kolinski.

At least 37 records · Page 2Linked to original sources

Tertiary structure prediction of the KIX domain of CBP using Monte Carlo simulations driven by restraints derived from multiple sequence alignments.

Using a recently developed protein folding algorithm, a prediction of the tertiary structure of the KIX domain of the CREB binding protein is described. The method incorporates predicted secondary and tertiary restraints derived from multiple sequence alignments in a reduced protein model whose conformational space is explored by Monte Carlo dynamics. Secondary structure restraints are provided by the PHD secondary structure prediction algorithm that was modified for the presence of predicted U-turns, i.e., regions where the chain reverses global direction. Tertiary restraints are obtained via a two-step process: First, seed side-chain contacts are identified from a correlated mutation analysis, and then, a threading-based algorithm expands the number of these seed contacts. Blind predictions indicate that the KIX domain is a putative three-helix bundle, although the chirality of the bundle could not be uniquely determined. The expected root-mean-square deviation for the correct chirality of the KIX domain is between 5.0 and 6.2 A. This is to be compared with the estimate of 12.9 A that would be expected by a random prediction, using the model of F. Cohen and M. Sternberg (J. Mol. Biol. 138:321-333, 1980).

Algorithms↗

Nativelike topology assembly of small proteins using predicted restraints in Monte Carlo folding simulations.

By incorporating predicted secondary and tertiary restraints derived from multiple sequence alignments into ab initio folding simulations, it has been possible to assemble native-like tertiary structures for a test set of 19 nonhomologous proteins ranging from 29 to 100 residues in length and representing all secondary structural classes. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Multiple sequence alignments also provide predicted tertiary restraints via a two-step process: First, seed side chain contacts are selected from a correlated mutation analysis, and then an inverse folding algorithm expands these seed contacts. The predicted secondary and tertiary restraints are incorporated into a lattice-based, reduced protein model for structure assembly and refinement. The resulting native-like topologies exhibit a coordinate root-mean-square deviation from native for the whole chain between 3.1 and 6.7 A, with values ranging from 2.6 to 4.1 A over approximately 80% of the structure. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

Computer simulations of de novo designed helical proteins.

In the context of reduced protein models, Monte Carlo simulations of three de novo designed helical proteins (four-member helical bundle) were performed. At low temperatures, for all proteins under consideration, protein-like folds having different topologies were obtained from random starting conformations. These simulations are consistent with experimental evidence indicating that these de novo designed proteins have the features of a molten globule state. The results of Monte Carlo simulations suggest that these molecules adopt four-helix bundle topologies. They also give insight into the possible mechanism of folding and association, which occurs in these simulations by on-site assembly of the helices. The low-temperature conformations of all three sequences have the features of a molten globule state.

Amino Acid Sequence↗

Reduced protein models and their application to the protein folding problem.

One of the most important unsolved problems of computational biology is prediction of the three-dimensional structure of a protein from its amino acid sequence. In practice, the solution to the protein folding problem demands that two interrelated problems be simultaneously addressed. Potentials that recognize the native state from the myriad of misfolded conformations are required, and the multiple minima conformational search problem must be solved. A means of partly surmounting both problems is to use reduced protein models and knowledge-based potentials. Such models have been employed to elucidate a number of general features of protein folding, including the nature of the energy landscape, the factors responsible for the uniqueness of the native state and the origin of the two-state thermodynamic behavior of globular proteins. Reduced models have also been used to predict protein tertiary and quaternary structure. When combined with a limited amount of experimental information about secondary and tertiary structure, molecules of substantial complexity can be assembled. If predicted secondary structure and tertiary restraints are employed, low resolution models of single domain proteins can be successfully predicted. Thus, simplified protein models have played an important role in furthering the understanding of the physical properties of proteins.

CREB-Binding Protein↗

Combined multiple sequence reduced protein model approach to predict the tertiary structure of small proteins.

By incorporating predicted secondary and tertiary restraints into ab initio folding simulations, low resolution tertiary structures of a test set of 20 nonhomologous proteins have been predicted. These proteins, which represent all secondary structural classes, contain from 37 to 100 residues. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Predicted tertiary restraints are obtained from multiple sequence alignments via a two-step process: First, "seed" side chain contacts are identified from a correlated mutation analysis, and then, the seed contacts are "expanded" by an inverse folding algorithm. These predicted restraints are then incorporated into a lattice based, reduced protein model. Depending upon fold complexity, the resulting nativelike topologies exhibit a coordinate root-mean-square deviation, cRMSD, from native between 3.1 and 6.7 A. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

MONSSTER: a method for folding globular proteins with a small number of distance restraints.

The MONSSTER (MOdeling of New Structures from Secondary and TEritary Restraints) method for folding of proteins using a small number of long-distance restraints (which can be up to seven times less than the total number of residues) and some knowledge of the secondary structure of regular fragments is described. The method employs a high-coordination lattice representation of the protein chain that incorporates a variety of potentials designed to produce protein-like behaviour. These include statistical preferences for secondary structure, side-chain burial interactions, and a hydrogen-bond potential. Using this algorithm, several globular proteins (1ctf, 2gbl, 2trx, 3fxn, 1mba, 1pcy and 6pti) have been folded to moderate-resolution, native-like compact states. For example, the 68 residue 1ctf molecule having ten loosely defined, long-range restraints was reproducibly obtained with a C alpha-backbone root-mean-square deviation (RMSD) from native of about 4. A. Flavodoxin with 35 restraints has been folded to structures whose average RMSD is 4.28 A. Furthermore, using just 20 restraints, myoglobin, which is a 146 residue helical protein, has been folded to structures whose average RMSD from native is 5.65 A. Plastocyanin with 25 long-range restraints adopts conformations whose average RMSD is 5.44 A. Possible applications of the proposed approach to the refinement of structures from NMR data, homology model-building and the determination of tertiary structure when the secondary structure and a small number of restraints are predicted are briefly discussed.

Algorithms↗

Improved method for prediction of protein backbone U-turn positions and major secondary structural elements between U-turns.

A new and more accurate method has been developed for predicting the backbone U-turn positions (where the chain reverses global direction) and the dominant secondary structure elements between U-turns in globular proteins. The current approach uses sequence-specific secondary structure propensities and multiple sequence information. The latter plays an important role in the enhanced success of this approach. Application to two sets (total 108) of small to medium-sized, single-domain proteins indicates that approximately 94% of the U-turn locations are correctly predicted within three residues, as are 88% of dominant secondary structure elements. These results are significantly better than our previous method (Kolinski et al., Proteins 27:290-308, 1997). The current study strongly suggests that the U-turn locations are primarily determined by local interactions. Furthermore, both global length constraints and local interactions contribute significantly to the determination of the secondary structure types between U-turns. Accurate U-turn predictions are crucial for accurate secondary structure predictions in the current method. Protein structure modeling, tertiary structure predictions, and possibly, fold recognition should benefit from the predicted structural data provided by this new method.

Amino Acid Sequence↗

Derivation and testing of pair potentials for protein folding. When is the quasichemical approximation correct?

Many existing derivations of knowledge-based statistical pair potentials invoke the quasichemical approximation to estimate the expected side-chain contact frequency if there were no amino acid pair-specific interactions. At first glance, the quasichemical approximation that treats the residues in a protein as being disconnected and expresses the side-chain contact probability as being proportional to the product of the mole fractions of the pair of residues would appear to be rather severe. To investigate the validity of this approximation, we introduce two new reference states in which no specific pair interactions between amino acids are allowed, but in which the connectivity of the protein chain is retained. The first estimates the expected number of side-chain contracts by treating the protein as a Gaussian random coil polymer. The second, more realistic reference state includes the effects of chain connectivity, secondary structure, and chain compactness by estimating the expected side-chain contrast probability by placing the sequence of interest in each member of a library of structures of comparable compactness to the native conformation. The side-chain contact maps are not allowed to readjust to the sequence of interest, i.e., the side chains cannot repack. This situation would hold rigorously if all amino acids were the same size. Both reference states effectively permit the factorization of the side-chain contact probability into sequence-dependent and structure-dependent terms. Then, because the sequence distribution of amino acids in proteins is random, the quasichemical approximation to each of these reference states is shown to be excellent. Thus, the range of validity of the quasichemical approximation is determined by the magnitude of the side-chain repacking term, which is, at present, unknown. Finally, the performance of these two sets of pair interaction potentials as well as side-chain contact fraction-based interaction scales is assessed by inverse folding tests both without and with allowing for gaps.

Models, Chemical↗

A method for the prediction of surface "U"-turns and transglobular connections in small proteins.

A simple method for predicting the location of surface loops/turns that change the overall direction of the chain that is, "U" turns, and assigning the dominant secondary structure of the intervening transglobular blocks in small, single-domain globular proteins has been developed. Since the emphasis of the method is on the prediction of the major topological elements that comprise the global structure of the protein rather than on a detailed local secondary structure description, this approach is complementary to standard secondary structure prediction schemes. Consequently, it may be useful in the early stages of tertiary structure prediction when establishment of the structural class and possible folding topologies is of interest. Application to a set of small proteins of known structure indicates a high level of accuracy. The prediction of the approximate location of the surface turns/loops that are responsible for the change in overall chain direction is correct in more than 95% of the cases. The accuracy for the dominant secondary structure assignment for the linear blocks between such surface turns/loops is in the range of 82%.

Algorithms↗

Method for low resolution prediction of small protein tertiary structure.

A new method for the de novo prediction of protein structures at low resolution has been developed. Starting from a multiple sequence alignment, protein secondary structure is predicted, and only those topological elements with high reliability are selected. Then, the multiple sequence alignment and the secondary structure prediction are combined to predict side chain contacts. Such contact map prediction is carried out in two stages. First, an analysis of correlated mutations is carried out to identify pairs of topological elements of secondary structure which are in contact. Then, inverse folding is used to select compatible fragments in contact, thereby enriching the number and identity of predicted side chain contacts. The final outcome of the procedure is a set of noisy secondary and tertiary restraints. These are used as a restrained potential in a Monte Carlo simulation of simplified protein models driven by statistical potentials. Low energy structures are then searched for by using simulated annealing techniques. Implementation of the restraints is carried out so as to take into account of their low resolution. Using this procedure, it has been possible to predict de novo the structure of three very different protein topologies: an alpha/beta protein, the bovine pancreatic trypsin inhibitor (6pti), an alpha-helical protein, calbindin (3icb), and an all beta- protein, the SH3 domain of spectrin (1shg). In all cases, low resolution folds have been obtained with a root mean square deviation (RMSD) of 4.5-5.5 A with respect to the native structure. Some misfolded topologies appear in the simulations, but it is possible to select the native one on energetic grounds. Thus, it is demonstrated that the methodology is general for all protein motifs. Work is in progress in order to test the methodology on a larger set of protein structures.

Amino Acid Sequence↗

Method for predicting the state of association of discretized protein models. Application to leucine zippers.

A method that employs a transfer matrix treatment combined with Monte Carlo sampling has been used to calculate the configurational free energies of folded and unfolded states of lattice models of proteins. The method is successfully applied to study the monomer-dimer equilibria in various coiled coils. For the short coiled coils, GCN4 leucine zipper, and its fragments, Fos and Jun, very good agreement is found with experiment. Experimentally, some subdomains of the GCN4 leucine zipper form stable dimeric structures, suggesting the regions of differential stability in the parent structure. Our calculations suggest that the stabilities of the subdomains are in general different from the values expected simply from the stability of the corresponding fragment in the wild type molecule. Furthermore, parts of the fragments structurally rearrange in some regions with respect to their corresponding wild type positions. Our results suggest for an Asn in the dimerization interface at least a pair of hydrophobic interacting helical turns at each side is required to stabilize the stable coiled coil. Finally, the specificity of heterodimer formation in the Fos-Jun system comes from the relative instability of Fos homodimers, resulting from unfavorable intra- and interhelical interactions in the interfacial coiled coil region.

Amino Acid Sequence↗

Folding simulations and computer redesign of protein A three-helix bundle motifs.

In solution, the B domain of protein A from Staphylococcus aureus (B domain) possesses a three-helix bundle structure. This simple motif has been previously reproduced by Kolinski and Skolnick (Proteins 18: 353-366, 1994) using a reduced representation lattice model of proteins with a statistical interaction scheme. In this paper, an improved version of the potential has been used, and the robustness of this result has been tested by folding from the random state a set of three-helix bundle proteins that are highly homologous to the B domain of protein A. Furthermore, an attempt to redesign the B domain native structure to its topological mirror image fold has been made by multiple mutations of the hydrophobic core and the turn region between helices I and II. A sieve method for scanning a large set of mutations to search for this desired property has been proposed. It has been shown that mutations of native B domain hydrophobic core do not introduce significant changes in the protein motif. Mutations in the turn region were also very conservative; nevertheless, a few mutants acquired the desired topological mirror image motif. A set of all atom models of the most probable mutant was reconstructed from the reduced models and refined using a molecular dynamics algorithm in the presence of water. The packing of all atom structures obtained corroborates the lattice model results. We conclude that the change in the handedness of the turn induced by the mutations, augmented by the repacking of hydrophobic core and the additional burial of the second helix N-cap side chain, are responsible for the predicted preferential adoption of the mirror image structure.

Computer Simulation↗

On the origin of the cooperativity of protein folding: implications from model simulations.

There is considerable experimental evidence that the cooperativity of protein folding resides in the transition from the molten globule to the native state. The objective of this study is to examine whether simplified models can reproduce this cooperativity and if so, to identify its origin. In particular, the thermodynamics of the conformational transition of a previously designed sequence (A. Kolinski, W. Galazka, and J. Skolnick, J. Chem. Phys. 103: 10286-10297, 1995), which adopts a very stable Greek-key beta-barrel fold has been investigated using the entropy Monte Carlo sampling (ESMC) technique of Hao and Scheraga (M.-H. Hao and H.A. Scheraga, J. Phys. Chem. 98: 9882-9883, 1994). Here, in addition to the original potential, which includes one body and pair interactions between side chains, the force field has been supplemented by two types of multi-body potentials describing side chain interactions. These potentials facilitate the protein-like pattern of side chain packing and consequently increase the cooperativity of the folding process. Those models that include an explicit cooperative side chain packing term exhibit a well-defined all-or-none transition from a denatured, random coil state to a high-density, well-defined, nativelike low-energy state. By contrast, models lacking such a term exhibit a conformational transition that is essentially continuous. Finally, an examination of the conformations at the free-energy barrier between the native and denatured states reveals that they contain a substantial amount of native-state secondary structure, about 50% of the native contacts, and have an average root mean square radius of gyration that is about 15% larger than native.

Amino Acids↗

Does a backwardly read protein sequence have a unique native state?

Amino acid sequences of native proteins are generally not palindromic. Nevertheless, the protein molecule obtained as a result of reading the sequence backwards, i.e. a retro-protein, obviously has the same amino acid composition and the same hydrophobicity profile as the native sequence. The important questions which arise in the context of retro-proteins are: does a retro-protein fold to a well defined native-like structure as natural proteins do and, if the answer is positive, does a retro-protein fold to a structure similar to the native conformation of the original protein? In this work, the fold of retro-protein A, originated from the retro-sequence of the B domain of Staphylococcal protein A, was studied. As a result of lattice model simulations, it is conjectured that the retro-protein A also forms a three-helix bundle structure in solution. It is also predicted that the topology of the retro-protein A three-helix bundle is that of the native protein A, rather than that corresponding to the mirror image of native protein A. Secondary structure elements in the retro-protein do not exactly match their counterparts in the original protein structure; however, the amino acid side chain contract pattern of the hydrophobic core is partly conserved.

Amino Acid Sequence↗

An algorithm for prediction of structural elements in small proteins.

A method for predicting the location of surface loops/turns and assigning the intervening secondary structure of the transglobular linkers in small, single domain globular proteins has been developed. Application to a set of 10 proteins of known structure indicates a high level of accuracy. The secondary structure assignment in the center of transglobular connections is correct in more than 85% of the cases. A similar error rate is found for loops. Since more global information about the fold is provided, it is complementary to standard secondary structure prediction approaches. Consequently, it may be useful in early stages of tertiary structure prediction when establishment of the structural class and possible folding topologies is of interest.

Algorithms↗

Prediction of the quaternary structure of coiled coils: GCN4 leucine zipper and its mutants.

A methodology for predicting coiled coil quaternary structure and for the dissection of the interactions responsible for the global fold is described. Application is made to the equilibrium between different oligomeric species of the wild type GCN4 leucine zipper and seven of its mutants that were studied by Harbury et al. Over the entire experimental concentration range, agreement with experiment is found in five cases, while in two other cases, agreement is found over a portion of the concentration range. These simulations suggest that the degree of chain association is determined by the balance between specific side chain packing preferences and the entropy reduction associated with side chain burial in higher order multimers.

Amino Acid Substitution↗

Prediction of quaternary structure of coiled coils. Application to mutants of the GCN4 leucine zipper.

Using a simplified protein model, the equilibrium between different oligomeric species of the wild-type GCN4 leucine zipper and seven of its mutants have been predicted. Over the entire experimental concentration range, agreement with experiment is found in five cases, while in two cases agreement is found over a portion of the concentration range. These studies demonstrate a methodology for predicting coiled coil quaternary structure and allow for the dissection of the interactions responsible for the global fold. In agreement with the conclusion of Harbury et al., the results of the simulations indicate that the pattern of hydrophobic and hydrophilic residues alone is insufficient to define a protein's three-dimensional structure. In addition, these simulations indicate that the degree of chain association is determined by the balance between specific side-chain packing preferences and the entropy reduction associated with side-chain burial in higher-order multimers.

Computer Simulation↗

Neural network system for the evaluation of side-chain packing in protein structures.

An artificial neural network system is used for pattern recognition in protein side-chain-side-chain contact maps. A back-propagation network was trained on a set of patterns which are popular in side-chain contact maps of protein structures. Several neural network architectures and different training parameters were tested to decide on the best combination for the neural network. The resulting network can distinguish between original (from protein structures) and randomized patterns with an accuracy of 84.5% and a Matthews' coefficient of 0.72 for the testing set. Applications of this system for protein structure evaluation and refinement are also proposed. Examples include structures obtained after the application of molecular dynamics to crystal structures, structures obtained from X-ray crystallography at various stages of refinement, structures obtained from a de novo folding algorithm and deliberately misfolded structures.

Amino Acid Sequence↗