[POLAROGRAPHY OF TERTIARY PROTEIN STRUCTURES].
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We present a global optimization strategy that incorporates predicted restraints in both a local optimization context and as directives for global optimization approaches, to predict protein tertiary structure for alpha-helical proteins. Specifically, neural networks are used to predict the secondary structure of a protein, restraints are defined as manifestations of the network with a predicted secondary structure and the secondary structure is formed using local minimizations on a protein energy surface, in the presence of the restraints. Those residues predicted to be coil, by the network, define a conformational sub-space that is subject to optimization using a global approach known as stochastic perturbation that has been found to be effective for Lennard-Jones clusters and homo-polypeptides. Our energy surface is an all-atom 'gas phase' molecular mechanics force field, that is combined with a new solvation energy function that penalizes hydrophobic group exposure. This energy function gives the crystal structure of four different alpha-helical proteins as the lowest energy structure relative to other conformations, with correct secondary structure but incorrect tertiary structure. We demonstrate this global optimization strategy by determining the tertiary structure of the A-chain of the alpha-helical protein, uteroglobin and of a four-helix bundle, DNA binding protein.
We report a new method for predicting protein tertiary structure from sequence and secondary structure information. The predictions result from global optimization of a potential energy function, including van der Waals, hydrophobic, and excluded volume terms. The optimization algorithm, which is based on the alphaBB method developed by Floudas and coworkers (Costas and Floudas, J Chem Phys 1994;100:1247-1261), uses a reduced model of the protein and is implemented in both distance and dihedral angle space, enabling a side-by-side comparison of methodologies. For a set of eight small proteins, representing the three basic types--all alpha, all beta, and mixed alpha/beta--the algorithm locates low-energy native-like structures (less than 6A root mean square deviation from the native coordinates) starting from an unfolded state. Serial and parallel implementations of this methodology are discussed.
This paper discusses the benefit of mapping paired cysteine mutation patterns as a guide to identifying the positions of protein disulfide bonds. This information can facilitate the computer modeling of protein tertiary structure. First, a simple, paired natural-cysteine-mutation map is presented that identifies the positions of putative disulfide bonds in protein families. The method is based on the observation that if, during the process of evolution, a disulfide-bonded cysteine residue is not conserved, then it is likely that its counterpart will also be mutated. For each target protein, protein databases were searched for the primary amino acid sequences of all known members of distinct protein families. Primary sequence alignment was carried out using PileUp algorithms in the GCG package. To search for correlated mutations, we listed only the positions where cysteine residues were highly conserved and emphasized the mutated residues. In proteins of known three-dimensional structure, a striking pattern of paired cysteine mutations correlated with the positions of known disulfide bridges. For proteins of unknown architecture, the mutation maps showed several positions where disulfide bridging might occur.
From the most recent Brookhaven Protein Co-ordinate Databank, 229 sequence-identical pentapeptide pairs, each found in two unrelated protein structures, were collected; 9115 such pairs differing in only one residue were also gathered. For both samples the main-chain fold was conserved about 20% of the time, despite the different atomic environments presented by the unrelated protein architectures. An analysis of the substituted residues as well as the composition of the sequence-similar pentapeptides allowed several suggestions regarding protein folding mechanisms. An examination of the most frequently observed residue substitutions and their correlation with structural changes in the oligopeptide pairs yields a possible guide for site-directed mutagenesis experiments, especially when no tertiary structural information is at hand.
Although freeze-induced perturbations of the protein native fold are common, the underlying mechanism is poorly understood owing to the difficulty of monitoring their structure in ice. In this report we propose that binding of the fluorescence probe 1-anilino-8-naphthalene sulfonate (ANS) to proteins in ice can provide a useful monitor of ice-induced strains on the native fold. Experiments conducted with copper-free azurin from Pseudomonas aeruginosa, as a model protein system, demonstrate that in frozen solutions the fluorescence of ANS is enhanced several fold and becomes blue shifted relative free ANS. From the enhancement factor it is estimated that, at -13 degrees C, on average at least 1.6 ANS molecules become immobilized within hydrophobic sites of apo-azurin, sites that are destroyed when the structure is largely unfolded by guanidinium hydrochloride. The extent of ANS binding is influenced by temperature of ice as well as by conditions that affect the stability of the globular structure. Lowering the temperature from -4 degrees C to -18 degrees C leads to an apparent increase in the number of binding sites, an indication that low temperature and /or a reduced amount of liquid water augment the strain on the protein tertiary structure. It is significant that ANS binding is practically abolished when the native fold is stabilized upon formation of the Cd(2+) complex or on addition of glycerol to the solution but is further enhanced in the presence of NaSCN, a known destabilizing agent. The results of the present study suggest that the ANS binding method may find practical utility in testing the effectiveness of various additives employed in protein formulations as well as to devise safer freeze-drying protocols of pharmaceutical proteins.
Explore the source record for details and available documents.
A new method for the de novo prediction of protein structures at low resolution has been developed. Starting from a multiple sequence alignment, protein secondary structure is predicted, and only those topological elements with high reliability are selected. Then, the multiple sequence alignment and the secondary structure prediction are combined to predict side chain contacts. Such contact map prediction is carried out in two stages. First, an analysis of correlated mutations is carried out to identify pairs of topological elements of secondary structure which are in contact. Then, inverse folding is used to select compatible fragments in contact, thereby enriching the number and identity of predicted side chain contacts. The final outcome of the procedure is a set of noisy secondary and tertiary restraints. These are used as a restrained potential in a Monte Carlo simulation of simplified protein models driven by statistical potentials. Low energy structures are then searched for by using simulated annealing techniques. Implementation of the restraints is carried out so as to take into account of their low resolution. Using this procedure, it has been possible to predict de novo the structure of three very different protein topologies: an alpha/beta protein, the bovine pancreatic trypsin inhibitor (6pti), an alpha-helical protein, calbindin (3icb), and an all beta- protein, the SH3 domain of spectrin (1shg). In all cases, low resolution folds have been obtained with a root mean square deviation (RMSD) of 4.5-5.5 A with respect to the native structure. Some misfolded topologies appear in the simulations, but it is possible to select the native one on energetic grounds. Thus, it is demonstrated that the methodology is general for all protein motifs. Work is in progress in order to test the methodology on a larger set of protein structures.
A new, automated, knowledge-based method for the construction of three-dimensional models of proteins is described. Geometric restraints on target structures are calculated from a consideration of homologous template structures and the wider knowledge base of unrelated protein structures. Three-dimensional structures are calculated from initial partly folded states by high-temperature molecular dynamics simulations followed slow cooling of the system (simulated annealing) using nonphysical potentials. Three-dimensional models for the biotinylated domain from the pyruvate carboxylase of yeast and the lipoylated H-protein from the glycine cleavage system of pea leaf were constructed, based on the known structures of two lipoylated domains of 2-oxo acid dehydrogenase multienzyme complexes. Despite their weak sequence similarity, the three proteins are predicted to have similar three-dimensional structures, representative of a new protein module. Implications for the mechanisms of posttranslational modification of these proteins and their catalytic function are discussed.
We report the tertiary structure predictions for 95 proteins ranging in size from 17 to 160 residues starting from known secondary structure. Predictions are obtained from global minimization of an empirical potential function followed by the application of a refined atomic overlap potential. The minimization strategy employed represents a variant of the Monte Carlo plus minimization scheme of Li and Scheraga applied to a reduced model of the protein chain. For all of the cases except beta-proteins larger than 75 residues, a native-like structure, usually 4-6 A root-mean-square deviation from the native, is located. For beta-proteins larger than 75 residues, the energy gap between native-like structures and the lowest energy structures produced in the simulation is large, so that low RMSD structures are not generated starting from an unfolded state. This is attributed to the lack of an explicit hydrogen bond term in the potential function, which we hypothesize is necessary to stabilize large assemblies of beta-strands.
Freeze-induced perturbations of the protein native fold are poorly understood owing to the difficulty of monitoring their structure in ice. Here, we report that binding of the fluorescence probe 1-anilino-8-naphthalene sulfonate (ANS) to proteins in ice can provide a general monitor of ice-induced alterations of their tertiary structure. Experiments conducted with copper-free azurin from Pseudomonas aeruginosa and mutants I7S, F110S, and C3A/C26A correlate the magnitude of the ice-induced perturbation, as inferred from the extent of ANS binding, to the plasticity of the globular fold, increasing with less stable globular folds as well as when the flexibility of the macromolecule is enhanced. The distortion of the native structure inferred from ANS binding was found to draw a parallel with the extent of irreversible denaturation by freeze-thawing, suggesting that these altered conformations play a direct role on freeze damage. ANS binding experiments, extended to a set of proteins including serum albumin, alpha-amylase, beta-galactosidase, alcohol dehydrogenase from horse liver, alcohol dehydrogenase from yeast, lactic dehydrogenase, and aldolase, confirmed that a stressed condition of the native fold in the frozen state appears to be general to most proteins and pointed out that oligomers tend to be more labile than monomers presumably because the globular fold can be further destabilized by subunit dissociation. The results of this study suggest that the ANS binding method may find practical utility in testing the effectiveness of various additives employed in protein formulations as well as to devise safer freeze-drying protocols of pharmaceutical proteins.
The purpose of this work was to obtain information about protein tertiary structure in solid state by using steady state tryptophan (Trp) fluorescence emission spectroscopy on protein powders. Beta-lactoglobulin (betaLg) and interferon alpha-2a (IFN) powder samples were studied by fluorescence spectroscopy using a front surface sample holder. Two different sets of dried betaLg samples were prepared by vacuum drying of solutions: one containing betaLg, and the other containing a mixture of betaLg and guanidine hydrochloride. Dried IFN samples were prepared by vacuum drying of IFN solutions and by vacuum drying of polyethylene glycol precipitated IFN. The results obtained from solid samples were compared with the emission scans of these proteins in solutions. The emission scans obtained from protein powders were slightly blue-shifted compared to the solution spectra due to the absence of water. The emission scans were red-shifted for betaLg samples dried from solutions containing GuHCl. The magnitude of the shifts in lambda(max) depended on the extent of drying of the samples, which was attributed to the crystallization of GuHCl during the drying process. The shifts in the lambda(max) of the Trp emission spectrum are associated with the changes in the tertiary structure of betaLg. In the case of IFN, the emission scans obtained from PEG-precipitated and dried sample were different compared to the emission scans obtained from IFN in solution and from vacuum dried IFN. The double peaks observed in this sample were attributed to the unfolding of the protein. In the presence of trehalose, the two peaks converged to form a single peak, which was similar to solution emission spectra, whereas no change was observed in the presence of mannitol. We conclude that Trp fluorescence spectroscopy provides a simple and reliable means to characterize Trp microenvironment in protein powders that is related to the tertiary conformation of proteins in the solid state. This study shows that the use of fluorescence spectroscopy of proteins can be extended from simple protein aqueous solutions to protein powders, precipitates, and semidried protein samples to gain understanding of protein tertiary structure in these physical states.
The analysis of disulphide bond containing proteins in the Protein Data Bank (PDB) revealed that out of 27,209 protein structures analyzed, 12,832 proteins contain at least one intra-chain disulphide bond and 811 proteins contain at least one inter-chain disulphide bond. The intra-chain disulphide bond containing proteins can be grouped into 256 categories based on the number of disulphide bonds and the disulphide bond connectivity patterns (DBCPs) that were generated according to the position of half-cystine residues along the protein chain. The PDB entries corresponding to these 256 categories represent 509 unique SCOP superfamilies. A simple web-based computational tool is made freely available at the website that allows flexible queries to be made on the database in order to retrieve useful information on the disulphide bond containing proteins in the PDB. The database is useful to identify the different SCOP superfamilies associated with a particular disulphide bond connectivity pattern or vice versa. It is possible to define a query based either on a single field or a combination of the following fields, i.e., PDB code, protein name, SCOP superfamily name, number of disulphide bonds, disulphide bond connectivity pattern and the number of amino acid residues in a protein chain and retrieve information that match the criterion. Thereby, the database may be useful to select suitable protein structural templates in order to model the more distantly related protein homologs/analogs using the comparative modeling methods.
BACKGROUND: Success in solving the protein structure prediction problem relies on the choice of an accurate potential energy function. for a single protein sequence, it has been shown that the potential energy function can be optimized for predictive success by maximizing the energy gap between the correct structure and the ensemble of random structures relative to the distribution of the energies of these random structures (the Z-score). Different methods have been described for implementing this procedure for an ensemble of database proteins. Here, we demonstrate a new approach. RESULTS: For a single protein sequence, the probability of success (i.e the probability that the folded state is the lowest energy state) is derived. We then maximize the average probability of success for a set of proteins to obtain the optimal potential energy function. This results in maximum attention being focused on the proteins whose structures are difficult but not impossible to predict. CONCLUSIONS: Using a lattice model of proteins, we show that the optimal interaction potentials obtained by our method are both more accurate and more likely to produce successful predictions than those obtained by other averaging procedures.
A 3-D model of a protein can be constructed from its amino acid sequence and the 3-D structures of one or more homologues by annealing three sets of fragments: the structurally conserved regions, structurally variable regions and the side chains. The method encoded in the computer program COMPOSER was assessed by generating 3-D models of eight proteins whose crystal structures are already known and for which 3-D structures of homologues are available. In the structurally conserved regions, differences between modelled and X-ray structures are smaller than the differences between the X-ray structures of the modelled protein and the homologues used to build the model. When several homologues are used, the contributions of the known structures are weighted, preferably by the square of sequence similarity; this is especially important when the similarities of the homologues to the modelled structure differ greatly. The 'collar' extension approach, in which a similar region of different length in a homologue is used to extend the framework, can result in a more accurate model. If known homologues comprise more than one related group of proteins and they are both distantly related to the unknown, then alignment of the sequence to be modelled with each group of homologues facilitates identification of structurally conserved regions of the unknown and leads to an improved model. Models have root mean square differences (r.m.s.d.s) with the structures defined by X-ray analysis of between 0.73 and 1.56 A for all C alpha atoms, for seven the eight models. For the model of mucor pepsin, where the closest homologue has 33% sequence identity and 20% of the residues are in structurally variable regions, the r.m.s.d. for the framework region is 1.71 A and the r.m.s.d. for all C alpha atoms is 3.47 A.
Protein folding codes embodying local interactions including surface and secondary structure propensities and residue-residue contacts are optimized for a set of training proteins by using spin-glass theory. A screening method based on these codes correctly matches the structure of a set of test proteins with proteins of similar topology with 100% accuracy, even with limited sequence similarity between the test proteins and the structural homologs and the absence of any structurally similar proteins in the training set.
A challenge in computational protein folding is to assemble secondary structure elements-helices and strands-into well-packed tertiary structures. Particularly difficult is the formation of beta-sheets from strands, because they involve large conformational searches at the same time as precise packing and hydrogen bonding. Here we describe a method, called Geocore-2, that (1) grows chains one monomer or secondary structure at a time, then (2) disconnects the loops and performs a fast rigid-body docking step to achieve canonical packings, then (3) in the case of intrasheet strand packing, adjusts the side-chain rotamers; and finally (4) reattaches loops. Computational efficiency is enhanced by using a branch-and-bound search in which pruning rules aim to achieve a hydrophobic core and satisfactory hydrogen bonding patterns. We show that the pruning rules reduce computational time by 10(3)- to 10(5)-fold, and that this strategy is computationally practical at least for molecules up to about 100 amino acids long.
Success in the protein structure prediction problem relies heavily on the choice of an appropriate potential function. One approach toward extracting these potentials from a database of known protein structures is to maximize the Z-score of the database proteins, which represents the ability of the potential to discriminate correct from random conformations. These optimization methods model the entire distribution of alternative structures, reducing their ability to concentrate on the lowest energy structures most competitive with the native state and resulting in an unfortunate tendency to underestimate the repulsive interactions. This leads to reduced accuracy and predictive ability. Using a lattice model, we demonstrate how we can weight the distribution to suppress the contributions of the high-energy conformations to the Z-score calculation. The result is a potential that is more accurate and more likely to yield correct predictions than other Z-score optimization methods as well as potentials of mean force.