Search PubMedSearch

Biomedical subjects

S Rackovsky

Publications and source records attributed to S Rackovsky.

At least 19 recordsLinked to original sources

On the existence and implications of an inverse folding code in proteins.

The existence of a code relating the set of possible sequences at a given position in a protein backbone to the local structure at that location is investigated. It is shown that only 73% of 4-C alpha structure fragments in a sample of 114 protein structures exhibit a preference for a particular set of sequences. The remaining structures can accommodate essentially any sequence. The structures that encode specific sequence distributions include the classical "secondary" structures, with the notable exception of planar (beta) bends. It is suggested that this has implications as to the mechanism of folding in proteins with extensive sheet/barrel structure. The possible role of structures that do not encode specific sequences as mutation hot spots is noted.

Biological Evolution

Protein sequence randomness and sequence/structure correlations.

We investigated protein sequence/structure correlation by constructing a space of protein sequences, based on methods developed previously for constructing a space of protein structures. The space is constructed by using a representation of the amino acids as vectors of 10 property factors that encode almost all of their physical properties. Each sequence is represented by a distribution of overlapping sequence fragments. A distance between any two sequences can be calculated. By attaching a weight to each factor, intersequence distances can be varied. We optimize the correlation between corresponding distances in the sequence and structure spaces. The optimal correlation between the sequence and structure spaces is significantly better than that which results from correlating randomly generated sequences, having the overall composition of the data base, with the structure space. However, sets of randomly generated sequences, each of which approximates the composition of the real sequence it replaces, produce correlations with the structure space that are as good as that observed for the actual protein sequences. A connection is proposed with previous studies of the protein folding code. It is shown that the most important property factors for the correlation of the sequence and structure spaces are related to helix/bend preference, side chain bulk, and beta-structure preference.

Amino Acid Sequence

Prediction of conformation of rat galanin in the presence and absence of water with the use of Monte Carlo methods and the ECEPP/3 force field.

The conformation of the 29-residue rat galanin neuropeptide was studied using the Monte Carlo with energy minimization (MCM) and electrostatically driven Monte Carlo (EDMC) methods. According to a previously elaborated procedure, the polypeptide chain was first treated in a united-residue approximation, in order to enable extensive exploration of the conformational space to be carried out (with the use of MCM). Then the low-energy united-residue conformations were converted to the all-atom representations, and EDMC simulations were carried out for the all-atom polypeptide chains, using the ECEPP/3 force field with hydration included. In order to estimate the effect of environment on galanin conformation, the low-energy conformations obtained as a result of these simulations were taken as starting structures for further EDMC runs that did not include hydration. The lowest-energy conformation obtained in aqueous solution calculations had a nonhelical N-terminal part packed against the nonpolar face of a residual helix that extended from Pro13 toward the C-terminus. One next lowest-energy structure was a nearly-all-helical conformation, but with a markedly higher energy. In contrast, all of the low-energy conformations in the absence of water were all-helical differing only by the extent to which the helix was kinked around Pro13. These results are in qualitative agreement with the available NMR and CD data of galanin in aqueous and nonaqueous solvents.

Amino Acid Sequence

Unfolding and refolding of the native structure of bovine pancreatic trypsin inhibitor studied by computer simulations.

A new procedure for studying the folding and unfolding of proteins, with an application to bovine pancreatic trypsin inhibitor (BPTI), is reported. The unfolding and refolding of the native structure of the protein are characterized by the dimensions of the protein, expressed in terms of the three principal radii of the structure considered as an ellipsoid. A dynamic equation, describing the variations of the principal radii on the unfolding path, and a numerical procedure to solve this equation are proposed. Expanded and distorted conformations are refolded to the native structure by a dimensional-constraint energy minimization procedure. A unique and reproducible unfolding pathway for an intermediate of BPTI lacking the [30,51] disulfide bond is obtained. The resulting unfolded conformations are extended; they contain near-native local structure, but their longest principal radii are more than 2.5 times greater than that of the native structure. The most interesting finding is that the majority of expanded conformations, generated under various conditions, can be refolded closely to the native structure, as measured by the correct overall chain fold, by the rms deviations from the native structure of only 1.9-3.1 A, and by the energy differences of about 10 kcal/mol from the native structure. Introduction of the [30,51] disulfide bond at this stage, followed by minimization, improves the closeness of the refolded structures to the native structure, reducing the rms deviations to 0.9-2.0 A. The unique refolding of these expanded structures over such a large conformational space implies that the folding is strongly dictated by the interactions in the amino acid sequence of BPTI. The simulations indicate that, under conditions that favor a compact structure as mimicked by the volume constraints in our algorithm, the expanded conformations have a strong tendency to move toward the native structure; therefore, they probably would be favorable folding intermediates. The results presented here support a general model for protein folding, i.e., progressive formation of partially folded structural units, followed by collapse to the compact native structure. The general applicability of the procedure is also discussed.

Animals

On the nature of the protein folding code.

This paper investigates quantitatively the characteristics of the local folding code. The overlapping four-residue fragments which make up the amino acid sequences of 114 proteins are divided into classes on the basis of the physical properties of their constituent amino acids. The distribution of structural types associated with each class of sequence fragment is determined and compared with an ensemble of random structural distributions of the same size selected from the actual protein structures. A criterion is proposed, based on the relative entropies of the two types of distribution, and on a hypothesis as to the characteristics of fragments which code for local structure, that makes it possible to identify those four-residue sequence elements which encode specific time-averaged structure. It is determined that, by this criterion, only 60-70% of the four-residue fragments encode specific structures. It is suggested that the remaining sequence fragments intrinsically encode susceptibility to conformational alteration under the influence of long-range interactions and that this susceptibility is required for correct folding of the molecule. This feature introduces an inherent indeterminacy into the local folding code. The implications of this observation for the prediction of protein structure by various methods are briefly discussed.

Amino Acids

Calculation of protein backbone geometry from alpha-carbon coordinates based on peptide-group dipole alignment.

An algorithm is proposed for the conversion of a virtual-bond polypeptide chain (connected C alpha atoms) to an all-atom backbone, based on determining the most extensive hydrogen-bond network between the peptide groups of the backbone, while maintaining all of the backbone atoms in energetically feasible conformations. Hydrogen bonding is represented by aligning the peptide-group dipoles. These peptide groups are not contiguous in the amino acid sequence. The first dipoles to be aligned are those that are both sufficiently close in space to be arranged in approximately linear arrays termed dipole paths. The criteria used in the construction of dipole paths are: to assure good alignment of the greatest possible number of dipoles that are close in space; to optimize the electrostatic interactions between the dipoles that belong to different paths close in space; and to avoid locally unfavorable amino acid residue conformations. The equations for dipole alignment are solved separately for each path, and then the remaining single dipoles are aligned optimally with the electrostatic field from the dipoles that belong to the dipole-path network. A least-squares minimizer is used to keep the geometry of the alpha-carbon trace of the resulting backbone close to that of the input virtual-bond chain. This procedure is sufficient to convert the virtual-bond chain to a real chain; in applications to real systems, however, the final structure is obtained by minimizing the total ECEPP/2 (empirical conformational energy program for peptides) energy of the system, starting from the geometry resulting from the solution of the alignment equations. When applied to model alpha-helical and beta-sheet structures, the algorithm, followed by the ECEPP/2 energy minimization, resulted in an energy and backbone geometry characteristic of these alpha-helical and beta-sheet structures. Application to the alpha-carbon trace of the backbone of the crystallographic 5PTI structure of bovine pancreatic trypsin inhibitor, followed by ECEPP/2 energy minimization with C alpha-distance constraints, led to a structure with almost as low energy and root mean square deviation as the ECEPP/2 geometry analog of 5PTI, the best agreement between the crystal and reconstructed backbone being observed for the residues involved in the dipole-path network.

Algorithms

Prediction of protein conformation on the basis of a search for compact structures: test on avian pancreatic polypeptide.

Based on the concept that hydrophobic interactions cause a polypeptide chain to adopt a compact structure, a method is proposed to predict the structure of a protein. The procedure is carried out in four stages: (1) use of a virtual-bond united-residue approximation with the side chains represented by spheres to search conformational space extensively using specially designed interactions to lead to a collapsed structure, (2) conversion of the lowest-energy virtual-bond united-residue chain to one with a real polypeptide backbone, with optimization of the hydrogen-bond network among the backbone groups, (3) perturbation of the latter structure by the electrostatically driven Monte Carlo (EDMC) procedure, and (4) conversion of the spherical representation of the side chains to real groups and perturbation of the whole molecule by the EDMC procedure using the empirical conformational energy program for peptides (ECEPP/2) energy function plus hydration. Application of this procedure to the 36-residue avian pancreatic polypeptide led to a structure that resembled the one determined by X-ray crystallography; it had an alpha-helix starting at residue 13, with the N-terminal portion of the chain in an extended conformation packed against the alpha-helix. Similar structures with slightly higher energies, but looser packing, were also obtained.

Amino Acids

Effects of compact volume and chain stiffness on the conformations of native proteins.

An investigation of the statistical properties of the native conformations of proteins, observed from crystal structures, is reported. Protein conformations were analyzed in terms of a bond vector correlation function and molecular volume. It was observed that, while the volume of a protein structure varies nearly linearly with the number of residues, the bond vector correlation function exhibits a universal feature for all sizes of proteins. To interpret the nature of the bond vector correlation function of native protein structures quantitatively, Monte Carlo simulations of realistic polypeptide chains of specific but arbitrary amino acid sequence were carried out. The molecule was constrained in an ellipsoidal volume determined by its chain length, and conformations with unacceptable nonbonded contacts between different amino acid residues were excluded. The interactions within a terminally blocked single residue, which correlate two nearest-neighbor peptide groups in a chain, were taken into account by an energetically biased sampling of its phi-psi space. The simulated chain correlation functions were found to be in good agreement with those of the crystal structures of beta-sheet-type and mixed-type (alpha+beta) proteins of similar length. On the basis of these calculations, it is concluded that the observed conformations of these native proteins may arise from two basic factors: the compactness of structures under hydrophobic interactions and the intrinsic stiffness of polypeptide chains due to the interactions within each terminally blocked residue.

Biophysical Phenomena

Quantitative organization of the known protein x-ray structures. I. Methods and short-length-scale results.

We address herein the problem of delineating the relationships between the known protein structures. In order to study this problem, methods have been developed to represent arbitrarily sized fragments of biopolymer backbone, and to compare distributions of such fragments. These methods are applied to a classification of 123 structures representing the entire set of known x-ray structures. The resulting data are analyzed (on the four-C alpha length scale) to determine both the large-scale organization of the set of known structures (i.e., the relationships between large groups of structures, each comprised of proteins that are structurally related) and its local structure (i.e., the quantitative degree of similarity between any two specific structures). It is shown that the set of structures forms a continuum of structural types, ranging from all-helical to all-sheet/barrel proteins. It is further demonstrated that the density of protein structures is not uniform across this continuum, but rather that structures cluster in certain regions, separated by regions of lower population. The properties of the various regions of the structural space are determined. The existence is demonstrated of strong quantitative correlations between the contents of different types of four-C alpha fragments within protein structures, which imply significant constraints on the types of architecture that can occur in proteins. Analysis of the distribution of structures demonstrates some hitherto unsuspected similarities and suggests that, in some circumstances, neither structural similarity nor sequence homology may be necessary conditions for evolutionary relationship between proteins. It is also suggested that these unsuspected similarities may imply similar folding mechanisms for structures of apparently different global architecture. Cases are also noted in which apparently similar structures may fold by different mechanisms. The connection between structure and dynamic properties is discussed, and a possible role of dynamics in the evolution of protein structures is suggested. The sensitivity of the methods presented herein to anomalies of structure refinement is demonstrated. It is suggested that the present results provide a framework for analyzing experimental results on structural similarity obtained using vibrational circular dichroism spectra, which are sensitive to local backbone structure.

Animals

Comparison of the predicted structure for the activated form of the P21 protein with the X-ray crystal structure.

The predicted conformation and position of the central transforming region (residues 55-67) of the p21 protein are compared with the conformation and position of this segment in a recently determined X-ray crystal structure of residues 1-166 of this protein in the activated state bound to a nonhydrolyzable GTP derivative. We previously predicted that this segment of the protein would adopt a roughly extended conformation from Ile 55-Thr 58, a reverse turn at Ala 59-Gln 61, followed by an alpha-helix from Glu 62-Met 67. We further predicted that this region of the activated protein occupies a position that is virtually identical to corresponding regions in the homologous purine nucleotide-binding proteins, bacterial elongation factor (EF-tu), and adenylate kinase (ADK). We find that there is a close correspondence between the conformation and position of our predicted structure and those found in the X-ray crystal structure. A mechanism for activation of the protein is proposed and is corroborated by X-ray crystallographic data.

Computer Simulation

The structure of the carboxyl terminus of the p21 protein. Structural relationship to the nucleotide-binding/transforming regions of the protein.

The carboxyl-terminal region of the ras oncogene-encoded p21 protein is critical to the protein's function, since membrane binding through the C-terminus is necessary for its cellular activity. X-ray crystal structures for truncated p21 proteins are available, but none of these include the C-terminal region of the protein (from residues 172-189). Using conformational energy analysis, we determined the preferred three-dimensional structures for this C-terminal octadecapeptide of the H-ras oncogene p21 protein and generated these structures onto the crystal structure of the remainder of the protein. The results indicate that, like other membrane-associated proteins, the membrane-binding C-terminus of p21 assumes a helical hairpin conformation. In several low-energy orientations, the C-terminal structure is in close proximity to other critical locales of p21. These include the central transforming region (around Gln 61) and the amino terminal transforming region (around Gly 12), indicating that extracellular signals can be transduced through the C-terminal helical hairpin to the effector regions of the protein. This finding is consistent with the results of recent genetic experiments.

Amino Acid Sequence

Correlation of the structure of the transmembrane domain of the neu oncogene-encoded p185 protein with its function.

The human homologue of the neu oncogene is frequently found in human tumors. Certain amino acid substitutions at position 664 in the transmembrane domain of the neu oncogene-encoded p185 protein product are known to cause malignant transformation of cells. Using conformational energy analysis based on ECEPP (empirical conformational energies for polypeptides program), we have previously determined the preferred three-dimensional structures for the transmembrane domain of the p185 protein with a transforming (glutamic acid) and a nontransforming (valine) substitution at the critical position 664 and found that the global minimum-energy conformation of this region in the nontransforming protein contains a sharp bend, whereas the global minimum-energy conformation for this region from the transforming protein is entirely alpha-helical. We now demonstrate that this result holds for other known nontransforming (glycine, histidine, tyrosine, and lysine) and transforming (glutamine) substitutions at position 664. Furthermore, a simple statistical thermodynamic analysis of the results indicates that approximately 85% of each of the nontransforming sequences exist with the bend at positions 664 and 665, while approximately 90% of each of the transforming sequences exist as an alpha-helix. About 9% of the nontransforming sequences exist as the alpha-helix. These results suggest that if the intracellular concentration of the normal protein is increased at least 10-fold, thereby increasing the alpha-helical form by this factor, cell transformation should result. This conclusion is directly supported by genetic experiments in which this level of overexpression of the normal protein was achieved with attendant cell transformation.

Amino Acid Sequence

Conformations of the central transforming region (Ile 55-Met 67) of the p21 protein and their relationship to activation of the protein.

The GTP-binding p21 protein, encoded by the ras-oncogene, becomes transforming if amino acid substitutions are made at critical positions in the polypeptide chain, e.g., at Gly 12, Gly 13, Ala 59, Gln 61 and Glu 63. Most of these substitutions occur in two phosphate-binding loop regions, Tyr 4-Thr 20, herein designated as segment 1, and Ile 55-Met 67, herein designated, as segment 2. These two segments are homologous to two corresponding regions in the two purine nucleotide binding proteins, bacterial elongation factor (EF-tu) (Val 12-Thr 28 corresponds to segment 1; His 78-Ile 92 corresponds to segment 2) and adenylate kinase (ADK) (Lys 9-Cys 25 corresponds to segment 1 and Tyr 95-Arg 107 corresponds to segment 2). We find that the conformations of the segment 1 region in the p21 protein, EF-tu and ADK are similar to one another and that the conformation of the segment 2 region of EF-tu is superimposable on that of segment 2 of ADK. Furthermore, the relative position of the two segments in EF-tu is strikingly similar to that of the two segments in ADK. In the originally proposed X-ray structure for the p21 protein, the conformation of segment 2 in the p21 protein is not similar to that found for the other two proteins, and its disposition relative to segment 1 and the remainder of the protein is also different from that observed for the other two proteins.(ABSTRACT TRUNCATED AT 250 WORDS)

Adenylate Kinase

Protein comparison and classification: a differential geometric approach.

A method is proposed for rapidly, quantitatively comparing protein structures of arbitrary sizes, based on the differential geometric representation. The method is applied to a group of 22 protein x-ray structures, and the resulting network of closest relationships is delineated. Several well-known fold types are automatically detected as groupings of related structures, even when the constituent proteins are of different lengths. A complete gradation of types is shown to be detected, ranging from all-helical to all-beta-structure proteins. A relationship among functionally similar proteins is shown in several cases, even where their three-dimensional structures differ. It is suggested that the positions of proteins within the network of relationships correspond with their folding mechanisms.

Protein Conformation

Substitutions of proline 76 in yeast iso-1-cytochrome c. Analysis of residues compatible and incompatible with folding requirements.

Fine-structure genetic mapping previously revealed numerous nonfunctional cyc1 mutations having alterations at or near the site corresponding to amino acid position 76 of iso-1-cytochrome c from the yeast Saccharomyces cerevisiae. DNA sequencing of the alterations in four of these cyc1 mutations indicated that the normal Pro-76 was replaced by Leu-76. Revertants containing at least partially functional iso-1-cytochromes c were isolated, and the alterations were analyzed by DNA sequencing and protein analysis. Specific activities of the altered iso-1-cytochromes c were estimated in vivo by growth of the strains in lactate medium; compared to normal iso-1-cytochrome c with Pro-76, the following activities were associated with the following replacements: approximately 90% for Val-76, approximately 60% for Thr-76, approximately 30% for Ser-76, approximately 20% for Ile-76, and 0% for Leu-76. In order to develop an understanding of the factors that determine whether or not an altered iso-1-cytochrome c will function, we undertook a theoretical analysis which led to the conclusion that the activity of the proteins was dependent on both short- and long-range interactions. Short-range interactions were estimated from studies on known protein structures which gave the likelihood that various amino acids would be found in a local backbone configuration similar to the native protein; long-range interactions with the rest of the molecule were analyzed by considering the size of the side chain. We believe this approach can be used to analyze a wide variety of mutant proteins.

Amino Acid Sequence

On the redox conformational change in cytochrome c.

The relationship between the crystal structures of oxidized and reduced tuna cytochrome c has been reexamined by a superposition method motivated by recent studies of the cytochrome c-cytochrome c peroxidase complex. It is shown that the observed structural changes precisely reflect the binding face suggested by chemical modification studies. It is further suggested that the large observed motion of lysine-27 and a smaller overall motion of the two binding edges constitute a redox binding-affinity switch and that the driving force for the conformational change of the protein is provided by the internal conformational change and charge redistribution of the heme, which cause it to tilt, under the influence of covalent and nonbonded interactions, within its protein envelope. A picture is presented of the molecule as an electron storage/transfer machine with three elements--a binding module, an electron storage module, and a conformational energy-storage module.

Animals