Search PubMed⌕ Search

Biomedical subjects

P Argos

Publications and source records attributed to P Argos.

At least 55 records · Page 3Linked to original sources

Hydrophobic regions on protein surfaces: definition based on hydration shell structure and a quick method for their computation.

The hydrophobic part of the solvent-accessible surface of a typical monomeric globular protein consists of a single, large interconnected region formed from faces of apolar atoms and constituting approximately 60% of the solvent-accessible surface area. Therefore, the direct delineation of the hydrophobic surface patches on an atom-wise basis is impossible. Experimental data indicate that, in a two-state hydration model, a protein can be considered to be unified with its first hydration shell in its interaction with bulk water. We show that, if the surface area occupied by water molecules bound at polar protein atoms as generated by AUTOSOL is removed, only about two-thirds of the hydrophobic part of the protein surface remains accessible to bulk solvent. Moreover, the organization of the hydrophobic part of the solvent-accessible surface experiences a drastic change, such that the single interconnected hydrophobic region disintegrates into many smaller patches, i.e. the physical definition of a hydrophobic surface region as unoccupied by first hydration shell water molecules can distinguish between hydrophobic surface clusters and small interconnecting channels. It is these remaining hydrophobic surface pieces that probably play an important role in intra- and intermolecular recognition processes such as ligand binding, protein folding and protein-protein association in solution conditions. These observations have led to the development of an accurate and quick analytical technique for the automatic determination of hydrophobic surface patches of proteins. This technique is not aggravated by the limiting assumptions of the methods for generating explicit water hydration positions. Formation of the hydrophobic surface regions owing to the structure of the first hydration shell can be computationally simulated by a small radial increment in solvent-accessible polar atoms, followed by calculation of the remaining exposed hydrophobic patches. We demonstrate that a radial increase of 0.35-0.50 A resembles the effect of tightly bound water on the organization of the hydrophobic part of the solvent-accessible surface.

Algorithms↗

Incorporation of non-local interactions in protein secondary structure prediction from the amino acid sequence.

Existing approaches to protein secondary structure prediction from the amino acid sequence usually rely on the statistics of local residue interactions within a sliding window and the secondary structural state of the central residue. The practically achieved accuracy limit of such single residue and single sequence prediction methods is 65% in three structural stages (alpha-helix, beta-strand and coil). Further improvement in the prediction quality is likely to require exploitation of various aspects of three-dimensional protein architecture. Here we make such an attempt and present an accurate algorithm for secondary structure prediction based on recognition of potentially hydrogen-bonded residues in a single amino acid sequence. The unique feature of our approach involves database-derived statistics on residue type occurrences in different classes of beta-bridges to delineate interacting beta-strands. The alpha-helical structures are also recognized on the basis of amino acid occurrences in hydrogen-bonded pairs (i,i + 4). The algorithm has a prediction accuracy of 68% in three structural stages, relies only on a single protein sequence as input and has the potential to be improved by 5-7% if homologous aligned sequences are also considered.

Algorithms↗

Intrahelical side chain-side chain contacts: the consequences of restricted rotameric states and implications for helix engineering and design.

Intrahelical side chain-side chain (sc-sc) interactions are assumed to play a crucial role in the formation and stability of alpha-helices, yet it was found that only 37.2% of all helical residues are involved in such close contacts, assuming a specific minimum contact distance. The majority (58.0%) of these were detected between residues with amino acid sequence spacing i, i + 4. The low frequency of intrahelical sc-sc contacts with sequence separations i, i + 1 and i, i + 3, each observed with only about one-third of the i, i + 4 counts, can be directly and generally attributed to the absence of the g- conformation in helices for the dihedral angle chi 1. However, if it was assumed that each side chain may maximally make only one sc-sc contact, as most commonly observed, the percentage of contacting pairs increased relative to the maximum possible pairs for a given sequence spacing by a factor of approximately 4, e.g. from 20.9 to 81.7% for i, i + 4 contacts. Stereochemical reasons are also given for the observation that i, i + 3 contacts are composed largely of ion or polar pairs, while hydrophobic residues dominate the i, i + 4 contacts. No significantly increased density of intrahelical sc-sc contacts with increasing helix length was found. Although there were generally fewer intrahelical contacts between buried helical residues when more contacts were made to the tertiary protein environment, the number of intrahelical contacts did not increase with increasing solvent exposure of the helices. Implications for helix design and the packing of helices are discussed.

Amino Acid Sequence↗

An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.

The sensitivity of most protein sequence alignment methods depends strongly on the quality of the comparison matrices used. These matrices, which assign weights or similarity scores to every possible amino acid substitution pair, are utilized to differentiate amongst the various possible alignments of two or more sequences. There are many ways to generate these exchange weights and new matrices are constantly published. There has been no overall assessment of these various matrices when applied in different alignment techniques and over many protein folds and families, both close and distant and with the use of several gap penalty values. In this work, a set of amino acid sequences matched by superposition of known protein tertiary topologies is used to test the alignment accuracy of the different method/matrix/penalty combinations. The comparisons show relatively similar results for the top scoring matrices, a preference for the global alignment method of Needleman and Wunsch, and the importance of matrix modification and optimized gap penalties. The relationship between the percentage identity in a resulting alignment and the level of correctness to be expected are given for the top-performing matrix, resulting in a better definition of the so-called "twilight zone". Estimates are made for the probability that two sequences, aligned at a certain level of residue percentage identity, are in fact unrelated.

Amino Acid Sequence↗

Comparison of atomic solvation parametric sets: applicability and limitations in protein folding and binding.

Atomic solvation parameters (ASP) are widely used to estimate the solvation contribution to the thermodynamic stability of proteins as well as the free energy of association for protein-ligand complexes. They are also included in several molecular mechanics computer programs. In this work, a total of eight atomic solvation parametric sets has been employed to calculate the solvation contribution to the free energy of folding delta Gs for 17 proteins. A linear correlation between delta Gs and the number of residues in each protein was found for each ASP set. The calculations also revealed a great variety in the absolute value and in the sign of delta Gs values such that certain ASP sets predicted the unfolded state to be more stable than the folded, whereas others yield precisely the opposite. Further, the solvation contribution to the free energy of association of helix pairs and to the disassociation of loops (connection between secondary structural elements in proteins) from the protein tertiary structures were computed for each of the eight ASP sets and discrepancies were evident among them.

Chemical Phenomena↗

A simple and fast approach to prediction of protein secondary structure from multiply aligned sequences with accuracy above 70%.

To improve secondary structure predictions in protein sequences, the information residing in multiple sequence alignments of substituted but structurally related proteins is exploited. A database comprised of 70 protein families and a total of 2,500 sequences, some of which were aligned by tertiary structural superpositions, was used to calculate residue exchange weight matrices within alpha-helical, beta-strand, and coil substructures, respectively. Secondary structure predictions were made based on the observed residue substitutions in local regions of the multiple alignments and the largest possible associated exchange weights in each of the three matrix types. Comparison of the observed and predicted secondary structure on a per-residue basis yielded a mean accuracy of 72.2%. Individual alpha-helix, beta-strand, and coil states were respectively predicted at 66.7, and 75.8% correctness, representing a well-balanced three-state prediction. The accuracy level, verified by cross-validation through jack-knife tests on all protein families, dropped, on average, to only 70.9%, indicating the rigor of the prediction procedure. On the basis of robustness, conceptual clarity, accuracy, and executable efficiency, the method has considerable advantage, especially with its sole reliance on amino acid substitutions within structurally related proteins.

Algorithms↗

Knowledge-based protein secondary structure assignment.

We have developed an automatic algorithm STRIDE for protein secondary structure assignment from atomic coordinates based on the combined use of hydrogen bond energy and statistically derived backbone torsional angle information. Parameters of the pattern recognition procedure were optimized using designations provided by the crystallographers as a standard-of-truth. Comparison to the currently most widely used technique DSSP by Kabsch and Sander (Biopolymers 22:2577-2637, 1983) shows that STRIDE and DSSP assign secondary structural states in 58 and 31% of 226 protein chains in our data sample, respectively, in greater agreement with the specific residue-by-residue definitions provided by the discoverers of the structures while in 11% of the chains, the assignments are the same. STRIDE delineates every 11th helix and every 32nd strand more in accord with published assignments.

Algorithms↗

Evidence on close packing and cavities in proteins.

The packing of a protein's constituent atoms and the attendant constraints placed upon them form the basis of many attempts to understand and predict protein structure, stability, folding and even function. Although the significance of packing is yet to be fully comprehended, recent experimental and theoretical investigations have increased our understanding through the description of mutational effects on structure and stability, determination of the limits of packing constraints for both protein folding and structure prediction, and delineation of packing guidelines on the basis of observed cavities in the native protein folds. These advances and allowing protein modellers, engineers and designers to tackle their problems from a more rational perspective.

Hydrogen Bonding↗

Increasing thermal stability of subtilisin from mutations suggested by strongly interacting side-chain clusters.

In this paper we present for seven subtilisin structures a systematic comparison of densely packed side-group clusters (defined as an ensemble of side chains with extensive internal atomic contacts as compared with those made with the surrounding protein environment and measured relative to the maximum possible for each residue type). Spatially consistent clusters are observed at structurally equivalent positions in the proteins, as revealed by careful multiple superpositioning of the respective backbone atoms. The clusters are positioned at strategic loop-connecting sites near the protein surfaces. The residues within consistent clusters displaying extensive association show varying conservation at structurally equivalent alignment sites. Suggestions for residue substitutions, as observed over the seven tertiary structures, were taken from the cluster positions and were shown to be consistent with a number of point mutations in one of the seven structures (savinase) that result in increased thermal stability.

Amino Acid Sequence↗

Detection of internal cavities in globular proteins.

We have undertaken a study of internal cavities in five protein structure groups, each containing different crystallographic structure determinations of the same protein, to understand better the nature of packing defects in protein tertiary architectures. Our results show that cavity detection and consistency of detection are highly dependent on probe and cavity size, cavity position within the globular protein and the local "quality' (r.m.s. deviation) of structural consistency within the group. The consistency of solvent placement within cavities has also been examined. We provide guidelines for estimating the likelihood of a given cavity to be an actual packing defect or to be a result of experimental error.

Animals↗

Protein structure prediction: recognition of primary, secondary, and tertiary structural features from amino acid sequence.

This review attempts a critical stock-taking of the current state of the science aimed at predicting structural features of proteins from their amino acid sequences. At the primary structure level, methods are considered for detection of remotely related sequences and for recognizing amino acid patterns to predict posttranslational modifications and binding sites. The techniques involving secondary structural features include prediction of secondary structure, membrane-spanning regions, and secondary structural class. At the tertiary structural level, methods for threading a sequence into a mainchain fold, homology modeling and assigning sequences to protein families with similar folds are discussed. A literature analysis suggests that, to date, threading techniques are not able to show their superiority over sequence pattern recognition methods. Recent progress in the state of ab initio structure calculation is reviewed in detail. The analysis shows that many structural features can be predicted from the amino acid sequence much better than just a few years ago and with attendant utility in experimental research. Best prediction can be achieved for new protein sequences that can be assigned to well-studied protein families. For single sequences without homologues, the folding problem has not yet been solved.

Amino Acid Sequence↗

The role of side-chain hydrogen bonds in the formation and stabilization of secondary structure in soluble proteins.

Intra-molecular side-chain:main-chain (sch:mch) and side-chain (sch:sch) hydrogen bonds observed in 44 well refined crystallographic protein structures with non-homologous sequences have been identified, classified and analysed to detect recurring structural patterns. Each observed bond was characterized by the position of its acceptor and donor groups relative to the N and C termini of the particular secondary structure in which they occur and according to their appearance within the same of sequentially separated secondary structures. The role of short-range hydrogen bonds in the formation and stabilization of a secondary structure and the importance of long-range hydrogen bonds as a cohesive force for different structural segments were also examined. It was found that the N terminus of alpha-helices is characterized by recurring sch:mch and sch:sch bonds with elements of the preceding coil segment, while at the C terminus a frequent intra-helix sch:mch hydrogen bond was frequently observed. The residues at or near the beta-strand termini often cross-linked, through hydrogen-bonding, non-sequential coil segments. Coil structures were characterized by recurring, internal sch:mch hydrogen bonding involving small polar side-chain groups situated at or near their N termini (coil N-capping). The significance of hydrogen bonds as formers and stabilizers of a protein fold and the association of its secondary structural units was also considered through an examination of bond density and distribution throughout the protein tertiary structure.

Amino Acids↗

Prediction of transmembrane segments in proteins utilising multiple sequence alignments.

A method for prediction of transmembrane segments from multiply aligned amino acid sequences is presented. For the calculations, two sets of propensity values were used: one for the middle, hydrophobic portion and one for the terminal regions of the transmembrane sequence spans. Average propensity values were calculated for each position along the alignment, with the contribution from each sequence weighted according to its dissimilarity relative to the other aligned sequences. Eight-residue segments were considered as potential cores of transmembrane segments and elongated if their middle propensity values were above a given threshold. End propensity values were also considered as stop signals. Only helices with length of 15 to 29 residues were allowed and corrections for strictly conserved charged residues were also made. The method is shown to be more successful than predictions based upon single sequences alone. In the test set of 28 families with 126 transmembrane segments, only five spans were not predicted or constituted false positives. The method is applied to sequence families for which data on transmembrane segments do not exist or are sparse or contradictory included voltage-gated potassium-channels, cytochrome c oxidases, NADH-ubiquinone oxidoreductase, beta-glucosides-specific phosphotransferase enzyme and major surface antigen of hepatitis B virus.

Algorithms↗

Folding the main chain of small proteins with the genetic algorithm.

Grid-free protein folding simulations were effected using the genetic algorithm, a backbone representation and standard dihedral angular conformations. The topological folding of idealized four-helix bundles was investigated in detail to differentiate among the important protein folding forces used as fitness criteria. Hydrophobic interactions were the most significant while local forces and hydrogen bonds were far less effective in promoting folding. Stable secondary structural regions were also important as nucleating centers. Using the fitness parameters optimized in idealized simulations together with standard secondary structure predictions derived from the amino acid sequence alone, the proper main-chain folding of the four-helix bundle proteins cytochrome b562, cytochrome c' and hemerythrin was achieved. In addition the backbone topology as predicted by the genetic algorithm for crambin, a mixed helix/strand protein with known structure, is presented and discussed.

Algorithms↗

Structural characteristics and stabilizing principles of bent beta-strands in protein tertiary architectures.

beta-Strands as constituents of beta-pleated sheets in protein tertiary structures often display considerable distortion from a purely extended conformation. The dislocation types are often characterized as "bulging," "twisting," and "bending." The former 2 properties have been extensively studied and classified. In this work an investigation of bent beta-structures is undertaken. The structural characteristics examined included the bending angles within and out of the principal strand plane, their distribution among various strand types such as parallel and antiparallel, the amino acid preferences at bend sites, and the usage of charged and polar residues for stabilization through interactive anchoring with other atoms of the beta-sheet within which the bent strand lies.

Amino Acids↗

Cavities and packing at protein interfaces.

An analysis of internal packing defects or "cavities" (both empty and water-containing) within protein structures has been undertaken and includes 3 cavity classes: within domains, between domains, and between protein subunits. We confirm several basic features common to all cavity types but also find a number of new characteristics, including those that distinguish the classes. The total cavity volume remains only a small fraction of the total protein volume and yet increases with protein size. Water-filled "cavities" possess a more polar surface and are typically larger. Their constituent waters are necessary to satisfy the local hydrogen bonding potential. Cavity-surrounding atoms are observed to be, on average, less flexible than their environments. Intersubunit and interdomain cavities are on average larger than the intradomain cavities, occupy a larger fraction of their resident surfaces, and are more frequently water-filled. We observe increased cavity volume at domain-domain interfaces involved with shear type domain motions. The significance of interfacial cavities upon subunit and domain shape complementarity and the protein docking problem, as well as in their structural and functional role in oligomeric proteins, will be discussed. The results concerning cavity size, polarity, solvation, general abundance, and residue type constituency should provide useful guidelines for protein modeling and design.

Amino Acids↗