Search PubMed⌕ Search

Biomedical subjects

P Argos

Publications and source records attributed to P Argos.

At least 37 records · Page 2Linked to original sources

The future of protein secondary structure prediction accuracy.

BACKGROUND: The accuracy of secondary structure prediction for a protein from knowledge of its sequence has been significantly improved by about 7% to the 70-75% range by inclusion of information residing in sequences similar to the query sequence. The scientific literature has been inconsistent, if not negative, regarding chances for further improvement from the vast knowledge to be provided by genome sequencing efforts. RESULTS: By applying a prediction technique that is particularly sensitive to added sequence information to a standard set of query sequences with related primary structures taken from chronologically successive releases of the SWISS-PROT database, it is shown that prediction accuracy can be expected to reach 80-85% with a large 10-fold increase in present sequence knowledge. CONCLUSIONS: Even with present prediction approaches, improvement in prediction accuracy can still be expected, albeit limited to no more than 10%.

Algorithms↗

Protein thermal stability: hydrogen bonds or internal packing?

Thermally stable proteins are of interest for several reasons. They can be used to improve the efficiency of many industrial processes and provide insight into the general mechanisms of protein folding and stabilization. Comparison of tertiary structural properties of several protein families with members of different thermostability should help to delineate the role of individual factors in achieving stability at high temperature. In this work, 16 protein families with at least one known thermophilic and one known mesophilic tertiary structure were examined for the number and type of hydrogen bonds and salt links, polar surface composition, internal cavities and packing densities, and secondary structural composition. The results show a consistent increase in the number of hydrogen bonds and in polar surface area fraction with increased thermostability.

Hydrogen Bonding↗

Prediction of membrane protein topology utilizing multiple sequence alignments.

A technique for prediction of protein membrane topology (intra- and extracellular sidedness) has been developed. Membrane-spanning segments are first predicted using an algorithm based upon multiply aligned amino acid sequences. The compositional differences in the protein segments exposed at each side of the membrane are then investigated. The ratios are calculated for Asn, Asp, Gly, Phe, Pro, Trp, Tyr, and Val, mostly found on the extracellular side, and for Ala, Arg, Cys, and Lys, mostly occurring on the intracellular side. The consensus over these 12 residue distributions is used for sidedness prediction. The method was developed with a set of 42 protein families for which all but one were correctly predicted with the new algorithm. This represents an improvement over previous techniques. The new method, applied to a set of 12 membrane protein families different from the test set and with recently determined topologies, performed well, with 11 of 12 sidedness assignments agreeing with experimental results. The method has also been applied to several membrane protein families for which the topology has yet to be determined. An electronic prediction service is available at the E-mail address tmap@embl-heidelberg.de and on WWW via http://www.embl-heidelberg.de.

Amino Acid Sequence↗

Correlation between side chain mobility and conformation in protein structures.

Thermal factors of protein atoms as determined by X-ray crystallographic techniques show a tendency to be larger in side chains with unfavourable local conformations rather than in those displaying conformational energy minima. It follows that side chain atoms are more mobile if they are in a non-rotameric configuration and that the stereochemistry of protein structures cannot be fully assessed or simulated without consideration of thermal factors that monitor flexibility in various regions of the protein. The observations should also prove useful in protein folding and design.

Amino Acids↗

Applying experimental data to protein fold prediction with the genetic algorithm.

Specific residue interactions as revealed from a few and readily available experiments can be quite important in shaping a protein's tertiary topology by complementing basic and general folding principles. This experimental information is employed in structure prediction (mainchain topology) based on sequence knowledge and the genetic algorithm with its ability to optimize simultaneously many parameters. Examples investigated include the distribution of cysteinyl S-S bonds, protein side-chain ligands to iron-sulfur cages, cofactor-ligands, crosslinks amongst side-chains, and conserved hydrophobic and catalytic residues. Such interactions yield an improvement in the predicted topology (0.4-6.6 A root mean square deviation in the positions of the backbone C alpha-atoms relative to those observed) compared with those resulting from simulations relying only on basic protein folding principles. For several examples the resultant topology depended critically on knowledge of the few and specific interactions such that the relationship between predicted and observed C alpha-positions was near random without their use. The combined methodology (experimental data and the genetic algorithm) should prove helpful in settings where experiment and theory can cooperate in successive steps to elucidate an unknown structure.

Algorithms↗

Hydrophobic patches on protein subunit interfaces: characteristics and prediction.

Hydrophobic patches, defined as clusters of neighboring apolar atoms deemed accessible on a given protein surface, have been investigated on protein subunit interfaces. The data were taken from known tertiary structures of multimeric protein complexes. Amino acid composition and preference, patch size distribution, and patch contact complementarity across associating subunits were examined and compared with hydrophobic patches found on the solvent-accessible surface of the multimeric complexes. The largest or second largest patch on the accessible surface of the entire subunit was involved in multimeric interfaces in 90% of the cases. These results should prove useful for subunit design and engineering as well as for prediction of subunit interface regions.

Amino Acids↗

A functional role for protein cavities in domain: domain motions.

Motions between individual domains are known to play an important role in protein function. Protein cavities at domain interfaces have been suggested to facilitate such movements. Consequently, the cavity morphology in a set of multi-domain proteins has been critically examined. The conformational changes were well characterised by atomic resolution tertiary structures prior to and after domain motions. The results showed that interdomain cavities play a number of specific functional roles by either facilitating, or being otherwise involved with, domain: domain motions. Correspondingly, a higher fraction of cavity surface is observed at domain interfaces as compared to that buried within individual domains. Furthermore, interdomain cavity-forming residues were found to be highly conserved in terms of amino acid residue sequence and volume within their aligned protein families, more so than residues exclusive to the domain interface and intradomain cavities. These results provide substantial evidence of cavities fulfilling a specific functional role in multi-domain proteins.

Amino Acid Sequence↗

Identifying the tertiary fold of small proteins with different topologies from sequence and secondary structure using the genetic algorithm and extended criteria specific for strand regions.

Grid-free protein folding simulations based on sequence and secondary structure knowledge (using mostly experimentally determined secondary structure information but also analysing results from secondary structure predictions) were investigated using the genetic algorithm, a backbone representation, and standard dihedral angular conformations. Optimal structures are selected according to basic protein building principles. Having previously applied this approach to proteins with helical topology, we have now developed additional criteria and weights for beta-strand-containing proteins, validated them on four small beta-strand-rich proteins with different topologies, and tested the general performance of the method on many further examples from known protein structures with mixed secondary structural type and less than 100 amino acid residues. Topology predictions close to the observed experimental structures were obtained in four test cases together with fitness values that correlated with the similarity of the predicted topology to the observed structures. Root-mean-square deviation values of C alpha atoms in the superposed predicted and observed structures, the latter of which had different topologies, were between 4.5 and 5.5 A(2.9 to 5.1 A without loops). Including 15 further protein examples with unique folds, root-mean-square deviation values ranged between 1.8 and 6.9 A with loop regions and averaged 5.3 A and 4.3 A, including and excluding loop regions, respectively.

Algorithms↗

Principles of helix-helix packing in proteins: the helical lattice superposition model.

The geometry of helix-helix packing in globular proteins is comprehensively analysed within the model of the superposition of two helix lattices which result from unrolling the helix cylinders onto a plane containing points representing each residue. The requirements for the helix geometry (the radius R, the twist angle omega and the rise per residue delta) under perfect match of the lattices are studied through a consistent mathematical model that allows consideration of all possible associations of all helix types (alpha-, pi- and 3(10)). The corresponding equations have three well-separated solutions for the interhelical packing angle, omega, as a function of the helix geometric parameters allowing optimal packing. The resulting functional relations also show unexpected behaviour. For a typically observed alpha-helix (omega = 99.1 degrees, delta = 1.45 A), the three optimal packing angles are omega a,b,c = -37.1 degrees, -97.4 degrees and +22.0 degrees with a periodicity of 180 degrees and respective helix radii Ra,b,c = 3.0 A, 3.5 A and 4.3 A. However, the resulting radii are very sensitive to variations in the twist angle omega. At omega triple = 96.9 degrees, all three solutions yield identical radii at delta = 1.45 A where Rtriple = 3.46 A. This radius is close to that of a poly(Ala) helix, indicating a great packing flexibility when alanine is involved in the packing core, and omega triple is close to the mean observed twist angle. In contrast, the variety of possible theoretical solutions is limited for the other two helix types. Besides the perfect matches, novel suboptimal "knobs into holes" hydrophobic packing patterns as a function of the helix radius are described. Alternative "knobs onto knobs" and mixed models can be applied in cases where salt bridges, hydrogen bonds, disulphide bonds and tight hydrophobic head-to-head contacts are involved in helix-helix associations. An analysis of the experimentally observed packings in proteins confirmed the conclusions of the theoretical model. Nonetheless, the observed alpha-helix packings showed deviations from the 180 degrees periodicity expected from the model. An investigation of the actual three-dimensional geometry of helix-helix packing revealed an explanation for the observed discrepancies where a decisive role was assigned to the defined orientation of the C alpha-C beta vectors of the side-chains. As predicted form the model, helices with different radii (differently sized side-chains in the packing core) were observed to utilize different packing cells (packing patterns). In agreement with the coincidence between Rtriple and the radius of a poly(Ala) helix, Ala was observed to show greatest propensity to build the packing core. The application of the helix lattice superposition model suggests that the packing of amino acid residues is best described by a "knobs into holes" scheme rather than "ridges into grooves". The various specific packing modes made salient by the model should be useful in protein engineering and design.

Algorithms↗

Prediction of secondary structural content of proteins from their amino acid composition alone. I. New analytic vector decomposition methods.

The predictive limits of the amino acid composition for the secondary structural content (percentage of residues in the secondary structural states helix, sheet, and coil) in proteins are assessed quantitatively. For the first time, techniques for prediction of secondary structural content are presented which rely on the amino acid composition as the only information on the query protein. In our first method, the amino acid composition of an unknown protein is represented by the best (in a least square sense) linear combination of the characteristic amino acid compositions of the three secondary structural types computed from a learning set of tertiary structures. The second technique is a generalization of the first one and takes into account also possible compositional couplings between any two sorts of amino acids. Its mathematical formulation results in an eigenvalue/eigenvector problem of the second moment matrix describing the amino acid compositional fluctuations of secondary structural types in various proteins of a learning set. Possible correlations of the principal directions of the eigenspaces with physical properties of the amino acids were also checked. For example, the first two eigenvectors of the helical eigenspace correlate with the size and hydrophobicity of the residue types respectively. As learning and test sets of tertiary structures, we utilized representative, automatically generated subsets of Protein Data Bank (PDB) consisting of non-homologous protein structures at the resolution thresholds < or = 1.8A, < or = 2.0A, < or = 2.5A, and < or = 3.0 A. We show that the consideration of compositional couplings improves prediction accuracy, albeit not dramatically. Whereas in the self-consistency test (learning with the protein to be predicted), a clear decrease of prediction accuracy with worsening resolution is observed, the jackknife test (leave the predicted protein out) yielded best results for the largest dataset (< or = 3.0A, almost no difference to the self-consistency test!), i.e., only this set, with more than 400 proteins, is sufficient for stable computation of the parameters in the prediction function of the second method. The average absolute error in predicting the fraction of helix, sheet, and coil from amino acid composition of the query protein are 13.7, 12.6, and 11.4%, respectively with r.m.s. deviations in the range of 8.6 divided by 11.8% for the 3.0 A dataset in a jackknife test. The absolute precision of the average absolute errors is in the range of 1 divided by 3% as measured for other representative subsets of the PDB. Secondary structural content prediction methods found in the literature have been clustered in accordance with their prediction accuracies. To our surprise, much more complex secondary structure prediction methods utilized for the same purpose of secondary structural content prediction achieve prediction accuracies very similar to those of the present analytic techniques, implying that all the information beyond the amino acid composition is, in fact, mainly utilized for positioning the secondary structural state in the sequence but not for determination of the overall number of residues in a secondary structural type. This result implies that higher prediction accuracies cannot be achieved relying solely on the amino acid composition of an unknown query protein as prediction input. Our prediction program SSCP has been made available as a World Wide Web and E-mail service.

Amino Acids↗

Prediction of secondary structural content of proteins from their amino acid composition alone. II. The paradox with secondary structural class.

The success rates reported for secondary structural class prediction with different methods are contradictory. On one side, the problem of recognizing the secondary structural class of a protein knowing only its amino acid composition appears completely solved by simply applying jury decision with an elliptically scaled distance function. Chou and coworkers repeatedly (see Crit. Rev. Biochem. Mol. Biol. 30:275-349, 1995) published prediction accuracies near 100%. On the other hand, traditional secondary structure prediction techniques achieve success rates of about 70% for the secondary structural state per residue and about 75% for structural class only with extensive input information (full sequence of the query protein, its amino acid composition and length, multiple alignments with homologous sequences). In this article, we resolve the paradox and consider (1) the question of the secondary structural class definition, (2) the role of the representativity of the test set of protein tertiary structure for the current state of the Protein Data Bank (PDB); and (3) we estimate the real impact of amino acid composition on secondary structural class. We formulate three objective criteria for a reasonable definition of secondary structural classes and show that only the criterion of Nakashima et al. (J. Biochem. 99:153-162, 1986) complies with all of them. Only this definition matches the distribution of secondary structural content in representative PDB subsets, whereas other criteria leave many proteins (up to 65% of all PDB entries) simply unassigned. We review critically specialized secondary-structural class prediction methods, especially those of Chou and coworkers, which claim almost 100% accuracy using only amino acid composition, and resolve the paradox that these prediction accuracies are better than those from secondary structure predictions from multiple alignments. We show (i) that these techniques rely on a preselection of test sets which removes irregular proteins and other proteins without any class assignment (about 35% of all PDB entries); and (ii) that even for preselected representative test sets, the success rate drops to 60% and lower for a 4-type classification (alpha, beta, alpha + beta, alpha/beta). The prediction accuracies fall to about 50% if the secondary structural class definition of Nakashima et al. is applied and only few irregular proteins are preselected and removed from automatically generated, representative subsets of the PDB. We have applied two new vector decomposition methods for secondary structural content prediction from amino acid composition alone, with and without consideration of amino acid compositional coupling in the learning set of tertiary structures respectively, to the problem of class prediction and achieve about 60% correct assignment among four classes (alpha, beta, mixed, irregular) as well as single sequence-based secondary structure prediction methods like GORIII and COMBI. Our results demonstrate that 60% correctness is the upper limit for a 4-type class prediction from amino acid composition alone for an unknown query protein and that consideration of compositional coupling does not improve the prediction success. The prediction program SSCP offering secondary structural class assignment for query compositions and sequences has been made available as a World Wide Web and E-mail service.

Amino Acids↗

Hydrophobic patches on the surfaces of protein structures.

A survey of hydrophobic patches on the surface of 112 soluble, monomeric proteins is presented. The largest patch on each individual protein averages around 400 A2 but can range from 200 to 1,200 A2. These areas are not correlated to the sizes of the proteins and only weakly to their apolar surface fraction. Ala, Lys, and Pro have dominating contributions to the apolar surface for smaller patches, while those of the hydrophobic amino acids become more important as the patch size increases. The hydrophilic amino acids expose an approximately constant fraction of their apolar area independent of patch size; the hydrophobic residue types reach similar exposure only in the larger patches. Though the mobility of residues on the surface is generally higher, it decreases for hydrophilic residues with increasing patch size. Several characteristics of hydrophobic patches catalogued here should prove useful in the design and engineering of proteins.

Amino Acids↗

A method for detecting hydrophobic patches on protein surfaces.

A method for the detection of hydrophobic patches on the surfaces of protein tertiary structures is presented. It delineates explicit contiguous pieces of surface of arbitrary size and shape that consist solely of carbon and sulphur atoms using a dot representation of the solvent-accessible surface. The technique is also useful in detecting surface segments with other characteristics, such as polar patches. Its potential as a tool in the study of protein-protein interactions and substrate recognition is demonstrated by applying the method to myoglobin, Leu/IIe/Val-binding protein, lipase, lysozyme, azurin, triose phosphate isomerase, carbonic anhydrase, and phosphoglycerate kinase. Only the largest patches, having sizes exceeding random expectation, are deemed meaningful. In addition to well-known hydrophobic patches on these proteins, a number of other patches are found, and their significance is discussed. The method is simple, fast, and robust. The program text is obtainable by anonymous ftp.

Models, Molecular↗

Topology prediction of membrane proteins.

A new method is described for prediction of protein membrane topology (intra- and extracellular sidedness) from multiply aligned amino acid sequences after determination of the membrane-spanning segments. The prediction technique relies on residue compositional differences in the protein segments exposed at each side of the membrane. Intra/extracellular ratios are calculated for the residue types Asn, Asp, Gly, Phe, Pro, Trp, Tyr, and Val, preferably found on the extracellular side, and for Ala, Arg, Cys, and Lys, mostly occurring on the intracellular side. The consensus over these 12 residue distributions is used for sidedness prediction. The method was developed with a test set of 42 protein families, for which all but one were correctly predicted with the new algorithm. This represents an improvement over predictions based on the widely used "positive-inside rule" and other techniques, where at least six mispredictions were observed for the same data set. Further, application of this and other methods to 12 protein families not in the test set still showed the better performance of the present technique, which was subsequently applied to another set of membrane protein families where the topology has yet to be determined.

Algorithms↗

Ribosome-mediated translational pause and protein domain organization.

Because regions on the messenger ribonucleic acid differ in the rate at which they are translated by the ribosome and because proteins can fold cotranslationally on the ribosome, a question arises as to whether the kinetics of translation influence the folding events in the growing nascent polypeptide chain. Translationally slow regions were identified on mRNAs for a set of 37 multidomain proteins from Escherichia coli with known three-dimensional structures. The frequencies of individual codons in mRNAs of highly expressed genes from E. coli were taken as a measure of codon translation speed. Analysis of codon usage in slow regions showed a consistency with the experimentally determined translation rates of codons; abundant codons that are translated with faster speeds compared with their synonymous codons were found to be avoided; rare codons that are translated at an unexpectedly higher rate were also found to be avoided in slow regions. The statistical significance of the occurrence of such slow regions on mRNA spans corresponding to the oligopeptide domain termini and linking regions on the encoded proteins was assessed. The amino acid type and the solvent accessibility of the residues coded by such slow regions were also examined. The results indicated that protein domain boundaries that mark higher-order structural organization are largely coded by translationally slow regions on the RNA and are composed of such amino acids that are stickier to the ribosome channel through which the synthesized polypeptide chain emerges into the cytoplasm. The translationally slow nucleotide regions on mRNA possess the potential to form hairpin secondary structures and such structures could further slow the movement of ribosome. The results point to an intriguing correlation between protein synthesis machinery and in vivo protein folding. Examination of available mutagenic data indicated that the effects of some of the reported mutations were consistent with our hypothesis.

Bacterial Proteins↗

Protein secondary structural types are differentially coded on messenger RNA.

Tricodon regions on messenger RNAs corresponding to a set of proteins from Escherichia coli were scrutinized for their translation speed. The fractional frequency values of the individual codons as they occur in mRNAs of highly expressed genes from Escherichia coli were taken as an indicative measure of the translation speed. The tricodons were classified by the sum of the frequency values of the constituent codons. Examination of the conformation of the encoded amino acid residues in the corresponding protein tertiary structures revealed a correlation between codon usage in mRNA and topological features of the encoded proteins. Alpha helices on proteins tend to be preferentially coded by translationally fast mRNA regions while the slow segments often code for beta strands and coil regions. Fast regions correspondingly avoid coding for beta strands and coil regions while the slow regions similarly move away from encoding alpha helices. Structural and mechanistic aspects of the ribosome peptide channel support the relevance of sequence fragment translation and subsequent conformation. A discussion is presented relating the observation to the reported kinetic data on the formation and stabilization of protein secondary structural types during protein folding. The observed absence of such strong positive selection for codons in non-highly expressed genes is compatible with existing theories that mutation pressure may well dominate codon selection in non-highly expressed genes.

Bacterial Proteins↗

Ab initio tertiary-fold prediction of helical and non-helical protein chains using a genetic algorithm.

In this study, protein structures were predicted using a genetic algorithm. Optimal structures were selected according to basic protein-building principles. Having previously applied this approach to proteins with helical topology, this paper reports preliminary results on 3 proteins with mixed topology (helical and non-helical, e.g. sheet). The method seems to be generally applicable to protein-fold prediction since topology predictions close to the observed experimental structures are obtained.

Algorithms↗