Search PubMed⌕ Search

Biomedical subjects

P Argos

Publications and source records attributed to P Argos.

At least 73 records · Page 4Linked to original sources

Recognition of distantly related proteins through energy calculations.

A new method to detect remote relationships between protein sequences and known three-dimensional structures based on direct energy calculations and without reliance on statistics has been developed. The likelihood of a residue to occupy a given position on the structural template was represented by an estimate of the stabilization free energy made after explicit prediction of the substituted side chain conformation. The profile matrix derived from these energy values and modified by increasing the residue self-exchange values successfully predicted compatibility of heat-shock protein and globin sequences with the three-dimensional structures of actin and phycocyanin, respectively, from a full protein sequence databank search. The high sensitivity of the method makes it a unique tool for predicting the three-dimensional fold for the rapidly growing number of protein sequences.

Amino Acid Sequence↗

The protein folding problem: finding a few minimums in a near infinite space.

Folding a protein from only a knowledge of its amino acid sequence is a formidable many-body problem. Since it is computationally impossible to test all possible atomic conformations to determine the global minimum representing the compact state, methods need to be developed to sample only a small part of the configurational space and yet delineate the free energy optimum (or nearly so). This article largely reviews such techniques as applied by the authors and their colleagues.

Algorithms↗

Sensitive methods for determining the relatedness of proteins with limited sequence homology.

Recently, considerable advances have been made in attempts to determine the relatedness of protein sequences distant in evolution, when little or no knowledge is available concerning the corresponding tertiary architectures. Several improvements have been made to existing techniques, and these include better amino acid substitution weights contained in scoring matrices, better understanding of the effect of different gap penalty values in delineating the optimal alignment of two sequences, improved assessment of the significance of suggested sequence similarities, consideration of high scoring alternative alignments, and advances in searching entire sequence databases with the profile technique utilizing multiple-sequence information. New approaches that search for similarity a query sequence against large data banks rely on highly conserved segmental motifs defined from an aligned family of sequences. A sensitive algorithm to find distant repeats within one primary structure has also been developed recently. Solution of the inverse protein-folding problem, which involves an estimation of the ability of a sequence to take on a known main-chain tertiary topology (despite little homology with the known sequence), is being facilitated by the recent explosion in the number of new algorithms.

Amino Acid Sequence↗

Easy adaptation of protein structure to sequence.

An investigation into the conservation of coarse, medium and fine grain structural properties has been performed over a data set of 175 protein tertiary structures in 34 different families, each characterized by a common core fold and a library of conserved sites formed for each family. It is shown that, while the conservation of coarse and medium grain properties correlates to the structural deviation between the proteins, fine grain properties are poorly conserved except in functional sites. This flexibility in fine grain properties suggests that folding can be viewed as an optimization process whereby side chains have freedom to position themselves as best as possible given environmental conformational constraints and that given a basic framework, the local structure is able to adapt easily to sequence variation. The conserved cores of the 34 families are used to estimate a minimal core size of 35% of the fold, consistent with buried residue considerations. Finally, conservation in side chain chi 1 torsion angles is combined with structural deviation, sequence deviation and resolution to suggest a set of example structure pairs suitable for testing automatic homology modelling programs.

Hydrogen Bonding↗

Conservation of amphipathic conformations in multiple protein structural alignments.

Protein amphipathic conformations, mainly alpha-helices and beta-strands, are believed to play an important role in protein folding, stability and function. The most popular method for characterizing such structures is the hydrophobic moment. We have analyzed the distribution of hydrophobic moment characteristics (peak magnitude, amphipathic indices and characteristic frequency) in a data bank containing several families of distant sequences multiply aligned by structural superposition. Sequence fragments were classified according to alpha-helix, beta-strand, non-alpha and non-beta conformations. This data bank provided an enhanced sample space compared with those previously reported in the literature. Precautions were taken to reduce over-representation of homologous sequences. Approximately 50% of all individual alpha-helices showed a hydrophobic moment peak in the expected position of the periodicity spectrum while only 38% of individual beta-strands fell in the expected range. False positives account for a surprisingly large 14 and 36% of the non-alpha and non-beta samples respectively. Conservation of hydrophobic moment characteristics and mainly the hydrophobic peak position in the expected periodicity range was examined in the multiple alignments of the distant sequences. Helices tend to conserve more frequently their hydrophobic moment than any other conformation and yet only 13% of all helical segments display such conservation in three-quarters or more of the familial sequences; the similar observation for beta-strands was even lower at 9%. Nonetheless, strongly hydrophobic positions within the structural segments were more conserved than expected.

Amino Acid Sequence↗

Intramolecular cavities in globular proteins.

An analysis of internal cavities in 121 protein chains has been undertaken to improve the characterization of their occurrence, morphology and role in protein tertiary structure, including an analysis of the optimal probe size for use in their detection. A number of basic cavity characteristics were elucidated. Cavities are non-artefactual and apparently independent of the method of structure determination, resolution and refinement of the data. Overall cavity volume increases with protein size and yet constitutes only a small fraction of the total protein volume but cavities are nearly always present in proteins > 100 residues in size. They are most commonly found in the protein core. 'Empty' and solvent-containing cavities have been compared and solvated cavities found to possess a more polar surface; the two classes are also seen to exhibit different amino acid type and secondary structural preferences. In general, residues that enclose cavities do not display any extra local mobility relative to their surrounding environments. Water-containing cavities do not impose volume restrictions upon their internal solvent beyond that of bulk solvent and permit good hydrogen bonding. These results should prove useful in protein modelling and design.

Amino Acids↗

A method to configure protein side-chains from the main-chain trace in homology modelling.

Protein homology modelling typically involves the prediction of side-chain conformations in the modelled protein while assuming a main-chain trace taken from a known tertiary structure of a protein with homologous sequence. It is generally believed that the need to examine all possible combinations of side-chain conformations poses the major obstacle to accurate homology modelling. Methods proposed heretofore use only discrete or limited searches of the side-chain torsion angle space to mitigate the combinatorial problem and also rely on simplified energy functions for calculational speed. The configurational constraints are typically based upon use of frequently observed torsion angles, fixed steps in torsion angles, or oligopeptide segments taken from tertiary structural databanks that are similar in sequence and conformation with the target structure. In the present work, a more fundamental approach is explored for several protein structures and it is demonstrated that the combinatorial barrier in side-chain placement hardly exists. Each side-group can be configured individually in the environment of only the backbone atoms using a systematic search procedure combined with extensive local energy minimization. Tests, using the main-chain or both the main-chain and remaining side-chain atoms to calculate low energy geometries for each residue, established the dominance of the main-chain contribution. The final structure is achieved by combining the individually placed side-chains followed by a full energy refinement of the structure. The prediction accuracy of the present homology modelling technique was assessed relative to other automated procedures and was found to yield improved predictions relative to the known side-chain conformations determined by X-ray crystallography.

Amino Acid Sequence↗

Rotamers: to be or not to be? An analysis of amino acid side-chain conformations in globular proteins.

Originally, rotamers were defined as side-chain torsion (chi-angle) combinations corresponding to the local minima of potential energy (van-der-Waals and torsion terms) for the side-chain of a terminally blocked amino acid. If at least one chi-angle differed by more than 20 degrees from that of a rotamer, the side-chain was considered as deviant both from energetic (increase in potential energy of no less than 1 to 2 kcal/mol) and geometric (precision of atom positioning is worse than 0.5 A) aspects. In this work the distribution of side-chain conformations in protein crystal structures is analysed. Large deviations from rotameric chi-values occur systematically and cannot be attributed merely to errors in crystal structure determination. The "rotamericity" (the fraction of residues within +/- 20 degrees of the chi-angles of a rotamer) not only remains substantially below 100% (70 to 95% for various amino acids) with improving crystallographic resolution but actually decreases for 8 out of 17 amino acid types after a critical resolution limit is crossed. This effect has been observed for external as well as for internal residues. The set of amino acid side-chain conformations in globular proteins cannot be considered as normally distributed around some rotamer points. Outliers occur systematically. The rotamericity of an amino acid depends essentially on the different environments the amino acid meets in real protein structures. Factors such as the backbone torsion angles of the residue itself, the secondary structure and tertiary contacts influence the rotamericity. The deviations in regions of regular main-chain structure from the average g-:t:g+ relationship in the chi 1-angle become much more evident if, in addition to the typical secondary structure assignments, the actual backbone torsion angles of the residue are taken into account. In alpha-helices the t:g+ distribution in the chi 1-angle correlates with physical properties describing volume, extension and flexibility of the side-chain. In beta-strands the factors influencing the t:g+ distribution in the chi 1-angle are the polarity and hydrophobicity of the side-chain. Nevertheless, a considerable number of residues do not comply with the statistical preferences observed for the side-chain conformation. Large deviations from the rotamer values are observed especially in cases when normally advantageous chi 1-values are not allowed and adjustments in chi 2 become necessary to accommodate the side-chain.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acids↗

A method to recognize distant repeats in protein sequences.

An automated algorithm is presented that delineates protein sequence fragments which display similarity. The method incorporates a selection of a number of local nonoverlapping sequence alignments with the highest similarity scores and a graph-theoretical approach to elucidate the consistent start and end points of the fragments comprising one or more ensembles of related subsequences. The procedure allows the simultaneous identification of different types of repeats within one sequence. A multiple alignment of the resulting fragments is performed and a consensus sequence derived from the ensemble(s). Finally, a profile is constructed from the multiple alignment to detect possible and more distant members within the sequence. The method tolerates mutations in the repeats as well as insertions and deletions. The sequence spans between the various repeats or repeat clusters may be of different lengths. The technique has been applied to a number of proteins where the repeating fragments have been derived from information additional to the protein sequences.

Algorithms↗

Profile sequence analysis and database searches on a transputer machine connected to a Macintosh computer.

An implementation of Profilesearch (a technique to search for relationships between a protein sequence and multiply aligned sequences) for a parallel computer is described. The number-crunching machine, consisting of 21 T800 transputers, is connected to a Macintosh IIcx host computer. The program utilizes a standard Macintosh application as its user-interface, resulting in a transparent and user-friendly environment for addressing the parallel computer. The program is independent of the number of available processors and exceeds the speed of a VAXstation 3200 with only one transputer in operation, thus allowing cheap and fast database searches with a PC front-end. For a larger number of processors, the speed increase is approximately linear with no obvious symptoms of saturation with the available maximum of 21 transputers. The program and environment are useful to search quickly and easily for similarities between a single sequence or sequence set and individual sequences contained in a large database. The alignment is determined by typical dynamic programming techniques.

Amino Acid Sequence↗

SRS--an indexing and retrieval tool for flat file data libraries.

SRS (Sequence Retrieval System) is an information indexing and retrieval system designed for libraries with a flat file format such as the EMBL nucleotide sequence databank, the SwissProt protein sequence databank or the Prosite library of protein subsequence consensus patterns. SRS supports the data structure of these libraries by providing special indices for implementing lists of subentities (e.g. feature tables) or hierarchically structured data-fields (e.g. taxonomic classification). A language (ODD) has been designed for the convenient specification of library format and organization, representation of individual data-fields within the system (design of indices) and structuring other data needed during retrieval. This ensures flexibility required for coping with different library formats, which are subject to continuous change. Queries and inspection of retrieved entries can be performed from a user interface with pull-down menus and windows. SRS supports various input and output formats but is particularly well adapted to the GCG programs.

Abstracting and Indexing↗

Transforming a set of biological flat file libraries to a fast access network.

SRS (Sequence Retrieval System), an indexing system for flat file libraries, provides fast access to individual library entries via retrieval by keywords from various data fields. SRS is now also able to build indices using cross-references that most libraries provide. Fifteen libraries of DNA and protein sequences and structures have been selected. These libraries interact with at least one other by means of cross-references. Indexing these cross-references allows a complete network of libraries to be built. In the network an entry from one library can be linked in principle to every other library. If two libraries are not directly cross-referenced, the linkage can be made with a succession of single links between neighbouring, cross-referenced libraries. A new operator has been added to the query language of SRS for convenient specification of links amongst complete libraries or entry sets generated by previous queries on particular libraries. All the information in the network can now be used to retrieve an entry in a specific library, e.g. the full information given in amino acid sequence entries from SwissProt can now be used to retrieve related tertiary structure entries from PDB. Furthermore, a search in a single library can be extended to a search in the complete library network, e.g. all entries in all databases pertaining to elastase can be found.

Algorithms↗

Quantification of secondary structure prediction improvement using multiple alignments.

The use of multiple sequence alignments for secondary structure predictions is analysed. Seven different protein families, containing only sequences of known structure, were considered to provide a range of alignment and prediction conditions. Using alignments obtained by spatial superposition of main chain atoms in known tertiary protein structures allowed a mean of 8% in secondary structure prediction accuracy, when compared to those obtained from the individual sequences. Substitution of these alignments by those determined directly from an automated sequence alignment algorithm showed variations in the prediction accuracy which correlated with the quality of the multiple alignments and distance of the primary sequence. Secondary structure predictions can be reliably improved using alignments from an automatic alignment procedure with a mean increase of 6.8%, giving an overall prediction accuracy of 68.5%, if there is a minimum of 25% sequence identity between all sequences in a family.

Amino Acid Sequence↗

Recognition of distantly related protein sequences using conserved motifs and neural networks.

A sensitive technique for protein sequence motif recognition based on neural networks has been developed. It involves three major steps. (1) At each appropriate alignment position of a set of N matched sequences, a set of N aligned oligopeptides is specified with preselected window length. N neural nets are subsequently and successively trained on N-1 amino acid spans after eliminating each ith oligopeptide. A test for recognition of each of the ith spans is performed. The average neural net recognition over N such trials is used as a measure of conservation for the particular windowed region of the multiple alignment. This process is repeated for all possible spans of given length in the multiple alignment. (2) The M most conserved regions are regarded as motifs and the oligopeptides within each are used to train intensively M individual neural networks. (3) The M networks are then applied in a search for related primary structures in a databank of known protein sequences. The oligopeptide spans in the database sequence with strongest neural net output for each of the M networks are saved and then scored according to the output signals and the proper combination that follows the expected N- to C-terminal sequence order. The motifs from the database with highest similarity scores can then be used to retrain the M neural nets, which can be subsequently utilized for further searches in the databank, thus providing even greater sensitivity to recognize distant familial proteins. This technique was successfully applied to the integrase, DNA-polymerase and immunoglobulin families.

Aldehyde Dehydrogenase↗

Anatomy and evolution of proteins displaying the viral capsid jellyroll topology.

In this paper the anatomy of 25 structures containing a jellyroll motif, consisting of eight antiparallel beta-strands forming a so-called beta-barrel, was investigated. This involved performing a careful structural alignment based on hydrogen bonds for the equivalent regions of the tertiary folds and a subsequent analysis of conserved amino acids, equivalenced residue-residue contacts, and various parameters describing the size, shape and other geometrical characteristics of these regions. It was found that the jellyroll motif is best viewed as a two-sheet wedge structure rather than a barrel. The more conserved parameters are discussed. A model of evolutionary development for the jellyroll fold in the various protein and viral structures is proposed.

Amino Acid Sequence↗

Prediction of protein folding pathways.

Recent 1H nuclear magnetic resonance (n.m.r.) hydrogen exchange experiments on five different proteins have delineated the secondary structures formed in trapped, partially folded intermediates. The early forming structural elements are identifiable through a technique described in this work to predict folding pathways. The method assumes that the sequential selection of structural fragments such as alpha-helices and beta-strands involved in the folding process is founded upon the maximal burial of solvent accessible surface from both the formation of internal structure and substructure association. The substructural elements were defined objectively by major changes in main-chain direction. The predicted folding pathways are in complete correspondence with the n.m.r. results in that the formed structural fragments found in the folding intermediates are those predicted earliest in the pathways. The technique was also applied to proteins of known tertiary structure and with fold similar to one of the five proteins examined by 1H n.m.r. The pathways for these structures also showed general consistency with the n.m.r. observations, suggesting conservation of a secondary structural framework or molten globule about which folding nucleates and proceeds.

Amino Acid Sequence↗

Optimal protocol and trajectory visualization for conformational searches of peptides and proteins.

Conformational searches by molecular dynamics and different types of Monte Carlo or build-up methods usually aim to find the lowest-energy conformation. However, this is often misleading, as the energy functions used in conformational calculations are imprecise. For instance, though positions of local minima defined by the repulsive part of the Lennard-Jones potential are usually altered only slightly by functional modification, the relative depths of the minima could change significantly. Thus, the purpose of conformational searches and, correspondingly, performance criteria should be reformulated and appropriate methods found to extract different local minima from the search trajectory and allow visualization in the search space. Attempts at convergence to the lowest-energy structure should be replaced with efforts to visit a maximum number of different local energy minima with energies within a certain range. We use this quantitative criterion consistently to evaluate performances of different search procedures. To utilize information generated in the course of simulation, a "stack" of low energy conformations is created and stored. It keeps track of variables and visit numbers for the best representatives of different conformational families. To visualize the search, projection of multidimensional walks onto a principal plane defined by a set of reference structures is used. With Met-enkephalin as a structural example and a Monte Carlo procedure combined with energy minimization (MCM) as a basic search method, we analyzed the influence on search efficiency of different characteristics as temperature schedules, the step size for variable modification, constrained random step and response mechanisms to search difficulties. Simulated annealing MCM had comparable efficiency with MCM at constant and elevated temperature (about 600 K). Constraining the randomized choice of side-chain chi angles to optimal values (rotamers) on every MCM step did not improve, but rather worsened, the search efficiency. Two low-energy Met-enkephalin conformations with parallel Tyr1 and Phe4 rings, a gamma-turn around the Gly2 residue, and Phe4 and Met5 side-chains forming together a compact hydrophobic cluster were found and are suggested as possible structural candidates for interaction with a receptor or a membrane.

Amino Acid Sequence↗