Search PubMed⌕ Search

Biomedical subjects

P Argos

Publications and source records attributed to P Argos.

At least 91 records · Page 5Linked to original sources

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence↗

Overseer: a nucleotide sequence searching tool.

Overseer is a computer program that searches databases of nucleic acid sequences for objects of interest to the user. Such objects may consist of any number of simpler building blocks such as repeats, palindromes or stem-loops, strings of particular bases with or without mismatches, etc. Written in standard Pascal, this program runs under Unix and VMS and should also run under other operating systems. A simple interface allows the user to generate interactively a file containing a description of the target to be found. The searching program runs non-interactively, processing the information from the file and searching the sequences. The results are output to a file. Search capabilities are quite flexible and the code is designed to be modified. Since the framework of the program is simple, adding new modules to search for new target types as the need arises is possible.

Algorithms↗

Searching for distantly related protein sequences in large databases by parallel processing on a transputer machine.

AliMac is an implementation of a sensitive sequence alignment algorithm on a parallel computer. The method achieves reliable alignments for very distantly related sequences from a combined use of amino acid exchange weights and physicochemical characteristics. The algorithm is computing intensive and its usage on conventional computers is limited to a relatively small number of sequences. The parallel implementation uses a Macintosh IIcx host computer and 21 transputers and achieves 22 times the speed of a VAX 8650 at a fraction of the cost. This paper describes the AliMac hardware and software and discusses problems and peculiarities of parallel implementations, especially with transputers. Finally, several popular sequence alignment algorithms are compared in their ability to detect distantly related sequences in searching large databases.

Algorithms↗

OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity.

A program OBSTRUCT has been developed to obtain the largest possible subset according to specific constraints from a set of protein sequences whose tertiary structures have been determined crystallographically. The user can request a range in sequence similarity level and/or structural resolution. The program optionally includes sequences with known three-dimensional folds elicited from NMR data.

Protein Conformation↗

A data bank merging related protein structures and sequences.

A data collection which merges protein structural and sequence information is described. Structural superpositions amongst proteins with similar main-chain fold were performed or collected from the literature. Sequences taken from the protein primary structure databases were associated with the multiple structural alignments providing they were at least 50% homologous in residue identity to one of the structural sequences and at least 50% of the structural sequence residues were alignable. Such restrictions allow reasonable confidence that the primary sequences share the conformation of the tertiary structural templates, except in the less conserved loop regions. Multiple structural superpositions were collected for 38 familial groups containing a total of 209 tertiary structures; 45 structures had no superposable mates and were used individually. Other information is also provided as main-chain and side-chain conformational angles, secondary structural assignments and the like. Wedding the primary and tertiary structural data resulted in an 8-fold increase of data bank sequence entries over those associated with the known three-dimensional architectures alone.

Amino Acid Sequence↗

Potential of genetic algorithms in protein folding and protein engineering simulations.

Genetic algorithms are very efficient search mechanisms which mutate, recombine and select amongst tentative solutions to a problem until a near optimal one is achieved. We introduce them as a new tool to study proteins. The identification and motivation for different fitness functions is discussed. The evolution of the zinc finger sequence motif from a random start is modelled. User specified changes of the lambda repressor structure were simulated and critical sites and exchanges for mutagenesis identified. Vast conformational spaces are efficiently searched as illustrated by the ab initio folding of a model protein of a four beta strand bundle. The genetic algorithm simulation which mimicked important folding constraints as overall hydrophobic packaging and a propensity of the betaphilic residues for trans positions achieved a unique fold. Cooperativity in the beta strand regions and a length of 3-5 for the interconnecting loops was critical. Specific interaction sites were considerably less effective in driving the fold.

Algorithms↗

Identification of proteins in sequence databases from amino acid composition data.

Having obtained the amino acid composition of a protein, chemists and molecular biologists may wish to identify the protein from this data alone. In general such data will have errors associated with them and the length of the protein may be known only approximately or not at all. In this paper a method is described which enables searching of protein sequence databases for sequences or fragments of sequences which have a composition similar to the one being sought. Such searches are generally quite discriminating as shown by the examples provided. This method has been implemented as part of the computer program Scrutineer and is being freely distributed. It is simple to use.

Algorithms↗

Side-chain clusters in protein structures and their role in protein folding.

A method has been developed to detect dense clusters of residue side-chains in proteins, where contact is based upon the percentage of the maximum possible for a given residue type. The clusters represent protein sites with the highest degree of interaction amongst their member residues, while contacts with the environment surrounding the cluster are lower in number. The method has been applied to three distinct structural sets of proteins to check for consistency: mixed alpha-helical/beta-sheet proteins, all beta-strand proteins, and all alpha-helical proteins. A number of cluster features generated from these sets are of general interest for protein folding. (1) A majority of the clusters, comprising three to four residues on average, are localized near the protein surfaces and not within the protein cores. (2) The clusters have preferences for the N- and C-terminal ends of alpha-helices and beta-strands in alpha/beta and alpha-proteins, while beta-proteins utilize the middle strand regions more often. A number of clusters connect three or more beta-strands and/or alpha-helices. (3) More than half of the clusters display residue pairs with oppositely charged atoms within 4.5 A of each other. (4) The residue composition of the clusters does not show correlation with hydrophobicity measures but rather with side-chain volume and surface. The highly preferred cluster residues are (in order of decreasing preference) Trp, His, Arg, Tyr, Glu, Gln and Phe. Clusters with extensive internal contacts in related haemoglobin and immunoglobulin tertiary structures show respective conservation. Several examples illustrate "strategic" folding positions in proteins that often bring together a number of sheets and/or helices, suggesting a folding model in which largely preformed secondary structures are joined together in a cluster induced collapse. Alternatively, the clusters may form at some stage in the folding process to reduce considerably the searchable conformational space and help maintain the proper folding pathway. The clusters also provide hints for site-directed mutagenesis and protein engineering experiments as they are also suggested to be important for structural stability.

Amino Acid Sequence↗

Homology between IRE-BP, a regulatory RNA-binding protein, aconitase, and isopropylmalate isomerase.

Iron-responsive elements (IREs) are regulatory RNA elements which serve as specific binding sites for the IRE-binding protein (IRE-BP). Interaction between IREs and IRE-BP induces repression of ferritin mRNA translation and transferrin receptor mRNA stabilization. We describe the identification of extensive amino acid sequence homology between IRE-BP and two known isomerases, aconitase and isopropylmalate (IPM) isomerase. We discuss the implications of this observation with regard to structure/function relationships of IRE-BP. The structural conservation between a regulatory RNA-binding protein and two enzymes involved in intermediary metabolism provides a surprising example of the functional flexibility in biological structures.

Aconitate Hydratase↗

Motif recognition and alignment for many sequences by comparison of dot-matrices.

Calculation of dot-matrices is a widespread tool in the search for sequence similarities. When sequences are distant, even this approach may fail to point out common regions. If several plots calculated for all members of a sequence set consistently displayed a similarity between them, this would increase its credibility. We present an algorithm to delineate dot-plot agreement. A novel procedure based on matrix multiplication is developed to identify common patterns and reliably aligned regions in a set of distantly related sequences. The algorithm finds motifs independent of input sequence lengths and reduces the dependence on gap penalties. When sequences share greater similarity, the same approach converts to a multiple sequence alignment procedure.

Algorithms↗

Suggestions for "safe" residue substitutions in site-directed mutagenesis.

The conserved topological structure observed in various molecular families such as globins or cytochromes c allows structural equivalencing of residues in every homologous structure and defines in a coherent way a global alignment in each sequence family. A search was performed for equivalent residue pairs in various topological families that were buried in protein cores or exposed at the protein surface and that had mutated but maintained similar unmutated environments. Amino acid residues with atoms in contact with the mutated residue pairs defined the environment. Matrices of preferred amino acid exchanges were then constructed and preferred or avoided amino acid substitutions deduced. Given the conserved atomic neighborhoods, such natural in vivo substitutions are subject to similar constrains as point mutations performed in site-directed mutagenesis experiments. The exchange matrices should provide guidelines for "safe" amino acid substitutions least likely to disturb the protein structure, either locally or in its overall folding pathway, and most likely to allow probing the structural and functional significance of the substituted site.

Amino Acid Sequence↗

Beta-COP, a 110 kd protein associated with non-clathrin-coated vesicles and the Golgi complex, shows homology to beta-adaptin.

We have cloned and sequenced beta-COP, a peripheral 110 kd Golgi membrane protein. beta-COP shows significant homology to beta-adaptin. It is present in a membrane-bound form and in a cytosolic complex of 13-14S, with a Stokes radius of approximately 10 nm and an estimated Mr of approximately 550,000. By immunofluorescence labeling, beta-COP is associated with the structures of the Golgi complex. Immunoelectron microscopy has localized beta-COP to non-clathrin-coated vesicles and cisternae of the Golgi complex. These coated vesicles accumulate in rat liver Golgi fractions treated with GTP gamma S and strongly label for beta-COP. Our data suggest that beta-COP is a component of a coat associated with vesicles and cisternae of the Golgi complex.

Amino Acid Sequence↗

Automated protein sequence pattern handling and PROSITE searching.

The protein sequence searching program Scrutineer has been modified to search for targets from a file. We are distributing a reformatted file of PROSITES which can be read by Scrutineer. In addition, Scrutineer still accepts targets typed in interactively but can now write them out in the format required as input. Since the input format is the same as the output format, target management and re-use is simple.

Amino Acid Sequence↗

Weighting aligned protein or nucleic acid sequences to correct for unequal representation.

Aligned sequences from the same family (e.g. the haemoglobins) are seldom representative of the entire family. This is because (1) the sequence databases are heavily skewed toward a small number of organisms and (2) only a minute fraction of all the different family members have been sequenced. For many applications, such as using alignments or profiles to perform database searches for distantly related family members, such unequal representation requires correction. An algorithm to perform appropriate weighting of individual sequences is presented along with examples illustrating its efficacy.

Algorithms↗

Protruding domain of tomato bushy stunt virus coat protein is a hitherto unrecognized class of jellyroll conformation.

The capsid protein of tomato bushy stunt virus (TBSV) has two antiparallel beta-sheet domains with the so-called jellyroll conformation. Contrary to previous analyses, we note that these domains are non-superimposable topologies. The TBSV shell (S) domain topology is common to many other proteins but the protruding (P) domain is a unique conformation so far found in no other protein. The TBSV capsid P domain did not arise from the S domain by a gene duplication event as previously assumed. It is proposed instead that the P domain was acquired from an as yet unidentified cellular protein. The four possible unique jellyroll topologies that might occur in proteins are discussed and illustrated.

Capsid↗

An investigation of oligopeptides linking domains in protein tertiary structures and possible candidates for general gene fusion.

Fifty-one examples of oligopeptides linking protein domains were extracted from the Brookhaven database of three-dimensional protein structures. In general, the peptides displayed specific characteristics in composition, conformation, hydrogen bonding, flexibility and the like. The entire database was then searched for pentapeptides that would optimize these natural linker properties. The oligopeptides found are suggested as general candidates to link protein molecules or domains through gene fusion.

Amino Acid Sequence↗

Evolution of protein cores. Constraints in point mutations as observed in globin tertiary structures.

The amino acid sequences of ten globin chain tertiary structures were aligned and structurally equivalenced by spatial superposition of main-chain C alpha atoms. A search was then performed for structurally equivalent residue pairs that were buried in the protein core and that had mutated but maintained similar unmutated environments. Residues with atoms in contact with such central residue pairs define their environments. Such examples of point mutations would represent in vivo site-directed mutagenesis as would be observed in evolution. A search for mutated but exposed equivalent central residues was also performed. The constraints placed on the characteristics of the mutated residues (e.g., side-chain volume, polarity, radius of gyration) allow suggestions for the evolutionary modes of protein core and surface development as well as residue substitution guidelines to maintain structural stability in protein engineering and design.

Amino Acid Sequence↗