Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “ARGOS”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

A fast and sensitive multiple sequence alignment algorithm.

A two-step multiple alignment strategy is presented that allows rapid alignment of a set of homologous sequences and comparison of pre-aligned groups of sequences. Examples are given demonstrating the improvement in the quality of alignments when comparing entire groups instead of single sequences. The modular design of computer programs based on this algorithm allows for storage of aligned sequences and successive alignment of any number of sequences.

Algorithms↗

Scrutineer: a computer program that flexibly seeks and describes motifs and profiles in protein sequence databases.

Scrutineer is an interactive, user-friendly program designed to search for motifs, patterns and profiles in the Swissprot, Protein Identification Resource (PIR) or SeqDb protein sequence databases. Basic capabilities include (i) searches for strings of amino acids with multiple choices at a given position; (ii) searches for strings including variable-length segments and delocalized constraints; (iii) searches over subsets of a database or particular regions within each sequence (e.g. N-terminal one-third); (iv) searches involving secondary structure predictions, physicochemical characteristics, and the like; and (v) searches using aligned sequences as targets with various optional weighting schemes. The various search criteria and hits can be combined and complex targets located. Once the data are loaded into virtual memory, all occurrences in PIR release 22.0 (3.7 x 10(6) amino acids) of a given short string of amino acids (e.g. a hexamer) are found in approximately 36 s. Scrutineer can also describe the entire database, user-specified hits, user-defined regions of sequence and all hits. The source code and accompanying manual are being freely distributed.

Algorithms↗

Automated protein sequence pattern handling and PROSITE searching.

The protein sequence searching program Scrutineer has been modified to search for targets from a file. We are distributing a reformatted file of PROSITES which can be read by Scrutineer. In addition, Scrutineer still accepts targets typed in interactively but can now write them out in the format required as input. Since the input format is the same as the output format, target management and re-use is simple.

Amino Acid Sequence↗

Overseer: a nucleotide sequence searching tool.

Overseer is a computer program that searches databases of nucleic acid sequences for objects of interest to the user. Such objects may consist of any number of simpler building blocks such as repeats, palindromes or stem-loops, strings of particular bases with or without mismatches, etc. Written in standard Pascal, this program runs under Unix and VMS and should also run under other operating systems. A simple interface allows the user to generate interactively a file containing a description of the target to be found. The searching program runs non-interactively, processing the information from the file and searching the sequences. The results are output to a file. Search capabilities are quite flexible and the code is designed to be modified. Since the framework of the program is simple, adding new modules to search for new target types as the need arises is possible.

Algorithms↗

Searching for distantly related protein sequences in large databases by parallel processing on a transputer machine.

AliMac is an implementation of a sensitive sequence alignment algorithm on a parallel computer. The method achieves reliable alignments for very distantly related sequences from a combined use of amino acid exchange weights and physicochemical characteristics. The algorithm is computing intensive and its usage on conventional computers is limited to a relatively small number of sequences. The parallel implementation uses a Macintosh IIcx host computer and 21 transputers and achieves 22 times the speed of a VAX 8650 at a fraction of the cost. This paper describes the AliMac hardware and software and discusses problems and peculiarities of parallel implementations, especially with transputers. Finally, several popular sequence alignment algorithms are compared in their ability to detect distantly related sequences in searching large databases.

Algorithms↗

OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity.

A program OBSTRUCT has been developed to obtain the largest possible subset according to specific constraints from a set of protein sequences whose tertiary structures have been determined crystallographically. The user can request a range in sequence similarity level and/or structural resolution. The program optionally includes sequences with known three-dimensional folds elicited from NMR data.

Protein Conformation↗

Profile sequence analysis and database searches on a transputer machine connected to a Macintosh computer.

An implementation of Profilesearch (a technique to search for relationships between a protein sequence and multiply aligned sequences) for a parallel computer is described. The number-crunching machine, consisting of 21 T800 transputers, is connected to a Macintosh IIcx host computer. The program utilizes a standard Macintosh application as its user-interface, resulting in a transparent and user-friendly environment for addressing the parallel computer. The program is independent of the number of available processors and exceeds the speed of a VAXstation 3200 with only one transputer in operation, thus allowing cheap and fast database searches with a PC front-end. For a larger number of processors, the speed increase is approximately linear with no obvious symptoms of saturation with the available maximum of 21 transputers. The program and environment are useful to search quickly and easily for similarities between a single sequence or sequence set and individual sequences contained in a large database. The alignment is determined by typical dynamic programming techniques.

Amino Acid Sequence↗

SRS--an indexing and retrieval tool for flat file data libraries.

SRS (Sequence Retrieval System) is an information indexing and retrieval system designed for libraries with a flat file format such as the EMBL nucleotide sequence databank, the SwissProt protein sequence databank or the Prosite library of protein subsequence consensus patterns. SRS supports the data structure of these libraries by providing special indices for implementing lists of subentities (e.g. feature tables) or hierarchically structured data-fields (e.g. taxonomic classification). A language (ODD) has been designed for the convenient specification of library format and organization, representation of individual data-fields within the system (design of indices) and structuring other data needed during retrieval. This ensures flexibility required for coping with different library formats, which are subject to continuous change. Queries and inspection of retrieved entries can be performed from a user interface with pull-down menus and windows. SRS supports various input and output formats but is particularly well adapted to the GCG programs.

Abstracting and Indexing↗

Transforming a set of biological flat file libraries to a fast access network.

SRS (Sequence Retrieval System), an indexing system for flat file libraries, provides fast access to individual library entries via retrieval by keywords from various data fields. SRS is now also able to build indices using cross-references that most libraries provide. Fifteen libraries of DNA and protein sequences and structures have been selected. These libraries interact with at least one other by means of cross-references. Indexing these cross-references allows a complete network of libraries to be built. In the network an entry from one library can be linked in principle to every other library. If two libraries are not directly cross-referenced, the linkage can be made with a succession of single links between neighbouring, cross-referenced libraries. A new operator has been added to the query language of SRS for convenient specification of links amongst complete libraries or entry sets generated by previous queries on particular libraries. All the information in the network can now be used to retrieve an entry in a specific library, e.g. the full information given in amino acid sequence entries from SwissProt can now be used to retrieve related tertiary structure entries from PDB. Furthermore, a search in a single library can be extended to a search in the complete library network, e.g. all entries in all databases pertaining to elastase can be found.

Algorithms↗

Similarity in gene organization and homology between proteins of animal picornaviruses and a plant comovirus suggest common ancestry of these virus families.

The amino acid sequences deduced from the nucleic acid sequences of several animal picornaviruses and cowpea mosaic virus (CPMV), a plant virus, were compared. Good homology was found between CPMV and the picornaviruses in the region of the picornavirus 2C (P2-X protein), VPg, 3C pro (proteinase) and 3D pol (RNA polymerase) regions. The CPMV B genome was found to have a similar gene organization to the picornaviruses. A comparison of the 3C pro (proteinase) regions of all of the available picornavirus sequences and CPMV allowed us to identify residues that are completely conserved; of these only two residues, Cys-147 and His-161 (poliovirus proteinase) could be the reactive residues of the active site of a proteinase with analogous mechanism to a known proteinase. We conclude that the proteinases encoded by these viruses are probably cysteine proteinases, mechanistically related, but not homologous to papain.

Amino Acid Sequence↗

Primary structural comparison of RNA-dependent polymerases from plant, animal and bacterial viruses.

Possible alignments for portions of the genomic codons in eight different plant and animal viruses are presented: tobacco mosaic, brome mosaic, alfalfa mosaic, sindbis, foot-and-mouth disease, polio, encephalomyocarditis, and cowpea mosaic viruses. Since in one of the viruses (polio) the aligned sequence has been identified as an RNA-dependent polymerase, this would imply the identification of the polymerases in the other viruses. A conserved fourteen-residue segment consisting of an Asp-Asp sequence flanked by hydrophobic residues has also been found in retroviral reverse transcriptases, a bacteriophage, influenza virus, cauliflower mosaic virus and hepatitis B virus, suggesting this span as a possible active site or nucleic acid recognition region for the polymerases. Evolutionary implications are discussed.

Animals↗

The primary structure of human hemopexin deduced from cDNA sequence: evidence for internal, repeating homology.

We have cloned and analyzed a cDNA containing the coding sequence for human hemopexin. We have first identified, by immunological screening of 30.000 colonies of a liver cDNA library in the expression vector pEX1, a clone carrying an insert 1170 base pairs long that shows 100% homology with a known human hemopexin peptide. The complete sequence coding for hemopexin was isolated from a liver cDNA library in the vector pAT218. The DNA insert of 1523 base pairs shows an open reading frame coding for 439 amino acids, a 3' noncoding region of 159 nucleotides long, followed by a poly(A) tail. The insert spans the entire coding region and from which the primary structure of the protein was deduced. By computer assisted analysis of the amino acid sequence, it was possible to recognize a core unit, of about 45 amino acids, which is repeated 8 or possibly even 10 fold along the polypeptide chain. This feature suggests that the gene might have evolved through a series of duplications. This characteristic, together with prediction of secondary structure, suggest a rough model for the tridimensional folding that allows some speculations on the function of hemopexin. Blot hybridization of total RNA from human liver with nick translated hemopexin cDNA detected a message of about 1600 nucleotides. Southern blot experiments to identify the hemopexin gene (s) suggest that it is not a large multi-gene family, but that there is only one or at most a few genes in the human genome.

Amino Acid Sequence↗

A sequence motif in many polymerases.

A 15-residue sequence motif has been found in many polymerases from various species and involving DNA and RNA dependence and product. The motif is characterized by a Tyr-Gly-Asp-(Thr)-Asp core flanked by hydrophobic spans five residues in length. An mRNA maturase segment is also suggested to display the motif pattern. The aspartates may be important in polymerase function by acting directly in catalysis and/or by binding magnesium.

Amino Acid Sequence↗

Homology between IRE-BP, a regulatory RNA-binding protein, aconitase, and isopropylmalate isomerase.

Iron-responsive elements (IREs) are regulatory RNA elements which serve as specific binding sites for the IRE-binding protein (IRE-BP). Interaction between IREs and IRE-BP induces repression of ferritin mRNA translation and transferrin receptor mRNA stabilization. We describe the identification of extensive amino acid sequence homology between IRE-BP and two known isomerases, aconitase and isopropylmalate (IPM) isomerase. We discuss the implications of this observation with regard to structure/function relationships of IRE-BP. The structural conservation between a regulatory RNA-binding protein and two enzymes involved in intermediary metabolism provides a surprising example of the functional flexibility in biological structures.

Aconitate Hydratase↗

Correlation between side chain mobility and conformation in protein structures.

Thermal factors of protein atoms as determined by X-ray crystallographic techniques show a tendency to be larger in side chains with unfavourable local conformations rather than in those displaying conformational energy minima. It follows that side chain atoms are more mobile if they are in a non-rotameric configuration and that the stereochemistry of protein structures cannot be fully assessed or simulated without consideration of thermal factors that monitor flexibility in various regions of the protein. The observations should also prove useful in protein folding and design.

Amino Acids↗

Applying experimental data to protein fold prediction with the genetic algorithm.

Specific residue interactions as revealed from a few and readily available experiments can be quite important in shaping a protein's tertiary topology by complementing basic and general folding principles. This experimental information is employed in structure prediction (mainchain topology) based on sequence knowledge and the genetic algorithm with its ability to optimize simultaneously many parameters. Examples investigated include the distribution of cysteinyl S-S bonds, protein side-chain ligands to iron-sulfur cages, cofactor-ligands, crosslinks amongst side-chains, and conserved hydrophobic and catalytic residues. Such interactions yield an improvement in the predicted topology (0.4-6.6 A root mean square deviation in the positions of the backbone C alpha-atoms relative to those observed) compared with those resulting from simulations relying only on basic protein folding principles. For several examples the resultant topology depended critically on knowledge of the few and specific interactions such that the relationship between predicted and observed C alpha-positions was near random without their use. The combined methodology (experimental data and the genetic algorithm) should prove helpful in settings where experiment and theory can cooperate in successive steps to elucidate an unknown structure.

Algorithms↗

Structural adaptation of enzymes to low temperatures.

A systematic comparative analysis of 21 psychrophilic enzymes belonging to different structural families from prokaryotic and eukaryotic organisms is reported. The sequences of these enzymes were multiply aligned to 427 homologous proteins from mesophiles and thermophiles. The net flux of amino acid exchanges from meso/thermophilic to psychrophilic enzymes was measured. To assign the observed preferred exchanges to different structural environments, such as secondary structure, solvent accessibility and subunit interfaces, homology modeling was utilized to predict the secondary structure and accessibility of amino acid residues for the psychrophilic enzymes for which no experimental three-dimensional structure is available. Our results show a clear tendency for the charged residues Arg and Glu to be replaced at exposed sites on alpha-helices by Lys and Ala, respectively, in the direction from 'hot' to 'cold' enzymes. Val is replaced by Ala at buried regions in alpha-helices. Compositional analysis of psychrophilic enzymes shows a significant increase in Ala and Asn and a decrease in Arg at exposed sites. Buried sites in beta-strands tend to be depleted of VAL: Possible implications of the observed structural variations for protein stability and engineering are discussed.

Alanine↗

An investigation of protein subunit and domain interfaces.

Protein structures were collected from the Brookhaven Database of tertiary architectures that displayed oligomeric association (24 molecules) or whose polypeptide folding revealed domains (34 proteins). The subunit and domain interfaces for these proteins were respectively examined from the following aspects: percentage water-accessible surface area buried by the respective associations, surface compositions and physical characteristics of the residues involved in the subunit and domain contacts, secondary structural state of the interface amino acids, preferred polar and non-polar interactions, spatial distribution of polar and non-polar residues on the interface surface, same residue interactions in the oligomeric contacts, and overall cross-section and shape of the contact surfaces. A general, consistent picture emerged for both the domain and subunit interfaces.

Amino Acids↗