Unique protein database imperiled.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Currently available sequence alignment programs are generally not capable of detecting functional and structural homologs in the twilight zone of sequence similarity, i.e. when the sequence identity falls below about 25%. Here we attempt to detect such weak similarities using an approach based on a notion of protein sequence similarity radically different from that used in sequential alignment. The approach defines protein sequence dissimilarity (or distance) as a weighted sum of differences of compositional properties such as singlet and doublet amino acid composition, molecular weight, isoelectric point (protein property search or PropSearch). With PropSearch, either single sequences can be used for a database query, or multiple sequences can be merged into an "average" sequence reflecting the average composition of a protein family. First, we show that members of structural protein families have a low mutual PropSearch distance when the weights are optimized to discriminate maximally between structural families. Second, we demonstrate the results of database searches using the PropSearch method. Such searches are very rapid when scanning a preprocessed database and do not require alignments. In cases in which conventional alignment tools fail to detect similarities, PropSearch can be used to generate hypotheses about possible structural or functional relationships between a new sequence and sequences in the database.
We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.
Explore the source record for details and available documents.
Previously, we discussed the calculation of the dipole moments of small proteins using the three-dimensional protein data-base. Our results demonstrate that the calculated dipole moments are in acceptable agreement with measured values. We, however, noted the difficulty of the calculation with larger proteins, in particular those consisting of several subunits. Hemoglobin (Hb) is a protein having a molecular weight of 64,000 that consists of four subunits, a typical case where the computation was found to be difficult. To circumvent the difficulties, we calculated the dipole moment of each subunit separately. The dipole moment of the whole protein was calculated by the vectorial summation of subunit moments. With this method, the calculated net dipole moment is in good agreement with the experimental value. Our calculation shows that the dipole moment vectors of subunits are, by and large, antiparallel in tetramers causing partial cancellation of the net dipole moment. In addition to normal HbA, the dipole moment of abnormal HbS was calculated using an approximate computational technique. Because of the loss of two negative changes as a result of the replacement of glutamic acid with valine in beta-chains, the dipole moment of HbS was found, experimentally and theoretically, to be significantly smaller than that of HbA.
MOTIVATION: Optimal sequence alignment based on the Smith-Waterman algorithm is usually too computationally demanding to be practical for searching large sequence databases. Heuristic programs like FASTA and BLAST have been developed which run much faster, but at the expense of sensitivity. RESULTS: In an effort to approximate the sensitivity of an optimal alignment algorithm, a new algorithm has been devised for the computation of a gapped alignment of two sequences. After scanning for high-scoring words and extensions of these to form fragments of similarity, the algorithm uses dynamic programming to build an accurate alignment based on the fragments initially identified. The algorithm has been implemented in a program called SALSA and the performance has been evaluated on a set of test sequences. The sensitivity was found to be close to the Smith-Waterman algorithm, while the speed was similar to FASTA (ktup = 2). AVAILABILITY: Searches can be performed from the SALSA homepage at http://dna.uio.no/salsa/ using a wide range of databases. Source code and precompiled executables are also available. CONTACT: torbjorn.rognes@labmed.uio.no
An improved method of high-resolution two-dimensional gel electrophoresis has been used to study the patterns of protein synthesis in wing imaginal discs of late instar larvae of Drosophila melanogaster. A total of one thousand and twenty five labelled polypeptides (787 acidic and 238 basic) have so far been separated and catalogued. For convenience, all these polypeptides have been numbered and their position fixed by its molecular weight and relative mobility. They are indicated on a reference protein map for further studies.
During the past year, the number of biological databases that can be queried via Internet has dramatically increased. This increase has resulted from the introduction of networking tools, such as Gopher and WAIS, that make it easy for research workers to index databases and make them available for on-line browsing. Biocomputing in the nineties will see the advent of more client/server options for the solution of problems in bioinformatics.
ESTHER (for esterases, alpha/betahydrolase enzyme and relatives) is a database of sequences phylogenetically related to cholinesterases. These sequences define a homogeneous group of enzymes (carboxylesterases, lipases and hormone-sensitive lipases) sharing a similar structure of a central beta-sheet surrounded by alpha-helices. Among these proteins a wide range of functions can be found (hydrolases, adhesion molecules, hormone precursors). The purpose of ESTHER is to help comparison of structures and functions of members of the family. Since the last release, new features have been added to the server. A BLAST comparison tool allows sequence homology searches within the database sequences. New sections are available: kinetics and inhibitors of cholinesterases, fasciculin-acetylcholinesterase interaction and a gene structure review. The mutation analysis compilation has been improved with three-dimensional images. A mailing list has been created.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).
The Protein Information Resource (PIR; http://www-nbrf.georgetown. edu/pir/) supports research on molecular evolution, functional genomics, and computational biology by maintaining a comprehensive, non-redundant, well-organized and freely available protein sequence database. Since 1988 the database has been maintained collaboratively by PIR-International, an international association of data collection centers cooperating to develop this resource during a period of explosive growth in new sequence data and new computer technologies. The PIR Protein Sequence Database entries are classified into superfamilies, families and homology domains, for which sequence alignments are available. Full-scale family classification supports comparative genomics research, aids sequence annotation, assists database organization and improves database integrity. The PIR WWW server supports direct on-line sequence similarity searches, information retrieval, and knowledge discovery by providing the Protein Sequence Database and other supplementary databases. Sequence entries are extensively cross-referenced and hypertext-linked to major nucleic acid, literature, genome, structure, sequence alignment and family databases. The weekly release of the Protein Sequence Database can be accessed through the PIR Web site. The quarterly release of the database is freely available from our anonymous FTP server and is also available on CD-ROM with the accompanying ATLAS database search program.
Databases of protein information from human embryonal lung fibroblasts (MRC-5) have been established using computer analyzed two-dimensional gel electrophoresis. One thousand four hundred and eighty-two cellular proteins (1060 with isoelectric focusing and 422 with nonequilibrium pH gradient electrophoresis, in the first dimension) ranging in molecular mass between 8 and 234 kDa were separated and numbered. Information entered in the database (in most cases for major proteins) includes: protein name, HeLa protein catalog number, mouse protein catalog number, proteins matched in transformed human epithelial amnion cells (AMA) and peripheral blood mononuclear cells (PBMC), transformation and/or proliferation sensitive proteins, synthesis in quiescent cells, cell cycle regulated proteins, mitochondrial and heat shock proteins, cytoskeletal proteins and proteins whose synthesis is affected by interferons. Additional information entered for a few transformation-sensitive proteins that have been selected for future studies includes levels of synthesis and amounts in fetal human tissues. A total of four hundred and seventy-six [35S]methionine labeled polypeptides (258 isoelectric focusing; 218, nonequilibrium pH gradient electrophoresis) secreted by MRC-5 fibroblasts were separated and recorded (J. E. Celis et al., Leukemia 1987, 1, 707-717). Information entered in this database includes molecular weight and transformation sensitive proteins. These databases, as well as those of epithelial and lymphoid cell proteins (J. E. Celis et al., Leukemia 1988, 9, 561-601), represent the initial stages of a systematic effort to establish comprehensive databases of human protein information. In the long run, these databases are expected to offer a useful framework in which to focus the human genome sequencing effort.
We describe a database of protein structure alignments as well as methods and tools that use this database to improve comparative protein modeling. The current version of the database contains 105 alignments of similar proteins or protein segments. The database comprises 416 entries, 78,495 residues, 1,233 equivalent entry pairs, and 230,396 pairs of equivalent alignment positions. At present, the main application of the database is to improve comparative modeling by satisfaction of spatial restraints implemented in the program MODELLER (Sali A, Blundell TL, 1993, J Mol Biol 234:779-815). To illustrate the usefulness of the database, the restraints on the conformation of a disulfide bridge provided by an equivalent disulfide bridge in a related structure are derived from the alignments; the prediction success of the disulfide dihedral angle classes is increased to approximately 80%, compared to approximately 55% for modeling that relies on the stereochemistry of disulfide bridges alone. The second example of the use of the database is the derivation of the probability density function for comparative modeling of the cis/trans isomerism of the proline residues; the prediction success is increased from 0% to 82.9% for cis-proline and from 93.3% to 96.2% for trans-proline. The database is available via electronic mail.
Explore the source record for details and available documents.
Explore the source record for details and available documents.