Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

The DExH/D protein family database.

DExH/D proteins are essential for all aspects of cellular RNA metabolism and processing, in the replication of many viruses and in DNA replication. DExH/D proteins are subject to current biological, biochemical and biophysical research which provides a continuous wealth of data. The DExH/D protein family database compiles this information and makes it available over the WWW (http://www.columbia.edu/ ej67/dbhome.htm ). The database can be fully searched by text based queries, facilitating fast access to specific information about this important class of enzymes.

Amino Acid Sequence↗

A protein class database organized with ProSite protein groups and PIR superfamilies.

A protein class (ProClass) database is developed as a "value-added" "second-generation" database organized according to family relationships. The database collects non-redundant protein sequence entries from SwissProt and PIR databases, and classifies them in families defined collectively by the ProSite protein groups and PIR superfamilies. The major objectives of the database are to maximize family information retrieval, to provide speedy family identification, and to help organizing existing protein sequence databases. The database has two sub-databases: PCFam (ProClass Family) to define protein families and provide links to ProSite patterns and PIR superfamilies, and PCSeq (ProClass Sequence) to describe sequence entries and provide links to PCFam, SwissProt, PIR, and ProSite databases. The current ProClass release has a total of 85,165 sequence entries, about half of which are classified in 3072 ProClass families; it also contains 10,431 newly established SwissProt-PIR links. The database can help reveal domain structures of related families, define new ProSite and PIR families, and provide family assignments for unclassified sequence entries. New ProSite and PIR family members are readily identified via database cross-reference, including 9437 SwissProt entries and 8522 PIR entries. False negative family members missed by both ProSite and PIR are detected using a neural network family identification system. The newly identified superfamily memberships are being incorporated into the current PIR database releases in a collaborative effort with the PIR. The ProClass database is accessible through anonymous FTP and on-line search on the World Wide Web.

Amino Acid Sequence↗

Percolation of annotation errors through hierarchically structured protein sequence databases.

Databases of protein sequences have grown rapidly in recent years as a result of genome sequencing projects. Annotating protein sequences with descriptions of their biological function ideally requires careful experimentation, but this work lags far behind. Instead, biological function is often imputed by copying annotations from similar protein sequences. This gives rise to annotation errors, and more seriously, to chains of misannotation. [Percolation of annotation errors in a database of protein sequences (2002)] developed a probabilistic framework for exploring the consequences of this percolation of errors through protein databases, and applied their theory to a simple database model. Here we apply the theory to hierarchically structured protein sequence databases, and draw conclusions about database quality at different levels of the hierarchy.

Amino Acid Sequence↗

Circularly permuted proteins in the protein structure database.

Some proteins are homologous to others after their sequence is circularly permuted. A few such proteins have been recognized, mainly by sequence comparison, but also by comparing their three-dimensional structures. Here we report the result of a systematic search for all protein pairs in the SCOP 90% id domain database that become structurally superimposable when the sequence of one of the pairs is circularly permuted. Using a reasonable set of criteria, we find that 47% of all protein domains are superimposable to at least one other protein domain in the database after their sequence is circularly permuted. Many of these are symmetric proteins, which superimpose to another protein both with and without a circular permutation of the sequence. However, 412 of the total 3035 domains are nonsymmetric, and these become structurally superimposable to another protein only after a circular permutation of the sequence. These include most known and many previously undetected circularly permuted proteins with remote homology.

Amino Acid Sequence↗

New algorithmic approaches to protein spot detection and pattern matching in two-dimensional electrophoresis gel databases.

Protein spot identification in two-dimensional electrophoresis gels can be supported by the comparison of gel images accessible in different World Wide Web two-dimensional electrophoresis (2-DE) gel protein databases. The comparison may be performed either by visual cross-matching between gel images or by automatic recognition of similar protein spot patterns. A prerequisite for the automatic point pattern matching approach is the detection of protein spots yielding the x(s),y(s) coordinates and integrated spot intensities i(s). For this purpose an algorithm is developed based on a combination of hierarchical watershed transformation and feature extraction methods. This approach reduces the strong over-segmentation of spot regions normally produced by watershed transformation. Measures for the ellipticity and curvature are determined as features of spot regions. The resulting spot lists containing x(s),y(s),i(s)-triplets are calculated for a source as well as for a target gel image accessible in 2-DE gel protein databases. After spot detection a matching procedure is applied. Both the matching of a local pattern vs. a full 2-DE gel image and the global matching between full images are discussed. Preset slope and length tolerances of pattern edges serve as matching criteria. The local matching algorithm relies on a data structure derived from the incremental Delaunay triangulation of a point set and a two-step hashing technique. For the incremental construction of triangles the spot intensities are considered in decreasing order. The algorithm needs neither landmarks nor an a priori image alignment. A graphical user interface for spot detection and gel matching is written in the Java programming language for the Internet. The software package called CAROL (http://gelmatching.inf.fu-berlin.de) is realized in a client-server architecture.

Algorithms↗

MINT: a Molecular INTeraction database.

Protein interaction databases represent unique tools to store, in a computer readable form, the protein interaction information disseminated in the scientific literature. Well organized and easily accessible databases permit the easy retrieval and analysis of large interaction data sets. Here we present MINT, a database (http://cbm.bio.uniroma2.it/mint/index.html) designed to store data on functional interactions between proteins. Beyond cataloguing binary complexes, MINT was conceived to store other types of functional interactions, including enzymatic modifications of one of the partners. Release 1.0 of MINT focuses on experimentally verified protein-protein interactions. Both direct and indirect relationships are considered. Furthermore, MINT aims at being exhaustive in the description of the interaction and, whenever available, information about kinetic and binding constants and about the domains participating in the interaction is included in the entry. MINT consists of entries extracted from the scientific literature by expert curators assisted by 'MINT Assistant', a software that targets abstracts containing interaction information and presents them to the curator in a user-friendly format. The interaction data can be easily extracted and viewed graphically through 'MINT Viewer'. Presently MINT contains 4568 interactions, 782 of which are indirect or genetic interactions.

Amino Acid Sequence↗

Expression of fibrinogen E-fragment and fibrin E-fragment is inhibited in the human infiltrating ductal carcinoma of the breast: the two-dimensional electrophoresis and MALDI-TOF-mass spectrometry analyses.

In the present study, total proteins from a tissue of an infiltrating ductal carcinoma of the breast (IDCA) were compared by the two-dimensional electrophoresis (2D-PAGE) to proteins from an adjacent non-neoplastic breast tissue. Analysis of multiple gels for each sample identified nine proteins present in the tumor sample that were less present in the matched normal adjacent breast tissue and four proteins present at higher levels in the normal tissue. The altered proteins were identified by matrix assisted laser desorption/ionization-time of flight (MALDI-TOF) mass spectrometry and search in protein databases. Protein disulfide isomerase, BiP protein, calreticulin, cathepsin D, inorganic pyrophosphatase, vimentin, apolipoprotein A1 precursor, tropomyosin 4 and beta5-tubulin were identified as being significantly over-expressed in the IDCA with regard to the normal tissue. The expression of fibrinogen E-fragment (known as anti-angiogenic factor) as well as of fibrin E, Pro2619 and actinG1 was found to be inhibited in the tumor sample. The identified proteins might play an important role during malignant transformation, breast cancer progression, and angiogenesis as well as in cellular signaling. This study demonstrates quantitative and qualitative changes in protein abundance between IDCA and normal tissue. The identification of these differentially expressed proteins could lead to a better understanding of the molecular events linked to breast cancer progression.

Breast↗