Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Protein database, human retinal pigment epithelium.

The retinal pigment epithelium (RPE) is a single cell layer adjacent to the rod and cone photoreceptors that plays key roles in retinal physiology and the biochemistry of vision. RPE cells were isolated from normal adult human donor eyes, subcellular fractions were prepared, and proteins were fractionated by electrophoresis. Following in-gel proteolysis, proteins were identified by peptide sequencing using liquid chromatography tandem electrospray mass spectrometry and/or by peptide mass mapping using matrix-assisted laser desorption ionization time-of-flight mass spectrometry. Preliminary analyses have identified 278 proteins and provide a starting point for building a database of the human RPE proteome.

Chromatography, Liquid↗

Access to DNA and protein databases on the Internet.

During the past year, the number of biological databases that can be queried via Internet has dramatically increased. This increase has resulted from the introduction of networking tools, such as Gopher and WAIS, that make it easy for research workers to index databases and make them available for on-line browsing. Biocomputing in the nineties will see the advent of more client/server options for the solution of problems in bioinformatics.

Amino Acid Sequence↗

The alpha/beta fold family of proteins database and the cholinesterase gene server ESTHER.

ESTHER (for esterases, alpha/betahydrolase enzyme and relatives) is a database of sequences phylogenetically related to cholinesterases. These sequences define a homogeneous group of enzymes (carboxylesterases, lipases and hormone-sensitive lipases) sharing a similar structure of a central beta-sheet surrounded by alpha-helices. Among these proteins a wide range of functions can be found (hydrolases, adhesion molecules, hormone precursors). The purpose of ESTHER is to help comparison of structures and functions of members of the family. Since the last release, new features have been added to the server. A BLAST comparison tool allows sequence homology searches within the database sequences. New sections are available: kinetics and inhibitors of cholinesterases, fasciculin-acetylcholinesterase interaction and a gene structure review. The mutation analysis compilation has been improved with three-dimensional images. A mailing list has been created.

Amino Acid Sequence↗

EXProt: a database for proteins with an experimentally verified function.

EXProt is a non-redundant protein database containing a selection of entries from genome annotation projects and public databases, aimed at including only proteins with an experimentally verified function. In EXProt release 2.0 we have collected entries from the Pseudomonas aeruginosa community annotation project (PseudoCAP), the Escherichia coli genome and proteome database (GenProtEC) and the translated coding sequences from the Prokaryotes division of EMBL nucleotide sequence database, which are described as having an experimentally verified function. Each entry in EXProt has a unique ID number and contains information about the species, amino acid sequence, functional annotation and, in most cases, links to references in MEDLINE/PubMed and to the entry in the original database. EXProt is indexed in SRS at CMBI (http://www.cmbi.kun.nl/srs/) and can be searched with BLAST and FASTA through the EXProt web page (http://www.cmbi.kun.nl/EXProt/).

Animals↗

Retrieval accuracy, statistical significance and compositional similarity in protein sequence database searches.

Protein sequence database search programs may be evaluated both for their retrieval accuracy--the ability to separate meaningful from chance similarities--and for the accuracy of their statistical assessments of reported alignments. However, methods for improving statistical accuracy can degrade retrieval accuracy by discarding compositional evidence of sequence relatedness. This evidence may be preserved by combining essentially independent measures of alignment and compositional similarity into a unified measure of sequence similarity. A version of the BLAST protein database search program, modified to employ this new measure, outperforms the baseline program in both retrieval and statistical accuracy on ASTRAL, a SCOP-based test set.

Data Interpretation, Statistical↗

Bacterial identification by protein mass mapping combined with an experimentally derived protein mass database.

A protein mass mapping approach using mass spectrometry (MS) combined with an experimentally derived protein mass database is presented for rapid and effective identification of bacterial species. A prototype mass database from the protein extracts of nine bacterial species has been created by off-line high-performance liquid chromatography (HPLC) matrix-assisted laser desorption/ionization (MALDI) MS, in which the microbiological parameter of bacterial growth time is considered. A numerical method using a statistical weight factor algorithm is devised for matching the protein masses of an unknown bacterial sample against the database. The sum of these weight factors produces a corresponding summed weight factor score for each bacterial species listed in the database, and the database species producing the highest score represents the identity of the respective unknown bacterium. The applicability and reliability of this protein mass mapping approach has been tested with seven bacterial species in a single-blind study by both direct MALDI MS and HPLC electrospray ionization MS methods, and identification results with 100% accuracy are obtained. Our studies have demonstrated that the protein mass database can be rapidly established and readily adopted with relatively less dependency on experimental factors. Furthermore, it is shown that a number of proteins can be detected using a protein sample amount equivalent to an extract of less than 1000 cells, demonstrating that this protein mass mapping approach can potentially be highly sensitive for rapid bacterial identification.

Bacteria↗

The PIR-International Protein Sequence Database.

The Protein Information Resource (PIR; http://www-nbrf.georgetown. edu/pir/) supports research on molecular evolution, functional genomics, and computational biology by maintaining a comprehensive, non-redundant, well-organized and freely available protein sequence database. Since 1988 the database has been maintained collaboratively by PIR-International, an international association of data collection centers cooperating to develop this resource during a period of explosive growth in new sequence data and new computer technologies. The PIR Protein Sequence Database entries are classified into superfamilies, families and homology domains, for which sequence alignments are available. Full-scale family classification supports comparative genomics research, aids sequence annotation, assists database organization and improves database integrity. The PIR WWW server supports direct on-line sequence similarity searches, information retrieval, and knowledge discovery by providing the Protein Sequence Database and other supplementary databases. Sequence entries are extensively cross-referenced and hypertext-linked to major nucleic acid, literature, genome, structure, sequence alignment and family databases. The weekly release of the Protein Sequence Database can be accessed through the PIR Web site. The quarterly release of the database is freely available from our anonymous FTP server and is also available on CD-ROM with the accompanying ATLAS database search program.

Amino Acid Sequence↗

The PMDB Protein Model Database.

The Protein Model Database (PMDB) is a public resource aimed at storing manually built 3D models of proteins. The database is designed to provide access to models published in the scientific literature, together with validating experimental data. It is a relational database and it currently contains >74,000 models for approximately 240 proteins. The system is accessible at http://www.caspur.it/PMDB and allows predictors to submit models along with related supporting evidence and users to download them through a simple and intuitive interface. Users can navigate in the database and retrieve models referring to the same target protein or to different regions of the same protein. Each model is assigned a unique identifier that allows interested users to directly access the data.

Databases, Protein↗

Constructing ontology-driven protein family databases.

MOTIVATION: Protein family databases provide a central focus for scientific communities as well as providing useful resources to aide research. However, such resources require constant curation and often become outdated and discontinued. We have developed an ontology-driven system for capturing and managing protein family data that addresses the problems of maintenance and sustainability. RESULTS: Using protein phosphatases and ABC transporters as model protein families, we constructed two protein family database resources around a central DAML+OIL ontology. Each resource contains specialist information about each protein family, providing specialized domain-specific resources based on the same template structure. The formal structure, combined with the extraction of biological data using GO terms, allows for automated update strategies. Despite the functional differences between the two protein families, the ontology model was equally applicable to both, demonstrating the generic nature of the system. AVAILABILITY: The protein phosphatase resource, PhosphaBase, is freely available on the internet (http://www.bioinf.man.ac.uk/phosphabase). The DAML+OIL ontology for the protein phosphatases and the ABC transporters is available on request from the authors. CONTACT: kwolstencroft@cs.man.ac.uk.

ATP-Binding Cassette Transporters↗

PFDB: a generic protein family database integrating the CATH domain structure database with sequence based protein family resources.

MOTIVATION: The PFDB (Protein Family Database) is a new database designed to integrate protein family-related data with relevant functional and genomic data. It currently manages biological data for three projects-the CATH protein domain database (Orengo et al., 1997; Pearl et al., 2001), the VIDA virus domains database (Albà et al., 2001) and the Gene3D database (Buchan et al., 2001). The PFDB has been designed to accommodate protein families identified by a variety of sequence based or structure based protocols and provides a generic resource for biological research by enabling mapping between different protein families and diverse biochemical and genetic data, including complete genomes. RESULTS: A characteristic feature of the PFDB is that it has a number of meta-level entities (for example aggregation, collection and inclusion) represented as base tables in the final design. The explicit representation of relationships at the meta-level has a number of advantages, including flexibility-both in terms of the range of queries that can be formulated and the ability to integrate new biological entities within the existing design. A potential drawback with this approach-poor performance caused by the number of joins across meta-level tables-is avoided by implementing the PFDB with materialized views using the mature relational database technology of Oracle 8i. The resultant database is both fast and flexible. This paper presents the principles on which the database has been designed and implemented, and describes the current status of the database and query facilities supported.

Database Management Systems↗

Computerized, comprehensive databases of cellular and secreted proteins from normal human embryonic lung MRC-5 fibroblasts: identification of transformation and/or proliferation sensitive proteins.

Databases of protein information from human embryonal lung fibroblasts (MRC-5) have been established using computer analyzed two-dimensional gel electrophoresis. One thousand four hundred and eighty-two cellular proteins (1060 with isoelectric focusing and 422 with nonequilibrium pH gradient electrophoresis, in the first dimension) ranging in molecular mass between 8 and 234 kDa were separated and numbered. Information entered in the database (in most cases for major proteins) includes: protein name, HeLa protein catalog number, mouse protein catalog number, proteins matched in transformed human epithelial amnion cells (AMA) and peripheral blood mononuclear cells (PBMC), transformation and/or proliferation sensitive proteins, synthesis in quiescent cells, cell cycle regulated proteins, mitochondrial and heat shock proteins, cytoskeletal proteins and proteins whose synthesis is affected by interferons. Additional information entered for a few transformation-sensitive proteins that have been selected for future studies includes levels of synthesis and amounts in fetal human tissues. A total of four hundred and seventy-six [35S]methionine labeled polypeptides (258 isoelectric focusing; 218, nonequilibrium pH gradient electrophoresis) secreted by MRC-5 fibroblasts were separated and recorded (J. E. Celis et al., Leukemia 1987, 1, 707-717). Information entered in this database includes molecular weight and transformation sensitive proteins. These databases, as well as those of epithelial and lymphoid cell proteins (J. E. Celis et al., Leukemia 1988, 9, 561-601), represent the initial stages of a systematic effort to establish comprehensive databases of human protein information. In the long run, these databases are expected to offer a useful framework in which to focus the human genome sequencing effort.

Cell Line, Transformed↗

Derivation of rules for comparative protein modeling from a database of protein structure alignments.

We describe a database of protein structure alignments as well as methods and tools that use this database to improve comparative protein modeling. The current version of the database contains 105 alignments of similar proteins or protein segments. The database comprises 416 entries, 78,495 residues, 1,233 equivalent entry pairs, and 230,396 pairs of equivalent alignment positions. At present, the main application of the database is to improve comparative modeling by satisfaction of spatial restraints implemented in the program MODELLER (Sali A, Blundell TL, 1993, J Mol Biol 234:779-815). To illustrate the usefulness of the database, the restraints on the conformation of a disulfide bridge provided by an equivalent disulfide bridge in a related structure are derived from the alignments; the prediction success of the disulfide dihedral angle classes is increased to approximately 80%, compared to approximately 55% for modeling that relies on the stereochemistry of disulfide bridges alone. The second example of the use of the database is the derivation of the probability density function for comparative modeling of the cis/trans isomerism of the proline residues; the prediction success is increased from 0% to 82.9% for cis-proline and from 93.3% to 96.2% for trans-proline. The database is available via electronic mail.

Amino Acid Sequence↗