Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

YPED: a proteomics database for protein expression analysis.

We have developed the Yale Protein Expression Database (YPED) to address the storage, retrieval, and integrated analysis of proteomics data generated by Yale's Keck Protein Chemistry and Mass Spectrometry Facility. YPED is Web-accessible and currently handles sample requisition, result reporting and sample comparison for ICAT, DIGE and MUDPIT samples. Sample descriptions are compatible with the evolving MIAPE standards. Peptides and proteins identified using Sequest or Mascot are validated with the Trans-Proteomic Pipeline developed at the Institute of Systems Biology and data from the resulting XML file are stored in the database. Researchers can view, subset and download their data through a secure Web interface.

Databases, Protein↗

Human protein reference database--2006 update.

Human Protein Reference Database (HPRD) (http://www.hprd.org) was developed to serve as a comprehensive collection of protein features, post-translational modifications (PTMs) and protein-protein interactions. Since the original report, this database has increased to >20 000 proteins entries and has become the largest database for literature-derived protein-protein interactions (>30 000) and PTMs (>8000) for human proteins. We have also introduced several new features in HPRD including: (i) protein isoforms, (ii) enhanced search options, (iii) linking of pathway annotations and (iv) integration of a novel browser, GenProt Viewer (http://www.genprot.org), developed by us that allows integration of genomic and proteomic information. With the continued support and active participation by the biomedical community, we expect HPRD to become a unique source of curated information for the human proteome and spur biomedical discoveries based on integration of genomic, transcriptomic and proteomic data.

Databases, Protein↗

Automated clustering of ensembles of alternative models in protein structure databases.

Experimentally determined protein structures have been classified in different public databases according to their structural and evolutionary relationships. Frequently, alternative structural models, determined using X-ray crystallography or NMR spectroscopy, are available for a protein. These models can present significant structural dissimilarity. Currently there is no classification available for these alternative structures. In order to classify them, we developed STRuster, an automated method for clustering ensembles of structural models according to their backbone structure. The method is based on the calculation of carbon alpha (Calpha) distance matrices. Two filters are applied in the calculation of the dissimilarity measure in order to identify both large and small (but significant) backbone conformational changes. The resulting dissimilarity value is used for hierarchical clustering and partitioning around medoids (PAM). Hierarchical clustering reflects the hierarchy of similarities between all pairs of models, while PAM groups the models into the 'optimal' number of clusters. The method has been applied to cluster the structures in each SCOP species level and can be easily applied to any other sets of conformers. The results are available at: http://bioinf.mpi-sb.mpg.de/projects/struster/.

Aldehyde-Lyases↗

The PRINTS protein fingerprint database in its fifth year.

PRINTS is a database of protein family 'fingerprints' offering a diagnostic resource for newly-determined sequences. By contrast with PROSITE, which uses single consensus expressions to characterise particular families, PRINTS exploits groups of motifs to build characteristic signatures. These signatures offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 800 fingerprints have been constructed and stored in PRINTS. The current version, 17.0, encodes approximately 4500 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via the UCL Bioinformatics World Wide Web (WWW) Server at http://www. biochem.ucl.ac.uk/bsm/dbbrowser/ . We have recently enhanced the usefulness of PRINTS by making available new, intuitive search software. This allows both individual query sequence and bulk data submission, permitting easy analysis of single sequences or complete genomes. Preliminary results indicate that use of the PRINTS system is able to assign additional functions not found by other methods, and hence offers a useful adjunct to current genome analysis protocols.

Animals↗

DBSubLoc: database of protein subcellular localization.

We have built a protein subcellular localization annotation database, the DBSubLoc database, which is available at http://www.bioinfo.tsinghua. edu.cn/dbsubloc.html. Annotations were taken from primary protein databases, model organism genome projects and literature texts, and then were analyzed to dig out the subcellular localization features of the proteins. The proteins are also classified into different categories. Based on sequence alignment, non-redundant subsets of the database have been built, which may provide useful information for subcellular localization prediction. The database now contains >60,000 protein sequences including approximately 30,000 protein sequences in the non-redundant data sets. Online download, search and Blast tools are also available.

Animals↗

Novel developments with the PRINTS protein fingerprint database.

The PRINTS database of protein family 'fingerprints' is a diagnostic resource that complements the PROSITE dictionary of sites and patterns. Unlike regular expressions, fingerprints exploit groups of conserved motifs within sequence alignments to build characteristic signatures of family membership. Thus fingerprints inherently offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 600 fingerprints have been constructed and stored in PRINTS, representing a 50% increase in the size of the database in the last year. The current version, 13.0, encodes approximately 3000 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via UCL's Bioinformatics World Wide Web (WWW) server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser / . We describe here progress with the database, its Web interface, and a recent exciting development: the integration of a novel colour alignment editor (http://www.biochem.ucl.ac.uk/bsm/dbbrowser++ +/CINEMA ), which allows visualisation and interactive manipulation of PRINTS alignments over the Internet.

Amino Acid Sequence↗

INTERACT: an object oriented protein-protein interaction database.

MOTIVATION: Protein-protein interactions provide vital information concerning the function of proteins, complexes and networks. Currently there is no widely accepted repository of this interaction information. Our aim is to provide a single database with the necessary architecture to fully store, query and analyse interaction data. RESULTS: An object oriented database has been created which provides scientists with a resource for examining existing protein-protein interactions and inferring possible interactions from the data stored. It also provides a basis for examining networks of interacting proteins, via analysis of the data stored. The database contains over a thousand interactions. CONTACT: k.eilbeck@stud.man.ac.uk

Computational Biology↗

An iterative calibration method with prediction of post-translational modifications for the construction of a two-dimensional electrophoresis database of mouse mammary gland proteins.

Protein databases serve as general reference resources providing an orientation on two-dimensional electrophoresis (2-DE) patterns of interest. The intention behind constructing a 2-DE database of the water soluble proteins from wild-type mouse mammary gland tissue was to create a reference before going on to investigate cancer-associated protein variations. This database shall be deemed to be a model system for mouse tissue, which is open for transgenic or knockout experiments. Proteins were separated and characterized in terms of their molecular weight (M(r)) and isoelectric point (pI) by high resolution 2-DE. The proteins were identified using prevalent proteomics methods. One method was peptide mass fingerprinting by matrix-assisted laser desorption/ionization-mass spectrometry. Another method was N-terminal sequencing by Edman degradation. By N-terminal sequencing M(r) and pI values were specified more accurately and so the calibration of the master gel was obtained more systematically and exactly. This permits the prediction of possible post-translational modifications of some proteins. The mouse mammary gland 2-DE protein database created presently contains 66 identified protein spots, which are clickable on the gel pattern. This relational database is accessible on the WWW under the URL: http://www.mpiib-berlin.mpg.de/2D-PAGE.

Animals↗

YPL.db2: the Yeast Protein Localization database, version 2.0.

The Yeast Protein Localization database (YPL.db(2)) is an archive of microscopic image data of protein localization patterns in the yeast Saccharomyces cerevisiae. The current version of YPL.db(2) harbours 500 sets of image data derived from high-resolution microscopic analyses of proteins tagged with the green fluorescent protein (GFP). Major functional improvements in YPL.db(2) over a previous release are a web-based experiment and image submission interface, facilitating standardized data entry by remote users through the Internet. The image display page provides image gallery and image scrolling features. In addition, fluorescence and transmission images can be superimposed, allowing image fading for precise correlation of the protein's localization in the cellular context. The reference structure database displaying 'prototypic' localization patterns was extended, and a feature to display and manipulate 3D-image datasets, using a freely available VRML plug-in, was included. Access to the Yeast Protein Localization database version 2.0 (YPL.db(2)) is available through http://YPL.uni-graz.at.

Databases, Protein↗

DIP: the database of interacting proteins.

The Database of Interacting Proteins (DIP; http://dip.doe-mbi.ucla.edu) is a database that documents experimentally determined protein-protein interactions. This database is intended to provide the scientific community with a comprehensive and integrated tool for browsing and efficiently extracting information about protein interactions and interaction networks in biological processes. Beyond cataloging details of protein-protein interactions, the DIP is useful for understanding protein function and protein-protein relationships, studying the properties of networks of interacting proteins, benchmarking predictions of protein-protein interactions, and studying the evolution of protein-protein interactions.

Databases, Factual↗

Increased coverage obtained by combination of methods for protein sequence database searching.

MOTIVATION: Sequence alignment methods that compare two sequences (pairwise methods) are important tools for the detection of biological sequence relationships. In genome annotation, multiple methods are often run and agreement between methods taken as confirmation. In this paper, we assess the advantages of combining search methods by comparing seven pairwise alignment methods, including three local dynamic programming algorithms (PRSS, SSEARCH and SCANPS), two global dynamic programming algorithms (GSRCH and AMPS) and two heuristic approximations (BLAST and FASTA), individually and by pairwise intersection and union of their result lists at equal p-value cut-offs. RESULTS: When applied singly, the dynamic programming methods SCANPS and SSEARCH gave significantly better coverage (p=0.01) compared to AMPS, GSRCH, PRSS, BLAST and FASTA. Results ranked by BLAST p-values gave significantly better coverage compared to ranking by BLAST e-values. Of 56 combinations of eight methods considered, 19 gave significant increases in coverage at low error compared to the parent methods at an equal p-value cutoff. The union of results by BLAST (p-value) and FASTA at an equal p-value cutoff gave significantly better coverage than either method individually. The best overall performance was obtained from the intersection of the results from SSEARCH and the GSRCH62 global alignment method. At an error level of five false positives, this combination found 444 true positives, a significant 12.4% increase over SSEARCH applied alone.

Algorithms↗

A hypergeometric probability model for protein identification and validation using tandem mass spectral data and protein sequence databases.

We present a new probability-based method for protein identification using tandem mass spectra and protein databases. The method employs a hypergeometric distribution to model frequencies of matches between fragment ions predicted for peptide sequences with a specific (M + H)+ value (at some mass tolerance) in a protein sequence database and an experimental tandem mass spectrum. The hypergeometric distribution constitutes null hypothesis-all peptide matches to a tandem mass spectrum are random. It is used to generate a score characterizing the randomness of a database sequence match to an experimental tandem mass spectrum and to determine the level of significance of the null hypothesis. For each tandem mass spectrum and database search, a peptide is identified that has the least probability of being a random match to the spectrum and the corresponding level of significance of the null hypothesis is determined. To check the validity of the hypergeometric model in describing fragment ion matches, we used chi2 test. The distribution of frequencies and corresponding hypergeometric probabilities are generated for each tandem mass spectrum. No proteolytic cleavage specificity is used to create the peptide sequences from the database. We do not use any empirical probabilities in this method. The scores generated by the hypergeometric model do not have a significant molecular weight bias and are reasonably independent of database size. The approach has been implemented in a database search algorithm, PEP_PROBE. By using a large set of tandem mass spectra derived from a set of peptides created by digestion of a collection of known proteins using four different proteases, a false positive rate of 5% is demonstrated.

Amino Acid Sequence↗

ProTherm: Thermodynamic Database for Proteins and Mutants.

The first release of the Thermodynamic Database for Proteins and Mutants (ProTherm) contains more than 3300 data of several thermodynamic parameters for wild type and mutant proteins. Each entry includes numerical data for unfolding Gibbs free energy change, enthalpy change, heat capacity change, transition temperature, activity etc., which are important for understanding the mechanism of protein stability. ProTherm also includes structural information such as secondary structure and solvent accessibility of wild type residues, and experimental methods and other conditions. A WWW interface enables users to search data based on various conditions with different sorting options for outputs. Further, ProTherm is cross-linked with NCBI PUBMED literature database, Protein Mutant Database, Enzyme Code and Protein Data Bank structural database. Moreover, all the mutation sites associated with each PDB structure are automatically mapped and can be directly viewed through 3DinSight developed in our laboratory. The database is available at the URL, http://www.rtc.riken.go.jp/protherm.htm l

Calorimetry↗

DisProt: the Database of Disordered Proteins.

The Database of Protein Disorder (DisProt) links structure and function information for intrinsically disordered proteins (IDPs). Intrinsically disordered proteins do not form a fixed three-dimensional structure under physiological conditions, either in their entireties or in segments or regions. We define IDP as a protein that contains at least one experimentally determined disordered region. Although lacking fixed structure, IDPs and regions carry out important biological functions, being typically involved in regulation, signaling and control. Such functions can involve high-specificity low-affinity interactions, the multiple binding of one protein to many partners and the multiple binding of many proteins to one partner. These three features are all enabled and enhanced by protein intrinsic disorder. One of the major hindrances in the study of IDPs has been the lack of organized information. DisProt was developed to enable IDP research by collecting and organizing knowledge regarding the experimental characterization and the functional associations of IDPs. In addition to being a unique source of biological information, DisProt opens doors for a plethora of bioinformatics studies. DisProt is openly available at http://www.disprot.org.

Databases, Protein↗

Towards a comprehensive database of proteins from the urine of patients with bladder cancer.

PURPOSE: To establish a comprehensive two-dimensional database of proteins from the urine of patients with bladder cancer. MATERIALS AND METHODS: Urines dialized against distilled H2O were freeze-dried and subjected to isoelectrofocusing two-dimensional gel electrophoresis. Coomassie brilliant blue stained dry gels were scanned and analyzed with the PDQUEST software. Proteins were identified by one or more of the following techniques: microsequencing (44 proteins), mass spectrometry, comigration and immunoblotting. RESULTS: The urine protein database, which includes all the polypeptides detected in the urines of 50 patients, lists 339 proteins (including variants) of which 124 have been identified. Of these, psoriasin, the psoriasis-associated fatty acid binding protein 5, the gelsolin fragments and prostaglandin D2 synthetase have not been previously described. CONCLUSIONS: The database provides a solid basis for further studies aiming at identifying tumor markers in the urine that may serve as prognostic factors in bladder cancer.

Amino Acid Sequence↗

Human cellular protein patterns and their link to genome DNA mapping and sequencing data: towards an integrated approach to the study of gene expression.

Analysis of cellular protein patterns by computer-aided two-dimensional gel electrophoresis together with recent advances in protein sequence analysis and expression systems have made possible the establishment of comprehensive two-dimensional gel protein databases that may link protein and DNA mapping and sequence information and that offer an integrated approach to the study of gene expression. With the integrated approach offered by two-dimensional gel protein databases it is now possible to reveal phenotype-specific protein(s), to microsequence them, to search for homology with previous identified proteins, to clone the cDNAs, to assign partial protein sequences to genes for which the full DNA sequence and the chromosome location are known, and to study the regulatory properties and function of groups of proteins that are coordinately expressed in a given biological process. Comprehensive two-dimensional gel protein databases will provide an integrated picture of the expression levels and properties of the thousands of protein components of organelles, pathways, and cytoskeletal systems, both under physiological and abnormal conditions, and are expected to lead to the identification of new regulatory networks. So far, about 20% (600 out of 2,980) of the total number of proteins recorded in the human keratinocyte protein database have been identified and we are actively gathering qualitative and quantitative biological data on all resolved proteins. Given the current improvements on microsequencing as well as the availability of specific antibodies, it seems feasible to expect that most known keratinocyte proteins will be identified in the very near future. This feast will reveal a wealth of new proteins that will become amenable to experimentation both at the biochemical and molecular biology level.

Amino Acid Sequence↗