Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Databases, Protein”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Bioinformatics analyses of circular dichroism protein reference databases.

MOTIVATION: Circular dichroism (CD) spectroscopy has become established as a key method for determining the secondary structure contents of proteins which has had a significant impact on molecular biology. Many excellent mathematical protocols have been developed for this purpose and their quality is above question. However, reference database sets of proteins, with CD spectra matched to secondary structure components derived from X-ray structures, provide the key resource for this task. These databases were created many years ago, before most CD spectrophotometers became standardized and before it was commonplace to validate X-ray structures prior to publication. The analyses presented here were undertaken to investigate the overall quality of these reference databases in light of their extensive usage in determining protein secondary structure content from CD spectra. RESULTS: The analyses show that there are a number of significant problems associated with the CD reference database sets in current use. There are disparities between CD spectra for the same protein collected by different groups. These include differences in magnitudes, peak positions or both. However, many current reference sets are now amalgamations of spectra from these groups, introducing inconsistencies that can lead to inaccuracies in the determination of secondary structure components from the CD spectra. A number of the X-ray structures used fall short on the validation criteria now employed as standard for structure determination. Many have substantial percentages of residues in the disallowed regions of the Ramachandran plot. Hence their calculated secondary structure components, used as a foundation for the reference databases, are likely to be in error. Additionally, the coverage of secondary structure space in the reference datasets is poorly correlated to the secondary structure components found in the Protein Data Bank. A conclusion is that a new reference CD database with cross-correlated, machine-independent CD spectra and validated X-ray structures that cover more secondary structure components, including diverse protein folds, is now needed. However, that reasonably accurate values for the secondary structure content of proteins can be determined from spectra is a testament to CD spectroscopy being a very powerful technique.

Circular Dichroism↗

Automated protein sequence database classification. I. Integration of compositional similarity search, local similarity search, and multiple sequence alignment.

MOTIVATION: Genome sequencing projects require the periodic application of analysis tools that can classify and multiply align related protein sequence domains. Full automation of this task requires an efficient integration of similarity and alignment techniques. RESULTS: We have developed a fully automated process that classifies entire protein sequence databases, resulting in alignment of the homologous sequences. The successive steps of the procedure are based on compositional and local sequence similarity searches followed by multiple sequence alignments. Global similarities are detected from the pairwise comparison of amino acid and dipeptide compositions of each protein. After the elimination of all but one sequence from each detected cluster of closely related proteins, the remaining sequences are compiled in a suffix tree which is self-compared to detect local sequence similarities. Sets of proteins which share similar sequence segments are then weighted according to their closeness and multiply aligned using a fast hierarchical dynamic programming algorithm. Computational strategies were devised to minimize computer processing time and memory space requirements. The accuracy of the sequence classifications has been evaluated for 12 462 primary structures distributed over 341 known families. The percentage of sequences with missed or incorrect family assignments was 6.8% on the test set. This low error level is only twice that of the manually constructed PROSITE database ( 3.4% ) and is substantially better than that found for the automatically built PRODOM database ( 34.9% ). AVAILABILITY: The resulting database, called DOMO, is available through database search routine SRS at Infobiogen (http://www.infobiogen.fr/srs5/), EBI (http://srs.ebi.ac.uk:5000/) and EMBL (http://www.embl-heidelberg.de/srs5/) World Wide Web sites. CONTACT: gracy@infobiogen.fr

Algorithms↗

Computer-assisted generation of a protein-interaction database for nuclear receptors.

With the increasing amount of biological data available, automated methods for information retrieval become necessary. We employed computer-assisted text mining to retrieve all protein-protein interactions for nuclear receptors from MEDLINE in a systematic way. A dictionary of protein names and of terms denoting interactions was generated, and trioccurrences of two protein names and one interaction term in one sentence were retrieved. Abstracts containing at least one such trioccurrence were manually checked by biologists to select the relevant interactions out of the automatically extracted data. In total, 4360 abstracts were retrieved containing data on protein interactions for nuclear receptors. The resulting database contains all reported protein interactions involving nuclear receptors from 1966 to September 2001. Remarkably, the annual increase in number of reported interactors for nuclear receptors has been following an exponential growth curve in the years 1991 to 2001. Apparent in the data set is the high complexity of protein interactions for nuclear receptors. The number of interactions correlates with the number of published papers for a given receptor, suggesting that the number of reported interactors is a reflection of the intensity of research dedicated to a given receptor. Indeed, comparison of the retrieved data to a systematic yeast two-hybrid-based interaction analysis suggests that most NRs are similar with respect to the number of interacting proteins. The data set obtained serves as a source for information on NR interactions, as well as a reference data set for the improvement of advanced text-mining methods.

Computers↗

Motif-based searching in TOPS protein topology databases.

MOTIVATION: TOPS cartoons are a schematic ion of protein three-dimensional structures in two dimensions, and are used for understanding and manual comparison of protein folds. Recently, an algorithm that produces the cartoons automatically from protein structures has been devised and cartoons have been generated to represent all the structures in the structural databank. There is now a need to be able to define target topological patterns and to search the database for matching domains. RESULTS: We have devised a formal language for describing TOPS diagrams and patterns, and have designed an efficient algorithm to match a pattern to a set of diagrams. A pattern-matching system has been implemented, and tested on a database derived from all the current entries in the Protein Data Bank (15,000 domains). Users can search on patterns selected from a library of motifs or, alternatively, they can define their own search patterns. AVAILABILITY: The system is accessible over the Web at http://tops.ebi.ac.uk/tops

Algorithms↗

LISA: an intranet-based flexible database for protein crystallography project management.

The increase in the number of projects carried out in protein crystallography laboratories has emphasized the need for effective management of project information and data. To meet this need, a flexible web-accessible database for protein crystallography project management (LISA) has been developed using the open-source software MySQL and PHP4. The database contains information about all aspects of structure-determination projects, including primer and plasmid sequences, protein expression, purification and crystallization results, structure coordinate files and resultant publications. The database web pages include links to relevant servers and contain, in addition, tools for processing stored information. The software package is freely available.

Crystallography↗

PANDIT: an evolution-centric database of protein and associated nucleotide domains with inferred trees.

PANDIT is a database of homologous sequence alignments accompanied by estimates of their corresponding phylogenetic trees. It provides a valuable resource to those studying phylogenetic methodology and the evolution of coding-DNA and protein sequences. Currently in version 17.0, PANDIT comprises 7738 families of homologous protein domains; for each family, DNA and corresponding amino acid sequence multiple alignments are available together with high quality phylogenetic tree estimates. Recent improvements include expanded methods for phylogenetic tree inference, assessment of alignment quality and a redesigned web interface, available at the URL http://www.ebi.ac.uk/goldman-srv/pandit.

Databases, Nucleic Acid↗

PASS2: a semi-automated database of protein alignments organised as structural superfamilies.

PASS2 is a nearly automated version of CAMPASS and contains sequence alignments of proteins grouped at the level of superfamilies. This database has been created to fall in correspondence with SCOP database (1.53 release) and currently consists of 110 multi-member superfamilies and 613 superfamilies corresponding to single members. In multi-member superfamilies, protein chains with no more than 25% sequence identity have been considered for the alignment and hence the database aims to address sequence alignments which represent 26 219 protein domains under the SCOP 1.53 release. Structure-based sequence alignments have been obtained by COMPARER and the initial equivalences are provided automatically from a MALIGN alignment and subsequently augmented using STAMP4.0. The final sequence alignments have been annotated for the structural features using JOY4.0. Several interesting links are provided to other related databases and genome sequence relatives. Availability of reliable sequence alignments of distantly related proteins, despite poor sequence identity and single-member superfamilies, permit better sampling of structures in libraries for fold recognition of new sequences and for the understanding of protein structure-function relationships of individual superfamilies. The database can be queried by keywords and also by sequence search, interfaced by PSI-BLAST methods. Structure-annotated sequence alignments and several structural accessory files can be retrieved for all the superfamilies including the user-input sequence. The database can be accessed from http://www.ncbs.res.in/%7Efaculty/mini/campass/pass.html.

Amino Acid Sequence↗

Organelle DB: a cross-species database of protein localization and function.

To efficiently utilize the growing body of available protein localization data, we have developed Organelle DB, a web-accessible database cataloging more than 25,000 proteins from nearly 60 organelles, subcellular structures and protein complexes in 154 organisms spanning the eukaryotic kingdom. Organelle DB is the first on-line resource devoted to the identification and presentation of eukaryotic proteins localized to organelles and subcellular structures. As such, Organelle DB is a strong resource of data from the human proteome as well as from the major model organisms Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, Caenorhabditis elegans and Mus musculus. In particular, Organelle DB is a central repository of yeast data, incorporating results--and actual fluorescent imagesfrom ongoing large-scale studies of protein localization in S.cerevisiae. Each protein in Organelle DB is presented with its sequence and, as available, a detailed description of its function; functions were extracted from relevant model organism databases, and links to these databases are provided within Organelle DB. To facilitate data interoperability, we have annotated all protein localizations using vocabulary from the Gene Ontology consortium. We also welcome new data for inclusion in Organelle DB, which may be freely accessed at http://organelledb.lsi.umich.edu.

Animals↗

Molecular cloning and expression of a novel keratinocyte protein (psoriasis-associated fatty acid-binding protein [PA-FABP]) that is highly up-regulated in psoriatic skin and that shares similarity to fatty acid-binding proteins.

Analysis by means of two-dimensional (2D) gel electrophoresis of the protein patterns of normal and psoriatic unfractionated non-cultured keratinocytes has revealed a few low-molecular-weight proteins that are highly up-regulated in psoriatic skin. These include psoriasin; calgranulin B, also known as MRP 14, L1, or calprotectin; calgranulin A or MRP 8; and cystatin A or stefin A. Here, we have cloned and sequenced the cDNA (clone 1592) encoding a new member of this group of low-molecular-weight proteins [isoelectric focusing (IEF) SSP 3007 in the keratinocyte 2D gel protein database] that we have termed PA-FABP (psoriasis-associated fatty acid-binding protein). The deduced sequence predicted a protein with molecular weight of 15,164 daltons and a calculated pI of 6.96, values that are close to those recorded in the keratinocyte 2D gel protein database. The protein comigrated with PA-FABP as determined by 2D gel analysis of [35S]-methionine-labeled proteins expressed by transformed human amnion (AMA) cells transfected with clone 1592 using the vaccinia virus expression system and reacted with a rabbit polyclonal antibody raised against 2D gel purified PA-FABP. Structural analysis of the amino acid sequence revealed 48%, 52%, and 56% identity to known low-molecular-weight fatty acid-binding proteins belonging to the FABP family. Northern blot analysis showed that PA-FABP mRNA is indeed highly up-regulated in psoriatic keratinocytes. The transcript is present in human cell lines of epithelial and lymphoid (Molt 4) origin but cannot be detected in normal or SV40 transformed MRC-5 fibroblasts. 2D gel protein analysis of normal primary keratinocytes cultured for at least 8 d under conditions that promoted incomplete terminal differentiation [serum-free keratinocyte (SFK) medium supplemented with epidermal growth factor (EGF), pituitary extract, and 10% fetal calf serum] revealed a strong up-regulation of PA-FABP, psoriasin, calgranulins A and B, and a few other proteins that are highly expressed in psoriatic skin. The levels of these proteins exceeded by far those observed in non-cultured normal keratinocytes implying that the cultured cells have followed an altered pattern of differentiation that resembles--at least in part--that of non-cultured psoriatic keratinocytes. The implications of these results for the study of psoriasis are discussed.

Amino Acid Sequence↗

Improving sensitivity in shotgun proteomics using a peptide-centric database with reduced complexity: protease cleavage and SCX elution rules from data mining of MS/MS spectra.

Correct identification of a peptide sequence from MS/MS data is still a challenging research problem, particularly in proteomic analyses of higher eukaryotes where protein databases are large. The scoring methods of search programs often generate cases where incorrect peptide sequences score higher than correct peptide sequences (referred to as distraction). Because smaller databases yield less distraction and better discrimination between correct and incorrect assignments, we developed a method for editing a peptide-centric database (PC-DB) to remove unlikely sequences and strategies for enabling search programs to utilize this peptide database. Rules for unlikely missed cleavage and nontryptic proteolysis products were identified by data mining 11 849 high-confidence peptide assignments. We also evaluated ion exchange chromatographic behavior as an editing criterion to generate subset databases. When used to search a well-annotated test data set of MS/MS spectra, we found no loss of critical information using PC-DBs, validating the methods for generating and searching against the databases. On the other hand, improved confidence in peptide assignments was achieved for tryptic peptides, measured by changes in DeltaCN and RSP. Decreased distraction was also achieved, consistent with the 3-9-fold decrease in database size. Data mining identified a major class of common nonspecific proteolytic products corresponding to leucine aminopeptidase (LAP) cleavages. Large improvements in identifying LAP products were achieved using the PC-DB approach when compared with conventional searches against protein databases. These results demonstrate that peptide properties can be used to reduce database size, yielding improved accuracy and information capture due to reduced distraction, but with little loss of information compared to conventional protein database searches.

Amino Acid Sequence↗

Influence of protein structure databases on the predictive power of statistical pair potentials.

A long standing goal in protein structure studies is the development of reliable energy functions that can be used both to verify protein models derived from experimental constraints as well as for theoretical protein folding and inverse folding computer experiments. In that respect, knowledge-based statistical pair potentials have attracted considerable interests recently mainly because they include the essential features of protein structures as well as solvent effects at a low computing cost. However, the basis on which statistical potentials are derived have been questioned. In this paper, we investigate statistical pair potentials derived from protein three-dimensional structures, addressing in particular questions related to the form of these potentials, as well as to the content of the database from which they are derived. We have shown that statistical pair potentials depend on the size of the proteins included in the database, and that this dependence can be reduced by considering only pairs of residue close in space (i.e., with a cutoff of 8 A). We have shown also that statistical potentials carry a memory of the quality of the database in terms of the amount and diversity of secondary structure it contains. We find, for example, that potentials derived from a database containing alpha-proteins will only perform best on alpha-proteins in fold recognition computer experiments. We believe that this is an overall weakness of these potentials, which must be kept in mind when constructing a database.

Chemical Phenomena↗

Pandit: a database of protein and associated nucleotide domains with inferred trees.

MOTIVATION: A large, high-quality database of homologous sequence alignments with good estimates of their corresponding phylogenetic trees will be a valuable resource to those studying phylogenetics. It will allow researchers to compare current and new models of sequence evolution across a large variety of sequences. The large quantity of data may provide inspiration for new models and methodology to study sequence evolution and may allow general statements about the relative effect of different molecular processes on evolution. RESULTS: The Pandit 7.6 database contains 4341 families of sequences derived from the seed alignments of the Pfam database of amino acid alignments of families of homologous protein domains (Bateman et al., 2002). Each family in Pandit includes an alignment of amino acid sequences that matches the corresponding Pfam family seed alignment, an alignment of DNA sequences that contain the coding sequence of the Pfam alignment when they can be recovered (overall, 82.9% of sequences taken from Pfam) and the alignment of amino acid sequences restricted to only those sequences for which a DNA sequence could be recovered. Each of the alignments has an estimate of the phylogenetic tree associated with it. The tree topologies were obtained using the neighbor joining method based on maximum likelihood estimates of the evolutionary distances, with branch lengths then calculated using a standard maximum likelihood approach.

Algorithms↗

Construction of HSC-2DPAGE: a two-dimensional gel electrophoresis database of heart proteins.

The dissemination of information relating to the characterisation of proteins from two-dimensional electrophoresis (2-DE) gel databases is essential for their effective utilisation in the study of protein expression in cell biology. Since the inception of the World Wide Web and the pioneering development of SWISS-2DPAGE as a tool for retrieving information on proteins separated by 2-DE, the Internet has become the method of choice for disseminating and accessing information on 2-DE protein databases. At Harefield we have established HSC-2DPAGE which is an advanced interface for accessing protein database relating to heart disease. The Web site currently includes databases of proteins from human, dog and rat ventricular tissue and a human endothelial cell line. The databases are searchable individually or as a whole by remote keyword searches. Each database is represented by both synthetic (computer generated) and real (scanned gel) clickable images upon which characterised protein spots are highlighted by hyperlinked symbols. The database conforms to all the rules proposed for federated 2-DE protein databases and individual protein entries are linked to other protein databases such as SWISS-PROT by active cross-references. This paper describes the construction of HSC-2DPAGE, its maintenance, and access via the Internet.

Animals↗

The Protein Information Resource (PIR) and the PIR-International Protein Sequence Database.

From its origin, the PIR has aspired to support research in computational biology and genomics through the compilation of a comprehensive, quality controlled and well-organized protein sequence information resource. The resource originated with the pioneering work of the late Margaret O. Dayhoff in the early 1960s. Since 1988, the Protein Sequence Database has been maintained collaboratively by PIR-International, an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. The work of the resource is widely distributed and is available on the World Wide Web, via FTP, E-mail server, CD-ROM and magnetic media. It is widely redistributed and incorporated into many other protein sequence data compilations including SWISS-PROT and theEntrezsystem of the NCBI.

Amino Acid Sequence↗

Comprehensive mass spectrometric analysis of the 20S proteasome complex.

The 20S proteasome is a multicatalytic protein complex that plays an important role in intracellular protein degradation from archaebacteria to eukaryotes. This complex is made up of two copies each of seven different alpha (alpha) and seven different beta (beta) subunits arranged into four stacked rings (alpha7beta7beta7alpha7). Although the proteasome's cylindrical structure is conserved, the subunit composition of the 20S protein complex varies during the evolution, and the number of subunits increases from archaebacteria to mammals. To fully characterize the 20S proteasome subunit composition and understand the subunit functions, we, the authors of this chapter, have developed and employed various mass spectrometry (MS)-based approaches to generate a comprehensive profile of the 20S proteasomes from rat liver and Tropanosoma brucei. We have identified 7 alpha and 10 beta subunits, including 7 essential and 3 nonessential beta subunits from rat 20S proteasome complex using two-dimensional (2-D) gel electrophoresis and tandem MS (MS/MS). In addition, multiple isoforms of most of the subunits were determined; indicating the composition of rat 20S proteasome complex was much more complicated than expected. Further analysis of the intact protein molecular weight of each subunit using LC-MS confirmed the heterogeneous population of the 20S proteasome and revealed that many of the experimental measured molecular weights do not correspond well with the theoretical values deduced from the sequences in protein databases. This finding is mostly due to the sequence errors in the protein databases and possible posttranslational modifications. Although the protein sequences of rat 20S proteasome are present in the databases, the sequences of the 20S proteasome from T. brucei were not available at the time when the analysis was carried out. To determine the subunit composition of the 20S proteasome from T. brucei, we developed a homology-based database searching tool to identify unknown proteins based on the novel sequences determined by de novo sequencing using MS/MS. As a result, 14 subunits (7 alpha and 7 beta) were identified on the 2-D gel, which was later confirmed by the full-length sequences. Using the same approach, we also identified and characterized an activator protein, PA26, from T. brucei. The purified recombinant PA26 self-assembles into a heptamer ring, which can bind and activate the 20S proteasome from T. brucei as well as rat. Compared to the human PA28 complex, PA26 may be the prototype activator protein involved in proteasomal protein degradation. Therefore, the MS-based strategy developed here for identification of the known and unknown protein complexes can be generalized for the study of other protein complexes.

Amino Acid Sequence↗

Columba: an integrated database of proteins, structures, and annotations.

BACKGROUND: Structural and functional research often requires the computation of sets of protein structures based on certain properties of the proteins, such as sequence features, fold classification, or functional annotation. Compiling such sets using current web resources is tedious because the necessary data are spread over many different databases. To facilitate this task, we have created COLUMBA, an integrated database of annotations of protein structures. DESCRIPTION: COLUMBA currently integrates twelve different databases, including PDB, KEGG, Swiss-Prot, CATH, SCOP, the Gene Ontology, and ENZYME. The database can be searched using either keyword search or data source-specific web forms. Users can thus quickly select and download PDB entries that, for instance, participate in a particular pathway, are classified as containing a certain CATH architecture, are annotated as having a certain molecular function in the Gene Ontology, and whose structures have a resolution under a defined threshold. The results of queries are provided in both machine-readable extensible markup language and human-readable format. The structures themselves can be viewed interactively on the web. CONCLUSION: The COLUMBA database facilitates the creation of protein structure data sets for many structure-based studies. It allows to combine queries on a number of structure-related databases not covered by other projects at present. Thus, information on both many and few protein structures can be used efficiently. The web interface for COLUMBA is available at http://www.columba-db.de.

Base Sequence↗

Mass spectrometric methods for generation of protein mass database used for bacterial identification.

The availability of a suitable database is critical in a proteomic approach for bacterial identification by mass spectrometry (MS). The major limitation of the present public proteome database is the lack of extensive low-mass bacterial protein entries with masses experimentally verified for most bacteria. Here, we present a method based on mass spectrometry to create protein mass tables specifically tailored for bacterial identification. Several issues related to the detection of bacterial proteins for the purpose of database creation are addressed. Three species of bacteria, namely, Escherichia coli, Bacillus megaterium, and Citrobacter freundii, which can be found in the ambient environment, were chosen for this study. Direct matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) MS analysis of each bacterial extract reveals 20-29 protein components in the mass range from 2000 to 20,000 Da. HPLC fractionation of bacterial extracts followed by off-line MALDI-TOF analysis of individual fractions detects 156-423 components. Analysis of the extracts by HPLC/electrospray ionization MS shows the number of detectable proteins in the range of 46-59. Although a number of components were common to the three detection schemes employed, some unique components were found using each of these techniques. In addition, for E. coli where a large proteome database exists in the public domain, a number of masses detected by the mass spectrometric methods do not match with the proteome database. Compared to the public proteome database, the mass tables generated in this work are demonstrated to be more useful for bacterial identification in an application where the bacteria of interest have limited protein entries in the public database. The implication of this work for future development of a comprehensive mass database is discussed.

Bacteria↗