Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

An efficient disk based data structure for rapid searching of quantitative two-dimensional gel databases.

Fast access of two-dimensional (2-D) gel quantitative databases is important for rapid searching for protein differences between sets of 2-D gels from an experiment. The GELLAB-II system organizes corresponding spots from the gels in the database into reference or "Rspot" sets. These Rspot numeric names index fixed regions in the paged composite gel database file. This is adequate for an existing database, but has several problems. (i) Building the initial database requires guessing how much disk space to pre-allocate for each corresponding spot (i.e. spots from different gels). If it ever runs out of pre-allocated space during this process, it must expand the size of each corresponding set of spots copying the old database data into the new in-place on the disk. (ii) When adding new gels or editing the database, if a new spot is created, the system may also go into this expansion mode. The time spent and wasted disk space can be appreciable--depending on the size of the database (order of 100 gel database). (iii) Because each set of corresponding spots is the same size, we waste space in most spot sets since they do not require the additional space a few spot sets require which contain additional fragmented spots. We present a new low-level disk object-based structure and algorithm, paged indexed buckets (PIB), which optimizes disk space usage while having similar retrieval speed to the original method.

Algorithms↗

Spermatocytes and round spermatids of rat testis: protein patterns.

Spermatogenesis is a process in the testis that involves meiotic cell division and spermiogenesis. The mechanisms of regulation and its associated proteins are mostly unknown. This publication shows the two-dimensional (2-D) gel electrophoresis protein map obtained from rat testis using nonlinear 3.5-10 immobilized pH gradients for the first-dimensional separation. Eighteen proteins were successfully identified in the SWISS-PROT protein database using amino acid analysis of proteins recovered from polyvinylidene difluoride (PVDF) membranes and verified for one of them by comparison with Anderson's rat liver reference map. Fourteen new polypeptides were identified and four were previously known. Two of these new proteins were closely related to the spermatogenetic process. T-complex protein 1 is expressed in large amounts in germ cells. Androgen-dependent sperm-coating glycoprotein is secreted by epididymal cells. In order to detect changes in protein expression during meiosis and spermiogenesis, spermatocytes and round spermatid cell populations were purified by centrifugal elutriation and compared. In this way several proteins not found in the spermatocyte 2-D images could be high-lighted. The sperm-coating glycoprotein was thus shown to be present in large amounts in round spermatids.

Adolescent↗

Protein folding: from the levinthal paradox to structure prediction.

This article is a personal perspective on the developments in the field of protein folding over approximately the last 40 years. In addition to its historical aspects, the article presents a view of the principles of protein folding with particular emphasis on the relationship of these principles to the problem of protein structure prediction. It is argued that despite much that is new, the essential elements of our current understanding of protein folding were anticipated by researchers many years ago. These elements include the recognition of the central importance of the polypeptide backbone as a determinant of protein conformation, hierarchical protein folding, and multiple folding pathways. Important areas of progress include a detailed characterization of the folding pathways of a number of proteins and a fundamental understanding of the physical chemical forces that determine protein stability. Despite these developments, fold prediction algorithms still encounter difficulties in identifying the correct fold for a given sequence. This may be due to the possibility that the free energy differences between at least a few alternate conformations of many proteins are not large. Significant progress in protein structure prediction has been due primarily to the explosive growth of sequence and structural databases. However, further progress is likely to depend in part on the ability to combine information available from databases with principles and algorithms derived from physical chemical studies of protein folding. An approach to the integration of the two areas is outlined with specific reference to the PrISM program that is a fully integrated sequence/structural-analysis/fold-recognition/homology model building software system.

Models, Molecular↗

Gene indexing: characterization and analysis of NLM's GeneRIFs.

We present an initial analysis of the National Library of Medicine's (NLM) Gene Indexing initiative. Gene Indexing occurs at the time of indexing for all 4600 journals and over 500,000 articles added to PubMed/MEDLINE each year. Gene Indexing links articles about the basic biology of a gene or protein within eight model organisms to a specific record in the NLM's LocusLink database of gene products. The result is an entry called a Gene Reference Into Function (GeneRIF) within the LocusLink database. We analyzed the numbers of GeneRIFs produced in the first year of GeneRIF production. 27,645 GeneRIFs were produced, pertaining to 9126 loci over eight model organisms. 60% of these were associated with human genes and 27% with mouse genes. About 80% discuss genes with an established MeSH Heading or other MeSH term. We developed a prototype functional alerting system for researchers based on the GeneRIFs, and a strategy to find all of the literature related to genes. We conclude that the Gene Indexing initiative adds considerable value to the life sciences research community.

Abstracting and Indexing↗

The RESID database of protein structure modifications: 2000 update.

The RESID Database contains supplemental information on post-translational modifications for the standardized annotations appearing in the PIR-International Protein Sequence Database. The RESID Database includes: systematic and frequently observed alternate names, Chemical s Service registry numbers, atomic formulas and weights, enzyme activities, indicators for N-terminal, C-terminal or peptide chain cross-link modifications, keywords, literature citations with database cross-references, structural diagrams and molecular models. Since 1995 updates of the RESID Database have appeared as often as weekly, and full releases appear quarterly. The database is freely accessible through the PIR Web site http://pir.georgetown.edu/pirwww/dbinfo/resid.html and by FTP.

Databases, Factual↗

Ensemble docking of multiple protein structures: considering protein structural variations in molecular docking.

One approach to incorporate protein flexibility in molecular docking is the use of an ensemble consisting of multiple protein structures. Sequentially docking each ligand into a large number of protein structures is computationally too expensive to allow large-scale database screening. It is challenging to achieve a good balance between docking accuracy and computational efficiency. In this work, we have developed a fast, novel docking algorithm utilizing multiple protein structures, referred to as ensemble docking, to account for protein structural variations. The algorithm can simultaneously dock a ligand into an ensemble of protein structures and automatically select an optimal protein structure that best fits the ligand by optimizing both ligand coordinates and the conformational variable m, where m represents the m-th structure in the protein ensemble. The docking algorithm was validated on 10 protein ensembles containing 105 crystal structures and 87 ligands in terms of binding mode and energy score predictions. A success rate of 93% was obtained with the criterion of root-mean-square deviation <2.5 A if the top five orientations for each ligand were considered, comparable to that of sequential docking in which scores for individual docking are merged into one list by re-ranking, and significantly better than that of single rigid-receptor docking (75% on average). Similar trends were also observed in binding score predictions and enrichment tests of virtual database screening. The ensemble docking algorithm is computationally efficient, with a computational time comparable to that for docking a ligand into a single protein structure. In contrast, the computational time for the sequential docking method increases linearly with the number of protein structures in the ensemble. The algorithm was further evaluated using a more realistic ensemble in which the corresponding bound protein structures of inhibitors were excluded. The results show that ensemble docking successfully predicts the binding modes of the inhibitors, and discriminates the inhibitors from a set of noninhibitors with similar chemical properties. Although multiple experimental structures were used in the present work, our algorithm can be easily applied to multiple protein conformations generated by computational methods, and helps improve the efficiency of other existing multiple protein structure(MPS)-based methods to accommodate protein flexibility.

Algorithms↗

A novel approach to remote homology detection: jumping alignments.

We describe a new algorithm for protein classification and the detection of remote homologs. The rationale is to exploit both vertical and horizontal information of a multiple alignment in a well-balanced manner. This is in contrast to established methods such as profiles and profile hidden Markov models which focus on vertical information as they model the columns of the alignment independently and to family pairwise search which focuses on horizontal information as it treats given sequences separately. In our setting, we want to select from a given database of "candidate sequences" those proteins that belong to a given superfamily. In order to do so, each candidate sequence is separately tested against a multiple alignment of the known members of the superfamily by means of a new jumping alignment algorithm. This algorithm is an extension of the Smith-Waterman algorithm and computes a local alignment of a single sequence and a multiple alignment. In contrast to traditional methods, however, this alignment is not based on a summary of the individual columns of the multiple alignment. Rather, the candidate sequence is at each position aligned to one sequence of the multiple alignment, called the "reference sequence." In addition, the reference sequence may change within the alignment, while each such jump is penalized. To evaluate the discriminative quality of the jumping alignment algorithm, we compare it to profiles, profile hidden Markov models, and family pairwise search on a subset of the SCOP database of protein domains. The discriminative quality is assessed by median false positive counts (med-FP-counts). For moderate med-FP-counts, the number of successful searches with our method is considerably higher than with the competing methods.

Algorithms↗

Reference points for comparisons of two-dimensional maps of proteins from different human cell types defined in a pH scale where isoelectric points correlate with polypeptide compositions.

A highly reproducible, commercial and nonlinear, wide-range immobilized pH gradient (IPG) was used to generate two-dimensional (2-D) gel maps of [35S]methionine-labeled proteins from noncultured, unfractionated normal human epidermal keratinocytes. Forty one proteins, common to most human cell types and recorded in the human keratinocyte 2-D gel protein database were identified in the 2-D gel maps and their isoelectric points (pI) were determined using narrow-range IPGs. The latter established a pH scale that allowed comparisons between 2-D gel maps generated either with other IPGs in the first dimension or with different human protein samples. Of the 41 proteins identified, a subset of 18 was defined as suitable to evaluate the correlation between calculated and experimental pI values for polypeptides with known composition. The variance calculated for the discrepancies between calculated and experimental pI values for these proteins was 0.001 pH units. Comparison of the values by the t-test for dependent samples (paired test) gave a p-level of 0.49, indicating that there is no significant difference between the calculated and experimental pI values. The precision of the calculated values depended on the buffer capacity of the proteins, and on average, it improved with increased buffer capacity. As shown here, the widely available information on protein sequences cannot, a priori, be assumed to be sufficient for calculating pI values because post-translational modifications, in particular N-terminal blockage, pose a major problem. Of the 36 proteins analyzed in this study, 18-20 were found to be N-terminally blocked and of these only 6 were indicated as such in databases. The probability of N-terminal blockage depended on the nature of the N-terminal group. Twenty six of the proteins had either M, S or A as N-terminal amino acids and of these 17-19 were blocked. Only 1 in 10 proteins containing other N-terminal groups were blocked.

Amino Acid Sequence↗

STRING: a database of predicted functional associations between proteins.

Functional links between proteins can often be inferred from genomic associations between the genes that encode them: groups of genes that are required for the same function tend to show similar species coverage, are often located in close proximity on the genome (in prokaryotes), and tend to be involved in gene-fusion events. The database STRING is a precomputed global resource for the exploration and analysis of these associations. Since the three types of evidence differ conceptually, and the number of predicted interactions is very large, it is essential to be able to assess and compare the significance of individual predictions. Thus, STRING contains a unique scoring-framework based on benchmarks of the different types of associations against a common reference set, integrated in a single confidence score per prediction. The graphical representation of the network of inferred, weighted protein interactions provides a high-level view of functional linkage, facilitating the analysis of modularity in biological processes. STRING is updated continuously, and currently contains 261 033 orthologs in 89 fully sequenced genomes. The database predicts functional interactions at an expected level of accuracy of at least 80% for more than half of the genes; it is online at http://www.bork.embl-heidelberg.de/STRING/.

Algorithms↗

Human 2-D PAGE databases for proteome analysis in health and disease: http://biobase.dk/cgi-bin/celis.

Human 2-D PAGE Databases established at the Danish Centre for Human Genome Research are now available on the World Wide Web (http://biobase.dk/cgi-bin/celis). The databanks, which offer a comprehensive approach to the analysis of the human proteome both in health and disease, contain data on known and unknown proteins recorded in various IEF and NEPHGE 2-D PAGE reference maps (non-cultured keratinocytes, non-cultured transitional cell carcinomas, MRC-5 fibroblasts and urine). One can display names and information on specific protein spots by clicking on the image of the gel representing the 2-D gel map in which one is interested. In addition, the database can be searched by protein name, keywords or organelle or cellular component. The entry files contain links to other databases such as Medline, Swiss-Prot, PIR, PDB, CySPID, OMIM, Methabolic pathways, etc. The on-line information is updated regularly.

Computer Communication Networks↗

Data management and preliminary data analysis in the pilot phase of the HUPO Plasma Proteome Project.

The pilot phase of the HUPO Plasma Proteome Project (PPP) is an international collaboration to catalog the protein composition of human blood plasma and serum by analyzing standardized aliquots of reference serum and plasma specimens using a variety of experimental techniques. Data management for this project included collection, integration, analysis, and dissemination of findings from participating organizations world-wide. Accomplishing this task required a communication and coordination infrastructure specific enough to support meaningful integration of results from all participants, but flexible enough to react to changing requirements and new insights gained during the course of the project and to allow participants with varying informatics capabilities to contribute. Challenges included integrating heterogeneous data, reducing redundant information to minimal identification sets, and data annotation. Our data integration workflow assembles a minimal and representative set of protein identifications, which account for the contributed data. It accommodates incomplete concordance of results from different laboratories, ambiguity and redundancy in contributed identifications, and redundancy in the protein sequence databases. Recommendations of the PPP for future large-scale proteomics endeavors are described.

Algorithms↗

The Yeast Protein Database (YPD): a curated proteome database for Saccharomyces cerevisiae.

The Yeast Protein Database (YPD) is a curated database for the proteome of Saccharomyces cerevisiae . It consists of approximately 6000 Yeast Protein Reports, one for each of the known or predicted yeast proteins. Each Yeast Protein Report is a one-page presentation of protein properties, annotation lines that summarize findings from the literature, and references. In the past year, the number of annotation lines has grown from 25 000 to approximately 35 000, and the number of articles curated has grown from approximately 3500 to >5000. Recently, new data types have been included in YPD: protein-protein interactions, genetic interactions, and regulators of gene expression. Finally, a new layer of information, the YPD Protein Minireviews, has recently been introduced. The Yeast Protein Database can be found on the Web at http://www.proteome.com/YPDhome. html

Databases, Factual↗

InterPro: an integrated documentation resource for protein families, domains and functional sites.

The exponential increase in the submission of nucleotide sequences to the nucleotide sequence database by genome sequencing centres has resulted in a need for rapid, automatic methods for classification of the resulting protein sequences. There are several signature and sequence cluster-based methods for protein classification, each resource having distinct areas of optimum application owing to the differences in the underlying analysis methods. In recognition of this, InterPro was developed as an integrated documentation resource for protein families, domains and functional sites, to rationalise the complementary efforts of the individual protein signature database projects. The member databases - PRINTS, PROSITE, Pfam, ProDom, SMART and TIGRFAMs - form the InterPro core. Related signatures from each member database are unified into single InterPro entries. Each InterPro entry includes a unique accession number, functional descriptions and literature references, and links are made back to the relevant member database(s). Release 4.0 of InterPro (November 2001) contains 4,691 entries, representing 3,532 families, 1,068 domains, 74 repeats and 15 sites of post-translational modification (PTMs) encoded by different regular expressions, profiles, fingerprints and hidden Markov models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (2,141,621 InterPro hits from 586,124 SWISS-PROT and TrEMBL protein sequences). The database is freely accessible for text- and sequence-based searches.

Algorithms↗

Information transfer between large and small two-dimensional polyacrylamide gel electrophoresis.

To determine the feasibility of data transfer, an interlaboratory comparison was conducted on colon carcinoma cell line (DLD-1) proteins resolved by two-dimensional polyacrylamide gel electrophoresis either on small (6 x 7 cm) or large (16x18 cm) gels. The gels were silver-stained and scanned by laser densitometry, and the image obtained was analyzed using Melanie software. The number of spots detected was 1337+/-161 vs. 2382+/-176 for small vs. large format gels, respectively. After gel calibration using landmarks determined using pl and Mr markers, large- and small-format gels were matched and 712+/-36 proteins were found on both types of gels. Having performed accurate gel matching it was possible to acquire additional information after accessing a 2-D PAGE reference database (http://www.expasy.ch/ cgibin/map2/def?DLD1_HUMAN). Thus, the difference in gel size is not an obstacle for data transfer. This will facilitate exchanges between laboratories or consultation concerning existing databases.

Adenocarcinoma↗

Development of a liquid chromatography-tandem mass spectrometry method using capillary liquid chromatography and nanoelectrospray ionization-quadrupole time-of-flight hybrid mass spectrometer for the detection of milk allergens.

Liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis of the tryptic digest of a cleaned-up food matrix extract was used for the detection of milk allergens. The emphasis of this study was on casein, which is the most abundant milk protein and is also considered the most allergenic. A sample cleanup method was developed using an ion exchange column and centriprep device. Cookies spiked with milk powder from 0 to 1250 ppm were extracted, cleaned up, and either digested directly by trypsin or further cleaned up by gel electrophoresis before digestion. The peptide mixture was analyzed on a capillary LC-quadrupole time-of-flight system. Two marker peptides from alphaS1-casein were identified and used for prescreening. The MS/MS data from the mass spectrometry system were processed with Masslynx v4.0 and submitted for database search using either ProteinLynx Global Server or Mascot for protein identification. The LC-MS/MS method, using casein enzyme-linked immunosorbent assay as a reference, was tested on the cookie matrix and was extended to other sample matrices. There were good agreements between the two. This LC-MS/MS method provides a valuable confirmatory method for the presence of casein. It also allows the simultaneous detection of other milk allergens.

Allergens↗

The EBI SRS server--recent developments.

MOTIVATION: The current data explosion is intractable without advanced data management systems. The numerous data sets become really useful when they are interconnected under a uniform interface--representing the domain knowledge. The SRS has become an integration system for both data retrieval and applications for data analysis. It provides capabilities to search multiple databases by shared attributes and to query across databases fast and efficiently. RESULTS: Here we present recent developments at the EBI SRS server (http://srs.ebi.ac.uk). The EBI SRS server contains today more than 130 biological databases and integrates more than 10 applications. It is a central resource for molecular biology data as well as a reference server for the latest developments in data integration. One of the latest additions to the EBI SRS server is the InterPro database-Integrated Resource of Protein Domains and Functional Sites. Distributed in XML format it became a turning point in low level XML-SRS integration. We present InterProScan as an example of data analysis applications, describe some advanced features of SRS6, and introduce the SRSQuickSearch JavaScript interfaces to SRS.

Computational Biology↗

Distance-scaled, finite ideal-gas reference state improves structure-derived potentials of mean force for structure selection and stability prediction.

The distance-dependent structure-derived potentials developed so far all employed a reference state that can be characterized as a residue (atom)-averaged state. Here, we establish a new reference state called the distance-scaled, finite ideal-gas reference (DFIRE) state. The reference state is used to construct a residue-specific all-atom potential of mean force from a database of 1011 nonhomologous (less than 30% homology) protein structures with resolution less than 2 A. The new all-atom potential recognizes more native proteins from 32 multiple decoy sets, and raises an average Z-score by 1.4 units more than two previously developed, residue-specific, all-atom knowledge-based potentials. When only backbone and C(beta) atoms are used in scoring, the performance of the DFIRE-based potential, although is worse than that of the all-atom version, is comparable to those of the previously developed potentials on the all-atom level. In addition, the DFIRE-based all-atom potential provides the most accurate prediction of the stabilities of 895 mutants among three knowledge-based all-atom potentials. Comparison with several physical-based potentials is made.

Computational Biology↗

At what scale should microarray data be analyzed?

INTRODUCTION: The hybridization intensities derived from microarray experiments, for example Affymetrix's MAS5 signals, are very often transformed in one way or another before statistical models are fitted. The motivation for performing transformation is usually to satisfy the model assumptions such as normality and homogeneity in variance. Generally speaking, two types of strategies are often applied to microarray data depending on the analysis need: correlation analysis where all the gene intensities on the array are considered simultaneously, and gene-by-gene ANOVA where each gene is analyzed individually. AIM: We investigate the distributional properties of the Affymetrix GeneChip signal data under the two scenarios, focusing on the impact of analyzing the data at an inappropriate scale. METHODS: The Box-Cox type of transformation is first investigated for the strategy of pooling genes. The commonly used log-transformation is particularly applied for comparison purposes. For the scenario where analysis is on a gene-by-gene basis, the model assumptions such as normality are explored. The impact of using a wrong scale is illustrated by log-transformation and quartic-root transformation. RESULTS: When all the genes on the array are considered together, the dependent relationship between the expression and its variation level can be satisfactorily removed by Box-Cox transformation. When genes are analyzed individually, the distributional properties of the intensities are shown to be gene dependent. Derivation and simulation show that some loss of power is incurred when a wrong scale is used, but due to the robustness of the t-test, the loss is acceptable when the fold-change is not very large.

Algorithms↗