Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Towards developing a protein infrared spectra databank (PISD) for proteomics research.

Fourier transform infrared (FTIR) spectroscopy is an attractive tool for proteomics research as it can be used to rapidly characterize protein secondary structure in aqueous solution. Recently, a number of secondary structure prediction methods based on reference sets of FTIR spectra from proteins with known structure from X-ray crystallography have been suggested. These prediction methods, often referred to as pattern recognition based approaches, demonstrated good prediction accuracy using some error measure, e.g., the standard error of prediction (SEP). However, to avoid possible adverse effects from differences in recording, the analysis has been mostly based on reference sets of FTIR spectra from proteins recorded in one laboratory only. As a result, these studies were based on reference sets of FTIR spectra from a limited number of proteins. Pattern recognition based approaches, however, rely on reference sets of FTIR spectra from as many proteins as possible representing all possible band shape variation to be related to the diversity of protein structural classes. Hence, if we want to build reliable pattern recognition based systems to support proteomics research, which are capable of making good predictions from spectral data of any unknown protein, one common goal should be to build a comprehensive protein infrared spectra databank (PISD) containing FTIR spectra of proteins of known structure. We have started the process of developing a comprehensive PISD composed of spectra recorded in different laboratories. As part of this work, here we investigate possible effects on prediction accuracy achieved by a neural network analysis when using reference sets composed of FTIR spectra from different laboratories. Surprisingly low magnitude of difference in SEPs throughout all our experiments suggests that FTIR spectra recorded in different laboratories may be safely combined into one reference set with only minor deterioration of prediction accuracy in the worst case.

Algorithms↗

Improved classification of mass spectrometry database search results using newer machine learning approaches.

Manual analysis of mass spectrometry data is a current bottleneck in high throughput proteomics. In particular, the need to manually validate the results of mass spectrometry database searching algorithms can be prohibitively time-consuming. Development of software tools that attempt to quantify the confidence in the assignment of a protein or peptide identity to a mass spectrum is an area of active interest. We sought to extend work in this area by investigating the potential of recent machine learning algorithms to improve the accuracy of these approaches and as a flexible framework for accommodating new data features. Specifically we demonstrated the ability of boosting and random forest approaches to improve the discrimination of true hits from false positive identifications in the results of mass spectrometry database search engines compared with thresholding and other machine learning approaches. We accommodated additional attributes obtainable from database search results, including a factor addressing proton mobility. Performance was evaluated using publically available electrospray data and a new collection of MALDI data generated from purified human reference proteins.

Amino Acid Sequence↗

SWISS-2DPAGE: a database of two-dimensional gel electrophoresis images.

This publication presents the SWISS-2DPAGE database which gathers data on proteins identified on various two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) maps. Each SWISS-2DPAGE entry contains data on one protein, including mapping procedures, physiological and pathological data and bibliographical references, as well as several 2-D PAGE images showing the protein location. Links are also provided to other databases such as SWISS-PROT, EMBL, PROSITE and OMIM. The database has been set up on a server which may be accessed from any computer connected to the internet and it also makes it possible to display the theoretical location of proteins, the positions of which are not yet known on the 2-D PAGE.

Databases, Factual↗

The mammalian protein-protein interaction database and its viewing system that is linked to the main FANTOM2 viewer.

Here, we describe the development of a mammalian protein-protein interaction (PPI) database and of a PPI Viewer application to display protein interaction networks (http://fantom21.gsc.riken.go.jp/PPI/). In the database, we stored the mammalian PPIs identified through our PPI assays (internal PPIs), as well as those we extracted and processed (external PPIs) from publicly available data sources, the DIP and BIND databases and MEDLINE abstracts by using FACTS, a new functional inference and curation system. We integrated the internal and external PPIs into the PPI database, which is linked to the main FANTOM2 viewer. In addition, we incorporated into the PPI Viewer information regarding the luciferase reporter activity of internal PPIs and the data confidence of external PPIs; these data enable visualization and evaluation of the reliability of each interaction. Using the described system, we successfully identified several interactions of biological significance. Therefore, the PPI Viewer is a useful tool for exploring FANTOM2 clone-related protein interactions and their potential effects on signaling and cellular communication.

Animals↗

Use of multiple profiles corresponding to a sequence alignment enables effective detection of remote homologues.

MOTIVATION: Position specific scoring matrices (PSSMs) corresponding to aligned sequences of homologous proteins are commonly used in homology detection. A PSSM is generated on the basis of one of the homologues as a reference sequence, which is the query in the case of PSI-BLAST searches. The reference sequence is chosen arbitrarily while generating PSSMs for reverse BLAST searches. In this work we demonstrate that the use of multiple PSSMs corresponding to a given alignment and variable reference sequences is more effective than using traditional single PSSMs and hidden Markov models. RESULTS: Searches for proteins with known 3-D structures have been made against three databases of protein family profiles corresponding to known structures: (1) One PSSM per family; (2) multiple PSSMs corresponding to an alignment and variable reference sequences for every family; and (3) hidden Markov models. A comparison of the performances of these three approaches suggests that the use of multiple PSSMs is most effective. CONTACT: ns@mbu.iisc.ernet.in.

Algorithms↗

Ongoing development of two-dimensional polyacrylamide gel electrophoresis data standards.

We present an approach toward standardizing two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) data in support of developing a globally relevant proteomics consensus in order to provide more efficient database querying and data comparisons through the establishment of the necessary definitions and interdisciplinary reference fields for both the 2-D PAGE community, particularly in the proteomics area, and the clinical and experimental biological research communities, in general. This article covers the need for unifying the 2-D PAGE data through a common data repository, and its usefulness in data standards and data interoperability.

Databases, Protein↗

Identification of cellular changes associated with increased production of human growth hormone in a recombinant Chinese hamster ovary cell line.

A proteomics approach was used to identify the proteins potentially implicated in the cellular response concomitant with elevated production levels of human growth hormone in a recombinant Chinese hamster ovary (CHO) cell line following exposure to 0.5 mM butyrate and 80 microM zinc sulphate in the production media. This involved incorporation of two-dimensional (2-D) gel electrophoresis and protein identification by a combination of N-terminal sequencing, matrix-assisted laser desorption/ionisation-time of flight mass spectrometry, amino acid analysis and cross species database matching. From these identifications a CHO 2-D reference map and annotated database have been established. Metabolic labelling and subsequent autoradiography showed the induction of a number of cellular proteins in response to the media additives butyrate and zinc sulphate. These were identified as GRP75, enolase and thioredoxin. The chaperone proteins GRP78, HSP90, GRP94 and HSP70 were not up-regulated under these conditions.

Amino Acids↗

Active Sequences Collection (ASC) database: a new tool to assign functions to protein sequences.

Active Sequences Collection (ASC) is a collection of amino acid sequences, with an unique feature: only short sequences are collected, with a demonstrated biological activity. The current version of ASC consists of three sections: DORRS, a collection of active RGD-containing peptides; TRANSIT, a collection of protein regions active as substrates of transglutaminase enzyme (TGase), and BAC, a collection of short peptides with demonstrated biological activity. Literature references for each entry are reported, as well as cross references to other databases, when available. The current version of ASC includes more than 800 different entries. The main scope of this collection is to offer a new tool to investigate the structural features of protein active sites, additionally to similarity searches against large protein databases or searching for known functional patterns. ASC database is available at the web address http://crisceb.unina2.it/ASC/ which also offers a dedicated query interface to compare user-defined protein sequences with the database, as well as an updating interface to allow contribution of new referenced active sequences.

Amino Acid Sequence↗

Immunodeficiency mutation databases (IDbases).

Primary immunodeficiencies (IDs) are a heterogenic group of inherited disorders of the immune system. Immunodeficiency patients have increased susceptibility to recurrent and persistent, even life-threatening infections. Mutations in a large number of genes can cause defects in different cellular functions and lead to impaired immune response. To date, approximately 150 IDs and more than 100 affected genes have been identified. ID-related genes are distributed throughout the genome, and diseases can be inherited in an X-linked, an autosomal recessive, or an autosomal dominant way. We have collected ID mutation data into locus-specific patient-related mutation databases, IDbases (http://bioinf.uta.fi/IDbases). Mutations are described at DNA, mRNA, and protein levels with links to reference sequences and reference articles. The mutation data has been collated into entries along with some clinical information. IDbases offer an easy way, e.g., to find recently identified mutations, to reveal genotype-phenotype correlations, and to discover a specific mutation or to examine the most common mutations in a single immunodeficiency related gene. At the moment we have databases for 107 ID genes with 4,140 public patient entries. An exhaustive statistical analysis of mutation data from the IDbases was made. Missense and nonsense mutations are the most common mutation types, and the most common single substitution is a nonsense mutation from tryptophan to a stop codon. Arginine is the most mutated as well as the most abundant mutant amino acid.

Amino Acid Sequence↗

A comprehensive two-dimensional gel protein database of noncultured unfractionated normal human epidermal keratinocytes: towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

A two-dimensional (2-D) gel database of cellular proteins from noncultured, unfractionated normal human epidermal keratinocytes has been established. A total of 2651 [35S]methionine-labeled cellular proteins (1868 isoelectric focusing, 783 nonequilibrium pH gradient electrophoresis) were resolved and recorded using computer-aided 2-D gel electrophoresis. The protein numbers in this database differ from those reported in an earlier version due to changes in the scanning hardware (Celis et al., Electrophoresis 1990, 11, 242-254). Annotation categories reported include: "protein name" (listing 207 known proteins in alphabetical order), "basal cell markers", "differentiation markers", "proteins highly up-regulated in psoriatic skin", "microsequenced proteins" and "human autoantigens". For reference, we have also included 2-D gel (isoelectric focusing) patterns of cultured normal and psoriatic keratinocytes, melanocytes, fibroblasts, dermal microvascular endothelial cells, peripheral blood mononuclear cells and sweat duct cells. The keratinocyte 2-D gel protein database will be updated yearly in the November issue of Electrophoresis.

Biopsy↗

Protein affinity map of chemical space.

Affinity fingerprinting is a quantitative method for mapping chemical space based on binding preferences of compounds for a reference panel of proteins. An effective reference panel of <20 proteins can be empirically selected which shows differential interaction with nearly all compounds. By using this map to iteratively sample the chemical space, identification of active ligands from a library of 30,000 candidate compounds has been accomplished for a wide spectrum of specific protein targets. In each case, <200 compounds were directly assayed against the target. Further, analysis of the fingerprint database suggests a strategy for effective selection of affinity chromatography ligands and scaffolds for combinatorial chemistry. With such a system, the large numbers of potential therapeutic targets emerging from genome research can be categorized according to ligand binding properties, complementing sequence based classification.

Chromatography, Affinity↗

Comparative assessment of large-scale data sets of protein-protein interactions.

Comprehensive protein protein interaction maps promise to reveal many aspects of the complex regulatory network underlying cellular function. Recently, large-scale approaches have predicted many new protein interactions in yeast. To measure their accuracy and potential as well as to identify biases, strengths and weaknesses, we compare the methods with each other and with a reference set of previously reported protein interactions.

Bias↗

Methodology and application of gc-ms to study altered organic binding media from objects of the Kunsthistorisches Museum, Vienna.

Within the Kunsthistorisches Museum (KHM), Vienna, three off-line GC-MS analytical procedures for the identification of natural organic media have been refined, tested, and validated on a series of reference materials (partly artificially aged) to apply this knowledge for investigations of original, historic works of art from the museum's collections. At first, a set of artificially aged mockups has been prepared and a reference database has been built up for the identification of drying oils, resins, waxes, proteins and polysaccharides. Some interesting observations concerning the alteration of the composition of these organic media during different ageing steps are presented in the following text. In addition, some selected examples for the application of the refined techniques for the analysis of real samples from various museum objects are shown.

Art↗

Construction, analysis, and beta-glucanase screening of a bacterial artificial chromosome library from the large-bowel microbiota of mice.

A metagenomic (community genomic) library consisting of 5,760 bacterial artificial chromosome clones was prepared in Escherichia coli DH10B from DNA extracted from the large-bowel microbiota of BALB/c mice. DNA inserts detected in 61 randomly chosen clones averaged 55 kbp (range, 8 to 150 kbp) in size. A functional screen of the library for beta-glucanase activity was conducted using lichenin agar plates and Congo red solution. Three clones with beta-glucanase activity were detected. The inserts of these three clones were sequenced and annotated. Open reading frames (ORF) that encoded putative proteins with identity to glucanolytic enzymes (lichenases and laminarinases) were detected by reference to databases. Other putative genes were detected, some of which might have a role in environmental sensing, nutrient acquisition, or coaggregation. The insert DNA from two clones probably originated from uncultivated bacteria because the ORF had low sequence identity with database entries, but the genes associated with the remaining clone resembled sequences reported in Bacteroides species.

Amino Acid Sequence↗

Plant protein annotation in the UniProt Knowledgebase.

The Swiss-Prot, TrEMBL, Protein Information Resource (PIR), and DNA Data Bank of Japan (DDBJ) protein database activities have united to form the Universal Protein Resource (UniProt) Consortium. UniProt presents three database layers: the UniProt Archive, the UniProt Knowledgebase (UniProtKB), and the UniProt Reference Clusters. The UniProtKB consists of two sections: UniProtKB/Swiss-Prot (fully manually curated entries) and UniProtKB/TrEMBL (automated annotation, classification and extensive cross-references). New releases are published fortnightly. A specific Plant Proteome Annotation Program (http://www.expasy.org/sprot/ppap/) was initiated to cope with the increasing amount of data produced by the complete sequencing of plant genomes. Through UniProt, our aim is to provide the scientific community with a single, centralized, authoritative resource for protein sequences and functional information that will allow the plant community to fully explore and utilize the wealth of information available for both plant and non-plant model organisms.

Amino Acid Sequence↗

Proteins of rat serum, urine, and cerebrospinal fluid: VI. Further protein identifications and interstrain comparison.

We have investigated the biological fluids--serum, cerebrospinal fluid, and urine--of three strains of rats; the present data extend our database (also available on-line) and may be of interest for pharmacological and toxicological investigation. Specifically, we have defined reference maps of the major protein components in cerebrospinal fluid and urine. Compartment-specific isoforms were recognized for transferrin and transthyretin. Mass spectrometric data established the cleavage site of the signal peptide and identified the N-terminal blocking group of prostaglandin D synthase from rat cerebrospinal fluid. A previously undescribed member of the family of low molecular mass rat urinary proteins was characterized as containing a sequence similar, but not identical, to the N-terminal region of rat urinary protein-2 (RUP-2), and divergent from RUP-1.

Amino Acid Sequence↗

A reference map and identification of porcine testis proteins using 2-DE and MS.

The development of the testis is essential for maturation of male mammals. A complete understanding of proteins expressed in the testis will provide biological information on many reproductive dysfunctions in males. The purposes of this study were to apply a proteomic approach to investigating protein composition and to establish a 2-D PAGE reference map for porcine testis proteins. MALDI-TOF MS was performed for protein identification. When 1 mg of total proteins was assayed by 2-D PAGE and stained with colloidal CBB, more than 400 proteins with a pI of pH 3-10 and M(r) of 10-200 kDa could be detected. Protein expression varied among individuals, with CV between 4.7 and 131.5%. A total of 447 protein spots were excised for identification, among which 337 spots were identified by searching the mass spectra against the NCBInr database. Identification of the remaining 110 spots was unsuccessful. A 2-D PAGE-based porcine testis protein database has been constructed on the basis of the results and will be published on the WWW. This database should be valuable for investigating the developmental biology and pathology of porcine testis.

Animals↗