Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Proteins of rat serum, urine, and cerebrospinal fluid: VI. Further protein identifications and interstrain comparison.

We have investigated the biological fluids--serum, cerebrospinal fluid, and urine--of three strains of rats; the present data extend our database (also available on-line) and may be of interest for pharmacological and toxicological investigation. Specifically, we have defined reference maps of the major protein components in cerebrospinal fluid and urine. Compartment-specific isoforms were recognized for transferrin and transthyretin. Mass spectrometric data established the cleavage site of the signal peptide and identified the N-terminal blocking group of prostaglandin D synthase from rat cerebrospinal fluid. A previously undescribed member of the family of low molecular mass rat urinary proteins was characterized as containing a sequence similar, but not identical, to the N-terminal region of rat urinary protein-2 (RUP-2), and divergent from RUP-1.

Amino Acid Sequence↗

A reference map and identification of porcine testis proteins using 2-DE and MS.

The development of the testis is essential for maturation of male mammals. A complete understanding of proteins expressed in the testis will provide biological information on many reproductive dysfunctions in males. The purposes of this study were to apply a proteomic approach to investigating protein composition and to establish a 2-D PAGE reference map for porcine testis proteins. MALDI-TOF MS was performed for protein identification. When 1 mg of total proteins was assayed by 2-D PAGE and stained with colloidal CBB, more than 400 proteins with a pI of pH 3-10 and M(r) of 10-200 kDa could be detected. Protein expression varied among individuals, with CV between 4.7 and 131.5%. A total of 447 protein spots were excised for identification, among which 337 spots were identified by searching the mass spectra against the NCBInr database. Identification of the remaining 110 spots was unsuccessful. A 2-D PAGE-based porcine testis protein database has been constructed on the basis of the results and will be published on the WWW. This database should be valuable for investigating the developmental biology and pathology of porcine testis.

Animals↗

The COG database: new developments in phylogenetic classification of proteins from complete genomes.

The database of Clusters of Orthologous Groups of proteins (COGs), which represents an attempt on a phylogenetic classification of the proteins encoded in complete genomes, currently consists of 2791 COGs including 45 350 proteins from 30 genomes of bacteria, archaea and the yeast Saccharomyces cerevisiae (http://www.ncbi.nlm.nih. gov/COG). In addition, a supplement to the COGs is available, in which proteins encoded in the genomes of two multicellular eukaryotes, the nematode Caenorhabditis elegans and the fruit fly Drosophila melanogaster, and shared with bacteria and/or archaea were included. The new features added to the COG database include information pages with structural and functional details on each COG and literature references, improvements of the COGNITOR program that is used to fit new proteins into the COGs, and classification of genomes and COGs constructed by using principal component analysis.

Animals↗

The power and the limitations of cross-species protein identification by mass spectrometry-driven sequence similarity searches.

Mass spectrometry-driven BLAST (MS BLAST) is a database search protocol for identifying unknown proteins by sequence similarity to homologous proteins available in a database. MS BLAST utilizes redundant, degenerate, and partially inaccurate peptide sequence data obtained by de novo interpretation of tandem mass spectra and has become a powerful tool in functional proteomic research. Using computational modeling, we evaluated the potential of MS BLAST for proteome-wide identification of unknown proteins. We determined how the success rate of protein identification depends on the full-length sequence identity between the queried protein and its closest homologue in a database. We also estimated phylogenetic distances between organisms under study and related reference organisms with completely sequenced genomes that allow substantial coverage of unknown proteomes.

Animals↗

SWISS-2DPAGE, ten years later.

The SWISS-2DPAGE database was established in 1993 and is maintained collaboratively by the Swiss Institute of Bioinformatics (SIB) and the Biomedical Proteomics Research Group (BPRG) of the Geneva University Hospital. During these years, SWISS-2DPAGE underwent constant modification and improvement. Current content includes about 4000 identified spots corresponding to 1200 different protein entries in 36 reference maps from human, mouse, Arabidopsis thaliana, Dictyostelium discoideum, Escherichia coli, Saccharomyces cerevisiae and Staphylococcus aureus origins. With a high level of annotation and integration with other relevant databases, SWISS-2DPAGE is a reference source in the proteomics world. Queries to SWISS-2DPAGE database currently reach 1000 hits per day.

Animals↗

Building a protein name dictionary from full text: a machine learning term extraction approach.

BACKGROUND: The majority of information in the biological literature resides in full text articles, instead of abstracts. Yet, abstracts remain the focus of many publicly available literature data mining tools. Most literature mining tools rely on pre-existing lexicons of biological names, often extracted from curated gene or protein databases. This is a limitation, because such databases have low coverage of the many name variants which are used to refer to biological entities in the literature. RESULTS: We present an approach to recognize named entities in full text. The approach collects high frequency terms in an article, and uses support vector machines (SVM) to identify biological entity names. It is also computationally efficient and robust to noise commonly found in full text material. We use the method to create a protein name dictionary from a set of 80,528 full text articles. Only 8.3% of the names in this dictionary match SwissProt description lines. We assess the quality of the dictionary by studying its protein name recognition performance in full text. CONCLUSION: This dictionary term lookup method compares favourably to other published methods, supporting the significance of our direct extraction approach. The method is strong in recognizing name variants not found in SwissProt.

Abstracting and Indexing↗

Overview of the HUPO Plasma Proteome Project: results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database.

HUPO initiated the Plasma Proteome Project (PPP) in 2002. Its pilot phase has (1) evaluated advantages and limitations of many depletion, fractionation, and MS technology platforms; (2) compared PPP reference specimens of human serum and EDTA, heparin, and citrate-anti-coagulated plasma; and (3) created a publicly-available knowledge base (www.bioinformatics.med.umich.edu/hupo/ppp; www.ebi.ac.uk/pride). Thirty-five participating laboratories in 13 countries submitted datasets. Working groups addressed (a) specimen stability and protein concentrations; (b) protein identifications from 18 MS/MS datasets; (c) independent analyses from raw MS-MS spectra; (d) search engine performance, subproteome analyses, and biological insights; (e) antibody arrays; and (f) direct MS/SELDI analyses. MS-MS datasets had 15 710 different International Protein Index (IPI) protein IDs; our integration algorithm applied to multiple matches of peptide sequences yielded 9504 IPI proteins identified with one or more peptides and 3020 proteins identified with two or more peptides (the Core Dataset). These proteins have been characterized with Gene Ontology, InterPro, Novartis Atlas, OMIM, and immunoassay-based concentration determinations. The database permits examination of many other subsets, such as 1274 proteins identified with three or more peptides. Reverse protein to DNA matching identified proteins for 118 previously unidentified ORFs. We recommend use of plasma instead of serum, with EDTA (or citrate) for anticoagulation. To improve resolution, sensitivity and reproducibility of peptide identifications and protein matches, we recommend combinations of depletion, fractionation, and MS/MS technologies, with explicit criteria for evaluation of spectra, use of search algorithms, and integration of homologous protein matches. This Special Issue of PROTEOMICS presents papers integral to the collaborative analysis plus many reports of supplementary work on various aspects of the PPP workplan. These PPP results on complexity, dynamic range, incomplete sampling, false-positive matches, and integration of diverse datasets for plasma and serum proteins lay a foundation for development and validation of circulating protein biomarkers in health and disease.

Algorithms↗

Workshop on two-dimensional gel protein databases.

A workshop on two-dimensional gel electrophoresis (2-DE) protein database, organized by the Committee on Data for Science and Technology (CODATA) of the International Council of Scientific Unions Task Group on Biological Macromolecules, was held at the CODATA Secretariat in Paris on March 9, 1992. Eleven scientists from eight different countries represented various aspects of 2-DE analysis--namely, cellular protein database development and protein microsequencing methodologies. The purpose of the workshop was to explore means of integrating the rapidly expanding body of information on 2-DE resolved proteins from different laboratories. A major proposal emanating from the workshop was the establishment of an intermediary or "relational" 2-DE gel protein database. This intermediary database, which would catalogue pertinent information on 2-DE resolved proteins (experimental source, 2-DE loci, biological information, etc.) could be an adjunct to, and accessed through, the existing international protein sequence databanks. It would function as a pointer for researchers to the individual 2-DE protein databases where primary and more specialized 2-DE data would be housed.

Databases, Factual↗

A comprehensive dictionary of protein accession codes for complete protein accession identifier alias resolving.

In mass spectrometry-based proteomics, protein identification results usually consist of peptide sequences and database-dependent accession identifiers of the matching proteins. Often certain annotations are only available in particular databases that in turn must be queried by a certain identifier. In order to simplify and unify the tracing of identified proteins back to their original annotation information, a system capable of set-oriented mapping the different accession identifiers of proteins derived from multiple sequence database sources has been developed. This allows unification of the access to protein information and tracing to other online resources providing additional information as well as resolving cross-references of protein identifications. The interface of seqDB is available via http://www.protein-ms.de following the link to seqDB.

Database Management Systems↗

The Diatom EST Database.

The Diatom EST database provides integrated access to expressed sequence tag (EST) data from two eukaryotic microalgae of the class Bacillariophyceae, Phaeodactylum tricornutum and Thalassiosira pseudonana. The database currently contains sequences of close to 30,000 ESTs organized into PtDB, the P.tricornutum EST database, and TpDB, the T.pseudonana EST database. The EST sequences were clustered and assembled into a non-redundant set for each organism, and these non-redundant sequences were then subjected to automated annotation using similarity searches against protein and domain databases. EST sequences, clusters of contiguous sequences, their annotation and analysis with reference to the publicly available databases, and a codon usage table derived from a subset of sequences from PtDB and TpDB can all be accessed in the Diatom EST Database. The underlying RDBMS enables queries over the raw and annotated EST data and retrieval of information through a user-friendly web interface, with options to perform keyword and BLAST searches. The EST data can also be retrieved based on Pfam domains, Cluster of Orthologous Groups (COG) and Gene Ontologies (GO) assigned to them by similarity searches. The Database is available at http://avesthagen.sznbowler.com.

DNA, Algal↗

A skeletal gene database.

Systematic organization of documented data coupled with ready accessibility is of great value to research. Catalogs and databases are created specifically to meet this purpose. The Skeletal Gene Database evolves as part of the Skeletal Genome Anatomy Project (SGAP), an ongoing multi-institute collaborative effort, to study the functional genome of bone and other skeletal tissues. The primary objective of the Skeletal Gene Database is to create a contemporary list of skeletal-related genes, offering the following information for each gene: gene name, protein name, cellular function, disease(s) caused by mutation of the corresponding gene, chromosomal location, LocusLink number, gene size, exon/intron numbers, messenger RNA (mRNA) coding region size, protein size/molecular weight, Online Mendelian Inheritance in Man (OMIM) number of the gene, UniGene assignment, and PubMed reference. The database includes genes already known and published in the literature as well as novel genes not yet characterized but known to be expressed in skeletal tissue. It will be posted on the web for easy access and swift referencing. The data will be updated in tempo with current and future research, thereby providing an invaluable service to the scientific community interested in obtaining information on bone-related genes.

Bone and Bones↗

Forensic determination of ricin and the alkaloid marker ricinine from castor bean extracts.

Liquid chromatography/mass spectrometry (LC/ MS) and matrix assisted laser desorption/ionization time-of-flight (MALDI-TOF) MS methods were developed for the presumptive identification of ricin toxin and the alkaloid marker ricinine from crude plant materials. Ricin is an extremely potent poison, which is of forensic interest due to its appearance in terrorism literature and its potential for use as a homicide agent. Difficulties arise in attempting to analyze ricin because it is a large heterogeneous protein with glycosylation. The general protein identification scheme developed uses LC/MS or MALDI-TOF for size classification followed by the use of the same instrumentation for the analysis of the tryptic digest. Fragments of the digest can be searched in an online database for tentative identification of the unknown protein and then followed by comparison to authentic reference materials. LC fractionation or molecular weight cutoff filtration was used for preparation of the intact toxin before analysis. Extracts from two types of castor beans were prepared using a terrorist handbook procedure and determined to contain 1% ricin. Additionally, a forensic sample suspected to contain ricin was analyzed using the presented identification scheme (data not shown). The identification of the alkaloid ricinine by GC/MS and LC/MS was shown to be a complementary technique for the determination of castor bean extracts.

Alkaloids↗

PRECIS: protein reports engineered from concise information in SWISS-PROT.

MOTIVATION: There have been several endeavours to address the problem of annotating sequence data computationally, but the task is non-trivial and few tools have emerged that gather useful information on a given sequence, or set of sequences, in a simple and convenient manner. As more genome projects bear fruit, the mass of uncharacterized sequence data accumulating in public repositories grows ever larger. There is thus a pressing need for tools to support the process of automatic analysis and annotation of newly determined sequences. With this in mind, we have developed PRECIS, which automatically creates protein reports from sets of SWISS-PROT entries, collating results into structured reports, detailing known biological and medical information, literature and database cross-references, and relevant keywords.

Abstracting and Indexing↗

EXProt--a database for EXPerimentally verified Protein functions.

EXProt (database for EXPerimentally verified Protein functions) is a new non-redundant database containing protein sequences for which the function has been experimentally verified. It is a selection of 3976 entries from the Prokaryotes section of the EMBL Nucleotide Sequence Database, Release 66, and 375 entries from the Pseudomonas Community Annotation Project (PseudoCAP). The entries in EXProt all have a unique ID number and provide information about the organism, protein sequence, functional annotation, link to entry in original database, and if known, gene name and link to references in PubMed/Medline. The EXProt web page (http://www.cmbi.nl/EXProt) provides further details of the database and a link to a BLAST search (blastp & blastx) of the database. The EXProt entries are indexed in SRS (http://www.cmbi.nl/srs/) and can be searched by means of keywords. Authors can be reached by email (exprot(cmbi.kun.nl).

Amino Acid Sequence↗

GLYCOSCIENCES.de: an Internet portal to support glycomics and glycobiology research.

The development of glycan-related databases and bioinformatics applications is considerably lagging behind compared with the wealth of available data and software tools in genomics and proteomics. Because the encoding of glycan structures is more complex, most of the bioinformatics approaches cannot be applied to glycan structures. No standard procedures exist where glycan structures found in various species, organs, tissues or cells can be routinely deposited. In this article the concepts of the GLYCOSCIENCES.de portal are described. It is demonstrated how an efficient structure-based cross-linking of various glycan-related data originating from different resources can be accomplished using a single user interface. The structure oriented retrieval options-exact structure, substructure, motif, composition and sugar components-are discussed. The types of available data-references, composition, spatial structures, nuclear magnetic resonance (NMR) shifts (experimental and estimated), theoretically calculated fragments and Protein Database (PDB) entries-are exemplified for Man(3.) The free availability and unrestricted use of glycan-related data is an absolute prerequisite to efficiently share distributed resources. Additionally, there is an urgent need to agree to a generally accepted exchange format as well as to a common software interface. An open access repository for glyco-related experimental data will secure that the loss of primary data will be considerably reduced.

Computational Biology↗

Numerical analysis of electrophoretic protein patterns of Providencia stuartii strains from urine, wound and other clinical sources.

Eighty-six strains of Providencia stuartii (mainly of human origin) were characterized by one-dimensional SDS-PAGE of cellular proteins. The strains came from various countries; 52 were from urine, 11 from wounds, five from blood (one of these also from urine), four from ear infections, two each from faeces and sputum, one from 'alimentation' and nine from unknown sources. The protein patterns, which contained 45 to 50 discrete bands, were highly reproducible. The patterns of 46 Prov. stuartii strains (selected to represent the full range of protein pattern diversity) plus those of the type strains of the four other Providencia species were used as the basis for two numerical analyses. In the first, which included all the protein bands, the Prov. stuartii strains formed 13 clusters at the 88% S level. In the second analysis, in which the principal protein bands (in the 33.8-40.7 kDa range) were excluded, 45 of the 46 Prov. stuartii strains formed a single cluster at the 82% S level, whilst the four Providencia reference strains remained unclustered. The 40 strains of Prov. stuartii not included in the cluster analysis were assigned to a protein type by calculating their similarity with the strains in the database used for the cluster analysis. We conclude that high resolution PAGE combined with computerized analysis of protein patterns provides the basis for typing clinical strains of Prov. stuartii. Reference strains of each of the 13 PAGE types identified are available from NCTC for inclusion in future studies.

Bacterial Proteins↗

Study of human laryngeal muscle protein using two-dimensional electrophoresis and mass spectrometry.

Proteomic analysis was performed to construct a protein database for human laryngeal muscle. Thyroarytenoid (TA) muscle specimens were obtained from six post mortem cases within 24 h of death. Isoelectric focusing was performed by using immobilized pH gradient strips followed by 12% sodium dodecyl sulfate-polyacrylamide gel electrophoresis. Silver stained gels were then analyzed using PDQuest software to locate, quantify and match spots. Proteins were identified by matrix-assisted laser desorption/ionization-mass spectrometry on the basis of peptide mass fingerprinting following in-gel digestion with trypsin. Comparison of protein distribution between broad and narrow pH range gels demonstrated that 75% of all protein spots from human TA muscle were located within the pH range 5-8, and between mass 15-120 kDa. Based on peptide mass fingerprinting, 75 proteins were identified and classified into six functional groups. These include membrane proteins (8.5%), cytoskeletal and myofibrillar proteins (14.6%), energy production proteins (28%), proteins associated with stress responses (8.5%), and protein associated with transcription regulation (10.9%). Approximately one-third (29%) were categorized as "other proteins". This data provides an initial reference map for comparative studies of protein expression in human and laryngeal muscle. Further development of this database will provide a valuable resource for molecular analysis of normal and pathologic conditions affecting human striated muscle.

Databases as Topic↗

Gene-protein database of Escherichia coli K-12: edition 3.

The first two editions of the E. coli Gene-Protein Index were published to provide identifications of protein spots resolved by two-dimensional gel electrophoresis as the products of known genes. This third edition has been expanded to include information about genes and proteins gained directly from two-dimensional gel analysis--including information about protein spots not yet characterized genetically or biochemically--and is therefore more properly called a cellular protein database. An alpha-numeric designation has been uniquely assigned to each of the 616 polypeptide spots in the current database. To this, information is linked about the polypeptide's identification (protein name, gene name, Enzyme Commission--EC number), location on reference gels (x-y coordinates), genetics (Genbank code, DNA sequence reference), biochemistry (molecular weight, isoelectric point), and physiology (steady state level of the protein as a function of media and temperature, membership in various regulons and stimulons).

Bacterial Proteins↗