Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Identification and kinetic analysis of a functional homolog of elongation factor 3, YEF3 in Saccharomyces cerevisiae.

Yeast and other fungi contain a soluble elongation factor 3 (EF-3) which is required for growth and protein synthesis. EF-3 contains two ABC cassettes, and binds and hydrolyses ATP. We identified a homolog of the YEF3 gene in the Saccharomyces cerevisiae genome database. This gene, designated YEF3B, is 84% identical in protein sequence to YEF3, which we will now refer to as YEF3A. YEF3B is not expressed during growth under laboratory conditions, and thus cannot rescue growth of YEF3A deletion strains. However, YEF3B can take the place of YEF3A in vivo when expressed from the YEF3A or ADH1 promoters. The products of the YEF3A and YEF3B genes, EF-3A and EF-3B, respectively, were expressed from the ADH1 promoter and purified. Both factors possessed basal and ribosomal-stimulated ATPase activity, and had similar affinity for yeast ribosomes (103 to 113 nM). K(m) values for ATP were similar, but the Kcat values differed significantly. Ribosome-dependent ATPase activity of EF-3A was more efficient than EF-3B, since the Kcat and Kcat/K(m) values for EF-3A were about two-fold higher; however, the difference in Kcat/K(m) values between the two factors was small for basal ATPase activity.

Base Sequence↗

Proteome analysis of chicken embryonic gonads: identification of major proteins from cultured gonadal primordial germ cells.

The domestic chicken (Gallus gallus) is an important model for research in developmental biology because its embryonic development occurs in ovo. To examine the mechanism of embryonic germ cell development, we constructed proteome map of gonadal primordial germ cells (gPGCs) from chicken embryonic gonads. Embryonic gonads were collected from 500 embryos at 6 days of incubation, and the gPGCs were cultured in vitro until colony formed. After 7-10 days in culture, gPGC colonies were separated from gonadal stroma cells (GSCs). Soluble extracts of cultured gPGCs were then fractionated by two-dimensional gel electrophoresis (pH 4-7). A number of protein spots, including those that displayed significant expression levels, were then identified by use of matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) mass spectrometry and LC-MS/MS. Of the 89 gPGC spots examined, 50 yielded mass spectra that matched avian proteins found in on-line databases. Proteome map of this type will serve as an important reference for germ cell biology and transgenic research.

Animals↗

The proteomics standards initiative.

The Proteomics Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparison, exchange and verification. Progress has been made in the development of common standards for data exchange in the fields of both mass spectrometry and protein-protein interaction. A proteomics-specific extension is being created for the emerging American Society for Tests and Measurements mass spectrometry standard, which will be supported by manufacturers of both hardware and software. A data model for proteomics experimentation is under development and discussions on a public repository for published proteomics data are underway. The Protein-Protein Interactions group expects to publish the Level 1 PSI data exchange format for protein-protein interactions soon and discussions as to the content of Level 2 have been initiated.

Biochemistry↗

Quantitative proteome analysis of breast cancer cell lines using 18O-labeling and an accurate mass and time tag strategy.

Proteome comparison of cell lines derived from cancer and normal breast epithelium provide opportunities to identify differentially expressed proteins and pathways associated with specific phenotypes. We employed 16O/18O peptide labeling, FT-ICR MS, and an accurate mass and time (AMT) tag strategy to simultaneously compare the relative abundance of hundreds of proteins in non-cancer and cancer cell lines derived from breast tissue. A cell line reference panel allowed relative protein abundance comparisons among multiple cell lines and across multiple experiments. A peptide database generated from multidimensional LC separations and MS/MS analysis was used for subsequent AMT tag-based peptide identifications. This peptide database represented a total of 2299 proteins, including 514 that were quantified in five cell lines using the AMT tag and 16O/18O strategies. Eighty-six proteins showed at least a threefold protein abundance change between cancer and non-cancer cell lines. Hierarchical clustering of protein abundance ratios revealed that several groups of proteins were differentially expressed between the cancer cell lines.

Breast Neoplasms↗

A human protein atlas for normal and cancer tissues based on antibody proteomics.

Antibody-based proteomics provides a powerful approach for the functional study of the human proteome involving the systematic generation of protein-specific affinity reagents. We used this strategy to construct a comprehensive, antibody-based protein atlas for expression and localization profiles in 48 normal human tissues and 20 different cancers. Here we report a new publicly available database containing, in the first version, approximately 400,000 high resolution images corresponding to more than 700 antibodies toward human proteins. Each image has been annotated by a certified pathologist to provide a knowledge base for functional studies and to allow queries about protein profiles in normal and disease tissues. Our results suggest it should be possible to extend this analysis to the majority of all human proteins thus providing a valuable tool for medical and biological research.

Antibodies↗

SNOW: standard nomenclature wizard to help searching for (bio) chemical standardized names.

UNLABELLED: When developing bioinformatical tools dealing with enzymatic activity, metabolism or enzymatic networks, the problem of the lack of a clear nomenclature for biochemical compounds often arises. This problem leads us to develop a small web-based tool (SNOW, Standard NOmenclature Wizard) which may help to find recommended and trivial names or the correct closest spelling for a query compound name, if it exists. AVAILABILITY: Web-based interface available at http://ibb.uab.es/snow/ SUPPLEMENTARY INFORMATION: http://ibb.uab.es/snow/snow_moreinfo.html

Algorithms↗

Update of the UMD-FBN1 mutation database and creation of an FBN1 polymorphism database.

Fibrillin is the major component of extracellular microfibrils. Mutations in the fibrillin gene on chromosome 15 (FBN1) were first described in the heritable connective disorder, Marfan syndrome (MFS). FBN1 has also been shown to harbor mutations related to a spectrum of conditions phenotypically related to MFS, called "type-1 fibrillinopathies." In 1995, in an effort to standardize the information regarding these mutations and to facilitate their mutational analysis and identification of structure/function and phenotype/genotype relationships, we created a human FBN1 mutation database, UMD-FBN1. This database gives access to a software package that provides specific routines and optimized multicriteria research and sorting tools. For each mutation, information is provided at the gene, protein, and clinical levels. This tool is now a worldwide reference and is frequently used by teams working in the field; more than 220,000 interrogations have been made to it since January 1998. The database has recently been modified to follow the guidelines on mutation databases of the HUGO Mutation Database Initiative (MDI) and the Human Genome Variation Society (HGVS), including their approved mutation nomenclature. The current update shows 559 entries, of which 421 are novel. UMD-FBN1 is accessible at www.umd.be/. We have also recently developed a FBN1 polymorphism database in order to facilitate diagnostics.

Animals↗

The PDBbind database: collection of binding affinities for protein-ligand complexes with known three-dimensional structures.

We have screened the entire Protein Data Bank (Release No. 103, January 2003) and identified 5671 protein-ligand complexes out of 19 621 experimental structures. A systematic examination of the primary references of these entries has led to a collection of binding affinity data (K(d), K(i), and IC(50)) for a total of 1359 complexes. The outcomes of this project have been organized into a Web-accessible database named the PDBbind database.

Databases, Protein↗

MITOP: database for mitochondria-related proteins, genes and diseases.

The MITOP database http://websvr.mips.biochem.mpg. de/proj/medgen/mitop/ consolidates information on both nuclear- and mitochondrial-encoded genes and their proteins. The five species files- Saccharomyces cerevisiae, Mus musculus, Caenorhabditis elegans, Neurospora crassa and Homo sapiens -include annotated data derived from a variety of online resources and the literature. A wide spectrum of search facilities is given in the interelated sections 'Gene catalogues', 'Protein catalogues', 'Homologies', 'Pathways and metabolism', and 'Human disease catalogue' including extensive references and hyperlinks for each entry. Precomputed FASTA searches using all the MITOP yeast protein entries and a list of the best EST hits with graphical cluster alignments related to the yeast reference sequence are presented. The MITOP orthologue tables with cross-listing to all the protein entries for each species in the database facilitate investigations into interspecies homology. A program (MITOPROT) is available to identify mitochondrial targeting sequences and graphical depictions of several important mitochondrial processes are included. The 'Human disease catalogue' lists a total of 101 disorders related to mitochondrial protein abnormalities, sorted by clinical criteria and age of onset.

Animals↗

MIPS: analysis and annotation of proteins from whole genomes in 2005.

The Munich Information Center for Protein Sequences (MIPS at the GSF), Neuherberg, Germany, provides resources related to genome information. Manually curated databases for several reference organisms are maintained. Several of these databases are described elsewhere in this and other recent NAR database issues. In a complementary effort, a comprehensive set of >400 genomes automatically annotated with the PEDANT system are maintained. The main goal of our current work on creating and maintaining genome databases is to extend gene centered information to information on interactions within a generic comprehensive framework. We have concentrated our efforts along three lines (i) the development of suitable comprehensive data structures and database technology, communication and query tools to include a wide range of different types of information enabling the representation of complex information such as functional modules or networks Genome Research Environment System, (ii) the development of databases covering computable information such as the basic evolutionary relations among all genes, namely SIMAP, the sequence similarity matrix and the CABiNet network analysis framework and (iii) the compilation and manual annotation of information related to interactions such as protein-protein interactions or other types of relations (e.g. MPCDB, MPPI, CYGD). All databases described and the detailed descriptions of our projects can be accessed through the MIPS WWW server (http://mips.gsf.de).

Animals↗

Sequencing emm-specific PCR products for routine and accurate typing of group A streptococci.

Rapid sequence analysis of specific PCR products was used to accurately deduce emm types corresponding to the majority of the known group A streptococcal (GAS) M serotypes. The study involved 95 M type reference GAS strains and a survey of 74 recent clinical isolates. A high percentage of agreement between M type serology and the previously published 5' sequences of the emm genes of M type reference strains was noted. The 5' sequences for six established M protein genes--the emm-32, emm-34, emm-38, emm-40, emm-42, and emm-71 genes--were determined to supplement the existing emm sequence database. Rapid sequence analysis differentiated serologically M-nontypeable strains and was used to establish the probable.

Amino Acid Sequence↗

MitoProteome: mitochondrial protein sequence database and annotation system.

MitoProteome is an object-relational mitochondrial protein sequence database and annotation system. The initial release contains 847 human mitochondrial protein sequences, derived from public sequence databases and mass spectrometric analysis of highly purified human heart mitochondria. Each sequence is manually annotated with primary function, subfunction and subcellular location, and extensively annotated in an automated process with data extracted from external databases, including gene information from LocusLink and Ensembl; disease information from OMIM; protein-protein interaction data from MINT and DIP; functional domain information from Pfam; protein fingerprints from PRINTS; protein family and family-specific signatures from InterPro; structure data from PDB; mutation data from PMD; BLAST homology data from NCBI NR; and proteins found to be related based on LocusLink and SWISS-PROT references and sequence and taxonomy data. By highly automating the processes of maintaining the MitoProteome Protein List and extracting relevant data from external databases, we are able to present a dynamic database, updated frequently to reflect changes in public resources. The MitoProteome database is publicly available at http://www. mitoproteome.org/. Users may browse and search MitoProteome, and access a complete compilation of data relevant to each protein of interest, cross-linked to external databases.

Computational Biology↗

The aMAZE LightBench: a web interface to a relational database of cellular processes.

The aMAZE LightBench (http://www.amaze.ulb. ac.be/) is a web interface to the aMAZE relational database, which contains information on gene expression, catalysed chemical reactions, regulatory interactions, protein assembly, as well as metabolic and signal transduction pathways. It allows the user to browse the information in an intuitive way, which also reflects the underlying data model. Moreover links are provided to literature references, and whenever appropriate, to external databases.

Biochemical Phenomena↗

Two-dimensional gel electrophoresis of Escherichia coli homogenates: the Escherichia coli SWISS-2DPAGE database.

Numerous Escherichia coli proteins have already been characterized by two-dimensional gel electrophoresis (2-D PAGE), using carrier ampholytes in the first dimension (VanBogelen, R. A., Sankar, P., Clark, R. L., Bogan, J. A. and Neidhardt, F. C., Electrophoresis 1992, 13, 1014-1054). We present here a reference protein map of E. coli obtained with immobilized pH gradients (IPG) and available in a SWISS-2DPAGE format. Out of the protein spots identified in the E. coli gene protein database by Neidhardt's group, 153 have been identified in the E. coli gene protein database by Neihardt's group, 153 have been identified on the E. coli SWISS-2DPAGE database map by gel comparison and most of them were confirmed either by the analysis of amino acid composition (AAC) and/or N-terminal microsequencing. Additionally, five as yet unsequenced proteins were found. The E. coli SWISS-2DPAGE database is part of the ExPASy molecular biology server accessible through the Word Wide Web network.

Amino Acid Sequence↗

SPiD: a subtilis protein interaction database.

MOTIVATION: Protein-protein interactions are a potential source of valuable clues in determining the functional role of as yet uncharacterized gene products in metabolic pathways. Graph-like structures emerging from the accumulation of interaction data make it difficult to maintain a consistent and global overview by hand. Bioinformatics tools are needed to perform this graph visualization while maintaining a link to the experimental data. RESULTS: "SPiD" is an online database for exploring networks of interacting proteins in Bacillus subtilis characterized by the two-hybrid system. Graphical displays of interaction networks are created dynamically as users interactively navigate through these networks. Third party applications can interface the database through a Common Object Request Broker Architecture (CORBA) tier. AVAILABILITY: SPiD is available through its web site at http://www-mig.versailles.inra.fr/bdsi/SPiD, and through an Interoperable Object Reference (IOR) and its associated Interface Definition Language (IDL). CONTACT: hoebeke@versailles.inra.fr

Bacillus subtilis↗

PSSM-based prediction of DNA binding sites in proteins.

BACKGROUND: Detection of DNA-binding sites in proteins is of enormous interest for technologies targeting gene regulation and manipulation. We have previously shown that a residue and its sequence neighbor information can be used to predict DNA-binding candidates in a protein sequence. This sequence-based prediction method is applicable even if no sequence homology with a previously known DNA-binding protein is observed. Here we implement a neural network based algorithm to utilize evolutionary information of amino acid sequences in terms of their position specific scoring matrices (PSSMs) for a better prediction of DNA-binding sites. RESULTS: An average of sensitivity and specificity using PSSMs is up to 8.7% better than the prediction with sequence information only. Much smaller data sets could be used to generate PSSM with minimal loss of prediction accuracy. CONCLUSION: One problem in using PSSM-derived prediction is obtaining lengthy and time-consuming alignments against large sequence databases. In order to speed up the process of generating PSSMs, we tried to use different reference data sets (sequence space) against which a target protein is scanned for PSI-BLAST iterations. We find that a very small set of proteins can actually be used as such a reference data without losing much of the prediction value. This makes the process of generating PSSMs very rapid and even amenable to be used at a genome level. A web server has been developed to provide these predictions of DNA-binding sites for any new protein from its amino acid sequence. AVAILABILITY: Online predictions based on this method are available at http://www.netasa.org/dbs-pssm/

Algorithms↗

euHCVdb: the European hepatitis C virus database.

The hepatitis C virus (HCV) genome shows remarkable sequence variability, leading to the classification of at least six major genotypes, numerous subtypes and a myriad of quasispecies within a given host. A database allowing researchers to investigate the genetic and structural variability of all available HCV sequences is an essential tool for studies on the molecular virology and pathogenesis of hepatitis C as well as drug design and vaccine development. We describe here the European Hepatitis C Virus Database (euHCVdb, http://euhcvdb.ibcp.fr), a collection of computer-annotated sequences based on reference genomes. The annotations include genome mapping of sequences, use of recommended nomenclature, subtyping as well as three-dimensional (3D) molecular models of proteins. A WWW interface has been developed to facilitate database searches and the export of data for sequence and structure analyses. As part of an international collaborative effort with the US and Japanese databases, the European HCV Database (euHCVdb) is mainly dedicated to HCV protein sequences, 3D structures and functional analyses.

Databases, Protein↗