Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

SPiD: a subtilis protein interaction database.

MOTIVATION: Protein-protein interactions are a potential source of valuable clues in determining the functional role of as yet uncharacterized gene products in metabolic pathways. Graph-like structures emerging from the accumulation of interaction data make it difficult to maintain a consistent and global overview by hand. Bioinformatics tools are needed to perform this graph visualization while maintaining a link to the experimental data. RESULTS: "SPiD" is an online database for exploring networks of interacting proteins in Bacillus subtilis characterized by the two-hybrid system. Graphical displays of interaction networks are created dynamically as users interactively navigate through these networks. Third party applications can interface the database through a Common Object Request Broker Architecture (CORBA) tier. AVAILABILITY: SPiD is available through its web site at http://www-mig.versailles.inra.fr/bdsi/SPiD, and through an Interoperable Object Reference (IOR) and its associated Interface Definition Language (IDL). CONTACT: hoebeke@versailles.inra.fr

Bacillus subtilis↗

Web-based two-dimensional database of Saccharomyces cerevisiae proteins using immobilized pH gradients from pH 6 to pH 12 and matrix-assisted laser desorption/ionization-time of flight mass spectrometry.

An image based two-dimensional (2-D) reference map of very alkaline yeast cell proteins was established by using immobilized pH gradients (IPG) up to pH 12 (IPG 6-12, IPG 9-12 and IPG 10-12) for 2-D electrophoresis and by using matrix-assisted laser desorption/ionization-time of flight mass spectrometry peptide mass fingerprinting for spot identification. Up to now 106 proteins with theoretical isoelectric points up to pH 11.15 and molecular mass between 7.5 and 115 kDa were localized and identified. Additionally, due to the improved resolution of steady-state isoelectric focussing with IPGs, even low copy number proteins with codon bias below 0.02 were detected and identified.

Databases as Topic↗

Database of mutations within the adenovirus 5 E1A oncogene.

The Ad5 E1A database is a listing of mutations affecting the early region 1A (E1A) proteins of human adenovirus type 5. The database contains the name of the mutation, the nucleic acid sequence changes, the resulting alterations in amino acid sequence and reference. Additional notes and references are provided on the effect of each mutation on E1A function. The database is contained within the Adenovirus 5 E1A page on the World Wide Web at: http://www.geocities.com/CapeCanaveral/Hangar /2541/

Adenovirus E1A Proteins↗

High-throughput screening of historic collections: observations on file size, biological targets, and file diversity.

At Pfizer Central Research, high-throughput screening has been an important source of new leads for drug discovery for a decade. Our experience with over 150 high-throughput screens can address questions about necessary file size, how well particular biological targets fare (with particular reference to protein-protein interactions), and what file diversity means in practice.

Databases, Factual↗

The human keratinocyte two-dimensional gel protein database (update 1992): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional gel database of human keratinocytes currently lists 2980 cellular proteins (2098 isoelectric focusing, IEF; and 882 nonequilibrium pH gradient electrophoresis, NEPHGE) many of which correspond to posttranslational modifications. About 20% of all recorded proteins have been identified (protein name, organelle components, etc.) and they are listed in alphabetical order together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Also, we have listed 145 microsequenced proteins that are recorded in this database. As an aid in localizing the polypeptides we have included blow-ups of the master images (IEF, NEPHGE) displaying all the protein numbers. In the long run, the master keratinocyte database is expected to link protein and DNA sequencing and mapping information (Human Genome Program) and to provide an integrated picture of the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease.

Cell Differentiation↗

Soy isoflavone analysis: quality control and a new internal standard.

Development of a database of the soy isoflavone content of foods requires accurate and precise evaluation of different food matrixes. To evaluate accuracy, we estimated recoveries of both internal and external standards in 5 different soyfoods weekly. Standards were evaluated daily for system quality assurance. To evaluate sample precision, we analyzed soybeans and soymilk bimonthly for within-day precision and over 4 d for day-to-day precision. CVs should be < or = 8%. We validated our methods for single and multiple recovery concentrations by using our new internal standard, 2,4,4'-trihydroxydeoxybenzoin, and the external standards daidzein, genistein, and genistin. Concentrations of 12 isoflavone isomers, 3 aglycones (daidzein, genistein, and glycitein), and 9 glucosides (daidzin, genistin, glycitin, acetyldaidzin, acetylgenistin, acetylglycitin, malonyldaidzin, malonylgenistin, and malonylglycitin) were measured in a variety of soybeans and soyfoods. The extraction methods used depended on soyfood type. The HPLC conditions for soy isoflavone analysis were improved, leading to good separation with a short analysis time (60 min/sample). A data bank of concentration and distribution of isoflavones in different soybean products was assembled. A wide range of isoflavone concentrations, from < 50 microg/g to > 20,000 microg/g, was found in different soy products. The glucoside forms are almost twice the molecular weight of the aglycones; reported isoflavone concentrations should be normalized to the aglycone mass (or an isoflavonoid equivalent) rather than a simple sum of all isomers.

Chromatography, High Pressure Liquid↗

Location on the human genetic linkage map of 26 genes involved in blood coagulation.

Several human genetic linkage maps have been constructed as part of the Human Genome Project. These maps show the positional order of closely linked, highly informative AC-repeat polymorphisms on each human chromosome, and are extremely useful in genetic linkage analysis of inheritable diseases. For a candidate gene approach the current linkage maps are less useful, since they consist mainly of anonymous markers rather than of specific genes. This situation also applies for inheritable disorders of blood coagulation. Numerous genes are involved in the blood coagulation cascade and its regulation, and can be considered as candidate genes for unexplained haemophilia and thrombophilia. We have selected 29 candidate genes that seem to be the ones most likely to be involved in thrombophilia. For 19 genes genotype data were already present in the CEPH database (version 7.0). We typed 7 additional genes in the CEPH reference families, i.e. the factor V, factor XII, protein C, protein S, prothrombin, thrombomodulin, and heparin cofactor II gene. The genotype data were used to integrate these 26 genes in the current genetic linkage map, and to identify closely linked AC-repeat polymorphisms. This information will benefit the investigation of inheritable disorders of blood coagulation, especially thrombophilia.

Base Sequence↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

Database of mutations that alter the large tumor antigen in simian virus 40.

The SV40 T antigen database is a listing of plasmids and/or viruses that express mutant forms of the virus-encoded large T antigen protein. The parental virus strain, nucleic acid sequence of the mutations, the effect of the mutation on the T antigen amino acid sequence, and key references are included in the listing. The database is available from the authors as a Macintosh FileMaker Pro file, and as a hard copy printout.

Antigens, Polyomavirus Transforming↗

Structure prediction and modelling.

Protein structure prediction from sequence remains a major goal in molecular biology. The methods described in this review concentrate on deriving structural information through the detection of similarities between a test sequence and a database of known structures. Such methods are often referred to as knowledge-based strategies reflecting the use of a structural database in the analyses. The past year has seen considerable advances in both the development of automated procedures and their application to protein sequences of outstanding biological interest.

Algorithms↗

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals↗

The Ribonuclease P database.

The Ribonuclease P Sequence database is a compilation of RNase P sequences, sequence alignments, secondary structures, three-dimensional models, and accessory information. In its initial form, the database contains information on RNase P RNA in bacteria and archaea, and RNase P protein in bacteria. The sequences themselves are presented phylogenetically ordered and aligned. The database also contains secondary structures of bacterial and archaeal RNAs, including specially annotated 'reference' secondary structures of Escherichia coli and Bacillus subtilis RNase P RNAs, a minimum phylogenetic consensus structure, and coordinates for models of three-dimensional structure.

Bacillus subtilis↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

Identification of proteassemblin, a mammalian homologue of the yeast protein, Ump1p, that is required for normal proteasome assembly.

We have identified a mammalian homologue of yeast Ump1p by searching for similar proteins in human and mouse expressed sequence tag (EST) databases. Ump1p is an accessory protein that is required for normal proteasome assembly in yeast (1). A mammalian homologue, which we refer to as "proteassemblin," is a constituent of proteasome assembly intermediates (preproteasomes), but not fully assembled 20S proteasomes, as is Ump1p in yeast. We also provide evidence that proteassemblin is a constituent of pre-immunoproteasomes that contain the precursor of the interferon-gamma-inducible subunit LMP2. By analogy with Ump1p, we hypothesize that proteassemblin is required for normal mammalian proteasome assembly.

Amino Acid Sequence↗

Novel odorant-binding proteins expressed in the taste tissue of the fly.

A taste tissue cDNA library of the fleshfly Boettcherisca peregrina was screened with a subtracted cDNA probe enriched with taste-receptor-tissue-specific cDNA. Seven genes were identified with sequence similarity to insect odorant-binding protein (OBP) genes. The predicted amino acid sequences of the genes contain the putative signal peptide sequence at the N-terminal and most of them conserve the six cysteines common to known insect OBPs. These genes show a high degree of sequence divergence with approximately 20% amino acid identity. The most striking feature was that all seven of these genes are expressed mainly in the taste tissues, such as the labellum and tarsus, unlike the known insect OBP genes expressed in olfactory tissue. The predicted amino acid sequences had the highest degree of sequence similarity to the Drosophila melanogaster OBPs named pheromone binding protein-related proteins (PBPRPs). These gene products are here referred to as gustatory PBP-related proteins (GPBPRPs) 1-7. Homologous GPBPRP genes were found also in D. melanogaster by database search and are shown to be expressed in Drosophila taste tissues.

Amino Acid Sequence↗

Cloning and expression of CIS6, chromosome assignment to 3p22 and 2p21 by in situ hybridization.

A family of negative regulators of JAK signaling pathway referred to as suppressor of cytokines signaling (SOCS) or cytokine-inducible SH2 protein (CIS) has been recently identified. In order to find additional members of this family, we have used a consensus amino acid sequence contained in the well-conserved central SH2 domain to search DNA databases. We isolated cDNA coding for the human homologue of SOCS-5, referred to as CIS6. Northern blot analysis revealed CIS6 mRNA expression in various tissues such as heart, muscle, spleen, and thymus and in all myeloma cell lines examined. The gene was assigned to human chromosome bands 2p21 and 3p22 by in situ hybridization. CIS6 is structurally related to other members of the CIS family and therefore could act as a negative regulator of signal transduction.

Amino Acid Sequence↗