Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The Ribonuclease P database.

The Ribonuclease P Sequence database is a compilation of RNase P sequences, sequence alignments, secondary structures, three-dimensional models, and accessory information. In its initial form, the database contains information on RNase P RNA in bacteria and archaea, and RNase P protein in bacteria. The sequences themselves are presented phylogenetically ordered and aligned. The database also contains secondary structures of bacterial and archaeal RNAs, including specially annotated 'reference' secondary structures of Escherichia coli and Bacillus subtilis RNase P RNAs, a minimum phylogenetic consensus structure, and coordinates for models of three-dimensional structure.

Bacillus subtilis↗

Twenty thousand ORFan microbial protein families for the biologist?

The genomes of most newly sequenced organisms contain a significant fraction of ORFs (open reading frames) that match no other sequence in the databases. We refer to these singleton ORFs as sequence ORFans. Because little can be learned about ORFans by homology, the origin and functions of ORFans remain a mystery. However, in this era of full genome sequencing, it seems that ORFans have been underemphasized. In this minireview, we draw attention to the increasing number of ORFans and to the consequences of this growth to biological research in the postgenomic era.

Animals↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

WorfDB: the Caenorhabditis elegans ORFeome Database.

WorfDB (Worm ORFeome DataBase; http://worfdb.dfci.harvard.edu) was created to integrate and disseminate the data from the cloning of complete set of approximately 19 000 predicted protein-encoding Open Reading Frames (ORFs) of Caenorhabditis elegans (also referred to as the 'worm ORFeome'). WorfDB serves as a central data repository enabling the scientific community to search for availability and quality of cloned ORFs. So far, ORF sequence tags (OSTs) obtained for all individual clones have allowed exon structure corrections for approximately 3400 ORFs originally predicted by the C. elegans sequencing consortium. In addition, we now have OSTs for approximately 4300 predicted genes for which no ESTs were available. The database contains this OST information along with data pertinent to the cloning process. WorfDB could serve as a model database for other metazoan ORFeome cloning projects.

Animals↗

Design and implementation of a prototype Human Protein Index.

This paper describes information-handling aspects of the TYCHO I analysis system (Clin, Chem. 27: 1807--1820, 1981), which analyzes two-dimensional electrophoresis gels, matches the individual protein spots with those in a reference pattern, and stores various information--including spot measurements, identifications, treatment profiles, set memberships, and comments--in a computerized database. This and additional information such as amino acid composition and cellular localization is then accessible from an interactive program that includes a pictorial user interface and presents much of the data in graphical form. Use of the TYCHO I system is illustrated by examples drawn from analyses of gel patterns from human leukocytes.

Blood Proteins↗

Identification of proteassemblin, a mammalian homologue of the yeast protein, Ump1p, that is required for normal proteasome assembly.

We have identified a mammalian homologue of yeast Ump1p by searching for similar proteins in human and mouse expressed sequence tag (EST) databases. Ump1p is an accessory protein that is required for normal proteasome assembly in yeast (1). A mammalian homologue, which we refer to as "proteassemblin," is a constituent of proteasome assembly intermediates (preproteasomes), but not fully assembled 20S proteasomes, as is Ump1p in yeast. We also provide evidence that proteassemblin is a constituent of pre-immunoproteasomes that contain the precursor of the interferon-gamma-inducible subunit LMP2. By analogy with Ump1p, we hypothesize that proteassemblin is required for normal mammalian proteasome assembly.

Amino Acid Sequence↗

Novel odorant-binding proteins expressed in the taste tissue of the fly.

A taste tissue cDNA library of the fleshfly Boettcherisca peregrina was screened with a subtracted cDNA probe enriched with taste-receptor-tissue-specific cDNA. Seven genes were identified with sequence similarity to insect odorant-binding protein (OBP) genes. The predicted amino acid sequences of the genes contain the putative signal peptide sequence at the N-terminal and most of them conserve the six cysteines common to known insect OBPs. These genes show a high degree of sequence divergence with approximately 20% amino acid identity. The most striking feature was that all seven of these genes are expressed mainly in the taste tissues, such as the labellum and tarsus, unlike the known insect OBP genes expressed in olfactory tissue. The predicted amino acid sequences had the highest degree of sequence similarity to the Drosophila melanogaster OBPs named pheromone binding protein-related proteins (PBPRPs). These gene products are here referred to as gustatory PBP-related proteins (GPBPRPs) 1-7. Homologous GPBPRP genes were found also in D. melanogaster by database search and are shown to be expressed in Drosophila taste tissues.

Amino Acid Sequence↗

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

Cloning and expression of CIS6, chromosome assignment to 3p22 and 2p21 by in situ hybridization.

A family of negative regulators of JAK signaling pathway referred to as suppressor of cytokines signaling (SOCS) or cytokine-inducible SH2 protein (CIS) has been recently identified. In order to find additional members of this family, we have used a consensus amino acid sequence contained in the well-conserved central SH2 domain to search DNA databases. We isolated cDNA coding for the human homologue of SOCS-5, referred to as CIS6. Northern blot analysis revealed CIS6 mRNA expression in various tissues such as heart, muscle, spleen, and thymus and in all myeloma cell lines examined. The gene was assigned to human chromosome bands 2p21 and 3p22 by in situ hybridization. CIS6 is structurally related to other members of the CIS family and therefore could act as a negative regulator of signal transduction.

Amino Acid Sequence↗

Does conformational free energy distinguish loop conformations in proteins?

Limitations in protein homology modeling often arise from the inability to adequately model loops. In this paper we focus on the selection of loop conformations. We present a complete computational treatment that allows the screening of loop conformations to identify those that best fit a molecular model. The stability of a loop in a protein is evaluated via computations of conformational free energies in solution, i.e., the free energy difference between the reference structure and the modeled one. A thermodynamic cycle is used for calculation of the conformational free energy, in which the total free energy of the reference state (i.e., gas phase) is the CHARMm potential energy. The electrostatic contribution of the solvation free energy is obtained from solving the finite-difference Poisson-Boltzmann equation. The nonpolar contribution is based on a surface area-based expression. We applied this computational scheme to a simple but well-characterized system, the antibody hypervariable loop (complementarity-determining region, CDR). Instead of creating loop conformations, we generated a database of loops extracted from high-resolution crystal structures of proteins, which display geometrical similarities with antibody CDRs. We inserted loops from our database into a framework of an antibody; then we calculated the conformational free energies of each loop. Results show that we successfully identified loops with a "reference-like" CDR geometry, with the lowest conformational free energy in gas phase only. Surprisingly, the solvation energy term plays a confusing role, sometimes discriminating "reference-like" CDR geometry and many times allowing "non-reference-like" conformations to have the lowest conformational free energies (for short loops). Most "reference-like" loop conformations are separated from others by a gap in the gas phase conformational free energy scale. Naturally, loops from antibody molecules are found to be the best models for long CDRs (> or = 6 residues), mainly because of a better packing of backbone atoms into the framework of the antibody model.

Antibodies↗

[Cloning and sequence analysis of a gene encoding amastin from Leishmania major].

OBJECTIVE: To clone a gene encoding surface protein from Leishmania major. METHODS: Using T. cruzi amastin DNA sequence as a reference, computer search was done on GenBank and dbEST databases by using BLAST path. A Leishmania major DNA library has been constructed and screened by in situ colony hybridization. RESULTS: A 309nt DNA fragment from Leishmania major was found in dbEST. Leishmania major DNA library was screened using specific primers synthesized according to 309 nt DNA sequence, and a full-length coding sequence for Leishmania major amastin was cloned. The coding sequence consisted of 552 nt, and translated into 183 amino acid residues. The homology is 23.5% at amino acid sequence level between Leishmania major and T. cruzi amastins. CONCLUSION: A full length amastin coding gene for Leishmania major has been cloned.

Amino Acid Sequence↗

Protein supplementation of human milk for promoting growth in preterm infants.

BACKGROUND: For term infants, human milk provides adequate nutrition to facilitate growth, as well as potential beneficial effects on immunity and the maternal-infant emotional state. However, the role of human milk in preterm infants is less well defined as it contains insufficient quantities of some nutrients to meet the estimated needs of the infant. Preterm infants require higher protein intakes than term infants to attain adequate growth rates, and have relatively higher protein turnover rates. Inadequate protein intakes may be partly responsible for low serum albumin and blood urea concentrations in preterm infants. OBJECTIVES: The main objective was to determine if addition of protein to human milk leads to improved growth and neurodevelopmental outcomes without significant adverse effects in preterm infants. SEARCH STRATEGY: The standard search strategy of the Cochrane Neonatal Review Group was used. This includes searches of the Oxford Database of Perinatal Trials, MEDLINE, previous reviews including cross references, abstracts, conferences and symposia proceedings, expert informants, and journal handsearching mainly in the English language. SELECTION CRITERIA: All trials utilizing random or quasi-random allocation to supplementation of human milk with protein or no supplementation in preterm infants who remained in hospital were eligible. DATA COLLECTION AND ANALYSIS: Data were extracting using the standard methods of the Cochrane Neonatal Review Group, with separate evaluation of trial quality and data extraction by each author and synthesis of data using relative risk and weighted mean difference. MAIN RESULTS: Protein supplementation of human milk results in increases in short term weight gain (WMD 3.6 g/kg/day, 95% CI 2.4 to 4.8 g/kg/day), linear growth (WMD 0.28 cm/week, 95% CI 0.18 to 0.38 cm/week) and head growth (WMD 0.15 cm/week, 95% CI 0.06 to 0.23 cm/week). There are insufficient data to evaluate long term neurodevelopmental and growth outcomes. There are too few infants studied to be certain that adverse effects of protein supplementation are not increased. Blood urea levels are increased (WMD 1.0 mmol/l, 95% CI 0.8 to 1.2 mmol/l). REVIEWER'S CONCLUSIONS: Protein supplementation of human milk in relatively well preterm infants results in increases in short term weight gain, linear and head growth. Urea levels are increased, which may reflect adequate rather than excessive dietary protein intake. Further research should be directed towards the evaluation of specific levels of protein intake in preterm infants and the clinical effects of supplementation with protein, including long term growth and neurodevelopmental outcomes. This may best be done in the context of refinement of available multicomponent fortifier preparations.

Dietary Proteins↗

ACNUC--a portable retrieval system for nucleic acid sequence databases: logical and physical designs and usage.

ACNUC is a database structure and retrieval software for use with either the GenBank or EMBL nucleic acid sequence data collections. The nucleotide and textual data furnished by both collections are each restructured into a database that allows sequence retrieval on a multi-criterion basis. The main selection criteria are: species (or higher order taxon), keyword, reference, journal, author, and organelle; all logical combinations of these criteria can be used. Direct access to sequence regions that code for a specific product (protein, tRNA or rRNA) is provided. A versatile extraction procedure copies selected sequences, or fragments of them, from the database to user files suitable to be analysed by user-supplied application programs. A detailed help mechanism is provided to aid the user at any time during the retrieval session. All software has been written in FORTRAN 77 which guarantees a high degree of transportability to minicomputers or mainframes.

Base Sequence↗

YPL.db: the Yeast Protein Localization database.

The Yeast Protein Localization database (YPL.db) contains information about the localization patterns of yeast proteins resulting from microscopic analyses. The data and parameters of the experiments to obtain the localization information, together with images from confocal or video microscopy, are stored in a relational database, building an archive of, and the documentation for, all experiments. The database can be queried based on gene name, protein localization, growth conditions and a number of additional parameters. All experiment parameters are selectable from predefined lists to ensure database integrity and conformity across different investigators. The database provides a structure reference resource to allow for better characterization of unknown or ambiguous localization patterns. Links to MIPS, YPD and SGD databases are provided to allow fast access to further information not contained in the localization database itself. YPL.db is available at http://ypl.tugraz.at.

Computer Graphics↗

Continuous and discontinuous domains: an algorithm for the automatic generation of reliable protein domain definitions.

An algorithm is presented for the fast and accurate definition of protein structural domains from coordinate data without prior knowledge of the number or type of domains. The algorithm explicitly locates domains that comprise one or two continuous segments of protein chain. Domains that include more than two segments are also located. The algorithm was applied to a nonredundant database of 230 protein structures and the results compared to domain definitions obtained from the literature, or by inspection of the coordinates on molecular graphics. For 70% of the proteins, the derived domains agree with the reference definitions, 18% show minor differences and only 12% (28 proteins) show very different definitions. Three screens were applied to identify the derived domains least likely to agree with the subjective definition set. These screens revealed a set of 173 proteins, 97% of which agree well with the subjective definitions. The algorithm represents a practical domain identification tool that can be run routinely on the entire structural database. Adjustment of parameters also allows smaller compact units to be identified in proteins.

Actins↗

Protein analysis by mass spectrometry and sequence database searching: a proteomic approach to identify human lymphoblastoid cell line proteins.

Lymphoblastoid cell lines correspond to in vitro EBV-immortalized lymphocyte B-cells. These cells display a suitable model for experiments dealing with changes in protein expression occurring upon B-cell differentiation, after drug treatment, or after inhibition of some transcription factors. For all these reasons we have undertaken an effort aimed at developing a hematopoietic cell line protein two-dimensional electrophoresis (2-DE) database, containing B-lymphoblastoid 2-DE maps. In this work, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF-MS) peptide mass fingerprinting analysis was adopted for protein identification. The peptide mass fingerprinting identification and the sequence coverage obtained on colloidal Coomassie blue (CBB) stained gel was close to that obtained using zinc-imidazole staining. Everything considered, CBB being more comfortable for subsequent spot manipulations, CBB staining was chosen for identification of a larger number of polypeptides. The results suggest that reticulation of the gel can interfere preventing the uptake of the enzyme during the in-gel digestion step. Consequently, low molecular mass proteins appear more difficult to identify by mass fingerprinting. Finally, the information provided in this study allows the construction of a new annoted reference map of human lymphoblastoid cell proteins. Among the identified proteins 60% were not yet positioned on 2-DE maps in three of the most important well-documented databases. The annoted map will be accessible via Internet on the LBPP server at URL:http:// www-smbh.univ-paris13.fr/lbtp/index.htm.

Acrylic Resins↗

CyanoBase, the genome database for Synechocystis sp. strain PCC6803: status for the year 2000.

CyanoBase provides an online resource for access to data on genomic information about the cyanobacterium Synechocystis sp. strain PCC6803. The database contains annotations for each protein-coding gene deduced from the entire nucleotide sequence of the genome, gene classification lists, and keyword and similarity search engines. Core portions of CyanoBase consist of annotations for each of the 3168 protein genes deduced from the entire nucleotide sequence of this genome. The contents of each gene were improved by updating with the results of similarity searches and by introducing references for analysis in bioinformatics. The database now contains repository facilities that store and provide experimental information, in addition to providing proposals for the function of each gene. This information should help to avoid unnecessary, overlapping experiments and should assist communication between scientists who wish to elucidate the function of putative genes on the cyanobacteria genome. The current URL of CyanoBase is http://www.kazusa.or.jp:8080/cyano/

Cyanobacteria↗