Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Identification of proteassemblin, a mammalian homologue of the yeast protein, Ump1p, that is required for normal proteasome assembly.

We have identified a mammalian homologue of yeast Ump1p by searching for similar proteins in human and mouse expressed sequence tag (EST) databases. Ump1p is an accessory protein that is required for normal proteasome assembly in yeast (1). A mammalian homologue, which we refer to as "proteassemblin," is a constituent of proteasome assembly intermediates (preproteasomes), but not fully assembled 20S proteasomes, as is Ump1p in yeast. We also provide evidence that proteassemblin is a constituent of pre-immunoproteasomes that contain the precursor of the interferon-gamma-inducible subunit LMP2. By analogy with Ump1p, we hypothesize that proteassemblin is required for normal mammalian proteasome assembly.

Amino Acid Sequence↗

Toxicogenomic approach for assessing toxicant-related disease.

The problems of identifying environmental factors involved in the etiology of human disease and performing safety and risk assessments of drugs and chemicals have long been formidable issues. Three principal components for predicting potential human health risks are: (1) the diverse structure and properties of thousands of chemicals and other stressors in the environment; (2) the time and dose parameters that define the relationship between exposure and disease; and (3) the genetic diversity of organisms used as surrogates to determine adverse chemical effects. The global techniques evolving from successful genomics efforts are providing new exciting tools with which to address these intractable problems of environmental health and toxicology. In order to exploit the scientific opportunities, the National Institute of Environmental Health Sciences has created the National Center for Toxicogenomics (NCT). The primary mission of the NCT is to use gene expression technology, proteomics and metabolite profiling to create a reference knowledge base that will allow scientists to understand mechanisms of toxicity and to be able to predict the potential toxicity of new chemical entities and drugs. A principal scientific objective underpinning the use of microarray analysis of chemical exposures is to demonstrate the utility of signature profiling of the action of drugs or chemicals and to utilize microarray methodologies to determine biomarkers of exposure and potential adverse effects. The initial approach of the NCT is to utilize proof-of-principle experiments in an effort to "phenotypically anchor" the altered patterns of gene expression to conventional parameters of toxicity and to define dose and time relationships in which the expression of such signature genes may precede the development of overt toxicity. The microarray approach is used in conjunction with proteomic techniques to identify specific proteins that may serve as signature biomarkers. The longer-range goal of these efforts is to develop a reference relational database of chemical effects in biological systems (CEBS) that can be used to define common mechanisms of toxicity, chemical and drug actions, to define cellular pathways of response, injury and, ultimately, disease. In order to implement this strategy, the NCT has created a consortium of research organizations and private sector companies to actively collaborative in populating the database with high quality primary data. The evolution of discrete databases to a knowledge base of toxicogenomics will be accomplished through establishing relational interfaces with other sources of information on the structure and activity of chemicals such as that of the National Toxicology Program (NTP) and with databases annotating gene identity, sequence, and function.

Animals↗

Incremental generation of summarized clustering hierarchy for protein family analysis.

MOTIVATION: Protein sequence clustering has been widely exploited to facilitate in-depth analysis of protein functions and families. For some applications of protein sequence clustering, it is highly desirable that a hierarchical structure, also referred to as dendrogram, which shows how proteins are clustered at various levels, is generated. However, as the sizes of contemporary protein databases continue to grow at rapid rates, it is of great interest to develop some summarization mechanisms so that the users can browse the dendrogram and/or search for the desired information more effectively. RESULTS: In this paper, the design of a novel incremental clustering algorithm aimed at generating summarized dendrograms for analysis of protein databases is described. The proposed incremental clustering algorithm employs a statistics-based model to summarize the distributions of the similarity scores among the proteins in the database and to control formation of clusters. Experimental results reveal that, due to the summarization mechanism incorporated, the proposed incremental clustering algorithm offers the users highly concise dendrograms for analysis of protein clusters with biological significance. Another distinction of the proposed algorithm is its incremental nature. As the sizes of the contemporary protein databases continue to grow at fast rates, due to the concern of efficiency, it is desirable that cluster analysis of a protein database can be carried out incrementally, when the protein database is updated. Experimental results with the Swiss-Prot protein database reveal that the time complexity for carrying out incremental clustering with k new proteins added into the database containing n proteins is O(n2betalogn), where beta congruent with 0.865, provided that k << n. AVAILABILITY: The Linux executable is available on the following supplementary page.

Algorithms↗

Novel odorant-binding proteins expressed in the taste tissue of the fly.

A taste tissue cDNA library of the fleshfly Boettcherisca peregrina was screened with a subtracted cDNA probe enriched with taste-receptor-tissue-specific cDNA. Seven genes were identified with sequence similarity to insect odorant-binding protein (OBP) genes. The predicted amino acid sequences of the genes contain the putative signal peptide sequence at the N-terminal and most of them conserve the six cysteines common to known insect OBPs. These genes show a high degree of sequence divergence with approximately 20% amino acid identity. The most striking feature was that all seven of these genes are expressed mainly in the taste tissues, such as the labellum and tarsus, unlike the known insect OBP genes expressed in olfactory tissue. The predicted amino acid sequences had the highest degree of sequence similarity to the Drosophila melanogaster OBPs named pheromone binding protein-related proteins (PBPRPs). These gene products are here referred to as gustatory PBP-related proteins (GPBPRPs) 1-7. Homologous GPBPRP genes were found also in D. melanogaster by database search and are shown to be expressed in Drosophila taste tissues.

Amino Acid Sequence↗

Functionally specified protein signatures distinctive for each of the different blue copper proteins.

BACKGROUND: Proteins having similar functions from different sources can be identified by the occurrence in their sequences, a conserved cluster of amino acids referred to as pattern, motif, signature or fingerprint. The wide usage of protein sequence analysis in par with the growth of databases signifies the importance of using patterns or signatures to retrieve out related sequences. Blue copper proteins are found in the electron transport chain of prokaryotes and eukaryotes. The signatures already existing in the databases like the type 1 copper blue, multiple copper oxidase, cyt b/b6, photosystem 1 psaA&B, psaG&K, and reiske iron sulphur protein are not specified signatures for blue copper proteins as the name itself suggests. Most profile and motif databases strive to classify protein sequences into a broad spectrum of protein families. This work describes the signatures designed based on the copper metal binding motifs in blue copper proteins. The common feature in all blue copper proteins is a trigonal planar arrangement of two nitrogen ligands [each from histidine] and one sulphur containing thiolate ligand [from cysteine], with strong interactions between the copper center and these ligands. RESULTS: Sequences that share such conserved motifs are crucial to the structure or function of the protein and this could provide a signature of family membership. The blue copper proteins chosen for the study were plantacyanin, plastocyanin, cucumber basic protein, stellacyanin, dicyanin, umecyanin, uclacyanin, cusacyanin, rusticyanin, sulfocyanin, halocyanin, azurin, pseudoazurin, amicyanin and nitrite reductase which were identified in both eukaryotes and prokaryotes. ClustalW analysis of the protein sequences of each of the blue copper proteins was the basis for designing protein signatures or peptides. The protein signatures and peptides identified in this study were designed involving the active site region involving the amino acids bound to the copper atom. It was highly specific for each kind of blue copper protein and the false picks were minimized. The set of signatures designed specifically for the BCP's was entirely different from the existing broad spectrum signatures as mentioned in the background section. CONCLUSIONS: These signatures can be very useful for the annotation of uncharacterized proteins and highly specific to retrieve blue copper protein sequences of interest from the non redundant databases containing a large deposition of protein sequences.

Amino Acid Sequence↗

PIRSF: family classification system at the Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics. To facilitate the sensible propagation and standardization of protein annotation and the systematic detection of annotation errors, PIR has extended its superfamily concept and developed the SuperFamily (PIRSF) classification system. Based on the evolutionary relationships of whole proteins, this classification system allows annotation of both specific biological and generic biochemical functions. The system adopts a network structure for protein classification from superfamily to subfamily levels. Protein family members are homologous (sharing common ancestry) and homeomorphic (sharing full-length sequence similarity with common domain architecture). The PIRSF database consists of two data sets, preliminary clusters and curated families. The curated families include family name, protein membership, parent-child relationship, domain architecture, and optional description and bibliography. PIRSF is accessible from the website at http://pir.georgetown.edu/pirsf/ for report retrieval and sequence classification. The report presents family annotation, membership statistics, cross-references to other databases, graphical display of domain architecture, and links to multiple sequence alignments and phylogenetic trees for curated families. PIRSF can be utilized to analyze phylogenetic profiles, to reveal functional convergence and divergence, and to identify interesting relationships between homeomorphic families, domains and structural classes.

Amino Acid Motifs↗

The apoptosis database.

The apoptosis database is a public resource for researchers and students interested in the molecular biology of apoptosis. The resource provides functional annotation, literature references, diagrams/images, and alternative nomenclatures on a set of proteins having 'apoptotic domains'. These are the distinctive domains that are often, if not exclusively, found in proteins involved in apoptosis. The initial choice of proteins to be included is defined by apoptosis experts and bioinformatics tools. Users can browse through the web accessible lists of domains, proteins containing these domains and their associated homologs. The database can also be searched by sequence homology using basic local alignment search tool, text word matches of the annotation, and identifiers for specific records. The resource is available at http://www.apoptosis-db.org and is updated on a regular basis.

Animals↗

Protein profiling of human pancreatic islets by two-dimensional gel electrophoresis and mass spectrometry.

Completion of the human genome sequence has provided scientists with powerful resources with which to explore the molecular events associated with disease states such as diabetes. Understanding the relative levels of expression of gene products, especially of proteins, and their post-translational modifications will be critical. However, though the pancreatic islets play a key role in glucose homeostasis, global protein expression data in human are decidedly lacking. We here report the two-dimensional protein map and database of human pancreatic islets. A high level of reproducibility was obtained among the gels and a total of 744 protein spots were detected. We have successfully identified 130 spots corresponding to 66 different protein entries and generated a reference map of human islets. The functionally characterized proteins include enzymes, chaperones, cellular structural proteins, cellular defense proteins, signaling molecules, and transport proteins. A number of proteins identified in this study (e.g., annexin A2, elongation factor 1-alpha 2, histone H2B.a/g/k, heat shock protein 90 beta, heat shock 27 kDa protein, cyclophilin B, peroxiredoxin 4, cytokeratins 7, 18, and 19) have not been previously described in the database of mouse pancreatic islets. In addition, altered expression of several proteins, like GRP78, GRP94, PDI, calreticulin, annexin, cytokeratins, profilin, heat shock proteins, and ORP150 have been associated with the development of diabetes. The data presented in this study provides a first-draft reference map of the human islet proteome, that will pave the way for further proteome analysis of pancreatic islets in both healthy and diabetic individuals, generating insights into the pathophysiology of this condition.

Adult↗

General and targeted statistical potentials for protein-ligand interactions.

We present a novel atom-atom potential derived from a database of protein-ligand complexes. First, we clarify the similarities and differences between two statistical potentials described in the literature, PMF and Drugscore. We highlight shortcomings caused by an important factor unaccounted for in their reference states, and describe a new potential, which we name the Astex Statistical Potential (ASP). ASP's reference state considers the difference in exposure of protein atom types towards ligand binding sites. We show that this new potential predicts binding affinities with an accuracy similar to that of Goldscore and Chemscore. We investigate the influence of the choice of reference state by constructing two additional statistical potentials that differ from ASP only in this respect. The reference states in these two potentials are defined along the lines of Drugscore and PMF. In docking experiments, the potential using the new reference state proposed for ASP gives better success rates than when these literature reference states were used; a success rate similar to the established scoring functions Goldscore and Chemscore is achieved with ASP. This is the case both for a large, general validation set of protein-ligand structures and for small test sets of actives against four pharmaceutically relevant targets. Virtual screening experiments for these targets show less discrimination between the different reference states in terms of enrichment. In addition, we describe how statistical potentials can be used in the construction of targeted scoring functions. Examples are given for cdk2, using four different targeted scoring functions, biased towards increasingly large target-specific databases. Using these targeted scoring functions, docking success rates as well as enrichments are significantly better than for the general ASP scoring function. Results improve with the number of structures used in the construction of the target scoring functions, thus illustrating that these targeted ASP potentials can be continuously improved as new structural data become available.

Binding Sites↗

Short tandem repeats are associated with diverse mRNAs encoding membrane-targeted proteins.

Within the genomes of multicellular organisms, short tandem repeating sequences (STRs) are ubiquitous, yet usage patterns remain obscure. The repeats (AC)n and (GU)n appear frequently in the untranslated regions (UTRs) of messenger RNAs (mRNAs). To investigate STR usage patterns, we used three approaches: (1) comparisons of individual mRNA database sequences including annotations and linked references, (2) statistical analysis of complete, UTR databases and (3) study of a large gene family, the aquaporins. Among 500 (AC)n- or (GU)n-containing mRNAs, 58 (12%) had known functions. Of these, 50 (86%) encoded proteins whose activities involved membranes or lipids, including integral membrane proteins, peripheral membrane proteins, ion channels, lipid enzymes, receptors and secreted proteins. A control sequence (AU)n also occurred in mRNAs, but only 5% encoded membrane-related functions. Investigation of all reported 3' UTR sequences, demonstrated that the STR (AC)n was 9 times more common in mRNAs encoding membrane functions than in the total UTR database (P < 0.001). Similarly, (GU)n was 8 times more common in membrane-function mRNAs than in the total database (P < 0.001). These observations suggest that (AC)n and (GU)n may be UTR signals for some mRNAs encoding membrane-targeted proteins.

3' Untranslated Regions↗

Cloning and expression of CIS6, chromosome assignment to 3p22 and 2p21 by in situ hybridization.

A family of negative regulators of JAK signaling pathway referred to as suppressor of cytokines signaling (SOCS) or cytokine-inducible SH2 protein (CIS) has been recently identified. In order to find additional members of this family, we have used a consensus amino acid sequence contained in the well-conserved central SH2 domain to search DNA databases. We isolated cDNA coding for the human homologue of SOCS-5, referred to as CIS6. Northern blot analysis revealed CIS6 mRNA expression in various tissues such as heart, muscle, spleen, and thymus and in all myeloma cell lines examined. The gene was assigned to human chromosome bands 2p21 and 3p22 by in situ hybridization. CIS6 is structurally related to other members of the CIS family and therefore could act as a negative regulator of signal transduction.

Amino Acid Sequence↗

PHProteomicDB: a module for two-dimensional gel electrophoresis database creation on personal web sites.

PHProteomicDB is a PHP-written module to help researchers in proteomics to share two-dimensional electrophoresis gel data using personal web sites. No technical or PHP knowledge is necessary except a few basics about web site management. PHProteomicDB has a user-friendly administration interface to enter and update data. It creates web pages on the fly displaying gel characteristics, gel pictures, and numbered gel spots with their related identifications pointing to their reference pages in protein databanks. The module is freely available at http://www.huvec.com/index.php3?rub=Download.

Animals↗

Does conformational free energy distinguish loop conformations in proteins?

Limitations in protein homology modeling often arise from the inability to adequately model loops. In this paper we focus on the selection of loop conformations. We present a complete computational treatment that allows the screening of loop conformations to identify those that best fit a molecular model. The stability of a loop in a protein is evaluated via computations of conformational free energies in solution, i.e., the free energy difference between the reference structure and the modeled one. A thermodynamic cycle is used for calculation of the conformational free energy, in which the total free energy of the reference state (i.e., gas phase) is the CHARMm potential energy. The electrostatic contribution of the solvation free energy is obtained from solving the finite-difference Poisson-Boltzmann equation. The nonpolar contribution is based on a surface area-based expression. We applied this computational scheme to a simple but well-characterized system, the antibody hypervariable loop (complementarity-determining region, CDR). Instead of creating loop conformations, we generated a database of loops extracted from high-resolution crystal structures of proteins, which display geometrical similarities with antibody CDRs. We inserted loops from our database into a framework of an antibody; then we calculated the conformational free energies of each loop. Results show that we successfully identified loops with a "reference-like" CDR geometry, with the lowest conformational free energy in gas phase only. Surprisingly, the solvation energy term plays a confusing role, sometimes discriminating "reference-like" CDR geometry and many times allowing "non-reference-like" conformations to have the lowest conformational free energies (for short loops). Most "reference-like" loop conformations are separated from others by a gap in the gas phase conformational free energy scale. Naturally, loops from antibody molecules are found to be the best models for long CDRs (> or = 6 residues), mainly because of a better packing of backbone atoms into the framework of the antibody model.

Antibodies↗

Testing statistical significance scores of sequence comparison methods with structure similarity.

BACKGROUND: In the past years the Smith-Waterman sequence comparison algorithm has gained popularity due to improved implementations and rapidly increasing computing power. However, the quality and sensitivity of a database search is not only determined by the algorithm but also by the statistical significance testing for an alignment. The e-value is the most commonly used statistical validation method for sequence database searching. The CluSTr database and the Protein World database have been created using an alternative statistical significance test: a Z-score based on Monte-Carlo statistics. Several papers have described the superiority of the Z-score as compared to the e-value, using simulated data. We were interested if this could be validated when applied to existing, evolutionary related protein sequences. RESULTS: All experiments are performed on the ASTRAL SCOP database. The Smith-Waterman sequence comparison algorithm with both e-value and Z-score statistics is evaluated, using ROC, CVE and AP measures. The BLAST and FASTA algorithms are used as reference. We find that two out of three Smith-Waterman implementations with e-value are better at predicting structural similarities between proteins than the Smith-Waterman implementation with Z-score. SSEARCH especially has very high scores. CONCLUSION: The compute intensive Z-score does not have a clear advantage over the e-value. The Smith-Waterman implementations give generally better results than their heuristic counterparts. We recommend using the SSEARCH algorithm combined with e-values for pairwise sequence comparisons.

Base Sequence↗

Knowledge-based chemoinformatic approaches to drug discovery.

The modern drug discovery process is steadily becoming more information driven. Structural, physicochemical and ADME-Tox property profiles of reference (successful) ligands, along with structural information of their target proteins, have been extremely useful for early-stage drug discovery. Recently, databases of known biologically active ligands (knowledge bases) have become more focused toward different protein-target classes. The number of new chemoinformatics tools used to analyze structures and properties of successful molecules has also increased enormously. Scientists in this area are exploring new physicochemical properties and appropriate drug sets to understand druglike properties. In this review, the various uses of the ligand knowledge bases in the drug discovery process have been critically reviewed.

Drug Design↗

[Cloning and sequence analysis of a gene encoding amastin from Leishmania major].

OBJECTIVE: To clone a gene encoding surface protein from Leishmania major. METHODS: Using T. cruzi amastin DNA sequence as a reference, computer search was done on GenBank and dbEST databases by using BLAST path. A Leishmania major DNA library has been constructed and screened by in situ colony hybridization. RESULTS: A 309nt DNA fragment from Leishmania major was found in dbEST. Leishmania major DNA library was screened using specific primers synthesized according to 309 nt DNA sequence, and a full-length coding sequence for Leishmania major amastin was cloned. The coding sequence consisted of 552 nt, and translated into 183 amino acid residues. The homology is 23.5% at amino acid sequence level between Leishmania major and T. cruzi amastins. CONCLUSION: A full length amastin coding gene for Leishmania major has been cloned.

Amino Acid Sequence↗

Confident protein identification using the average peptide score method coupled with search-specific, ab initio thresholds.

Perhaps the greatest difficulty in interpreting large sets of protein identifications derived from mass spectrometric methods is whether or not to trust the results. For such experiments, the level of confidence in each protein identification made needs to be far greater than the often used 95% significance threshold to avoid the identification of many false-positives. To provide higher confidence results, we have developed an innovative scoring strategy coupling the recently published Average Peptide Score (APS) method with pre-filtering of peptide identifications, using a simple peptide quality filter. Iterative generation of these filters in conjunction with reversed database searching is used to determine the correct levels at which the APS and peptide quality thresholds should be set to return virtually zero false-positive reports. This proceeds without the need to reference a known dataset.

Computational Biology↗