Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “reference protein database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Screening of Bifidobacterium strains isolated from human faeces for antagonistic activities against potentially bacterial pathogens.

As probiotic bacteria, strains belonging to the genus Bifidobacterium colonise the gastro-intestinal tract of humans and animals at the time of birth, and they are found in young as well as in adult individuals in great numbers. Moreover, they can interact with the development of enteric infections by the production of antimicrobial metabolites. In this work 281 strains of bifidobacteria were anaerobically isolated from human faecal samples, supplied by volunteers of different ages (youngs, adults, elders), and preliminarly described by microscopic observation. All strains were screened by the fructose 6-phosphate phosphoketolase (F6PPK) test in order to confirm their classification within the genus Bifidobacterium. Selected strains were used to evaluate their antagonistic activities against Escherichia coli, Salmonella thyphimurium, Staphylococcus lentus, Enterococcus faecalis, Acinetobacter calcoaceticus, Sphingomonas paucimobilis, Listeria monocytogenes, Yersinia enterocolitica, Bacillus cereus, Clostridium sporogenes. Experiments were performed in vitro by different methods based on the observation of growth inhibition in Petri dishes. The strains that showed the highest inhibiting activities were compared by SDS-PAGE for total cell proteins, using type strains of human origin as references. Representative isolates were metabolically characterised by the BIOLOG system; a specific database was created with strains obtained from our collection and a statistical evaluation for metabolic patterns was carried out.

Adolescent↗

The human cornea proteome: bioinformatic analyses indicate import of plasma proteins into the cornea.

Increased biochemical knowledge of normal and diseased corneas is essential for the understanding of corneal homeostasis and pathophysiology. In a recent study, we characterized the proteome of the normal human cornea and identified 141 distinct proteins. This dataset represents the most comprehensive protein study of the cornea to date and provides a useful reference for further studies of normal and diseased human corneas. The list of identified proteins is available at the Cornea Protein Database. In the present paper, we review the utilized procedures for extraction and fractionation of corneal proteins and discuss the potential roles of the identified proteins in relation to homeostasis, diseases, and wound-healing of the cornea. In addition, we compare the list of identified proteins with high quality gene expression libraries (cDNA libraries) and Serial Analysis of Gene Expression (SAGE) data. Of the 141 proteins, 86 (61%) were recognized in cDNA libraries from the corneas of dogs and rabbits, or humans with keratoconus, and 98 (69.5%) were recognized in SAGE data of mouse and human corneas. However, the percentages of identified genes in each of the protein functional groups differed markedly. Thus, exceptionally few of the traditional blood/plasma proteins and immune defense proteins that were identified in the human cornea were recognized in the gene expression libraries of the cornea. This observation strongly indicates that these abundant corneal proteins are not expressed in the cornea but originate from the surrounding pericorneal tissue.

Animals↗

Generalization of a targeted library design protocol: application to 5-HT7 receptor ligands.

Herein a general concept for the design of targeted libraries for proteins with binding sites that are divided into subsites is laid out, including several practical aspects and their solutions. The design is based on a chemogenomic classification of the subsites followed by collection of bioactive molecular fragments and virtual library generation. The general process is outlined and applied to the assembly of a library of 500 molecules targeting the serotonin type 7 (5-HT7) receptor, a class A G-Protein Coupled Receptor (GPCR). Utilizing commercially available building blocks of similar size and composition, a reference library was created. Control sets of known ligands for the 5-HT7 receptor, other GPCRs, and nuclear receptors were collected from literature sources. Principal component analysis of molecular descriptors for the two libraries and the literature sets, displayed a focusing of the targeted library to the region in the chemical space defined by the literature actives, suggesting a denser coverage of the bioactive region than for the more diverse reference library. Additional computational validations, including PCA class predictions, 3D pharmacophore modeling, and docking calculations all indicated an enrichment factor of 5-HT7 ligand-like molecules in the range of 2-4 for the targeted library compared to the reference library.

Binding Sites↗

The Ribonuclease P database.

The Ribonuclease P Sequence database is a compilation of RNase P sequences, sequence alignments, secondary structures, three-dimensional models, and accessory information. In its initial form, the database contains information on RNase P RNA in bacteria and archaea, and RNase P protein in bacteria. The sequences themselves are presented phylogenetically ordered and aligned. The database also contains secondary structures of bacterial and archaeal RNAs, including specially annotated 'reference' secondary structures of Escherichia coli and Bacillus subtilis RNase P RNAs, a minimum phylogenetic consensus structure, and coordinates for models of three-dimensional structure.

Bacillus subtilis↗

Twenty thousand ORFan microbial protein families for the biologist?

The genomes of most newly sequenced organisms contain a significant fraction of ORFs (open reading frames) that match no other sequence in the databases. We refer to these singleton ORFs as sequence ORFans. Because little can be learned about ORFans by homology, the origin and functions of ORFans remain a mystery. However, in this era of full genome sequencing, it seems that ORFans have been underemphasized. In this minireview, we draw attention to the increasing number of ORFans and to the consequences of this growth to biological research in the postgenomic era.

Animals↗

STRING: known and predicted protein-protein associations, integrated and transferred across organisms.

A full description of a protein's function requires knowledge of all partner proteins with which it specifically associates. From a functional perspective, 'association' can mean direct physical binding, but can also mean indirect interaction such as participation in the same metabolic pathway or cellular process. Currently, information about protein association is scattered over a wide variety of resources and model organisms. STRING aims to simplify access to this information by providing a comprehensive, yet quality-controlled collection of protein-protein associations for a large number of organisms. The associations are derived from high-throughput experimental data, from the mining of databases and literature, and from predictions based on genomic context analysis. STRING integrates and ranks these associations by benchmarking them against a common reference set, and presents evidence in a consistent and intuitive web interface. Importantly, the associations are extended beyond the organism in which they were originally described, by automatic transfer to orthologous protein pairs in other organisms, where applicable. STRING currently holds 730,000 proteins in 180 fully sequenced organisms, and is available at http://string.embl.de/.

Databases, Protein↗

The InterPro database, an integrated documentation resource for protein families, domains and functional sites.

Signature databases are vital tools for identifying distant relationships in novel sequences and hence for inferring protein function. InterPro is an integrated documentation resource for protein families, domains and functional sites, which amalgamates the efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. Each InterPro entry includes a functional description, annotation, literature references and links back to the relevant member database(s). Release 2.0 of InterPro (October 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification encoded by a total of 6804 different regular expressions, profiles, fingerprints and Hidden Markov Models. Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1,000,000 hits from 462,500 proteins in SWISS-PROT and TrEMBL). The database is accessible for text- and sequence-based searches at http://www.ebi.ac.uk/interpro/. Questions can be emailed to interhelp@ebi.ac.uk.

Databases, Factual↗

WorfDB: the Caenorhabditis elegans ORFeome Database.

WorfDB (Worm ORFeome DataBase; http://worfdb.dfci.harvard.edu) was created to integrate and disseminate the data from the cloning of complete set of approximately 19 000 predicted protein-encoding Open Reading Frames (ORFs) of Caenorhabditis elegans (also referred to as the 'worm ORFeome'). WorfDB serves as a central data repository enabling the scientific community to search for availability and quality of cloned ORFs. So far, ORF sequence tags (OSTs) obtained for all individual clones have allowed exon structure corrections for approximately 3400 ORFs originally predicted by the C. elegans sequencing consortium. In addition, we now have OSTs for approximately 4300 predicted genes for which no ESTs were available. The database contains this OST information along with data pertinent to the cloning process. WorfDB could serve as a model database for other metazoan ORFeome cloning projects.

Animals↗

Antigenic and molecular characterization of isolates of the Italy 02 infectious bronchitis virus genotype.

As part of an epidemiological surveillance of infectious bronchitis virus (IBV) in Spain, four Spanish field isolates showed high S1 spike sequence similarities with an IBV sequence from the GenBank database named Italy 02. Given that little was known about this new emergent IBV strain we have characterized the four isolates by sequencing the entire S1 part of the spike protein gene and have compared them with many reference IBV serotypes. In addition, cross-virus neutralization assays were conducted with the main IBV serotypes present in Europe. The four Spanish field strains and the Italy 02 S1 sequence from the NCBI database were established as a new genotype that showed maximum amino acid identities with the 4/91 serotype (81.7% to 83.7%), the D274 group that included D207, D274 and D3896 strains (79.8% to 81.7%), and the B1648 serotype (79.3% to 80%). Furthermore, on the basis of these results, it was demonstrated that the Italy 02 genotype had been circulating in Spain since as early as 1997. Based on the average ratio of synonymous:non-synonymous (dS/dN) amino acid substitutions within Italy 02 sequences, no positive selection pressures were related with changes observed in the S1 gene. Moreover, phylogenetic analysis of the S1 gene suggested that the Italy 02 genotype has undergone a recombination event. Virus neutralization assays demonstrated that little antigenic relatedness (less than 35%) exists between Italy 02 and some of the reference IBV serotypes, and indicated that Italy 02 is likely to be a new serotype.

Amino Acid Sequence↗

Comparative analysis of coiled-coil prediction methods.

In this study we compare commonly used coiled-coil prediction methods against a database derived from proteins of known structure. We find that the two older programs COILS and PairCoil/MultiCoil are significantly outperformed by two recent developments: Marcoil, a program built on hidden Markov models, and PCOILS, a new COILS version that uses profiles as inputs; and to a lesser extent by a PairCoil update, PairCoil2. Overall Marcoil provides a slightly better performance over the reference database than PCOILS and is considerably faster, but it is sensitive to highly charged false positives, whereas the weighting option of PCOILS allows the identification of such sequences.

Amino Acid Sequence↗

Improving sensitivity in shotgun proteomics using a peptide-centric database with reduced complexity: protease cleavage and SCX elution rules from data mining of MS/MS spectra.

Correct identification of a peptide sequence from MS/MS data is still a challenging research problem, particularly in proteomic analyses of higher eukaryotes where protein databases are large. The scoring methods of search programs often generate cases where incorrect peptide sequences score higher than correct peptide sequences (referred to as distraction). Because smaller databases yield less distraction and better discrimination between correct and incorrect assignments, we developed a method for editing a peptide-centric database (PC-DB) to remove unlikely sequences and strategies for enabling search programs to utilize this peptide database. Rules for unlikely missed cleavage and nontryptic proteolysis products were identified by data mining 11 849 high-confidence peptide assignments. We also evaluated ion exchange chromatographic behavior as an editing criterion to generate subset databases. When used to search a well-annotated test data set of MS/MS spectra, we found no loss of critical information using PC-DBs, validating the methods for generating and searching against the databases. On the other hand, improved confidence in peptide assignments was achieved for tryptic peptides, measured by changes in DeltaCN and RSP. Decreased distraction was also achieved, consistent with the 3-9-fold decrease in database size. Data mining identified a major class of common nonspecific proteolytic products corresponding to leucine aminopeptidase (LAP) cleavages. Large improvements in identifying LAP products were achieved using the PC-DB approach when compared with conventional searches against protein databases. These results demonstrate that peptide properties can be used to reduce database size, yielding improved accuracy and information capture due to reduced distraction, but with little loss of information compared to conventional protein database searches.

Amino Acid Sequence↗

Proteome analysis reveals elevated serum levels of clusterin in patients with preeclampsia.

Preeclampsia is a pregnancy-specific syndrome and a major cause of maternal mortality. The pathophysiology of preeclampsia is unknown, and no proteome analysis of preeclampsia has been reported. We sought to identify proteins associated with preeclampsia using a proteomic technique and performed two-dimensional electrophoresis (2-DE) on sera from six patients with preeclampsia and six normal pregnant women, followed by comparison of the SYPRO Ruby-stained 2-DE profiles. A group of overexpressed spots was identified in the limited study set. Overexpressed spots were identified as clusterin by matrix-assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) followed by peptide mass fingerprinting, a protein database search, and Western blot analysis. Additionally, sera of 80 preeclamptic women and 80 normal pregnant women were processed by immunoassay methods to confirm changes in clusterin concentrations quantitatively. Immunoassays showed that clusterin levels in the 80 preeclamptic women were significantly higher than those in the 80 controls (mean +/- SD; 1.62 +/- 0.46 times reference level in preeclamptic women vs. 1.30 +/- 0.46 times reference level in controls, P < 0.001). Proteomic analysis of serum proteins is a promising tool for studying preeclampsia pathophysiology and identifying proteins associated with preeclampsia.

Blood Proteins↗

What is a desirable statistical energy function for proteins and how can it be obtained?

Can one obtain a physical energy function for proteins from statistical analysis of protein structures? A direct answer to this question is likely "no." Aless demanding question is whether one can produce a statistical energy function that has the desirable features of a physical-based energy function. Such a desirable energy function would be founded on a physical basis with few or no adjustable parameters, reproduce the known physical characters of amino acid residues, be mostly database independent and transferable, and, more importantly, reasonably accurate in various applications. In this review, we show how such a desirable energy function can be obtained via introducing a simple physical-based reference state called DFIRE (Distance-scaled, Finite, Ideal-gas REference state).

Computer Simulation↗

Identification of proteassemblin, a mammalian homologue of the yeast protein, Ump1p, that is required for normal proteasome assembly.

We have identified a mammalian homologue of yeast Ump1p by searching for similar proteins in human and mouse expressed sequence tag (EST) databases. Ump1p is an accessory protein that is required for normal proteasome assembly in yeast (1). A mammalian homologue, which we refer to as "proteassemblin," is a constituent of proteasome assembly intermediates (preproteasomes), but not fully assembled 20S proteasomes, as is Ump1p in yeast. We also provide evidence that proteassemblin is a constituent of pre-immunoproteasomes that contain the precursor of the interferon-gamma-inducible subunit LMP2. By analogy with Ump1p, we hypothesize that proteassemblin is required for normal mammalian proteasome assembly.

Amino Acid Sequence↗

Toxicogenomic approach for assessing toxicant-related disease.

The problems of identifying environmental factors involved in the etiology of human disease and performing safety and risk assessments of drugs and chemicals have long been formidable issues. Three principal components for predicting potential human health risks are: (1) the diverse structure and properties of thousands of chemicals and other stressors in the environment; (2) the time and dose parameters that define the relationship between exposure and disease; and (3) the genetic diversity of organisms used as surrogates to determine adverse chemical effects. The global techniques evolving from successful genomics efforts are providing new exciting tools with which to address these intractable problems of environmental health and toxicology. In order to exploit the scientific opportunities, the National Institute of Environmental Health Sciences has created the National Center for Toxicogenomics (NCT). The primary mission of the NCT is to use gene expression technology, proteomics and metabolite profiling to create a reference knowledge base that will allow scientists to understand mechanisms of toxicity and to be able to predict the potential toxicity of new chemical entities and drugs. A principal scientific objective underpinning the use of microarray analysis of chemical exposures is to demonstrate the utility of signature profiling of the action of drugs or chemicals and to utilize microarray methodologies to determine biomarkers of exposure and potential adverse effects. The initial approach of the NCT is to utilize proof-of-principle experiments in an effort to "phenotypically anchor" the altered patterns of gene expression to conventional parameters of toxicity and to define dose and time relationships in which the expression of such signature genes may precede the development of overt toxicity. The microarray approach is used in conjunction with proteomic techniques to identify specific proteins that may serve as signature biomarkers. The longer-range goal of these efforts is to develop a reference relational database of chemical effects in biological systems (CEBS) that can be used to define common mechanisms of toxicity, chemical and drug actions, to define cellular pathways of response, injury and, ultimately, disease. In order to implement this strategy, the NCT has created a consortium of research organizations and private sector companies to actively collaborative in populating the database with high quality primary data. The evolution of discrete databases to a knowledge base of toxicogenomics will be accomplished through establishing relational interfaces with other sources of information on the structure and activity of chemicals such as that of the National Toxicology Program (NTP) and with databases annotating gene identity, sequence, and function.

Animals↗

Incremental generation of summarized clustering hierarchy for protein family analysis.

MOTIVATION: Protein sequence clustering has been widely exploited to facilitate in-depth analysis of protein functions and families. For some applications of protein sequence clustering, it is highly desirable that a hierarchical structure, also referred to as dendrogram, which shows how proteins are clustered at various levels, is generated. However, as the sizes of contemporary protein databases continue to grow at rapid rates, it is of great interest to develop some summarization mechanisms so that the users can browse the dendrogram and/or search for the desired information more effectively. RESULTS: In this paper, the design of a novel incremental clustering algorithm aimed at generating summarized dendrograms for analysis of protein databases is described. The proposed incremental clustering algorithm employs a statistics-based model to summarize the distributions of the similarity scores among the proteins in the database and to control formation of clusters. Experimental results reveal that, due to the summarization mechanism incorporated, the proposed incremental clustering algorithm offers the users highly concise dendrograms for analysis of protein clusters with biological significance. Another distinction of the proposed algorithm is its incremental nature. As the sizes of the contemporary protein databases continue to grow at fast rates, due to the concern of efficiency, it is desirable that cluster analysis of a protein database can be carried out incrementally, when the protein database is updated. Experimental results with the Swiss-Prot protein database reveal that the time complexity for carrying out incremental clustering with k new proteins added into the database containing n proteins is O(n2betalogn), where beta congruent with 0.865, provided that k << n. AVAILABILITY: The Linux executable is available on the following supplementary page.

Algorithms↗

Novel odorant-binding proteins expressed in the taste tissue of the fly.

A taste tissue cDNA library of the fleshfly Boettcherisca peregrina was screened with a subtracted cDNA probe enriched with taste-receptor-tissue-specific cDNA. Seven genes were identified with sequence similarity to insect odorant-binding protein (OBP) genes. The predicted amino acid sequences of the genes contain the putative signal peptide sequence at the N-terminal and most of them conserve the six cysteines common to known insect OBPs. These genes show a high degree of sequence divergence with approximately 20% amino acid identity. The most striking feature was that all seven of these genes are expressed mainly in the taste tissues, such as the labellum and tarsus, unlike the known insect OBP genes expressed in olfactory tissue. The predicted amino acid sequences had the highest degree of sequence similarity to the Drosophila melanogaster OBPs named pheromone binding protein-related proteins (PBPRPs). These gene products are here referred to as gustatory PBP-related proteins (GPBPRPs) 1-7. Homologous GPBPRP genes were found also in D. melanogaster by database search and are shown to be expressed in Drosophila taste tissues.

Amino Acid Sequence↗