Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,297 records · Page 72Linked to original sources

The comparative proteomics of ubiquitination in mouse.

Ubiquitination is a common posttranslational modification in eukaryotic cells, influencing many fundamental cellular processes. Defects in ubiquitination and the processes it mediates are involved in many human disease states. The ubiquitination of a substrate involves four classes of enzymes:a ubiquitin-activating enzyme (E1), a ubiquitin-conjugating enzyme (E2), a ubiquitin protein ligase (E3), and a de-ubiquitinating enzyme (DUB). A substantial number of E1s (four), E2s (13), E3s (97), and DUBs (six) that were previously unknown in the mouse are included in the FANTOM2 Representative Transcript and Protein Set (RTPS). Many of the genes encoding these proteins will constitute promising candidates for involvement in disease. In addition, the RTPS provides the basis for the most comprehensive survey of ubiquitination-associated proteins across eukaryotes undertaken to date. Comparisons of these proteins across human and other organisms suggest that eukaryotic evolution has been associated with an increase in the number and diversity of E3s (possessing either zinc-finger RING, F-box, or HECT domains) and DUBs (containing the ubiquitin thiolesterase family 2 domain). These increases in numbers are too large to be accounted for by the presence of fragmentary proteins in the data sets examined. Much of this innovation appears to have been associated with the emergence of multicellular organisms, and subsequently of vertebrates, increasing the opportunity for complex regulation of ubiquitination-mediated cellular and developmental processes.

Animals↗

pSTIING: a 'systems' approach towards integrating signalling pathways, interaction and transcriptional regulatory networks in inflammation and cancer.

pSTIING (http://pstiing.licr.org) is a new publicly accessible web-based application and knowledgebase featuring 65 228 distinct molecular associations (comprising protein-protein, protein-lipid, protein-small molecule interactions and transcriptional regulatory associations), ligand-receptor-cell type information and signal transduction modules. It has a particular major focus on regulatory networks relevant to chronic inflammation, cell migration and cancer. The web application and interface provide graphical representations of networks allowing users to combine and extend transcriptional regulatory and signalling modules, infer molecular interactions across species and explore networks via protein domains/motifs, gene ontology annotations and human diseases. pSTIING also supports the direct cross-correlation of experimental results with interaction information in the knowledgebase via the CLADIST tool associated with pSTIING, which currently analyses and clusters gene expression, proteomic and phenotypic datasets. This allows the contextual projection of co-expression patterns onto prior network information, facilitating the identification of functional modules in physiologically relevant systems.

Amino Acid Motifs↗

Two-dimensional gel protein database of Saccharomyces cerevisiae (update 1999).

By proving the opportunity to visualize several hundred proteins at a time, two-dimensional (2-D) gel electrophoresis is an important tool for proteome research. In order to take advantage of the full potential of this technique for yeast studies, we have undertaken a systematic identification of yeast proteins resolved by this technique. We report here the identification of 92 novel protein spots on the yeast 2-D protein map. These identifications extend the number of protein spots identified on our yeast reference map to 401. These spots correspond to the products of 279 different genes. They have been essentially identified by three methods: gene overexpression, amino acid composition and mass spectrometry. These data can be accessed on the Yeast Protein Map server (htpp://www.ibgc.u-bordeaux2.fr/YPM).

Databases, Factual↗

Risk of Cardiovascular Disease Mortality in Patients With Diagnosed Cancer and Associated Genetic and Proteomic Mechanisms: A UK Biobank-Based Cohort Study.

BACKGROUND: Previous studies have identified a link between cancer and cardiovascular disease; however, the underlying genetic and proteomic mechanisms remain unclear. Therefore, this study aimed to investigate the association between cancer diagnosis and cardiovascular mortality and to explore the potential mechanisms involved. METHODS: A total of 379 944 participants without cardiovascular disease at baseline, including 65 047 individuals with cancer, were recruited from the UK Biobank database. The primary end point was cardiovascular death. Multivariate Cox regression was performed to evaluate the risk of cardiovascular death in populations with and without cancer. Genome-wide association studies, phenome-wide association studies, and proteomic analyses were applied to investigate the underlying genetic and proteomic mechanisms. RESULTS: Multivariate Cox regression analysis showed an increased risk of cardiovascular death in the group with cancer (hazard ratio, 1.50 [95% CI, 1.40-1.61]) after multivariable adjustment. Proteomic analysis confirmed a strong association between cancer and cardiovascular disease, primarily involving pathways related to complement and coagulation cascades, and various inflammatory processes. In contrast, genome-wide association studies and phenome-wide association studies revealed only a limited number of shared genetic variations between cancer and cardiovascular conditions, such as hypertension and cardiac dysrhythmias. CONCLUSIONS: Cardiovascular risk is increased in patients with cancer and may be related to altered expression of inflammation- and coagulation-related proteins. In clinical practice, it is recommended to emphasize the management of endocrine, kidney, and inflammation-related risk factors in the population with cancer.

Humans↗

Primary and secondary metabolism, and post-translational protein modifications, as portrayed by proteomic analysis of Streptomyces coelicolor.

The newly sequenced genome of Streptomyces coelicolor is estimated to encode 7825 theoretical proteins. We have mapped approximately 10% of the theoretical proteome experimentally using two-dimensional gel electrophoresis and matrix-assisted laser desorption ionization time-of-flight (MALDI-TOF) mass spectrometry. Products from 770 different genes were identified, and the types of proteins represented are discussed in terms of their annotated functional classes. An average of 1.2 proteins per gene was observed, indicating extensive post-translational regulation. Examples of modification by N-acetylation, adenylylation and proteolytic processing were characterized using mass spectrometry. Proteins from both primary and certain secondary metabolic pathways are strongly represented on the map, and a number of these enzymes were identified at more than one two-dimensional gel location. Post-translational modification mechanisms may therefore play a significant role in the regulation of these pathways. Unexpectedly, one of the enzymes for synthesis of the actinorhodin polyketide antibiotic appears to be located outside the cytoplasmic compartment, within the cell wall matrix. Of 20 gene clusters encoding enzymes characteristic of secondary metabolism, eight are represented on the proteome map, including three that specify the production of novel metabolites. This information will be valuable in the characterization of the new metabolites.

Acetylation↗

Unlocking the mysteries of virus-host interactions: does functional genomics hold the key?

The interactions between viruses and the cells they infect are complex and multifaceted. While viruses strive to usurp cellular functions to their advantage, the cell strives to thwart these efforts by mounting a variety of defensive responses. These responses may include the induction of interferon, stress response, or apoptotic pathways, all of which are accompanied by changes in gene expression. Some viruses consistently win this tug of war, whereas others succumb to cellular defense mechanisms. The viral and cellular factors that determine the outcome are for the most part still unknown. With the advent of functional genomics, potent new technologies are now available to probe the complexities of virus-host interactions in ever increasing depth and detail. We describe here our efforts to use microarrays, proteomics, and bioinformatics to focus in on the changes in gene expression and protein production that occur in a virus-infected cell and to use these technologies to unlock the mysteries of virus-host interactions.

Animals↗

Automatic target selection for structural genomics on eukaryotes.

A central goal of structural genomics is to experimentally determine representative structures for all protein families. At least 14 structural genomics pilot projects are currently investigating the feasibility of high-throughput structure determination; the National Institutes of Health funded nine of these in the United States. Initiatives differ in the particular subset of "all families" on which they focus. At the NorthEast Structural Genomics consortium (NESG), we target eukaryotic protein domain families. The automatic target selection procedure has three aims: 1) identify all protein domain families from currently five entirely sequenced eukaryotic target organisms based on their sequence homology, 2) discard those families that can be modeled on the basis of structural information already present in the PDB, and 3) target representatives of the remaining families for structure determination. To guarantee that all members of one family share a common foldlike region, we had to begin by dissecting proteins into structural domain-like regions before clustering. Our hierarchical approach, CHOP, utilizing homology to PrISM, Pfam-A, and SWISS-PROT chopped the 103,796 eukaryotic proteins/ORFs into 247,222 fragments. Of these fragments, 122,999 appeared suitable targets that were grouped into >27,000 singletons and >18,000 multifragment clusters. Thus, our results suggested that it might be necessary to determine >40,000 structures to minimally cover the subset of five eukaryotic proteomes.

Algorithms↗

Multi-domain proteins in the three kingdoms of life: orphan domains and other unassigned regions.

Comparative studies of the proteomes from different organisms have provided valuable information about protein domain distribution in the kingdoms of life. Earlier studies have been limited by the fact that only about 50% of the proteomes could be matched to a domain. Here, we have extended these studies by including less well-defined domain definitions, Pfam-B and clustered domains, MAS, in addition to Pfam-A and SCOP domains. It was found that a significant fraction of these domain families are homologous to Pfam-A or SCOP domains. Further, we show that all regions that do not match a Pfam-A or SCOP domain contain a significantly higher fraction of disordered structure. These unstructured regions may be contained within orphan domains or function as linkers between structured domains. Using several different definitions we have re-estimated the number of multi-domain proteins in different organisms and found that several methods all predict that eukaryotes have approximately 65% multi-domain proteins, while the prokaryotes consist of approximately 40% multi-domain proteins. However, these numbers are strongly dependent on the exact choice of cut-off for domains in unassigned regions. In conclusion, all eukaryotes have similar fractions of multi-domain proteins and disorder, whereas a high fraction of repeating domain is distinguished only in multicellular eukaryotes. This implies a role for repeats in cell-cell contacts while the other two features are important for intracellular functions.

Amino Acid Sequence↗

Oligosaccharides, neoglycoproteins and humanized plastics: their biocatalytic synthesis and possible medical applications.

Glycobiology has become one of the fastest growing branches of the biological sciences. Glycomics, which is the study of an organism's entire array of oligosaccharides, is now emerging as the third informatics wave after genomics and proteomics. For example, it is possible to see this progress in the KEGG (Kyoto Encyclopedia of Genes and Genomes) database (http://www.genome.jp/kegg/pathway/map/map01110.html). The interest in this area stems from the realization that carbohydrates, especially oligosaccharides, and their interactions with proteins, play diverse informative roles in all organisms, and that more than half of all proteins are glycosylated. When the biological and pharmaceutical importance of glycoconjugates is considered, it is surprising how little glycobiotechnology has developed. This review reports the latest developments in the biocatalytic synthesis of oligosaccharides and glycoconjugates, with special attention paid to the glycosyltransferase approach. The second part of the review takes the 'conceptual approach' and covers possible medical applications of synthesized glycoconjugates. Various new examples of the conjugation of glyco-informative saccharide sequence to known pharmaceuticals or biomaterials are cited.

Catalysis↗

A sequence-profile-based HMM for predicting and discriminating beta barrel membrane proteins.

MOTIVATION: Membrane proteins are an abundant and functionally relevant subset of proteins that putatively include from about 15 up to 30% of the proteome of organisms fully sequenced. These estimates are mainly computed on the basis of sequence comparison and membrane protein prediction. It is therefore urgent to develop methods capable of selecting membrane proteins especially in the case of outer membrane proteins, barely taken into consideration when proteome wide analysis is performed. This will also help protein annotation when no homologous sequence is found in the database. Outer membrane proteins solved so far at atomic resolution interact with the external membrane of bacteria with a characteristic beta barrel structure comprising different even numbers of beta strands (beta barrel membrane proteins). In this they differ from the membrane proteins of the cytoplasmic membrane endowed with alpha helix bundles (all alpha membrane proteins) and need specialised predictors. RESULTS: We develop a HMM model, which can predict the topology of beta barrel membrane proteins using, as input, evolutionary information. The model is cyclic with 6 types of states: two for the beta strand transmembrane core, one for the beta strand cap on either side of the membrane, one for the inner loop, one for the outer loop and one for the globular domain state in the middle of each loop. The development of a specific input for HMM based on multiple sequence alignment is novel. The accuracy per residue of the model is 83% when a jack knife procedure is adopted. With a model optimisation method using a dynamic programming algorithm seven topological models out of the twelve proteins included in the testing set are also correctly predicted. When used as a discriminator, the model is rather selective. At a fixed probability value, it retains 84% of a non-redundant set comprising 145 sequences of well-annotated outer membrane proteins. Concomitantly, it correctly rejects 90% of a set of globular proteins including about 1200 chains with low sequence identity (<30%) and 90% of a set of all alpha membrane proteins, including 188 chains.

Algorithms↗

Computational identification of strain-, species- and genus-specific proteins.

BACKGROUND: The identification of unique proteins at different taxonomic levels has both scientific and practical value. Strain-, species- and genus-specific proteins can provide insight into the criteria that define an organism and its relationship with close relatives. Such proteins can also serve as taxon-specific diagnostic targets. DESCRIPTION: A pipeline using a combination of computational and manual analyses of BLAST results was developed to identify strain-, species-, and genus-specific proteins and to catalog the closest sequenced relative for each protein in a proteome. Proteins encoded by a given strain are preliminarily considered to be unique if BLAST, using a comprehensive protein database, fails to retrieve (with an e-value better than 0.001) any protein not encoded by the query strain, species or genus (for strain-, species- and genus-specific proteins respectively), or if BLAST, using the best hit as the query (reverse BLAST), does not retrieve the initial query protein. Results are manually inspected for homology if the initial query is retrieved in the reverse BLAST but is not the best hit. Sequences unlikely to retrieve homologs using the default BLOSUM62 matrix (usually short sequences) are re-tested using the PAM30 matrix, thereby increasing the number of retrieved homologs and increasing the stringency of the search for unique proteins. The above protocol was used to examine several food- and water-borne pathogens. We find that the reverse BLAST step filters out about 22% of proteins with homologs that would otherwise be considered unique at the genus and species levels. Analysis of the annotations of unique proteins reveals that many are remnants of prophage proteins, or may be involved in virulence. The data generated from this study can be accessed and further evaluated from the CUPID (Core and Unique Protein Identification) system web site (updated semi-annually) at http://pir.georgetown.edu/cupid. CONCLUSION: CUPID provides a set of proteins specific to a genus, species or a strain, and identifies the most closely related organism.

Algorithms↗

STEM: a software tool for large-scale proteomic data analyses.

We describe the software, STEM (STrategic Extractor for Mascot's results), which efficiently processes large-scale mass spectrometry-based proteomics data. V (View)-mode evaluates the Mascot peptide identification dataset, removes unreliable candidates and redundant assignments, and integrates the results with key information in the experiment. C (Comparison)-mode compares peptide coverage among multiple datasets and displays proteins commonly/specifically found therein, and processes data for quantitative studies that utilize conventional isotope tags or tags having a smaller mass difference. STEM significantly improves throughput of proteomics study.

Algorithms↗

Analysis of the low molecular weight serum peptidome using ultrafiltration and a hybrid ion trap-Fourier transform mass spectrometer.

Advances in proteomics are continuing to expand the ability to analyze the serum proteome. In recent years, it has been realized that in addition to the circulating proteins, human serum also contains a large number of peptides. Many of these peptides are believed to be fragments of larger proteins that have been at least partially degraded by various enzymes such as metalloproteases. Identifying these peptides from a small amount of serum/plasma is difficult due to the complexity of the sample, the low levels of these peptides, and the difficulties in getting a protein identification from a single peptide. In this study, we modified previously published protocols for using centrifugal ultrafiltration, and unlike past studies did not digest the filtrate with trypsin with the intent of identifying endogenous peptides with this method. The filtrate fraction was concentrated and analyzed by a reversed phase-high performance liquid chromatography system connected to a nanospray ionization hybrid ion trap-Fourier transform mass spectrometer (LTQ-FTMS). The mass accuracy of this instrument allows confidence for identifying the protein precursors by a single peptide. The utility of this approach was demonstrated by the identification of over 300 unique peptides with 2 ppm or better mass accuracy per serum sample. With confident identifications, the origin and function of native serum peptides can be more seriously explored. Interestingly, over 34 peptide ladders were observed from over 17 serum proteins. This indicates that a cascade of proteolytic processes affects the serum peptidome. To examine whether this result was an artifact of serum, matched plasma and serum samples were analyzed with similar peptide ladders found in each.

Amino Acid Sequence↗

IMGT, the international ImMunoGeneTics database.

IMGT, the international ImMunoGeneTics database (http://imgt.cines. fr:8104 ), is a high-quality integrated database specialising in Immunoglobulins (Ig), T cell Receptors (TcR) and Major Histocompatibility Complex (MHC) molecules of all vertebrate species, created in 1989 by Marie-Paule Lefranc, Université Montpellier II, CNRS, Montpellier, France (lefranc@ligm.igh.cnrs.fr ). At present, IMGT includes two databases: IMGT/LIGM-DB, a comprehensive database of Ig and TcR from human and other vertebrates, with translation for fully annotated sequences, and IMGT/HLA-DB, a database of the human MHC referred to as HLA (Human Leucocyte Antigens). The IMGT server provides a common access to expertized genomic, proteomic, structural and polymorphic data of Ig and TcR molecules of all vertebrates. By its high quality and its easy data distribution, IMGT has important implications in medical research (repertoire in autoimmune diseases, AIDS, leukemias, lymphomas), therapeutic approaches (antibody engineering), genome diversity and genome evolution studies. IMGT is freely available at http://imgt.cines.fr:8104. The IMGT Index is provided at the IMGT Marie-Paule page (http://imgt.cines.fr:8104/textes/IMGTindex.html).

Amino Acid Sequence↗

Peptide-mass fingerprinting and the ideal covering set for protein characterisation.

The rules that govern the dynamics of protein characterisation by peptide-mass fingerprinting (PMF) were investigated through multiple interrogations of a nonredundant protein database. This was achieved by analysing the efficiency of identifying each entry in the entire database via perfect in silico digestion with a series of 20 pseudo-endoproteinases cutting at the carboxy terminal of each amino acid residue, and the multiple cutters: trypsin, chymotrypsin and Glu-C. The distribution of peptide fragment masses generated by endoproteinase digestion was examined with a view to designing better approaches to protein characterisation by PMF. On average, and for both common and rare cutters, the combination of approximately two fragments was sufficient to identify most database entries. However, the rare cutters left more entries unidentified in the database. Total coverage of the entire database could not be achieved with one enzymatic cutter alone, nor when all 23 cutters were used together. Peptide fragments of > 5000 Da had little effect on the outcome of PMF to correctly characterise database entries, while those with low mass (near to 350 Da in the case of trypsin) were found to be of most utility. The most frequently occurring fragments were also found in this lower mass region. The maximum size of uncut database entries (those not containing a specific amino acid residue) ranged from 52,908 Da to 258,314 Da, while the failure rate for a single cutter in identifying database entries varied from 10,865 (8.4%) to 23,290 (18.1%). PMF is likely to be a mainstay of any high-throughput protein screening strategy for large-scale proteome analysis. A better understanding of the merits and limitations of this technique will allow researchers to optimise their protein characterisation procedures.

Amino Acid Sequence↗

A proteomic analysis of leaf sheaths from rice.

The proteins extracted from the leaf sheaths of rice seedlings were separated by 2-D PAGE, and analyzed by Edman sequencing and mass spectrometry, followed by database searching. Image analysis revealed 352 protein spots on 2-D PAGE after staining with Coomassie Brilliant Blue. The amino acid sequences of 44 of 84 proteins were determined; for 31 of these proteins, a clear function could be assigned, whereas for 12 proteins, no function could be assigned. Forty proteins did not yield amino acid sequence information, because they were N-terminally blocked, or the obtained sequences were too short and/or did not give unambiguous results. Fifty-nine proteins were analyzed by mass spectrometry; all of these proteins were identified by matching to the protein database. The amino acid sequences of 19 of 27 proteins analyzed by mass spectrometry were similar to the results of Edman sequencing. These results suggest that 2-D PAGE combined with Edman sequencing and mass spectrometry analysis can be effectively used to identify plant proteins.

Amino Acid Sequence↗

Detection and analysis of urinary peptides by on-line liquid chromatography and mass spectrometry: application to patients with renal Fanconi syndrome.

Urinary proteomics has become a topical and potentially valuable field of study in relation to normal and abnormal renal function. Filtered bioactive peptides present in high concentration in the nephron of patients with tubular proteinuria may have downstream effects on renal tubular function. In renal Fanconi syndromes, such as Dent's disease, peptides implicated in altered tubular function or injury have recently been measured in urine by immunochemical methods. However, the limited availability of antibodies means that only certain peptides can be detected in this way. We have used nanoflow liquid chromatography and tandem mass spectrometry (nanoLC-MS/MS) as a complementary technique to analyse urinary peptides. Urine was desalted by solid-phase extraction (SPE) and its peptides were then separated from neutral and acidic compounds by strong cation-exchange chromatography (SCX), which was also used to fractionate the peptide mixture. Fractions from the SCX step were separated further by reversed-phase LC and analysed on-line by MS/MS. Extraction by SPE showed a good recovery of small peptides. We detected over 100 molecular species in urine samples from three individuals with Dent's disease. In addition to plasma and known urinary proteins, we identified some novel proteins and potentially bioactive peptides in urine from these patients, which were not present in normal urine. These data show that nanoLC-MS/MS complements existing techniques for the identification of polypeptides in urine. This approach is a potentially powerful tool to discover new markers and/or causative factors in renal disease; in addition, its sensitivity may also make it applicable to the direct ultramicroanalysis of renal tubule fluid.

Biomarkers↗

AbMiner: a bioinformatic resource on available monoclonal antibodies and corresponding gene identifiers for genomic, proteomic, and immunologic studies.

BACKGROUND: Monoclonal antibodies are used extensively throughout the biomedical sciences for detection of antigens, either in vitro or in vivo. We, for example, have used them for quantitation of proteins on "reverse-phase" protein lysate arrays. For those studies, we quality-controlled > 600 available monoclonal antibodies and also needed to develop precise information on the genes that encode their antigens. Translation among the various protein and gene identifier types proved non-trivial because of one-to-many and many-to-one relationships. To organize the antibody, protein, and gene information, we initially developed a relational database in Filemaker for our own use. When it became apparent that the information would be useful to many other researchers faced with the need to choose or characterize antibodies, we developed it further as AbMiner, a fully relational web-based database under MySQL, programmed in Java. DESCRIPTION: AbMiner is a user-friendly, web-based relational database of information on > 600 commercially available antibodies that we validated by Western blot for protein microarray studies. It includes many types of information on the antibody, the immunogen, the vendor, the antigen, and the antigen's gene. Multiple gene and protein identifier types provide links to corresponding entries in a variety of other public databases, including resources for phosphorylation-specific antibodies. AbMiner also includes our quality-control data against a pool of 60 diverse cancer cell types (the NCI-60) and also protein expression levels for the NCI-60 cells measured using our high-density "reverse-phase" protein lysate microarrays for a selection of the listed antibodies. Some other available database resources give information on antibody specificity for one or a couple of cell types. In contrast, the data in AbMiner indicate specificity with respect to the antigens in a pool of 60 diverse cell types from nine different tissues of origin. CONCLUSION: AbMiner is a relational database that provides extensive information from our own laboratory and other sources on more than 600 available antibodies and the genes that encode the antibodies' antigens. The data will be made freely available at http://discover.nci.nih.gov/abminer.

Antibodies, Monoclonal↗