Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Useful mass spectrometry programs freely available on the internet.

The intention with this chapter is to give an overview of a broad range of freely available programs on the internet which are useful for analyses of mass spectrometry data. The list is by no means a full list of free proteomics tools on the net and I apologize if there are other good tools on the internet that have been missed. The presented programs should cover the needs for the most general tasks in data analysis of proteomics data and have to a limited extend been tested. The links provided in this chapter will over time become invalid. In such cases it is worth while to try a World Wide Web search using the program packages names. This will often reveal the updated links. Some of the links presented in this chapter will also be maintained at http://yass.sdu.dk.

Computational Biology↗

Bioinformatics in proteomics.

Several genome sequencing projects have recently been completed and the majority of human coding regions have been sequenced. In the next step many of the further studies will concentrate on proteins. Proteomics methods are essential for studying protein expression, activity, regulation and modifications. Bioinformatics is an integral part of proteomics research. The recent developments and applications in proteomics are discussed including mass spectrometry data analysis and interpretation, analysis and storage of the gel images to databases, gel comparison, and advanced methods to study e.g. protein co-expression, protein-protein interactions, as well as metabolic and cellular pathways. The significance of informatics in proteomics will gradually increase because of the advent of high-throughput methods relying on powerful data analysis.

Computational Biology↗

Scoring hidden Markov models to discriminate beta-barrel membrane proteins.

A new method is presented for identification of beta-barrel membrane proteins. It is based on a hidden Markov model (HMM) with an architecture obeying these proteins' construction principles. Once the HMM is trained, log-odds score relative to a null model is used to discriminate beta-barrel membrane proteins from other proteins. The method achieves only 10% false positive and false negative rates in a six-fold cross-validation procedure. The results compare favorably with existing methods. This method is proposed to be a valuable tool to quickly scan proteomes of entirely sequenced organisms for beta-barrel membrane proteins.

Algorithms↗

Cyano2Dbase updated: linkage of 234 protein spots to corresponding genes through N-terminal microsequencing.

The cyanobacterium Synechocystis sp. strain PCC6803 is an interesting model organism for preoteome study because it is a photosynthetic procaryote and its genomic sequence has already been determined at our institute. We thus initiated characterization of this organism from a proteomic viewpoint by exploiting two-dimensional (2-D) gel electrophoresis coupled with N-terminal protein sequencing. In a previous study, we linked 130 protein spots on two dimensional gels with the genes that encoded them. As an extension of the previous study, the number of protein spots linked to their corresponding genes was increased to 227 in this study by separately analyzing cyanobacterial proteins in four different fractions (soluble, insoluble, thylakoid membrane, and secretory protein fractions). The resultant updated 2-D protein-gene linkage database, named Cyano2Dbase, will serve as an indispensable tool in future cyanobacterial proteomic studies. From the data compiled in the Cyano2Dbase, we can extract many items of information concerning translation, posttranslational processing including characteristics of cyanobacterial signal sequences and modification of cyanobacterial proteins. The Cyano2Dbase is available to the public through the World Wide Web (http://www.kazusa.or.jp/tech/sazuka/cyano/pr oteome.html).

Amino Acid Sequence↗

The use of common ontologies and controlled vocabularies to enable data exchange and deposition for complex proteomic experiments.

Controlled vocabularies provide a roadmap through complex biological data. Proteomic data is increasing in volume and is currently poorly served by public repositories due to the large number of different formats in which the data is generated and stored. The Human Proteome Organization Proteome Standards Initiative is establishing standards for data transfer and deposition. These standards utilize ontologies and controlled vocabularies to describe experimental procedures and common processes such as sample preparation This paper will discuss the development of such ontologies by the user community and their current utilization in the fields of protein:proein interactions and mass spectrometry.

Computational Biology↗

Application of genomics and proteomics for identification of bacterial gene products as potential vaccine candidates.

The ability of bioinformatics to characterize genomic sequences from pathogenic bacteria for prediction of genes that may encode vaccine candidates, e.g. surface localized proteins, has been evaluated. By applying appropriate tools for genomic mining to the published sequence of Haemophilus influenzae Rd genome, it was possible to identify a putative vaccine candidate, the outer membrane lipoprotein, P6. Proteomics complements genomics by offering abilities to rapidly identify the products of predicted genes, e.g. proteins in outer membrane preparations. The ability to identify the P6 protein uniquely from entries in a sequence database from the expected peptide-mass fingerprint of P6 demonstrates the power of proteomics. The application of proteomics for identification of vaccine candidates for another pathogenic bacterium, Helicobacter pylori using two different approaches is described. The first involves rapid identification of a series of monoclonal antibody reactive proteins from N-terminal sequence tags. The other approach involves identification of proteins in outer membrane preparations by 2-D electrophoresis followed by trypsin digestion and peptide mass map analysis. Our combined studies demonstrate that utilization of genome sequences by application of bioinformatics through genomics and proteomics can expedite the vaccine discovery process by rapidly providing a set of potential candidates for further testing.

Antibodies, Bacterial↗

AgBase: a unified resource for functional analysis in agriculture.

Analysis of functional genomics (transcriptomics and proteomics) datasets is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation. To facilitate systems biology in these species we have established the curated, web-accessible, public resource 'AgBase' (www.agbase.msstate.edu). We have improved the structural annotation of agriculturally important genomes by experimentally confirming the in vivo expression of electronically predicted proteins and by proteogenomic mapping. Proteogenomic data are available from the AgBase proteogenomics link. We contribute Gene Ontology (GO) annotations and we provide a two tier system of GO annotations for users. The 'GO Consortium' gene association file contains the most rigorous GO annotations based solely on experimental data. The 'Community' gene association file contains GO annotations based on expert community knowledge (annotations based directly from author statements and submitted annotations from the community) and annotations for predicted proteins. We have developed two tools for proteomics analysis and these are freely available on request. A suite of tools for analyzing functional genomics datasets using the GO is available online at the AgBase site. We encourage and publicly acknowledge GO annotations from researchers and provide an online mechanism for agricultural researchers to submit requests for GO annotations.

Agriculture↗

De Novo Interpretation of MS/MS Spectra and Protein Identification via Database Searching.

Peptide sequencing via tandem mass spectrometry(MS/MS)is one of the most powerful tools in proteomics to identify proteins. A new algorithm was developed for de novo interpretation of MS/MS spectra using graph theory and dynamic alignment between real spectra and theoretical spectra. The trustworthy peptides from de novo interpretation were used in protein identification via database searching. A high throughput statistical analysis of SwissProt and TrEMBL protein databases showed that it's enough to identify a protein in database with three sequence tags of four amino acid residues, two sequence tags of five amino acid residues or one sequence tag of eight amino acid residues.

Journal Article↗

The proteome: structure, function and evolution.

This paper reports two studies to model the inter-relationships between protein sequence, structure and function. First, an automated pipeline to provide a structural annotation of proteomes in the major genomes is described. The results are stored in a database at Imperial College, London (3D-GENOMICS) that can be accessed at www.sbg.bio.ic.ac.uk. Analysis of the assignments to structural superfamilies provides evolutionary insights. 3D-GENOMICS is being integrated with related proteome annotation data at University College London and the European Bioinformatics Institute in a project known as e-protein (http://www.e-protein.org/). The second topic is motivated by the developments in structural genomics projects in which the structure of a protein is determined prior to knowledge of its function. We have developed a new approach PHUNCTIONER that uses the gene ontology (GO) classification to supervise the extraction of the sequence signal responsible for protein function from a structure-based sequence alignment. Using GO we can obtain profiles for a range of specificities described in the ontology. In the region of low sequence similarity (around 15%), our method is more accurate than assignment from the closest structural homologue. The method is also able to identify the specific residues associated with the function of the protein family.

Computational Biology↗

A complementary bioinformatics approach to identify potential plant cell wall glycosyltransferase-encoding genes.

Plant cell wall (CW) synthesizing enzymes can be divided into the glycan (i.e. cellulose and callose) synthases, which are multimembrane spanning proteins located at the plasma membrane, and the glycosyltransferases (GTs), which are Golgi localized single membrane spanning proteins, believed to participate in the synthesis of hemicellulose, pectin, mannans, and various glycoproteins. At the Carbohydrate-Active enZYmes (CAZy) database where e.g. glucoside hydrolases and GTs are classified into gene families primarily based on amino acid sequence similarities, 415 Arabidopsis GTs have been classified. Although much is known with regard to composition and fine structures of the plant CW, only a handful of CW biosynthetic GT genes-all classified in the CAZy system-have been characterized. In an effort to identify CW GTs that have not yet been classified in the CAZy database, a simple bioinformatics approach was adopted. First, the entire Arabidopsis proteome was run through the Transmembrane Hidden Markov Model 2.0 server and proteins containing one or, more rarely, two transmembrane domains within the N-terminal 150 amino acids were collected. Second, these sequences were submitted to the SUPERFAMILY prediction server, and sequences that were predicted to belong to the superfamilies NDP-sugartransferase, UDP-glycosyltransferase/glucogen-phosphorylase, carbohydrate-binding domain, Gal-binding domain, or Rossman fold were collected, yielding a total of 191 sequences. Fifty-two accessions already classified in CAZy were discarded. The resulting 139 sequences were then analyzed using the Three-Dimensional-Position-Specific Scoring Matrix and mGenTHREADER servers, and 27 sequences with similarity to either the GT-A or the GT-B fold were obtained. Proof of concept of the present approach has to some extent been provided by our recent demonstration that two members of this pool of 27 non-CAZy-classified putative GTs are xylosyltransferases involved in synthesis of pectin rhamnogalacturonan II (J. Egelund, B.L. Petersen, A. Faik, M.S. Motawia, C.E. Olsen, T. Ishii, H. Clausen, P. Ulvskov, and N. Geshi, unpublished data).

Amino Acid Motifs↗

Proteomic analysis of normal human urinary proteins isolated by acetone precipitation or ultracentrifugation.

BACKGROUND: Proteomic techniques have recently become available for large-scale protein analysis. The utility of these techniques in identification of urinary proteins is poorly defined. We constructed a proteome map of normal human urine as a reference protein database by using two differential fractionated techniques to isolate the proteins. METHODS: Proteins were isolated from urine obtained from normal human volunteers by acetone precipitation or ultracentrifugation, separated by two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) and identified by matrix-assisted laser desorption ionization-time-of-flight (MALDI-TOF) mass spectrometry followed by peptide mass fingerprinting. RESULTS: A total of 67 protein forms of 47 unique proteins were identified, including transporters, adhesion molecules, complement, chaperones, receptors, enzymes, serpins, cell signaling proteins and matrix proteins. Acetone precipitated more acidic and hydrophilic proteins, whereas ultracentrifugation fractionated more basic, hydrophobic, and membrane proteins. Bioinformatic analysis predicted glycosylation to be the most common explanation for multiple forms of the same protein. CONCLUSIONS: Combining two differential isolation techniques magnified protein identification from human urine. Proteomic analysis of urinary proteins is a promising tool to study renal physiology and pathophysiology and to determine biomarkers of renal disease.

Acetone↗

PDZ-binding kinase promotes ovarian cancer cell proliferation and invasion via CCNB1 regulation.

BACKGROUND: Ovarian cancer is one of the most lethal gynecological malignancies, characterized by late diagnosis, frequent recurrence, and high mortality. PDZ-binding kinase (PBK), a serine/threonine kinase of the mitogen-activated protein kinase kinase (MAPKK) family, has been implicated in the tumorigenesis of multiple cancers, yet its role in ovarian cancer remains incompletely characterized. This study aimed to investigate the effect of PBK on the proliferation and invasion of ovarian cancer cells. METHODS: The expression of PBK and cyclin B1 (CCNB1) in normal ovarian tissues and ovarian cancer tissues was analyzed using online databases including Gene Expression Profiling Interactive Analysis 2 (GEPIA2), Clinical Proteomic Tumor Analysis Consortium (CPTAC), and Kaplan-Meier Plotter. Clinical tissue specimens were collected to detect the expression of PBK and CCNB1 by immunohistochemistry. Quantitative real-time polymerase chain reaction (PCR) was performed to detect PBK messenger RNA (mRNA) expression levels in clinical specimens and cell lines. Western blot was used to detect PBK protein expression in ovarian cancer cell lines. ES2 and A2780 cells with higher PBK expression were selected to construct PBK knockdown cell lines using lentiviral interference vectors. Cell Counting Kit-8 (CCK-8) assay, colony formation assay, and 5-ethynyl-2'-deoxyuridine (EdU) assay were performed to explore the effect of PBK knockdown on cell proliferation. Transwell assay was used to investigate the effect on cell invasion. The Cancer Genome Atlas (TCGA) and Kyoto Encyclopedia of Genes and Genomes (KEGG) databases were utilized to analyze PBK-related pathways and predict CCNB1 as the gene most closely related to PBK. RESULTS: PBK was significantly overexpressed in ovarian cancer tissues and cell lines compared with normal controls, and high PBK expression was associated with poor overall survival (OS) and progression-free survival (PFS). Knockdown of PBK expression inhibited the proliferation, colony formation, and invasion of ovarian cancer cells. Bioinformatics analysis revealed that CCNB1 was significantly overexpressed in ovarian cancer and high CCNB1 expression was associated with poor OS. CCNB1 was also significantly highly expressed in ovarian cancer tissues as validated by immunohistochemistry and was associated with lymph node metastasis. PBK and CCNB1 expression showed a significant positive correlation in TCGA ovarian cancer datasets. Knockdown of PBK inhibited CCNB1 expression in ovarian cancer cells. CONCLUSIONS: PBK promotes ovarian cancer cell proliferation and invasion. PBK knockdown leads to CCNB1 downregulation. These findings suggest that CCNB1 contributes to PBK-mediated oncogenic effects and identify the PBK-CCNB1 axis as a potential therapeutic target for ovarian cancer treatment.

PDZ-binding kinase (PBK)↗

VEMS 3.0: algorithms and computational tools for tandem mass spectrometry based identification of post-translational modifications in proteins.

Protein and peptide mass analysis and amino acid sequencing by mass spectrometry is widely used for identification and annotation of post-translational modifications (PTMs) in proteins. Modification-specific mass increments, neutral losses or diagnostic fragment ions in peptide mass spectra provide direct evidence for the presence of post-translational modifications, such as phosphorylation, acetylation, methylation or glycosylation. However, the commonly used database search engines are not always practical for exhaustive searches for multiple modifications and concomitant missed proteolytic cleavage sites in large-scale proteomic datasets, since the search space is dramatically expanded. We present a formal definition of the problem of searching databases with tandem mass spectra of peptides that are partially (sub-stoichiometrically) modified. In addition, an improved search algorithm and peptide scoring scheme that includes modification specific ion information from MS/MS spectra was implemented and tested using the Virtual Expert Mass Spectrometrist (VEMS) software. A set of 2825 peptide MS/MS spectra were searched with 16 variable modifications and 6 missed cleavages. The scoring scheme returned a large set of post-translationally modified peptides including precise information on modification type and position. The scoring scheme was able to extract and distinguish the near-isobaric modifications of trimethylation and acetylation of lysine residues based on the presence and absence of diagnostic neutral losses and immonium ions. In addition, the VEMS software contains a range of new features for analysis of mass spectrometry data obtained in large-scale proteomic experiments. Windows binaries are available at http://www.yass.sdu.dk/.

Algorithms↗

Mining genomes and mapping proteomes: identification and characterization of protein subunit vaccines.

Currently, there is an extensive and unprecedented effort to obtain the complete nucleotide sequence of the complex genomes of many micro-organisms. In this post-genomic era, based on the availability of the entire genome sequence of an organism, three new disciplines of molecular biology have emerged: genomics, transcriptional profiling and proteomics. All these technologies have the potential to accelerate the process of identifying protective protein antigens as subunit vaccine targets as well as validating and extending the range of available candidate antigens. The progress of these technologies has led to the origination of the science of bioinformatics for management and critical evaluation of the large amount of information generated. Although genomics, transcriptional profiling and proteomics are each based on different principles, there is considerable synergy between them. Appropriate application of any one, or a combination of two or more of these approaches, coupled with bioinformatics, would allow identification of a short-list of vaccine candidates from the entire list of several hundreds to thousands of proteins encoded by the genome. These candidates would then require usual channelling through the subsequent process involving recombinant expression, purification and testing for immunogenicity and protective efficacy.

Animals↗

High-quality protein knowledge resource: SWISS-PROT and TrEMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc.), a minimal level of redundancy and a high level of integration with other databases. Together with its automatically annotated supplement TrEMBL, it provides a comprehensive and high-quality view of the current state of knowledge about proteins. Ongoing developments include the further improvement of functional and automatic annotation in the databases including evidence attribution with particular emphasis on the human, archaeal and bacterial proteomes and the provision of additional resources such as the International Protein Index (IPI) and XML format of SWISS-PROT and TrEMBL to the user community.

Amino Acid Sequence↗