Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

CoSMoS: Conserved Sequence Motif Search in the proteome.

BACKGROUND: With the ever-increasing number of gene sequences in the public databases, generating and analyzing multiple sequence alignments becomes increasingly time consuming. Nevertheless it is a task performed on a regular basis by researchers in many labs. RESULTS: We have now created a database called CoSMoS to find the occurrences and at the same time evaluate the significance of sequence motifs and amino acids encoded in the whole genome of the model organism Escherichia coli K12. We provide a precomputed set of multiple sequence alignments for each individual E. coli protein with all of its homologues in the RefSeq database. The alignments themselves, information about the occurrence of sequence motifs together with information on the conservation of each of the more than 1.3 million amino acids encoded in the E. coli genome can be accessed via the web interface of CoSMoS. CONCLUSION: CoSMoS is a valuable tool to identify highly conserved sequence motifs, to find regions suitable for mutational studies in functional analyses and to predict important structural features in E. coli proteins.

Amino Acid Motifs↗

HUPO Brain Proteome Project Pilot Studies: bioinformatics at work.

The data acquisition phase of initial pilot studies (human and mouse brain samples) of the Human Proteome Organisation (HUPO) Brain Proteome Project (BPP) is now complete and the data generated by the participating laboratories has been submitted to the central Data Collection Center. The BPP Bioinformatics Group met on 8th April 2005 at the European Bioinformatics Institute (Hinxton, UK) to discuss strategies for the reanalysis of the pooled data from all the participating laboratories. A summary of the results of the data reprocessing will be presented at the 4th HUPO World Congress that will be held in August/September 2005.

Animals↗

Homology-based functional proteomics by mass spectrometry: application to the Xenopus microtubule-associated proteome.

The application of functional proteomics to important model organisms with unsequenced genomes is restricted because of the limited ability to identify proteins by conventional mass spectrometry (MS) methods. Here we applied MS and sequence-similarity database searching strategies to characterize the Xenopus laevis microtubule-associated proteome. We identified over 40 unique, and many novel, microtubule-bound proteins, as well as two macromolecular protein complexes involved in protein translation. This finding was corroborated by electron microscopy showing the presence of ribosomes on spindles assembled from frog egg extracts. Taken together, these results suggest that protein translation occurs on the spindle during meiosis in the Xenopus oocyte. These findings were made possible due to the application of sequence-similarity methods, which extended mass spectrometric protein identification capabilities by 2-fold compared to conventional methods.

Animals↗

Estimating probabilities of peptide database identifications to LC-FTICR-MS observations.

BACKGROUND: The field of proteomics involves the characterization of the peptides and proteins expressed in a cell under specific conditions. Proteomics has made rapid advances in recent years following the sequencing of the genomes of an increasing number of organisms. A prominent technology for high throughput proteomics analysis is the use of liquid chromatography coupled to Fourier transform ion cyclotron resonance mass spectrometry (LC-FTICR-MS). Meaningful biological conclusions can best be made when the peptide identities returned by this technique are accompanied by measures of accuracy and confidence. METHODS: After a tryptically digested protein mixture is analyzed by LC-FTICR-MS, the observed masses and normalized elution times of the detected features are statistically matched to the theoretical masses and elution times of known peptides listed in a large database. The probability of matching is estimated for each peptide in the reference database using statistical classification methods assuming bivariate Gaussian probability distributions on the uncertainties in the masses and the normalized elution times. RESULTS: A database of 69,220 features from 32 LC-FTICR-MS analyses of a tryptically digested bovine serum albumin (BSA) sample was matched to a database populated with 97% false positive peptides. The percentage of high confidence identifications was found to be consistent with other database search procedures. BSA database peptides were identified with high confidence on average in 14.1 of the 32 analyses. False positives were identified on average in just 2.7 analyses. CONCLUSION: Using a priori probabilities that contrast peptides from expected and unexpected proteins was shown to perform better in identifying target peptides than using equally likely a priori probabilities. This is because a large percentage of the target peptides were similar to unexpected peptides which were included to be false positives. The use of triplicate analyses with a "2 out of 3" reporting rule was shown to have excellent rejection of false positives.

Journal Article↗

ImmunoTar-integrative prioritization of cell surface targets for cancer immunotherapy.

MOTIVATION: Cancer remains a leading cause of mortality globally. Recent improvements in survival have been facilitated by the development of targeted and less toxic immunotherapies, such as chimeric antigen receptor (CAR)-T cells and antibody-drug conjugates (ADCs). These therapies, effective in treating both pediatric and adult patients with solid and hematological malignancies, rely on the identification of cancer-specific surface protein targets. While technologies like RNA sequencing and proteomics exist to survey these targets, identifying optimal targets for immunotherapies remains a challenge in the field. RESULTS: To address this challenge, we developed ImmunoTar, a novel computational tool designed to systematically prioritize candidate immunotherapeutic targets. ImmunoTar integrates user-provided RNA-sequencing or proteomics data with quantitative features from multiple public databases, selected based on predefined criteria, to generate a score representing the gene's suitability as an immunotherapeutic target. We validated ImmunoTar using three distinct cancer datasets, demonstrating its effectiveness in identifying both known and novel targets across various cancer phenotypes. By compiling diverse data into a unified platform, ImmunoTar enables comprehensive evaluation of surface proteins, streamlining target identification and empowering researchers to efficiently allocate resources, thereby accelerating the development of effective cancer immunotherapies. AVAILABILITY AND IMPLEMENTATION: Code and data to run and test ImmunoTar are available at https://github.com/sacanlab/immunotar.

Humans↗

Potential for false positive identifications from large databases through tandem mass spectrometry.

The biomedical research community at large is increasingly employing shotgun proteomics for large-scale identification of proteins from enzymatic digests. Typically, the approach used to identify proteins and peptides from tandem mass spectral data is based on the matching of experimentally generated tandem mass spectra to the theoretical best match from a protein database. Here, we present the potential difficulties of using such an approach without statistical consideration of the false positive rate, especially when large databases, as are encountered in eukaryotes are considered. This is illustrated by searching a dataset generated from a multidimensional separation of a eukaryotic tryptic digest against an in silico generated random protein database, which generated a significant number of positive matches, even when previously suggested score filtering criteria are used.

Algorithms↗

Top-down proteomic analysis of the soluble sub-proteome of the obligate thermophile, Geobacillus thermoleovorans T80: insights into its cellular processes.

We report the first analysis of the soluble sub-proteome of the obligate thermophile, Geobacillus thermoleovorans T80, utilizing a robust multidimensional protein identification protocol. A total of 1,336 proteins were initially identified utilizing automated MS/MS identification software. Intensive manual curation resulted in a final list containing a total of 294 unique proteins. Physiochemical characterization and functional classification of the soluble sub-proteome was carried out. The strategy has allowed us to gain an insight into the cellular processes of this obligate thermophile, identifying a variety of proteins known to play a role in stress response. Included within these were a number of sigma factors such as sigma(A) that initiate transcription of the heat shock operons controlled by the HrcA-CIRCE complex within gram positive bacteria. In addition, it has enabled us to assign a degree of functionality to 29 out of 36 gene products detected in this study that were hitherto described as being only hypothetical conserved proteins.

Bacillaceae↗

Efficiency improvement of peptide identification for an organism without complete genome sequence, using expressed sequence tag database and tandem mass spectral data.

We compared peptide identification by database (DB) search methods with de novo sequencing results for proteomics study in an organism without genome sequence information. When the former was done by searching the Expressed Sequence Tag (EST) DB of the sample organism or the NCBI nonredundant (nr) protein DB of green plants using either the MASCOT or SEQUEST software program, it was confirmed that the former is as accurate as the latter. Peptides identified from EST DB were twice as many as those from the nr protein DB, in spite of the fact that the EST DB has less data (26 222 EST) than the NCBI nr protein DB (224 238). This study demonstrates that EST DB with tandem mass spectra can be used reliably for high-throughput proteomics studies in an organism without genome information.

Algorithms↗

Datamining methodology for LC-MALDI-MS based peptide profiling.

This report will provide a brief overview of the application of data mining in proteomic peptide profiling used for medical biomarker research. Mass spectrometry based profiling of peptides and proteins is frequently used to distinguish disease from non-disease groups and to monitor and predict drug effects. It has the promising potential to enter clinical laboratories as a general purpose diagnostic tool. Data mining methodologies support biomedical science to manage the vast data sets obtained from these instrumentations. Here we will review the typical workflow of peptide profiling, together with typical data mining methodology. Mass spectrometric experiments in peptidomics raise numerous questions in the fields of signal processing, statistics, experimental design and discriminant analysis.

Animals↗

Contribution of proteomics to tumor immunology.

In the postgenome era, the global analysis of gene and protein expression is allowed at RNA and protein level by microarrays and proteomics. The application of these complementary approaches provides new opportunities for tumor immunology investigation. Indeed, applied to the study at the molecular level of the differentiation and maturation of dendritic cells which are the most potent antigen presenting cells, microarray analysis has identified important changes in a large number of genes, whereas the proteomic analysis provided information that could not be obtained at the RNA level, such as the separation of different isoforms and the characterization of post-translational modifications. On the other hand, proteomics allows serological screening of tumor antigens. Indeed, two-dimensional (2-D) polyacrylamide gel electrophoresis allows simultaneously separation several thousand individual proteins from tumor tissue or tumor cell lines. Proteins eliciting humoral response in cancer are identified by 2-D Western blot using cancer patient sera, followed by mass spectrometry analysis and database search. Applied to different types of cancer, the proteome based approach has allowed us to define several tumor antigens. The common occurrence of autoantibodies to certain of these proteins in different cancers may be useful in cancer screening and diagnosis as well as for immunotherapy.

Amino Acid Sequence↗

Analysis of proteomic components of sera from patients with uremia by two dimensional electrophoresis and matrix assisted laser desorption/ ionization time of flight mass spectrometry.

The different sera proteomic components between uremia patients and normal subjects were studied through two-dimensional gel electrophoresis technique. Immobilized pH gradient two-dimensional polyacrylamide gel electrophoresis (2DE), silver staining, ImageMaster 2D 5.0 analysis software, matrix assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF-TOF-MS) and IPI human database searching were used to separate and identify the proteome of the sera from the patients with uremia. The results showed that satisfactory 2DE patterns of the serum proteins were obtained. Twenty-six protein spots showed significant difference in quantity in uremia patients, and 20 protein spots were identified by MALDI-TOF-TOF-MS. It was concluded that good reproducibility could be obtained by applying immobilized pH gradient 2DE to separate and identify the proteome in serum, which provided the foundation for the further study on uremia toxins pertaining to protein.

Adult↗

Pathway mapping tools for analysis of high content data.

The complexity of human biology requires a systems approach that uses computational approaches to integrate different data types. Systems biology encompasses the complete biological system of metabolic and signaling pathways, which can be assessed by measuring global gene expression, protein content, metabolic profiles, and individual genetic, clinical, and phenotypic data. High content screening assays can also be used to generate systems biology knowledge. In this review, we will summarize the pathway databases and describe biological network tools used predominantly with this genomics, proteomics, and metabolomics data but which are equally as applicable for high content screening data analysis. We describe in detail the integrated data-mining tools applicable to building biological networks developed by GeneGo, namely, MetaCore and MetaDrug.

Computational Biology↗

Cluster-C, an algorithm for the large-scale clustering of protein sequences based on the extraction of maximal cliques.

Although the characterization of proteins cannot solely rely upon sequence similarity, it has been widely proved that all-vs-all massive sequence comparisons may be an effective approach and a good basis for the prediction of biochemical functions or for the delineation of common shared properties. The program Cluster-C presented here enables a stand-alone and efficient construction of protein families within whole proteomes. The algorithm, which is based on the detection of cliques, ensures a high level of connectivity within the clusters. As opposed to the single transitive linkage method, Cluster-C allows a large number of sequences to be classified in such a way that the multidomain proteins do not produce a chain-grouping effect resulting in meaningless clusters. Moreover, some proteins can be present in several different but relevant clusters, which is of help in the determination of their functional domains. In the present analysis we used the Z-value, an evaluation of the significance of the similarity score, as the criterion for connecting sequences (the user can freely define the threshold of the similarity criterion). The clusters built with a rather low threshold (Z= 14) include more than 97% of the sequences and are consistent with known protein families and PROSITE patterns.

Algorithms↗

GénoPlante-Info (GPI): a collection of databases and bioinformatics resources for plant genomics.

Génoplante is a partnership program between public French institutes (INRA, CIRAD, IRD and CNRS) and private companies (Biogemma, Bayer CropScience and Bioplante) that aims at developing genome analysis programs for crop species (corn, wheat, rapeseed, sunflower and pea) and model plants (Arabidopsis and rice). The outputs of these programs form a wealth of information (genomic sequence, transcriptome, proteome, allelic variability, mapping and synteny, and mutation data) and tools (databases, interfaces, analysis software), that are being integrated and made public at the public bioinformatics resource centre of Génoplante: GénoPlante-Info (GPI). This continuous flood of data and tools is regularly updated and will grow continuously during the coming two years. Access to the GPI databases and tools is available at http://genoplante-info.infobiogen.fr/.

Alleles↗

Overcoming the dynamic range problem in mass spectrometry-based shotgun proteomics.

Protein profiling using mass spectrometry technology has emerged as a powerful method for analyzing large-scale protein-expression patterns in cells and tissues. However, a number of challenges are present in proteomics research, one of the greatest being the high degree of protein complexity and huge dynamic range of proteins expressed in the complex biological mixtures, which exceeds six orders of magnitude in cells and ten orders of magnitude in body fluids. Since many important signaling proteins have low expression levels, methods to detect the low-abundance proteins in a complex sample are required. This review will focus on the fundamental fractionation and mass spectrometry techniques currently used for large-scale shotgun proteomics research.

Animals↗

Genome sequence and global gene expression of Q54, a new phage species linking the 936 and c2 phage species of Lactococcus lactis.

The lytic lactococcal phage Q54 was previously isolated from a failed sour cream production. Its complete genomic sequence (26,537 bp) is reported here, and the analysis indicated that it represents a new Lactococcus lactis phage species. A striking feature of phage Q54 is the low level of similarity of its proteome (47 open reading frames) with proteins in databases. A global gene expression study confirmed the presence of two early gene modules in Q54. The unusual configuration of these modules, combined with results of comparative analysis with other lactococcal phage genomes, suggests that one of these modules was acquired through recombination events between c2- and 936-like phages. Proteolytic cleavage and cross-linking of the major capsid protein were demonstrated through structural protein analyses. A programmed translational frameshift between the major tail protein (MTP) and the receptor-binding protein (RBP) was also discovered. A "shifty stop" signal followed by putative secondary structures is likely involved in frameshifting. To our knowledge, this is only the second report of translational frameshifting (+1) in double-stranded DNA bacteriophages and the first case of translational coupling between an MTP and an RBP. Thus, phage Q54 represents a fascinating member of a new species with unusual characteristics that brings new insights into lactococcal phage evolution.

Amino Acid Sequence↗

Protein informatics towards function identification.

The study of structural genomics and structural proteomics has determined the tertiary structures of many hypothetical proteins, whose molecular functions could not be understood using conventional methods. In order to infer the geometrical location of the functional site, the biochemical function and the biological function of the hypothetical protein, much effort has been made in protein informatics. The importance of heterogeneous databases and various descriptors of amino acid sequences, tertiary structures and pathways on the proteome scale has been emphasised.

Binding Sites↗