Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Tracing the origin of functional and conserved domains in the human proteome: implications for protein evolution at the modular level.

BACKGROUND: The functional repertoire of the human proteome is an incremental collection of functions accomplished by protein domains evolved along the Homo sapiens lineage. Therefore, knowledge on the origin of these functionalities provides a better understanding of the domain and protein evolution in human. The lack of proper comprehension about such origin has impelled us to study the evolutionary origin of human proteome in a unique way as detailed in this study. RESULTS: This study reports a unique approach for understanding the evolution of human proteome by tracing the origin of its constituting domains hierarchically, along the Homo sapiens lineage. The uniqueness of this method lies in subtractive searching of functional and conserved domains in the human proteome resulting in higher efficiency of detecting their origins. From these analyses the nature of protein evolution and trends in domain evolution can be observed in the context of the entire human proteome data. The method adopted here also helps delineate the degree of divergence of functional families occurred during the course of evolution. CONCLUSION: This approach to trace the evolutionary origin of functional domains in the human proteome facilitates better understanding of their functional versatility as well as provides insights into the functionality of hypothetical proteins present in the human proteome. This work elucidates the origin of functional and conserved domains in human proteins, their distribution along the Homo sapiens lineage, occurrence frequency of different domain combinations and proteome-wide patterns of their distribution, providing insights into the evolutionary solution to the increased complexity of the human proteome.

Animals↗

Advancement of biomarker discovery and validation through the HUPO plasma proteome project.

The Human Proteome Organization (HUPO) Plasma Proteome Project has mounted a Pilot Phase focussed on key problems essential for standardization of specimen collection, specimen handling, choice of fractionation and analysis technologies, and search engines and databases for protein identifications. This international collaboration will lay the groundwork for many large-scale clinical and epidemiological studies of health and disease.

Biomarkers↗

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics↗

Towards alignment independent quantitative assessment of homology detection.

Identification of homologous proteins provides a basis for protein annotation. Sequence alignment tools reliably identify homologs sharing high sequence similarity. However, identification of homologs that share low sequence similarity remains a challenge. Lowering the cutoff value could enable the identification of diverged homologs, but also introduces numerous false hits. Methods are being continuously developed to minimize this problem. Estimation of the fraction of homologs in a set of protein alignments can help in the assessment and development of such methods, and provides the users with intuitive quantitative assessment of protein alignment results. Herein, we present a computational approach that estimates the amount of homologs in a set of protein pairs. The method requires a prevalent and detectable protein feature that is conserved between homologs. By analyzing the feature prevalence in a set of pairwise protein alignments, the method can estimate the number of homolog pairs in the set independently of the alignments' quality. Using the HomoloGene database as a standard of truth, we implemented this approach in a proteome-wide analysis. The results revealed that this approach, which is independent of the alignments themselves, works well for estimating the number of homologous proteins in a wide range of homology values. In summary, the presented method can accompany homology searches and method development, provides validation to search results, and allows tuning of tools and methods.

Animals↗

Computational analysis of shotgun proteomics data.

Proteomics technology is progressing at an incredible rate. The latest generation of tandem mass spectrometers can now acquire tens of thousands of fragmentation spectra in a matter of hours. Furthermore, quantitative proteomics methods have been developed that incorporate a stable isotope-labeled internal standard for every peptide within a complex protein mixture for the measurement of relative protein abundances. These developments have opened the doors for 'shotgun' proteomics, yet have also placed a burden on the computational approaches that manage the data. With each new method that is developed, the quantity of data that can be derived from a single experiment increases. To deal with this increase, new computational approaches are being developed to manage the data and assess false positives. This review discusses current approaches for analyzing proteomics data by mass spectrometry and identifies present computational limitations and bottlenecks.

Algorithms↗

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software↗

Modeling and designing a proteomics application on PROTEUS.

OBJECTIVES: Biomedical applications, such as analysis and management of mass spectrometry proteomics experiments, involve heterogeneous platforms and knowledge, massive data sets, and complex algorithms. Main requirements of such applications are semantic modeling of the experiments and data analysis, as well as high performance computational platforms. In this paper we propose a software platform allowing to model and execute biomedical applications on the Grid. METHODS: Computational Grids offer the required computational power, whereas ontologies and workflow help to face the heterogeneity of biomedical applications. In this paper we propose the use of domain ontologies and workflow techniques for modeling biomedical applications, whereas Grid middleware is responsible for high performance execution. As a case study, the modeling of a proteomics experiment is discussed. RESULTS: The main result is the design and first use of PROTEUS, a Grid-based problem-solving environment for biomedical and bioinformatics applications. CONCLUSION: To manage the complexity of biomedical experiments, ontologies help to model applications and to identify appropriate data and algorithms, workflow techniques allow to combine the elements of such applications in a systematic way. Finally, translation of workflow into execution plans allows the exploitation of the computational power of Grids. Along this direction, in this paper we present PROTEUS discussing a real case study in the proteomics domain.

Algorithms↗

Comparative analysis of two-dimensional protein patterns in malignant and normal human breast tissue.

Malignant and normal human breast tissue were compared by evaluating two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) maps of frozen tissue samples. Image analyzing software was used to scan and process 34 gels. Eight (8/34) of these gels (4 malignant breast tumor samples, 4 normal tissue samples) were selected on the basis of gel and image quality to build a database to identify and measure the expression of a previously unidentified proteome. Growth factor receptor proteins (GFRs), including ERBB2 (HER2) and ERBB3 (HER3), were expressed in the malignant tissue samples. Growth factor receptor proteins were not expressed in the normal tissue. Also, expression of PS2-protein (pS2) was detected in neither malignant nor normal tissue. In benign breast samples a higher intensity of protein expression could be observed for maspin, desmoglein 3 and keratin 8 than in malignant samples. Other proteins expressed in malignant breast tissue include mitogen-activated protein kinase 3 (MK03), heat shock protein 27 kDa (HS27), growth factor receptor-bound protein (GRB2), cathepsin D, G1/S specific cyclin E1 (CGEI), glucose transporter type 5 (GTR5), and a number of as yet unidentified proteins.

Breast↗

Development of an open source laboratory information management system for 2-D gel electrophoresis-based proteomics workflow.

BACKGROUND: In the post-genome era, most research scientists working in the field of proteomics are confronted with difficulties in management of large volumes of data, which they are required to keep in formats suitable for subsequent data mining. Therefore, a well-developed open source laboratory information management system (LIMS) should be available for their proteomics research studies. RESULTS: We developed an open source LIMS appropriately customized for 2-D gel electrophoresis-based proteomics workflow. The main features of its design are compactness, flexibility and connectivity to public databases. It supports the handling of data imported from mass spectrometry software and 2-D gel image analysis software. The LIMS is equipped with the same input interface for 2-D gel information as a clickable map on public 2DPAGE databases. The LIMS allows researchers to follow their own experimental procedures by reviewing the illustrations of 2-D gel maps and well layouts on the digestion plates and MS sample plates. CONCLUSION: Our new open source LIMS is now available as a basic model for proteome informatics, and is accessible for further improvement. We hope that many research scientists working in the field of proteomics will evaluate our LIMS and suggest ways in which it can be improved.

Computational Biology↗

Expanded protein information at SGD: new pages and proteome browser.

The recent explosion in protein data generated from both directed small-scale studies and large-scale proteomics efforts has greatly expanded the quantity of available protein information and has prompted the Saccharomyces Genome Database (SGD; http://www.yeastgenome.org/) to enhance the depth and accessibility of protein annotations. In particular, we have expanded ongoing efforts to improve the integration of experimental information and sequence-based predictions and have redesigned the protein information web pages. A key feature of this redesign is the development of a GBrowse-derived interactive Proteome Browser customized to improve the visualization of sequence-based protein information. This Proteome Browser has enabled SGD to unify the display of hidden Markov model (HMM) domains, protein family HMMs, motifs, transmembrane regions, signal peptides, hydropathy plots and profile hits using several popular prediction algorithms. In addition, a physico-chemical properties page has been introduced to provide easy access to basic protein information. Improvements to the layout of the Protein Information page and integration of the Proteome Browser will facilitate the ongoing expansion of sequence-specific experimental information captured in SGD, including post-translational modifications and other user-defined annotations. Finally, SGD continues to improve upon the availability of genetic and physical interaction data in an ongoing collaboration with BioGRID by providing direct access to more than 82,000 manually-curated interactions.

Computer Graphics↗

Quantify this! Report on a round table discussion on quantitative mass spectrometry in proteomics.

Following the success of the first round table in 2001, the Swiss Proteomic Society has organized two additional specific events during its last two meetings: a proteomic application exercise in 2002 and a round table in 2003. Such events have as their main objective to bring together, around a challenging topic in mass spectrometry, two groups of specialists, those who develop and commercialize mass spectrometry equipment and software, and expert MS users for peptidomics and proteomics studies. The first round table (Geneva, 2001) entitled "Challenges in Mass Spectrometry" was supported by brief oral presentations that stressed critical questions in the field of MS development or applications (Stöcklin and Binz, Proteomics 2002, 2, 825-827). Topics such as (i) direct analysis of complex biological samples, (ii) status and perspectives for MS investigations of noncovalent peptide-ligant interactions; (iii) is it more appropriate to have complementary instruments rather than a universal equipment, (iv) standardization and improvement of the MS signals for protein identification, (v) what would be the new generation of equipment and finally (vi) how to keep hardware and software adapted to MS up-to-date and accessible to all. For the SPS'02 meeting (Lausanne, 2002), a full session alternative event "Proteomic Application Exercise" was proposed. Two different samples were prepared and sent to the different participants: 100 micro g of snake venom (a complex mixture of peptides and proteins) and 10-20 micro g of almost pure recombinant polypeptide derived from the shrimp Penaeus vannamei carrying an heterogeneous post-translational modification (PTM). Among the 15 participants that received the samples blind, eight returned results and most of them were asked to present their results emphasizing the strategy, the manpower and the instrumentation used during the congress (Binz et. al., Proteomics 2003, 3, 1562-1566). It appeared that for the snake venom extract, the quality of the results was not particularly dependant on the strategy used, as all approaches allowed Lication of identification of a certain number of protein families. The genus of the snake was identified in most cases, but the species was ambiguous. Surprisingly, the precise identification of the recombinant almost pure polypeptides appeared to be much more complicated than expected as only one group reported the full sequence. Finally the SPS'03 meeting reported here included a round table on the difficult and challenging task of "Quantification by Mass Spectrometry", a discussion sustained by four selected oral presentations on the use of stable isotopes, electrospray ionization versus matrix-assisted laser desorption/ionization approaches to quantify peptides and proteins in biological fluids, the handling of differential two-dimensional liquid chromatography tandem mass spectrometry data resulting from high throughput experiments, and the quantitative analysis of PTMs. During these three events at the SPS meetings, the impressive quality and quantity of exchanges between the developers and providers of mass spectrometry equipment and software, expert users and the audience, were a key element for the success of these fruitful events and will have definitively paved the way for future round tables and challenging exercises at SPS meetings.

Animals↗

Fractionation of peptides in proteomics with the use of pI-based approach and ZipTip pipette tips.

The aim of the work was to explore the utility of the in-solution isoelectric focusing (sIEF) fractionation method. That method was proved to be the alternative separation method of mixtures of protein tryptic digests in proteomics. Analysis of the identification of peptides was performed with the use of matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF/TOF-MS). For that research, previously designed the miniaturized multi-chamber fractionation sIEF device (75 microl volume for each chamber) based on polyacrylamide gel membranes with immobilines technology was utilized. To evaluate the efficiency and accuracy of sIEF fractionation combined with MS/MS peptides identification, bovine serum albumin (BSA) digest and mixture of five proteins digest were used. First, fractionation of bovine serum albumin digest sample was performed using sIEF method. Studies performed for that simple mixture of peptides proved the ability of the sIEF device to focus peptides mostly in one chamber. Additionally performed the correlation analysis between pI(calc) and pI(exp) values for identified peptides proved the possibility to obtain experimentally useful high correlation. That information was found to have a potential value for construction of additional constraint during false positives evaluation process among identified proteins. Next, studies on the sIEF fractionation were combined with the evaluation of practical use of ZipTip pipette tips to fractionate peptides in the case of simple mixture of proteins. For this, five proteins digest samples were used. The analysis without prior any fractionation enabled to identify very limited number of proteins. The significant improvement was obtained when one used sIEF alone or with combination with ZipTips fractionation prior to MS analysis. The proposed approach based on in-solution isoelectric focusing proved to be an efficient and accurate alternative fractionation method of protein digests and can be considered as the first useful dimension in two-dimensional proteomics separations. Moreover, analytical information from that pI-based fractionation method can be considered as the additional source of database matching constraint. It can also be a valuable tool for analytical and bioinformatic studies of peptides fractionation in proteomics.

Amino Acid Sequence↗

Improving gene annotation using peptide mass spectrometry.

Annotation of protein-coding genes is a key goal of genome sequencing projects. In spite of tremendous recent advances in computational gene finding, comprehensive annotation remains a challenge. Peptide mass spectrometry is a powerful tool for researching the dynamic proteome and suggests an attractive approach to discover and validate protein-coding genes. We present algorithms to construct and efficiently search spectra against a genomic database, with no prior knowledge of encoded proteins. By searching a corpus of 18.5 million tandem mass spectra (MS/MS) from human proteomic samples, we validate 39,000 exons and 11,000 introns at the level of translation. We present translation-level evidence for novel or extended exons in 16 genes, confirm translation of 224 hypothetical proteins, and discover or confirm over 40 alternative splicing events. Polymorphisms are efficiently encoded in our database, allowing us to observe variant alleles for 308 coding SNPs. Finally, we demonstrate the use of mass spectrometry to improve automated gene prediction, adding 800 correct exons to our predictions using a simple rescoring strategy. Our results demonstrate that proteomic profiling should play a role in any genome sequencing project.

Algorithms↗

Proteomic approaches to antigen discovery.

Proteomics has been widely applied to develop two-dimensional polyacrylamide gel electrophoresis maps and databases, evaluate gene expression profiles under different environmental conditions, assess global changes associated with specific mutations, and define drug targets of bacterial pathogens. When coupled to immunological assays, proteomics may also be used to identify B-cell and T-cell antigens within complex protein mixtures. This chapter describes the proteomic approaches developed by our laboratories to accelerate the antigen discovery program for Mycobacterium tuberculosis. As presented or with minor modifications, these techniques may be universally applied to other bacterial pathogens or used to identify bacterial proteins possessing other immunological properties.

Animals↗

Open source system for analyzing, validating, and storing protein identification data.

This paper describes an open-source system for analyzing, storing, and validating proteomics information derived from tandem mass spectrometry. It is based on a combination of data analysis servers, a user interface, and a relational database. The database was designed to store the minimum amount of information necessary to search and retrieve data obtained from the publicly available data analysis servers. Collectively, this system was referred to as the Global Proteome Machine (GPM). The components of the system have been made available as open source development projects. A publicly available system has been established, comprised of a group of data analysis servers and one main database server.

Computational Biology↗