Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Database of two-dimensional polyacrylamide gel electrophoresis of proteins labeled with CyDye DIGE Fluor saturation dye.

CyDye DIGE Fluor saturation dye (saturation dye, GE Healthcare Amersham Biosciences) enables highly sensitive 2-D PAGE. As the dye reacts with all reduced cysteine thiols, 2-D PAGE can be performed with a lower amount of protein, compared with CyDye DIGE Fluor minimal dye (GE Healthcare Amersham Biosciences), the sensitivity of which is equivalent to that of silver staining. We constructed a 2-D map of the saturation dye-labeled proteins of a liver cancer cell line (HepG2) and identified by MS 92 proteins corresponding to 123 protein spots. Functional classification revealed that the identified proteins had chaperone, protein binding, nucleotide binding, metal ion binding, isomerase activity, and motor activity. The functional distribution and the cysteine contents of the proteins were similar to those in the most comprehensive 2-D database of hepatoma cells (Seow et al.., Electrophoresis 2000, 21, 1787-1813), where silver staining was used for protein visualization. Hierarchical clustering on the basis of the quantitative expression profiles of the 123 characterized spots labeled with two charge- and mass-matched saturation dyes (Cy3 and Cy5) discriminated between nine hepatocellular carcinoma cell lines and primary cultured hepatocytes from five individuals, suggesting the utility of saturation dye and our database for proteomic studies of liver cancer.

Cell Line, Tumor↗

PRIDE: a public repository of protein and peptide identifications for the proteomics community.

PRIDE, the 'PRoteomics IDEntifications database' (http://www.ebi.ac.uk/pride) is a database of protein and peptide identifications that have been described in the scientific literature. These identifications will typically be from specific species, tissues and sub-cellular locations, perhaps under specific disease conditions. Any post-translational modifications that have been identified on individual peptides can be described. These identifications may be annotated with supporting mass spectra. At the time of writing, PRIDE includes the full set of identifications as submitted by individual laboratories participating in the HUPO Plasma Proteome Project and a profile of the human platelet proteome submitted by the University of Ghent in Belgium. By late 2005 PRIDE is expected to contain the identifications and spectra generated by the HUPO Brain Proteome Project. Proteomics laboratories are encouraged to submit their identifications and spectra to PRIDE to support their manuscript submissions to proteomics journals. Data can be submitted in PRIDE XML format if identifications are included or mzData format if the submitter is depositing mass spectra without identifications. PRIDE is a web application, so submission, searching and data retrieval can all be performed using an internet browser. PRIDE can be searched by experiment accession number, protein accession number, literature reference and sample parameters including species, tissue, sub-cellular location and disease state. Data can be retrieved as machine-readable PRIDE or mzData XML (the latter for mass spectra without identifications), or as human-readable HTML.

Databases, Protein↗

Molecular mechanisms in cancer: what should clinicians know?

Normal cells are influenced by a variety of environmental and host influences that can produce pro-carcinogenic mutations. Either a single or a series of mutations might result in cellular transformation. Like normal cells, most cancer cells use multiple redundant intracellular signaling pathways to ensure the maintenance and viability of functions critical to their survival. Thus, cellular pathways that are integral to cell function, survival, proliferation, and receptor expression are potential targets for therapeutic intervention. One example of this is the epidermal growth factor receptor signaling pathway. Other potential targets are molecules that mediate processes through which tumors produce angiogenic and invasion factors that stimulate host blood vessel growth into tumors and allow tumor growth and metastasis, such as the vascular endothelial growth factor. Targeting of downstream events that result in cellular apoptosis is another potential strategy. Continued investigations may result in the development of proteomic profiling databases through which a patient might be matched with molecular signatures in a library and upon which individualized cancer therapies might be selected. In this way, clinicians might recommend combinations of molecularly targeted agents and other therapies on the basis of an individual patient's proteomic profile.

Humans↗

Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.

We demonstrate an integrated approach to build, test, and refine a model of a cellular pathway, in which perturbations to critical pathway components are analyzed using DNA microarrays, quantitative proteomics, and databases of known physical interactions. Using this approach, we identify 997 messenger RNAs responding to 20 systematic perturbations of the yeast galactose-utilization pathway, provide evidence that approximately 15 of 289 detected proteins are regulated posttranscriptionally, and identify explicit physical interactions governing the cellular response to each perturbation. We refine the model through further iterations of perturbation and global measurements, suggesting hypotheses about the regulation of galactose utilization and physical interactions between this and a variety of other metabolic pathways.

Computational Biology↗

Using annotated peptide mass spectrum libraries for protein identification.

A system for creating a library of tandem mass spectra annotated with corresponding peptide sequences was described. This system was based on the annotated spectra currently available in the Global Proteome Machine Database (GPMDB). The library spectra were created by averaging together spectra that were annotated with the same peptide sequence, sequence modifications, and parent ion charge. The library was constructed so that experimental peptide tandem mass spectra could be compared with those in the library, resulting in a peptide sequence identification based on scoring the similarity of the experimental spectrum with the contents of the library. A software implementation that performs this type of library search was constructed and successfully used to obtain sequence identifications. The annotated tandem mass spectrum libraries for the Homo sapiens, Mus musculus, and Saccharomyces cerevisiae proteomes and search software were made available for download and use by other groups.

Amino Acid Sequence↗

Proteomic analysis on multi-drug resistant cells HL-60/DOX of acute myeloblastic leukemia.

Multi-drug resistance (MDR) is an important factor that causes treatment failure in acute leukemia. However, the full development mechanisms of MDR still await [corrected] investigation. The purpose of this study is to investigate differentially expressed proteins in the multi-drug resistant acute myeloblastic leukemia (AML) cell line HL-60/DOX and the drug sensitive cell line HL-60, and to identify new potential multi-drug resistant related molecules with the proteomic approach. Two-dimensional gel electrophoresis (2-DE) maps of the proteins, extracted from two AML cell lines, HL-60/DOX and HL-60, were established respectively. The extracted proteins were digested by enzymes and identified with the matrix assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF-MS). The data of the peptide mass fingerprinting (PMF) was matched with databases of proteomics available on the Internet. Results showed that 16 proteins were identified to be differentially expressed between HL-60/DOX and HL-60 cells. They involved the protein disulfide isomerase precursor (PDI), the proteasomes alpha1 and other proteins which are related to drug resistance or cell metabolism, but their functional significances are required further investigation. Nevertheless, it is clear that this proteomic approach for studing the biology and development of MDR is a prerequisite in leukemia.

Drug Resistance, Multiple↗

Improving sensitivity in shotgun proteomics using a peptide-centric database with reduced complexity: protease cleavage and SCX elution rules from data mining of MS/MS spectra.

Correct identification of a peptide sequence from MS/MS data is still a challenging research problem, particularly in proteomic analyses of higher eukaryotes where protein databases are large. The scoring methods of search programs often generate cases where incorrect peptide sequences score higher than correct peptide sequences (referred to as distraction). Because smaller databases yield less distraction and better discrimination between correct and incorrect assignments, we developed a method for editing a peptide-centric database (PC-DB) to remove unlikely sequences and strategies for enabling search programs to utilize this peptide database. Rules for unlikely missed cleavage and nontryptic proteolysis products were identified by data mining 11 849 high-confidence peptide assignments. We also evaluated ion exchange chromatographic behavior as an editing criterion to generate subset databases. When used to search a well-annotated test data set of MS/MS spectra, we found no loss of critical information using PC-DBs, validating the methods for generating and searching against the databases. On the other hand, improved confidence in peptide assignments was achieved for tryptic peptides, measured by changes in DeltaCN and RSP. Decreased distraction was also achieved, consistent with the 3-9-fold decrease in database size. Data mining identified a major class of common nonspecific proteolytic products corresponding to leucine aminopeptidase (LAP) cleavages. Large improvements in identifying LAP products were achieved using the PC-DB approach when compared with conventional searches against protein databases. These results demonstrate that peptide properties can be used to reduce database size, yielding improved accuracy and information capture due to reduced distraction, but with little loss of information compared to conventional protein database searches.

Amino Acid Sequence↗

Proteomics analysis of the interactome of N-myc downstream regulated gene 1 and its interactions with the androgen response program in prostate cancer cells.

NDRG1 is known to play important roles in both androgen-induced cell differentiation and inhibition of prostate cancer metastasis. However, the proteins associated with NDRG1 function are not fully enumerated. Using coimmunoprecipitation and mass spectrometry analysis, we identified 58 proteins that interact with NDRG1 in prostate cancer cells. These proteins include nuclear proteins, adhesion molecules, endoplasmic reticulum (ER) chaperons, proteasome subunits, and signaling proteins. Integration of our data with protein-protein interaction data from the Human Proteome Reference Database allowed us to build a comprehensive interactome map of NDRG1. This interactome map consists of several modules such as a nuclear module and a cell membrane module; these modules explain the reported versatile functions of NDRG1. We also determined that serine 330 and threonine 366 of NDRG1 were phosphorylated and demonstrated that the phosphorylation of NDRG1 was prominently mediated by protein kinase A (PKA). Further, we showed that NDRG1 directly binds to beta-catenin and E-cadherin. However, the phosphorylation of NDRG1 did not interrupt the binding of NDRG1 to E-cadherin and beta-catenin. Finally, we showed that the inhibition of NDRG1 expression by RNA interference decreased the ER inducible chaperon GRP94 expression, directly proving that NDRG1 is involved in the ER stress response. Intriguingly, we observed that many members of the NDRG1 interactome are androgen-regulated and that the NDRG1 interactome links to the androgen response network through common interactions with beta-catenin and heat shock protein 90. Therefore we overlaid the transcriptomic expression changes in the NDRG1 interactome in response to androgen treatment and built a dual dynamic picture of the NDRG1 interactome in response to androgen. This interactome map provides the first road map for understanding the functions of NDRG1 in cells and its roles in human diseases, such as prostate cancer, which can progress from androgen-dependent curable stages to androgen-independent incurable stages.

Androgens↗

Peptide mass fingerprint sequence coverage from differently stained proteins on two-dimensional electrophoresis patterns by matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS).

Identification of proteins separated by two-dimensional electrophoresis (2-DE) is a necessary task to overcome the purely descriptive character of 2-DE and a prerequisite to the construction of 2-DE databases in proteome projects. Matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS) has a sensitivity for peptide detection in the lower fmol range, which should be sufficient for an analysis of even weakly silver-stained protein spots by peptide mass fingerprinting. Unfortunately, proteins are modified by the silver staining procedure, leading to low sequence coverage. Omission of glutaraldehyde increased the sequence coverage, but this improved sequence coverage is still clearly below the sequence coverage starting with Coomassie Brilliant Blue (CBB) R-250-stained spots. Other factors additionally seem to modify proteins during silver staining. By decreasing the protein amount, the advantage of very sensitive detection on the gel is lost during identification, because the resulting low sequence coverage is not sufficient for secure identification. Low-quantity proteins can be identified better starting with CBB G-250 or Zn-imidazol-stained proteins. In contrast, for high-quantity CBB R-250-stained spots, a sequence coverage of up to 90% can be obtained by using only one cleaving enzyme, and up to 80% was reached for medium-quantity spots after combination of tryptic digest with Asp-N- and Glu-C digest.

Electrophoresis, Gel, Two-Dimensional↗

A role for oligonucleotide-based RNA-knock down technologies in functional genomics.

Functional genomics is inundating the pharmaceutical industry with large numbers of potential gene targets from several sources such as gene expression profiling experiments (DNA microchips, proteomics) or database mining. Oligonucleotide-based RNA-knock down technologies such as antisense or RNA interference can aid in the filtering and prioritization of target candidates in the drug discovery process.

Drug Industry↗

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases↗

Advancing proteomic discovery through optimized multi-stage scoring and deep learning-enhanced open search.

MOTIVATION: Protein search engines are essential for interpreting mass spectrometry data into biological insight. Current tools often face limitations in sensitivity when analyzing complex modern datasets, and lack a unified framework that effectively integrates deep learning features for both restricted and open searches, especially for scenarios aimed at discovering unknown modifications. RESULTS: We present pFind+, a high-performance search engine for data-dependent acquisition (DDA) proteomics, extending pFind. It introduces an enhanced raw scoring that delivers substantially improved pre-filtering ability, while recovering most of the computational overhead through a tailored acceleration strategy. Coupled with an enhanced rescoring framework that effectively integrates deep learning features, pFind+ uniquely supports high-sensitivity, DL-enhanced open search, enabling comprehensive PTM discovery while incorporating hardware-aware inference optimizations for practical deployment. Evaluations across diverse datasets demonstrate its superior sensitivity, with gains of 12.7%-29.3% (average 17.9%) in restricted search and 8.0%-38.4% (average 25.8%) in open search over the best existing tools.

Deep Learning↗

[Recent trends in protein structural studies].

Since the 1980's, structural studies of proteins have changed remarkably. It is currently possible to predict the entire amino acid sequence of a protein by the rapid and highly sensitive analysis of the nucleotide sequence of genomic DNA or cDNA encoding the protein. In the near future, the entire sequence of a protein may be predicted from a partial sequence just by searching a variety of databases now being constructed for many biological species. The predicted protein sequence, however, is the backbone structure of the precursor protein without post-translational modifications. Therefore, the major objectives of recent structural studies of proteins are directed to 1) rapid and sensitive confirmation of the predicted sequence and identification of those modifications present in mature proteins by newly developed mass spectrometry, 2) determination of the 3D structures of intact and mutant proteins isolated or expressed in cultured E. coli, yeast or animal cells using X-ray crystallography or NMR analysis, and 3) rapid prediction of the 3D structures of proteins utilizing protein databases. The "PROTEOME" project was proposed in 1998 to bring together all the data on the structure and function of mature proteins under international cooperation. The present paper summarizes such recent trends in protein structural studies.

Mass Spectrometry↗

A database of unique protein sequence identifiers for proteome studies.

In proteome studies, identification of proteins requires searching protein sequence databases. The public protein sequence databases (e.g., NCBInr, UniProt) each contain millions of entries, and private databases add thousands more. Although much of the sequence information in these databases is redundant, each database uses distinct identifiers for the identical protein sequence and often contains unique annotation information. Users of one database obtain a database-specific sequence identifier that is often difficult to reconcile with the identifiers from a different database. When multiple databases are used for searches or the databases being searched are updated frequently, interpreting the protein identifications and associated annotations can be problematic. We have developed a database of unique protein sequence identifiers called Sequence Globally Unique Identifiers (SEGUID) derived from primary protein sequences. These identifiers serve as a common link between multiple sequence databases and are resilient to annotation changes in either public or private databases throughout the lifetime of a given protein sequence. The SEGUID Database can be downloaded (http://bioinformatics.anl.gov/SEGUID/) or easily generated at any site with access to primary protein sequence databases. Since SEGUIDs are stable, predictions based on the primary sequence information (e.g., pI, Mr) can be calculated just once; we have generated approximately 500 different calculations for more than 2.5 million sequences. SEGUIDs are used to integrate MS and 2-DE data with bioinformatics information and provide the opportunity to search multiple protein sequence databases, thereby providing a higher probability of finding the most valid protein identifications.

Amino Acid Sequence↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗