Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Applications of InterPro in protein annotation and genome analysis.

The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.

Amino Acid Sequence↗

The Dictyostelium discoideum proteome--the SWISS-2DPAGE database of the multicellular aggregate (slug).

The cellular slime mold Dictyostelium discoideum is a eukaryotic microorganism which has developmental life stages attractive to the cell and molecular biologist. By displaying the two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) protein map of different developmental stages, the key molecules can be identified and characterised, allowing a detailed understanding of the D. discoideum proteome. Here we describe the preparation of reference gel of the D. discoideum multicellular aggregate, the slug. Proteins were separated by 2-D PAGE with immobilised pH gradients (pH 3.5-10) in the first dimension and sodium dodecyl sulfate (SDS)-PAGE in the second dimension. Micropreparative gels were electroblotted onto polyvinylidene difluoride (PVDF) membranes and 150 spots were visualised by amido black staining. Protein spots were excised and 31 were putatively identified by matching their amino acid composition, estimated isoelectric point (pI) and molecular weight (M(r)) against the SWISS-PROT database with the ExPASy AAcompID tool (http:// expasy.hcuge.ch/ch2d/aacompi.html). A total of 25 proteins were identified by matching against database entries for D. discoideum, and another six by cross-species matching against database entries for Saccharomyces cerevisiae proteins. This map will be available in the SWISS-2DPAGE database.

Animals↗

Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.

We demonstrate an integrated approach to build, test, and refine a model of a cellular pathway, in which perturbations to critical pathway components are analyzed using DNA microarrays, quantitative proteomics, and databases of known physical interactions. Using this approach, we identify 997 messenger RNAs responding to 20 systematic perturbations of the yeast galactose-utilization pathway, provide evidence that approximately 15 of 289 detected proteins are regulated posttranscriptionally, and identify explicit physical interactions governing the cellular response to each perturbation. We refine the model through further iterations of perturbation and global measurements, suggesting hypotheses about the regulation of galactose utilization and physical interactions between this and a variety of other metabolic pathways.

Computational Biology↗

Peptide mass fingerprint sequence coverage from differently stained proteins on two-dimensional electrophoresis patterns by matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS).

Identification of proteins separated by two-dimensional electrophoresis (2-DE) is a necessary task to overcome the purely descriptive character of 2-DE and a prerequisite to the construction of 2-DE databases in proteome projects. Matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS) has a sensitivity for peptide detection in the lower fmol range, which should be sufficient for an analysis of even weakly silver-stained protein spots by peptide mass fingerprinting. Unfortunately, proteins are modified by the silver staining procedure, leading to low sequence coverage. Omission of glutaraldehyde increased the sequence coverage, but this improved sequence coverage is still clearly below the sequence coverage starting with Coomassie Brilliant Blue (CBB) R-250-stained spots. Other factors additionally seem to modify proteins during silver staining. By decreasing the protein amount, the advantage of very sensitive detection on the gel is lost during identification, because the resulting low sequence coverage is not sufficient for secure identification. Low-quantity proteins can be identified better starting with CBB G-250 or Zn-imidazol-stained proteins. In contrast, for high-quantity CBB R-250-stained spots, a sequence coverage of up to 90% can be obtained by using only one cleaving enzyme, and up to 80% was reached for medium-quantity spots after combination of tryptic digest with Asp-N- and Glu-C digest.

Electrophoresis, Gel, Two-Dimensional↗

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases↗

Advancing proteomic discovery through optimized multi-stage scoring and deep learning-enhanced open search.

MOTIVATION: Protein search engines are essential for interpreting mass spectrometry data into biological insight. Current tools often face limitations in sensitivity when analyzing complex modern datasets, and lack a unified framework that effectively integrates deep learning features for both restricted and open searches, especially for scenarios aimed at discovering unknown modifications. RESULTS: We present pFind+, a high-performance search engine for data-dependent acquisition (DDA) proteomics, extending pFind. It introduces an enhanced raw scoring that delivers substantially improved pre-filtering ability, while recovering most of the computational overhead through a tailored acceleration strategy. Coupled with an enhanced rescoring framework that effectively integrates deep learning features, pFind+ uniquely supports high-sensitivity, DL-enhanced open search, enabling comprehensive PTM discovery while incorporating hardware-aware inference optimizations for practical deployment. Evaluations across diverse datasets demonstrate its superior sensitivity, with gains of 12.7%-29.3% (average 17.9%) in restricted search and 8.0%-38.4% (average 25.8%) in open search over the best existing tools.

Deep Learning↗

[Recent trends in protein structural studies].

Since the 1980's, structural studies of proteins have changed remarkably. It is currently possible to predict the entire amino acid sequence of a protein by the rapid and highly sensitive analysis of the nucleotide sequence of genomic DNA or cDNA encoding the protein. In the near future, the entire sequence of a protein may be predicted from a partial sequence just by searching a variety of databases now being constructed for many biological species. The predicted protein sequence, however, is the backbone structure of the precursor protein without post-translational modifications. Therefore, the major objectives of recent structural studies of proteins are directed to 1) rapid and sensitive confirmation of the predicted sequence and identification of those modifications present in mature proteins by newly developed mass spectrometry, 2) determination of the 3D structures of intact and mutant proteins isolated or expressed in cultured E. coli, yeast or animal cells using X-ray crystallography or NMR analysis, and 3) rapid prediction of the 3D structures of proteins utilizing protein databases. The "PROTEOME" project was proposed in 1998 to bring together all the data on the structure and function of mature proteins under international cooperation. The present paper summarizes such recent trends in protein structural studies.

Mass Spectrometry↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗

Glycome project: concept, strategy and preliminary application to Caenorhabditis elegans.

Glycans play a central role as potential mediators between complex cell societies, because all living organisms consist of cells covered with diverse carbohydrate chains reflecting various cell types and states. However, we have no idea how diverse these carbohydrate chains actually are. The main purpose of this article is to persuade life scientists to realize the fundamental importance of taking some action by becoming involved in "glycomics". "Glycome" is a term meaning the whole set of glycans produced by individual organisms, as the third bioinformative macromolecules to be elucidated next to the genome and proteome. Here a basic strategy is presented. The essence of the project includes the following: (a) glycopeptides, but not glycans released from their core proteins, are targeted for linkage to genome databases; (b) Caenorhabditis elegans is used as the first model organism for this project, since its genome project has already been completed; (c) four essential attributes are adopted to characterize each glycopeptide: (i) cosmid identification number (ID), (ii) molecular weight (M(r)), (iii) retention (Rs) of pyridylaminated (PA) oligosaccharides in 2-D mapping, and (iv) dissociation constants (Kd's) of PA-oligosaccharides for a set of lectins. Thus, the obtained ID, M(r), R and Kd's construct the glycome database, which will be open as the previous genome and proteome databases. For the project to proceed the "glyco-catch" method is proposed, where a group of target glycopeptides are captured by means of lectin-affinity chromatography after protease digestion. Already glycopeptides from asialofetuin and ovalbumin were successfully captured by galectin-agarose and Con A-agarose, respectively. Further, to examine the practical validity of the method, we extracted membrane proteins from C. elegans with 1% Triton X-100, and isolated specific glycopeptides by use of the same galectin column. One of the glycopeptides was successfully identified in the C. elegans genome database. Finally, for determination of Kd between glycopeptides and lectins, a recently reinforced frontal affinity chromatography (FAC) is proposed as an alternative to define glycan structures in place of determining every covalent structure.

Animals↗

Construction of a Francisella tularensis two-dimensional electrophoresis protein database.

We have started the construction of a two-dimensional database of the proteome of Francisella tularensis, a bacterium that is responsible for the highly pathogenic disease tularemia. The genome of this intracellular pathogen is not completely sequenced yet and, currently, information about only 66 proteins is available from NCBI database. We have analyzed the F. tularensis live vaccine strain by two-dimensional gel electrophoresis with immobilized pH 3-10 gradient in the first dimension and 9-16% gradient or tricine SDS-PAGE in the second dimension. In both cases about 2000 spots were detected. Furthermore, we compared the protein pattern of the nonvirulent F. tularensis live vaccine strain with protein profiles of two wild type clinical isolates and more than 50 differentially expressed proteins were counted. The separated proteins are going to be identified by peptide mass fingerprinting. However, due to the lack of complete genome sequence data only eight proteins were unambiguously identified. Among them, acid phosphatase and the most basic isoform of a hypothetical 23 kDa protein are characteristic only for virulent strains.

Bacterial Proteins↗

Mining the human proteome: experience with the human lymphoid protein database.

We have undertaken an effort in the past five years aimed at developing a database of lymphoid proteins detectable by two-dimensional (2-D) polyacrylamide gel electrophoresis. The database contains 2-D patterns and derived information pertaining to: (i) polypeptide constituents of unstimulated and stimulated mature T cells and immature thymocytes; (ii) cultured T cells and cell lines that have been manipulated by transfection with a variety of constructs or by treatment with specific agents; (iii) single cell-derived T and B cell clones; (iv) cells obtained from patients with lymphoproliferative disorders and leukemia; and (v) a variety of other relevant cell populations. The database has experienced a substantial expansion in 2-D patterns it contains, numbering currently 9167 individual 2-D patterns. This number represents a fraction of the 30,682 2-D patterns maintained in our databases. The capacity to design and undertake experiments, produce high-quality 2-D patterns, and to undertake simple or rudimentary analyses of 2-D patterns to meet the basic needs of the experiments for which the 2-D gels were produced has exceeded the capacity to fully and uniformly integrate information generated from any gel image or experiment, across all images and experiments. While only a fraction of the information in the 2-D patterns contained in the lymphoid database has been mined, novel findings derived from querying the database point to the merits of this protein based approach. Additional resources have recently been put into place to mine more effectively data pertaining to protein expression in lymphoid cells.

Cell Cycle↗

Establishment of a root proteome reference map for the model legume Medicago truncatula using the expressed sequence tag database for peptide mass fingerprinting.

We have established a proteome reference map for Medicago truncatula root proteins using two-dimensional gel electrophoresis combined with peptide mass fingerprinting to aid the dissection of nodulation and root developmental pathways by proteome analysis. M. truncatula has been chosen as a model legume for the study of nodulation-related genes and proteins. Over 2,500 root proteins could be displayed reproducibly across an isoelectric focussing range of 4-7. We analysed 485 proteins by peptide mass fingerprinting, and 179 of those were identified by matching against the current M. truncatula expressed sequence tag (EST) database containing DNA sequences of approximately 105,000 ESTs. Matching the EST sequences to available plant DNA sequences by BLAST searches enabled us to predict protein function. The use of the EST database for peptide identification is discussed. The majority of identified proteins were metabolic enzymes and stress response proteins, and 44% of proteins occurred as isoforms, a result that could not have been predicted from sequencing data alone. We identified two nodulins in uninoculated root tissue, supporting evidence for a role of nodulins in normal plant development. This proteome map will be updated continuously (http://semele.anu.edu.au/2d/2d.html) and will be a powerful tool for investigating the molecular mechanisms of root symbioses in legumes.

Databases as Topic↗

A proteomic analysis of organelles from Arabidopsis thaliana.

We introduce the use of Arabidopsis thaliana callus culture as a system for proteomic analysis of plant organelles using liquid-grown callus. This callus is relatively homogeneous, reproducible and cytoplasmically rich, and provides organelles in sufficient quantities for proteomic studies. A database was generated of mitochondrial, endoplasmic reticulum (ER), Golgi/prevacuolar compartment and plasma membrane (PM) markers using two-dimensional sodium dodecyl sulphate-polyacrylamide gel electrophoresis (2-D SDS-PAGE) and peptide sequencing or mass spectrometric methods. The major callus membrane-associated proteins were characterised as being integral or peripheral by Triton X-114 phase partitioning. The database was used to define specific proteins at the Arabidopsis callus plasma membrane. This database of organelle proteins provides the basis for future characterisation of the expression and localisation of novel plant proteins.

Amino Acid Sequence↗

Separation and characterization of needle and xylem maritime pine proteins.

Two-dimensional gel electrophoresis (2-DE) and image analysis are currently used for proteome analysis in maritime pine (Pinus pinaster Ait.). This study presents a database of expressed proteins extracted from needles and xylem, two important tissues for growth and wood formation. Electrophoresis was carried out by isoelectric focusing (IEF) in the first dimension and sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) in the second. Silver staining made it possible to detect an average of 900 and 600 spots on 2-DE gels from needles and xylem, respectively. A total of 28 xylem and 35 needle proteins were characterized by internal peptide microsequencing. Out of these 63 proteins, 57 (90%) could be identified based on amino acid similarity with known proteins, of which 24 (42%) have already been described in conifers. Overall comparison of both tissues indicated that 29% and 36% of the spots were specific to xylem and needles, respectively, while the other spots were of identical molecular weight and isoelectric point. The homology of spot location in 2-DE patterns was further validated by sequence analysis of proteins present in both tissues. A proteomic database of maritime pine is accessible on the internet (http://www.pierroton.inra.fr/genetics/2D/).

Acrylic Resins↗

The Proteomic Landscape of CTNNB1 Mutated Low-Grade Early-Stage Endometrial Carcinomas.

Endometrial carcinoma is the most frequent gynecologic malignancy in western countries. In recent years, mutations in CTNNB1 have been associated with worse prognosis in low-risk carcinomas. However, there is a lack of understanding of the proteomic implications of CTNNB1 mutations in this type of tumor. In this study, we performed shotgun proteomics using Formalin-Fixed Paraffin-Embedded (FFPE) tissue samples of CTNNB1 mutated and wild-type low-risk endometrial carcinomas. A publicly available proteomic and transcriptomic database was used to validate results. Differential protein expression and Gene Set Enrichment Analysis revealed dysregulation of pathways associated with cell keratinization, immune response modulation, and intracellular calcium regulation. CTNNB1 mutated tumors showed immune dysregulation at multiple levels including cytokine secretion, cell adhesion, and lymphocyte activation. These results were supported by tissue multiplex immunofluorescence analysis, demonstrating reduced CD8 tumor-infiltrating lymphocytes and different immune spatial interaction patterns. Intracellular calcium dysfunction was associated with key transcript dysregulation. We found an increased expression of CAMK2A and ROR2, suggesting a potential role for non-canonical Wnt pathway activation in CTNNB1 mutated tumors.

Humans↗

Proteomics and immunohistochemistry define some of the steps involved in the squamous differentiation of the bladder transitional epithelium: a novel strategy for identifying metaplastic lesions.

Here, we present a novel strategy for dissecting some of the steps involved in the squamous differentiation of the bladder urothelium leading to squamous cell carcinomas (SCCs). First, we used proteomic technologies and databases (http://biobase.dk/cgi-bin/celis) to reveal proteins that were expressed specifically by fresh normal urothelium and three SCCs showing no urothelial components. Thereafter, antibodies against some of the differentially expressed proteins as well as a few known keratinocyte markers were used to stain serial cryostat sections (immunowalking) of biopsies obtained from bladder cystectomies of two of the SCC-bearing patients (884-1 and 864-1). Because bladder cancer is a field disease, we surmised that the urothelium of these patients may exhibit a spectrum of abnormalities ranging from early metaplastic stages to invasive disease. Immunohistochemical analysis revealed three types of non-keratinizing metaplastic lesions (types 1-3) that did not express keratins 7, 8, 18, and 20 (expressed by normal urothelium) and could be distinguished based on their staining with keratin 19 antibodies. Type 1 lesions showed staining of all cell layers in the epithelium (with differences in the staining intensity of the basal compartment), whereas type 2 lesions exhibited mainly basal cell staining. Type 3 lesions did not stain with keratin 19 antibodies. In cystectomy 884-1, type 3 lesions exhibited the same immunophenotype as the SCC and may be regarded as precursors to the tumor. Basal cells in these lesions did not express keratin 13, suggesting that the tumor, which was also keratin 13 negative, may have arisen from the expansion of these cells. Similar results were observed with cystectomy 864-1, which showed carcinoma in situ of the SCC type. SCC 864-1 exhibited both keratin 19-negative and -positive cells, implying that the tumor arose from the expansion of the basal cell compartment of type 2 and 3 lesions. Besides providing with a novel strategy for revealing metaplastic lesions, our studies have shown that it is feasible to apply powerful proteomic technologies to the analysis of complex biological samples under conditions that are as close as possible to the in vivo situation.

Biomarkers, Tumor↗

Heme oxygenase 1 (HO-1) is a drug target for reversing cisplatin resistance in non-small cell lung cancer.

INTRODUCTION: Platinum-based drugs, the most widely used chemotherapeutic drugs in clinical oncology, have long faced the problem of drug resistance, which is urgently in need of resolution. Identifying biomarkers of drug resistance may help reduce platinum resistance and improve therapeutic efficacy. OBJECTIVES: This study aims to identify potential biomarkers associated with the development of cisplatin resistance in non-small cell lung cancer (NSCLC) and explore mechanisms to overcome chemoresistance. METHODS: NSCLC cisplatin resistance cell lines were constructed, and transcriptome sequencing was performed. Results were validated using Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA) databases. Molecular docking, proteomics sequencing, and in vitro and in vivo experiments were conducted to evaluate the role of Heme Oxygenase 1 (HO-1) in cisplatin resistance. RESULTS: NSCLC cisplatin resistance cell lines, GEO and TCGA data identified HMOX1, downstream of Nrf2, as a key drug resistance gene induced by cisplatin. Activation of the Nrf2/HO-1 pathway was found to induce ferroptosis resistance, a critical mechanism of cisplatin resistance. Candidate compounds SB 202190 and Nordihydroguaiaretic acid (NDGA) effectively reactivated ferroptosis by inhibiting HO-1, thereby increasing cisplatin sensitivity. CONCLUSION: The Nrf2/HO-1 pathway is a significant contributor to cisplatin resistance in NSCLC. Targeting HO-1 with SB 202190 and NDGA presents a promising strategy to overcome resistance and improve chemotherapy outcomes.

Cisplatin↗