Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Separation and characterization of needle and xylem maritime pine proteins.

Two-dimensional gel electrophoresis (2-DE) and image analysis are currently used for proteome analysis in maritime pine (Pinus pinaster Ait.). This study presents a database of expressed proteins extracted from needles and xylem, two important tissues for growth and wood formation. Electrophoresis was carried out by isoelectric focusing (IEF) in the first dimension and sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) in the second. Silver staining made it possible to detect an average of 900 and 600 spots on 2-DE gels from needles and xylem, respectively. A total of 28 xylem and 35 needle proteins were characterized by internal peptide microsequencing. Out of these 63 proteins, 57 (90%) could be identified based on amino acid similarity with known proteins, of which 24 (42%) have already been described in conifers. Overall comparison of both tissues indicated that 29% and 36% of the spots were specific to xylem and needles, respectively, while the other spots were of identical molecular weight and isoelectric point. The homology of spot location in 2-DE patterns was further validated by sequence analysis of proteins present in both tissues. A proteomic database of maritime pine is accessible on the internet (http://www.pierroton.inra.fr/genetics/2D/).

Acrylic Resins↗

Prediction of error associated with false-positive rate determination for peptide identification in large-scale proteomics experiments using a combined reverse and forward peptide sequence database strategy.

In recent years, a variety of approaches have been developed using decoy databases to empirically assess the error associated with peptide identifications from large-scale proteomics experiments. We have developed an approach for calculating the expected uncertainty associated with false-positive rate determination using concatenated reverse and forward protein sequence databases. After explaining the theoretical basis of our model, we compare predicted error with the results of experiments characterizing a series of mixtures containing known proteins. In general, results from characterization of known proteins show good agreement with our predictions. Finally, we consider how these approaches may be applied to more complicated data sets, as when peptides are separated by charge state prior to false-positive determination.

Algorithms↗

Immunoproteomics of outer membrane proteins and extracellular proteins of Shigella flexneri 2a 2457T.

Shigella flexneri 2a is an important pathogen causing bacillary dysentery in humans. In order to investigate any potential vaccine candidate proteins present in outer membrane proteins (OMPs) and extracellular proteins of S. flexneri 2a 2457T, we use the proteome mapping and database analyzing techniques. A subproteome map and database of OMPs were established first. One hundred and nine of the total 126 marked spots were cut out and processed to MALDI-TOF-MS and PMF. Eighty-seven spots were identified and they represented 55 OMP entries. Furthermore, immunoproteomics analysis of OMPs and extracellular proteins were performed. Total of 34 immunoreactive spots were identified, in which 22 and 12 were from OMPs and extracellular proteins, respectively. Eight novel antigens were found and some of these antigens may be potential vaccine candidate proteins. These results are useful for future studying of pathogenicity, vaccine, and novel antibacterial drugs. Maps and tables of all identified proteins are available on the Internet at www.proteomics.com.cn.

Bacterial Outer Membrane Proteins↗

The Proteomic Landscape of CTNNB1 Mutated Low-Grade Early-Stage Endometrial Carcinomas.

Endometrial carcinoma is the most frequent gynecologic malignancy in western countries. In recent years, mutations in CTNNB1 have been associated with worse prognosis in low-risk carcinomas. However, there is a lack of understanding of the proteomic implications of CTNNB1 mutations in this type of tumor. In this study, we performed shotgun proteomics using Formalin-Fixed Paraffin-Embedded (FFPE) tissue samples of CTNNB1 mutated and wild-type low-risk endometrial carcinomas. A publicly available proteomic and transcriptomic database was used to validate results. Differential protein expression and Gene Set Enrichment Analysis revealed dysregulation of pathways associated with cell keratinization, immune response modulation, and intracellular calcium regulation. CTNNB1 mutated tumors showed immune dysregulation at multiple levels including cytokine secretion, cell adhesion, and lymphocyte activation. These results were supported by tissue multiplex immunofluorescence analysis, demonstrating reduced CD8 tumor-infiltrating lymphocytes and different immune spatial interaction patterns. Intracellular calcium dysfunction was associated with key transcript dysregulation. We found an increased expression of CAMK2A and ROR2, suggesting a potential role for non-canonical Wnt pathway activation in CTNNB1 mutated tumors.

Humans↗

MALDI-TOF/TOF de novo sequence analysis of 2-D PAGE-separated proteins from Halorhodospira halophila, a bacterium with unsequenced genome.

Because protein identifications rely on matches with sequence databases, high-throughput proteomics is currently largely restricted to those species for which comprehensive sequence databases are available. The identification of proteins derived from organisms with unsequenced genomes mainly depends on homology searching. Here, we report the use of a simplified, gel-based, chemical derivatization strategy for de novo sequence analysis using a MALDI-TOF/TOF mass spectrometer. This approach allows the determination of de novo peptide sequences of up to 20 amino acid residues in length. The protocol was applied on a proteomic study of 2-D PAGE-separated proteins from Halorhodospira halophila, an extremophilic eubacterium with yet unsequenced genome. Using three different homology-based search algorithms, we were able to identify more than 30 proteins from this organism using subpicomole quantities of protein.

Amino Acid Sequence↗

Proteomics and immunohistochemistry define some of the steps involved in the squamous differentiation of the bladder transitional epithelium: a novel strategy for identifying metaplastic lesions.

Here, we present a novel strategy for dissecting some of the steps involved in the squamous differentiation of the bladder urothelium leading to squamous cell carcinomas (SCCs). First, we used proteomic technologies and databases (http://biobase.dk/cgi-bin/celis) to reveal proteins that were expressed specifically by fresh normal urothelium and three SCCs showing no urothelial components. Thereafter, antibodies against some of the differentially expressed proteins as well as a few known keratinocyte markers were used to stain serial cryostat sections (immunowalking) of biopsies obtained from bladder cystectomies of two of the SCC-bearing patients (884-1 and 864-1). Because bladder cancer is a field disease, we surmised that the urothelium of these patients may exhibit a spectrum of abnormalities ranging from early metaplastic stages to invasive disease. Immunohistochemical analysis revealed three types of non-keratinizing metaplastic lesions (types 1-3) that did not express keratins 7, 8, 18, and 20 (expressed by normal urothelium) and could be distinguished based on their staining with keratin 19 antibodies. Type 1 lesions showed staining of all cell layers in the epithelium (with differences in the staining intensity of the basal compartment), whereas type 2 lesions exhibited mainly basal cell staining. Type 3 lesions did not stain with keratin 19 antibodies. In cystectomy 884-1, type 3 lesions exhibited the same immunophenotype as the SCC and may be regarded as precursors to the tumor. Basal cells in these lesions did not express keratin 13, suggesting that the tumor, which was also keratin 13 negative, may have arisen from the expansion of these cells. Similar results were observed with cystectomy 864-1, which showed carcinoma in situ of the SCC type. SCC 864-1 exhibited both keratin 19-negative and -positive cells, implying that the tumor arose from the expansion of the basal cell compartment of type 2 and 3 lesions. Besides providing with a novel strategy for revealing metaplastic lesions, our studies have shown that it is feasible to apply powerful proteomic technologies to the analysis of complex biological samples under conditions that are as close as possible to the in vivo situation.

Biomarkers, Tumor↗

pFind: a novel database-searching software system for automated peptide and protein identification via tandem mass spectrometry.

SUMMARY: Research in proteomics requires powerful database-searching software to automatically identify protein sequences in a complex protein mixture via tandem mass spectrometry. In this paper, we describe a novel database-searching software system called pFind (peptide/protein Finder), which employs an effective peptide-scoring algorithm that we reported earlier. The pFind server is implemented with the C++ STL, .Net and XML technologies. As a result, high speed and good usability of the software are achieved.

Algorithms↗

Heme oxygenase 1 (HO-1) is a drug target for reversing cisplatin resistance in non-small cell lung cancer.

INTRODUCTION: Platinum-based drugs, the most widely used chemotherapeutic drugs in clinical oncology, have long faced the problem of drug resistance, which is urgently in need of resolution. Identifying biomarkers of drug resistance may help reduce platinum resistance and improve therapeutic efficacy. OBJECTIVES: This study aims to identify potential biomarkers associated with the development of cisplatin resistance in non-small cell lung cancer (NSCLC) and explore mechanisms to overcome chemoresistance. METHODS: NSCLC cisplatin resistance cell lines were constructed, and transcriptome sequencing was performed. Results were validated using Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA) databases. Molecular docking, proteomics sequencing, and in vitro and in vivo experiments were conducted to evaluate the role of Heme Oxygenase 1 (HO-1) in cisplatin resistance. RESULTS: NSCLC cisplatin resistance cell lines, GEO and TCGA data identified HMOX1, downstream of Nrf2, as a key drug resistance gene induced by cisplatin. Activation of the Nrf2/HO-1 pathway was found to induce ferroptosis resistance, a critical mechanism of cisplatin resistance. Candidate compounds SB 202190 and Nordihydroguaiaretic acid (NDGA) effectively reactivated ferroptosis by inhibiting HO-1, thereby increasing cisplatin sensitivity. CONCLUSION: The Nrf2/HO-1 pathway is a significant contributor to cisplatin resistance in NSCLC. Targeting HO-1 with SB 202190 and NDGA presents a promising strategy to overcome resistance and improve chemotherapy outcomes.

Cisplatin↗

Focusing of gene expression as the basis of stem cell differentiation.

In a prior report (Stem Cells Dev 14(4):354-366, 2005), we employed two-dimensional gel electrophoresis followed by advanced proteomics and the Database for Annotation, Visualization and Integrated Discovery (DAVID) to compare the protein expression profiles of mesenchymal stem cells to that of fully differentiated osteoblasts. These data were reported to advance technical approaches to define the basis of differentiation, but also led us to suggest that osteogenic differentiation of stem cells may result from the focusing of gene expression in functional clusters (e.g., calcium-regulated signaling proteins or adherence proteins) rather than simply from the induced expression of new genes, as many have assumed. Here, we have employed these analytical techniques to compare protein expression by mesenchymal stem cells directly with that of cells derived from them after induced osteogenic differentiation. Our results support the concept of gene focusing as the basis of differentiation. Specifically, induced differentiation results in a decrease in the number of mesenchymal cell markers and calcium-mediated signaling molecules expressed by their differentiated progeny. This effect was seen in parallel to increased expression of specific extracellular matrix (ECM) molecules and their receptors. These results strongly imply that changes in the ECM have a direct impact on stem cell differentiation, and that osteogenic differentiation of stem cells directed by matrix clues results from focusing of the expression of genes involved in Ca2+-dependent signaling pathways.

Calcium Signaling↗

Large-scale functional genomic analysis of sporulation and meiosis in Saccharomyces cerevisiae.

We have used a single-gene deletion mutant bank to identify the genes required for meiosis and sporulation among 4323 nonessential Saccharomyces cerevisiae annotated open reading frames (ORFs). Three hundred thirty-four sporulation-essential genes were identified, including 78 novel ORFs and 115 known genes without previously described sporulation defects in the comprehensive Saccharomyces Genome (SGD) or Yeast Proteome (YPD) phenotype databases. We have further divided the uncharacterized sporulation-essential genes into early, middle, and late stages of meiosis according to their requirement for IME1 induction and nuclear division. We believe this represents a nearly complete identification of the genes uniquely required for this complex cellular pathway. The set of genes identified in this phenotypic screen shows only limited overlap with those identified by expression-based studies.

Genes, Fungal↗

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational↗

An SVM scorer for more sensitive and reliable peptide identification via tandem mass spectrometry.

Tandem mass spectrometry (MS/MS) has become increasingly important and indispensable in high-throughput proteomics for identifying complex protein mixtures. Database searching is the standard method to accomplish this purpose. A key sub-routine, peptide identification, is used to generate a list of candidate peptides from a protein database according to an experimental MS/MS spectrum, and then validate these candidate peptides for protein identification. Although currently there are many algorithms for peptide identification, most of them either lack an effective validation module or only validate the first-ranked peptide, thus leading to a low identification reliability or sensitivity. This paper proposes a new algorithm, named pepReap, to overcome the above drawbacks. It consists of a two-layered scoring scheme based on machine learning. The first layer is a rough scoring function which uses some simple and heuristic factors to measure the degree of the matches between an experimental MS/MS spectrum and the candidate peptides; thus a ranked list of candidate peptides is generated at a relatively low computational cost. The second layer is a fine scoring function which re-ranks the candidate peptides generated in the first layer and determines which one among them is the true positive. The fine scoring function was designed based on support vector machines (SVMs) using more comprehensive factors, such as the correlations between ions, the mass matching errors of fragment and peptide ions, etc. Consequently, the SVM classifier serves as not only a scorer but also a validation module. Experimental comparison with the popular SEQUEST algorithm coupled with threshold validation criteria on a reported dataset demonstrates that the pepReap algorithm achieves higher performance in terms of identification sensitivity with comparable precision.

Algorithms↗

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein↗

Large-scale identification of cytosolic mouse brain proteins by chromatographic prefractionation.

Proteomic studies on mouse brain protein expression are still holding center stage as the generation of a reference database for the brain proteome, a need for designing expressional studies at the protein level. We therefore decided to extend the amount of identified brain proteins by the use of prefractionation. In order to reduce the complexity of mouse brain proteome we applied chromatographic prefractionations, ion-exchange and hydrophobic interaction chromatography, prior to 2-DE, followed by mass spectrometric identification (2-DE MALDI-MS). We analyzed about 17,000 protein spots in cytosolic fractions of mouse brain and identified about 10,000 spots. A total of 1841 proteins showing different pI or M(r), representing probably post-translational modifications or splice variants, were products of 789 different genes. Numerous proteins were clearly identified as metabolic, antioxidant, cytoskeleton, signaling, transcription/translation, nucleic acid-binding, proteolysis-related proteins. We additionally provided evidence for the existence of hypothetical proteins predicted from nucleic acid sequences. Moreover, observed pIs of proteins are listed thus enabling localization of proteins in a gel, information that cannot be obtained from theoretical pI's in databases. The results represent so far the largest database of mouse brain proteins and provide valuable information for the design of proteomic studies in the mouse.

Amino Acid Sequence↗

Mouse proteome analysis.

A general overview of the protein sequence set for the mouse transcriptome produced during the FANTOM2 sequencing project is presented here. We applied different algorithms to characterize protein sequences derived from a nonredundant representative protein set (RPS) and a variant protein set (VPS) of the mouse transcriptome. The functional characterization and assignment of Gene Ontology terms was done by analysis of the proteome using InterPro. The Superfamily database analyses gave a detailed structural classification according to SCOP and provide additional evidence for the functional characterization of the proteome data. The MDS database analysis revealed new domains which are not presented in existing protein domain databases. Thus the transcriptome gives us a unique source of data for the detection of new functional groups. The data obtained for the RPS and VPS sets facilitated the comparison of different patterns of protein expression. A comparison of other existing mouse and human protein sequence sets (e.g., the International Protein Index) demonstrates the common patterns in mammalian proteomes. The analysis of the membrane organization within the transcriptome of multiple eukaryotes provides valuable statistics about the distribution of secretory and transmembrane proteins

Animals↗

Assessing the impact of alternative splicing on domain interactions in the human proteome.

We have constructed a database of alternatively spliced protein forms (ASP), consisting of 13,384 protein isoform sequences of 4422 human genes (www.bioinformatics.ucla.edu/ASP). We identified fifty protein domain types that were selectively removed by alternative splicing at much higher frequencies than average (p-value < 0.01). These include many well-known protein-interaction domains (e.g., KRAB; ankyrin repeats; Kelch) including some that have been previously shown to be regulated functionally by alternative splicing (e.g., collagen domain). We present a number of novel examples (Kruppel transcription factors; Pbx2; Enc1) from the ASP database, illustrating how this pattern of alternative splicing changes the structure of a biological pathway, by redirecting protein interaction networks at key switch points. Our bioinformatics analysis indicates that a major impact of alternative splicing is removal of protein-protein interaction domains that mediate key linkages in protein interaction networks. ASP expands the available dataset of human alternatively spliced protein forms from 1989 human genes (SwissProt release 42) to 5413 (nonredundant set, ASP + SwissProt), a nearly 3-fold increase. ASP will enhance the existing pool of protein sequences that are searched by mass spectroscopy software during the identification of peptide fragments.

Alternative Splicing↗