Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Integrated genomic and proteomic analyses of a systematically perturbed metabolic network.

We demonstrate an integrated approach to build, test, and refine a model of a cellular pathway, in which perturbations to critical pathway components are analyzed using DNA microarrays, quantitative proteomics, and databases of known physical interactions. Using this approach, we identify 997 messenger RNAs responding to 20 systematic perturbations of the yeast galactose-utilization pathway, provide evidence that approximately 15 of 289 detected proteins are regulated posttranscriptionally, and identify explicit physical interactions governing the cellular response to each perturbation. We refine the model through further iterations of perturbation and global measurements, suggesting hypotheses about the regulation of galactose utilization and physical interactions between this and a variety of other metabolic pathways.

Computational Biology↗

Peptide mass fingerprint sequence coverage from differently stained proteins on two-dimensional electrophoresis patterns by matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS).

Identification of proteins separated by two-dimensional electrophoresis (2-DE) is a necessary task to overcome the purely descriptive character of 2-DE and a prerequisite to the construction of 2-DE databases in proteome projects. Matrix assisted laser desorption/ionization-mass spectrometry (MALDI-MS) has a sensitivity for peptide detection in the lower fmol range, which should be sufficient for an analysis of even weakly silver-stained protein spots by peptide mass fingerprinting. Unfortunately, proteins are modified by the silver staining procedure, leading to low sequence coverage. Omission of glutaraldehyde increased the sequence coverage, but this improved sequence coverage is still clearly below the sequence coverage starting with Coomassie Brilliant Blue (CBB) R-250-stained spots. Other factors additionally seem to modify proteins during silver staining. By decreasing the protein amount, the advantage of very sensitive detection on the gel is lost during identification, because the resulting low sequence coverage is not sufficient for secure identification. Low-quantity proteins can be identified better starting with CBB G-250 or Zn-imidazol-stained proteins. In contrast, for high-quantity CBB R-250-stained spots, a sequence coverage of up to 90% can be obtained by using only one cleaving enzyme, and up to 80% was reached for medium-quantity spots after combination of tryptic digest with Asp-N- and Glu-C digest.

Electrophoresis, Gel, Two-Dimensional↗

A role for oligonucleotide-based RNA-knock down technologies in functional genomics.

Functional genomics is inundating the pharmaceutical industry with large numbers of potential gene targets from several sources such as gene expression profiling experiments (DNA microchips, proteomics) or database mining. Oligonucleotide-based RNA-knock down technologies such as antisense or RNA interference can aid in the filtering and prioritization of target candidates in the drug discovery process.

Drug Industry↗

DARKIN: a zero-shot benchmark for phosphosite-dark kinase association using protein language models.

MOTIVATION: Protein language models (pLMs) have emerged as powerful tools for capturing the intricate information encoded in protein sequences, facilitating various downstream protein prediction tasks. With numerous pLMs available, there is a critical need for diverse benchmarks to systematically evaluate their performance across biologically relevant tasks. Here, we introduce DARKIN, a zero-shot classification benchmark designed to assign phosphosites to understudied kinases, termed dark kinases. Kinases, which catalyze phosphorylation, are central to cellular signaling pathways. While phosphoproteomics enables the large-scale identification of phosphosites, determining the cognate kinase responsible for the phosphorylation event remains an experimental challenge. RESULTS: In DARKIN, we prepared training, validation, and test folds that respect the zero-shot nature of this classification problem, incorporating stratification based on kinase groups and sequence similarity. We evaluated multiple pLMs using two zero-shot classifiers: a novel, training-free k-NN-based method, and a bilinear classifier. Our findings indicate that ESM, ProtT5-XL, and SaProt exhibit superior performance on this task. DARKIN provides a challenging benchmark for assessing pLM efficacy and fosters deeper exploration of under-characterized (dark) kinases by offering a biologically relevant test bed. AVAILABILITY AND IMPLEMENTATION: The DARKIN benchmark data and the scripts for generating additional splits are publicly available at: https://github.com/tastanlab/darkin.

Protein Kinases↗

Advancing proteomic discovery through optimized multi-stage scoring and deep learning-enhanced open search.

MOTIVATION: Protein search engines are essential for interpreting mass spectrometry data into biological insight. Current tools often face limitations in sensitivity when analyzing complex modern datasets, and lack a unified framework that effectively integrates deep learning features for both restricted and open searches, especially for scenarios aimed at discovering unknown modifications. RESULTS: We present pFind+, a high-performance search engine for data-dependent acquisition (DDA) proteomics, extending pFind. It introduces an enhanced raw scoring that delivers substantially improved pre-filtering ability, while recovering most of the computational overhead through a tailored acceleration strategy. Coupled with an enhanced rescoring framework that effectively integrates deep learning features, pFind+ uniquely supports high-sensitivity, DL-enhanced open search, enabling comprehensive PTM discovery while incorporating hardware-aware inference optimizations for practical deployment. Evaluations across diverse datasets demonstrate its superior sensitivity, with gains of 12.7%-29.3% (average 17.9%) in restricted search and 8.0%-38.4% (average 25.8%) in open search over the best existing tools.

Deep Learning↗

[Recent trends in protein structural studies].

Since the 1980's, structural studies of proteins have changed remarkably. It is currently possible to predict the entire amino acid sequence of a protein by the rapid and highly sensitive analysis of the nucleotide sequence of genomic DNA or cDNA encoding the protein. In the near future, the entire sequence of a protein may be predicted from a partial sequence just by searching a variety of databases now being constructed for many biological species. The predicted protein sequence, however, is the backbone structure of the precursor protein without post-translational modifications. Therefore, the major objectives of recent structural studies of proteins are directed to 1) rapid and sensitive confirmation of the predicted sequence and identification of those modifications present in mature proteins by newly developed mass spectrometry, 2) determination of the 3D structures of intact and mutant proteins isolated or expressed in cultured E. coli, yeast or animal cells using X-ray crystallography or NMR analysis, and 3) rapid prediction of the 3D structures of proteins utilizing protein databases. The "PROTEOME" project was proposed in 1998 to bring together all the data on the structure and function of mature proteins under international cooperation. The present paper summarizes such recent trends in protein structural studies.

Mass Spectrometry↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

Large-scale open bioinformatics data resources.

The data explosion in bioinformatics is relentless. More and more genomes are being sequenced and many new types of datasets are being generated in large-scale projects. Integration and true open access to the data are still difficult issues, although they are gradually being addressed. Notably, certain fields have good standardization and interoperability, while others lag behind. This review summarizes the latest developments in genome and sequences databases, transcriptomics data (ESTs, ORESTES, full-length cDNAs), proteomics data (protein databases, protein structures, family and domain classification) as well as loosely integrated fields, such as microarray experiments, mutation databases and databases of regulatory regions and elements. The review attempts to resist simply summarizing what data are available, and aims to provide a critical look at some of the integration and access issues associated with several of these resources.

Computational Biology↗

Toward computer-based cleavage site prediction of cysteine endopeptidases.

Identification of relevant substrates is essential for elucidation of in vivo functions of peptidases. The recent availability of the complete genome sequences of many eukaryotic organisms holds the promise of identifying specific peptidase substrates by systematic proteome analyses in combination with computer-based screening of genome databases. Currently available proteomics and bioinformatics tools are not sufficient for reliable endopeptidase substrate predictions. To address these shortcomings the bioinformatics tool 'PEPS' (Prediction of Endopeptidase Substrates) has been developed and is presented here. PEPS uses individual rule-based endopeptidase cleavage site scoring matrices (CSSM). The efficiency of PEPS in predicting putative caspase 3, cathepsin B and cathepsin L cleavage sites is demonstrated in comparison to established algorithms. Mortalin, a member of the heat shock protein family HSP70, was identified by PEPS as a putative cathepsin L substrate. Comparative proteome analyses of cathepsin L-deficient and wild-type mouse fibroblasts showed that mortalin is enriched in the absence of cathepsin L. These results indicate that CSSM/PEPS can correctly predict relevant peptidase substrates.

Animals↗

Glycome project: concept, strategy and preliminary application to Caenorhabditis elegans.

Glycans play a central role as potential mediators between complex cell societies, because all living organisms consist of cells covered with diverse carbohydrate chains reflecting various cell types and states. However, we have no idea how diverse these carbohydrate chains actually are. The main purpose of this article is to persuade life scientists to realize the fundamental importance of taking some action by becoming involved in "glycomics". "Glycome" is a term meaning the whole set of glycans produced by individual organisms, as the third bioinformative macromolecules to be elucidated next to the genome and proteome. Here a basic strategy is presented. The essence of the project includes the following: (a) glycopeptides, but not glycans released from their core proteins, are targeted for linkage to genome databases; (b) Caenorhabditis elegans is used as the first model organism for this project, since its genome project has already been completed; (c) four essential attributes are adopted to characterize each glycopeptide: (i) cosmid identification number (ID), (ii) molecular weight (M(r)), (iii) retention (Rs) of pyridylaminated (PA) oligosaccharides in 2-D mapping, and (iv) dissociation constants (Kd's) of PA-oligosaccharides for a set of lectins. Thus, the obtained ID, M(r), R and Kd's construct the glycome database, which will be open as the previous genome and proteome databases. For the project to proceed the "glyco-catch" method is proposed, where a group of target glycopeptides are captured by means of lectin-affinity chromatography after protease digestion. Already glycopeptides from asialofetuin and ovalbumin were successfully captured by galectin-agarose and Con A-agarose, respectively. Further, to examine the practical validity of the method, we extracted membrane proteins from C. elegans with 1% Triton X-100, and isolated specific glycopeptides by use of the same galectin column. One of the glycopeptides was successfully identified in the C. elegans genome database. Finally, for determination of Kd between glycopeptides and lectins, a recently reinforced frontal affinity chromatography (FAC) is proposed as an alternative to define glycan structures in place of determining every covalent structure.

Animals↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

Data mining crystallization databases: knowledge-based approaches to optimize protein crystal screens.

Protein crystallization is a major bottleneck in protein X-ray crystallography, the workhorse of most structural proteomics projects. Because the principles that govern protein crystallization are too poorly understood to allow them to be used in a strongly predictive sense, the most common crystallization strategy entails screening a wide variety of solution conditions to identify the small subset that will support crystal nucleation and growth. We tested the hypothesis that more efficient crystallization strategies could be formulated by extracting useful patterns and correlations from the large data sets of crystallization trials created in structural proteomics projects. A database of crystallization conditions was constructed for 755 different proteins purified and crystallized under uniform conditions. Forty-five percent of the proteins formed crystals. Data mining identified the conditions that crystallize the most proteins, revealed that many conditions are highly correlated in their behavior, and showed that the crystallization success rate is markedly dependent on the organism from which proteins derive. Of the proteins that crystallized in a 48-condition experiment, 60% could be crystallized in as few as 6 conditions and 94% in 24 conditions. Consideration of the full range of information coming from crystal screening trials allows one to design screens that are maximally productive while consuming minimal resources, and also suggests further useful conditions for extending existing screens.

Archaeal Proteins↗

The human proteome organization (HUPO) and environmental health.

The Human Proteome Organization, or HUPO, was formed to promote research and large-scale analysis of the human proteome. By consolidating national proteome organizations into an international body, HUPO will coordinate international initiatives, biological resources, protocols, standards and data for studying the human proteome. HUPO has identified five key areas to advance study of the human proteome, specifically in bioinformatics, new technologies, the plasma proteome, cell models, and a public antibody initiative. Consideration of three major issue areas may help develop HUPO's strategy for human proteome study. First is the need to distinguish the value of high throughput platforms from discovery platforms in proteomics. Second is the importance for international planning on integrating both transcriptome and proteome data and databases. Last is that effects of the environment from chemical, physical, and biological exposures alter the expression and structure of the proteome, which become manifest in long-term adverse health effects and disease. Environmental health research stands to greatly benefit from the shared resources, data, and vision of the HUPO organization as a valuable resource in exploiting knowledge of the human proteome toward improving public health.

Computational Biology↗

Construction of a Francisella tularensis two-dimensional electrophoresis protein database.

We have started the construction of a two-dimensional database of the proteome of Francisella tularensis, a bacterium that is responsible for the highly pathogenic disease tularemia. The genome of this intracellular pathogen is not completely sequenced yet and, currently, information about only 66 proteins is available from NCBI database. We have analyzed the F. tularensis live vaccine strain by two-dimensional gel electrophoresis with immobilized pH 3-10 gradient in the first dimension and 9-16% gradient or tricine SDS-PAGE in the second dimension. In both cases about 2000 spots were detected. Furthermore, we compared the protein pattern of the nonvirulent F. tularensis live vaccine strain with protein profiles of two wild type clinical isolates and more than 50 differentially expressed proteins were counted. The separated proteins are going to be identified by peptide mass fingerprinting. However, due to the lack of complete genome sequence data only eight proteins were unambiguously identified. Among them, acid phosphatase and the most basic isoform of a hypothetical 23 kDa protein are characteristic only for virulent strains.

Bacterial Proteins↗

LIFEdb: a database for functional genomics experiments integrating information from external sources, and serving as a sample tracking system.

We have implemented LIFEdb (http://www.dkfz.de/LIFEdb) to link information regarding novel human full-length cDNAs generated and sequenced by the German cDNA Consortium with functional information on the encoded proteins produced in functional genomics and proteomics approaches. The database also serves as a sample-tracking system to manage the process from cDNA to experimental read-out and data interpretation. A web interface enables the scientific community to explore and visualize features of the annotated cDNAs and ORFs combined with experimental results, and thus helps to unravel new features of proteins with as yet unknown functions.

Automation↗

A comparative proteomics resource: proteins of Arabidopsis thaliana.

Using an integrative genome annotation pipeline (iGAP) for proteome-wide protein structure and functional domain assignment, we analyzed all the proteins of Arabidopsis thaliana. Three-dimensional structures at the level of the domain are assigned by fold recognition and threading based on a novel fold library that extends common domain classifications. iGAP is being applied to proteins from all available proteomes as part of a comparative proteomics resource. The database is accessible from the web.

Arabidopsis↗

Exploring the proteome of Plasmodium.

With the entire genomic sequence of several species of Plasmodium soon to be available, researchers are now focusing on methods to study gene and protein expression at the whole organism level. Traditional methods of characterising and identifying large numbers of proteins from a complex protein mixture have relied predominantly on two-dimensional gel electrophoresis combined with N-terminal sequencing or mass spectrometry of individually prepared proteins. New proteomics methods are now available that are based on resolving small peptides derived from complex protein mixtures by high-resolution liquid chromatography and directly identifying them by tandem mass spectrometry (LC/LC/MS/MS) and sophisticated computer search algorithms against whole genome sequence databases. These newer proteomic methods have the potential to accelerate the reproducible identification of large numbers of proteins from various life cycle stages of Plasmodium and may help to better understand parasite biology and lead to the identification of new targets of vaccines and drugs.

Animals↗