Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Sociomics! Using the IssueCrawler to map, monitor and engage with the global proteomics research network.

We invite comment upon an experiment to locate proteomics on the WWW using a software tool called the IssueCrawler. We call our research "sociomics" because, like the bioscience omics, it is a semi-automated, computerised approach to the global analysis of data, whose computerised results can be integrated towards the development of a new "systems sociology" approach to the study of society. Our findings are that proteomics on the web is a scale-free network whose nodes display considerable "dynamic range".

Cluster Analysis↗

SYSTOMONAS--an integrated database for systems biology analysis of Pseudomonas.

To provide an integrated bioinformatics platform for a systems biology approach to the biology of pseudomonads in infection and biotechnology the database SYSTOMONAS (SYSTems biology of pseudOMONAS) was established. Besides our own experimental metabolome, proteome and transcriptome data, various additional predictions of cellular processes, such as gene-regulatory networks were stored. Reconstruction of metabolic networks in SYSTOMONAS was achieved via comparative genomics. Broad data integration is realized using SOAP interfaces for the well established databases BRENDA, KEGG and PRODORIC. Several tools for the analysis of stored data and for the visualization of the corresponding results are provided, enabling a quick understanding of metabolic pathways, genomic arrangements or promoter structures of interest. The focus of SYSTOMONAS is on pseudomonads and in particular Pseudomonas aeruginosa, an opportunistic human pathogen. With this database we would like to encourage the Pseudomonas community to elucidate cellular processes of interest using an integrated systems biology strategy. The database is accessible at http://www.systomonas.de.

Bacterial Proteins↗

Intregrated analysis of the human cardiac transcriptome, proteome and phosphoproteome.

Altered expression of different classes of genes has been shown to differentiate between failing and nonfailing human hearts. However, characterization of proteins and the post-translational modifications that regulate their functions is required for understanding both the physiology of cardiac muscle and the mechanisms leading to pathological states associated with cardiac diseases. We present in this paper, an analysis of the human cardiac transcriptome, proteome and phosphoproteome. Data from two sources (i) experiments performed in our laboratory and (ii) bioinformatics searches of public databases (SWISS-PROT, NCBI, Cardiac Gene Expression Knowledge Base, Gene Ontology Consortium and Affymetrix) are reported in a relational database that allows user-designed specific queries. Microarray experiments were performed with Affymetrix Hu95Av2. Cardiac proteins were digested with trypsin. An 11 step cation exchange procedure produced fractions for analysis in separate reversed phase high-performance liquid chromatography-tandem mass spectrometry (MS/MS) experiments. Immobilized metal affinity chromatography was used to select the phosphopeptides from the same tryptic peptide mixture. They were then further investigated by MS/MS. Gel-free approaches were used to detect 267 proteins and 47 phosphopeptides. Our human cardiac database contains 447 entries. We propose the use of this platform, built with data derived from nonfailing hearts, as a template for initiating the effort to characterize the human cardiac proteome and its associated post-translational modifications.

Chromatography, Affinity↗

A Proteogenomic Approach to Discover Novel lncRNA-Derived Microproteins and Their Potential Clinical Utility in Hepatocellular Carcinoma.

Microproteins (i.e., peptides) are increasingly recognized for their functions in versatile biological contexts, but their clinical relevance and utility remain largely unexplored. Proteogenomic approaches can accelerate microprotein discovery in clinical samples by integrating proteomic data with genomics and transcriptomics evidence. However, long noncoding RNA (lncRNA)-derived microproteins (lncPeps) remain largely unidentified, resulting in unmatchable MS/MS spectra. To solve this problem, we have used high-quality Ribo-seq translatomic datasets to generate an extensive database of human liver lncRNA-derived open reading frames (lncORFs), which we subsequently applied to proteomics data of tumor-adjacent normal tissue pairs from hepatocellular carcinoma (HCC) patients. Using the new database, we discovered 104 novel lncPeps, including 46 lncPeps differentially expressed between tumor and nontumor tissues, and 13 lncPeps with significant correlation with prognosis. Remarkably, combining the expression of lncPeps with canonical proteins in a LASSO regression model improved predictive performance for recurrence, increasing the AUC by 0.005 to 0.085 across three recurrence time points. These findings suggest that the discovery of lncPeps contributes to our understanding of the molecular heterogeneity and progression of HCC and broadens the range of potential biomarker candidates and treatment targets for the disease.

Humans↗

The UCSC Genome Browser Database: update 2006.

The University of California Santa Cruz Genome Browser Database (GBD) contains sequence and annotation data for the genomes of about a dozen vertebrate species and several major model organisms. Genome annotations typically include assembly data, sequence composition, genes and gene predictions, mRNA and expressed sequence tag evidence, comparative genomics, regulation, expression and variation data. The database is optimized to support fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. The Genome Browser displays a wide variety of annotations at all scales from single nucleotide level up to a full chromosome. The Table Browser provides direct access to the database tables and sequence data, enabling complex queries on genome-wide datasets. The Proteome Browser graphically displays protein properties. The Gene Sorter allows filtering and comparison of genes by several metrics including expression data and several gene properties. BLAT and In Silico PCR search for sequences in entire genomes in seconds. These tools are highly integrated and provide many hyperlinks to other databases and websites. The GBD, browsing tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Amino Acid Sequence↗

PHProteomicDB: a module for two-dimensional gel electrophoresis database creation on personal web sites.

PHProteomicDB is a PHP-written module to help researchers in proteomics to share two-dimensional electrophoresis gel data using personal web sites. No technical or PHP knowledge is necessary except a few basics about web site management. PHProteomicDB has a user-friendly administration interface to enter and update data. It creates web pages on the fly displaying gel characteristics, gel pictures, and numbered gel spots with their related identifications pointing to their reference pages in protein databanks. The module is freely available at http://www.huvec.com/index.php3?rub=Download.

Animals↗

Annotating the human proteome.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner.

Databases, Factual↗

Nomenclature and structural biology of allergens.

Purified allergens are named using the systematic nomenclature of the Allergen Nomenclature Sub-Committee of the World Health Organization and International Union of Immunological Societies. The system uses abbreviated Linnean genus and species names and an Arabic number to indicate the chronology of allergen purification. Most major allergens from mites, animal dander, pollens, insects, and foods have been cloned, and more than 40 three-dimensional allergen structures are in the Protein Database. Allergens are derived from proteins with a variety of biologic functions, including proteases, ligand-binding proteins, structural proteins, pathogenesis-related proteins, lipid transfer proteins, profilins, and calcium-binding proteins. Biologic function, such as the proteolytic enzyme allergens of dust mites, might directly influence the development of IgE responses and might initiate inflammatory responses in the lung that are associated with asthma. Intrinsic structural or biologic properties might also influence the extent to which allergens persist in indoor and outdoor environments or retain their allergenicity in the digestive tract. Analyses of the protein family database suggest that the universe of allergens comprises more than 120 distinct protein families. Structural biology and proteomics define recombinant allergen targets for diagnostic and therapeutic purposes and identify motifs, patterns, and structures of immunologic significance.

Air Pollution, Indoor↗

Analyzing proteomes and protein function using graphical comparative analysis of tandem mass spectrometry results.

Although generating large amounts of proteomic data using tandem mass spectrometry has become routine, there is currently no single set of comprehensive tools for the rigorous analysis of tandem mass spectrometry results given the large variety of possible experimental aims. Currently available applications are typically designed for displaying proteins and posttranslational modifications from the point of view of the mass spectrometrist and are not versatile enough to allow investigators to develop biological models of protein function, protein structure, or cell state. In addition, storage and dissemination of mass spectrometry-based proteomic data are problems facing the scientific community. To address these issues, we have developed a relational database model that efficiently stores and manages large amounts of tandem mass spectrometry results. We have developed an integrated suite of multifunctional analysis software for interpreting, comparing, and displaying these results. Our system, Bioinformatic Graphical Comparative Analysis Tools (BIGCAT), allows sophisticated analysis of tandem mass spectrometry results in a biologically intuitive format and provides a solution to many data storage and dissemination issues.

Amino Acid Sequence↗

Rank information: a structure-independent measure of evolutionary trace quality that improves identification of protein functional sites.

Protein functional sites are key targets for drug design and protein engineering, but their large-scale experimental characterization remains difficult. The evolutionary trace (ET) is a computational approach to this problem that has been useful in a variety of case studies, but its proteomic scale application is partially hindered because automated retrieval of input sequences from databases often includes some with errors that degrade functional site identification. To recognize and purge these sequences, this study introduces a novel and structure-free measure of ET quality called rank information (RI). It is shown that RI decreases in response to errors in sequences, alignments, or functional classifications. Conversely, an automated procedure to increase RI by selectively removing sequences improves functional site identification so as to nearly match manually curated traces in kinases and in a test set of 79 diverse proteins. Thus we conclude that RI partially reflects the evolutionary consistency of sequence, structure, and function. In practice, as the size of the proteome continues to grow exponentially, it provides a novel and structure-free measure of ET quality that increases its accuracy for large-scale automated annotation of protein functional sites.

Algorithms↗

Identification and functional analysis of 'hypothetical' genes expressed in Haemophilus influenzae.

The progress in genome sequencing has led to a rapid accumulation in GenBank submissions of uncharacterized 'hypothetical' genes. These genes, which have not been experimentally characterized and whose functions cannot be deduced from simple sequence comparisons alone, now comprise a significant fraction of the public databases. Expression analyses of Haemophilus influenzae cells using a combination of transcriptomic and proteomic approaches resulted in confident identification of 54 'hypothetical' genes that were expressed in cells under normal growth conditions. In an attempt to understand the functions of these proteins, we used a variety of publicly available analysis tools. Close homologs in other species were detected for each of the 54 'hypothetical' genes. For 16 of them, exact functional assignments could be found in one or more public databases. Additionally, we were able to suggest general functional characterization for 27 more genes (comprising approximately 80% total). Findings from this analysis include the identification of a pyruvate-formate lyase-like operon, likely to be expressed not only in H.influenzae but also in several other bacteria. Further, we also observed three genes that are likely to participate in the transport and/or metabolism of sialic acid, an important component of the H.influenzae lipo-oligosaccharide. Accurate functional annotation of uncharacterized genes calls for an integrative approach, combining expression studies with extensive computational analysis and curation, followed by eventual experimental verification of the computational predictions.

Amino Acid Sequence↗

Two-dimensional electrophoresis resources available from ExPASy.

This paper describes the set of two-dimensional electrophoresis (2-DE) resources currently available from the ExPASy proteomics Web server. These resources include the SWISS-2DPAGE database, 2-DE software packages, 2-DE technical and educational services, as well as indexes and search engines for 2-DE related sites over the Internet.

Databases, Factual↗

Protein production and crystallization at the joint center for structural genomics.

By definition, structural genomics centers must be able to address a large number of diverse protein targets. The methods developed should permit parallel and cost-effective processing while allowing for the diverse nature of proteins. Our approach to this problem is a multi-tiered effort where targets are characterized and categorized by behavior and processed in parallel by appropriate methods. The Joint Center for Structural Genomics (JCSG) has applied this tactic to create a fully integrated and scaleable structure determination pipeline. Highlights of the development of the current pipeline for protein production and crystallization are presented here.

Crystallization↗

Integrated analysis reveals the impact of obesity on triple-negative breast cancer.

Triple-negative breast cancer (TNBC) is a highly aggressive and heterogeneous breast cancer subtype with limited therapeutic options. While the prevalence of overweight/obese (OW/OB) women continues to rise, the impact of obesity on molecular features of TNBC remains incompletely understood. We investigated clinicopathological and molecular data (including genomic, transcriptomic, proteomic and metabolomic profiling) using our original multi-omics database of TNBC (N = 465) for associations with patient body mass index (BMI). Multi-omics profiling revealed that OW/OB patients exhibited worse survival as well as elevated inflammation of tumor microenvironment, higher expression of immune checkpoints, and dysregulated lipid metabolism. Our in vivo experiments demonstrated that tumors in obese mice displayed faster growth rates, a higher proportion of PD-1+CD8+ T cells and enhanced responsiveness to anti-PD-1 treatment. In addition, we analyzed data from four independent clinical trials and discovered that OW/OB patients demonstrated higher pathological complete response rates and longer progression-free survival following anti-PD-1-based immunotherapy. In conclusion, our study systematically revealed that obesity is associated with coordinated immune-metabolic remodeling in TNBC, characterized by checkpoint enrichment and lipid dysregulation, which may help explain the enhanced anti-PD-1 responsiveness and should be taken into account in the field of precision medicine.

Immunity↗

Chemoproteomics-driven drug discovery: addressing high attrition rates.

The advent of multiple high-throughput technologies has brought drug discovery round almost full circle, from pharmacological testing of compounds in vivo to engineered molecular target assays and back to integrated phenotypic screens in cells and organisms. In the past, primary screens to identify new pharmacological agents involved administering compounds to an animal and monitoring a pharmacologic endpoint. For example, antihypertensive agents were identified by dosing spontaneously hypertensive rats with compounds and observing whether their blood pressure dropped. In taking this phenomenological approach, scientists were focused on the final goal, in this example lowering of blood pressure, rather than developing an understanding of the target, or targets, the compounds were impacting. With the evolution of rational target-based approaches, scientists were able to study the direct interaction of compounds with their intended targets, expecting that this would lead to more-selective and safer therapeutics. With the industrialization of screening, referred to as HTS, hundreds of thousands of compounds were screened in robot-driven assays against targets of interest (with this goal in mind). However, an unintentional outcome of the migration from in vivo primary screens to highly target-specific HTS assays was a reduction in biological context caused by the separation of the target from other cellular proteins and processes that might impact its function. Recognition of the potential consequences of this over-simplification drove the modification of HTS processes and equipment to be compatible with cellular assays.

Animals↗

Modelling gene networks at different organisational levels.

Approaches to modelling gene regulation networks can be categorized, according to increasing detail, as network parts lists, network topology models, network control logic models, or dynamic models. We discuss the current state of the art for each of these approaches. There is a gap between the parts list and topology models on one hand, and control logic and dynamic models on the other hand. The first two classes of models have reached a genome-wide scale, while for the other model classes high throughput technologies are yet to make a major impact.

Algorithms↗

The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry.

Donald Hunt has made seminal contributions to the fields of proteomics, immunology, epigenetics, and glycobiology. The foundation of every important work to come out of the Hunt Laboratory is de novo peptide sequencing. For decades, he taught hundreds of students, postdocs, engineers, and scientists to directly interpret mass spectral data. To honor his legacy and ensure that the art of de novo sequencing is not lost, we have adapted his teaching materials into "The Hunt Lab Guide to De Novo Peptide Sequence Analysis by Tandem Mass Spectrometry". In addition to the de novo sequencing tutorials, we present two freely available software tools that facilitate manual interpretation of mass spectra and validation of search results. The first, "Hunt Lab Peptide Fragment Calculator", calculates precursor and fragment mass-to-charge ratios for any peptide. The second program, "Predator Protein Fragment Calculator", was inspired in part by the fragment calculator developed in the Hunt Lab. Its capabilities are enhanced to facilitate interpretation of mass spectral data derived from intact proteins. We hope that the combination of these educational tools will continue to benefit students and researchers by empowering them to interpret data on their own.

Tandem Mass Spectrometry↗