Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Fourier-transform mass spectrometry for automated fragmentation and identification of 5-20 kDa proteins in mixtures.

When presented with a mixture of intact proteins, electrospray ionization with Fourier-transform mass spectrometry (ESI-FTMS) has the capability to obtain direct fragmentation information from isolated ions. However, the automation of this capability has not been achieved to date. We have developed software for unattended acquisition of protein tandem mass spectrometry (MS/MS) data and batch processing of the resulting files for identification of whole proteins. Mixtures of both protein standards (8-29 kDa) and Methanococcus jannaschii cytosolic proteins (up to six components + 20 kDa) were infused via an autosampler, and MS/MS data were acquired without human intervention. The acquisition software recognizes ESI charge state patterns, generates protein-specific isolation waveforms on-the-fly, and fragments ions using two different infrared laser times. In addition to protein standards, five wild-type proteins (7-14 kDa) were identified automatically with 100% sequence coverage from the M. jannaschii database. The software underpins a measurement platform for sample-dependent acquisition of MS/MS data for whole proteins, a critical step to realize proteomics with 100% sequence coverage in a higher throughput setting.

Archaeal Proteins↗

Functional genomics of hepatocellular carcinoma.

The majority of DNA-microarray based gene expression profiling studies on human hepatocellular carcinoma (HCC) has focused on identifying genes associated with clinicopathological features of HCC patients. Although notable success has been achieved, this approach still faces significant challenges due to the heterogeneous nature of HCC (and other cancers) as well as the many confounding factors embedded in gene expression profile data. However, these limitations are being overcome by improved bioinformatics and sophisticated analyses. Also, application of cross comparison of multiple gene expression data sets from human tumors and animal models are facilitating the identification of critical regulatory modules in the expression profiles. The success of this new experimental approach, comparative functional genomics, suggests that integration of independent data sets will enhance our ability to identify key regulatory elements in tumor development. Furthermore, integrating gene expression profiles with data from DNA sequence information in promoters, array-based CGH, and expression of non-coding genes (i.e., microRNAs) will further increase the reliability and significance of the biological and clinical inferences drawn from the data. The pace of current progress in the cancer profiling field, combined with the advances in high-throughput technologies in genomics and proteomics, as well as in bioinformatics, promises to yield unprecedented biological insights from the integrative (or systems) analysis of the combined cancer genomics database. The predicted beneficial impact of this "new integrative biology" on diagnosis, treatment and prevention of liver cancer and indeed cancer in general is enormous.

Animals↗

Protein production in Escherichia coli for structural studies by X-ray crystallography.

The arrival of genomic sequences to the database has provided a seemingly unlimited supply of targets for protein structure determination and the possibility of solving the structure of an entire proteome. Based on our experience with the proteomes of Pyrobaculum aerophilum and Mycobacterium tuberculosis, we have developed a simple strategy for the production of proteins for structural studies by X-ray crystallography. Our scheme demonstrates a strong protein target commitment and includes the expression of genes from these organisms in Escherichia coli. These proteins are expressed with affinity tags and purified for characterization and crystallization. We have identified protein solubility and crystallization as the two major bottlenecks in the process toward the determination of protein structures by X-ray diffraction. Strategies to overcome these bottlenecks are discussed.

Cloning, Molecular↗

The protein-protein interaction map of Helicobacter pylori.

With the availability of complete DNA sequences for many prokaryotic and eukaryotic genomes, and soon for the human genome itself, it is important to develop reliable proteome-wide approaches for a better understanding of protein function. As elementary constituents of cellular protein complexes and pathways, protein-protein interactions are key determinants of protein function. Here we have built a large-scale protein-protein interaction map of the human gastric pathogen Helicobacter pylori. We have used a high-throughput strategy of the yeast two-hybrid assay to screen 261 H. pylori proteins against a highly complex library of genome-encoded polypeptides. Over 1,200 interactions were identified between H. pylori proteins, connecting 46.6% of the proteome. The determination of a reliability score for every single protein-protein interaction and the identification of the actual interacting domains permitted the assignment of unannotated proteins to biological pathways.

Amino Acid Sequence↗

Therapeutic targets: progress of their exploration and investigation of their characteristics.

Modern drug discovery is primarily based on the search and subsequent testing of drug candidates acting on a preselected therapeutic target. Progress in genomics, protein structure, proteomics, and disease mechanisms has led to a growing interest in and effort for finding new targets and more effective exploration of existing targets. The number of reported targets of marketed and investigational drugs has significantly increased in the past 8 years. There are 1535 targets collected in the therapeutic target database compared with approximately 500 targets reported in a 1996 review. Knowledge of these targets is helpful for molecular dissection of the mechanism of action of drugs and for predicting features that guide new drug design and the search for new targets. This article summarizes the progress of target exploration and investigates the characteristics of the currently explored targets to analyze their sequence, structure, family representation, pathway association, tissue distribution, and genome location features for finding clues useful for searching for new targets. Possible "rules" to guide the search for druggable proteins and the feasibility of using a statistical learning method for predicting druggable proteins directly from their sequences are discussed.

Adrenergic beta-Antagonists↗

Proteomics in Drosophila melanogaster.

To be able to understand cellular mechanisms, we require fully integrated data sets combining information about gene expression, protein expression, post-translational modification states, sub-cellular location and complex formation. Proteomics is a very powerful technique that can be applied to interrogate changes at the protein level. Studying this effectively requires specialised facilities within research institutes. Here, we describe the setting up and operation of such a facility, providing a resource for the Arabidopsis and Drosophila research communities.

Animals↗

Identification of Glioblastoma Cell Surface Proteins and Assessment of Their Expression Across Patient-Derived Stem-Like Cell Cultures.

Glioblastoma (GBM) is the most common primary brain cancer in adults and remains fatal, with a median survival of a few months. There is an urgent need to develop novel therapeutic strategies against this aggressive malignancy. Modern cancer research increasingly focuses on personalized therapies tailored toward unique molecular features of each tumor or patient. In this context, cell surface proteins (CSPs) represent an attractive class of therapeutic targets due to their accessibility and central roles in physiological and pathological processes, making them among the most targeted proteins in current drug development. In this study, promising CSPs were identified through an untargeted proteomics approach using high-resolution mass spectrometry on patient-derived GBM stem-like cell (GSC) cultures, complemented by RNA-seq data and computational database analyses. From this primary discovery, five CSPs, namely PTK7, PTPRZ1, OSMR, CSPG4, and IGDCC4, were selected for detailed investigation. A targeted UHPLC-multiple reaction monitoring (MRM) method was developed and optimized to assess their expression and evaluate their abundance variations across different GSC cultures and cell passage levels. Beyond confirming these CSPs as potential therapeutic targets in GBM, our study demonstrates the value of three-dimensional GSC cultures as robust models for biomarker research and target assessment.

Humans↗

A proteomics study of in vitro cyst germination and appressoria formation in Phytophthora infestans.

A proteomics study using two-dimensional gel electrophoresis (2-DE) and mass spectrometry was performed on Phytophthora infestans. Proteins from cysts, germinated cysts and appressoria grown in vitro were isolated and separated by 2-DE. Statistical quantitative analysis of the protein spots from five independent experiments of each developmental stage revealed significant up-regulation of ten spots on gels from germinated cysts compared to cysts. Five spots were significantly up-regulated on gels from appressoria compared to germinated cysts and one of these up-regulated spots was not detectable on gels from cysts. In addition, one spot was significantly down-regulated and another spot not detectable on the gels from appressoria. The corresponding proteins to 13 of these spots were identified with high confidence using tandem mass spectrometry and database searches. The functions of the proteins that were up-regulated in germinated cysts and appressoria can be grouped into the following categories: protein synthesis (e.g. a DEAD box RNA helicase), amino acid metabolism, energy metabolism and reactive oxygen species scavenging. The spot not detected in appressoria was identified as the P. infestans crinkling- and necrosis-inducing protein CRN2. The identified proteins are most likely involved in the establishment of the infection of the host plant.

Algal Proteins↗

Genomic data for alternate production strategies. I. Identification of major contaminating species for Cobalt(+2) immobilized metal affinity chromatography.

Recent advances in technology have allowed for the identification of complex protein mixtures in a rapid fashion. This report highlights the use of 2D gel electrophoresis, mass spectrometry, and database analysis to determine contaminating species of the Escherichia coli genome that are present during immobilized metal affinity chromatography (IMAC), highlighting Co(2+) as the affinity ligand. Four proteins (triosephosphate isomerase, alpha galactosidase, Hsp90, and glucosamine 6-phosphate synthase) constitute the majority of E. coli proteins that bind and potentially may coelute during chromatography. Results are discussed within the context of changes that when implemented could lead to an increase in IMAC efficiency, not by altering column conditions, but rather by changing the nature of the nuisance proteins that principally reduce column capacity and extend processing times. Such a study illustrates the use of proteome data to aid in bioprocess design.

Amino Acid Sequence↗

Survey of current protein family databases and their application in comparative, structural and functional genomics.

The last two decades have witnessed significant expansions in the databases storing information on the sequences and structures of proteins. This has led to the creation of many excellent protein family resources, which classify proteins according to their evolutionary relationship. These have allowed extensive insights into evolution and particularly how protein function mutates and evolves over time. Such analyses have greatly assisted the inheritance of functional annotations between experimentally characterised and uncharacterised genes. Moreover, the development of bioinformatics tools acts as a companion to the new technologies emerging in biology, such as transcriptomics and proteomics. The latter enable researchers to analyse gene expression profiles and interactions on a genome-wide scale, generating vast datasets of proteins, many of which include experimentally uncharacterised proteins. Protein family/function databases can be used to help interpret this data and allow us to benefit more fully from these technologies. This review aims to summarise the most popular sequence- and structure-based protein family databases. We also cover their application to comparative genomics and the functional annotation of the genomes.

Biological Evolution↗

Detection of hypothetical proteins in human fetal perireticular nucleus.

There is a legion of hypothetical proteins (HP) in prokaryotic and eukaryotic proteomes and the aim of this study was to describe HP in the perireticular nucleus (PN), a key structure in human brain development. Tissue from four PNs was homogenized and extracted proteins were run on two-dimensional gel electrophoresis followed by in-gel digestion and mass spectrometrical identification of proteins. Several databases were used for obtaining bioinformatic information and searching for functional and structural domains. Five spots represented HP: KIAA0423 protein (Q9Y4F4), hypothetical protein KIAA0153 (Q14166), hypothetical protein DKFZp564A2416 (Q9NTW4), hypothetical protein DKFZp564H1122 (Q9H0W9), and hypothetical protein DKFZp564D1378 (Q9H0R4). These structures were predicted to serve in cell cycle, DNA-condensation, neurogenesis, or apoptosis. The existence of formerly HP proteins in the PN of human fetal brain is shown, thus extending knowledge of the brain proteome and proposing the method used as a suitable analytical tool for searching HP.

Apoptosis↗

A systematic approach to modeling, capturing, and disseminating proteomics experimental data.

Both the generation and the analysis of proteome data are becoming increasingly widespread, and the field of proteomics is moving incrementally toward high-throughput approaches. Techniques are also increasing in complexity as the relevant technologies evolve. A standard representation of both the methods used and the data generated in proteomics experiments, analogous to that of the MIAME (minimum information about a microarray experiment) guidelines for transcriptomics, and the associated MAGE (microarray gene expression) object model and XML (extensible markup language) implementation, has yet to emerge. This hinders the handling, exchange, and dissemination of proteomics data. Here, we present a UML (unified modeling language) approach to proteomics experimental data, describe XML and SQL (structured query language) implementations of that model, and discuss capture, storage, and dissemination strategies. These make explicit what data might be most usefully captured about proteomics experiments and provide complementary routes toward the implementation of a proteome repository.

Database Management Systems↗

Cancer proteomics: from identification of novel markers to creation of artifical learning models for tumor classification.

Studies of global protein expression in human tumors have led to the identification of various polypeptide markers, potentially useful as diagnostic tools. Many changes in gene expression recorded between benign and malignant human tumors are due to post-translational modifications, not detected by analyses of RNA. Proteome analyses have also yielded information about tumor heterogeneity and the degree of relatedness between primary tumors and their metastases. Results from our own studies have shown a similar pattern of changes in protein expression in different epithelial tumors, such as decreases in tropomyosin and cytokeratin expression and increases in proliferating cell nuclear antigen (PCNA) and heat shock protein expression. Such information has been used to create artificial learning models for tumor classification. The artificial learning approach has potential to improve tumor diagnosis and cancer treatment prediction.

Biomarkers, Tumor↗

IMGT, the international ImMunoGeneTics database: a high-quality information system for comparative immunogenetics and immunology.

IMGT, the international ImMunoGeneTics database (http://imgt.cines.fr), is a high quality integrated information system specializing in Immunoglobulins (IG), T cell Receptors (TR) and Major Histocompatibility Complex (MHC) of human and other vertebrates, created in 1989, by LIGM, at the Université Montpellier II, CNRS, Montpellier, France. IMGT provides a common access to standardized data, which include nucleotide and protein sequences, oligonucleotide primers, gene maps, genetic polymorphisms, specificities, 2D and 3D structures. IMGT includes several databases (IMGT/LIGM-DB, IMGT/HLA-DB, IMGT/3Dstructure-DB), Web resources ('IMGT Marie-Paule page') which comprise IMGT Scientific Chart, IMGT Repertoire, IMGT Bloc-notes, IMGT Education, IMGT Aide-mémoire and IMGT Index, and interactive tools (IMGT/V-QUEST, IMGT/JunctionAnalysis). These expertly annotated data on the genome, proteome, genetics and structure of the IG, TR and MHC are of high value for comparative genome evolution studies of the adaptative immune response.

Animals↗

A cell-free protein synthesis system for high-throughput proteomics.

We report a cell-free system for the high-throughput synthesis and screening of gene products. The system, based on the eukaryotic translation apparatus of wheat seeds, has significant advantages over other commonly used cell-free expression systems. To maximize the yield and throughput of the system, we optimized the mRNA UTRs, designed an expression vector for large-scale protein production, and developed a new strategy to construct PCR-generated DNAs for high-throughput production of many proteins in parallel. The resulting system achieves high-yield expression and can maintain productive translation for 14 days. Additionally, in the integration of a PCR-directed system for template creation, at least 50 genes can be translated in parallel, yielding between 0.1 and 2.3 mg of protein by one person within 2 days. Assessment of correct protein folding by the products of this high-throughput protein-expression system were performed by enzymatic assays of kinases and by NMR spectroscopic analysis. The cell-free system, reported here, bypasses many of the time-consuming cloning steps of conventional expression systems and lends itself to a robotic automation for the high-throughput expression of proteins.

Cell-Free System↗

GAPP: a fully automated software for the confident identification of human peptides from tandem mass spectra.

This paper introduces the genome annotating proteomic pipeline (GAPP), a totally automated publicly available software pipeline for the identification of peptides and proteins from human proteomic tandem mass spectrometry data. The pipeline takes as its input a series of MS/MS peak lists from a given experimental sample and produces a series of database entries corresponding to the peptides observed within the sample, along with related confidence scores. The pipeline is capable of finding any peptides expected, including those that cross intron-exon boundaries, and those due to single nucleotide polymorphisms (SNPs), alternate splicing, and post-translational modifications (PTMs). GAPP can therefore be used to re-annotate genomes, and this is supported through the inclusion of a Distributed Annotation System (DAS) server, which allows the peptides identified by the pipeline to be displayed in their genomic context within the Ensembl genome browser. GAPP is freely available via the web, at www. gapp.info.

Alternative Splicing↗

Global protein expression pattern of Bradyrhizobium japonicum bacteroids: a prelude to functional proteomics.

As a prelude to using functional proteomics towards understanding the process of symbiotic nitrogen fixation between the legume soybean and the soil bacteria Bradyrhizobium japonicum, we examined the total protein expression pattern of the nodule bacteria, often referred to as bacteroids. A partial proteome map was constructed by separating the total bacteroid proteins using high-resolution 2-DE. Of the several hundred protein spots analyzed using PMF, 180 spots were tentatively identified by searching the available database for B. japonicum, (http://www.kazusa.or.jp/index.html). The data showed that the bacteroid expressed a dominant and elaborate protein network for nitrogen and carbon metabolism, which is closely dependent on the plant supplied metabolites, and seems aptly supported by a selective group of bacteroid transporter proteins. However, they seem to lack a defined fatty acid and nucleic acid metabolism. Interestingly, the proteins related to protein synthesis, scaffolding and degradation were among the most predominant spots of the bacteroid proteome. In addition, several proteins, which showed fairly good expression, were identified to be involved with cellular detoxification, stress regulation and signaling communication components. This preliminary proteomic data matches very well with several biochemical and genetic reports, and clearly shows the inter-connection between several metabolic pathways that meet the needs of the bacteroid. It is expected that in the future this will allow us to develop testable hypotheses about the roles of several of these proteins in context to the metabolic pathway connections and metabolite fluxes.

Amino Acids↗

Domain combinations in archaeal, eubacterial and eukaryotic proteomes.

There is a limited repertoire of domain families that are duplicated and combined in different ways to form the set of proteins in a genome. Proteins are gene products, and at the level of genes, duplication, recombination, fusion and fission are the processes that produce new genes. We attempt to gain an overview of these processes by studying the evolutionary units in proteins, domains, in the protein sequences of 40 genomes. The domain and superfamily definitions in the Structural Classification of Proteins Database are used, so that we can view all pairs of adjacent domains in genome sequences in terms of their superfamily combinations. We find 783 out of the 859 superfamilies in SCOP in these genomes, and the 783 families occur in 1307 pairwise combinations. Most families are observed in combination with one or two other families, while a few families are very versatile in their combinatorial behaviour; 209 families do not make combinations with other families. This type of pattern can be described as a scale-free network. We also study the N to C-terminal orientation of domain pairs and domain repeats. The phylogenetic distribution of domain combinations is surveyed, to establish the extent of common and kingdom-specific combinations. Of the kingdom-specific combinations, significantly more combinations consist of families present in all three kingdoms than of families present in one or two kingdoms. Hence, we are led to conclude that recombination between common families, as compared to the invention of new families and recombination among these, has also been a major contribution to the evolution of kingdom-specific and species-specific functions in organisms in all three kingdoms. Finally, we compare the set of the domain combinations in the genomes to those in the RCSB Protein Data Bank, and discuss the implications for structural genomics.

Animals↗