Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

InterDom: a database of putative interacting protein domains for validating predicted protein interactions and complexes.

Advances in proteomics technology have enabled new proteins to be discovered at an unprecedented speed, and high throughput experimental methods have been developed to detect protein interactions and complexes en masse. Such bottom-up, data-driven approach has resulted in data that may be uninformative or potentially errorful, requiring further validation and annotation. The InterDom database focuses on providing supporting evidence for the detected protein interactions based on putative protein domain interactions. Using an integrative approach, InterDom derives potential domain interactions by combining data from multiple sources, ranging from domain fusions, protein interactions and complexes, to scientific literature. The InterDom database is available at http://InterDom.lit.org.sg.

Databases, Protein↗

PHOG-BLAST--a new generation tool for fast similarity search of protein families.

BACKGROUND: The need to compare protein profiles frequently arises in various protein research areas: comparison of protein families, domain searches, resolution of orthology and paralogy. The existing fast algorithms can only compare a protein sequence with a protein sequence and a profile with a sequence. Algorithms to compare profiles use dynamic programming and complex scoring functions. RESULTS: We developed a new algorithm called PHOG-BLAST for fast similarity search of profiles. This algorithm uses profile discretization to convert a profile to a finite alphabet and utilizes hashing for fast search. To determine the optimal alphabet, we analyzed columns in reliable multiple alignments and obtained column clusters in the 20-dimensional profile space by applying a special clustering procedure. We show that the clustering procedure works best if its parameters are chosen so that 20 profile clusters are obtained which can be interpreted as ancestral amino acid residues. With these clusters, only less than 2% of columns in multiple alignments are out of clusters. We tested the performance of PHOG-BLAST vs. PSI-BLAST on three well-known databases of multiple alignments: COG, PFAM and BALIBASE. On the COG database both algorithms showed the same performance, on PFAM and BALIBASE PHOG-BLAST was much superior to PSI-BLAST. PHOG-BLAST required 10-20 times less computer memory and computation time than PSI-BLAST. CONCLUSION: Since PHOG-BLAST can compare multiple alignments of protein families, it can be used in different areas of comparative proteomics and protein evolution. For example, PHOG-BLAST helped to build the PHOG database of phylogenetic orthologous groups. An essential step in building this database was comparing protein complements of different species and orthologous groups of different taxons on a personal computer in reasonable time. When it is applied to detect weak similarity between protein families, PHOG-BLAST is less precise than rigorous profile-profile comparison method, though it runs much faster and can be used as a hit pre-selecting tool.

Algorithms↗

Scaling and normalization effects in NMR spectroscopic metabonomic data sets.

Considerable confusion appears to exist in the metabonomics literature as to the real need for, and the role of, preprocessing the acquired spectroscopic data. A number of studies have presented various data manipulation approaches, some suggesting an optimum method. In metabonomics, data are usually presented as a table where each row relates to a given sample or analytical experiment and each column corresponds to a single measurement in that experiment, typically individual spectral peak intensities or metabolite concentrations. Here we suggest definitions for and discuss the operations usually termed normalization (a table row operation) and scaling (a table column operation) and demonstrate their need in 1H NMR spectroscopic data sets derived from urine. The problems associated with "binned" data (i.e., values integrated over discrete spectral regions) are also discussed, and the particular biological context problems of analytical data on urine are highlighted. It is shown that care must be exercised in calculation of correlation coefficients for data sets where normalization to a constant sum is used. Analogous considerations will be needed for other biofluids, other analytical approaches (e.g., HPLC-MS), and indeed for other "omics" techniques (i.e., transcriptomics or proteomics) and for integrated studies with "fused" data sets. It is concluded that data preprocessing is context dependent and there can be no single method for general use.

Algorithms↗

Protein profiling of the human epidermis from the elderly reveals up-regulation of a signature of interferon-gamma-induced polypeptides that includes manganese-superoxide dismutase and the p85beta subunit of phosphatidylinositol 3-kinase.

Aging of the human skin is a complex process that consists of chronological and extrinsic aging, the latter caused mainly by exposure to ultraviolet radiation (photoaging). Here we present studies in which we have used proteomic profiling technologies and two-dimensional (2D) PAGE database resources to identify proteins whose expression is deregulated in the epidermis of the elderly. Fresh punch biopsies from the forearm of 20 pairs of young and old donors (21-30 and 75-92 years old, respectively) were dissected to yield an epidermal fraction that consisted mainly of differentiated cells. One- to two-mm3 epidermal pieces were labeled with [35S]methionine for 18 h, lysed, and subjected to 2D PAGE (isoelectric focusing and non-equilibrium pH gradient electrophoresis) and phosphorimage autoradiography. Proteins were identified by matching the gels with the master 2D gel image of human keratinocytes (proteomics.cancer.dk). In selected cases 2D PAGE immunoblotting and/or mass spectrometry confirmed the identity. Quantitative analysis of 172 well focused and abundant polypeptides showed that the level of most proteins (148) remains unaffected by the aging process. Twenty-two proteins were consistently deregulated by a factor of 1.5 or more across the 20 sample pairs. Among these we identified a group of six polypeptides (Mx-A, manganese-superoxide dismutase, tryptophanyl-tRNA synthetase, the p85beta subunit of phosphatidylinositol 3-kinase, and proteasomal proteins PA28-alpha and SSP 0107) that is induced by interferon-gamma in primary human keratinocytes and that represents a specific protein signature for the effect of this cytokine. Changes in the expression of the eukaryotic initiation factor 5A, NM23 H2, cyclophilin A, HSP60, annexin I, and plasminogen activator inhibitor 2 were also observed. Two proteins exhibited irregular behavior from individual to individual. Besides arguing for a role of interferon-gamma in the aging process, the biological activities associated with the deregulated proteins support the contention that aging is linked with increased oxidative stress that could lead to apoptosis in vivo.

Adult↗

Improved sensitivity proteomics by postharvest alkylation and radioactive labelling of proteins.

We describe approaches to improve the detection of proteins by postharvest alkylation and subsequent radioactive labeling with either [3H]iodoacetamide or 125I. Database protein sequence analysis suggested that cysteine is not suitable for detection of the entire proteome, but that cysteine alkylating reagents can increase the number of proteins able to be detected by iodination chemistry. Proteins were alkylated with beta-(4-hydroxyphenyl)ethyl iodoacetamide, or with 1,5-l-AEDANS (the Hudson Weber reagent). Subsequent iodination using the Iodo-Gen system was found to be most efficient. The enhanced sensitivity obtainable by using these approaches is expected to be sufficient for visualization of the lowest copy number proteins from human cells, such as from clinical samples. However, we argue that significantly improved methods of protein separation will be necessary to resolve the large number of proteins expected to be detectable with this sensitivity.

Acetamides↗

Annotation of novel neuropeptide precursors in the migratory locust based on transcript screening of a public EST database and mass spectrometry.

BACKGROUND: For holometabolous insects there has been an explosion of proteomic and peptidomic information thanks to large genome sequencing projects. Heterometabolous insects, although comprising many important species, have been far less studied. The migratory locust Locusta migratoria, a heterometabolous insect, is one of the most infamous agricultural pests. They undergo a well-known and profound phase transition from the relatively harmless solitary form to a ferocious gregarious form. The underlying regulatory mechanisms of this phase transition are not fully understood, but it is undoubtedly that neuropeptides are involved. However, neuropeptide research in locusts is hampered by the absence of genomic information. RESULTS: Recently, EST (Expressed Sequence Tag) databases from Locusta migratoria were constructed. Using bioinformatical tools, we searched these EST databases specifically for neuropeptide precursors. Based on known locust neuropeptide sequences, we confirmed the sequence of several previously identified neuropeptide precursors (i.e. pacifastin-related peptides), which consolidated our method. In addition, we found two novel neuroparsin precursors and annotated the hitherto unknown tachykinin precursor. Besides one of the known tachykinin peptides, this EST contained an additional tachykinin-like sequence. Using neuropeptide precursors from Drosophila melanogaster as a query, we succeeded in annotating the Locusta neuropeptide F, allatostatin-C and ecdysis-triggering hormone precursor, which until now had not been identified in locusts or in any other heterometabolous insect. For the tachykinin precursor, the ecdysis-triggering hormone precursor and the allatostatin-C precursor, translation of the predicted neuropeptides in neural tissues was confirmed with mass spectrometric techniques. CONCLUSION: In this study we describe the annotation of 6 novel neuropeptide precursors and the neuropeptides they encode from the migratory locust, Locusta migratoria. By combining the manual annotation of neuropeptides with experimental evidence provided by mass spectrometry, we demonstrate that the genes are not only transcribed but also translated into precursor proteins. In addition, we show which neuropeptides are cleaved from these precursor proteins and how they are post-translationally modified.

Amino Acid Sequence↗

POGs/PlantRBP: a resource for comparative genomics in plants.

POGs/PlantRBP (http://plantrbp.uoregon.edu/) is a relational database that integrates data from rice, Arabidopsis, and maize by placing the complete Arabidopsis and rice proteomes and available maize sequences into 'putative orthologous groups' (POGs). Annotation efforts will focus on predicted RNA binding proteins (RBPs): i.e. those with known RNA binding domains or otherwise implicated in RNA function. POGs form the heart of the database, and were assigned using a mutual-best-hit-strategy after performing BLAST comparisons of the predicted Arabidopsis and rice proteomes. Each POG entry includes orthologs in Arabidopsis and rice, annotated with domain organization, gene models, phylogenetic trees, and multiple intracellular targeting predictions. A graphical display maps maize sequences on to their most similar rice gene model. The database can be queried using any combination of gene name, accession, domain, and predicted intracellular location, or using BLAST. Useful features of the database include the ability to search for proteins with both a specified domain content and intracellular location, the concurrent display of mutual best hits and phylogenetic trees which facilitates evaluation of POG assignments, the association of maize sequences with POGs, and the display of targeting predictions and domain organization for all POG members, which reveals consistency, or lack thereof, of those predictions.

Amino Acid Sequence↗

Proteome analysis of bacterial pathogens.

Combining two-dimensional electrophoresis with mass spectrometry resulted in a powerful technology ideally suited to recognize and identify proteins of pathogenic microorganisms. This classical proteome analysis is now complemented by capillary chromatography/mass spectrometry combinations, miniaturization by chip technology and protein interaction investigations. Comparative proteomics is used to reveal vaccine candidates and pathogenicity factors. Immunoproteomics identifies specific and nonspecific antigens. For the management of the huge data amounts, bioinformatics is a valuable instrument for the construction of complex protein databases.

Animals↗

Quantitative proteome analysis: methods and applications.

With the completion of a rapidly increasing number of complete genomic sequences, much attention is currently focused on how the information contained in sequence databases might be interpreted in terms of the structure, function, and control of biological systems. Quantitative proteome analysis, the global analysis of protein expression, has been proposed as a method to study steady-state gene expression and perturbation-induced changes. Here, we discuss the rationale for quantitative proteome analysis, highlight the limitations in the current standard technology, and introduce a new experimental approach to quantitative proteome analysis.

Amino Acid Sequence↗

Multivariate approaches in plant science.

The objective of proteomics is to get an overview of the proteins expressed at a given point in time in a given tissue and to identify the connection to the biochemical status of that tissue. Therefore sample throughput and analysis time are important issues in proteomics. The concept of proteomics is to encircle the identity of proteins of interest. However, the overall relation between proteins must also be explained. Classical proteomics consist of separation and characterization, based on two-dimensional electrophoresis, trypsin digestion, mass spectrometry and database searching. Characterization includes labor intensive work in order to manage, handle and analyze data. The field of classical proteomics should therefore be extended to also include handling of large datasets in an objective way. The separation obtained by two-dimensional electrophoresis and mass spectrometry gives rise to huge amount of data. We present a multivariate approach to the handling of data in proteomics with the advantage that protein patterns can be spotted at an early stage and consequently the proteins selected for sequencing can be selected intelligently. These methods can also be applied to other data generating protein analysis methods like mass spectrometry and near infrared spectroscopy and examples of application to these techniques are also presented. Multivariate data analysis can unravel complicated data structures and may thereby relieve the characterization phase in classical proteomics. Traditionally statistical methods are not suitable for analysis of the huge amounts of data, where the number of variables exceed the number of objects. Multivariate data analysis, on the other hand, may uncover the hidden structures present in these data. This study takes its starting point in the field of classical proteomics and shows how multivariate data analysis can lead to faster ways of finding interesting proteins. Multivariate analysis has shown interesting results as a supplement to classical proteomics and added a new dimension to the field of proteomics.

Algorithms↗

Automated annotation of microbial proteomes in SWISS-PROT.

Large-scale sequencing of prokaryotic genomes demands the automation of certain annotation tasks currently manually performed in the production of the SWISS-PROT protein knowledgebase. The HAMAP project, or 'High-quality Automated and Manual Annotation of microbial Proteomes', aims to integrate manual and automatic annotation methods in order to enhance the speed of the curation process while preserving the quality of the database annotation. Automatic annotation is only applied to entries that belong to manually defined orthologous families and to entries with no identifiable similarities (ORFans). Many checks are enforced in order to prevent the propagation of wrong annotation and to spot problematic cases, which are channelled to manual curation. The results of this annotation are integrated in SWISS-PROT, and a website is provided at http://www.expasy.org/sprot/hamap/.

Amino Acid Sequence↗

Proteomic analysis by multidimensional protein identification technology.

Multidimensional chromatography coupled to mass spectrometry is an emerging technique for the analysis of complex protein mixtures. One approach in this general category, multidimensional protein identification technology (MudPIT), couples biphasic or triphasic microcapillary columns to high-performance liquid chromatography, tandem mass spectrometry, and database searching. The integration of each of these components is critical to the implementation of MudPIT in a laboratory. MudPIT can be used for the analysis of complex peptide mixtures generated from biofluids, tissues, cells, organelles, or protein complexes. The information described in this chapter will provide researchers with details for sample preparation, column assembly, and chromatography parameters for complex peptide mixture analysis.

Animals↗

Correspondence regarding Schwend and Gustafsson, "False positives in MALDI-TOF detection of ERbeta in mitochondria".

Recently, Schwend and Gustafsson tried to use the MALDI-TOF methods to confirm one of the results reported by Yang et al., which provided definitive evidences to demonstrate the localization of estrogen receptor beta (ERbeta) in the mitochondria of multiple cell types, using immunocytochemistry, immunoblot, and proteomic approaches. Analysis of the data with the MASCOT database algorithm provided no evidence for the presence of ERbeta in the mouse live mitochondria, in which very low ERbeta expression has been detected in their own report. On the other hand, our MALDI-TOF analysis using human heart mitochondrial protein has identified 7 and 8 sequences that could be potentially from ERbeta and ERbeta3, respectively, but not from ATP synthases. Further, none of the sequences identified by us as those of ERbeta and ERbeta3 shares m/z targeted by Schwend and Gustafsson in their measurements. Therefore, the claim by Gustafsson's laboratory about false positives in MALDI-TOF detection of ERbeta in mitochondria has no relevance to our report.

Animals↗

Profiling excretory/secretory proteins of Trichinella spiralis muscle larvae by two-dimensional gel electrophoresis and mass spectrometry.

Infection of mammalian skeletal muscle with the intracellular parasite Trichinella spiralis results in profound alterations in the host cell and a realignment of host cell gene expression. The role of parasite excretory/secretory (E/S) products in mediating these effects is unknown, largely due to the difficulty in identifying and assigning function to individual proteins. In this study, we have used two-dimensional electrophoresis to analyse the profile of muscle larva excreted/secreted proteins and have coupled this to protein identification using MALDI-TOF mass spectrometry. Interpretation of the peptide mass fingerprint data has relied primarily on the interrogation of a custom-made Trichinella EST database and the NemaGene cluster database for T. spiralis. Our results suggest that this proteomic approach is a useful tool to study protein expression in Trichinella spp. and will contribute to the identification of excreted/secreted proteins.

Animals↗

Proteomics--post-genomic cartography to understand gene function.

The completion of the genomic sequences of numerous organisms from human and mouse to Caenorhabditis elegans and many microorganisms, and the definition of their genes provides a database to interpret cellular protein-expression patterns and relate them to protein function. Proteomics technologies that are dependent on mass spectrometry and involve two-dimensional gel electrophoresis are providing the main window into the world of differential protein-expression analysis. In this article, the limitations and expectations of this research field are examined and the future of the analytical needs of proteomics is explored.

Animals↗

Proteome analysis based on motif statistics.

MOTIVATION: Even for the amino acid motifs collected in the Prosite database there may be chance occurences as opposed to those occurences where the motif is involved in fold or function of a protein. With recent mathematical advances in assessing the significance of observing such a motif a particular number of times, we can now study the over- or under-representation of particular motifs in a complete genome and attempt to make functional deductions. RESULTS: We demonstrate that statistical over- or under-representation of motifs in complete proteomes may be an indicator of whether, in that organism, we are looking at chance occurrences of the motif or whether the occurrences are sufficiently numerous to suggest a systematic, and thus functionally important occurrence. This has important implications on databank annotations. AVAILABILITY: The complete dataset comprising the plotted statistics of 266 Prosite motifs on 42 proteomes is available at http://algo.inria.fr/nicodeme/proteomes/proteocomp.html. The software used to compute this data has been described by Nicodème (2000, 2001). They are available either by web access as mentioned in these articles or by direct request from Pierre Nicodème.

Amino Acid Motifs↗

Capillary array reversed-phase liquid chromatography-based multidimensional separation system coupled with MALDI-TOF-TOF-MS detection for high-throughput proteome analysis.

A high-throughput on-line capillary array-based two-dimensional liquid chromatography (2D-LC) system coupled with MALDI-TOF-TOF-MS proteomics analyzer for comprehensive proteomic analyses has been developed, in which one capillary strong-cation exchange (SCX) chromatographic column was used as the first separation dimension and 18 parallel capillary reversed-phase liquid chromatographic (RPLC) columns were integrated as the second separation dimension. Peptides bound to the SCX phase were "stepped" off using multiple salt pulses followed by sequentially loading of each subset of peptides onto the corresponding precolumns. After salt fractionation, by directing identically split solvent-gradient flows into 18 channels, peptide fractions were concurrently back-flushed from the precolumns and separated simultaneously with 18 capillary RP columns. LC effluents were directly deposited onto the MALDI target plates through an array of capillary tips at a 15-s interval, and then alpha-cyano-4-hydroxycinnamic acid (CHCA) matrix solution was added to each sample spot for subsequent MALDI experiments. This new system allows an 18-fold increase in throughput compared with serial-based 2D-LC system. The high efficiency of the overall system was demonstrated by the analysis of a tryptic digest of proteins extracted from normal human liver tissue. A total of 462 proteins was identified, which proved the system's promising potential for high-throughput analysis and application in proteomics.

Capillary Action↗