Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Proteomic identification of a novel protein regulated in CA1 and CA3 hippocampal regions during intermittent hypoxia.

The CA1 and CA3 regions of the hippocampus markedly differ in their susceptibility to hypoxia in general, and more particularly to the intermittent hypoxia (IH) that characterizes sleep apnea. We used proteomic analysis to build a database of proteins expressed in normoxic CA1 and CA3. The current hippocampus protein database identifies 106 proteins. A hypothetical protein with accession number AK006737 (gimid R:12839969) was strongly upregulated in the CA1, but not CA3 hippocampal region. Bioinformatic analysis revealed that the unknown protein contained a high stringency protein kinase e binding site. Domain analysis demonstrated the presence of a conserved sequence indicative of macrophage scavenger receptors. Using proteomic analysis we have previously demonstrated that acute (6 h) IH-mediated CA1 injury results from complex interactions between pathways involving increased metabolism, induction of stress-induced proteins and apoptosis, and ultimately disruption of structural proteins and cell integrity. The current findings identify a hypothetical protein that may play a key role in the response of CA1 to IH. These findings provide initial insights into mechanisms underlying differences in susceptibility to hypoxia in neural tissue and demonstrate how proteomic analysis can be used to generate new hypotheses, which define neuronal adaptation to IH.

Animals↗

Bioinformatics in proteomics.

Proteomics technologies are under continuous improvements and new technologies are introduced. Nowadays high throughput acquisition of proteome data is possible. The young and rapidly emerging field of bioinformatics in proteomics is introducing new algorithms to handle large and heterogeneous data sets and to improve the knowledge discovery process. For example new algorithms for image analysis of two dimensional gels have been developed within the last five years. Within mass spectrometry data analysis algorithms for peptide mass fingerprinting (PMF) and peptide fragmentation fingerprinting (PFF) have been developed. Local proteomics bioinformatics platforms emerge as data management systems and knowledge bases in Proteomics. We review recent developments in bioinformatics for proteomics with emphasis on expression proteomics.

Animals↗

Molecular-level description of proteins from saccharomyces cerevisiae using quadrupole FT hybrid mass spectrometry for top down proteomics.

For improved detection of diverse posttranslational modifications (PTMs), direct fragmentation of protein ions by top down mass spectrometry holds promise but has yet to be achieved on a large scale. Using lysate from Saccharomyces cerevisiae, 117 gene products were identified with 100% sequence coverage revealing 26 acetylations, 1 N-terminal dimethylation, 1 phosphorylation, 18 duplicate genes, and 44 proteolytic fragments. The platform for this study combined continuous-elution gel electrophoresis, reversed-phase liquid chromatography, automated nanospray coupled with a quadrupole-FT hybrid mass spectrometer, and a new search engine for querying a custom database. The proteins identified required no manual validation, ranged from 5 to 39 kDa, had codon biases from 0.93 to 0.083, and were primarily associated with glycolysis and protein synthesis. Illustrations of gene-specific identifications, PTM detection and subsequent PTM localization (using either electron capture dissociation or known PTM data stored in a database) show how larger scale proteome projects incorporating top down may proceed in the future using commercial Q-FT instruments.

Acetylation↗

Bioinformatics for study of autoimmunity.

Recent years have witnessed an explosive growth in available biological data pertaining to autoimmunity research. This includes a tremendous quantity of sequence data (biological structures, genetic and physical maps, pathways, etc.) generated by genome and proteome projects plus extensive clinical and epidemiological data. Autoimmunity research stands to greatly benefit from this data so long as appropriate strategies are available to enable full access to and utilization of this data. The quantity and complexity of this biological data necessitates use of advanced bioinformatics strategies for its efficient retrieval, analysis and interpretation. Major progress has been made in development of specialized tools for storage, analysis and modeling of immunological data, and this has led to development of a whole new field know as immunoinformatics. With advances in novel high-throughput immunology technologies immunoinformatics is transforming understanding of how the immune system functions. This paper reviews advances in the field of immunoinformatics pertinent to autoimmunity research including databases, tools in genomics and proteomics, tools for study of B- and T-cell epitopes, integrative approaches, and web servers.

Allergy and Immunology↗

Prolinks: a database of protein functional linkages derived from coevolution.

The advent of whole-genome sequencing has led to methods that infer protein function and linkages. We have combined four such algorithms (phylogenetic profile, Rosetta Stone, gene neighbor and gene cluster) in a single database--Prolinks--that spans 83 organisms and includes 10 million high-confidence links. The Proteome Navigator tool allows users to browse predicted linkage networks interactively, providing accompanying annotation from public databases. The Prolinks database and the Proteome Navigator tool are available for use online at http://dip.doe-mbi.ucla.edu/pronav.

ATP Synthetase Complexes↗

TassDB: a database of alternative tandem splice sites.

Subtle alternative splice events at tandem splice sites are frequent in eukaryotes and substantially increase the complexity of transcriptomes and proteomes. We have developed a relational database, TassDB (TAndem Splice Site DataBase), which stores extensive data about alternative splice events at GYNGYN donors and NAGNAG acceptors. These splice events are of subtle nature since they mostly result in the insertion/deletion of a single amino acid or the substitution of one amino acid by two others. Currently, TassDB contains 114 554 tandem splice sites of eight species, 5209 of which have EST/mRNA evidence for alternative splicing. In addition, human SNPs that affect NAGNAG acceptors are annotated. The database provides a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download large datasets. This database should facilitate further experimental studies and large-scale bioinformatics analyses of tandem splice sites. The database is available at http://helios.informatik.uni-freiburg.de/TassDB/.

Alternative Splicing↗

Probability-based validation of protein identifications using a modified SEQUEST algorithm.

Database-searching algorithms compatible with shotgun proteomics match a peptide tandem mass spectrum to a predicted mass spectrum for an amino acid sequence within a database. SEQUEST is one of the most common software algorithms used for the analysis of peptide tandem mass spectra by using a cross-correlation (XCorr) scoring routine to match tandem mass spectra to model spectra derived from peptide sequences. To assess a match, SEQUEST uses the difference between the first- and second-ranked sequences (ACn). This value is dependent on the database size, search parameters, and sequence homologies. In this report, we demonstrate the use of a scoring routine (SEQUEST-NORM) that normalizes XCorr values to be independent of peptide size and the database used to perform the search. This new scoring routine is used to objectively calculate the percent confidence of protein identifications and posttranslational modifications based solely on the XCorr value.

Algorithms↗

The proteome of Salmonella enterica serovar typhimurium: current progress on its determination and some applications.

Salmonella typhimurium (official designation Salmonella enterica serovar Typhimurium) is an enteric pathogen and a principal cause of gastroenteritis in humans. A comprehensive description of the proteins of Salmonella and their patterns of expression under different environmental conditions would greatly increase our understanding of the virulence of this organism at the molecular level and provide insights into many other aspects of Salmonella biology. While a variety of two-dimensional studies of Salmonella have been previously carried out to address specific questions, little systematic information is available at the protein level on the numbers of Salmonella polypeptides that have homologues in other organisms, their abundance, and the frequency of post-translational modifications. To test the feasibility of determining the proteome of Salmonella, the identities of 53 randomly sequenced cell envelope proteins have been determined by N-terminal sequencing of spots from two-dimensional gels. In addition to confirming the existence of previously hypothetical proteins predicted from genomic sequencing projects, we found that approximately 20% of the proteins had no matches in sequence databases. The results suggest that proteome analysis is an efficient way to identify novel proteins from prokaryotes and that the analysis provides a useful approach to the study of Salmonella virulence.

Amino Acid Sequence↗

UniPep--a database for human N-linked glycosites: a resource for biomarker discovery.

There has been considerable recent interest in proteomic analyses of plasma for the purpose of discovering biomarkers. Profiling N-linked glycopeptides is a particularly promising method because the population of N-linked glycosites represents the proteomes of plasma, the cell surface, and secreted proteins at very low redundancy and provides a compelling link between the tissue and plasma proteomes. Here, we describe UniPep http://www.unipep.org--a database of human N-linked glycosites--as a resource for biomarker discovery.

Computational Biology↗

[A novel approach for peptide identification by tandem mass spectrometry].

High throughput scoring algorithms that are used to find the match of a tandem mass spectrum to a predicted mass spectrum of a peptide within a database have been applied in shotgun proteomics. However, these algorithms could produce a significant number of incorrect peptide identifications. Here a novel approach was developed to scoring tandem mass spectra against a peptide database, in which fragment ion probabilities, number of enzymatic termini of candidate peptides, matching quality and match pattern between experimental and theoretical spectrum were considered. Benchmarking the novel scorer on a large set of experimental MS/MS spectra, it is demonstrated that PepSearch performs significantly better than the widely used software SEQUEST. The PepSearch software is available at http://compbio.sibsnet.org/projects/pepsearch.

Databases, Protein↗

Comparative proteomics of Cannabis sativa plant tissues.

Comparative proteomics of leaves, flowers, and glands of Cannabis sativa have been used to identify specific tissue-expressed proteins. These tissues have significantly different levels of cannabinoids. Cannabinoids accumulate primarily in the glands but can also be found in flowers and leaves. Proteins extracted from glands, flowers, and leaves were separated using two-dimensional gel electrophoresis. Over 800 protein spots were reproducibly resolved in the two-dimensional gels from leaves and flowers. The patterns of the gels were different and little correlation among the proteins could be observed. Some proteins that were only expressed in flowers were chosen for identification by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry and peptide mass fingerprint database searching. Flower and gland proteomes were also compared, with the finding that less then half of the proteins expressed in flowers were also expressed in glands. Some selected gland protein spots were identified: F1D9.26-unknown prot. (Arabidopsis thaliana), phospholipase D beta 1 isoform 1a (Gossypium hirsutum), and PG1 (Hordeum vulgare). Western blotting was employed to identify a polyketide synthase, an enzyme believed to be involved in cannabinoid biosynthesis, resulting in detection of a single protein.

Blotting, Western↗

ELISA: structure-function inferences based on statistically significant and evolutionarily inspired observations.

UNLABELLED: The problem of functional annotation based on homology modeling is primary to current bioinformatics research. Researchers have noted regularities in sequence, structure and even chromosome organization that allow valid functional cross-annotation. However, these methods provide a lot of false negatives due to limited specificity inherent in the system. We want to create an evolutionarily inspired organization of data that would approach the issue of structure-function correlation from a new, probabilistic perspective. Such organization has possible applications in phylogeny, modeling of functional evolution and structural determination. ELISA (Evolutionary Lineage Inferred from Structural Analysis, http://romi.bu.edu/elisa) is an online database that combines functional annotation with structure and sequence homology modeling to place proteins into sequence-structure-function "neighborhoods". The atomic unit of the database is a set of sequences and structural templates that those sequences encode. A graph that is built from the structural comparison of these templates is called PDUG (protein domain universe graph). We introduce a method of functional inference through a probabilistic calculation done on an arbitrary set of PDUG nodes. Further, all PDUG structures are mapped onto all fully sequenced proteomes allowing an easy interface for evolutionary analysis and research into comparative proteomics. ELISA is the first database with applicability to evolutionary structural genomics explicitly in mind. AVAILABILITY: The database is available at http://romi.bu.edu/elisa.

Amino Acid Sequence↗

An evaluation, comparison, and accurate benchmarking of several publicly available MS/MS search algorithms: sensitivity and specificity analysis.

MS/MS and associated database search algorithms are essential proteomic tools for identifying peptides. Due to their widespread use, it is now time to perform a systematic analysis of the various algorithms currently in use. Using blood specimens used in the HUPO Plasma Proteome Project, we have evaluated five search algorithms with respect to their sensitivity and specificity, and have also accurately benchmarked them based on specified false-positive (FP) rates. Spectrum Mill and SEQUEST performed well in terms of sensitivity, but were inferior to MASCOT, X!Tandem, and Sonar in terms of specificity. Overall, MASCOT, a probabilistic search algorithm, correctly identified most peptides based on a specified FP rate. The rescoring algorithm, PeptideProphet, enhanced the overall performance of the SEQUEST algorithm, as well as provided predictable FP error rates. Ideally, score thresholds should be calculated for each peptide spectrum or minimally, derived from a reversed-sequence search as demonstrated in this study based on a validated data set. The availability of open-source search algorithms, such as X!Tandem, makes it feasible to further improve the validation process (manual or automatic) on the basis of "consensus scoring", i.e., the use of multiple (at least two) search algorithms to reduce the number of FPs. complement.

Algorithms↗

Applications of machine learning and high-dimensional visualization in cancer detection, diagnosis, and management.

Recent technical advances in combinatorial chemistry, genomics, and proteomics have made available large databases of biological and chemical information that have the potential to dramatically improve our understanding of cancer biology at the molecular level. Such an understanding of cancer biology could have a substantial impact on how we detect, diagnose, and manage cancer cases in the clinical setting. One of the biggest challenges facing clinical oncologists is how to extract clinically useful knowledge from the overwhelming amount of raw molecular data that are currently available. In this paper, we discuss how the exploratory data analysis techniques of machine learning and high-dimensional visualization can be applied to extract clinically useful knowledge from a heterogeneous assortment of molecular data. After an introductory overview of machine learning and visualization techniques, we describe two proprietary algorithms (PURS and RadViz) that we have found to be useful in the exploratory analysis of large biological data sets. We next illustrate, by way of three examples, the applicability of these techniques to cancer detection, diagnosis, and management using three very different types of molecular data. We first discuss the use of our exploratory analysis techniques on proteomic mass spectroscopy data for the detection of ovarian cancer. Next, we discuss the diagnostic use of these techniques on gene expression data to differentiate between squamous and adenocarcinoma of the lung. Finally, we illustrate the use of such techniques in selecting from a database of chemical compounds those most effective in managing patients with melanoma versus leukemia.

Artificial Intelligence↗

Proteome analysis of the plant pathogen Xylella fastidiosa reveals major cellular and extracellular proteins and a peculiar codon bias distribution.

The bacteria Xylella fastidiosa is the causative agent of a number of economically important crop diseases, including citrus variegated chlorosis. Although its complete genome is already sequenced, X. fastidiosa is very poorly characterized by biochemical approaches at the protein level. In an initial effort to characterize protein expression in X. fastidiosa we used one- and two-dimensional gel electrophoresis and mass spectrometry to identify the products of 142 genes present in a whole cell extract and in an extracellular fraction of the citrus isolated strain 9a5c. Of particular interest for the study of pathogenesis are adhesion and secreted proteins. Homologs to proteins from three different adhesion systems (type IV fimbriae, mrk pili and hsf surface fibrils) were found to be coexpressed, the last two being detected only as multimeric complexes in the high molecular weight region of one-dimensional electrophoresis gels. Using a procedure to extract secreted proteins as well as proteins weakly attached to the cell surface we identified 30 different proteins including toxins, adhesion related proteins, antioxidant enzymes, different types of proteases and 16 hypothetical proteins. These data suggest that the intercellular space of X. fastidiosa colonies is a multifunctional microenvironment containing proteins related to in vivo bacterial survival and pathogenesis. A codon usage analysis of the most expressed proteins from the whole cell extract revealed a low biased distribution, which we propose is related to the slow growing nature of X. fastidiosa. A database of the X. fastidiosa proteome was developed and can be accessed via the internet (URL: www.proteome.ibi.unicamp.br).

Antioxidants↗

Bioinformatics challenges in proteomics.

A little after the genomic revolution had been celebrated, it seemed as if a competition began to found new -omics disciplines that ultimately all have the same goal, the understanding of biological function. There are many similar definitions for proteomics that can be summarized as follows: proteomics is a large-scale study of structure and function of proteins in an organism or cell. Importantly, the proteome is much more variable than the genome through its interactions with the genome and secondary modifications. It differs depending on the tissue and stage in life-cycle. Hence, proteomics is a very diverse discipline that uses a variety of experimental set-ups and targets in order to elucidate function. Its dissociation from other disciplines can only remain artificial. The bioinformatics applied to proteomics are equally varied. In this review we will focus mainly on a few areas of bioinformatics that seem to us as particularly noteworthy or characteristic for proteomics research, for example in 2DE analysis or mass spectrometry. Another important task of bioinformatics is the prediction of functional properties. We will summarize the approaches taken in order to predict protein networks, which are based on the extensive integration of several kinds of -omics data. We will give a short overview of a demanding field in computational biology, the analysis and prediction of protein 3D structures. In order to provide a broader perspective we will close this review with a generalized description of activities and databases in the realm of proteomics.

Animals↗