Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Identification of the proliferation/differentiation switch in the cellular network of multicellular organisms.

The protein-protein interaction networks, or interactome networks, have been shown to have dynamic modular structures, yet the functional connections between and among the modules are less well understood. Here, using a new pipeline to integrate the interactome and the transcriptome, we identified a pair of transcriptionally anticorrelated modules, each consisting of hundreds of genes in multicellular interactome networks across different individuals and populations. The two modules are associated with cellular proliferation and differentiation, respectively. The proliferation module is conserved among eukaryotic organisms, whereas the differentiation module is specific to multicellular organisms. Upon differentiation of various tissues and cell lines from different organisms, the expression of the proliferation module is more uniformly suppressed, while the differentiation module is upregulated in a tissue- and species-specific manner. Our results indicate that even at the tissue and organism levels, proliferation and differentiation modules may correspond to two alternative states of the molecular network and may reflect a universal symbiotic relationship in a multicellular organism. Our analyses further predict that the proteins mediating the interactions between these modules may serve as modulators at the proliferation/differentiation switch.

Animals↗

Recent advances in computational genomics.

In the post-genomic era, the new discipline of functional genomics is now facing the challenge of associating a function (as well as estimating its relevance to industrial applications) to about 100,000 microbial, plant or animal genes of known sequence but unknown function. Besides the design of databases, computational methods are increasingly becoming intimately linked with the various experimental approaches. Consequently, bioinformatics is rapidly evolving into independent fields addressing the specific problems of interpreting i) genomic sequences, ii) protein sequences and 3D-structures, as well as iii) transcriptome and macromolecular interaction data. It is thus increasingly difficult for the biologist to choose the computational approaches that perform best in these various areas. This paper attempts to review the most useful developments of the last 2 years.

Computational Biology↗

Who tangos with GOA?-Use of Gene Ontology Annotation (GOA) for biological interpretation of '-omics' data and for validation of automatic annotation tools.

The number of large-scale experimental datasets generated from high-throughput technologies has grown rapidly. Biological knowledge resources such as the Gene Ontology Annotation (GOA) database, which provides high-quality functional annotation to proteins within the UniProt Knowledgebase, can play an important role in the analysis of such data. The integration of GOA with analytical tools has proved to aid the clustering, annotation and biological interpretation of such large expression datasets. GOA is also useful in the development and validation of automated annotation tools, in particular text-mining systems. The increasing interest in GOA highlights the great potential of this freely available resource to assist both the biological research and bioinformatics communities.

Animals↗

An SVM-based algorithm for identification of photosynthesis-specific genome features.

This paper presents a novel algorithm for identification and functional characterization of "key" genome features responsible for a particular biochemical process of interest. The central idea is that individual genome features are identified as "key" features if the discrimination accuracy between two classes of genomes with respect to a given biochemical process is sufficiently affected by the inclusion or exclusion of these features. In this paper, genome features are defined by high-resolution gene functions. The discrimination procedure utilizes the Support Vector Machine classification technique. The application to the oxygenic photosynthetic process resulted in 126 highly confident candidate genome features. While many of these features are well-known components in the oxygenic photosynthetic process, others are completely unknown, even including some hypothetical proteins. It is obvious that our algorithm is capable of discovering features related to a targeted biochemical process.

Algorithms↗

Dependence of molecular properties on proteomic family for marketed oral drugs.

An association of drugs with their proteomic family reveals that molecular properties of drugs targeting proteases, lipid and peptide G-protein-coupled receptors (GPCRs), and nuclear hormone receptors significantly exceed limits for some properties in the "rule of five", while drugs targeting cytochrome P450s, biogenic amine GPCRs, and transporters have significantly lower values for certain properties. Also, the variation in drug targets appears to be a factor explaining increasing molecular weight over time.

Administration, Oral↗

Scansite 2.0: Proteome-wide prediction of cell signaling interactions using short sequence motifs.

Scansite identifies short protein sequence motifs that are recognized by modular signaling domains, phosphorylated by protein Ser/Thr- or Tyr-kinases or mediate specific interactions with protein or phospholipid ligands. Each sequence motif is represented as a position-specific scoring matrix (PSSM) based on results from oriented peptide library and phage display experiments. Predicted domain-motif interactions from Scansite can be sequentially combined, allowing segments of biological pathways to be constructed in silico. The current release of Scansite, version 2.0, includes 62 motifs characterizing the binding and/or substrate specificities of many families of Ser/Thr- or Tyr-kinases, SH2, SH3, PDZ, 14-3-3 and PTB domains, together with signature motifs for PtdIns(3,4,5)P(3)-specific PH domains. Scansite 2.0 contains significant improvements to its original interface, including a number of new generalized user features and significantly enhanced performance. Searches of all SWISS-PROT, TrEMBL, Genpept and Ensembl protein database entries are now possible with run times reduced by approximately 60% when compared with Scansite version 1.0. Scansite 2.0 allows restricted searching of species-specific proteins, as well as isoelectric point and molecular weight sorting to facilitate comparison of predictions with results from two-dimensional gel electrophoresis experiments. Support for user-defined motifs has been increased, allowing easier input of user-defined matrices and permitting user-defined motifs to be combined with pre-compiled Scansite motifs for dual motif searching. In addition, a new series of Sequence Match programs for non-quantitative user-defined motifs has been implemented. Scansite is available via the World Wide Web at http://scansite.mit.edu.

Algorithms↗

Peeling the yeast protein network.

A set of highly connected proteins (or hubs) plays an important role for the integrity of the protein interaction network of Saccharomyces cerevisae by connecting the network's intrinsic modules. The importance of the hubs' central placement is further confirmed by their propensity to be lethal. However, although highly emphasized, little is known about the topological coherence among the hubs. Applying a core decomposition method which allows us to identify the inherent layer structure of the protein interaction network, we find that the probability of nodes both being essential and evolutionary conserved successively increases toward the innermost cores. While connectivity alone is often not a sufficient criterion to assess a protein's functional, evolutionary and topological relevance, we classify nodes as globally and locally central depending on their appearance in the inner or outer cores. The observation that globally central proteins participate in a substantial number of protein complexes which display an elevated degree of evolutionary conservation allows us to hypothesize that globally central proteins serve as the evolutionary backbone of the proteome. Even though protein interaction data are extensively flawed, we find that our results are very robust against inaccurately determined protein interactions.

Databases, Protein↗

An XML-based system for synthesis of data from disparate databases.

Diverse data sets have become key building blocks of translational biomedical research. Data types captured and referenced by sophisticated research studies include high throughput genomic and proteomic data, laboratory data, data from imagery, and outcome data. In this paper, the authors present the application of an XML-based data management system to support integration of data from disparate data sources and large data sets. This system facilitates management of XML schemas and on-demand creation and management of XML databases that conform to these schemas. They illustrate the use of this system in an application for genotype-phenotype correlation analyses. This application implements a method of phenotype-genotype correlation based on phylogenetic optimization of large data sets of mouse SNPs and phenotypic data. The application workflow requires the management and integration of genomic information and phenotypic data from external data repositories and from the results of phenotype-genotype correlation analyses. Our implementation supports the process of carrying out a complex workflow that includes large-scale phylogenetic tree optimizations and application of Maddison's concentrated changes test to large phylogenetic tree data sets. The data management system also allows collaborators to share data in a uniform way and supports complex queries that target data sets.

Animals↗

Proteome analyses using accurate mass and elution time peptide tags with capillary LC time-of-flight mass spectrometry.

We describe the application of capillary liquid chromatography (LC) time-of-flight (TOF) mass spectrometric instrumentation for the rapid characterization of microbial proteomes. Previously (Lipton et al., Proc. Natl. Acad. Sci. U.S.A. 2002, 99, 11049) the peptides from a series of growth conditions of Deinococcus radiodurans have been characterized using capillary LC MS/MS and accurate mass measurements which are captured as an accurate mass and time (AMT) tag database. Using this AMT tag database, detected peptides can be assigned using measurements obtained on a TOF due to the additional use of elution time data as a constraint. When peptide matches are obtained using AMT tags (i.e., using both constraints) unique matches of a mass spectral peak occurs 88% of the time. Not only are AMT tag matches unique in most cases, the coverage of the proteome is high; approximately 3500 unique peptide AMT tags are found on average per capillary LC run. From the results of the AMT tag database search, approximately 900 ORFs detected using LC-TOFMS, with approximately 500 ORFs covered by at least two AMT tags. These results indicate that AMT database searches with modest mass and elution time criteria can provide proteomic information for approximately one thousand proteins in a single run of <3 h. The advantage of this method over using MS/MS based techniques is the large number of identifications that occur in a single experiment as well as the basis for improved quantitation. For MS/MS experiments, the number of peptide identifications is severely restricted because of the time required to dissociate the peptides individually. These results demonstrate the utility of the AMT tag approach using capillary LC-TOF MS instruments, and also show that AMT tags developed using other instrumentation can be effectively utilized.

Amino Acid Sequence↗

Annotating proteins from endoplasmic reticulum and Golgi apparatus in eukaryotic proteomes.

The sub-cellular localization of a native protein constitutes one coarse-grained aspect of its function. Transport between compartments is often regulated through short sequence motifs. Here, we analyzed experimentally characterized endoplasmic reticulum (ER)/ Golgi retrieval motifs and investigated the accuracy of homology-transfer. Only the C-terminal ER retrieval motifs KDEL, HDEL and AIAKE were sufficiently specific. However, even unspecific motifs may help, provided we know the probability for localization given the motif. We provided such estimates. We also rigorously estimated the accuracy and coverage for inferring ER and Golgi localization through homology-transfer by sequence similarity. In entire proteomes, we could thereby annotate 3304 ER (3182 membrane) and 1853 Golgi (759 membrane) proteins. We identified another putative 5157 globular and 3941 membrane ER or Golgi proteins. Each experimental annotation yielded, on average, one to three high-accuracy and five to six low-accuracy homology-transfers in the six proteomes. These numbers will increase with each new experimental annotation.

Amino Acid Motifs↗

Improved 2D nano-LC/MS for proteomics applications: a comparative analysis using yeast proteome.

The most commonly used method for protein identification with two-dimensional (2D) online liquid chromatography-mass spectrometry (LC/MS) involves the elution of digest peptides from a strong cation exchange column by an injected salt step gradient of increasing salt concentration followed by reversed phase separation. However, in this approach ion exchange chromatography does not perform to its fullest extent, primarily because the injected volume of salt solution is not optimized to the SCX column. To improve the performance of strong cation exchange chromatography, we developed a new method for 2D online nano-LC/MS that replaces the injected salt step gradient with an optimized semicontinuous pumped salt gradient. The viability of this method is demonstrated in the results of a comparative analysis of a complex tryptic digest of the yeast proteome using the injected salt solution method and the semicontinuous pump salt method. The semicontinuous pump salt method compares favorably with the commonly used injection method and also with an offline 2D-LC method.

Chromatography, Ion Exchange↗

Global protein function prediction from protein-protein interaction networks.

Determining protein function is one of the most challenging problems of the post-genomic era. The availability of entire genome sequences and of high-throughput capabilities to determine gene coexpression patterns has shifted the research focus from the study of single proteins or small complexes to that of the entire proteome. In this context, the search for reliable methods for assigning protein function is of primary importance. There are various approaches available for deducing the function of proteins of unknown function using information derived from sequence similarity or clustering patterns of co-regulated genes, phylogenetic profiles, protein-protein interactions (refs. 5-8 and Samanta, M.P. and Liang, S., unpublished data), and protein complexes. Here we propose the assignment of proteins to functional classes on the basis of their network of physical interactions as determined by minimizing the number of protein interactions among different functional categories. Function assignment is proteome-wide and is determined by the global connectivity pattern of the protein network. The approach results in multiple functional assignments, a consequence of the existence of multiple equivalent solutions. We apply the method to analyze the yeast Saccharomyces cerevisiae protein-protein interaction network. The robustness of the approach is tested in a system containing a high percentage of unclassified proteins and also in cases of deletion and insertion of specific protein interactions.

Algorithms↗

Widespread occurrence of alternative splicing at NAGNAG acceptors contributes to proteome plasticity.

Splice acceptors with the genomic NAGNAG motif may cause NAG insertion-deletions in transcripts, occur in 30% of human genes and are functional in at least 5% of human genes. We found five significant biases indicating that their distribution is nonrandom and that they are evolutionarily conserved and tissue-specific. Because of their subtle effects on mRNA and protein structures, these splice acceptors are often overlooked or underestimated, but they may have a great impact on biology and disease.

Alternative Splicing↗

Standard mixtures for proteome studies.

Mixtures of moderate complexity were formed from 23 peptides and 12 proteins digested with trypsin, all individually characterized. These mixtures were analyzed with replicates in full and windowed m/z ranges using online high-performance reverse phase liquid chromatography coupled via electrospray ionization to an ion trap mass spectrometer. The resulting spectra were searched using SEQUEST against databases of different sizes and contents and confidences of the observed identifications were evaluated by our earlier statistical model. These data were then combined with biologically derived spectral data, searched, and further evaluated. All peptides but one and all proteins were identified with high confidence. Additionally, the presence and behavior of quadruply charged peptides was analyzed. The properties of the proposed peptide and protein mixtures as well as the performance of the statistical model were carefully investigated. These mixtures mimic the complexity seen in large-scale proteomics experiments, and are proposed to serve as quality assessment standards for future proteome studies.

Animals↗

ProteomeCommons.org IO Framework: reading and writing multiple proteomics data formats.

MOTIVATION: Effective use of proteomics data, specifically mass spectrometry data, relies on the ability to read and write the many mass spectrometer file formats. Even with mass spectrometer vendor-specific libraries and vendor-neutral file formats, such as mzXML and mzData it can be difficult to extract raw data files in a form suitable for batch processing and basic research. Introduced here are the ProteomeCommons.org Input and Output Framework, abbreviated to IO Framework, which is designed to abstractly represent mass spectrometry data. This project is a public, open-source, free-to-use framework that supports most of the mass spectrometry data formats, including current formats, legacy formats and proprietary formats that require a vendor-specific library in order to operate. The IO Framework includes an on-line tool for non-programmers and a set of libraries that developers may use to convert between various proteomics file formats. AVAILABILITY: The current source-code and documentation for the ProteomeCommons.org IO Framework is freely available at http://www.proteomecommons.org/current/531/

Algorithms↗