Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Protein interactions: two methods for assessment of the reliability of high throughput observations.

High throughput methods for detecting protein interactions require assessment of their accuracy. We present two forms of computational assessment. The first method is the expression profile reliability (EPR) index. The EPR index estimates the biologically relevant fraction of protein interactions detected in a high throughput screen. It does so by comparing the RNA expression profiles for the proteins whose interactions are found in the screen with expression profiles for known interacting and non-interacting pairs of proteins. The second form of assessment is the paralogous verification method (PVM). This method judges an interaction likely if the putatively interacting pair has paralogs that also interact. In contrast to the EPR index, which evaluates datasets of interactions, PVM scores individual interactions. On a test set, PVM identifies correctly 40% of true interactions with a false positive rate of approximately 1%. EPR and PVM were applied to the Database of Interacting Proteins (DIP), a large and diverse collection of protein-protein interactions that contains over 8000 Saccharomyces cerevisiae pairwise protein interactions. Using these two methods, we estimate that approximately 50% of them are reliable, and with the aid of PVM we identify confidently 3003 of them. Web servers for both the PVM and EPR methods are available on the DIP website (dip.doe-mbi.ucla.edu/Services.cgi).

Algorithms↗

Fold recognition via a tree.

Recently, we developed a pairwise structural alignment algorithm using realistic structural and environmental information (SAUCE). In this paper, we at first present an automatic fold hierarchical classification based on SAUCE alignments. This classification enables us to build a fold tree containing different levels of multiple structural profiles. Then a tree-based fold search algorithm is described. We applied this method to a group of structures with sequence identity less than 35% and did a series of leave one out tests. These tests are approximately comparable to fold recognition tests on superfamily level. Results show that fold recognition via a fold tree can be faster and better at detecting distant homologues than classic fold recognition methods.

Algorithms↗

Bioinformatics for glycomics: status, methods, requirements and perspectives.

The term 'glycomics' describes the scientific attempt to identify and study all the glycan molecules - the glycome - synthesised by an organism. The aim is to create a cell-by-cell catalogue of glycosyltransferase expression and detected glycan structures. The current status of databases and bioinformatics tools, which are still in their infancy, is reviewed. The structures of glycans as secondary gene products cannot be easily predicted from the DNA sequence. Glycan sequences cannot be described by a simple linear one-letter code as each pair of monosaccharides can be linked in several ways and branched structures can be formed. Few of the bioinformatics algorithms developed for genomics/proteomics can be directly adapted for glycomics. The development of algorithms, which allow a rapid, automatic interpretation of mass spectra to identify glycan structures is currently the most active field of research. The lack of generally accepted ways to normalise glycan structures and exchange glycan formats hampers an efficient cross-linking and the automatic exchange of distributed data. The upcoming glycomics should accept that unrestricted dissemination of scientific data accelerates scientific findings and initiates a number of new initiatives to explore the data.

Animals↗

MFAML: a standard data structure for representing and exchanging metabolic flux models.

SUMMARY: MFAML is a standard data structure designed for the formal representation and effective exchange of metabolic flux models. It allows for the explicit description of stationary states of a metabolic system by defining environmental/genetic conditions of the system, e.g. flux measurements, balancing constraints and physiological objectives as well as basic information on metabolites and reactions. In addition, a library of MFAML comprising a model parser and a converter provides an open framework for establishing the pipeline from metabolic modeling to metabolic flux analysis. AVAILABILITY: MFAML (version 1) is fully described and available at http://mbel.kaist.ac.kr/mfaml/.

Computer Simulation↗

Representations of molecular pathways: an evaluation of SBML, PSI MI and BioPAX.

MOTIVATION: Analysis and simulation of pathway data is of high importance in bioinformatics. Standards for representation of information about pathways are necessary for integration and analysis of data from various sources. Recently, a number of representation formats for pathway data, SBML, PSI MI and BioPAX, have been proposed. RESULTS: In this paper we compare these formats and evaluate them with respect to their underlying models, information content and possibilities for easy creation of tools. The evaluation shows that the main structure of the formats is similar. However, SBML is tuned towards simulation models of molecular pathways while PSI MI is more suitable for representing details about particular interactions and experiments. BioPAX is the most general and expressive of the formats. These differences are apparent in allowed information and the structure for representation of interactions. We discuss the impact of these differences both with respect to information content in existing databases and computational properties for import and analysis of data.

Computational Biology↗

Metabolic pathway analysis in trypanosomes and malaria parasites.

Identification of novel drug targets is required for the development of new classes of drugs to overcome drug resistance and replace less efficacious treatments. In theory, knowledge of the entire genome of a pathogen identifies every potential drug target in any given microbe. In practice, the sheer complexity and the inadequate or inaccurate annotation of genomic information makes target identification and selection somewhat more difficult. Analysis of metabolic pathways provides a useful conceptual framework for the identification of potential drug targets and also for improving our understanding of microbial responses to nutritional, chemical and other environmental stresses. A number of metabolic databases are available as tools for such analyses. The strengths and weaknesses of this approach are discussed.

Animals↗

The challenges and rewards of integrating diverse neuroscience information.

The design of database models and schemas for storing, cross-referencing, and retrieving neuroscience information faces issues that are similar but more complex than most of the other biomedical disciplines, such as genomics and proteonomics. Specifically, the visualization and manipulation of very large and diverse image data, such as digital brain atlases and functional magnetic resonance images, play a unique role in neuroscience while much of the associated information is textually recorded. Nongraphical information can include the annotation of large brain structures ranging from anatomical regions to intracellular structures, the description of cellular functional properties, and their various interrelationships, such as fiber connections. It is necessary that the heterogeneous and distributed types of data be cross-referenced to each other so that this diverse information can be efficiently retrieved, shared, and exchanged among the different neuroscientific disciplines. Continued advances in computers and Internet technologies appear to indicate that increasingly large data sets will be maintained on local or regional file servers and that informational interoperability will be achieved using a networked information system infrastructure. The authors and others have proposed and implemented models of semantically organized information systems that utilize centrally stored and highly structured archival information to index, cross-reference, and retrieve diverse, Web-based data sets.

Brain↗

A direct comparison of protein interaction confidence assignment schemes.

BACKGROUND: Recent technological advances have enabled high-throughput measurements of protein-protein interactions in the cell, producing large protein interaction networks for various species at an ever-growing pace. However, common technologies like yeast two-hybrid may experience high rates of false positive detection. To combat false positive discoveries, a number of different methods have been recently developed that associate confidence scores with protein interactions. Here, we perform a rigorous comparative analysis and performance assessment among these different methods. RESULTS: We measure the extent to which each set of confidence scores correlates with similarity of the interacting proteins in terms of function, expression, pattern of sequence conservation, and homology to interacting proteins in other species. We also employ a new metric, the Signal-to-Noise Ratio of protein complexes embedded in each network, to assess the power of the different methods. Seven confidence assignment schemes, including those of Bader et al., Deane et al., Deng et al., Sharan et al., and Qi et al., are compared in this work. CONCLUSION: Although the performance of each assignment scheme varies depending on the particular metric used for assessment, we observe that Deng et al. yields the best performance overall (in three out of four viable measures). Importantly, we also find that utilizing any of the probability assignment schemes is always more beneficial than assuming all observed interactions to be true or equally likely.

Caenorhabditis elegans Proteins↗

Protein molecular function prediction by Bayesian phylogenomics.

We present a statistical graphical model to infer specific molecular function for unannotated protein sequences using homology. Based on phylogenomic principles, SIFTER (Statistical Inference of Function Through Evolutionary Relationships) accurately predicts molecular function for members of a protein family given a reconciled phylogeny and available function annotations, even when the data are sparse or noisy. Our method produced specific and consistent molecular function predictions across 100 Pfam families in comparison to the Gene Ontology annotation database, BLAST, GOtcha, and Orthostrapper. We performed a more detailed exploration of functional predictions on the adenosine-5'-monophosphate/adenosine deaminase family and the lactate/malate dehydrogenase family, in the former case comparing the predictions against a gold standard set of published functional characterizations. Given function annotations for 3% of the proteins in the deaminase family, SIFTER achieves 96% accuracy in predicting molecular function for experimentally characterized proteins as reported in the literature. The accuracy of SIFTER on this dataset is a significant improvement over other currently available methods such as BLAST (75%), GeneQuiz (64%), GOtcha (89%), and Orthostrapper (11%). We also experimentally characterized the adenosine deaminase from Plasmodium falciparum, confirming SIFTER's prediction. The results illustrate the predictive power of exploiting a statistical model of function evolution in phylogenomic problems. A software implementation of SIFTER is available from the authors.

Adenosine Deaminase↗

Genome-derived vaccines.

Vaccine research entered a new era when the complete genome of a pathogenic bacterium was published in 1995. Since then, more than 97 bacterial pathogens have been sequenced and at least 110 additional projects are now in progress. Genome sequencing has also dramatically accelerated: high-throughput facilities can draft the sequence of an entire microbe (two to four megabases) in 1 to 2 days. Vaccine developers are using microarrays, immunoinformatics, proteomics and high-throughput immunology assays to reduce the truly unmanageable volume of information available in genome databases to a manageable size. Vaccines composed by novel antigens discovered from genome mining are already in clinical trials. Within 5 years we can expect to see a novel class of vaccines composed by genome-predicted, assembled and engineered T- and Bcell epitopes. This article addresses the convergence of three forces--microbial genome sequencing, computational immunology and new vaccine technologies--that are shifting genome mining for vaccines onto the forefront of immunology research.

Animals↗

Struct2net: integrating structure into protein-protein interaction prediction.

UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the structure-based interaction prediction technique threads these two sequences to all the protein complexes in the PDB and then chooses the best potential match. Based on this match, structural information is incorporated into logistic regression to evaluate the probability of these two proteins interacting. This paper also describes a random forest classifier which can effectively combine the structure-based prediction results and other functional annotations together to predict protein interactions. Experimental results indicate that the predictive power of the structure-based method is better than many other information sources. Also, combining the structure-based method with other information sources allows us to achieve a better performance than when structure information is not used. We also tested our method on a set of approximately 1000 yeast genes and, interestingly, the predicted interaction network is a scale-free network. Our method predicted some potential interactions involving yeast homologs of human disease-related proteins. SUPPLEMENTARY INFORMATION: http://theory.csail.mit.edu/struct2net

Algorithms↗

Analysis of the wheat and Puccinia triticina (leaf rust) proteomes during a susceptible host-pathogen interaction.

Wheat leaf rust is caused by the fungus Puccinia triticina. The genetics of resistance follows the gene-for-gene hypothesis, and thus the presence or absence of a single host resistance gene renders a plant resistant or susceptible to a leaf rust race bearing the corresponding avirulence gene. To investigate some of the changes in the proteomes of both host and pathogen during disease development, a susceptible line of wheat infected with a virulent race of leaf rust were compared to mock-inoculated wheat using 2-DE (with IEF pH 4-8) and MS. Up-regulated protein spots were excised and analyzed by MALDI-QqTOF MS/MS, followed by cross-species protein identification. Where possible MS/MS spectra were matched to homologous proteins in the NCBI database or to fungal ESTs encoding putative proteins. Searching was done using the MASCOT search engine. Remaining unmatched spectra were then sequenced de novo and queried against the NCBInr database using the BLAST and MS BLAST tools. A total of 32 consistently up-regulated proteins were examined from the gels representing the 9-day post-infection proteome in susceptible plants. Of these 7 are host proteins, 22 are fungal proteins of known or hypothetical function and 3 are unknown proteins of putative fungal origin.

Amino Acid Sequence↗

Discover true association rates in multi-protein complex proteomics data sets.

Experimental processes to collect and process proteomics data are increasingly complex, while the computational methods to assess the quality and significance of these data remain unsophisticated. These challenges have led to many biological oversights and computational misconceptions. We developed a complete empirical Bayes model to analyze multi-protein complex (MPC) proteomics data derived from peptide mass spectrometry detections of purified protein complex pull-down experiments. Our model considers not only bait-prey associations, but also prey-prey associations missed in previous work. Using our model and a yeast MPC proteomics data set, we estimated that there should be an average of 28 true associations per MPC, almost ten times as high as was previously estimated. For data sets generated to mimic a real proteome, our model achieved on average 80% sensitivity in detecting true associations, as compared with the 3% sensitivity in previous work, while maintaining a comparable false discovery rate of 0.3%.

Algorithms↗

The Database of Quantitative Cellular Signaling: management and analysis of chemical kinetic models of signaling networks.

MOTIVATION: Analysis of cellular signaling interactions is expected to pose an enormous informatics challenge, perhaps even larger than analyzing the genome. The complex networks arising from signaling processes are traditionally represented as block diagrams. A key step in the evolution toward a more quantitative understanding of signaling is to explicitly specify the kinetics of all chemical reaction steps in a pathway. Technical advances in proteomics and high-throughput protein interaction assays promise a flood of such quantitative data. While annotations, molecular information and pathway connectivity have been compiled in several databases, and there are several proposals for general cell model description languages, there is currently little experience with databases of chemical kinetics and reaction level models of signaling networks. RESULTS: The Database of Quantitative Cellular Signaling is a repository of models of signaling pathways. It is intended both to serve the growing field of chemical-reaction level simulation of signaling networks, and to anticipate issues in large-scale data management for signaling chemistry. AVAILABILITY: The Database of Quantitative Cellular Signaling is available at http://doqcs.ncbs.res.in. Links to the signaling model simulator, GENESIS/Kinetikit are at http://www.ncbs.res.in/~bhalla/kkit/index.html and are also provided from within the database. The database source code is available under the GNU Public License.

Abstracting and Indexing↗

Compositional characterization of the cytoskeleton of NK-like cells.

The cytoskeleton is a dynamic structure that contributes to cell function in terms of shape, movement, transport and secretion. It also provides a platform for regional activities such as signaling, biosynthesis and energy production. The present manuscript describes a method for cytoskeleton isolation based on capture with magnetic microbeads and its application to the analysis of the NK like cell line, YTS. The isolated proteins were separated by SDS-PAGE and the peptides from the in gel digested proteins were analyzed by on line nano-LC-MSMS. Approximately 76% of the 126 isolated proteins were either components of the cytoskeleton or proteins that were known to be capable of associating with the cytoskeleton. The enrichment was confirmed by western blot for actin and alpha-actinin. The isolation was dependent on intact actin microfilaments as pretreatment of cells with cytochalasin D resulted in a marked reduction in the number of proteins isolated. The method allowed for the identification of several proteins that have not been previously described in lymphoid cells (EPLIN, SETA). A number of other scaffolding and lipid raft associated proteins were described suggesting a link between the cytoskeleton and these structures. The approach may have application to the proteomic examination of the cytoskeleton in a variety of cell types.

Actin Cytoskeleton↗