Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Systems for the detection and analysis of protein-protein interactions.

The analysis of protein-protein interactions is important for developing a better understanding of the functional annotations of proteins that are involved in various biochemical reactions in vivo. The discovery that a protein with an unknown function binds to a protein with a known function could provide a significant clue to the cellular pathway concerning the unknown protein. Therefore, information on protein-protein interactions obtained by the comprehensive analysis of all gene products is available for the construction of interactive networks consisting of individual protein-protein interactions, which, in turn, permit elaborate biological phenomena to be understood. Systems for detecting protein-protein interactions in vitro and in vivo have been developed, and have been modified to compensate for limitations. Using these novel approaches, comprehensive and reliable information on protein-protein interactions can be determined. Systems that permit this to be achieved are described in this review.

Chromatography, Affinity↗

Automatic classification and pattern discovery in high-throughput protein crystallization trials.

Conceptually, protein crystallization can be divided into two phases search and optimization. Robotic protein crystallization screening can speed up the search phase, and has a potential to increase process quality. Automated image classification helps to increase throughput and consistently generate objective results. Although the classification accuracy can always be improved, our image analysis system can classify images from 1,536-well plates with high classification accuracy (85%) and ROC score (0.87), as evaluated on 127 human-classified protein screens containing 5,600 crystal images and 189,472 non-crystal images. Data mining can integrate results from high-throughput screens with information about crystallizing conditions, intrinsic protein properties, and results from crystallization optimization. We apply association mining, a data mining approach that identifies frequently occurring patterns among variables and their values. This approach segregates proteins into groups based on how they react in a broad range of conditions, and clusters cocktails to reflect their potential to achieve crystallization. These results may lead to crystallization screen optimization, and reveal associations between protein properties and crystallization conditions. We also postulate that past experience may lead us to the identification of initial conditions favorable to crystallization for novel proteins.

Algorithms↗

Many commonly used siRNAs risk off-target activity.

Using small interfering RNA (siRNA) to induce sequence specific gene silencing is fast becoming a standard tool in functional genomics. As siRNAs in some cases tolerate mismatches with the mRNA target, knockdown of genes other than the intended target could make results difficult to interpret. In an investigation of 359 published siRNA sequences, we have found that about 75% of them have a risk of eliciting non-specific effects. A possible cause for this is the popular BLAST search engine, which is inappropriate for such short oligos as siRNAs. Furthermore, we used new special purpose hardware to do a transcriptome-wide screening of all possible siRNAs, and show that many unique siRNAs exist per target even if several mismatches are allowed. Hence, we argue that the risk of off-target effects is unnecessary and should be avoided in future siRNA design.

Base Sequence↗

Using GO-PseAA predictor to predict enzyme sub-class.

Enzyme function is much less conserved than anticipated, i.e., the requirement for sequence similarity that implies similarity in enzymatic function is much higher than the requirement that implies similarity in protein structure. This is because the function of an enzyme is an extremely complicated problem that may involve very subtle structural details as well as many other physical chemistry factors. Accordingly, if simply based on the sequence similarity approach, it would hardly get a decent success rate in predicting enzyme sub-class even for a dataset consisting of samples with 50% sequence identity. To cope with such a situation, the GO-PseAA predictor was adopted to identify the sub-class for each of the six main enzyme families. It has been observed that, even for the much more stringent datasets in which none of the enzymes has 25% sequence identity to any others, the overall success rates are 73-95%, suggesting that the GO-PseAA predictor can catch the core features of the statistical samples concerned and may become a useful high throughput tool in proteomics and bioinformatics.

Algorithms↗

Systems approaches to understanding cell signaling and gene regulation.

The age of 'omics' is upon us, and scientific papers that reflect this are starting to appear at an ever-increasing rate. The amount of information generated in any 'omics' program is daunting and often overwhelms plant scientists whose main interests relate to cell or developmental biology. For this revolution in data generation to have any impact in plant signaling studies, we must have great confidence in both the quality of the data and our ability to represent it in ways that are meaningful to general plant biologists. Systems biology has begun to address these issues and to provide examples in which the analysis of large data sets has led to biological insights into cell signaling and gene regulation.

Arabidopsis↗

Integrative bioinformatics for functional genome annotation: trawling for G protein-coupled receptors.

G protein-coupled receptors (GPCR) are amongst the best studied and most functionally diverse types of cell-surface protein. The importance of GPCRs as mediates or cell function and organismal developmental underlies their involvement in key physiological roles and their prominence as targets for pharmacological therapeutics. In this review, we highlight the requirement for integrated protocols which underline the different perspectives offered by different sequence analysis methods. BLAST and FastA offer broad brush strokes. Motif-based search methods add the fine detail. Structural modelling offers another perspective which allows us to elucidate the physicochemical properties that underlie ligand binding. Together, these different views provide a more informative and a more detailed picture of GPCR structure and function. Many GPCRs remain orphan receptors with no identified ligand, yet as computer-driven functional genomics starts to elaborate their functions, a new understanding of their roles in cell and developmental biology will follow.

Algorithms↗

The use of peptidomics in endocrine research.

In 2002, the Nobel Prize for chemistry was awarded to the inventors of two novel ionization techniques in mass spectrometry, MALDI and ESI. These techniques, often in combination with data from genomic databases, represent an extremely powerful tool in analytical (bio)chemistry, with many applications, e.g., in the field of proteomics. Peptides, which are small proteins, have, despite their importance as controlling agents in numerous physiological processes, as yet been much less intensively studied by these novel techniques than larger proteins. The term peptidomics, i.e., the study of all peptides expressed by a certain cell, organ or organism was only introduced in 2001. In neuroendocrinology, spectacular progress could already be realized and the future looks bright. In this minireview we discuss the different methodologies that are used in peptidomics and give an overview of the wide range of applications.

Animals↗

Progress in bioinformatics and the importance of being earnest.

In silico biology has gathered momentum as, worldwide, scientists have united in a common quest to sequence, store and analyse complete genomes. This year, a pivotal achievement of this cooperative endeavour was realised in the release of a public draft of the human genome, and with it the promises to improve our understanding of diverse aspects of biology and to yield a healthier future with safe personalized medicines. Key to these goals will be the need to elucidate and characterise the genes and gene products encoded not just in the human genome, but in many genomes. These tasks are underpinned by the concepts and processes of genome and gene/protein evolution, regulation of gene expression, mechanisms of protein folding, the manifestation of protein function, and so on, all of which must be understood in the context of complex, dynamic biological systems. Our use of computers to model such concepts and systems must be placed in the context of the current limits of our understanding of them:- it is important to recognise, for example, that we don't have a common understanding either of what constitutes a gene or a protein function; we can't invariably say that a particular sequence or fold has arisen via divergent or convergent evolution; and we don't fully understand the rules of protein folding. Accepting what we can't do in silico is essential in appreciating what we can do. Without this understanding, it is easy to be misled, as notions of what particular computational approaches can achieve are sometimes rather optimistic. There are valuable lessons to be learned here from the field of Artificial Intelligence, principal among which is the realisation that capturing and representing complex knowledge is time consuming, expensive and hard. Thus, we argue here that if bioinformatics is to tackle biological complexity in earnest, it would be wise to absorb the experience distilled from decades of artificial intelligence research, and to approach the road ahead with caution, rigour and pragmatism.

Artificial Intelligence↗

Shotgun annotation of histone modifications: a new approach for streamlined characterization of proteins by top down mass spectrometry.

Eukaryotic histones serve as prototypical examples of posttranslational complexity with diverse modifications (PTMs) on many different residues that comprise a "Histone Code". To help crack this code more efficiently, we demonstrate a new strategy for protein characterization wherein complete PTM descriptions are obtained by database retrieval instead of manual interpretation of information-rich data from high-resolution tandem mass spectrometry (MS/MS). A database of nearly 50 000 modified histone H4 sequences was created and queried with 91 fragment ions from electron capture dissociation of a histone form +112 Da (versus unmodified mass) selectively accumulated in a quadrupole Fourier transform hybrid mass spectrometer. The correct form atop the retrieval list indicated dimethylation at Lys20, acetylation at the N terminus, and acetylation at Lys16 (resolved from trimethylation, Deltam = 0.036 Da). A statistical evaluation reveals the critical role of mass accuracy and that PTM "isomers" are retrieved as next-best matches. The applicability of shotgun annotation to forms of H4 with up to six PTMs is demonstrated, with extensibility to other histones (e.g., H2A, H2B, H3) and other protein classes projected.

Amino Acid Sequence↗

Prediction of functional class of the SARS coronavirus proteins by a statistical learning method.

The complete genome of severe acute respiratory syndrome coronavirus (SARS-CoV) reveals the existence of putative proteins unique to SARS-CoV. Identification of their function facilitates a mechanistic understanding of SARS infection and drug development for its treatment. The sequence of the majority of these putative proteins has no significant similarity to those of known proteins, which complicates the task of using sequence analysis tools to probe their function. Support vector machines (SVM), useful for predicting the functional class of distantly related proteins, is employed to ascribe a possible functional class to SARS-CoV proteins. Testing results indicate that SVM is able to predict the functional class of 73% of the known SARS-CoV proteins with available sequences and 67% of 18 other novel viral proteins. A combination of the sequence comparison method BLAST and SVMProt can further improve the prediction accuracy of SMVProt such that the functional class of two additional SARS-CoV proteins is correctly predicted. Our study suggests that the SARS-CoV genome possibly contains a putative voltage-gated ion channel, structural proteins, a carbon-oxygen lyase, oxidoreductases acting on the CH-OH group of donors, and an ATP-binding cassette transporter. A web version of our software, SVMProt, is accessible at http://jing.cz3.nus.edu.sg/cgi-bin/svmprot.cgi .

Adenosine Triphosphate↗

Protein crystallization in the structural genomics era.

There are five broad areas where noteworthy advances have occurred in the field of macromolecular crystallization in the past 10 years, though some areas have seen the major part of those advances in only the last two years. This is largely a consequence of the international structural genomics initiative and its early results. The five areas are: (1) Physical studies and characterization of the protein crystallization process; (2) Development of new practical approaches and procedures; (3) The implementation of protein engineering by genetic means to enhance both purification and crystallization; (4) The creation of new screening conditions based on information and databases emerging from structural genomics; and (5) Development and implementation of automation, robotics, and mass screening of crystallization conditions using very small amounts of protein. A brief summary is provided here of the progress in the past few years and the influence of the structural genomics project.

Automation↗

A proposed framework for the description of plant metabolomics experiments and their results.

The study of the metabolite complement of biological samples, known as metabolomics, is creating large amounts of data, and support for handling these data sets is required to facilitate meaningful analyses that will answer biological questions. We present a data model for plant metabolomics known as ArMet (architecture for metabolomics). It encompasses the entire experimental time line from experiment definition and description of biological source material, through sample growth and preparation to the results of chemical analysis. Such formal data descriptions, which specify the full experimental context, enable principled comparison of data sets, allow proper interpretation of experimental results, permit the repetition of experiments and provide a basis for the design of systems for data storage and transmission. The current design and example implementations are freely available (http://www.armet.org/). We seek to advance discussion and community adoption of a standard for metabolomics, which would promote principled collection, storage and transmission of experiment data.

Database Management Systems↗

Looking ahead with structural genomics.

Structural genomics initiatives aim to create a library of all existing protein folds. We take a look at the progress that has been made and what more needs to be done.

Computational Biology↗

Detecting remotely related proteins by their interactions and sequence similarity.

The function of an uncharacterized protein is usually inferred either from its homology to, or its interactions with, characterized proteins. Here, we use both sequence similarity and protein interactions to identify relationships between remotely related protein sequences. We rely on the fact that homologous sequences share similar interactions, and, therefore, the set of interacting partners of the partners of a given protein is enriched by its homologs. The approach was bench-marked by assigning the fold and functional family to test sequences of known structure. Specifically, we relied on 1,434 proteins with known folds, as defined in the Structural Classification of Proteins (SCOP) database, and with known interacting partners, as defined in the Database of Interacting Proteins (DIP). For this subset, the specificity of fold assignment was increased from 54% for position-specific iterative BLAST to 75% for our approach, with a concomitant increase in sensitivity for a few percentage points. Similarly, the specificity of family assignment at the e-value threshold of 10(-8) was increased from 70% to 87%. The proposed method would be a useful tool for large-scale automated discovery of remote relationships between protein sequences, given its unique reliance on sequence similarity and protein-protein interactions.

Computational Biology↗