Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

A two-dimensional proteome map of Shigella flexneri.

Shigella flexneri is a Gram-negative facultatively intracellular pathogen responsible for bacillary dysentery in humans. In this study, extracellular proteins from the culture medium and whole cell proteins in cellular extracts of S. flexneri 2a strain 2457T were examined by two-dimensional (2-D) gel electrophoresis using immobilized pH gradient (IPG) technology. Proteins were identified by matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) in combination with Mascot search program. In total, among the 488 proteins spots processed, 388 proteins were identified. The identified proteins represented 169 genes. By comparing results of Mascot search against databases of Escherichia coli and genomes of S. flexneri 2a, one S. flexneri-specific protein was identified and one possible gap was found in 2457T genome sequences. Although this proteome map is still incomplete, it is already a useful reference for future studies involving pathogenicity, vaccine development, design of novel antibacterial drugs, etc. Proteome maps and a table of all identified proteins are available on the internet at www.proteomics.com.cn.

Bacterial Proteins↗

Comparative analysis of human intronless proteins.

The availability of the complete genome sequences of Homo sapiens together with those of taxonomically diverse organisms provides an opportunity to carry out cross-species comparison. Comparisons of protein sequences from different organisms are significant source of information as these could help in answering questions regarding the fraction of proteins that are shared by humans and organisms representing the three domains of life, viz., archaea, bacteria, and eukaryota. In the present study, a comparative analysis of the proteins encoded by intronless genes in humans was undertaken. We identified 1125 human intronless proteins that are solely present in eukaryotic lineage. More than two-thirds of these eukaryotic specific proteins appear to be mammalia specific while a small fraction of proteins are conserved in bilateria and coelomata, indicating that diversification of these proteins occurred after the divergence of the major lineages of the eukaryotic crown group. A large fraction of mammalia specific proteins are enriched in proteins responsible for transport and binding, cell envelope, and housekeeping function particularly translation. Another 228 intronless proteins are observed that do not exhibit homology to any of the proteins in the database. The distribution of human intronless proteins suggests that lineage specific expansion is one of the most important sources of organizational diversity in crown-group eukaryotes. The presence of these eukaryotic as well as human specific intronless proteins provides the foundation for rapid analysis of some of the basic processes involved in human genome.

Animals↗

The application of systems biology to drug discovery.

Recent advances in the 'omics' technologies, scientific computing and mathematical modeling of biological processes have started to fundamentally impact the way we approach drug discovery. Recent years have witnessed the development of genome-scale functional screens, large collections of reagents, protein microarrays, databases and algorithms for data and text mining. Taken together, they enable the unprecedented descriptions of complex biological systems, which are testable by mathematical modeling and simulation. While the methods and tools are advancing, it is their iterative and combinatorial application that defines the systems biology approach.

Animals↗

Comparative interactomics.

The behavior, morphology and response to stimuli in biological systems are dictated by the interactions between their components. These interactions, as we observe them now, are therefore shaped by genetic variations and selective pressure. Similar to what has been achieved by comparing genome structures and protein sequences, we hope to obtain valuable information about systems' evolution by comparing the organization of interaction networks and by analyzing their variation and conservation. Equally, significantly we can learn whether and how to extend the network information obtained experimentally in well-characterized model systems to different organisms. We conclude from our analysis that, despite the recent completion of several high throughput experiments aimed at the description of complete interactomes, the available interaction information is not yet of sufficient coverage and quality to draw any biologically meaningful conclusion from the comparison of different interactomes. Thus, the transfer of network information obtained from simple organism to evolutionary distant species should be carried out and considered with caution. By using smaller higher-confidence datasets, a larger fraction of interactions is shown to be conserved; this suggests that with the development of more accurate experimental and informatic approaches, we will soon be in the position to study the network evolution.

Animals↗

eHiTS: a new fast, exhaustive flexible ligand docking system.

The flexible ligand docking problem is divided into two subproblems: pose/conformation search and scoring function. For successful virtual screening the search algorithm must be fast and able to find the optimal binding pose and conformation of the ligand. Statistical analysis of experimental data of bound ligand conformations is presented with conclusions about the sampling requirements for docking algorithms. eHiTS is an exhaustive flexible-docking method that systematically covers the part of the conformational and positional search space that avoids severe steric clashes, producing highly accurate docking poses at a speed practical for virtual high-throughput screening. The customizable scoring function of eHiTS combines novel terms (based on local surface point contact evaluation) with traditional empirical and statistical approaches. Validation results of eHiTS are presented and compared to three other docking software on a set of 91 PDB structures that are common to the validation sets published for the other programs.

Algorithms↗

Functional Analysis of MS-Based Proteomics Data: From Protein Groups to Networks.

Mass spectrometry-based proteomics allows the quantification of thousands of proteins, protein variants, and their modifications, in many biological samples. These are derived from the measurement of peptide relative quantities, and it is not always possible to distinguish proteins with similar sequences due to the absence of protein-specific peptides. In such cases, peptide signals are reported in protein groups that can correspond to several genes. Here, we show that multi-gene protein groups have a limited impact on GO-term enrichment, but selecting only one gene per group affects network analysis. We thus present the Cytoscape app Proteo Visualizer (https://apps.cytoscape.org/apps/ProteoVisualizer) that is designed for retrieving protein interaction networks from STRING using protein groups as input and thus allows visualization and network analysis of bottom-up MS-based proteomics data sets.

Proteomics↗

Metalloproteomics: high-throughput structural and functional annotation of proteins in structural genomics.

A high-throughput method for measuring transition metal content based on quantitation of X-ray fluorescence signals was used to analyze 654 proteins selected as targets by the New York Structural GenomiX Research Consortium. Over 10% showed the presence of transition metal atoms in stoichiometric amounts; these totals as well as the abundance distribution are similar to those of the Protein Data Bank. Bioinformatics analysis of the identified metalloproteins in most cases supported the metalloprotein annotation; identification of the conserved metal binding motif was also shown to be useful in verifying structural models of the proteins. Metalloproteomics provides a rapid structural and functional annotation for these sequences and is shown to be approximately 95% accurate in predicting the presence or absence of stoichiometric metal content. The project's goal is to assay at least 1 member from each Pfam family; approximately 500 Pfam families have been characterized with respect to transition metal content so far.

Binding Sites↗

Bioinformatics in glycobiology.

In comparison with genes and proteins, attention paid to oligosaccharides that modify proteins is still marginal. Accordingly, bioinformatics is so far poorly involved in glycobiology. Some initiatives have been taken, however, to collect in databases all glycobiology-relevant information or to design specific data mining algorithms to infer predictions or identify oligosaccharide structures. In this review, we make a non-exhaustive survey of the available glycobiology-related bioinformatic resources, focussing mainly on those resources that are available through the World Wide Web. Some well-curated databases are identified, but the development of specialised algorithms appears to be limited.

Algorithms↗

Chemoproteomics as a basis for post-genomic drug discovery.

The large number of small organic compounds now available for drug-lead screening has led to numerous methods for classifying molecular similarity and diversity, the aim being to restore a balance between the quantity and drug-like quality of compounds in small-molecule libraries. Whereas structural and physicochemical attributes continue to be emphasized in compound selection for drug-lead screening, chemoproteomics--the use of biological information to guide chemistry--offers a highly efficient alternative to small-molecule characterization that can accelerate drug discovery in the post-genomic era.

Databases, Factual↗

Viewing and annotating sequence data with Artemis.

Artemis is a widely used software tool for annotating and viewing sequence data. No database is required to use Artemis. Instead, individual sequence data files can be analysed with little or no formatting, making it particularly suited to the study of small genomes and chromosomes, and straightforward for a novice user to get started. Since its release in 1999, Artemis has been used to annotate a diverse collection of prokaryotic and eukaryotic genomes, ranging from Streptomyces coelicolor to, more recently, a large proportion of the Plasmodium falciparum genome. Artemis allows annotated genomes to be easily browsed and makes it simple to add useful biological information to raw sequence data. This paper gives an overview of some of the features of Artemis and includes how it facilitates manual gene prediction and can provide an overview of entire chromosomes or small compact genomes--useful for uncovering unusual features such as pathogenicity islands.

Animals↗

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics↗

Exploring protein fold space by secondary structure prediction using data distribution method on Grid platform.

MOTIVATION: Since the newly developed Grid platform has been considered as a powerful tool to share resources in the Internet environment, it is of interest to demonstrate an efficient methodology to process massive biological data on the Grid environments at a low cost. This paper presents an efficient and economical method based on a Grid platform to predict secondary structures of all proteins in a given organism, which normally requires a long computation time through sequential execution, by means of processing a large amount of protein sequence data simultaneously. From the prediction results, a genome scale protein fold space can be pursued. RESULTS: Using the improved Grid platform, the secondary structure prediction on genomic scale and protein topology derived from the new scoring scheme for four different model proteomes was presented. This protein fold space was compared with structures from the Protein Data Bank, database and it showed similarly aligned distribution. Therefore, the fold space approach based on this new scoring scheme could be a guideline for predicting a folding family in a given organism.

Computing Methodologies↗

VAMP: visualization and analysis of array-CGH, transcriptome and other molecular profiles.

MOTIVATION: Microarray-based CGH (Comparative Genomic Hybridization), transcriptome arrays and other large-scale genomic technologies are now routinely used to generate a vast amount of genomic profiles. Exploratory analysis of this data is crucial in helping to understand the data and to help form biological hypotheses. This step requires visualization of the data in a meaningful way to visualize the results and to perform first level analyses. RESULTS: We have developed a graphical user interface for visualization and first level analysis of molecular profiles. It is currently in use at the Institut Curie for cancer research projects involving CGH arrays, transcriptome arrays, SNP (single nucleotide polymorphism) arrays, loss of heterozygosity results (LOH), and Chromatin ImmunoPrecipitation arrays (ChIP chips). The interface offers the possibility of studying these different types of information in a consistent way. Several views are proposed, such as the classical CGH karyotype view or genome-wide multi-tumor comparison. Many functionalities for analyzing CGH data are provided by the interface, including looking for recurrent regions of alterations, confrontation to transcriptome data or clinical information, and clustering. Our tool consists of PHP scripts and of an applet written in Java. It can be run on public datasets at http://bioinfo.curie.fr/vamp AVAILABILITY: The VAMP software (Visualization and Analysis of array-CGH,transcriptome and other Molecular Profiles) is available upon request. It can be tested on public datasets at http://bioinfo.curie.fr/vamp. The documentation is available at http://bioinfo.curie.fr/vamp/doc.

Algorithms↗

GOlorize: a Cytoscape plug-in for network visualization with Gene Ontology-based layout and coloring.

UNLABELLED: We have implemented a graph layout algorithm that exposes Gene Ontology (GO) class structure on the network nodes. It can be used in conjunction with BiNGO plug-in to Cytoscape, which finds the GO categories over-represented in a given network. Our plug-in, named GOlorize, first highlights the class members with category-specific color-coding and then constructs an enhanced visualization of the network using a class-directed layout algorithm. AVAILABILITY: http://www.cytoscape.org/plugins2.php. SUPPLEMENTARY INFORMATION: Installation instructions and tutorial at http://www.cytoscape.org/plugins/GOlorize/GOlorizeUserGuide.pdf.

Algorithms↗