Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

Human Protein Atlas charts a diverse terrain.

A complete set of high-quality antibodies against the human proteome would constitute a remarkable resource for the analysis of protein abundances in tissues, facilitating systematic protein expression profiling in health and disease. In two recent papers, Swedish researchers describe a concerted effort towards producing a 'Human Protein Atlas' by combining high-throughput antibody generation with immunohistochemical profiling of tissue microarrays.

Database Management Systems↗

Exhaustive assignment of compositional bias reveals universally prevalent biased regions: analysis of functional associations in human and Drosophila.

BACKGROUND: Compositionally biased (CB) regions are stretches in protein sequences made from mainly a distinct subset of amino acid residues; such regions are frequently associated with a structural role in the cell, or with protein disorder. RESULTS: We derived a procedure for the exhaustive assignment and classification of CB regions, and have applied it to thirteen metazoan proteomes. Sequences are initially scanned for the lowest-probability subsequences (LPSs) for single amino-acid types; subsequently, an exhaustive search for lowest probability subsequences (LPSs) for multiple residue types is performed iteratively until convergence, to define CB region boundaries. We analysed > 40,000 CB regions with > 20 million residues; strikingly, nine single-/double- residue biases are universally abundant, and are consistently highly ranked across both vertebrates and invertebrates. To home in subpopulations of CB regions of interest in human and D. melanogaster, we analysed CB region lengths, conservation, inferred functional categories and predicted protein disorder, and filtered for coiled coils and protein structures. In particular, we found that some of the universally abundant CB regions have significant associations to transcription and nuclear localization in Human and Drosophila, and are also predicted to be moderately or highly disordered. Focussing on Q-based biased regions, we found that these regions are typically only well conserved within mammals (appearing in 60-80% of orthologs), with shorter human transcription-related CB regions being unconserved outside of mammals; they are also preferentially linked to protein domains such as the homeodomain and glucocorticoid-receptor DNA-binding domain. In general, only approximately 40-50% of residues in these human and Drosophila CB regions have predicted protein disorder. CONCLUSION: This data is of use for the further functional characterization of genes, and for structural genomics initiatives.

Animals↗

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein↗

A quantitative analysis of secondary RNA structure using domination based parameters on trees.

BACKGROUND: It has become increasingly apparent that a comprehensive database of RNA motifs is essential in order to achieve new goals in genomic and proteomic research. Secondary RNA structures have frequently been represented by various modeling methods as graph-theoretic trees. Using graph theory as a modeling tool allows the vast resources of graphical invariants to be utilized to numerically identify secondary RNA motifs. The domination number of a graph is a graphical invariant that is sensitive to even a slight change in the structure of a tree. The invariants selected in this study are variations of the domination number of a graph. These graphical invariants are partitioned into two classes, and we define two parameters based on each of these classes. These parameters are calculated for all small order trees and a statistical analysis of the resulting data is conducted to determine if the values of these parameters can be utilized to identify which trees of orders seven and eight are RNA-like in structure. RESULTS: The statistical analysis shows that the domination based parameters correctly distinguish between the trees that represent native structures and those that are not likely candidates to represent RNA. Some of the trees previously identified as candidate structures are found to be "very" RNA like, while others are not, thereby refining the space of structures likely to be found as representing secondary RNA structure. CONCLUSION: Search algorithms are available that mine nucleotide sequence databases. However, the number of motifs identified can be quite large, making a further search for similar motif computationally difficult. Much of the work in the bioinformatics arena is toward the development of better algorithms to address the computational problem. This work, on the other hand, uses mathematical descriptors to more clearly characterize the RNA motifs and thereby reduce the corresponding search space. These preliminary findings demonstrate that graph-theoretic quantifiers utilized in fields such as computer network design hold significant promise as an added tool for genomics and proteomics.

Algorithms↗

The wheat (Triticum aestivum L.) leaf proteome.

The wheat leaf proteome was mapped and partially characterized to function as a comparative template for future wheat research. In total, 404 proteins were visualized, and 277 of these were selected for analysis based on reproducibility and relative quantity. Using a combination of protein and expressed sequence tag database searching, 142 proteins were putatively identified with an identification success rate of 51%. The identified proteins were grouped according to their functional annotations with the majority (40%) being involved in energy production, primary, or secondary metabolism. Only 8% of the protein identifications lacked ascertainable functional annotation. The 51% ratio of successful identification and the 8% unclear functional annotation rate are major improvements over most previous plant proteomic studies. This clearly indicates the advancement of the plant protein and nucleic acid sequence and annotation data available in the databases, and shows the enhanced feasibility of future wheat leaf proteome research.

Computational Biology↗

Proteomic analysis of Acinetobacter lwoffii K24 by 2-D gel electrophoresis and electrospray ionization quadrupole-time of flight mass spectrometry.

The MS/MS analysis by Electrospray ionization quadrupole-time of flight mass spectrometry (ESI-Q-TOF MS) was applied to identify proteins in proteome analysis of bacteria whose genomes are not known. The protein identification by ESI-Q-TOF MS was performed sequentially by database search and then de novo sequencing using MS/MS spectra. Soil bacteria having unanalyzed genome, Acinetobacter lwoffii K24 is an aniline degrading bacterium. In this report, we present the results of a comparison between the proteome profile of A. lwoffii K24 cultured in aniline- or succinate-containing media. Protein analysis was performed using two-dimensional gel electrophoresis (2-DE) with pH 3-10 immobilized pH gradient (IPG) strips followed by ESI-Q-TOF MS. More than 780 protein spots were detected by 2-DE from the soluble proteome. Forty-eight of these proteins were expressed exclusively in aniline cultured bacteria, and 81 proteins increased and 162 proteins decreased in aniline-cultured versus succinate cultured A. lwoffii K24. Internal amino acid sequences of 43 major protein spots were successfully determined by ESI-Q-TOF MS to try to identify the bacterial proteins responding to aniline culture condition. Since the A. lwoffii K24 genome is not yet sequenced, many proteins were found to be hypothetical. Comparative proteome analysis of the insoluble protein fractions showed that one novel protein that was strongly induced by succinate-cultured A. lwoffii K24 was repressed under aniline culture conditions. These results suggest that comprehensive analysis of bacterial proteomes by 2-DE and amino acid sequence analysis by ESI-Q-TOF MS is useful for understanding induced novel proteins of biodegrading bacteria.

Acinetobacter↗

Mass spectrometric analysis of the Schistosoma mansoni tegumental sub-proteome.

Schistosoma mansoni is a parasitic worm that lives in the blood vessels of its host. We mapped the S. mansoni tegumental outer-surface structure proteome by 1D SDS-PAGE and LC-MS/MS and an EST-database from the ongoing genome-sequencing project. We identified 740 proteins of which 43 were tegument-specific. Many of these proteins show no homology to any nonschistosomal protein, demonstrating that the schistosomal outer-surface comprises specific and unique proteins, likely to be critical for parasite survival.

Animals↗

Diagnosis of cellular states of microbial organisms using proteomics.

Two-dimensional (2-D) polyacrylamide gel electrophoresis has much to contribute to experimental analysis of the proteomes of microbial organisms, since this method separates most cellular proteins and allows synthesis rates to be determined quantitatively. Databases generated using 2-D gels can grow to be very large from even just a few experiments, since each sample provides the data for a field (or column) in the database for several hundreds to even thousands of records (or rows), each of which represents a single polypeptide species. The value of such databases for generating an encyclopedia of how each of the cell's proteins behave in different conditions (protein phenotypes) has been recognized for some time. The potential exists, however, to glean even more valuable information from such databases. Because the measurements of each protein are made in the context of all other proteins, a comprehensive glimpse of the cell's physiological state is theoretically achievable with each 2-D gel. By examining enough conditions (and 2-D gels), expression patterns of subsets of proteins (proteomic signatures) can be found that correlate with the cell's state. This type of information can provide a unique contribution to proteomic analysis, and should be a major focus of such analyses.

Genome, Bacterial↗

Combining functional and topological properties to identify core modules in protein interaction networks.

Advances in large-scale technologies in proteomics, such as yeast two-hybrid screening and mass spectrometry, have made it possible to generate large Protein Interaction Networks (PINs). Recent methods for identifying dense sub-graphs in such networks have been based solely on graph theoretic properties. Therefore, there is a need for an approach that will allow us to combine domain-specific knowledge with topological properties to generate functionally relevant sub-graphs from large networks. This article describes two alternative network measures for analysis of PINs, which combine functional information with topological properties of the networks. These measures, called weighted clustering coefficient and weighted average nearest-neighbors degree, use weights representing the strengths of interactions between the proteins, calculated according to their semantic similarity, which is based on the Gene Ontology terms of the proteins. We perform a global analysis of the yeast PIN by systematically comparing the weighted measures with their topological counterparts. To show the usefulness of the weighted measures, we develop an algorithm for identification of functional modules, called SWEMODE (Semantic WEights for MODule Elucidation), that identifies dense sub-graphs containing functionally similar proteins. The proposed method is based on the ranking of nodes, i.e., proteins, according to their weighted neighborhood cohesiveness. The highest ranked nodes are considered as seeds for candidate modules. The algorithm then iterates through the neighborhood of each seed protein, to identify densely connected proteins with high functional similarity, according to the chosen parameters. Using a yeast two-hybrid data set of experimentally determined protein-protein interactions, we demonstrate that SWEMODE is able to identify dense clusters containing proteins that are functionally similar. Many of the identified modules correspond to known complexes or subunits of these complexes.

Algorithms↗

Efficient peptide mapping and its application to identify embryo proteins in rice proteome analysis.

Using direct N-terminal analysis, only 31 N-terminally unblocked proteins out of 100 rice embryo proteins could be identified. To obtain protein sequence information for the remaining 69 blocked proteins, we developed a simple, efficient and rapid method. Using this method, we determined the peptide maps of 20 proteins per day in 10 pmol amounts. Applying this method to rice proteome analysis, we determined the internal sequences of all 69 blocked proteins. A total of 28 proteins out of 100 analyzed showed sequence similarity to the proteins with known functions in the SWISS-PROT and NCBI databases. Alternatively, we also used peptide mass fingerprinting determined by matrix assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) to identify the rice proteins separated by two-dimensional electrophoresis (2-DE). Although peptide-mass fingerprinting is a high-throughput method, we could not easily identify all the rice proteins or genes by this method, because the complete database information on rice, is not yet available and many proteins are post-translationally modified. Therefore, at present, the improved peptide mapping method as we report here is considered to be very useful in rice proteome analysis, especially for blocked proteins.

Databases, Protein↗

Phosphoproteomics toolbox: computational biology, protein chemistry and mass spectrometry.

Protein phosphorylation is important for regulation of most biological functions and up to 50% of all proteins are thought to be modified by protein kinases. Increased knowledge about potential phosphorylation of a protein may increase our understanding of the molecular processes in which it takes part. Despite the importance of protein phosphorylation, identification of phosphoproteins and localization of phosphorylation sites is still a major challenge in proteomics. However, high-throughput methods for identification of phosphoproteins are being developed, in particular within the fields of bioinformatics and mass spectrometry. In this review, we present a toolbox of current technology applied in phosphoproteomics including computational prediction, chemical approaches and mass spectrometry-based analysis, and propose an integrated strategy for experimental phosphoproteomics.

Computational Biology↗

SUBA: the Arabidopsis Subcellular Database.

Knowledge of protein localisation contributes towards our understanding of protein function and of biological inter-relationships. A variety of experimental methods are currently being used to produce localisation data that need to be made accessible in an integrated manner. Chimeric fluorescent fusion proteins have been used to define subcellular localisations with at least 1100 related experiments completed in Arabidopsis. More recently, many studies have employed mass spectrometry to undertake proteomic surveys of subcellular components in Arabidopsis yielding localisation information for approximately 2600 proteins. Further protein localisation information may be obtained from other literature references to analysis of locations (AmiGO: approximately 900 proteins), location information from Swiss-Prot annotations (approximately 2000 proteins); and location inferred from gene descriptions (approximately 2700 proteins). Additionally, an increasing volume of available software provides location prediction information for proteins based on amino acid sequence. We have undertaken to bring these various data sources together to build SUBA, a SUBcellular location database for Arabidopsis proteins. The localisation data in SUBA encompasses 10 distinct subcellular locations, >6743 non-redundant proteins and represents the proteins encoded in the transcripts responsible for 51% of Arabidopsis expressed sequence tags. The SUBA database provides a powerful means by which to assess protein subcellular localisation in Arabidopsis (http://www.suba.bcs.uwa.edu.au).

Arabidopsis Proteins↗

Proteomic data exchange and storage: the need for common standards and public repositories.

The ever increasing volumes of proteomic data now being produced by laboratories across the world have resulted in major issues in data storage and accessibility. The further demands of multilaboratory initiatives has highlighted issues when collaborators cannot import data generated within the same project but generated by different hardware types and processed by laboratory-specific work flows and analyses packages. There is an increasing need for common data standards that will allow the interchange of data between different instrumentation, search engines, and between laboratory databases. This could then lead to the establishment of data repositories from where benchmark datasets could be accessed and reanalyzed. The Human Proteome Organization is currently supporting efforts to establish such standards. The work of the Proteomics Standards Initiative has lead to the development of the mzData XML interchange standard and is now broadening its scope to produce a spectral analysis output format, mzIdent. Accompanying controlled vocabularies allow the accurate, while systematic, representation of metadata throughout both schema.

Benchmarking↗

Proteomics: posttranslational modifications, immune responses and current analytical tools.

The publication of the human genome sequence enables most of the still unknown protein sequences to be added to the current databases. A sequence alone does not, however, give information about the possible expression level of the corresponding protein, neither does it inform about the possible posttranslational modifications, like phosphorylation, glycosylation or changes in individual amino acids. Thus, the human proteome project, a large scale analysis of the functions of gene products, will have an enormous impact on our understanding of the biochemistry of proteins, processes and pathways they are involved in. The diversity in proteins is considerably expanded by various posttranslational modifications. These also pose problems to the investigators, but their careful analysis often pays back because they can reveal important properties in proteins or peptides--like an increased antigenicity leading to (auto)immune responses or an active form of a signaling protein. Immune tolerance usually exists towards self-proteins, but in specific cases it may be broken by posttranslational modifications in the proteins. Novel mass spectrometric, affinity and display techniques offer valuable tools for the large-scale analysis of proteomes. In the present paper we discuss their use for the detection of posttranslational modifications, functional interactions and possible disease-associated abnormalities in proteins.

Electrophoresis, Gel, Two-Dimensional↗

The role of informatics in glycobiology research with special emphasis on automatic interpretation of MS spectra.

This paper reviews the current status of bioinformatics applications and databases in glycobiology, which are based on bioinformatics approaches as well as informatics for glycobiology where an explicit encoding of glycan structures is required. The availability of the complete sequence of the human genome has accelerated the systematic identification of so far unidentified glycogenes considerably in many areas of glycobiology using well-established bioinfomatics tools. Although there has been an immense development of new glyco-related data collections as well as informatics tools and several efforts have been started to cross-link and reference the various data deposited in distributed databases, informatics for glycobiology and glycomics is still poorly developed compared to the genomics and proteomics area. The development of algorithms for the automatic interpretation of MS spectra - currently, a severe bottleneck, which hampers the rapid and reliable interpretation of MS data in high-throughput glycomics projects - is reviewed. A comprehensive list of web resources is given. Several lines of progression are discussed. There is an urgent need for the development of decentralised input facilities of experimentally determined glycan structures. Simultaneously, agreements of standards for the structural description of glycans as well as formats for the related data have to be established. The integration of glycomics with genomics/proteomics has to increase.

Computational Biology↗

Comparative proteomic analysis of extracellular proteins of enterohemorrhagic and enteropathogenic Escherichia coli strains and their ihf and ler mutants.

Enterohemorrhagic and enteropathogenic Escherichia coli (EHEC and EPEC, respectively) strains are closely related human pathogens that are responsible for food-borne epidemics in many countries. Integration host factor (IHF) and the locus of enterocyte effacement-encoded regulator (Ler) are needed for the expression of virulence genes in EHEC and EPEC, including the elicitation of actin rearrangements for attaching and effacing lesions. We applied a proteomic approach, using two-dimensional polyacrylamide gel electrophoresis in combination with matrix-assisted laser desorption ionization-time of flight mass spectrometry and a protein database search, to analyze the extracellular protein profiles of EHEC EDL933, EPEC E2348/69, and their ihf and ler mutants. Fifty-nine major protein spots from the extracellular proteomes were identified, including six proteins of unknown function. Twenty-six of them were conserved between EHEC EDL933 and EPEC E2348/69, while some of them were strain-specific proteins. Four common extracellular proteins (EspA, EspB, EspD, and Tir) were regulated by both IHF and Ler in EHEC EDL933 and EPEC E2348/69. TagA in EHEC EDL933 and EspC and EspF in EPEC E2348/69 were present in the wild-type strains but absent from their respective ler and ihf mutants, while FliC was overexpressed in the ihf mutant of EPEC E2348/69. Two dominant forms of EspB were found in EHEC EDL933 and EPEC E2348/69, but the significance of this is unknown. These results show that proteomics is a powerful platform technology for accelerating the understanding of EPEC and EHEC pathogenesis and identifying markers for laboratory diagnoses of these pathogens.

Escherichia coli↗