Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Proteomics techniques and their application to hematology.

The recent sequencing of a number of genomes has raised the level of opportunities for studies on proteins. This area of research has been described with the all-embracing term, proteomics. In proteomics, the use of mass spectrometric techniques enables genomic databases to be used to establish the identity of proteins with relatively little data, compared to the era before genome sequencing. The use of related analytical techniques also offers the opportunity to gain information on regulation, via posttranslational modification, and potential new diagnostic and prognostic indicators. Relative quantification of proteins and peptides in cellular and extracellular material remains a challenge for proteomics and mass spectrometry. This review presents an analysis of the present and future impact of these proteomic technologies with emphasis on relative quantification for hematologic research giving an appraisal of their potential benefits.

Animals↗

Do we want our data raw? Including binary mass spectrometry data in public proteomics data repositories.

With the human Plasma Proteome Project (PPP) pilot phase completed, the largest and most ambitious proteomics experiment to date has reached its first milestone. The correspondingly impressive amount of data that came from this pilot project emphasized the need for a centralized dissemination mechanism and led to the development of a detailed, PPP specific data gathering infrastructure at the University of Michigan, Ann Arbor as well as the protein identifications database project at the European Bioinformatics Institute as a general proteomics data repository. One issue that crept up while discussing which data to store for the PPP concerns whether the raw, binary data coming from the mass spectrometers should be stored, or rather the more compact and already significantly processed peak lists. As this debate is not restricted to the PPP but relates to the proteomics community in general, we will attempt to detail the relative merits and caveats associated with centralized storage and dissemination of raw data and/or peak lists, building on the extensive experience gained during the PPP pilot phase. Finally, some suggestions are made for both immediate and future storage of MS data in public repositories.

Computational Biology↗

GWFASTA: server for FASTA search in eukaryotic and microbial genomes.

Similarity searches are a powerful method for solving important biological problems such as database scanning, evolutionary studies, gene prediction, and protein structure prediction. FASTA is a widely used sequence comparison tool for rapid database scanning. Here we describe the GWFASTA server that was developed to assist the FASTA user in similarity searches against partially and/or completely sequenced genomes. GWFASTA consists of more than 60 microbial genomes, eight eukaryote genomes, and proteomes of annotatedgenomes. Infact, it provides the maximum number of databases for similarity searching from a single platform. GWFASTA allows the submission of more than one sequence as a single query for a FASTA search. It also provides integrated post-processing of FASTA output, including compositional analysis of proteins, multiple sequences alignment, and phylogenetic analysis. Furthermore, it summarizes the search results organism-wise for prokaryotes and chromosome-wise for eukaryotes. Thus, the integration of different tools for sequence analyses makes GWFASTA a powerful toolfor biologists.

Computer Systems↗

Analysis of the Candida albicans proteome. II. Protein information technology on the Net (update 2002).

Candida albicans is an important fungal model organism of noteworthy clinical interest in modern medicine. Different initiatives addressing its sequencing and physical mapping have been carried out. The C. albicans genome sequence is currently near to completion at Stanford University, heralding new challenges in proteomic research and functional analyses of its gene products. This review presents an update of the most relevant data resources that are available through the World Wide Web to scientists working in the area of the analysis of the C. albicans proteome. An overview of the current status of the main universal protein sequence databases and specialized data collections for C. albicans is given. Various issues of the single public C. albicans 2D-PAGE database are also described, highlighting the significance of setting up graphical query interface-based databanks to visualize 2D-PAGE images through the Net. Finally, we also emphasize the pressing need to create a "cyber-bioknowledge library" that will integrate all the databases developed at the different levels for the understanding of life processes as well as bioinformatic tools for interpreting this deluge of data generated through the Internet.

Candida albicans↗

Mass spectrometry. From genomics to proteomics.

Large-scale DNA sequencing has stimulated the development of proteomics by providing a sequence infrastructure for protein analysis. Rapid and automated protein identification can be achieved by searching protein and nucleotide sequence databases directly with data generated by mass spectrometry. A high-throughput and large-scale approach to identifying proteins has been the result. These technological changes have advanced protein expression studies and the identification of proteins in complexes, two types of studies that are essential in deciphering the networks of proteins that are involved in biological processes.

Databases, Factual↗

The Universal Protein Resource (UniProt): an expanding universe of protein information.

The Universal Protein Resource (UniProt) provides a central resource on protein sequences and functional annotation with three database components, each addressing a key need in protein bioinformatics. The UniProt Knowledgebase (UniProtKB), comprising the manually annotated UniProtKB/Swiss-Prot section and the automatically annotated UniProtKB/TrEMBL section, is the preeminent storehouse of protein annotation. The extensive cross-references, functional and feature annotations and literature-based evidence attribution enable scientists to analyse proteins and query across databases. The UniProt Reference Clusters (UniRef) speed similarity searches via sequence space compression by merging sequences that are 100% (UniRef100), 90% (UniRef90) or 50% (UniRef50) identical. Finally, the UniProt Archive (UniParc) stores all publicly available protein sequences, containing the history of sequence data with links to the source databases. UniProt databases continue to grow in size and in availability of information. Recent and upcoming changes to database contents, formats, controlled vocabularies and services are described. New download availability includes all major releases of UniProtKB, sequence collections by taxonomic division and complete proteomes. A bibliography mapping service has been added, and an ID mapping service will be available soon. UniProt databases can be accessed online at http://www.uniprot.org or downloaded at ftp://ftp.uniprot.org/pub/databases/.

Databases, Protein↗

Proteomics.

Explore the source record for details and available documents.

Databases, Protein↗

C. elegans: an invaluable model organism for the proteomics studies of the cholesterol-mediated signaling pathway.

With the availability of its complete genome sequence and unique biological features relevant to human disease, Caenorhabditis elegans has become an invaluable model organism for the studies of proteomics, leading to the elucidation of nematode gene function. A journey from the genome to proteome of C. elegans may begin with preparation of expressed proteins, which enables a large-scale analysis of all possible proteins expressed under specific physiological conditions. Although various techniques have been used for proteomic analysis of C. elegans, systematic high-throughput analysis is still to come in order to accommodate studies of post-translational modification and quantitative analysis. Given that no integrated C. elegans protein expression database is available, it is about time that a global C. elegans proteome project is launched through which datasets of transcriptomes, protein-protein interaction and functional annotation can be integrated. As an initial target of a pilot project of the C. elegans proteome project, the cholesterol-mediated signaling pathway will be an excellent example since, like in other organisms, it is one of the key controlling pathways in cell growth and development in C. elegans. As this field tends to broaden to functional proteomics, there is a high demand to develop the versatile proteome informatics tools that can mange many different data in an integrative manner.

Animals↗

From biological databases to platforms for biomedical discovery.

The use of high-throughput DNA sequencing and proteomic methods has led to an unprecedented increase in the amount of genomic and proteomic data. Application of computing technologies and development of computational tools to analyze and present these data has not kept pace with the accumulation of information. Here, we discuss the use of different database systems to store biological information and mention some of the key emerging computing technologies that are likely to have a key role in the future of bioinformatics.

Algorithms↗

Comparative proteome analysis of Helicobacter pylori.

Helicobacter pylori, the causative agent of gastritis, ulcer and stomach carcinoma, infects approximately half of the worlds population. After sequencing the complete genome of two strains, 26695 and J99, we have approached the demanding task of investigating the functional part of the genetic information containing macromolecules, the proteome. The proteins of three strains of H. pylori, 26695 and J99, and a prominent strain used in animal models SS1, were separated by a high-resolution two-dimensional electrophoresis technique with a resolution power of 5000 protein spots. Up to 1800 protein species were separated from H. pylori which had been cultivated for 5 days on agar plates. Using matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) peptide mass fingerprinting we have identified 152 proteins, including nine known virulence factors and 28 antigens. The three strains investigated had only a few protein spots in common. We observe that proteins with an amino acid exchange resulting in a net change of only one charge are shifted in the two-dimensional electrophoresis (2-DE) pattern. The expression of 27 predicted conserved hypothetical open reading frames (ORFs) and six unknown ORFs were confirmed. The growth conditions of the bacteria were shown to have an effect on the presence of certain proteins. A preliminary immunoblotting study using human sera revealed that this approach is ideal for identifying proteins of diagnostic or therapeutic value. H. pylori 2-DE patterns with their identified protein species were added to the dynamic 2D-PAGE database (http://www.mpiib-berlin.mpg.de/2D-PAGE/). This basic knowledge of the proteome in the public domain will be an effective instrument for the identification of new virulence or pathogenic factors, and antigens of potentially diagnostic or curative value against H. pylori.

Bacterial Proteins↗

[Differential proteomic analysis of human lung adenocarcinoma cell line A-549 and of normal cell line HBE].

To explore the differential proteomic expressions between human lung adenocarcinoma cell line A-549 and normal cell line HBE, a series of methods, including immobilized pH gradient-two dimensional polyacrylamide gel electrophoresis, silver staining, PDQuest 2-DE software analysis, peptide mass fingerprinting based on matrix-assisted laser desorption/ionization time of flight mass spectrometry (MALDI-TOF-MS) and SWISS-PROT database searching, were used to separate and identify the differential proteomic expressions between A-549 and HBE. The results showed that the good 2-DE pattern including high resolution and reproducibility was obtained. After silver staining, the 2-DE image analysis by PDQuest 2-DE software detected average (890 +/- 38) spots in A-549, and (757 +/- 27) spots in HBE. The average positional deviation of the matched spots between A-549 and HBE 2-DE maps was (2.85 +/- 0.48) mm in IEF direction, and (2.69 +/- 0.37) mm in SDS-PAGE direction. The differential proteomic expression analysis found that there were 535 matched spots between A-549 and HBE 2-DE maps, 355 spots that were not matched in A-549, 222 spots that were not matched in HBE. 18 differential spots (8 spots in A-549 and 10 spots in HBE) were cut off from silver staining gel at random, digested in gel with TPCK-trypsin, measured with MALDI-TOF-MS and searched in the SWISS-PROT database with PeptIdent software. 18 protein were preliminarily identified. These proteins were related to cell signal transduction, cell metabolism, proliferation and differentiation etc. There was a significant difference at protein level between human lung adenocarcinoma cell line A-549 and normal cell line HBE. It suggests that the differential expression analysis of proteomes may be useful to further study of the related proteins and the molecular markers of lung adenocarcinoma.

Adenocarcinoma↗

Multiple separations facilitate identification of protein variants by mass spectrometry.

Identification of variant proteins from complex biological samples promises to contribute much to our understanding the etiology of pathological states. Characterization of variants, either due to genetic mutations in protein sequences or to post-translational modifications, is considerably more difficult than the simple protein identifications typical of most current proteomic investigations. Identification of a few peptides by database retrieval is not adequate when the goal is to have a complete understanding of the modifications of the protein. Although one advantage of mass spectrometry is its ability to obtain specific responses to several components, the complexity of biological samples is often overwhelming, resulting in spectra lacking useful information. For complex mixtures, isolation procedures before mass spectrometric analysis may need to include a variety of chromatographic and electrophoretic separation techniques. In this report, we illustrate how several preparative steps were essential for obtaining information about modified human lens beta-crystallins. The preparative techniques prior to mass spectrometry included size exclusion chromatography, reversed phase chromatography, two-dimensional gel electrophoresis, in situ digestion of the proteins and peptide trapping and washing before a final reversed phase high performance liquid chromatographic separation on-line to the mass spectrometer. This approach for isolation and analysis, when customized for other proteins, should find application in many studies where protein variants of complex mixtures are to be identified.

Chromatography, Gel↗

Functional and topological characterization of protein interaction networks.

The elucidation of the cell's large-scale organization is a primary challenge for post-genomic biology, and understanding the structure of protein interaction networks offers an important starting point for such studies. We compare four available databases that approximate the protein interaction network of the yeast, Saccharomyces cerevisiae, aiming to uncover the network's generic large-scale properties and the impact of the proteins' function and cellular localization on the network topology. We show how each database supports a scale-free, topology with hierarchical modularity, indicating that these features represent a robust and generic property of the protein interactions network. We also find strong correlations between the network's structure and the functional role and subcellular localization of its protein constituents, concluding that most functional and/or localization classes appear as relatively segregated subnetworks of the full protein interaction network. The uncovered systematic differences between the four protein interaction databases reflect their relative coverage for different functional and localization classes and provide a guide for their utility in various bioinformatics studies.

Algorithms↗

Protein production and crystallization at SECSG -- an overview.

Using a high degree of automation, the Southeast Collaboratory for Structural Genomics (SECSG) has developed high throughput pipelines for protein production, and crystallization using a two-tiered approach. Primary, or tier-1, protein production focuses on producing proteins for members of large Pfam families that lack a representative structure in the Protein Data Bank. Target genomes are Pyrococcus furiosus and Caenorhabditis elegans. Selected human proteins are also under study. Tier-2 protein production, or target rescue, focuses on those tier-1 proteins, which either fail to crystallize or give poorly diffracting crystals. This two tier approach is more efficient since it allows the primary protein production groups to focus on the production of new targets while the tier-2 efforts focus on providing additional sample for further studies and modified protein for structure determination. Both efforts feed the SECSG high throughput crystallization pipeline, which is capable of screening over 40 proteins per week. Details of the various pipelines in use by the SECSG for protein production and crystallization, as well as some examples of target rescue are described.

Animals↗

A description scheme of biological processes based on elementary bricks of action.

With the fast growth of high-throughput strategies in Biology, there is a strong need to accelerate knowledge acquisition and organization of molecular functions. Unfortunately, although we know that there is a correlation between protein molecules and their functions, we are unable to clearly identify this link. Here, we revisit the current views of protein functions as well as their annotation, and we show that they are incompatible with unambiguous interpretations and the use of this knowledge. We describe herein a description scheme for biological processes based on elementary bricks of action that may be associated with biological molecules. To retrieve the descriptive quality found in annotations of other kinds of biological data, it was decided to develop a scheme involving four levels of abstraction: Basic Elements of Action, Biological Activities, Biological Functionalities and Biological Roles. This multi-level organization is a generic method; it allows for a description of biological processes by using a limited number of elementary bricks of action. Moreover, by using this description of biological processes, it should now be possible to clearly identify unambiguous relationships between the organization of biological processes and the structural or functional organizations of biological molecules.

Algorithms↗

Genomics in environmental health research--opportunities and challenges.

Environmental health research impacts both environmental health regulatory policy and the practice of medicine. However, this area of medical research has not garnered public support and attention of medical researchers because of its emphasis on prevention and public health. Also, the pervasiveness of a scientific culture wedded to old problems and outdated technologies and models systems has not been helpful in generating enthusiasm for the field. While the emphasis on prevention is both laudable and appropriate, the adoption of cutting-edge technologies to exploit the new scientific opportunities, made possible by the nation's investment in genomics, is essential if the discipline expects to be competitive with other highly deserving programs. The new 'omics' era of environmental health research, ushered in over the past decade, characterized by the linkage of genomics, proteomics and metabolomics to conventional toxicology and pathology databases, holds great promise for elucidating mechanisms of gene-environment interaction in human health and disease. These combined approaches will allow one to monitor multiple molecular events, pathways and interactive networks simultaneously-a requirement for elucidating toxic mechanisms. But, before embracing the 'omics' technologies as the 'be all-end all;' they need to be validated for their predictive capacities in large-scale multi-institutional studies, such as those described in this article.

Animals↗