Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Strategic shotgun proteomics approach for efficient construction of an expression map of targeted protein families in hepatoma cell lines.

An expression map of the most abundant proteins in human hepatoma HepG2 cells was established by a combination of complementary shotgun proteomics approaches. Two-dimensional liquid chromatography (LC)-nano electrospray ionization (ESI) tandem mass spectrometry (MS/MS) as well as one-dimensional LC-matrix-assisted laser desorption/ionization MS/MS were evaluated and shown that additional separation introduced at the peptide level was not as efficient as simple prefractionation of protein extracts in extending the range and total number of proteins identified. Direct LC-nanoESI MS/MS analyses of peptides from total solubilized fraction and the excised gel bands from one-dimensional sodium dodecyl sulfate-polyacrylamide gel electrophoresis fractionated insolubilized fraction afforded the best combination in efficient construction of a nonredundant cell map. Compiling data from multiple variations of rapid shotgun proteomics analyses is nonetheless useful to increase sequence coverage and confidence of hits especially for those proteins identified primarily by a single or two peptide matches. While the returned hit score in general reflects the abundance of the respective proteins, it is not a reliable index for differential expression. Using another closely related hepatoma Hep3B as a comparative basis, 16 proteins with more than two-fold difference in expression level as defined by spot intensity in two-dimensional gel electrophoresis analysis were identified which notably include members of the heat shock protein (Hsp) and heterogeneous nuclear ribonucleoprotein (hnRPN) families. The observed higher expression level of hnRNP A2/B1 and Hsp90 in Hep3B led to a search for reported functional roles mediated in concert by both these multifunctional cellular chaperones. In agreement with the proposed model for telomerase and telomere bound proteins in promoting their interactions, data was obtained which demonstrated that the expression proteomics data could be correlated with longer telomeric length in tumorigenic Hep3B. This biological significance constitutes the basis for further delineation of the dynamic interactions and modifications of the two protein families and demonstrated how proteomic and biological investigation could be mutually substantiated in a productive cycle of hypothesis and pattern driven research.

Carcinoma, Hepatocellular↗

Informatic tools for proteome profiling.

In recent years, the practice of proteomics research has experienced a dramatic shift within the pharmaceutical and biotechnology industry with the widespread implementation of novel applications. The areas of interest extend all the way from discovery of novel drug, vaccine, and diagnostic targets, characterization of protein-based products, toxicology, and identification of surrogate markers of activity in clinical research, to the ability to provide information on the mechanisms of drug action. The power of two-dimensional gel electrophoresis as well as advances in mass spectrometric techniques combined with sequence database correlation have enabled speed and accuracy in identification of proteins in complex mixtures. This article surveys currently available software and informatic tools related to these methods for proteome profiling. The broad acceptance of these technologies, however, has not been accompanied by significant advances in the informatics and software tools necessary to support the analysis and management of the massive amounts of data generated in the process. In this context, this article also discusses the importance of relational databases for protein identification data management.

Biotechnology↗

Proteome analysis. Novel proteins identified at the peribacteroid membrane from Lotus japonicus root nodules.

The peribacteroid membrane (PBM) forms the structural and functional interface between the legume plant and the rhizobia. The model legume Lotus japonicus was chosen to study the proteins present at the PBM by proteome analysis. PBM was purified from root nodules by an aqueous polymer two-phase system. Extracted proteins were subjected to a global trypsin digest. The peptides were separated by nanoscale liquid chromatography and analyzed by tandem mass spectrometry. Searching the nonredundant protein database and the green plant expressed sequence tag database using the tandem mass spectrometry data identified approximately 94 proteins, a number far exceeding the number of proteins reported for the PBM hitherto. In particular, a number of membrane proteins like transporters for sugars and sulfate; endomembrane-associated proteins such as GTP-binding proteins and vesicle receptors; and proteins involved in signaling, for example, receptor kinases, calmodulin, 14-3-3 proteins, and pathogen response-related proteins, including a so-called HIR protein, were detected. Several ATPases and aquaporins were present, indicating a more complex situation than previously thought. In addition, the unexpected presence of a number of proteins known to be located in other compartments was observed. Two characteristic protein complexes obtained from native gel electrophoresis of total PBM proteins were also analyzed. Together, the results identified specific proteins at the PBM involved in important physiological processes and localized proteins known from nodule-specific expressed sequence tag databases to the PBM.

Cell Membrane↗

Prediction of protein subcellular locations using fuzzy k-NN method.

MOTIVATION: Protein localization data are a valuable information resource helpful in elucidating protein functions. It is highly desirable to predict a protein's subcellular locations automatically from its sequence. RESULTS: In this paper, fuzzy k-nearest neighbors (k-NN) algorithm has been introduced to predict proteins' subcellular locations from their dipeptide composition. The prediction is performed with a new data set derived from version 41.0 SWISS-PROT databank, the overall predictive accuracy about 80% has been achieved in a jackknife test. The result demonstrates the applicability of this relative simple method and possible improvement of prediction accuracy for the protein subcellular locations. We also applied this method to annotate six entirely sequenced proteomes, namely Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, Oryza sativa, Arabidopsis thaliana and a subset of all human proteins. AVAILABILITY: Supplementary information and subcellular location annotations for eukaryotes are available at http://166.111.30.65/hying/fuzzy_loc.htm

Algorithms↗

IgStrand: A universal residue numbering scheme for the immunoglobulin-fold (Ig-fold) to study Ig-proteomes and Ig-interactomes.

The Immunoglobulin fold (Ig-fold) is found in proteins from all domains of life and represents the most populous fold in the human genome, with current estimates ranging from 2 to 3% of protein coding regions. That proportion is much higher in the surfaceome where Ig and Ig-like domains orchestrate cell-cell recognition, adhesion and signaling. The ability of Ig-domains to reliably fold and self-assemble through highly specific interfaces represents a remarkable property of these domains, making them key elements of molecular interaction systems: the immune system, the nervous system, the vascular system and the muscular system. We define a universal residue numbering scheme, common to all domains sharing the Ig-fold in order to study the wide spectrum of Ig-domain variants constituting the Ig-proteome and Ig-Ig interactomes at the heart of these systems. The "IgStrand numbering scheme" enables the identification of Ig structural proteomes and interactomes in and between any species, and comparative structural, functional, and evolutionary analyses. We review how Ig-domains are classified today as topological and structural variants and highlight the "Ig-fold irreducible structural signature" shared by all of them. The IgStrand numbering scheme lays the foundation for the systematic annotation of structural proteomes by detecting and accurately labeling Ig-, Ig-like and Ig-extended domains in proteins, which are poorly annotated in current databases and opens the door to accurate machine learning. Importantly, it sheds light on the robust Ig protein folding algorithm used by nature to form beta sandwich supersecondary structures. The numbering scheme powers an algorithm implemented in the interactive structural analysis software iCn3D to systematically recognize Ig-domains, annotate them and perform detailed analyses comparing any domain sharing the Ig-fold in sequence, topology and structure, regardless of their diverse topologies or origin. The scheme provides a robust fold detection and labeling mechanism that reveals unsuspected structural homologies among protein structures beyond currently identified Ig- and Ig-like domain variants. Indeed, multiple folds classified independently contain a common structural signature, in particular jelly-rolls. Examples of folds that harbor an "Ig-extended" architecture are given. Applications in protein engineering around the Ig-architecture are straightforward based on the universal numbering.

Humans↗

The PeptideAtlas project.

The completion of the sequencing of the human genome and the concurrent, rapid development of high-throughput proteomic methods have resulted in an increasing need for automated approaches to archive proteomic data in a repository that enables the exchange of data among researchers and also accurate integration with genomic data. PeptideAtlas (http://www.peptideatlas.org/) addresses these needs by identifying peptides by tandem mass spectrometry (MS/MS), statistically validating those identifications and then mapping identified sequences to the genomes of eukaryotic organisms. A meaningful comparison of data across different experiments generated by different groups using different types of instruments is enabled by the implementation of a uniform analytic process. This uniform statistical validation ensures a consistent and high-quality set of peptide and protein identifications. The raw data from many diverse proteomic experiments are made available in the associated PeptideAtlas repository in several formats. Here we present a summary of our process and details about the Human, Drosophila and Yeast PeptideAtlas builds.

Animals↗

Review of proteomics with applications to genetic epidemiology.

Mapping of the human genome has the potential to transform the traditional methods of genetic epidemiology. The complete draft sequence of the 3.3 billion nucleotides comprising the genome is now available over the Internet, including the location and nearly complete sequence of the 26,000 to 31,000 protein-encoding genes. However, aside from water, almost everything in the human body is either made of, or by, proteins. Although the DNA code provides the instructions for their amino acid sequence, there are an estimated 1.5 million proteins. Thus, the correlation between DNA sequence and protein is low, reflecting alternate splicing as well as post-translational modification. The purpose of this article is to explore ways in which the emerging field of proteomics, the study of proteins in a cell, may inform our approach to gene mapping. This article reviews the various technical approaches currently available for proteomics. Technologies are available to quantify protein expression (and compare normal versus disease states), identify proteins through comparison with sequence information in databases or direct sequencing (which can then be mapped to chromosomal locations to ensure appropriate markers), elucidate protein-protein interactions (which may underlie disease), determine localization of proteins within the cell (abnormal trafficking of proteins could have an inherited basis), and characterize modifications of proteins (which is relevant to modifier gene candidates). Several examples are presented to illustrate the potential application of proteomics to the field of genetic epidemiology, and we conclude with various considerations regarding design and analysis.

Breast Neoplasms↗

Multi-dimensional HPLC/MS of the nucleolar proteome using HPLC-chip/MS.

The proteome of the human nucleolus was investigated in a single analysis using off-line strong cation exchange chromatography and microfraction collection combined with HPLC-chip/MS. The analysis was conducted either as a 1-D workflow with HPLC-chip alone or as a 2-D workflow. Two hundred and six unique proteins were identified in the International Protein Index human database corresponding to 2024 unique tryptic peptides identified in the 2-D analysis. In contrast, only 34 proteins and 151 corresponding tryptic peptides were found by applying a 1-D separation strategy. This clearly indicated that the complexity of the samples required the combination of more than one orthogonal separation technique. Stringent database search criteria, including reversal of sequences and therefore better exclusion of false-positive identifications, were applied for reliable protein identification.

Cell Nucleolus↗

Improving reproducibility and sensitivity in identifying human proteins by shotgun proteomics.

Identifying proteins in cell extracts by shotgun proteomics involves digesting the proteins, sequencing the resulting peptides by data-dependent mass spectrometry (MS/MS), and searching protein databases to identify the proteins from which the peptides are derived. Manual analysis and direct spectral comparison reveal that scores from two commonly used search programs (Sequest and Mascot) validate less than half of potentially identifiable MS/MS spectra (class positive) from shotgun analyses of the human erythroleukemia K562 cell line. Here we demonstrate increased sensitivity and accuracy using a focused search strategy along with a peptide sequence validation script that does not rely exclusively on XCorr or Mowse scores generated by Sequest or Mascot, but uses consensus between the search programs, along with chemical properties and scores describing the nature of the fragmentation spectrum (ion score and RSP). The approach yielded 4.2% false positive and 8% false negative frequencies in peptide assignments. The protein profile is then assembled from peptide assignments using a novel peptide-centric protein nomenclature that more accurately reports protein variants that contain identical peptide sequences. An Isoform Resolver algorithm ensures that the protein count is not inflated by variants in the protein database, eliminating approximately 25% of redundant proteins. Analysis of soluble proteins from a human K562 cells identified 5130 unique proteins, with approximately 100 false positive protein assignments.

Cell Line, Tumor↗

[Proteomics and cardiovascular disease].

The description of the human genome has opened new venues for the study and understanding of pathophysiological phenomena. In the 20th century, individual cell components were studied. The 21st century began with a global analysis of cell components. Thanks to the development of new technologies such as DNA chips, or two-dimensional electrophoresis, we can now study the expression of thousands of genes, or the proteins they encode, in a few hours. Genomics has opened the way for proteomics. Improved knowledge of genes does not provide information about cell functions, because any cell expresses all genes simultaneously. Instead, there is selective gene expression depending on the cell type and the stimuli to which it is exposed. The result of this is the proteome, an ensemble of proteins that are responsible for cell functions at any given moment, which are the object of the study of proteomics. The description of the proteome of cardiac cells has begun and some new proteins have been found to be dysregulated in different cardiomyopathies. These proteins are involved either in energy production or in the stress response, or belong to the cell proteasome or cytoskeleton. They may be potential risk markers or new therapeutic targets in the future. In this sense, chemogenomics is a new methodology for the development of new drugs using genomic and proteomic data.

Animals↗

Computational prediction of cancer-gene function.

Most cancer genes remain functionally uncharacterized in the physiological context of disease development. High-throughput molecular profiling and interaction studies are increasingly being used to identify clusters of functionally linked gene products related to neoplastic cell processes. However, in vivo determination of cancer-gene function is laborious and inefficient, so accurately predicting cancer-gene function is a significant challenge for oncologists and computational biologists alike. How can modern computational and statistical methods be used to reliably deduce the function(s) of poorly characterized cancer genes from the newly available genomic and proteomic datasets? We explore plausible solutions to this important challenge.

Computational Biology↗

Utilization of a new biotinylation reagent in the development of a nondiscriminatory investigative approach for the study of cell surface proteins.

In order to circumvent the various problems encountered during the study of membrane-bound proteins, we designed and synthesized a novel membrane-impermeable biotinylation reagent incorporating chemical properties compatible with this goal. We then developed a nondiscriminatory analytical procedure for such studies which overcomes possible selectivity, contamination and solubility problems. The necessary steps (labeling, limited in situ proteolysis, affinity purification) are all conducted in mild or near native conditions. This versatile method could provide an accurate picture of the cell surface proteome.

Animals↗

Structural proteomics: from the molecule to the system.

Over the next few years, structural proteomics will grapple with the problem of visualizing increasingly elaborate structures, from the atomic details of protein structures up to subcellular structures and the whole cell. A recent EU workshop addressed the question of what experimental and theoretical approaches, technologies and infrastructures this will demand.

Databases, Protein↗

Strategies for the physiome project.

The physiome is the quantitative description of the functioning organism in normal and pathophysiological states. The human physiome can be regarded as the virtual human. It is built upon the morphome, the quantitative description of anatomical structure, chemical and biochemical composition, and material properties of an intact organism, including its genome, proteome, cell, tissue, and organ structures up to those of the whole intact being. The Physiome Project is a multicentric integrated program to design, develop, implement, test and document, archive and disseminate quantitative information, and integrative models of the functional behavior of molecules, organelles, cells, tissues, organs, and intact organisms from bacteria to man. A fundamental and major feature of the project is the databasing of experimental observations for retrieval and evaluation. Technologies allowing many groups to work together are being rapidly developed. Internet II will facilitate this immensely. When problems are huge and complex, a particular working group can be expert in only a small part of the overall project. The strategies to be worked out must therefore include how to pull models composed of many submodules together even when the expertise in each is scattered amongst diverse institutions. The technologies of bioinformatics will contribute greatly to this effort. Developing and implementing code for large-scale systems has many problems. Most of the submodules are complex, requiring consideration of spatial and temporal events and processes. Submodules have to be linked to one another in a way that preserves mass balance and gives an accurate representation of variables in nonlinear complex biochemical networks with many signaling and controlling pathways. Microcompartmentalization vitiates the use of simplified model structures. The stiffness of the systems of equations is computationally costly. Faster computation is needed when using models as thinking tools and for iterative data analysis. Perhaps the most serious problem is the current lack of definitive information on kinetics and dynamics of systems, due in part to the almost total lack of databased observations, but also because, though we are nearly drowning in new information being published each day, either the information required for the modeling cannot be found or has never been obtained. "Simple" things like tissue composition, material properties, and mechanical behavior of cells and tissues are not generally available. The development of comprehensive models of biological systems is a key to pharmaceutics and drug design, for the models will become gradually better predictors of the results of interventions, both genomic and pharmaceutic. Good models will be useful in predicting the side effects and long term effects of drugs and toxins, and when the models are really good, to predict where genomic intervention will be effective and where the multiple redundancies in our biological systems will render a proposed intervention useless. The Physiome Project will provide the integrating scientific basis for the Genes to Health initiative, and make physiological genomics a reality applicable to whole organisms, from bacteria to man.

Computer Simulation↗

Two-dimensional gel electrophoresis maps of the proteome and phosphoproteome of primitively cultured rat mesangial cells.

Mesangial cells (MC) play an important role in maintaining the structure and function of the glomerulus. The proliferation of MC is a prominent feature of many kinds of glomerular disease. The first reference 2-DE maps of rat mesangial cells (RMC), stained with silver staining or Pro-Q Diamond dye, have been established here to describe the proteome and phosphoproteome of RMC, respectively. A total of 157 selected protein spots, corresponding to 118 unique proteins, have been identified by MALDI-TOF-MS or LC-ESI-IT-MS/MS, in which 37 protein spots representing 28 unique proteins have also been stained with Pro-Q Diamond, indicating that they are in phosphorylated forms. All the identified proteins were bioinformatically annotated in detail according to their physiochemical characteristics, subcellular location, and function. Most of the separated or identified protein spots are distributed in the area of mass 10-70 kDa and pI 5.0-8.0. The identified proteins include mainly cytoplasmic and nuclear proteins and some mitochondrial, endoplasmic reticulum, and membrane proteins. These proteins are classified into different functional groups such as structure and mobility proteins (21.2%), metabolic enzymes (16.9%), protein folding and metabolism proteins (13.6%), signaling proteins (14.4%), heat-shock proteins (7.6%), and other functional proteins (12.7%). While structure and mobility proteins are mostly represented by protein spots with high abundance, signaling proteins are mostly represented by protein spots with relatively low abundance. Such a 2-DE database for RMC, especially with many signaling proteins and phosphoproteins characterized, will provide a valuable resource for comparative proteomics analysis of normal and pathologic conditions affecting MC function or pathologic progress.

Animals↗

Exploring the charge space of protein-protein association: a proteomic study.

The rate of association of a protein complex is a function of an intrinsic basal rate and of the magnitude of electrostatic steering. In the present study we analyze the contribution of electrostatics towards the association rate of proteins in a database of 68 transient hetero-protein-protein complexes. Our calculations are based on an upgraded version of the computer algorithm PARE, which was shown to successfully predict the impact of mutations on k(on) by calculating the difference in Columbic energy of interaction of a pair of proteins. HyPare (http://bip.weizmann.ac.il/HyPare), automatically calculates the impact of mutations on a per-residue basis for all residues of a protein-protein interaction, achieving a precision similar to that of PARE. Our calculations show that electrostatics play a marginal role (<10 fold) in determining the rate of association for about half of the complexes in the database. Strong electrostatic steering, which results in an increase of over 100-fold in k(on), was calculated for about 25% of the complexes. Applying HyPare to all 68 complexes in the database shows that a small number of residues are hotspots for association. About 40% of the hotspots are calculated to increase the rate of association upon mutation, and thus increase binding affinity. This is a much higher ratio than found for hotspots for dissociation, where the large majority cause weaker binding. About 40% of the hotspots are located outside the physical boundary of the binding site, making them ideal candidates for protein engineering. Our data shows that a majority of protein-protein complexes are not optimized for fast association. Hotspots are not evenly distributed between all types of amino acids. About 75% of all hotspots are of charged residues. This is understandable, as a charge-reverse mutant changes the total charge by 2. The small number of hydrophobic residues that are hotspots upon mutation probably relates to their location and surrounding. For 18 out of the 68 complexes in the database, experimental values of k(on) are available. For these, a basal rate of association was calculated to be in the range of 10(4)M(-1)s(-1) to 10(7)M(-1)s(-1). Some of these rates were verified independently from experimental mutant data. The basal rates were correlated with the size of the proteins and the shape of the interface.

Algorithms↗

Further advances in the development of a data interchange standard for proteomics data.

The Protein Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparison, exchange and verification. Significant progress was made in advancing the design and implementation of a draft standard for exchanging experimental data from proteomics experiments involving mass spectrometry at the 51st Annual Conference of the American Society for Mass Spectrometry. In collaboration with the American Society for Tests and Measurements, the PSI propose to publish this first draft at the forthcoming HUPO 2nd World Congress in Montreal, 8-11 October 2003.

Computational Biology↗