Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Charting the proteomes of organisms with unsequenced genomes by MALDI-quadrupole time-of-flight mass spectrometry and BLAST homology searching.

MALDI-quadrupole time-of-flight mass spectrometry was applied to identify proteins from organisms whose genomes are still unknown. The identification was carried out by successively searching a sequence database-first with a peptide mass fingerprint, then with a packet of noninterpreted MS/MS spectra, and finally with peptide sequences obtained by automated interpretation of the MS/MS spectra. A "MS BLAST" homology searching protocol was developed to overcome specific limitations imposed by mass spectrometric data, such as the limited accuracy of de novo sequence predictions. This approach was tested in a small-scale proteomic project involving the identification of 15 bands of gel-separated proteins from the methylotrophic yeast Pichia pastoris, whose genome has not yet been sequenced and which is only distantly related to other fungi.

Algorithms↗

In silico identification of potential therapeutic targets in the human pathogen Helicobacter pylori.

Availability of genome sequences of pathogens has provided a tremendous amount of information that can be useful in drug target and vaccine target identification. One of the recently adopted strategies is based on a subtractive genomics approach, in which the subtraction dataset between the host and pathogen genome provides information for a set of genes that are likely to be essential to the pathogen but absent in the host. This approach has been used successfully in recent times to identify essential genes in Pseudomonas aeruginosa. We have used the same methodology to analyse the whole genome sequence of the human gastric pathogen Helicobacter pylori. Our analysis revealed that out of the 1590 coding sequences of the pathogen, 40 represent essential genes that have no human homolog. We have further analysed these 40 genes by the protein sequence databases to list some 10 genes whose products are possibly exposed on the pathogen surface. This preliminary work reported here identifies a small subset of the Helicobacter proteome that might be investigated further for identifying potential drug and vaccine targets in this pathogen.

Anti-Bacterial Agents↗

Towards developing a protein infrared spectra databank (PISD) for proteomics research.

Fourier transform infrared (FTIR) spectroscopy is an attractive tool for proteomics research as it can be used to rapidly characterize protein secondary structure in aqueous solution. Recently, a number of secondary structure prediction methods based on reference sets of FTIR spectra from proteins with known structure from X-ray crystallography have been suggested. These prediction methods, often referred to as pattern recognition based approaches, demonstrated good prediction accuracy using some error measure, e.g., the standard error of prediction (SEP). However, to avoid possible adverse effects from differences in recording, the analysis has been mostly based on reference sets of FTIR spectra from proteins recorded in one laboratory only. As a result, these studies were based on reference sets of FTIR spectra from a limited number of proteins. Pattern recognition based approaches, however, rely on reference sets of FTIR spectra from as many proteins as possible representing all possible band shape variation to be related to the diversity of protein structural classes. Hence, if we want to build reliable pattern recognition based systems to support proteomics research, which are capable of making good predictions from spectral data of any unknown protein, one common goal should be to build a comprehensive protein infrared spectra databank (PISD) containing FTIR spectra of proteins of known structure. We have started the process of developing a comprehensive PISD composed of spectra recorded in different laboratories. As part of this work, here we investigate possible effects on prediction accuracy achieved by a neural network analysis when using reference sets composed of FTIR spectra from different laboratories. Surprisingly low magnitude of difference in SEPs throughout all our experiments suggests that FTIR spectra recorded in different laboratories may be safely combined into one reference set with only minor deterioration of prediction accuracy in the worst case.

Algorithms↗

Subtle modification of isotope ratio proteomics; an integrated strategy for expression proteomics.

Use of minor modification of isotope ratio to code samples for expression proteomics is being investigated. Alteration of (13)C abundance to approximately 2% yields a measurable effect on peptide isotopic distribution and inferred isotope ratio. Elevation of (13)C abundance to 4% leads to extension of isotopic distribution and background peaks across every unit of the mass range. Assessment of isotope ratio measurement variability suggests substantial contributions from natural measurement variability. A better understanding of this variable will allow assessment of the contribution of sequence dependence. Both variables must be understood before meaningful mixing experiments for relative expression proteomics are performed. Subtle modification of isotope ratio ( approximately 1-2% increase in (13)C) had no effect upon either the ability of data-dependent acquisition software or database searching software to trigger tandem mass spectrometry or match MSMS data to peptide sequences. More severe modification of isotope ratio caused a significant drop in performance of both functionalities. Development of software for deconvolution of isotope ratio concomitant with protein identification using LC-MSMS, or any other proteomics strategy, is underway (Isosolv). The identified peptide sequence is then be used to provide elemental composition for accurate isotope ratio decoding and the potential to control for specific amino acid biases should these prove significant. It is suggested that subtle modification of isotope ratio proteomics (SMIRP) offers a convenient approach to in vivo isotope coding of plants and might ultimately be extended to mammals including humans.

Carbon Isotopes↗

Predicting protein subcellular localization: past, present, and future.

Functional characterization of every single protein is a major challenge of the post-genomic era. The large-scale analysis of a cell's proteins, proteomics, seeks to provide these proteins with reliable annotations regarding their interaction partners and functions in the cellular machinery. An important step on this way is to determine the subcellular localization of each protein. Eukaryotic cells are divided into subcellular compartments, or organelles. Transport across the membrane into the organelles is a highly regulated and complex cellular process. Predicting the subcellular localization by computational means has been an area of vivid activity during recent years. The publicly available prediction methods differ mainly in four aspects: the underlying biological motivation, the computational method used, localization coverage, and reliability, which are of importance to the user. This review provides a short description of the main events in the protein sorting process and an overview of the most commonly used methods in this field.

Computational Biology↗

Multiple approaches to data-mining of proteomic data based on statistical and pattern classification methods.

The data-mining challenge presented is composed of two fundamental problems. Problem one is the separation of forty-one subjects into two classifications based on the data produced by the mass spectrometry of protein samples from each subject. Problem two is to find the specific differences between protein expression data of two sets of subjects. In each problem, one group of subjects has a disease, while the other group is nondiseased. Each problem was approached with the intent to introduce a new and potentially useful tool to analyze protein expression from mass spectrometry data. A variety of methodologies, both conventional and nonconventional were used in the analysis of these problems. The results presented show both overlap and discrepancies. What is important is the breadth of the techniques and the future direction this analysis will create.

Artificial Intelligence↗

A correlation algorithm for the automated quantitative analysis of shotgun proteomics data.

Quantitative shotgun proteomic analyses are facilitated using chemical tags such as ICAT and metabolic labeling strategies with stable isotopes. The rapid high-throughput production of quantitative "shotgun" proteomic data necessitates the development of software to automatically convert mass spectrometry-derived data of peptides into relative protein abundances. We describe a computer program called RelEx, which uses a least-squares regression for the calculation of the peptide ion current ratios from the mass spectrometry-derived ion chromatograms. RelEx is tolerant of poor signal-to-noise data and can automatically discard nonusable chromatograms and outlier ratios. We apply a simple correction for systematic errors that improves the accuracy of the quantitative measurement by 32 +/- 4%. Our automated approach was validated using labeled mixtures composed of known molar ratios and demonstrated in a real sample by measuring the effect of osmotic stress on protein expression in Saccharomyces cerevisiae.

Algorithms↗

Proteomic dataset of Sca-1+ progenitor cells.

Embryonic stem cells (ES cells) can differentiate into endothelial cells and smooth muscle cells (SMCs), which participate in vascular angiogenesis. In this study, we differentiated mouse ES cells into Sca-1(+) cells, which have the potential to serve as vascular progenitor cells, and mapped their proteome by 2-DE using a pH 3-10 non-linear gradient and 12% SDS-polyacrylamide gels. A subset of 300 protein spots was analysed and mapped, with 241 protein spots being identified by their PMF using MALDI-TOF MS or by partial amino acid sequencing using MS/MS. Our protein map is the first of Sca-1(+) progenitor cells and will facilitate the identification of proteins differentially expressed during stem cell differentiation. The proteome of adult arterial SMCs is described in an accompanying paper (in this issue, DOI 10.1002/pmic.200402045). All data are made accessible on our website http://www.vascular-proteomics.com.

Animals↗

Review: on the analysis and interpretation of correlations in metabolomic data.

A remarkable inherent feature of cellular metabolism is that the concentrations of a small but significant number of metabolites are strongly correlated when measurements of biological replicates are performed. This review seeks to summarize the recent efforts to elucidate the origin of these observed correlations and points out several aspects concerning their interpretation. It is argued that correlations between metabolites differ profoundly from their transcriptomic and proteomic counterparts, and a straightforward interpretation in terms of the underlying biochemical pathways will unavoidably fail. It is demonstrated that the comparative correlations analysis offers a way to exploit the observed correlations to obtain additional information about the physiological state of the system.

Animals↗

Analysing proteomic data.

The rapid growth of proteomics has been made possible by the development of reproducible 2D gels and biological mass spectrometry. However, despite technical improvements 2D gels are still less than perfectly reproducible and gels have to be aligned so spots for identical proteins appear in the same place. Gels can be warped by a variety of techniques to make them concordant. When gels are manipulated to improve registration, information is lost, so direct methods for gel registration which make use of all available data for spot matching are preferable to indirect ones. In order to identify proteins from gel spots a property or combination of properties that are unique to that protein are required. These can then be used to search databases for possible matches. Molecular mass, pI, amino acid composition and short sequence tags can all be used in database searches. Currently the method of choice for protein identification is mass spectrometry. Proteins are eluted from the gels and cleaved with specific endoproteases to produce a series of peptides of different molecular mass. In peptide mass fingerprinting, the peptide profile of the unknown protein is compared with theoretical peptide libraries generated from sequences in the different databases. Tandem mass spectroscopy (MS/MS) generates short amino acid sequence tags for the individual peptides. These partial sequences combined with the original peptide masses are then used for database searching, greatly improving specificity. Increasingly protein identification from MS/MS data is being fully or partially automated. When working with organisms, which do not have sequenced genomes (the case with most helminths), protein identification by database searching becomes problematical. A number of approaches to cross species protein identification have been suggested, but if the organism being studied is only distantly related to any organism with a sequenced genome then the likelihood of protein identification remains small. The dynamic nature of the proteome means that there really is no such thing as a single representative proteome and a complete set of metadata (data about the data) is going to be required if the full potential of database mining is to be realised in the future.

Animals↗

Fungal degradation of wood: initial proteomic analysis of extracellular proteins of Phanerochaete chrysosporium grown on oak substrate.

Two-dimensional (2-D) gel electrophoresis was used to separate the extracellular proteins produced by the white-rot fungus Phanerochaete chrysosporium. Solid-substrate cultures grown on red oak wood chips yielded extracellular protein preparations which were not suitable for 2-D gel analysis. However, pre-washing the wood chips with water helped decrease the amount of brown material which caused smearing on the acidic side of the isoelectric focusing gel. The 2-D gels from these wood-grown cultures revealed more than 45 protein spots. These spots were subjected to in-gel digestion with trypsin followed by either peptide fingerprint analysis by matrix assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF/MS) or by liquid chromatography (LC)/MS/MS sequencing. Data from both methods were analyzed by Protein Prospector and the local P. chrysosporium annotated database. MALDI-TOF/MS only identified two proteins out of 25 analyzed. This was most likely due to problems associated with glycosylation. Protein sequencing by LC/MS/MS of the same 25 proteins resulted in identification of 16 proteins. Most of the proteins identified act on either cellulose or hemicellulose or their hydrolysis products. Thus far no lignin peroxidase, Mn peroxidase or laccases have been detected.

Cellulose↗

Structural proteomics of minimal organisms: conservation of protein fold usage and evolutionary implications.

BACKGROUND: Determining the complete repertoire of protein structures for all soluble, globular proteins in a single organism has been one of the major goals of several structural genomics projects in recent years. RESULTS: We report that this goal has nearly been reached for several "minimal organisms"--parasites or symbionts with reduced genomes--for which over 95% of the soluble, globular proteins may now be assigned folds, overall 3-D backbone structures. We analyze the structures of these proteins as they relate to cellular functions, and compare conservation of fold usage between functional categories. We also compare patterns in the conservation of folds among minimal organisms and those observed between minimal organisms and other bacteria. CONCLUSION: We find that proteins performing essential cellular functions closely related to transcription and translation exhibit a higher degree of conservation in fold usage than proteins in other functional categories. Folds related to transcription and translation functional categories were also overrepresented in minimal organisms compared to other bacteria.

Bacterial Proteins↗

Effects of formaldehyde inhalation on lung of rats.

OBJECTIVE: To analyze protein changes in the lung of Wistar rats exposed to gaseous formaldehyde (FA) at 32-37 mg/m3 for 4 h/day for 15 days using proteomics technique. METHODS: Lung samples were solubilized and separated by two-dimensional electrophoresis (2-DE), and gel patterns were scanned and analyzed for detection of differently expressed protein spots. These protein spots were identified by MALDI-TOF-MS and NCBInr protein database searching. RESULTS: Four proteins were altered significantly in 32-37 mg/m3 FA group, with 3 proteins up-regulated, 1 protein down-regulated. The 4 proteins were identified as aldose reductase, LIM protein, glyceraldehyde-3-phosphate dehydrogenase, and chloride intracellular channel 3. CONCLUSION: The four proteins are related to cell proliferation induced by FA and defense reaction of anti-oxidation. Proteomics is a powerful tool in research of environmental health, and has prospects in search for protein markers for disease diagnosis and monitoring.

Administration, Inhalation↗

Proteomic analysis of the human pathogen Trypanosoma cruzi.

Trypanosoma cruzi, the protozoan that causes Chagas disease, possesses a complex life cycle involving different developmental stages. Experimental conditions for two-dimensional electrophoresis (2-DE) analysis of T. cruzi trypomastigote, amastigote and epimastigote proteomes were optimized. Comparative proteome analysis of the cell-cycle stages were carried out, revealing that few proteins included in the 2-DE maps displayed significant differential expression among the three developmental forms of the parasite. In order to identify landmark proteins, spots from the trypomastigote 2-DE map were subjected to matrix-assisted laser desorption/ionization-time of flight mass spectrometry peptide mass fingerprinting, resulting in 26 identifications that corresponded to 19 different proteins. Among the identified polypeptides, there were heat shock proteins (HSP; chaperones, HSP 60, HSP 70 and HSP 90), elongation factors, glycolytic pathway enzymes (enolase, pyruvate kinase and 2,3 bisphosphoglycerate mutase) and structural proteins (KMP 11, tubulin and paraflagellar rod components). The relative expression of the identified proteins in the 2-DE maps of the T. cruzi developmental stages is also presented.

Animals↗

Proteomic profiling and neurodegeneration in Alzheimer's disease.

Quantitative proteome analysis of Alzheimer's disease (AD) brains was performed using 2-D gels to identify disease specific changes in protein expression. The task of characterizing the proteome and its components is now practically achievable because of the development and integration of four important tools: protein, EST, and complete genome sequence databases, mass spectrometry, matching software for protein sequences and protein separation technology. Mass spectrometry (MS) instrumentation has undergone a tremendous change over the past decade, culminating in the development of highly sensitive, robust instruments that can reliably analyze biomolecules, particularly proteins and peptides; we identified 35 proteins from over 100 protein spots on a 2-D gel. Using this current technology, protein-expression profiling, which is actually a specialized form of mining, is an important principal application of proteomics. The information obtained has tremendous potential as a means of determining the pathogenesis, and detecting disease markers and potential targets for drug therapy in AD.

Aged↗

Proteomic analysis of glutamine-treated human intestinal epithelial HCT-8 cells under basal and inflammatory conditions.

Glutamine (Gln) promotes intestinal growth and maintains gut structure and function, especially in situations of injury and during inflammation. Several mechanisms could contribute to Gln protective effects on gut. Proteomics enable us to characterize differentially expressed proteins in tissues in response to modifications of the biological or nutritional environment. Gln effects on the human intestinal epithelial HCT-8 cell line proteome were assessed under basal and proinflammatory conditions. The 2-DE gels were obtained and compared. Proteins were identified by MS and using databases. About 1200 spots were detected in both 2- and 10-mM Gln concentrations. Under basal conditions, 24 proteins were differentially expressed in response to Gln. Half of these proteins were implicated in protein biosynthesis or proteolysis and 20% in membrane trafficking. Under proinflammatory conditions, 27 proteins were up- or down-regulated by Gln 10 mM. From these proteins, 40% were involved in protein biosynthesis or proteolysis, 16% in membrane trafficking, 8% in cell cycle and apoptosis mechanisms and 8% in nucleic acid metabolism. This study provides the first holistic picture of proteome modulation by Gln in a human enterocytic cell line under basal and proinflammatory conditions, and supports further evaluation of nutritional modulation of intestinal proteome in humans.

Cell Line↗

A streamlined approach to high-throughput proteomics.

Proteomics has rapidly become an important tool for life science research, allowing the integrated analysis of global protein expression from a single experiment. To accommodate the complexity and dynamic nature of any proteome, researchers must use a combination of disparate protein biochemistry techniques, often a highly involved and time-consuming process. Whilst highly sophisticated, individual technologies for each step in studying a proteome are available, true high-throughput proteomics that provides a high degree of reproducibility and sensitivity has been difficult to achieve. The development of high-throughput proteomic platforms, encompassing all aspects of proteome analysis and integrated with genomics and bioinformatics technology, therefore represents a crucial step for the advancement of proteomics research. ProteomIQ (Proteome Systems) is the first fully integrated, start-to-finish proteomics platform to enter the market. Sample preparation and tracking, centralized data acquisition and instrument control, and direct interfacing with genomics and bioinformatics databases are combined into a single suite of integrated hardware and software tools, facilitating high reproducibility and rapid turnaround times. This review will highlight some features of ProteomIQ, with particular emphasis on the analysis of proteins separated by 2D polyacrylamide gel electrophoresis.

Automation↗

Dynamic spectrum quality assessment and iterative computational analysis of shotgun proteomic data: toward more efficient identification of post-translational modifications, sequence polymorphisms, and novel peptides.

In mass spectrometry-based proteomics, frequently hundreds of thousands of MS/MS spectra are collected in a single experiment. Of these, a relatively small fraction is confidently assigned to peptide sequences, whereas the majority of the spectra are not further analyzed. Spectra are not assigned to peptides for diverse reasons. These include deficiencies of the scoring schemes implemented in the database search tools, sequence variations (e.g. single nucleotide polymorphisms) or omissions in the database searched, post-translational or chemical modifications of the peptide analyzed, or the observation of sequences that are not anticipated from the genomic sequence (e.g. splice forms, somatic rearrangement, and processed proteins). To increase the amount of information that can be extracted from proteomic MS/MS datasets we developed a robust method that detects high quality spectra within the fraction of spectra unassigned by conventional sequence database searching and computes a quality score for each spectrum. We also demonstrate that iterative search strategies applied to such detected unassigned high quality spectra significantly increase the number of spectra that can be assigned from datasets and that biologically interesting new insights can be gained from existing data.

Alternative Splicing↗