Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

A uniform proteomics MS/MS analysis platform utilizing open XML file formats.

The analysis of tandem mass (MS/MS) data to identify and quantify proteins is hampered by the heterogeneity of file formats at the raw spectral data, peptide identification, and protein identification levels. Different mass spectrometers output their raw spectral data in a variety of proprietary formats, and alternative methods that assign peptides to MS/MS spectra and infer protein identifications from those peptide assignments each write their results in different formats. Here we describe an MS/MS analysis platform, the Trans-Proteomic Pipeline, which makes use of open XML file formats for storage of data at the raw spectral data, peptide, and protein levels. This platform enables uniform analysis and exchange of MS/MS data generated from a variety of different instruments, and assigned peptides using a variety of different database search programs. We demonstrate this by applying the pipeline to data sets generated by ThermoFinnigan LCQ, ABI 4700 MALDI-TOF/TOF, and Waters Q-TOF instruments, and searched in turn using SEQUEST, Mascot, and COMET.

Archaeal Proteins↗

ConPred_elite: a highly reliable approach to transmembrane topology predication.

The function of transmembrane (TM) proteins is closely correlated to their TM topology; large quantities of highly reliable TM topology data are becoming increasingly required. We present a new consensus approach for TM topology prediction (ConPred_elite) that can predict the whole topology with accuracies of 0.98 for prokaryotic and 0.95 for eukaryotic proteins on a dataset of experimentally-characterized TM topologies. The predicted yield on the dataset is 30.4% for prokaryotic and 21.5% for eukaryotic proteins. Applying ConPred_elite to predicted TM proteins extracted from 29 prokaryotic and 10 eukaryotic proteomes, we obtained 3871 and 7271 highly reliable TM topologies (yields, 19.8 and 13.3%), respectively. The predicted TM topology data may contribute to further research into a comprehensive functional classification and identification of TM proteins based on information of the topology.

Archaeal Proteins↗

SNOSID, a proteomic method for identification of cysteine S-nitrosylation sites in complex protein mixtures.

Reversible addition of NO to Cys-sulfur in proteins, a modification termed S-nitrosylation, has emerged as a ubiquitous signaling mechanism for regulating diverse cellular processes. A key first-step toward elucidating the mechanism by which S-nitrosylation modulates a protein's function is specification of the targeted Cys (SNO-Cys) residue. To date, S-nitrosylation site specification has been laboriously tackled on a protein-by-protein basis. Here we describe a high-throughput proteomic approach that enables simultaneous identification of SNO-Cys sites and their cognate proteins in complex biological mixtures. The approach, termed SNOSID (SNO Site Identification), is a modification of the biotin-swap technique [Jaffrey, S. R., Erdjument-Bromage, H., Ferris, C. D., Tempst, P. & Snyder, S. H. (2001) Nat. Cell. Biol. 3, 193-197], comprising biotinylation of protein SNO-Cys residues, trypsinolysis, affinity purification of biotinylated-peptides, and amino acid sequencing by liquid chromatography tandem MS. With this approach, 68 SNO-Cys sites were specified on 56 distinct proteins in S-nitrosoglutathione-treated (2-10 microM) rat cerebellum lysates. In addition to enumerating these S-nitrosylation sites, the method revealed endogenous SNO-Cys modification sites on cerebellum proteins, including alpha-tubulin, beta-tubulin, GAPDH, and dihydropyrimidinase-related protein-2. Whereas these endogenous SNO proteins were previously recognized, we extend prior knowledge by specifying the SNO-Cys modification sites. Considering all 68 SNO-Cys sites identified, a machine learning approach failed to reveal a linear Cys-flanking motif that predicts stable transnitrosation by S-nitrosoglutathione under test conditions, suggesting that undefined 3D structural features determine S-nitrosylation specificity. SNOSID provides the first effective tool for unbiased elucidation of the SNO proteome, identifying Cys residues that undergo reversible S-nitrosylation.

Animals↗

The biology of the post-genomic era: the proteomics.

The complete identification of coding sequences in a number of species has led to announce the beginning of the post-genomic era, new tools have become available to study complex phenomena in biological systems. Rapid advances in genomic sequencing and bioinformatics have established the field of genomics to investigate thousands genes' activity through mRNA display. However, recent studies have demonstrated a lack of correlation between the transcriptional profiles and the actual protein levels in cells, so investigation of the expressed part of the genome is also required to link genomic data to biological function. It is possible that evolutional development occured by increasing complexity of regulation processes at the level of RNA and protein molecules instead of simple increase in gene number, so investigation of proteins and protein complexes became important fields of our post-genomic era. High-resolution two-dimensional gels combined with sensitive mass spectrometry can reveal virtually all proteins present in cells opening new insights into functions of cells, tissues and whole organisms.

Animals↗

Assessing matrix assisted laser desorption/ ionization-time of flight-mass spectrometry as a means of rapid embryo protein identification in rice.

Rice embryo proteins were separated by two-dimensional gel electrophoresis (2-DE). A total of 105 spots were digested with trypsin and the resultant peptides were analyzed by matrix assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS). Raw mass spectra were fully-automatically processed and searched with selected monoisotopic masses against SWISS-PROT/TrEMBL and NCBInr databases. High quality mass spectra were obtained from 53 spots, of which 36 spots were identified including 29 not registered in databases. Fifty percent of the rice embryo proteins resolved in 2-DE could not be identified, indicating more efficient sample preparation techniques need to be developed in the future. At least four to five matching peptides were found to be essential for unambiguous identification of rice embryo proteins; peptide matching of less than four lead to ambiguous results. The suitability of peptide mass fingerprinting method as a means of rapid embryo protein identification in rice was discussed.

Databases, Protein↗

Column chromatographic prefractionation leads to the detection of 543 different gene products in human fetal brain.

In a previous publication a large series of proteins were identified in fetal human brain by the use of two-dimensional electrophoresis (2-DE) with subsequent matrix-assisted laser desorption/ionization-time of flight (MALDI-TOF) and MALDI-tandem time-of-flight (TOF/TOF) analysis. Further identification of many more different spots by traditional 2-DE without additional step such as narrow immobilized ph gradient (IPG) strips or prefractionation seems unlikely and we therefore decided to separate extracted brain proteins by ion-exchange chromatography using a TSK gel DEAE-5PW column followed by 2-DE of individual fractions and analysis by MALDI-TOF/TOF with LIFT technology in fetal brain of the early second trimester. About 1880 protein spots corresponding to 543 different gene products were identified. These proteins included housekeeping, signaling, cytoskeletal, metabolic, antioxidant, and neuron/synaptosomal specific proteins. Among these, 314 gene products (314/543, 57.8%), which have never been detected in traditional 2-DE of human fetal brain, were observed by this method. This updated map of fetal brain proteins may serve as data base and reference map for fetal brain proteins, and the methodology applied may be used as a valuable analytical tool for the basis of protein expressional studies in health and disease.

Brain↗

Automated protein identification by tandem mass spectrometry: issues and strategies.

Protein identification by tandem mass spectrometry (MS/MS) is key to most proteomics projects and has been widely explored in bioinformatics research. Obtaining good and trustful identification results has important implications for biological and clinical work. Although well matured, automated software identification of proteins from MS/MS data still faces a number of obstacles due to the complexity of the proteome or procedural issues of mass spectrometry data acquisition. Expected or unexpected modifications of the peptide sequences, polymorphisms, errors in databases, missed or non-specific cleavages, unusual fragmentation patterns, and single MS/MS spectra of multiple peptides of the same m/z are so many pitfalls for identification algorithms. A lot of research work has been carried out in recent years that yielded new strategies to handle a number of these issues. Multiple MS/MS identification algorithms are now available or have been theoretically described. The difficulty resides in choosing the most adapted method for each type of spectra being identified. This review presents an overview of the state-of-the-art bioinformatics approaches to the identification of proteins by MS/MS to help the reader doing the spade work of finding the right tools among the many possibilities offered.

Automation↗

Profiling of maternal and developmental-stage specific mRNA transcripts in Atlantic halibut Hippoglossus hippoglossus.

cDNA libraries were constructed from the following developmental stages (tissues) of the Atlantic halibut (Hippoglossus hippoglossus): 2-cell stage (embryos), 1 day-old yolk sac larvae (trunk) and juvenile (fast skeletal muscle). A total of 4249 high quality expressed sequence tags from the three libraries were clustered into a partial transcriptome of 2124 putative genes. A large proportion of the gene clusters (48.3%) had no significant matches against known proteins. The most abundant ESTs of nuclear transcripts in the 2-cell library included sequences with high identity to zebrafish H1M, a linker histone-like protein involved in primordial germ cell specification, zinc finger protein, rRNA external transcribed spacer, thymosin beta-4, cyclin B1 and several predicted peptides from the Tetraodon nigroviridis genome assembly with unknown functions. 170 and 123 ESTs represented ribosomal proteins in the larval and juvenile libraries respectively, compared with only two sequences in the 2-cell library, which may reflect an abundance of maternally inherited pre-formed ribosomes in the yolk. Even though some clusters were common to all three libraries, most putative genes showed a developmental-stage specific distribution with 72% (2-cell embryo), 59% (larval) and 57% (juvenile) sequences having no significant matches against the 8400 adult halibut sequences in the EMBL nucleotide database. Comparison between the predicted halibut peptide data set and the human, zebrafish, and pufferfishes (T. nigroviridis and Takifugu rubripes) proteomes revealed that, as expected, the halibut sequences were more similar to the other two fish species than to human proteins. However, no clear bias towards the pufferfishes was observed, suggesting significant sequence variation between orthologues within the clade Acanthomorpha. The sequence information generated in the present study will represent a significant new resource for future studies on normal and abnormal development in Atlantic halibut.

Animals↗

The Phanerochaete chrysosporium secretome: database predictions and initial mass spectrometry peptide identifications in cellulose-grown medium.

The white rot basidiomycete, Phanerochaete chrysosporium, employs an array of extracellular enzymes to completely degrade the major polymers of wood: cellulose, hemicellulose and lignin. Towards the identification of participating enzymes, 268 likely secreted proteins were predicted using SignalP and TargetP algorithms. To assess the reliability of secretome predictions and to evaluate the usefulness of the current database, we performed shotgun LC-MS/MS on cultures grown on standard cellulose-containing medium. A total of 182 unique peptide sequences were matched to 50 specific genes, of which 24 were among the secretome subset. Underscoring the rich genetic diversity of P. chrysosporium, identifications included 32 glycosyl hydrolases. Functionally interconnected enzyme groups were recognized. For example, the multiple endoglucanases and processive exocellobiohydrolases observed quite probably attack cellulose in a synergistic manner. In addition, a hemicellulolytic system included endoxylanases, alpha-galactosidase, acetyl xylan esterase, and alpha-l-arabinofuranosidase. Glucose and cellobiose metabolism likely involves cellobiose dehydrogenase, glucose oxidase, and various inverting glycoside hydrolases, all perhaps enhanced by an epimerase. To evaluate the completeness of the current database, mass spectroscopy analysis was performed on a larger and more inclusive dataset containing all possible ORFs. This allowed identification of a previously undetected hypothetical protein and a putative acid phosphatase. The expression of several genes was supported by RT-PCR amplification of their cDNAs.

Cellulose↗

Predicting oligomeric assemblies: N-mers a primer.

Multi-protein complexes play key roles in many biological processes. However, since the structures of these assemblies are hard to resolve experimentally, the detailed mechanism of how they work cooperatively in the cell has remained elusive. Similarly, recent advances on in silico prediction of protein-protein interactions have so far avoided this difficult problem. In this paper, we present a general algorithm to predict molecular assemblies of homo-oligomers. Given the number of N-mers and the 3D structure of one monomer, the method samples all the possible symmetries that N-mers can be assembled. Based on a scoring function that clusters the low free energy structures at each binding interface, the algorithm predicts the complex structure as well as the symmetry of the protein assembly. The method is quite general and does not involve any free parameters. The algorithm has been implemented as a public server and integrated to the protein-protein complex prediction server ClusPro. Using this application, we validated predictions for trimers, tetramers (discriminating between dimer of dimers and 4-fold symmetry structures), pentamers and hexamers (discriminating between trimer of dimers, dimer of trimers, and 6-fold symmetry structures), for a total of 107 assemblies. For 85% of the multimers, the server predicts the complex structure within an average rms deviation of 2A from the full crystal. For complexes that involve more than one binding interface, the cluster size at each surface provides a strong indication as to which interface forms first. With improving scoring functions and computer power, our multimer docking approach could be used as a framework to address the more general problem of multi-protein assemblies.

Algorithms↗

Plasmodium post-genomics: an update.

The concept behind the first Molecular Approaches to Malaria meeting, held 1-5 February 2000 in Lorne, Australia, was ahead of its time; to convene a meeting of malaria researchers, database developers and genomics scientists, and to discuss how genomic sciences and their relevant disciplines could be applied to solve important problems in malaria research. The success of the second Molecular Approaches to Malaria meeting, held 1-5 February 2004 in the same place, together with the influence of genomics on malaria research, is testament to the vision that the organizers had at the first meeting. This review attempts to capture some of the current efforts in the post-genomics era of malaria research and highlights the approaches discussed at the Molecular Approaches to Malaria 2004 meeting.

Animals↗

Effective detection of remote homologues by searching in sequence dataset of a protein domain fold.

Profile matching methods are commonly used in searches in protein sequence databases to detect evolutionary relationships. We describe here a sensitive protocol, which detects remote similarities by searching in a specialized database of sequences belonging to a fold. We have assessed this protocol by exploring the relationships we detect among sequences known to belong to specific folds. We find that searches within sequences adopting a fold are more effective in detecting remote similarities and evolutionary connections than searches in a database of all sequences. We also discuss the implications of using this strategy to link sequence and structure space.

Databases, Protein↗

Reliability measures for membrane protein topology prediction algorithms.

We have developed reliability scores for five widely used membrane protein topology prediction methods, and have applied them both on a test set of 92 bacterial plasma membrane proteins with experimentally determined topologies and on all predicted helix bundle membrane proteins in three fully sequenced genomes: Escherichia coli, Saccharomyces cerevisiae and Caenorhabditis elegans. We show that the reliability scores work well for the TMHMM and MEMSAT methods, and that they allow the probability that the predicted topology is correct to be estimated for any protein. We further show that the available test set is biased towards high-scoring proteins when compared to the genome-wide data sets, and provide estimates for the expected prediction accuracy of TMHMM across the three genomes. Finally, we show that the performance of TMHMM is considerably better when limited experimental information (such as the in/out location of a protein's C terminus) is available, and estimate that at least ten percentage points in overall accuracy in whole-genome predictions can be gained in this way.

Algorithms↗

The FLEXGene repository: exploiting the fruits of the genome projects by creating a needed resource to face the challenges of the post-genomic era.

Thanks to the results of the multiple completed and ongoing genome sequencing projects and to the newly available recombination-based cloning techniques, it is now possible to build gene repositories with no precedent in their composition, formatting, and potential. This new type of gene repository is necessary to address the challenges imposed by the post-genomic era, i.e., experimentation on a genome-wide scale. We are building the FLEXGene (Full Length EXpression-ready) repository. This unique resource will contain clones representing the complete ORFeome of different organisms, including Homo sapiens as well as several pathogens and model organisms. It will consist of a comprehensive, characterized (sequence-verified), and arrayed gene repository. This resource will allow full exploitation of the genomic information by enabling genome-wide scale experimentation at the level of functional/phenotypic assays as well as at the level of protein expression, purification, and analysis. Here we describe the rationale and construction of this resource and focus on the data obtained from the Saccharomyces cerevisiae project.

Animals↗

Proteomic sensitivity to dietary manipulations in rainbow trout.

Changes in dietary protein sources due to substitution of fish meal by other protein sources can have metabolic consequences in farmed fish. A proteomics approach was used to study the protein profiles of livers of rainbow trout that have been fed two diets containing different proportions of plant ingredients. Both diets control (C) and soy (S) contained fish meal and plant ingredients and synthetic amino acids, but diet S had a greater proportion of soybean meal. A feeding trial was performed for 12 weeks at the end of which, growth and protein metabolism parameters were measured. Protein growth rates were not different in fish fed different diets; however, protein consumption and protein synthesis rates were higher in the fish fed the diet S. Fish fed diet S had lower efficiency of retention of synthesised protein. Ammonia excretion was increased as well as the activities of hepatic glutamate dehydrogenase and aspartate amino transferase (ASAT). No differences were found in free amino acid pools in either liver or muscle between diets. Protein extraction followed by high-resolution two-dimensional electrophoresis, coupled with gel image analysis, allowed identification and expression of hundreds of protein. Individual proteins of interest were then subjected to further analysis leading to protein identification by trypsin digest fingerprinting. During this study, approximately 800 liver proteins were analysed for expression pattern, of which 33 were found to be differentially expressed between diets C and S. Seventeen proteins were positively identified after database searching. Proteins were identified from diverse metabolic pathways, demonstrating the complex nature of gene expression responses to dietary manipulation revealed by proteomic characterisation.

Amino Acids↗