Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “proteomics database”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Method for qualitative comparisons of protein mixtures based on enzyme-catalyzed stable-isotope incorporation.

Determining which proteins are unique among one or several protein populations is an often-encountered task in proteomics. To this purpose, we present a new method based on trypsin-catalyzed incorporation of the stabile isotope (18)O in the C-termini of tryptic peptides, followed by LC-MALDI MS analysis. The analytical strategy was designed such that proteins unique to a given population out of several can be assigned in a single experiment by the isotopic signal intensity distributions of their tryptic peptides in the recorded mass spectra. The method is demonstrated for protein-protein interaction analysis, in which the differential isotope labeling was used to distinguish endogenous human brain proteins interacting with a recombinant bait protein from nonbiospecific background binders.

Brain↗

Structural and functional characterization of gene products encoded in the human genome by homology detection.

Availability of the human genome data has enabled the exploration of a huge amount of biological information encoded in it. There are extensive ongoing experimental efforts to understand the biological functions of the gene products encoded in the human genome. However, computational analysis can aid immensely in the interpretation of biological function by associating known functional/structural domains to the human proteins. In this article we have discussed the implications of such associations. The association of structural domains to human proteins could help in prioritizing the targets for structure determination in the structural genomics initiatives. The protein kinase family is one of the most frequently occurring protein domain families in the human proteome while P-loop hydrolase, which comprises many GTPases and ATPases, is a highly represented superfamily. Using the superfamily relationships between families of unknown and known structures we could increase structural information content of the human genome by about 5%. We could also make new associations of domain families to 33 human proteins that are potentially linked to genetically inherited diseases.

Databases, Genetic↗

CEBS object model for systems biology data, SysBio-OM.

MOTIVATION: To promote a systems biology approach to understanding the biological effects of environmental stressors, the Chemical Effects in Biological Systems (CEBS) knowledge base is being developed to house data from multiple complex data streams in a systems friendly manner that will accommodate extensive querying from users. Unified data representation via a single object model will greatly aid in integrating data storage and management, and facilitate reuse of software to analyze and display data resulting from diverse differential expression or differential profile technologies. Data streams include, but are not limited to, gene expression analysis (transcriptomics), protein expression and protein-protein interaction analysis (proteomics) and changes in low molecular weight metabolite levels (metabolomics). RESULTS: To enable the integration of microarray gene expression, proteomics and metabolomics data in the CEBS system, we designed an object model, Systems Biology Object Model (SysBio-OM). The model is comprehensive and leverages other open source efforts, namely the MicroArray Gene Expression Object Model (MAGE-OM) and the Proteomics Experiment Data Repository (PEDRo) object model. SysBio-OM is designed by extending MAGE-OM to represent protein expression data elements (including those from PEDRo), protein-protein interaction and metabolomics data. SysBio-OM promotes the standardization of data representation and data quality by facilitating the capture of the minimum annotation required for an experiment. Such standardization refines the accuracy of data mining and interpretation. The open source SysBio-OM model, which can be implemented on varied computing platforms is presented here. AVAILABILITY: A universal modeling language depiction of the entire SysBio-OM is available at http://cebs.niehs.nih.gov/SysBioOM/. The Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. The database and interface are being built to implement the model and will be available for public use at http://cebs.niehs.nih.gov.

Database Management Systems↗

Proteomic analysis of proteins associated with lipid droplets of basal and lipolytically stimulated 3T3-L1 adipocytes.

Adipocytes hold the body's major energy reserve as triacylglycerols packaged in large lipid droplets. Perilipins, the most abundant proteins on these lipid droplets, play a critical role in facilitating both triacylglycerol storage and hydrolysis. The stimulation of lipolysis by beta-adrenergic agonists triggers rapid phosphorylation of perilipin and translocation of hormone-sensitive lipase to the surfaces of lipid droplets and more gradual fragmentation and dispersion of micro-lipid droplets. Because few lipid droplet-associated proteins have been identified in adipocytes, we isolated lipid droplets from basal and lipolytically stimulated 3T3-L1 adipocytes and identified the component proteins by mass spectrometry. Structural proteins identified in both preparations include perilipin, S3-12, vimentin, and TIP47; in contrast, adipophilin, caveolin-1, and tubulin selectively localized to droplets in lipolytically stimulated cells. Lipid metabolic enzymes identified in both preparations include hormone-sensitive lipase, lanosterol synthase, NAD(P)-dependent steroid dehydrogenase-like protein, acyl-CoA synthetase, long chain family member (ACSL) 1, and CGI-58. 17-beta-Hydroxysteroid dehydrogenase, type 7, was identified only in basal preparations, whereas ACSL3 and 4 and two short-chain reductase/dehydrogenases were identified on droplets from lipolytically stimulated cells. Additionally, both preparations contained FSP27, ribophorin I, EHD2, diaphorase I, and ancient ubiquitous protein. Basal preparations contained CGI-49, whereas lipid droplets from lipolytically stimulated cells contained several Rab GTPases and tumor protein D54. A close association of mitochondria with lipid droplets was suggested by the identification of pyruvate carboxylase, prohibitin, and a subunit of ATP synthase in the preparations. Thus, adipocyte lipid droplets contain specific structural proteins as well as lipid metabolic enzymes; the structural reorganization of lipid droplets in response to the hormonal stimulation of lipolysis is accompanied by increases in the relative mass of several proteins and the recruitment of additional proteins.

3T3-L1 Cells↗

Proteomics analysis of phosphotyrosyl-proteins in human lumbar cerebrospinal fluid.

Cerebrospinal fluid (CSF) is a secretion product of several different central nervous system (CNS) structures, including the choroid plexus in the ventricles. Pathological CNS processes are reflected in the protein composition of CSF. To elucidate the molecular events that occur in the homeostatic and pathological processes of the CNS, the high-throughput characterization of differentially expressed proteins, and of post-translationally modified proteins, is needed for proteomics studies of CSF. Among the post-translational modifications of proteins, phosphorylation is the most common and important mechanism for the reversible regulation of protein function. In this study, CSF phosphotyrosyl (p-Tyr)-proteins were detected with antibodies and were analyzed with proteomics methods. Three different combination methods--1D gel electrophoresis and Western blotting, immunoprecipitation and 2D gel electrophoresis, and 2D gel electrophoresis and Western blotting--were used to detect p-Tyr-proteins in human lumbar CSF samples. Six protein spots, representing four proteins on a 2D Western blot, were identified as p-Tyr-proteins with the 2D gel electrophoresis and Western blotting method. Those four p-Tyr-proteins are kallikrein-6 precursor, complement C4 gamma-chain, gelsolin, and ceruloplasmin precursor. Additionally, four other nonphosphorylated CSF proteins--beta-2-glycoprotein I precursor, fibulin-1 precursor, EGF-containing fibulin-like extracellur matrix protein 1 precursor, and angiotensinogen precursor--were characterized for the first time.

Blotting, Western↗

PSORTb v.2.0: expanded prediction of bacterial protein subcellular localization and insights gained from comparative proteome analysis.

MOTIVATION: PSORTb v.1.1 is the most precise bacterial localization prediction tool available. However, the program's predictive coverage and recall are low and the method is only applicable to Gram-negative bacteria. The goals of the present work are as follows: increase PSORTb's coverage while maintaining the existing precision level, expand it to include Gram-positive bacteria and then carry out a comparative analysis of localization. RESULTS: An expanded database of proteins of known localization and new modules using frequent subsequence-based support vector machines was introduced into PSORTb v.2.0. The program attains a precision of 96% for Gram-positive and Gram-negative bacteria and predictive coverage comparable to other tools for whole proteome analysis. We show that the proportion of proteins at each localization is remarkably consistent across species, even in species with varying proteome size. AVAILABILITY: Web-based version: http://www.psort.org/psortb. Standalone version: Available through the website under GNU General Public License. CONTACT: psort-mail@sfu.ca, brinkman@sfu.ca SUPPLEMENTARY INFORMATION: http://www.psort.org/psortb/supplementaryinfo.html.

Algorithms↗

A proteome-wide analysis of domain architectures of prokaryotic single-spanning transmembrane proteins.

We performed a proteome-wide survey of the domain architectures in single-spanning transmembrane (TM) proteins (single-spannings) from 87 sequenced prokaryotic (Bacterial and Archaean) genomes by assigning Pfam domains to their N-tail and C-tail loops. Out of 14,625 single-spannings, 3,516 sequences have at least one domain assigned, and no domains were assigned to 7,850, with the remaining 3,259 with less reliable assignment. In the domain-assigned sequences, 3116 sequences are with at most two domains, and the other 400 sequences with more than two. The assigned domains distribute over 651 Pfam families, which account for 11.4% of the total Pfam-A families. Among the 651 families are mostly soluble-protein-originated ones, but only 21 families are unique to TM proteins. The occurrence frequency of the individual domain families follows a power-law, that is, 264 families occur only once, 106 just twice, and the families appeared more than 30 times are counted by only 39. It is found that the great majority of the sequences having one or two domains are of the type II topology with the C-tail loop containing domains on it. On the contrary, the N-tail loop of the same type topology seldom carries domains. Importantly, the assigned domains are always found on the tail loops longer than 60 residues, even for the small domains with less than 30 residues. There are still as many as 5,800 sequences without assigned domains in spite of having at least one long tail, on which no less than 1,000 novel domain families are expected most likely to lie concealed unknown yet. We also investigated the domain arrangement preference and the domain family combination patterns in 'singlets' (single-spannings with one assigned domain) and 'doublets' (with two domains).

Computational Biology↗

GARBAN: genomic analysis and rapid biological annotation of cDNA microarray and proteomic data.

SUMMARY: Genomic Analysis and Rapid Biological ANnotation (GARBAN) is a new tool that provides an integrated framework to analyze simultaneously and compare multiple data sets derived from microarray or proteomic experiments. It carries out automated classifications of genes or proteins according to the criteria of the Gene Ontology Consortium at a level of depth defined by the user. Additionally, it performs clustering analysis of all sets based on functional categories or on differential expression levels. GARBAN also provides graphical representations of the biological pathways in which all the genes/proteins participate. AVAILABILITY: http://garban.tecnun.es.

Algorithms↗

Proteome analysis reveals elevated serum levels of clusterin in patients with preeclampsia.

Preeclampsia is a pregnancy-specific syndrome and a major cause of maternal mortality. The pathophysiology of preeclampsia is unknown, and no proteome analysis of preeclampsia has been reported. We sought to identify proteins associated with preeclampsia using a proteomic technique and performed two-dimensional electrophoresis (2-DE) on sera from six patients with preeclampsia and six normal pregnant women, followed by comparison of the SYPRO Ruby-stained 2-DE profiles. A group of overexpressed spots was identified in the limited study set. Overexpressed spots were identified as clusterin by matrix-assisted laser desorption/ionization-time of flight-mass spectrometry (MALDI-TOF-MS) followed by peptide mass fingerprinting, a protein database search, and Western blot analysis. Additionally, sera of 80 preeclamptic women and 80 normal pregnant women were processed by immunoassay methods to confirm changes in clusterin concentrations quantitatively. Immunoassays showed that clusterin levels in the 80 preeclamptic women were significantly higher than those in the 80 controls (mean +/- SD; 1.62 +/- 0.46 times reference level in preeclamptic women vs. 1.30 +/- 0.46 times reference level in controls, P < 0.001). Proteomic analysis of serum proteins is a promising tool for studying preeclampsia pathophysiology and identifying proteins associated with preeclampsia.

Blood Proteins↗

EchoBASE: an integrated post-genomic database for Escherichia coli.

EchoBASE (http://www.ecoli-york.org) is a relational database designed to contain and manipulate information from post-genomic experiments using the model bacterium Escherichia coli K-12. Its aim is to collate information from a wide range of sources to provide clues to the functions of the approximately 1500 gene products that have no confirmed cellular function. The database is built on an enhanced annotation of the updated genome sequence of strain MG1655 and the association of experimental data with the E.coli genes and their products. Experiments that can be held within EchoBASE include proteomics studies, microarray data, protein-protein interaction data, structural data and bioinformatics studies. EchoBASE also contains annotated information on 'orphan' enzyme activities from this microbe to aid characterization of the proteins that catalyse these elusive biochemical reactions.

Databases, Genetic↗

The normal human amniotic fluid supernatant proteome.

Proteomic analysis combining two-dimentional electrophoresis (2DE) and mass spectrometry (MS) has the potential for a wide range of applications in biological and medical sciences, as protein screening in tissues obtained from healthy and diseased conditions can determine drug targets and diagnostic markers. Conventionally, amniotic fluid (AF) samples are routinely used for prenatal diagnosis of a wide range of fetal abnormalities. Proteomics have already been applied in the analysis of tissues from fetuses with Down's syndrome, in order to detect differences in their protein profile as compared to the normal profiles and to determine possible diagnostic tools. A detailed protein 2DE for the normal human AF has not been reported. In the present study, the 2D protein database of the normal human AF supernatant (AFS) was constructed. Ten AFS samples from women carrying normal fetuses were analysed by 2DE. A mean of 412 spots per gel were analyzed and protein identification was carried out by MALDI-MS and MALDI-MS-MS. A 2D protein map comprising of 136 different gene products was constructed. The majority of the identified proteins are regulatory proteins, enzymes, secreted proteins, carriers and immunoglobulins. Twelve hypothetical proteins were also included. The normal AFS proteome map is a valuable tool for the study of aberrant protein expression and the search for proteins as possible markers for the prediction of abnormal fetuses.

Amniotic Fluid↗

High-throughput expression, purification, and characterization of recombinant Caenorhabditis elegans proteins.

Modern proteomics approaches include techniques to examine the expression, localization, modifications, and complex formation of proteins in cells. In order to address issues of protein function in vitro using classical biochemical and biophysical approaches, high-throughput methods of cloning the appropriate reading frames, and expressing and purifying proteins efficiently are an important goal of modern proteomics approaches. This process becomes more difficult as functional proteomics efforts focus on the proteins from higher organisms, since issues of correctly identifying intron-exon boundaries and efficiently expressing and solubilizing the (often) multi-domain proteins from higher eukaryotes are challenging. Recently, 12,000 open-reading-frame (ORF) sequences from Caenorhabditis elegans have become available for functional proteomics studies [Nat. Gen. 34 (2003) 35]. We have implemented a high-throughput screening procedure to express, purify, and analyze by mass spectrometry hexa-histidine-tagged C. elegans ORFs in Escherichia coli using metal affinity ZipTips. We find that over 65% of the expressed proteins are of the correct mass as analyzed by matrix-assisted laser desorption MS. Many of the remaining proteins indicated to be "incorrect" can be explained by high-throughput cloning or genome database annotation errors. This provides a general understanding of the expected error rates in such high-throughput cloning projects. The ZipTip purified proteins can be further analyzed under both native and denaturing conditions for functional proteomics efforts.

Animals↗

Exploitation of molecular profiling techniques for GM food safety assessment.

Several strategies have been developed to identify unintended alterations in the composition of genetically modified (GM) food crops that may occur as a result of the genetic modification process. These include comparative chemical analysis of single compounds in GM food crops and their conventional non-GM counterparts, and profiling methods such as DNA/RNA microarray technologies, proteomics and metabolite profiling. The potential of profiling methods is obvious, but further exploration of specificity, sensitivity and validation is needed. Moreover, the successful application of profiling techniques to the safety evaluation of GM foods will require linked databases to be built that contain information on variations in profiles associated with differences in developmental stages and environmental conditions.

Consumer Product Safety↗

RIO: analyzing proteomes by automated phylogenomics using resampled inference of orthologs.

BACKGROUND: When analyzing protein sequences using sequence similarity searches, orthologous sequences (that diverged by speciation) are more reliable predictors of a new protein's function than paralogous sequences (that diverged by gene duplication). The utility of phylogenetic information in high-throughput genome annotation ("phylogenomics") is widely recognized, but existing approaches are either manual or not explicitly based on phylogenetic trees. RESULTS: Here we present RIO (Resampled Inference of Orthologs), a procedure for automated phylogenomics using explicit phylogenetic inference. RIO analyses are performed over bootstrap resampled phylogenetic trees to estimate the reliability of orthology assignments. We also introduce supplementary concepts that are helpful for functional inference. RIO has been implemented as Perl pipeline connecting several C and Java programs. It is available at http://www.genetics.wustl.edu/eddy/forester/. A web server is at http://www.rio.wustl.edu/. RIO was tested on the Arabidopsis thaliana and Caenorhabditis elegans proteomes. CONCLUSION: The RIO procedure is particularly useful for the automated detection of first representatives of novel protein subfamilies. We also describe how some orthologies can be misleading for functional inference.

Animals↗

Post-genomics of microsporidia, with emphasis on a model of minimal eukaryotic proteome: a review.

The genome sequence of the microsporidian parasite Encephalitozoon cuniculi Levaditi, Nicolau et Schoen, 1923 contains about 2,000 genes that are representative of a non-redundant potential proteome composed of 1,909 protein chains. The purpose of this review is to relate some advances in the characterisation of this proteome through bioinformatics and experimental approaches. The reduced diversity of the set of E. cuniculi proteins is perceptible in all the compilations of predicted domains, orthologs, families and superfamilies, available in several public databases. The phyletic patterns of orthologs for seven eukaryotic organisms support an extensive gene loss in the fungal clade, with additional deletions in E. cuniculi. Most microsporidial orthologs are the smallest ones among eukaryotes, justifying an interest in the use of these compacted proteins to better discriminate between essential and non-essential regions. The three components of the E. cuniculi mRNA capping apparatus have been especially well characterized and the three-dimensional structure of the cap methyltransferase has been elucidated following the crystallisation of the microsporidial enzyme Ecm1. So far, our mass spectrometry-based analyses of the E. cuniculi spore proteome has led to the identification of about 170 proteins, one-quarter of these having no clearly predicted function. Immunocytochemical studies are in progress to determine the subcellular localisation of microsporidia-specific proteins. Post-translational modifications such as phosphorylation and glycosylation are expected to be soon explored.

Animals↗

Bacterial pathogen genomics and vaccines.

Infectious diseases remain a major cause of deaths and disabilities in the world, the majority of which are caused by bacteria. Although immunisation is the most cost effective and efficient means to control microbial diseases, vaccines are not yet available to prevent many major bacterial infections. Examples include dysentery (shigellosis), gonorrhoea, trachoma, gastric ulcers and cancer (Helicobacter pylori). Improved vaccines are needed to combat some diseases for which current vaccines are inadequate. Tuberculosis, for example, remains rampant throughout most countries in the world and represents a global emergency heightened by the pandemic of HIV. The availability of complete genome sequences has dramatically changed the opportunities for developing novel and improved vaccines and facilitated the efficiency and rapidity of their development. Complete genomic databases provide an inclusive catalogue of all potential candidate vaccines for any bacterial pathogen. In conjunction with adjunct technologies, including bioinformatics, random mutagenesis, microarrays, and proteomics, a systematic and comprehensive approach to identifying vaccine discovery can be undertaken. Genomics must be used in conjunction with population biology to ensure that the vaccine can target all pathogenic strains of a species. A proof in principle of the utility of genomics is provided by the recent exploitation of the complete genome sequence of Neisseria meningitidis group B.

Animals↗

Proteomic analysis of individual human embryos to identify novel biomarkers of development and viability.

OBJECTIVE: To develop a method to analyze the proteome of individual human blastocysts and identify differentially expressed proteins prior to implantation. DESIGN: Experimental study. SETTING: Research environment. PATIENT(S): Couples undergoing infertility treatment donated with consent cryopreserved human cleavage-stage embryos for research. INTERVENTION(S): Individual embryos were extracted and analyzed by time-of-flight mass spectrometry. MAIN OUTCOME MEASURE(S): The protein expression profiles of individual embryos. RESULT(S): Differential protein expression profiles were observed between early and expanded blastocysts, as well as between developing blastocysts and degenerate embryos. Significantly, several up-regulated and down-regulated proteins were detected in degenerating embryos. A search in the protein databases highlighted several candidates, including an inhibitor of Tcf-4 (transcription factor mediating Wnt signaling) and an apoptotic protease-activating factor. CONCLUSION(S): This is the first study to successfully analyze the proteome of individual human embryos. This study has shown that protein expression profiles relate to morphology, with degenerating embryos exhibiting significant up-regulation of several potential biomarkers that might be involved in apoptotic and growth-inhibiting pathways.

Biomarkers↗

C. elegans ORFeome version 3.1: increasing the coverage of ORFeome resources with improved gene predictions.

The first version of the Caenorhabditis elegans ORFeome cloning project, based on release WS9 of Wormbase (August 1999), provided experimental verifications for approximately 55% of predicted protein-encoding open reading frames (ORFs). The remaining 45% of predicted ORFs could not be cloned, possibly as a result of mispredicted gene boundaries. Since the release of WS9, gene predictions have improved continuously. To test the accuracy of evolving predictions, we attempted to PCR-amplify from a highly representative worm cDNA library and Gateway-clone approximately 4200 ORFs missed earlier and for which new predictions are available in WS100 (May 2003). In this set we successfully cloned 63% of ORFs with supporting experimental data ("touched" ORFs), and 42% of ORFs with no supporting experimental evidence ("untouched" ORFs). Approximately 2000 full-length ORFs were cloned in-frame, 13% of which were corrected in their exon/intron structure relative to WS100 predictions. In total, approximately 12,500 C. elegans ORFs are now available as Gateway Entry clones for various reverse proteomics (ORFeome v3.1). This work illustrates why the cloning of a complete C. elegans ORFeome, and likely the ORFeomes of other multicellular organisms, needs to be an iterative process that requires multiple rounds of experimental validation together with gradually improving gene predictions.

Animals↗