Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Comparative proteomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Structural genomics of minimal organisms and protein fold space.

The initial aim of the Berkeley Structural Genomics Center is to obtain a near-complete structural complement of two minimal organisms, closely related pathogens Mycoplasma genitalium and M. pneumoniae. The former has fewer than 500 genes and the latter fewer than 700 genes. To achieve this goal, the current protein targets have been selected starting with those predicted to be most tractable and likely to yield new structural and functional information. During the past 3 years, the semi-automated structural genomics pipeline has been set up from cloning, expression, purification, and ultimately to structural determination. The results from the pipeline substantially increased the coverage of the protein fold space of M. pneumoniae and M. genitalium. Furthermore, about 1/2 of the structures of 'unique' protein sequences revealed new and novel folds, and over 2/3 of the structures of previously annotated 'hypothetical proteins' inferred their molecular functions.

Bacterial Proteins↗

On-column protein refolding for crystallization.

One major bottleneck in protein production in Escherichia coli for structural genomics projects is the formation of insoluble protein aggregates (inclusion bodies). The efficient refolding of proteins from inclusion bodies is becoming an important tool that can provide soluble native proteins for structural and functional studies. Here we report an on-column refolding method established at the Berkeley Structural Genomics Center (BSGC). Our method is a combination of an 'artificial chaperone-assisted refolding' method previously proposed and affinity chromatography to take advantage of a chromatographic step: less time-consuming, no filtration or concentration, with the additional benefit of protein purification. It can be easily automated and formatted for high-throughput process.

Chromatography, Affinity↗

Protein production and crystallization at the joint center for structural genomics.

By definition, structural genomics centers must be able to address a large number of diverse protein targets. The methods developed should permit parallel and cost-effective processing while allowing for the diverse nature of proteins. Our approach to this problem is a multi-tiered effort where targets are characterized and categorized by behavior and processed in parallel by appropriate methods. The Joint Center for Structural Genomics (JCSG) has applied this tactic to create a fully integrated and scaleable structure determination pipeline. Highlights of the development of the current pipeline for protein production and crystallization are presented here.

Crystallization↗

Automatic classification and pattern discovery in high-throughput protein crystallization trials.

Conceptually, protein crystallization can be divided into two phases search and optimization. Robotic protein crystallization screening can speed up the search phase, and has a potential to increase process quality. Automated image classification helps to increase throughput and consistently generate objective results. Although the classification accuracy can always be improved, our image analysis system can classify images from 1,536-well plates with high classification accuracy (85%) and ROC score (0.87), as evaluated on 127 human-classified protein screens containing 5,600 crystal images and 189,472 non-crystal images. Data mining can integrate results from high-throughput screens with information about crystallizing conditions, intrinsic protein properties, and results from crystallization optimization. We apply association mining, a data mining approach that identifies frequently occurring patterns among variables and their values. This approach segregates proteins into groups based on how they react in a broad range of conditions, and clusters cocktails to reflect their potential to achieve crystallization. These results may lead to crystallization screen optimization, and reveal associations between protein properties and crystallization conditions. We also postulate that past experience may lead us to the identification of initial conditions favorable to crystallization for novel proteins.

Algorithms↗

Genome-scale identification of conditionally essential genes in E. coli by DNA microarrays.

Identifying the genes required for the growth or viability of an organism under a given condition is an important step toward understanding the roles these genes play in the physiology of the organism. Currently, the combination of global transposon mutagenesis with PCR-based mapping of transposon insertion sites is the most common method for determining conditional gene essentiality. In order to accelerate the detection of essential gene products, here we test the utility and reliability of a DNA microarray technology-based method for the identification of conditionally essential genes of the bacterium, Escherichia coli, grown in rich medium under aerobic or anaerobic growth conditions using two different DNA microarray platforms. Identification and experimental verification of five hypothetical E. coli genes essential for anaerobic growth directly demonstrated the utility of the method. However, the two different DNA microarray platforms yielded largely non-overlapping results after a two standard deviations cutoff and were subjected to high false positive background levels. Thus, further methodological improvements are needed prior to the use of DNA microarrays to reliably identify conditionally essential genes on genome-scale.

Aerobiosis↗

ConPred_elite: a highly reliable approach to transmembrane topology predication.

The function of transmembrane (TM) proteins is closely correlated to their TM topology; large quantities of highly reliable TM topology data are becoming increasingly required. We present a new consensus approach for TM topology prediction (ConPred_elite) that can predict the whole topology with accuracies of 0.98 for prokaryotic and 0.95 for eukaryotic proteins on a dataset of experimentally-characterized TM topologies. The predicted yield on the dataset is 30.4% for prokaryotic and 21.5% for eukaryotic proteins. Applying ConPred_elite to predicted TM proteins extracted from 29 prokaryotic and 10 eukaryotic proteomes, we obtained 3871 and 7271 highly reliable TM topologies (yields, 19.8 and 13.3%), respectively. The predicted TM topology data may contribute to further research into a comprehensive functional classification and identification of TM proteins based on information of the topology.

Archaeal Proteins↗

Quantitative and qualitative comparisons of Cryptosporidium faecal purification procedures for the isolation of oocysts suitable for proteomic analysis.

With the recent publication of the Cryptosporidium genome, investigation of the proteins expressed by Cryptosporidium parvum will provide complementary information on the biology of this complex organism. Proteomic studies on this apicomplexan parasite have been hampered due to the inability to culture or isolate high numbers required for 2D gel analysis. Neonatal calves are a common source of Cryptosporidium oocysts and we report on the development of a sucrose-Percoll purification procedure which produced the high yield and purity (free from faecal and bacterial contaminants) that is required for successful proteomic studies from neonatal calves. We report on the development of quantitative and qualitative flow cytometric methods which were confirmed by epifluorescence microscopy. A comparison of five common purification procedures was carried out to determine the efficiency of the sucrose-Percoll gradient. 2D-PAGE results strongly support the sucrose-Percoll procedure as the most suitable method for applications like proteomics which require the recovery of high numbers of isolated oocysts with minimal faecal and bacterial contaminants.

Animals↗

Biotechnology and vaccines: application of functional genomics to Neisseria meningitidis and other bacterial pathogens.

Since its introduction, vaccinology has been very effective in preventing infectious diseases. However, in several cases, the conventional approach to identify protective antigens, based on biochemical, immunological and microbiological methods, has failed to deliver successful vaccine candidates against major bacterial pathogens. The recent development of powerful biotechnological tools applied to genome-based approaches has revolutionized vaccine development, biological research and clinical diagnostics. The availability of a genome provides an inclusive virtual catalogue of all the potential antigens from which it is possible to select the molecules that are likely to be more effective. Here, we describe the use of "reverse vaccinology", which has been successful in the identification of potential vaccines candidates against Neisseria meningitidis serogroup B and review the use of functional genomics approaches as DNA microarrays, proteomics and comparative genome analysis for the identification of virulence factors and novel vaccine candidates. In addition, we describe the potential of these powerful technologies in understanding the pathogenesis of various bacteria.

Bacterial Vaccines↗

A review of quantitative methods for proteomic studies.

An overview is provided of six strategies for relative or absolute quantitation of protein abundances that are widely used in proteomic studies. Strengths and limitations are discussed. Four of these involve stable isotope labeling and isotope ratio measurements by mass spectrometry. In another, mass spectra are used to deconvolute overlapping peptide HPLC peaks to provide relative quantitation based on peak areas. The sixth provides relative abundances of proteins based on 2-D gel arrays. It should be noted that these strategies measure peptide and protein abundances, and cannot directly assess changes in regulation or expression.

Amino Acid Sequence↗

Molecular regulation of histone H3 trimethylation by COMPASS and the regulation of gene expression.

The Set1-containing complex COMPASS, which is the yeast homolog of the human MLL complex, is required for mono-, di-, and trimethylation of lysine 4 of histone H3. We have performed a comparative global proteomic screen to better define the role of COMPASS in histone trimethylation. We report that both Cps60 and Cps40 components of COMPASS are required for proper histone H3 trimethylation, but not for proper regulation of telomere-associated gene silencing. Purified COMPASS lacking Cps60 can mono- and dimethylate but is not capable of trimethylating H3(K4). Chromatin immunoprecipitation (ChIP) studies indicate that the loss subunits of COMPASS required for histone trimethylation do not affect the localization of Set1 to chromatin for the genes tested. Collectively, our results suggest a molecular requirement for several components of COMPASS for proper histone H3 trimethylation and regulation of telomere-associated gene expression, indicating multiple roles for different forms of histone methylation by COMPASS.

DNA-Binding Proteins↗

Glycosylation changes in Alzheimer's disease as revealed by a proteomic approach.

Glycosylation influences the biological activity of proteins and affects their folding and stability. Because aberrant glycosylation is associated with Alzheimer's disease (AD), we applied proteome analysis together with Pro-Q Emerald 300 glycoprotein staining to investigate changes in glycosylated cytosolic proteins in AD and control brain. Frontal cortex proteins from 10 AD patients and 7 non-demented controls were subjected to separation by two-dimensional gel electrophoresis and subsequently stained with carbohydrate-specific Pro-Q Emerald 300 dye. Changes in glycosylation of separated proteins were quantified, and proteins of interest identified by mass spectrometry. Approximately 30% of all detectable proteins in the human frontal cortex appeared glycosylated, including heat shock cognate 71 stress protein and beta isoform of creatine kinase. The glycosylation of collapsin response mediator protein 2 (CRMP-2) and an unknown protein was reduced in AD, while the glycosylation of glial fibrillary acidic protein was increased. CRMP-2 regulates the assembly and polymerization of microtubules and is associated with neurofibrillary tangles in AD. Aberrant glycosylations in AD may help understand the mechanisms of neurodegenerative diseases.

Aged↗

Computational peptide dissection of Melan-a/MART-1 oncoprotein antigenicity.

We have mapped the linear antigenic determinant of a commercial MAb raised in the mouse against the melanoma-associated-antigen Melan-A/MART-1. The B cell epitope on the Melan-A/MART-1 oncoprotein is located in the 15-mer amino acid sequence 101-115 PPAYEKLSAEQSPPP, within residues 102-106. The definition of the antigenic sequence on Melan-A/MART-1 oncoprotein was reached following analyses of MHC II binding potential and similarity level to the mouse proteome, that put into evidence the 15-mer amino acid sequence 101-115 PPAYEKLSAEQSPPP as the top scoring peptide in binding H2-A(d) molecules and the epitopic sequence residues 102-106 (i.e., the peptide sequence PAYEK) as having low-similarity level to the mouse proteome. Dot-blot epitope mapping immunoassay identified proline residue 102 as critical, based on its effect on antibody recognition. The present study adds to previous companion reports in validating the hypothesis that low-similarity to the host's proteome and binding potential to MHC II molecules are essential concurring factors in the modulation of the pool of epitopic sequences.

Amino Acid Sequence↗

3-phenylpropionate catabolism and the Escherichia coli oxidative stress response.

Cells have devised a variety of protection systems against the toxic effects of dioxygen. Dioxygenases are part of this defence mechanism. In Escherichia coli, the positive regulator HcaR, a member of the LysR family of regulators, controls expression of the neighbouring genes, hcaA1, hcaA2, hcaC, hcaB and hcaD, coding for the 3-phenylpropionate dioxygenase complex and 3-phenylpropionate-2',3'-dihydrodiol dehydrogenase, that oxidizes 3-phenylpropionate to 3-(2,3-dihydroxyphenyl) propionate. Differences between expression of hcaR and expression of its target, hcaA, suggest that HcaR is involved in control of other cellular processes or that other regulatory proteins modulate hcaA expression. Protein expression profiling was used to identify other HcaR targets. Two-dimensional gel electrophoresis was used to compare the proteomes of wild-type E. coli and strains in which hcaR was disrupted. Several polypeptides whose production was up- or downregulated in the hcaR mutant were involved in the oxidative stress response. Subsequent experiments demonstrated that hcaR disruption was involved in regulation of genes involved in the oxidative stress response. Modification of the stress response also occurred in an hcaA1A2CD mutant strain. Using gel retardation, the HcaR binding site was estimated to be located about -70 to -55 bp upstream of the hcaA transcription start site. The expression of hcaR was repressed in the absence of oxygen by the ArcA/ArcB two-component system.

Amino Acid Sequence↗

Genome duplication led to highly selective expansion of the Arabidopsis thaliana proteome.

Multiple ancient genome duplications in Arabidopsis thaliana provide unique opportunities to assess factors that influence the fates of duplicated genes. We have found that genes retained in duplicate following one round of genome duplication are significantly more likely to be retained in duplicate again after a subsequent genome duplication. Genes retained in duplicate form a functionally biased set and include a significant over-representation of genes involved in the regulation of transcription.

Arabidopsis↗

Do transposable elements really contribute to proteomes?

Recent studies indicate that the initial classification of transposable elements (TEs) as 'useless', 'selfish' or 'junk' pieces of DNA is not an accurate one. TEs seem to have complex regulatory functions and contribute to the coding regions of many genes. Because this contribution had been documented only at transcript level, we searched for evidence that would also support the translation of TE cassettes. Our findings suggest that the proportion of proteins with TE-encoded fragments (approximately 0.1%), although probably underestimated, is much less than what the data at transcript level suggest (approximately 4%). In all cases, the TE cassettes are derived from old TEs, consistent with the idea that incorporation (exaptation) of TE fragments into functional proteins requires long evolutionary periods. We therefore argue that functional proteins are unlikely to contain TE cassettes derived from young TEs, the role of which is probably limited to regulatory functions.

Animals↗

From genes to photosynthesis in Arabidopsis thaliana.

Although photosynthesis in higher plants is of cyanobacterial descent, it differs strikingly in organization and regulation from the prokaryotic process. Genomics, proteomics, and comparative genome analysis are now providing powerful new tools for the molecular dissection of photosynthesis in higher plants. Mutant screens and reverse genetics identify an increasing number of gene-function relationships that have a bearing on photosynthesis, revealing a marked interdependency between photosynthesis and other cellular processes. Photosynthesis-related functions are mostly located in the chloroplast, but can also be located in other compartments of the plant cell. The analysis by DNA-array hybridization of mRNA expression patterns both in the chloroplast and the nucleus, under various environmental conditions and/or in different genetic backgrounds that affect the function of the plastid, is rapidly improving our understanding of how photosynthesis is regulated, and it reveals that plastid-to-nucleus signaling plays a central role in its control.

Arabidopsis↗

Identification of protein coding regions of rice genes using alternative spectral rotation measure and linear discriminant analysis.

An improved method, called Alternative Spectral Rotation (ASR) measure, for predicting protein coding regions in rice DNA has been developed. The method is based on the Spectral Rotation (SR) measure proposed by Kotlar and Lavner, and its accuracy is higher than that of the SR measure and the Spectral Content (SC) measure proposed by Tiwari et al. In order to increase the identifying accuracy, we chose three different coding characters, namely the asymmetric, purine, and stop-codon variables as parameters, and an approving result was presented by the method of Linear Discriminant Analysis (LDA).

Codon↗

A model for random sampling and estimation of relative protein abundance in shotgun proteomics.

Proteomic analysis of complex protein mixtures using proteolytic digestion and liquid chromatography in combination with tandem mass spectrometry is a standard approach in biological studies. Data-dependent acquisition is used to automatically acquire tandem mass spectra of peptides eluting into the mass spectrometer. In more complicated mixtures, for example, whole cell lysates, data-dependent acquisition incompletely samples among the peptide ions present rather than acquiring tandem mass spectra for all ions available. We analyzed the sampling process and developed a statistical model to accurately predict the level of sampling expected for mixtures of a specific complexity. The model also predicts how many analyses are required for saturated sampling of a complex protein mixture. For a yeast-soluble cell lysate 10 analyses are required to reach a 95% saturation level on protein identifications based on our model. The statistical model also suggests a relationship between the level of sampling observed for a protein and the relative abundance of the protein in the mixture. We demonstrate a linear dynamic range over 2 orders of magnitude by using the number of spectra (spectral sampling) acquired for each protein.

Data Collection↗