Search PubMed⌕ Search

Biomedical subjects

John R Yates

Publications and source records attributed to John R Yates.

At least 19 recordsLinked to original sources

A Donald F. Hunt Story (John's Version).

A personal narrative of my time in the Hunt laboratory and beyond is provided. The impact of the Hunt laboratory on the analysis of peptides and proteins by tandem mass spectrometry is described in the context of the time.

History, 20th Century↗

Novel essential DNA repair proteins Nse1 and Nse2 are subunits of the fission yeast Smc5-Smc6 complex.

The structural maintenance of chromosomes (SMC) family of proteins play essential roles in genomic stability. SMC heterodimers are required for sister-chromatid cohesion (Cohesin: Smc1 & Smc3), chromatin condensation (Condensin: Smc2 & Smc4), and DNA repair (Smc5 & Smc6). The SMC heterodimers do not function alone and must associate with essential non-SMC subunits. To gain further insight into the essential and DNA repair roles of the Smc5-6 complex, we have purified fission yeast Smc5 and identified by mass spectrometry the co-precipitating proteins, Nse1 and Nse2. We show that both Nse1 and Nse2 interact with Smc5 in vivo, as part of the Smc5-6 complex. Nse1 and Nse2 are essential proteins and conserved from yeast to man. Loss of Nse1 and Nse2 function leads to strikingly similar terminal phenotypes to those observed for Smc5-6 inactivation. In addition, cells expressing hypomorphic alleles of Nse1 and Nse2 are, like Smc5-6 mutants, hypersensitive to DNA damage. Epistasis analysis suggests that like Smc5-6, Nse1, and Nse2 function together with Rhp51 in the homologous recombination repair of DNA double strand breaks. The results of this study strongly suggest that Nse1 and Nse2 are novel non-SMC subunits of the fission yeast Smc5-6 DNA repair complex.

Alleles↗

Nuclear membrane proteins with potential disease links found by subtractive proteomics.

To comprehensively identify integral membrane proteins of the nuclear envelope (NE), we prepared separately NEs and organelles known to cofractionate with them from liver. Proteins detected by multidimensional protein identification technology in the cofractionating organelles were subtracted from the NE data set. In addition to all 13 known NE integral proteins, 67 uncharacterized open reading frames with predicted membrane-spanning regions were identified. All of the eight proteins tested targeted to the NE, indicating that there are substantially more integral proteins of the NE than previously thought. Furthermore, 23 of these mapped within chromosome regions linked to a variety of dystrophies.

Algorithms↗

Identification of Plasmodium falciparum antigens by antigenic analysis of genomic and proteomic data.

The recent explosion in genomic sequencing has made available a wealth of data that can now be analyzed to identify protein antigens, potential targets for vaccine development. Here we present, in the context of Plasmodium falciparum, a strategy that rapidly identifies target antigens from large and complex genomes. Sixteen antigenic proteins recognized by volunteers immunized with radiation-attenuated P. falciparum sporozoites, but not by mock immunized controls, were identified. Several of these were more antigenic than previously identified and well characterized P. falciparum-derived protein antigens. The data suggest that immune responses to Plasmodium are dispersed on a relatively large number of parasite antigens. These studies have implications for our understanding of immunodominance and breadth of responses to complex pathogens.

Adult↗

Proteomic characterization of the Chlamydomonas reinhardtii chloroplast ribosome. Identification of proteins unique to th e70 S ribosome.

We have conducted a proteomic analysis of the 70 S ribosome from the Chlamydomonas reinhardtii chloroplast. Twenty-seven orthologs of Escherichia coli large subunit proteins were identified in the 50 S subunit, as well as an ortholog of the spinach plastid-specific ribosomal protein-6. Several of the large subunit proteins of C. reinhardtii have short extension or insertion sequences, but overall the large subunit proteins are very similar to those of spinach chloroplast and E. coli. Two proteins of 38 and 41 kDa, designated RAP38 and RAP41, were identified from the 70 S ribosome that were not found in either of the ribosomal subunits. Phylogenetic analysis identified RAP38 and RAP41 as paralogs of spinach CSP41, a chloroplast RNA-binding protein with endoribonuclease activity. Overall, the chloroplast ribosome of C. reinhardtii is similar to those of spinach chloroplast and E. coli, but the C. reinhardtii ribosome has proteins associated with the 70 S complex that are related to non-ribosomal proteins in other species. In addition, the 30 S subunit contains unusually large orthologs of E. coli S2, S3, and S5 and a novel S1-type protein (Yamaguchi, K. et al., (2002) Plant Cell 14, 2957-2974). These additional proteins and domains likely confer functions used to regulate chloroplast translation in C. reinhardtii.

Amino Acid Sequence↗

Similarity among tandem mass spectra from proteomic experiments: detection, significance, and utility.

Liquid chromatography paired with tandem mass spectrometry is a standard technique for identifying peptides from complex protein mixtures. Most fragment ion spectra acquired by this technique are unique, but some are repeated. Similarities among the spectra from 1D and 2D liquid chromatography experiments were calculated by the dot product algorithm. Similar spectra were grouped, and the degree of duplication was calculated for each sample. In 1D liquid chromatography data from 1D gel bands, 18% of the fragment ion spectra were duplicates. A six-cycle 2D liquid chromatographic separation of more than 200 proteins produced 28% duplicate spectra. A rat hippocampal homogenate analyzed by a 12-cycle 2D liquid chromatographic separation contained 25% duplicate spectra. Removal of these duplicate spectra, however, resulted in fewer peptides being successfully identified by SEQUEST. We propose a modification for peptide identification algorithms that would improve their performance and accuracy by explicitly recognizing and making use of spectral similarity.

Algorithms↗

Cleavage N-terminal to proline: analysis of a database of peptide tandem mass spectra.

Fragmentation at the Xxx-Pro bond was analyzed for a group of peptide mass spectra that were acquired in a Finnigan ion trap mass spectrometer and were generated from proteins digested by enzymes and identified by the Sequest algorithm. Cleavage with formation of a + b + y ions occurred more readily at the Xxx-Pro bond than at other locations in these peptides, and the importance of this cleavage varied by the identity of Xxx. The most abundant Xxx-Pro relative bond cleavage ratios were observed when Xxx was Val, His, Asp, Ile, and Leu, whereas the least abundant cleavage ratios occurred when Xxx was Gly or Pro. Rationalization for these cleavage ratios at Xxx-Pro may include contribution of the Asp or His side chain to enhanced cleavage or the conformation of Pro, Gly, and the aliphatic residues Val, Ile, and Leu at the Xxx location in the Xxx-Pro bond. Although unusual fragmentation behavior has been noted for Pro-containing peptides, this analysis suggests that fragmentation at the Xxx-Pro bond is predictable and that this information may be used to improve the identification of proteins if it is incorporated into peptide sequencing algorithms.

Amino Acid Sequence↗

Wnt proteins are lipid-modified and can act as stem cell growth factors.

Wnt signalling is involved in numerous events in animal development, including the proliferation of stem cells and the specification of the neural crest. Wnt proteins are potentially important reagents in expanding specific cell types, but in contrast to other developmental signalling molecules such as hedgehog proteins and the bone morphogenetic proteins, Wnt proteins have never been isolated in an active form. Although Wnt proteins are secreted from cells, secretion is usually inefficient and previous attempts to characterize Wnt proteins have been hampered by their high degree of insolubility. Here we have isolated active Wnt molecules, including the product of the mouse Wnt3a gene. By mass spectrometry, we found the proteins to be palmitoylated on a conserved cysteine. Enzymatic removal of the palmitate or site-directed and natural mutations of the modified cysteine result in loss of activity, and indicate that the lipid is important for signalling. The purified Wnt3a protein induces self-renewal of haematopoietic stem cells, signifying its potential use in tissue engineering.

Amino Acid Sequence↗

A method for the comprehensive proteomic analysis of membrane proteins.

We describe a method that allows for the concurrent proteomic analysis of both membrane and soluble proteins from complex membrane-containing samples. When coupled with multidimensional protein identification technology (MudPIT), this method results in (i) the identification of soluble and membrane proteins, (ii) the identification of post-translational modification sites on soluble and membrane proteins, and (iii) the characterization of membrane protein topology and relative localization of soluble proteins. Overlapping peptides produced from digestion with the robust nonspecific protease proteinase K facilitates the identification of covalent modifications (phosphorylation and methylation). High-pH treatment disrupts sealed membrane compartments without solubilizing or denaturing the lipid bilayer to allow mapping of the soluble domains of integral membrane proteins. Furthermore, coupling protease protection strategies to this method permits characterization of the relative sidedness of the hydrophilic domains of membrane proteins.

Amino Acid Sequence↗

Large-scale protein identification using mass spectrometry.

Recent achievements in genomics have created an infrastructure of biological information. The enormous success of genomics promptly induced a subsequent explosion in proteomics technology, the emerging science for systematic study of proteins in complexes, organelles, and cells. Proteomics is developing powerful technologies to identify proteins, to map proteomes in cells, to quantify the differential expression of proteins under different states, and to study aspects of protein-protein interaction. The dynamic nature of protein expression, protein interactions, and protein modifications requires measurement as a function of time and cellular state. These types of studies require many measurements and thus high throughput protein identification is essential. This review will discuss aspects of mass spectrometry with emphasis on methods and applications for large-scale protein identification, a fundamental tool for proteomics.

Chromatography, Liquid↗

Protein pathway and complex clustering of correlated mRNA and protein expression analyses in Saccharomyces cerevisiae.

The mRNA and protein expression in Saccharomyces cerevisiae cultured in rich or minimal media was analyzed by oligonucleotide arrays and quantitative multidimensional protein identification technology. The overall correlation between mRNA and protein expression was weakly positive with a Spearman rank correlation coefficient of 0.45 for 678 loci. To place the data sets in a proper biological context, a clustering approach based on protein pathways and protein complexes was implemented. Protein expression levels were transcriptionally controlled for not only single loci but for entire protein pathways (e.g., Met, Arg, and Leu biosynthetic pathways). In contrast, the protein expression of loci in several protein complexes (e.g., SPT, COPI, and ribosome) was posttranscriptionally controlled. The coupling of the methods described provided insight into the biology of S. cerevisiae and a clustering strategy by which future studies should be based.

Amino Acid Sequence↗

Statistical characterization of ion trap tandem mass spectra from doubly charged tryptic peptides.

Collision-induced dissociation (CID) is a common ion activation technique used to energize mass-selected peptide ions during tandem mass spectrometry. Characteristic fragment ions form from the cleavage of amide bonds within a peptide undergoing CID, allowing the inference of its amino acid sequence. The statistical characterization of these fragment ions is essential for improving peptide identification algorithms and for understanding the complex reactions taking place during CID. An examination of 1465 ion trap spectra from doubly charged tryptic peptides reveals several trends important to understanding this fragmentation process. While less abundant than y ions, b ions are present in sufficient numbers to aid sequencing algorithms. Fragment ions exhibit a characteristic series-specific relationship between their masses and intensities. Each residue influences fragmentation at adjacent amide bonds, with Pro quantifiably enhancing cleavage at its N-terminal amide bond and His increasing the formation of b ions at its C-terminal amide bond. Fragment ions corresponding to a formal loss of ammonia appear preferentially in peptides containing Gln and Asn. These trends are partially responsible for the complexity of peptide tandem mass spectra.

Mass Spectrometry↗

The Set2 histone methyltransferase functions through the phosphorylated carboxyl-terminal domain of RNA polymerase II.

The histone methyltransferase Set2, which specifically methylates lysine 36 of histone H3, has been shown to repress transcription upon tethering to a heterologous promoter. However, the mechanism of targeting and the consequence of Set2-dependent methylation have yet to be demonstrated. We sought to identify the protein components associated with Set2 to gain some insights into the in vivo function of this protein. Mass spectrometry analysis of the Set2 complex, purified using a tandem affinity method, revealed that RNA polymerase II (pol II) is associated with Set2. Immunoblotting and immunoprecipitation using antibodies against subunits of pol II confirmed that the phosphorylated form of pol II is indeed an integral part of the Set2 complex. Gst-Set2 preferentially binds to CTD synthetic peptides phosphorylated at serine 2, and to a lesser extent, serine 5 phosphorylated peptides, but has no affinity for unphosphorylated CTD, suggesting that Set2 associates with the elongating form of the pol II. Furthermore, we show that set2Delta ppr2Delta double mutants (PPR2 encodes TFIIS, a transcription elongation factor) are synthetically hypersensitive to 6-azauracil, and that deletions in the CTD reduce in vivo levels of H3 lysine 36 methylation. Collectively, these results suggest that Set2 is involved in regulating transcription elongation through its direct contact with pol II.

Antimetabolites↗

Review of proteomics with applications to genetic epidemiology.

Mapping of the human genome has the potential to transform the traditional methods of genetic epidemiology. The complete draft sequence of the 3.3 billion nucleotides comprising the genome is now available over the Internet, including the location and nearly complete sequence of the 26,000 to 31,000 protein-encoding genes. However, aside from water, almost everything in the human body is either made of, or by, proteins. Although the DNA code provides the instructions for their amino acid sequence, there are an estimated 1.5 million proteins. Thus, the correlation between DNA sequence and protein is low, reflecting alternate splicing as well as post-translational modification. The purpose of this article is to explore ways in which the emerging field of proteomics, the study of proteins in a cell, may inform our approach to gene mapping. This article reviews the various technical approaches currently available for proteomics. Technologies are available to quantify protein expression (and compare normal versus disease states), identify proteins through comparison with sequence information in databases or direct sequencing (which can then be mapped to chromosomal locations to ensure appropriate markers), elucidate protein-protein interactions (which may underlie disease), determine localization of proteins within the cell (abnormal trafficking of proteins could have an inherited basis), and characterize modifications of proteins (which is relevant to modifier gene candidates). Several examples are presented to illustrate the potential application of proteomics to the field of genetic epidemiology, and we conclude with various considerations regarding design and analysis.

Breast Neoplasms↗

Mutations in the CACNA1F and NYX genes in British CSNBX families.

X-linked congenital stationary night blindness (CSNBX) is a genetically and phenotypically heterogeneous non-progressive disorder, characterised by impaired night vision but grossly normal retinal appearance. Other more variable features include reduction in visual acuity, myopia, nystagmus and strabismus. Genetic mapping studies by other groups, and our own studies of British patients, identified key recombination events indicating the presence of at least 2 disease genes on Xp11. Two causative genes (CACNA1F and NYX) for CSNBX have now been identified through positional cloning strategies. In this report, we present the results of comprehensive mutation screening in 14 CSNBX families, three with mutations in the CACNA1F gene and 10 with mutations in the NYX gene. In one family we failed to identify the mutation after testing RP2, RPGR, NYX and CACNA1F. NYX gene mutations are a more frequent cause of CSNBX, although there is evidence for founder mutations. Our report of patient population mutation screening for both CSNBX genes, and our exclusion of RP2 and RGPR, indicates that mutations in CACNA1F and NYX are likely to account for all CSNBX.

Calcium Channels↗

Analysis of the Shewanella oneidensis proteome by two-dimensional gel electrophoresis under nondenaturing conditions.

Proteomes are dynamic, i.e., the protein components of living cells change in response to various stimuli. Protein changes can involve shifts in the abundance of protein components, in the interactions of protein components, and in the activity of protein components. Two-dimensional gel electrophoresis (2-DE) coupled with peptide mass spectrometry is useful for the analysis of relative protein abundance, but the denaturing conditions of classical 2-DE do not allow analysis of protein interactions or protein function. We have developed a nondenaturing 2-DE method that allows analysis of protein interactions and protein functions, as demonstrated in our analysis of the cytosol and crude membrane fractions of the facultative anaerobe Shewanella oneidensis MR-1. Our experiments demonstrate that enzymatic activity is retained under the sample and protein separation methods described, as shown by positive malate dehydrogenase activity results. We have also found protein interactions within both the soluble and membrane fractions. The method described will be useful for the characterization of the functional proteomes of microbial systems.

Bacterial Proteins↗

High-throughput functional affinity purification of mannose binding proteins from Oryza sativa.

We have used affinity chromatography in combination with mass spectrometry to isolate, identify, and assign a preliminary functional annotation to a large number of both known and novel proteins from rice. Rice (Oryza sativa) leaf, root, and seed tissue extracts were fractionated by column affinity chromatography using alpha-D-mannose as the ligand. Bound fractions were eluted and subjected to one-dimensional electrophoresis, followed by high-performance liquid chromatography-tandem mass spectrometric analysis of separated proteins. This multiplexed technology resulted in the isolation and identification of 136 distinct mannose binding proteins from rice. A comparative analysis demonstrates very little overlap of identified proteins between the respective tissues, and confirms the correctly compartmentalized presence of a significant number of proteins from largely tissue-specific biochemical pathways. Over 30% of the identified proteins with a previously annotated function are directly involved in sugar metabolism, including several highly expressed known rice lectins. Direct comparison of the peptide sequences identified in this study to those peptides identified in the most comprehensive survey of the rice proteome to date indicates that our current data represents a significant enrichment of proteins unique to this dataset. Nearly 15% of the identified proteins, identified on the basis of exact peptide matching to sequences in the rice genomic database, represent proteins without a previously known functional annotation, indicating the potential of this combined chromatographic approach to assign a preliminary function to novel proteins in a high-throughput fashion.

Binding, Competitive↗

Investigative proteomics: identification of an unknown plant virus from infected plants using mass spectrometry.

We describe the identification of a previously uncharacterized plant virus that is capable of infecting Nicotiana spp. and Arabidopsis thaliana. Protein extracts were first prepared from leaf tissue of uninfected tobacco plants, and the proteins were visualized with two-dimensional electrophoresis (2-DE). Matching gels were then run using protein extracts of a tobacco plant infected with tobacco mosaic virus (TMV). After visual comparison, the proteins spots that were differentially expressed in infected plant tissues were cut from the gels and analyzed by high performance liquid chromatography-tandem mass spectrometry (HPLC-MS/MS). Tandem mass spectrometry data of individual peptides was searched with SEQUEST. Using this approach we demonstrated a successful proof-of-concept experiment by identifying TMV proteins present in the total protein extract. The same procedure was then applied to tobacco plants infected with a laboratory viral isolate of unknown identity. Several of the differentially expressed protein spots were identified as proteins of potato virus X (PVX), thus successfully identifying the causative agent of the uncharacterized viral infection. We believe this demonstrates that HPLC-MS/MS can be used to successfully characterize unknown viruses in infected plants.

Amino Acid Sequence↗