Search PubMed⌕ Search

Biomedical subjects

Steven A Carr

Publications and source records attributed to Steven A Carr.

15 recordsLinked to original sources

Expanding the human proteome with microproteins and peptideins.

A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1-3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of 'peptideins' as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome.

Humans↗

High-quality peptide evidence for annotating non-canonical open reading frames as human proteins.

A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.

GENCODE↗

The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease.

To pursue a systematic approach to the discovery of functional connections among diseases, genetic perturbation, and drug action, we have created the first installment of a reference collection of gene-expression profiles from cultured human cells treated with bioactive small molecules, together with pattern-matching software to mine these data. We demonstrate that this "Connectivity Map" resource can be used to find connections among small molecules sharing a mechanism of action, chemicals and physiological processes, and diseases and drugs. These results indicate the feasibility of the approach and suggest the value of a large-scale community Connectivity Map project.

Alzheimer Disease↗

mSin1 is necessary for Akt/PKB phosphorylation, and its isoforms define three distinct mTORC2s.

The mammalian target of rapamycin (mTOR) is a serine/threonine kinase that participates in at least two distinct multiprotein complexes, mTORC1 and mTORC2 . These complexes play important roles in the regulation of cell growth, proliferation, survival, and metabolism. mTORC2 is a hydrophobic motif kinase for the cell-survival protein Akt/PKB and, here, we identify mSin1 as a component of mTORC2 but not mTORC1. mSin1 is necessary for the assembly of mTORC2 and for its capacity to phosphorylate Akt/PKB. Alternative splicing generates at least five isoforms of the mSin1 protein , three of which assemble into mTORC2 to generate three distinct mTORC2s. Even though all mTORC2s can phosphorylate Akt/PKB in vitro, insulin regulates the activity of only two of them. Thus, we propose that cells contain several mTORC2 flavors that may phosphorylate Akt/PKB in response to different signals.

Adaptor Proteins, Signal Transducing↗

PEPPeR, a platform for experimental proteomic pattern recognition.

Quantitative proteomics holds considerable promise for elucidation of basic biology and for clinical biomarker discovery. However, it has been difficult to fulfill this promise due to over-reliance on identification-based quantitative methods and problems associated with chromatographic separation reproducibility. Here we describe new algorithms termed "Landmark Matching" and "Peak Matching" that greatly reduce these problems. Landmark Matching performs time base-independent propagation of peptide identities onto accurate mass LC-MS features in a way that leverages historical data derived from disparate data acquisition strategies. Peak Matching builds upon Landmark Matching by recognizing identical molecular species across multiple LC-MS experiments in an identity-independent fashion by clustering. We have bundled these algorithms together with other algorithms, data acquisition strategies, and experimental designs to create a Platform for Experimental Proteomic Pattern Recognition (PEPPeR). These developments enable use of established statistical tools previously limited to microarray analysis for treatment of proteomics data. We demonstrate that the proposed platform can be calibrated across 2.5 orders of magnitude and can perform robust quantification of ratios in both simple and complex mixtures with good precision and error characteristics across multiple sample preparations. We also demonstrate de novo marker discovery based on statistical significance of unidentified accurate mass components that changed between two mixtures. These markers were subsequently identified by accurate mass-driven MS/MS acquisition and demonstrated to be contaminant proteins associated with known proteins whose concentrations were designed to change between the two mixtures. These results have provided a real world validation of the platform for marker discovery.

Algorithms↗

Systematic identification of human mitochondrial disease genes through integrative genomics.

The majority of inherited mitochondrial disorders are due to mutations not in the mitochondrial genome (mtDNA) but rather in the nuclear genes encoding proteins targeted to this organelle. Elucidation of the molecular basis for these disorders is limited because only half of the estimated 1,500 mitochondrial proteins have been identified. To systematically expand this catalog, we experimentally and computationally generated eight genome-scale data sets, each designed to provide clues as to mitochondrial localization: targeting sequence prediction, protein domain enrichment, presence of cis-regulatory motifs, yeast homology, ancestry, tandem-mass spectrometry, coexpression and transcriptional induction during mitochondrial biogenesis. Through an integrated analysis we expand the collection to 1,080 genes, which includes 368 novel predictions with a 10% estimated false prediction rate. By combining this expanded inventory with genetic intervals linked to disease, we have identified candidate genes for eight mitochondrial disorders, leading to the discovery of mutations in MPV17 that result in hepatic mtDNA depletion syndrome. The integrative approach promises to better define the role of mitochondria in both rare and common human diseases.

Base Sequence↗

Protein biomarker discovery and validation: the long and uncertain path to clinical utility.

Better biomarkers are urgently needed to improve diagnosis, guide molecularly targeted therapy and monitor activity and therapeutic response across a wide spectrum of disease. Proteomics methods based on mass spectrometry hold special promise for the discovery of novel biomarkers that might form the foundation for new clinical blood tests, but to date their contribution to the diagnostic armamentarium has been disappointing. This is due in part to the lack of a coherent pipeline connecting marker discovery with well-established methods for validation. Advances in methods and technology now enable construction of a comprehensive biomarker pipeline from six essential process components: candidate discovery, qualification, verification, research assay optimization, biomarker validation and commercialization. Better understanding of the overall process of biomarker discovery and validation and of the challenges and strategies inherent in each phase should improve experimental study design, in turn increasing the efficiency of biomarker development and facilitating the delivery and deployment of novel clinical tests.

Animals↗

Mapping posttranslational modifications of proteins by MS-based selective detection: application to phosphoproteomics.

This chapter outlines general principals that apply to the analysis of posttranslational modifications of proteins, with an emphasis on phosphoproteins. Mass spectrometry (MS)-based approaches for selective detection and site-specific analysis of posttranslationally modified peptides are described, and an MS-based method that relies on production and detection of fragment ions specific for the modification(s) of interest and that was developed in the authors' laboratory is described in detail. The method is applicable to selective detection of N- and O-linked carbohydrates in glycoproteins, O-linked sulfate, and N- and O-linked lipids. Detailed procedures for application of this strategy to phosphorylation-site mapping are presented here.

Amino Acid Sequence↗

Phosphorylation by cyclin B-Cdk underlies release of mitotic exit activator Cdc14 from the nucleolus.

Budding yeast protein phosphatase Cdc14 is sequestered in the nucleolus in an inactive state during interphase by the anchor protein Net1. Upon entry into anaphase, the Cdc14 early anaphase release (FEAR) network initiates dispersal of active Cdc14 throughout the cell. We report that the FEARnetwork promotes phosphorylation of Net1 by cyclin-dependent kinase (Cdk) complexed with cyclin B1 or cyclin B2. These phosphorylations appear to be required for FEAR and sustain the proper timing of late mitotic events. Thus, a regulatory circuit exists to ensure that the arbiter of the mitotic state, Cdk, sets in motion events that culminate in exit from mitosis.

Anaphase↗

TULA: an SH3- and UBA-containing protein that binds to c-Cbl and ubiquitin.

Downregulation of protein tyrosine kinases is a major function of the multidomain protein c-Cbl. This effect of c-Cbl is critical for both negative regulation of normal physiological stimuli and suppression of cellular transformation. In spite of the apparent importance of these effects of c-Cbl, their own regulation is poorly understood. To search for possible novel regulators of c-Cbl, we purified a number of c-Cbl-associated proteins by affinity chromatography and identified them by mass spectrometry. Among them, we identified the UBA- and SH3-containing protein T-cell Ubiquitin LigAnd (TULA), which can also bind to ubiquitin. Functional studies in a model system based on co-expression of TULA, c-Cbl, and EGF receptor in 293T cells demonstrate that TULA is capable of inhibiting c-Cbl-mediated downregulation of EGF receptor. Furthermore, modulation of TULA concentration in Jurkat T-lymphoblastoid cells demonstrates that TULA upregulates the activity of both Zap kinase and NF-AT transcription factor. Therefore, our study indicates that TULA counters the inhibitory effect of c-Cbl on protein tyrosine kinases and, thus, may be involved in the regulation of biological effects of c-Cbl. Finally, our results suggest that TULA-mediated inhibition of the effects of c-Cbl on protein tyrosine kinases is caused by TULA-induced ubiquitylation and degradation of c-Cbl.

Amino Acid Substitution↗

Selective detection of glycopeptides on ion trap mass spectrometers.

Generation of carbohydrate-specific marker ions during LC-ESMS of digested glycoproteins has been demonstrated to be a highly selective and sensitive approach for detection of glycopeptides. In principle, any mass spectrometer can produce and selectively detect carbohydrate marker ions provided that the instrument is capable of collisional excitation in the region prior to the first mass analyzer sufficient to form abundant oxonium ions. This approach has yet to be demonstrated on 3D ion trap mass spectrometers, which have become widely used for proteomic applications. Here we report the successful development and optimization of carbohydrate marker ion detection on a LCQ Deca 3D ion trap utilizing this scan function. Human alpha-1 acid glycoprotein and a therapeutic monoclonal antibody were chosen to illustrate this methodology. Marker ion detection during LC-ESMS facilitated collection of glycopeptide-containing fractions. Analysis of the glycopeptides in these fractions by MS identified the specific glycosylation sites and enabled the prediction of the family of glycoforms at each attachment site. Using these optimized conditions, marker ion detection and glycopeptide analysis could be achieved with as little as 10 pmol of a glycoprotein.

Amino Acid Sequence↗

Improved sensitivity for phosphopeptide mapping using capillary column HPLC and microionspray mass spectrometry: comparative phosphorylation site mapping from gel-derived proteins.

Reversible protein phosphorylation regulates many cellular processes. Understanding how phosphorylation controls a given pathway usually involves specific knowledge of which amino acid residues are phosphorylated on a given protein. This is often a nontrivial task. In addition to the difficulties involved in purifying sufficient amounts of any given protein, most phosphoproteins contain multiple, substoichiometric sites of phosphorylation. In this paper, we describe substantial improvements made to our previously reported multidimensional electrospray MS-based phosphopeptide mapping technique that have resulted in a 20-fold increase in sensitivity for the overall process. Chief among these improvements are the incorporation of capillary chromatography and a microionspray source for the mass spectrometer into the first dimension of the analysis. In the first dimension of the process, phosphopeptides present in the proteolytic digest of a protein are selectively detected and collected into fractions during on-line LC/ESMS, which monitors for phosphopeptide specific marker ions. The phosphopeptide containing fractions are then analyzed in the second dimension by either MALDI-PSD or nano-ES with precursor ion scanning. The relative merits and limitations of these two techniques for phosphopeptide detection are demonstrated. The enhancement in sensitivity of the method under the new experimental conditions makes it suitable for phosphorylation mapping (from selective detection through sequencing) on gel-separated phosphoproteins where the level of phosphorylation at any given site is <200 fmol. Furthermore, this method detects serine, threonine, and tyrosine phosphorylation equally well. We have successfully employed this new configuration to map 11 in vivo sites of phosphorylation on the Saccharomyces cerevisiae protein kinase YAK1. YAK1 peptides containing all five YAK1 PKA consensus sites are phosphorylated, suggesting that YAK1 is an in vivo substrate for PKA. In addition, four peptides containing cdk sites and the autophosphorylation site at Tyr530 were found to be phosphorylated. Because the first dimension of this method generates a phosphorylation profile that can be used for a semiquantitative evaluation of site specific phosphoxylation, we evaluated its ability to detect site-specific changes in the phosphorylation profile of a protein in response to altered cellular conditions. This comparative phosphopeptide mapping strategy allowed us to detect a change in phosphorylation stoichiometry on the motor protein myosin-V in response to treatment with either mitotic or interphase Xenopus egg extracts and to identify the single functionally significant phosphorylation site that regulates myosin-V cargo binding.

Chromatography, High Pressure Liquid↗

N-Terminal peptide labeling strategy for incorporation of isotopic tags: a method for the determination of site-specific absolute phosphorylation stoichiometry.

Determining the phosphorylation stoichiometry at specific sites in a phosphoprotein is a very challenging task. We describe here a novel mass spectrometry based method that is capable of measuring the absolute phosphorylation stoichiometry at specific sites without the need for specific internal standards, phospho-site antibodies or radioactivity. The method is based on a gentle chemical labeling strategy which specifically and differentially labels the N-terminus of all peptides in a sample with either a D(5)- or D(0)-propionyl group and measures the ratio of the abundance of the D(5)/D(0) peptide pairs simultaneously using mass spectrometry. Using matrix-assisted laser desorption/ionization (MALDI), the method can measure absolute stoichiometry to within at least 10% and can be applied to both in vitro and in vivo phosphorylated peptides and proteins. Furthermore, this method can potentially be applied to the quantitative study of other types of protein post-translational modifications, and the profiling of protein expression on the proteome level.

Amino Acid Sequence↗

Mass spectrometry-based methods for phosphorylation site mapping of hyperphosphorylated proteins applied to Net1, a regulator of exit from mitosis in yeast.

Prior to anaphase in Saccharomyces cerevisiae, Cdc14 protein phosphatase is sequestered within the nucleolus and inhibited by Net1, a component of the RENT complex in budding yeast. During anaphase the RENT complex disassembles, allowing Cdc14 to migrate to the nucleus and cytoplasm where it catalyzes exit from mitosis. The mechanism of Cdc14 release appears to involve the polo-like kinase Cdc5, which is capable of promoting the dissociation of a recombinant Net1.Cdc14 complex in vitro by phosphorylation of Net1. We report here the phosphorylation site mapping of recombinant Net1 (Net1N) and a mutant Net1N allele (Net1N-19m) with 19 serines or threonines mutated to alanine. A variety of chromatographic and mass spectrometric-based strategies were used, including immobilized metal-affinity chromatography, alkaline phosphatase treatment, matrix-assisted laser-desorption post-source decay, and a multidimensional electrospray mass spectrometry-based approach. No one approach was able to identify all phosphopeptides in the tryptic digests of these proteins. Most notably, the presence of a basic residue near the phosphorylated residue significantly hampered the ability of alkaline phosphatase to hydrolyze the phosphate moiety. A major goal of research in proteomics is to identify all proteins and their interactions and post-translational modification states. The failure of any single method to identify all sites in highly phosphorylated Net1N, however, raises significant concerns about how feasible it is to map phosphorylation sites throughout the proteome using existing technologies.

Amino Acid Sequence↗

Place of pattern in proteomic biomarker discovery.

The role of pattern in biomarker discovery and clinical diagnosis is examined in its historical context. The use of MS-derived pattern is treated as a logical extension of prior applications of non-MS-derived pattern. Criticisms pertaining to specific technology platforms and analytic methodologies are considered separately from the larger issues of pattern utility and deployment in biomarker discovery. We present a hybrid strategy that marries the desirable attributes of high-information content MS pattern with the capability to obtain identity, and explore the key steps in establishing a data analysis pipeline for pattern-based biomarker discovery.

Biomarkers↗