Search PubMed⌕ Search

Biomedical subjects

Sandra Orchard

Publications and source records attributed to Sandra Orchard.

At least 19 recordsLinked to original sources

Expanding the human proteome with microproteins and peptideins.

A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1-3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of 'peptideins' as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome.

Humans↗

High-quality peptide evidence for annotating non-canonical open reading frames as human proteins.

A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.

GENCODE↗

Autumn 2005 Workshop of the Human Proteome Organisation Proteomics Standards Initiative (HUPO-PSI) Geneva, September, 4-6, 2005.

The autumn workshop of the Proteomics Standards Initiative of the Human Proteomics Organisation met to further advance the development of the existing standards in the fields of molecular interactions and mass spectrometry. In addition, new areas were addressed, in particular developing standards for the description and exchange of data from gel electrophoresis experiments. The General Proteomics Standards group is now working closely with the FuGE (Functional Genomics Experiment) efforts to define a general standard in which to encode data that will enable a systems biology approach to data analysis. Common to all these efforts is the field of protein modifications, and work has been initiated to establish an ontology in this field that can be used by both workers in the field of proteomics and the wider scientific community.

Databases, Protein↗

Annotating the human proteome.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner.

Databases, Factual↗

InterPro, progress and status in 2005.

InterPro, an integrated documentation resource of protein families, domains and functional sites, was created to integrate the major protein signature databases. Currently, it includes PROSITE, Pfam, PRINTS, ProDom, SMART, TIGRFAMs, PIRSF and SUPERFAMILY. Signatures are manually integrated into InterPro entries that are curated to provide biological and functional information. Annotation is provided in an abstract, Gene Ontology mapping and links to specialized databases. New features of InterPro include extended protein match views, taxonomic range information and protein 3D structure data. One of the new match views is the InterPro Domain Architecture view, which shows the domain composition of protein matches. Two new entry types were introduced to better describe InterPro entries: these are active site and binding site. PIRSF and the structure-based SUPERFAMILY are the latest member databases to join InterPro, and CATH and PANTHER are soon to be integrated. InterPro release 8.0 contains 11 007 entries, representing 2573 domains, 8166 families, 201 repeats, 26 active sites, 21 binding sites and 20 post-translational modification sites. InterPro covers over 78% of all proteins in the Swiss-Prot and TrEMBL components of UniProt. The database is available for text- and sequence-based searches via a webserver (http://www.ebi.ac.uk/interpro), and for download by anonymous FTP (ftp://ftp.ebi.ac.uk/pub/databases/interpro).

Databases, Protein↗

Further steps towards data standardisation: the Proteomic Standards Initiative HUPO 3(rd) annual congress, Beijing 25-27(th) October, 2004.

The increasing volume of proteomics data currently being generated by increasingly high-throughput methodologies has led to an increasing need for methods by which such data can be accurately described, stored and exchanged between experimental researchers and data repositories. Work by the Proteomics Standards Initiative of the Human Proteome Organisation has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. The progress of these efforts, and the direct benefits already accruing from them, were described at a plenary session of the 3(rd) Annual HUPO congress. Parallel sessions allowed the three work groups to present their progress to interested parties and to collect feedback from groups already implementing the available formats.

China↗

Further steps in standardisation. Report of the second annual Proteomics Standards Initiative Spring Workshop (Siena, Italy 17-20th April 2005).

The spring workshop of the HUPO-PSI convened in Siena to further progress the data standards which are already making an impact on data exchange and deposition in the field of proteomics. Separate work groups pushed forward existing XML standards for the exchange of Molecular Interaction data (PSI-MI, MIF) and Mass Spectrometry data (PSI-MS, mzData) whilst significant progress was made on PSI-MS' mzIdent, which will allow the capture of data from analytical tools such as peak list search engines. A new focus for PSI (GPS, gel electrophoresis) was explored; as was the need for a common representation of protein modifications by all workers in the field of proteomics and beyond. All these efforts are contextualised by the work of the General Proteomics Standards workgroup; which in addition to the MIAPE reporting guidelines, is continually evolving an object model (PSI-OM) from which will be derived the general standard XML format for exchanging data between researchers, and for submission to repositories or journals.

Mass Spectrometry↗

Human Proteome Organisation Proteomics Standards Initiative. Pre-Congress Initiative.

The plenary session of the Proteomics Standards Initiative of the Human Proteome Organisation discussed the current status of the ongoing work in the fields of molecular interactions, mass spectrometry and the description of protein modifications. In addition, new areas are being opened up, in particular developing standards for the description and exchange of data from gel electrophoresis experiments. The General Proteomics Standards group is now working closely with the Functional Genomics Experiment efforts to define a general standard in which to encode data that will enable a systems biology approach to data analysis.

Databases, Genetic↗

The use of common ontologies and controlled vocabularies to enable data exchange and deposition for complex proteomic experiments.

Controlled vocabularies provide a roadmap through complex biological data. Proteomic data is increasing in volume and is currently poorly served by public repositories due to the large number of different formats in which the data is generated and stored. The Human Proteome Organization Proteome Standards Initiative is establishing standards for data transfer and deposition. These standards utilize ontologies and controlled vocabularies to describe experimental procedures and common processes such as sample preparation This paper will discuss the development of such ontologies by the user community and their current utilization in the fields of protein:proein interactions and mass spectrometry.

Computational Biology↗

IntAct: an open source molecular interaction database.

IntAct provides an open source database and toolkit for the storage, presentation and analysis of protein interactions. The web interface provides both textual and graphical representations of protein interactions, and allows exploring interaction networks in the context of the GO annotations of the interacting proteins. A web service allows direct computational access to retrieve interaction networks in XML format. IntAct currently contains approximately 2200 binary and complex interactions imported from the literature and curated in collaboration with the Swiss-Prot team, making intensive use of controlled vocabularies to ensure data consistency. All IntAct software, data and controlled vocabularies are available at http://www.ebi.ac.uk/intact.

Animals↗

Common interchange standards for proteomics data: Public availability of tools and schema.

The Proteomics Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparision, exchange and verification. To this end, a Level 1 Molecular Interaction XML data exchange format has been developed which has been accepted for publication and is freely available at the PSI website (http.//psidev.sf.net/). Several major protein interaction databases are already making data available in this format. A draft XML interchange format for mass spectrometry data has been written and is currently undergoing evaluation whilst work is ongoing to develop a proteomics data integration model, MIAPE.

Computational Biology↗

Advances in the development of common interchange standards for proteomic data.

The generation of proteomics data is increasingly high-throughput and high volume. Both experimental design and the technologies used to produce and subsequently analyze the data are becoming ever more complex. An increasing need for methods by which such data can be accurately described, stored and exchanged between experimenters and data repositories has been recognised. Work by the Proteomics Standards Initiative of the Human Proteome Organisation has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. At a recent workshop in Nice, participants gathered to review the progress made to date and assist in pushing the process still further forward.

Humans↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

Current status of proteomic standards development.

The generation of proteomic data is becoming ever more high throughput. Both the technologies and experimental designs used to generate and analyze data are becoming increasingly complex. The need for methods by which such data can be accurately described, stored and exchanged between experimenters and data repositories has been recognized. Work by the Proteome Standards Initiative of the Human Proteome Organization has laid the foundation for the development of standards by which experimental design can be described and data exchange facilitated. The Minimum Information About a Proteomic Experiment data model describes both the scope and purpose of a proteomics experiment and encompasses the development of more specific interchange formats such as the mzData model of mass spectrometry. The eXtensible Mark-up Language-MI data interchange format, which allows exchange of molecular interaction data, has already been published and major databases within this field are supplying data downloads in this format.

Databases, Protein↗

The proteomics standards initiative.

The Proteomics Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparison, exchange and verification. Progress has been made in the development of common standards for data exchange in the fields of both mass spectrometry and protein-protein interaction. A proteomics-specific extension is being created for the emerging American Society for Tests and Measurements mass spectrometry standard, which will be supported by manufacturers of both hardware and software. A data model for proteomics experimentation is under development and discussions on a public repository for published proteomics data are underway. The Protein-Protein Interactions group expects to publish the Level 1 PSI data exchange format for protein-protein interactions soon and discussions as to the content of Level 2 have been initiated.

Biochemistry↗

Further advances in the development of a data interchange standard for proteomics data.

The Protein Standards Initiative (PSI) aims to define community standards for data representation in proteomics and to facilitate data comparison, exchange and verification. Significant progress was made in advancing the design and implementation of a draft standard for exchanging experimental data from proteomics experiments involving mass spectrometry at the 51st Annual Conference of the American Society for Mass Spectrometry. In collaboration with the American Society for Tests and Measurements, the PSI propose to publish this first draft at the forthcoming HUPO 2nd World Congress in Montreal, 8-11 October 2003.

Computational Biology↗