Search PubMed⌕ Search

Biomedical subjects

Ronald C Beavis

Publications and source records attributed to Ronald C Beavis.

11 recordsLinked to original sources

Use of peptide retention time prediction for protein identification by off-line reversed-phase HPLC-MALDI MS/MS.

A new algorithm, sequence-specific retention calculator, was developed to predict retention time of tryptic peptides during RP HPLC fractionation on C18, 300-A pore size columns. Correlations of up to approximately 0.98 R2 value were obtained for a test library of approximately 2000 peptides and approximately 0.95-0.97 for a variety of real samples. The algorithm was applied in conjunction with an exclusion protocol based on mass (15 ppm tolerance) and retention time (2-min tolerance for 0.66% acetonitrile/min gradient), MART criteria to significantly reduce the instrument time required for complete MS/MS analysis of a digest separated by RP HPLC. This was confirmed by reanalyzing the set of HPLC-MALDI MS/MS data with no loss in protein identifications, despite the number of virtually executed MS/MS analyses being decreased by 57%.

Cell Line, Tumor↗

General framework for developing and evaluating database scoring algorithms using the TANDEM search engine.

MOTIVATION: Tandem mass spectrometry (MS/MS) identifies protein sequences using database search engines, at the core of which is a score that measures the similarity between peptide MS/MS spectra and a protein sequence database. The TANDEM application was developed as a freely available database search engine for the proteomics research community. To extend TANDEM as a platform for further research on developing improved database scoring methods, we modified the software to allow users to redefine the scoring function and replace the native TANDEM scoring function while leaving the remaining core application intact. Redefinition is performed at run time so multiple scoring functions are available to be selected and applied from a single search engine binary. We introduce the implementation of the pluggable scoring algorithm and also provide implementations of two TANDEM compatible scoring functions, one previously described scoring function compatible with PeptideProphet and one very simple scoring function that quantitative researchers may use to begin their development. This extension builds on the open-source TANDEM project and will facilitate research into and dissemination of novel algorithms for matching MS/MS spectra to peptide sequences. The pluggable scoring schema is also compatible with related search applications P3 and Hunter, which are part of the X! suite of database matching algorithms. The pluggable scores and the X! suite of applications are all written in C++. AVAILABILITY: Source code for the scoring functions is available from http://proteomics.fhcrc.org

Algorithms↗

Using the global proteome machine for protein identification.

This chapter describes the use of an open-source, freely available informatics system for the identification of proteins using tandem mass spectra of peptides derived from an enzymatic digest of a mixture of mature proteins. The chapter describes the use of features of the Global Proteome Machine (GPM) interface that assist in making comprehensive assignments between spectra and sequences, including the detection of point mutations, posttranslational modifications, and experimental artifacts. The use of this interface to validate results using the GPM Database is also described. This data repository allows analysts to compare their own results to those obtained by other scientists to determine the degree to which their data are consistent with previous measurements.

Amino Acid Sequence↗

The use of proteotypic peptide libraries for protein identification.

This paper describes an algorithm to apply proteotypic peptide sequence libraries to protein identifications performed using tandem mass spectrometry (MS/MS). Proteotypic peptides are those peptides in a protein sequence that are most likely to be confidently observed by current MS-based proteomics methods. Libraries of proteotypic peptide sequences were compiled from the Global Proteome Machine Database for Homo sapiens and Saccharomyces cerevisiae model species proteomes. These libraries were used to scan through collections of tandem mass spectra to discover which proteins were represented by the data sets, followed by detailed analysis of the spectra with the full protein sequences corresponding to the discovered proteotypic peptides. This algorithm (Proteotypic Peptide Profiling, or P3) resulted in sequence-to-spectrum matches comparable to those obtained by conventional protein identification algorithms using only full protein sequences, with a 20-fold reduction in the time required to perform the identification calculations. The proteotypic peptide libraries, the open source code for the implementation of the search algorithm and a website for using the software have been made freely available. Approximately 4% of the residues in the H. sapiens proteome were required in the proteotypic peptide library to successfully identify proteins.

Algorithms↗

TANDEM: matching proteins with tandem mass spectra.

SUMMARY: Tandem mass spectra obtained from fragmenting peptide ions contain some peptide sequence specific information, but often there is not enough information to sequence the original peptide completely. Several proprietary software applications have been developed to attempt to match the spectra with a list of protein sequences that may contain the sequence of the peptide. The application TANDEM was written to provide the proteomics research community with a set of components that can be used to test new methods and algorithms for performing this type of sequence-to-data matching. AVAILABILITY: The source code and binaries for this software are available at http://www.proteome.ca/opensource.html, for Windows, Linux and Macintosh OSX. The source code is made available under the Artistic License, from the authors.

Algorithms↗

A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes.

This paper investigates the use of survival functions and expectation values to evaluate the results of protein identification experiments. These functions are standard statistical measures that can be used to reduce various protein identification scoring schemes to a common, easily interpretably representation. The relative merits of scoring systems were explored using this approach, as well as the effects of altering primary identification parameters. We would advocate the widespread use of these simple statistical measures to simplify and standardize the reporting of the confidence of protein identification results, allowing the users of different identification algorithms to compare their results in a straightforward and statistically significant manner. A method is described for measuring these distributions using information that is being discarded by most protein identification search engines, resulting in accurate survival functions that are specific to any combination of scoring algorithms, sequence databases, and mass spectra.

Mass Spectrometry↗

A method for reducing the time required to match protein sequences with tandem mass spectra.

An algorithm for reducing the time necessary to match a large set of peptide tandem mass spectra with a list of protein sequences is described. This algorithm breaks the process into multiple steps. A rapid survey step identifies all protein sequences that are reasonable candidates for a match with a set of tandem mass spectra. These candidates are then used as models, which are refined by detailed analysis of the set of tandem mass spectra for evidence of incomplete enzymatic hydrolysis, non-specific hydrolysis and chemical modifications of amino acid residues resulting from either post-translational modifications or sample handling. Compared with current one-step methods for matching proteins to mass spectra, this multiple-step method can decrease the time required for the calculation by several orders of magnitude.

Algorithms↗

Informatics and data management in proteomics.

Proteomics has become dominated by large amounts of experimental data and interpreted results. This experimental data cannot be effectively used without understanding the fundamental structure of its information content and representing that information in such a way that knowledge can be extracted from it. This review explores the structure of this information with regard to three fundamental issues: the extraction of relevant information from raw data, the scale of the projects involved and the statistical significance of protein identification results.

Algorithms↗

RADARS, a bioinformatics solution that automates proteome mass spectral analysis, optimises protein identification, and archives data in a relational database.

RADARS, a rapid, automated, data archiving and retrieval software system for high-throughput proteomic mass spectral data processing and storage, is described. The majority of mass spectrometer data files are compatible with RADARS, for consistent processing. The system automatically takes unprocessed data files, identifies proteins via in silico database searching, then stores the processed data and search results in a relational database suitable for customized reporting. The system is robust, used in 24/7 operation, accessible to multiple users of an intranet through a web browser, may be monitored by Virtual Private Network, and is secure. RADARS is scalable for use on one or many computers, and is suited to multiple processor systems. It can incorporate any local database in FASTA format, and can search protein and DNA databases online. A key feature is a suite of visualisation tools (many available gratis), allowing facile manipulation of spectra, by hand annotation, reanalysis, and access to all procedures. We also described the use of Sonar MS/MS, a novel, rapid search engine requiring 40 MB RAM per process for searches against a genomic or EST database translated in all six reading frames. RADARS reduces the cost of analysis by its efficient algorithms: Sonar MS/MS can identifiy proteins without accurate knowledge of the parent ion mass and without protein tags. Statistical scoring methods provide close-to-expert accuracy and brings robust data analysis to the non-expert user.

Amino Acid Sequence↗

Open source system for analyzing, validating, and storing protein identification data.

This paper describes an open-source system for analyzing, storing, and validating proteomics information derived from tandem mass spectrometry. It is based on a combination of data analysis servers, a user interface, and a relational database. The database was designed to store the minimum amount of information necessary to search and retrieve data obtained from the publicly available data analysis servers. Collectively, this system was referred to as the Global Proteome Machine (GPM). The components of the system have been made available as open source development projects. A publicly available system has been established, comprised of a group of data analysis servers and one main database server.

Computational Biology↗

Lectin affinity as an approach to the proteomic analysis of membrane glycoproteins.

The aim was to determine the proportion of membrane glycoproteins captured using concanavalin A or wheat germ agglutinin lectin affinity chromatography. Digests of the isolated proteins were separated by reversed-phase liquid chromatography and analyzed by matrix-assisted laser desorption tandem mass spectrometry. The two lectins identified different groups of proteins with a broad range of molecular mass and p/ values, including a number of proteins that overlapped the two groups. Approximately 30% of the proteins were positively identified as containing domains that were predicted using standard bioinformatics methods to be characteristic of integral membrane proteins. This approach represents an effective method of surveying the membrane protein pool of mammalian cells for subsequent proteomic analysis.

Cell Membrane↗