Search PubMedSearch

Biomedical subjects

Lloyd M Smith

Publications and source records attributed to Lloyd M Smith.

2 recordsLinked to original sources

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Improved Detection of Differentially Abundant Proteins through FDR-Control of Peptide-Identity-Propagation.

The goal of proteomics is to identify and quantify peptides and proteins within a biological sample. Almost all algorithms for the identification of peptides in LC-MS/MS data employ two steps: peptide/spectrum matching and peptide-identity-propagation (PIP), also known as match-between-runs. PIP can routinely account for up to 40% of all results, with that proportion rising as high as 75% in single-cell proteomics. Unlike peptide identities derived through peptide/spectrum matches, for which error estimation has been strictly enforced for decades, peptide identities derived through PIP have not historically been subject to statistical evaluation. As an indispensable component of label-free quantification, PIP needs a statistically rigorous method for estimating its false-discovery rate (FDR). We present a method for FDR control of PIP, called PIP-ECHO, and devise a rigorous protocol for evaluating FDR control of any PIP method. Using three different benchmark data sets, we evaluate PIP-ECHO alongside the PIP procedures implemented by FlashLFQ, IonQuant, and MaxQuant. These analyses show that only PIP-ECHO can accurately control the FDR of PIP at 1% across all data sets. When analyzing a spike-in data set, PIP-ECHO increases both the accuracy and sensitivity of differential expression analysis, yielding substantially more differentially abundant proteins than either MaxQuant or IonQuant.

Proteomics