Search PubMedSearch

SEARCH · Search PubMed

Results for “peptide identification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

On-filter fractionation by empFASP improves identification of membrane peptides in proteomic experiments.

Membrane proteins remain among the most analytically challenging targets in bottom-up proteomics due to their limited solubility and low abundance of protease-accessible sites within transmembrane domains. In addition, hydrophobic peptides are frequently lost during detergent removal and the on-filter processing steps. Here, we present empFASP, a straightforward on-filter-fractionation-based modification of the enhanced filter-aided sample preparation (eFASP) workflow that enhances recovery of membrane-embedded peptides otherwise lost during digestion and cleanup. The method combines controlled on-filter inversion with sequential ethyl acetate extraction at defined pH values, enabling recovery of peptide material retained on the filter and redistributed into detergent micelles. Compared with SP3 and SP4 in HEK293T lysates, empFASP increased unique hydrophobic peptide identifications by up to 48% and increased the proportion of detected transmembrane peptides. Application to mouse mitochondrial membranes and phosphatidylethanolamine-deficient and PE-containing Escherichia coli membranes showed that the additional fractions of empFASP contribute complementary recovery of hydrophobic and membrane-associated peptides, with the strongest gains observed at the peptide level. Because empFASP requires no specialized reagents or instrumentation, it can be readily implemented in standard proteomics workflows to improve coverage of membrane-embedded regions. SIGNIFICANCE: The empFASP (enhanced membrane peptide) workflow offers a practical solution to one of the persistent limitations in membrane proteomics-the underrepresentation of hydrophobic and transmembrane peptides in standard digests. By integrating simple pH-controlled extractions into an on-filter format, empFASP recovers peptides otherwise lost through adsorption or detergent micelle retention, substantially improving coverage of the membrane proteome. This method expands the analytical reach of bottom-up proteomics without requiring specialized instrumentation, making it immediately applicable for studies of membrane topology, protein-lipid interactions, and the structural consequences of altered membrane composition.

Proteomics

The use of synthetic peptide combinatorial libraries for the identification of bioactive peptides.

The systematic preparation of synthetic peptide combinatorial libraries (SPCLs), each composed of tens of millions of peptides that can be screened in existing diagnostically or pharmacologically relevant in vitro assay systems, is reviewed. The identification of optimal peptide sequences has been achieved through the screening in solution of SPCLs, each element of which is composed of more than 100,000 nonsupport-bound peptides in equimolar representation, along with an iterative synthesis and screening process. Examples are presented in which an SPCL, composed in total of 52,128,400 acetylated hexa-peptides, is used along with an iterative selection process to precisely identify the antigenic determinant of a peptide recognized by a monoclonal antibody using competitive enzyme-linked immunosorbent assay. This same library was also used to develop highly potent antimicrobial peptides in bacterial growth inhibition assays. A separate non-acetylated SPCL was used to screen and identify high affinity peptide ligands using an opiate radio-receptor binding assay.

Amino Acid Sequence

Intestinal surface peptide hydrolases: identification and characterization of three enzymes from rat brush border.

Peptide hydrolases were solubilized from rat small intestinal brush border by papain and separated by Sephadex G-200 chromatography, velocity gradient ultracentrifugation and polyacrylamide disc electrophoresis and designated according to approximate molecular size from sedimentation studies. Peptidases I (apparent Mr 230 000) and II (apparent Mr 160 000) are oligopeptidases with maximum specificity for tripeptides with identical pH optima (7.5) and similar apparent Km with L-Leu-Gly (I, 0.60 MM; II, 0.76 mM). L-Leucyl-beta-naphthylamide is a competitive inhibitor of both enzymes. Concentration of peptidase II produced partial conversion to peptidase I on polyacrylamide disc electrophoresis. The third peptide hydrolase (III, Mr 120 000) is a dipeptidase with pH optimum 8.5 and apparent Km for L-Leu-Gly of 0.65 mM. These peptide hydrolases were inhibited appreciably (37-59%) by 0.2 M glycine/NaOH, Tris - HCl or Tris - glycine buffers. EDTA (5 mM) completely inhibited these enzymes but all activity was restored by dialysis against buffer without divalent ions. Subsequent addition of Mg2+, Mn2+, Co2+ or Zn2+ (1-2 mM) inhibited peptidases I and II variably (4-81%) depending upon the substrate and buffer used. In contrast peptidase III was activated slightly by metal ions (5-20%). These peptide hydrolases are strategically located at the intestinal lumen-cell interface and possess biochemical characteristics making them ideally suited to play a pivotal role in the final stage of protein digestion.

Aminopeptidases

Investigation of the lectin-like binding domains in pertussis toxin using synthetic peptide sequences. Identification of a sialic acid binding site in the S2 subunit of the toxin.

Synthetic peptides corresponding to selected sequences in the S2 and S3 subunits of pertussis toxin were prepared and evaluated for their ability to inhibit the binding of biotinylated pertussis toxin and three biotinylated sialic acid specific plant lectins to fetuin and asialofetuin. The screening results indicated that two regions in the S2 subunit corresponding to amino acids 78-98 and 123-154 inhibited pertussis toxin binding to fetuin at submillimolar concentrations, while S3 sequences corresponding to amino acids 87-108 and 134-154 inhibited pertussis toxin-biotin binding to asialofetuin albeit with lower affinity. These results confirm earlier findings, which suggest that the S2 subunit is responsible for binding sialylated glycoconjugates. This was further confirmed by the ability of S2 peptides to inhibit the binding of the lectins from Maackia amurensis and wheat germ to fetuin. Two additional peptides from the S2 subunit of pertussis toxin corresponding to sequences 9-23 and 1-23 were found to contain within their sequences a 6-amino acid fragment which has strong homology with a sequence in wheat germ agglutinin that has been shown to be a component of the sialic acid binding site as determined by x-ray crystallography. One of these sequences from S2 (9-23) was biotinylated and evaluated for its ability to bind to carbohydrate. Through a series of experiments using fetuin, asialofetuin, asialoagalactofetuin, and simple saccharides, the biotinylated peptide was shown to bind with high affinity to sialic acid-containing glycoconjugates indicating that these sequences within the S2 subunit of pertussis toxin also play an important role in binding sialic acid.

Amino Acid Sequence

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

Ligand-independent oligomerization of natriuretic peptide receptors. Identification of heteromeric receptors and a dominant negative mutant.

Activation of many single-transmembrane receptors requires ligand-induced receptor oligomerization. We have examined the oligomerization of the atrial natriuretic peptide receptor, NPR-A, using epitope-tagged receptor in a co-immunoprecipitation assay. Unlike other single-transmembrane receptors, NPR-A oligomerized in a ligand-independent fashion. Extracellular receptor sequences were both necessary and sufficient for oligomer formation. NPR-A was also able to oligomerize with the related natriuretic peptide receptor, NPR-B. A truncated NPR-A lacking most of the cytoplasmic domain blocked activation of the full-length receptor, presumably through formation of an inactive heteromer. These results indicate that oligomerization of this single-transmembrane receptor is important for the transduction of a conformational change across the plasma membrane but are not consistent with models in which natriuretic peptide receptor oligomerization serves merely to bring intracellular domains together.

Amino Acid Sequence

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics

Strategies for identification of peptide growth factors.

Numerous peptides are known that have specific functions as growth factors in different tissues. These bioactive peptides are characterized by their ability to bind to high-affinity receptors, by their classification into superfamilies that share homology and function and by their synthesis as large precursor molecules that are processed to active forms. In some cases the precursors themselves also have biological activity. Modulation of growth factor activity at the level of the receptor or effector molecules has great therapeutic potential. This article will outline some of the strategies that have been successful in detecting and identifying growth factors and demonstrating their biological activity.

Animals

Large-scale discovery platform enables identification of peptides targeting drug-resistant candidiasis.

Natural products have an unparalleled track record as sources of clinical drugs. Among them, nonribosomal peptides (NRPs) stand as one of the most therapeutically significant classes, encompassing numerous approved anti-infective and anticancer agents. Yet, discovering bioactive NRPs remains profoundly challenging due to their complex biosynthesis and chemical architecture. Here, we present NPDiscover, a pathogen-oriented, scalable bioinformatics platform that integrates genome mining, metabolomics, and machine learning to identify NRPs active against drug-resistant pathogens. Applying NPDiscover to Actinobacteria datasets, we discovered edaphochelin A, a previously unreported NRP that kills multi-drug-resistant Candida auris and Candida glabrata by disrupting respiratory chain proteins. Structural elucidation via nuclear magnetic resonance and mass spectrometry, alongside in vitro and in vivo validation, confirmed its efficacy, safety, and a mode of action distinct from existing antifungals-establishing edaphochelin A as a compelling drug candidate and NPDiscover as a powerful engine for scalable natural product discovery.

CP: biotechnology

Characterization of a phagocyte cytochrome b558 91-kilodalton subunit functional domain: identification of peptide sequence and amino acids essential for activity.

The phagocyte NADPH oxidase is a multicomponent membrane-bound electron transport chain that catalyzes the reduction of O2 to superoxide. Cytochrome b558, the terminal electron donor to O2, is an integral membrane heterodimer containing 91- and 22-kDa subunits (gp91-phox and p22-phox, respectively). Synthetic peptides, whose amino acid sequences correspond to a gp91-phox carboxyl-terminal domain, inhibit superoxide production by blocking assembly of the oxidase from membrane and cytosol components. In this study, we examined the amino acid sequence requirements of a series of synthetic truncated gp91-phox peptides for inhibition of human neutrophil NADPH oxidase activation. RGVHFIF, corresponding to gp91-phox residues 559-565, was the minimum sequence capable of inhibiting superoxide generation. Contributions of individual amino acids to overall RGVHFIF inhibitory activity were determined by comparing the abilities of alanine-substituted RGVHFIF peptides to inhibit superoxide production. Substitution of alanine for arginine, valine, isoleucine, or either of the phenylalanines (but not glycine or histidine) within RGVHFIF resulted in loss of inhibitory activity. Synthetic gp91-phox carboxyl-terminal peptides are likely to be competitive inhibitors of the corresponding carboxyl-terminal domain of native gp91-phox by virtue of amino acid identity. We conclude that properties of arginine valine, isoleucine, and phenylalanine side chains within an RGVHFIF-containing domain of gp91-phox contribute significantly to cytochrome b558-mediated activation of the oxidase.

Amino Acid Sequence

[Studies on cytochrome c oxidase, I. Purification and characterization of bovine myocardial enzyme and identification of peptide chains in the complex].

As part of the preliminary work for the structural elucidation of cytochrome c oxidase, the enzyme complex was isolated from bovine heart muscle and characterised chemically. The enzyme contains 10-11 nmol haem a, and 12-13 nmol copper per mg protein. The solubilised active enzyme also contains 5% phospholipid, comprising about 2 mol each of cardiolipin and phosphatidylethanolamine per mol haem a. In addition, the preparation contains a small number of detergent molecules (Tween-80). Eight polypeptide components were isolated by preparative dodecylsulphate gel electrophoresis, gel filtration on Biogel P-60, and counter current distribution. The apparent molecular weights of these components were I - 36 000, II - 28 000 (21 000), III - 19 000, IV - 14 000, V - 12 500, VI - 11 000, VII - 10 000 and VIII - 6000. At least seven intact polypeptide chains contribute to the structure of the enzyme complex of the terminal oxidase. On the basis of amino acid analysis and end group determination, they can be divided into two groups. The high molecular weight peptides I -III are hydrophobic and their amino acid compositions differ markedly from those of known enzyme proteins, especially with respect to their contents of leucine and methionine. Components I and II have formyl methionine at their N-termini. They are therefore possibly mitochondrial membrane components from complex 4 of the respiratory chain. Polypeptides IV - VII resemble functional enzyme subunits in their amino acid composition. Some of them possess free N-termini (alanine). The low molecular weight component VIII is heterogeneous and contains the N-terminal amino acids isoleucine, serine and phenylelanine in non-stoichiometric amounts. Analysis gives a minimal protein molecular weight of 130 000 (65 000 per haem a) for the two haem and two copper-containing "monomers". The molecular weight of the moiety preliminarily defined as enzymatic is about 48 000. The chemical characterisation provides data for the strategy of the subsequent sequence analysis of the polypeptides.

Amino Acids

Identification by peptide analysis of the spectrin-binding protein in human erythrocytes.

One-dimensional and two-dimensional peptide-mapping techniques are used to identify the protein which gives rise to the 72,000 dalton alpha-chymotryptic fragment previously shown to be the membrane attachment site for spectrin. Peptide maps of the 72,000 dalton fragment are very different from maps of Bands 1, 2, 2.9, 3, 3.1, 4.1, and 4.2 and very similar to maps of the apparently closely homologous polypeptides, Bands 2.1, 2.2, 2.3, and 2.6. Limited proteolysis of erythrocyte membranes is shown to generate Band 3', another polypeptide which has been associated with spectrin-binding activity. Peptide maps of Band 3' are very similar to maps of Band 2.1, suggesting that Band 3' is also a proteolytic fragment of Band 2.1. It is concluded that Band 2.1 and possibly some or all of the other, related polypeptides which electrophorese in the 2 region is (are) the spectrin-binding protein(s) of the human erythrocyte.

Carrier Proteins

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics

Strategy for Simultaneous Multiomic Survey of N-Glycomic and Extracellular Matrix Proteome by Mass Spectrometry Imaging.

Recent advances in spatially resolved molecular profiling have positioned matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI) as a powerful platform for multiomic tissue analyses. However, conventional workflows that sequentially target distinct molecular classes are time- and resource-intensive, requiring repeated sequential sample preparation, imaging, and data integration. Here, we evaluate streamlined strategies for simultaneous or combined acquisition of N-glycan and collagen-derived peptide information using PNGase F and collagenase. In-solution studies demonstrate that simultaneous enzymatic digestion yields comparable peptide identifications and glycan profiles relative to traditional sequential workflows, with minimal impact on enzymatic specificity. On the basis of these findings, we developed and optimized MALDI-MSI protocols enabling either simultaneous enzyme application or sequential enzyme treatment with unified matrix deposition and single-pass imaging. While direct coapplication reduced image uniformity, a hybrid approach that used sequential enzyme deposition with combined imaging preserved spatial fidelity and spectral quality while significantly reducing processing and computational demands. Application to human tissues, including vertebral bone and ocular samples, highlights the utility of this workflow for fragile specimens and exploratory multiomic surveys. Collectively, these results establish a framework for integrated glycomic and proteomic imaging targeting the extracellular microenvironment, expanding multiomic MALDI-MSI analyses.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

Immunoglobulin D myeloma and amyloidosis: immunochemical and structural studies of Bence Jones and amyloid fibrillar proteins.

Urinary Bence Jones protein and amyloid fibril protein isolated from the subcutaneous tissue of a patient with IgD myeloma and associated amyloidosis were subjected to physicochemical and immunochemical identification. Peptide maps and amino-terminal tetrapeptide composition obtained from the two proteins were comparable. Immunochemical cross-reactivity between the two proteins, with other lambda-type amyloid and Bence Jones proteins, and with a serum component was demonstrated. The results suggest that the source of the amyloid fibril protein is an intact circulating light polypeptide chain as well as smaller amino-terminal fragments.

Aged

Mass Spectrometry-Based Proteomics for the Masses: Peptide and Protein Identification in the Hunt Laboratory During the 2000's.

There has been a rapid increase in the number of individuals utilizing mass spectrometry-based proteomics to study complex biological systems and questions since the start of the 2000's. Building off the advancements in ionization and liquid chromatography scientists continued to push towards technology that would enable in-depth analysis of biological specimen. Donald F Hunt and the Hunt laboratory were major contributors to this effort with their work on improving upon existing Fourier Transform MS, development of electron transfer dissociation, and continued work on ion-ion reactions to improve intact protein analysis. Collaboration with other instrumentation laboratories and instrument companies led to the sharing of technology and eventual commercialization providing greater access. Additionally, the Hunt laboratory spread the gospel of MS-based proteomics through collaborations that lasted decades with other scientists who were experts in immunology, cellular signaling, epigenetics, and other fascinating fields. This article attempts to highlight the many contributions of Don and the Hunt laboratory to peptide and protein identification since the year 2000.

Humans