Search PubMedSearch

PubMed · 42363597

DeepMASS v.2: An enhanced deep learning platform for large-scale discovery and structural annotation of unknown plant metabolites.

Abstract

Determining the structures of unknown metabolites remains a fundamental bottleneck in plant metabolomics, as the vast chemical diversity of plant secondary metabolites far exceeds the coverage of existing spectral libraries. Here, we present DeepMASS v.2, a substantially enhanced platform for annotating unknown metabolites from liquid chromatography-tandem mass spectrometry data, designed to address this challenge at scale. DeepMASS v.2 leverages a semantic spectral representation model trained on millions of spectra from GNPS, NIST, and in-house resources. By integrating Spec2Vec-based embeddings with HNSW (hierarchical navigable small world) graph retrieval and a unified chemical space defined by molecular fingerprints, DeepMASS v.2 identifies structurally related neighbors of unknown spectra and ranks candidate structures according to their proximity to the predicted structural neighborhoods within chemical space. Benchmarking against Critical Assessment of Small Molecule Identification datasets and a curated natural product collection demonstrated that DeepMASS v.2 outperforms state-of-the-art in silico annotation tools, including SIRIUS, CFM-ID, MetFrag, and MS-Finder. Importantly, DeepMASS v.2 maintains strong performance for metabolites absent from spectral libraries, highlighting its capacity to annotate genuinely unknown compounds. Application of DeepMASS v.2 to large-scale plant metabolomics datasets demonstrated its ability to expand accessible metabolome coverage. Implemented as an intuitive web platform, DeepMASS v.2 provides the community with a scalable, interpretable, and high-throughput solution for structural annotation, enabling more comprehensive characterization of plant chemical diversity and accelerating natural product discovery in molecular plant science. The DeepMASS v.2 web server is publicly available at http://deepmass.cn.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Siyu Jiang, Qiong Yang, Ziyao Xiong, Kairong Li, Qinliang Dai, Meifeng Su, Yaqing Lyu, Yanchun Peng, Ran Du, Jianbin Yan, Hongchao Ji. 2026-06-26. DeepMASS v.2: An enhanced deep learning platform for large-scale discovery and structural annotation of unknown plant metabolites.. https://doi.org/10.1016/j.xplc.2026.101976

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Column switching liquid chromatography dual mass spectrometry system for simultaneous untargeted metabolomics and targeted exposomics.

Exposome-wide association studies (ExWAS) require the detection of metabolites and exposures with diverse chemical properties across wide concentration ranges, a task that typically demands multiple analytical methods. To address this challenge, we develop an integrated column-switching two-dimensional liquid chromatography-dual mass spectrometry (2DLC-dual-MS) system. This system employs a 2DLC setup to sequentially separate polar and non-polar compounds with log P ranging from -8 to 15. The separated fractions are directed via a three-way valve to a high-resolution MS (HRMS) and a triple quadrupole MS (TQMS), enabling simultaneous untargeted metabolome analysis and targeted quantification of 601 exposures. The method is particularly suited for the concurrent analysis of metabolome and exposome in human blood, where their concentrations typically differ by 2-3 orders of magnitude. In a demonstration application on lung adenocarcinoma ExWAS, the system exhibits good stability over more than 300 consecutive injections for both metabolome and exposome analysis, confirming its robustness for ExWAS applications.

Metabolomics

Computational metabolomics at scale: from open data to insight.

Metabolomics data are currently generated at scale thanks to the evolution of technologies that have led to marked improvements in the number of metabolites detected, spanning all chemical classes. These data are increasingly submitted to public repositories for data reuse, integration, and interpretation. Despite the availability of public resources and associated computational tools, the field still lacks a widely adopted, consistent data and analytics infrastructure capable of transforming this wealth of information into scientific insight. Indeed, the metabolomics field is just now scratching the surface of being able to harness the power of new computational technologies. In this review, we summarize discussions from the "Dagstuhl-Seminar 24181 Computational Metabolomics: Towards Molecules, Models, and their Meaning" with a focus on public data availability, open data standards, data and knowledge integration, and education. Our goal is to raise awareness and adoption of the latest open science resources while highlighting key areas needing further development.

Metabolomics

hypeR-GEM: connecting metabolite signatures to enzyme-coding genes via genome-scale metabolic models.

MOTIVATION: Enrichment analysis is a cornerstone of "omics" data interpretation, enabling researchers to connect analysis results to biological processes and generate testable hypotheses. Enrichment analysis in metabolomics poses distinct challenges for interpretation and multi-omics integration due to the lack of well-defined and consistent connections to well-curated gene-centered biological knowledge repositories. To address these challenges, we developed hypeR-GEM, a methodology and associated R package that adapts gene set enrichment analysis to metabolomics. hypeR-GEM leverages genome-scale metabolic models (GEMs) to infer reaction-based links between metabolites and enzyme-coding genes, enabling the mapping of metabolite signatures to gene signatures and their subsequent annotation via gene set enrichment analysis. RESULTS: We validated hypeR-GEM using paired metabolomics-proteomics and metabolomics-transcriptomics datasets by assessing whether genes mapped from metabolites significantly overlapped with differentially expressed proteins or transcripts. We further evaluated whether pathways enriched via hypeR-GEM-mapped genes corresponded to those derived from paired proteomic or transcriptomic data. In most datasets analyzed, both the predicted enzyme-coding genes and the associated enriched pathways showed significant concordance with independently derived omics signatures, supporting the utility and robustness of hypeR-GEM. Finally, we applied hypeR-GEM to the analysis of age-associated metabolic signatures from the New England Centenarian Study. The results revealed consistent enrichment of lipid-related pathways, aligning with the well-established role of lipid metabolism in aging, and highlighted additional pathways not captured in the metabolites' annotation, demonstrating hypeR-GEM's practical utility in a real-world use case. AVAILABILITY AND IMPLEMENTATION: The hypeR-GEM R package, documentation, and workflow examples are freely available at https://github.com/montilab/hypeR-GEM and archived at https://doi.org/10.5281/zenodo.20586748.

Metabolomics