Search PubMedSearch

PubMed · 42323537

MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.

Abstract

BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yizhuo Shen, Jianpeng Sheng, Pengcheng Yan. 2026-06-20. MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.. https://doi.org/10.1186/s12859-026-06522-z

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction.

SUMMARY: Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health. AVAILABILITY: Data is available at https://gdo-meta2db.llnl.gov/ and https://zenodo.org/records/17315984.

Metadata

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata

HTSinfer: inferring metadata from bulk Illumina RNA-Seq libraries.

SUMMARY: The Sequencing Read Archive is one of the largest and fastest-growing repositories of sequencing data, containing tens of petabytes of sequenced reads. Its data is used by a wide scientific community, often beyond the primary study that generated them. Such analyses rely on accurate metadata concerning the type of experiment and library, as well as the organism from which the sequenced reads were derived. These metadata are typically entered manually by contributors in an error-prone process, and are frequently incomplete. In addition, easy-to-use computational tools that verify the consistency and completeness of metadata describing the libraries to facilitate data reuse, are largely unavailable. Here, we introduce HTSinfer, a Python-based tool to infer metadata directly and solely from bulk RNA-sequencing data generated on Illumina platforms. HTSinfer leverages genome sequence information and diagnostic genes to rapidly and accurately infer the library source and library type, as well as the relative read orientation, 3' adapter sequence and read length statistics. HTSinfer is written in a modular manner, published under a permissible free and open-source license and encourages contributions by the community, enabling easy addition of new functionalities, e.g. for the inference of additional metrics, or the support of different experiment types or sequencing platforms. AVAILABILITY AND IMPLEMENTATION: HTSinfer is released under the Apache License 2.0. Latest code is available via GitHub at https://github.com/zavolanlab/htsinfer, while releases are published on Bioconda. A snapshot of the HTSinfer version described in this article was deposited at Zenodo at 10.5281/zenodo.13985958.

Metadata