Search PubMedSearch

PubMed · 40973030

UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.

Abstract

MOTIVATION: One of the key applications of Unique Molecular Identifiers (UMIs) in high-throughput sequencing is to correct for PCR amplification bias and removal of PCR duplicates, thereby improving quantification in DNA-seq and RNA-seq applications. Accurately grouping error-bearing UMIs that originate from the same input molecule through a UMI deduplication method is a critical step in this process. However, many existing UMI deduplication tools rely on simple Hamming distance comparisons or suboptimal clustering algorithms, often resulting in erroneous UMI groupings, particularly in error-prone long-read sequencing or ultra-high-depth short-read sequencing. RESULTS: We introduce UMI-nea, a tool that utilizes Levenshtein distance comparisons and a novel clustering approach to optimize multithreading workflows. Compared against three other indel-aware UMI deduplication tools, UMI-nea achieves more accurate UMI groupings with efficient run time. It demonstrates robust performance across diverse sequencing platforms, depths, and UMI lengths. Additionally, UMI-nea incorporates a data-guided adaptive UMI filter, further enhancing quantification accuracy. AVAILABILITY AND IMPLEMENTATION: UMI-nea is available on github https://github.com/Qiaseq-research/UMI-nea.git or Zenodo https://doi.org/10.5281/zenodo.16745758. Sequencing data are stored at https://qiagenpublic.blob.core.windows.net/umi-nea-datasets/.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jixin Deng, Jingxiao Zhang, Song Tian, John DiCarlo, Hong Xu, Samuel J Rulli, Jonathan M Shaffer, Vikas Gupta, Toeresin Karakoyun. 2025-09-01. UMI-nea: a fast, robust tool for reference-free UMI deduplication and accurate quantification.. https://doi.org/10.1093/bioinformatics%2Fbtaf514

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing

RUMINA: high-throughput deduplication of unique molecular identifiers for amplicon and whole-genome sequencing with enhanced error correction.

MOTIVATION: Unique molecular identifiers (UMIs) are widely used in next-generation sequencing to enable accurate molecular counting and error correction. However, challenges remain in accurately collapsing UMI clusters, especially when read counts are low or sparse read clusters arise from barcode sequencing errors. RESULTS: We present RUMINA, a Rust-based pipeline for UMI-aware deduplication and error correction, optimized for both amplicon and shotgun sequencing. RUMINA supports multiple UMI cluster strategies, alongside majority-rule read selection independent of mapping quality, as well as discrete handling of 1-2 read clusters, paired-end merging, and read-length stratification. Benchmarking using simulated HIV population sequencing data and real-world iCLIP and TCR datasets showed that RUMINA improves ultra-low frequency SNV detection (0.01%-1%), reduces false positives, enhances reproducibility, and processes sequencing data up to 10-fold faster than existing tools. By integrating UMI- and sequence-level correction in a high-performance framework, RUMINA offers a fast, scalable, and robust solution for UMI-enabled sequencing workflows. AVAILABILITY AND IMPLEMENTATION: RUMINA is implemented in Rust and distributed as open-source code and precompiled binaries. Source code and installation instructions are available at https://github.com/greninger-lab/rumina. Documentation associated with this manuscript is available at https://github.com/greninger-lab/rumina_paper.

High-Throughput Nucleotide Sequencing

Enzymes in high-throughput RNA sequencing: Applications and challenges.

High-throughput RNA sequencing provides genome-wide information on the dynamics of RNA in each cell and how the dynamics responds to environmental changes. Next-generation sequencing by the Illumina platform currently provides the highest information output as compared to other platforms. A key component of next generation sequencing of each RNA is the successful end-to-end reverse-transcription into a cDNA strand. This can be highly challenging given the propensity of each RNA to adopt ordered structures and to contain post-transcriptional modifications. While many reverse transcriptase (RT) enzymes have been developed over the years to maximize read-through of an RNA, their processivity and efficiency varies, raising the question of how to select the RT for the experiment at hand. Here, we use tRNA as a model for genome-wide sequencing, as tRNA has a stable secondary and tertiary structure and has a high density and wide variety of post-transcriptional modifications, presenting one of the most challenging problems of sequencing RNA. We compare the efficiency of end-to-end cDNA synthesis of tRNA among several recent RT enzymes and provide a general sequencing workflow that is applicable to most of these enzymes.

High-Throughput Nucleotide Sequencing