Search PubMedSearch

SEARCH · Search PubMed

Results for “sequencing artifacts”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Precision ID mtDNA Whole Genome Panel and sequencing of telogen hairs - perspectives for validation and implementation in casework.

Shed hair is a commonly encountered type of forensic evidence. Shed telogen hairs generally contain insufficient or highly degraded nuclear DNA for STR profiling; however, mtDNA analysis of telogen hair and hair shafts remains possible. We validated whole mitochondrial genome (mtGenome) sequencing using the Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific) and subsequently implemented the panel for the analysis of telogen hair, buccal, and casework samples. We analysed 90 diluted DNA samples containing 3-3,600 mtDNA copies, shed telogen hairs and their corresponding mtDNA from buccal swabs from 91 individuals, and 11 archived DNA extracts from hair samples in criminal cases. Complete mtGenome sequences were consistently recovered in 99% of samples across DNA dilution series at DNA input levels as low as 47 mtDNA copies, demonstrating the assay's robustness under low-template conditions. We obtained complete and reproducible mtGenome sequences with ≥ 327 mtDNA copies/µL from telogen hair samples. After applying ISFG recommendations and excluding low-confidence discrepancies associated with high-strand bias, heteroplasmic variants and sequencing artifacts, mtGenome sequence concordance increased from 93.4% to 100%. None of the 16 negative controls produced complete mtDNA sequences. Six negative controls showed low-level mtDNA signal (2-8 variants), consisting predominantly of common polymorphisms. These samples did not yield complete mtGenome sequences and showed no correspondence to any of the analysed samples. Finally, archived telogen hair samples from criminal cases presented complete mtGenome sequences with an average read depth of 1,037x.Our findings highlight the reliability of mtDNA analysis of telogen hairs using the Precision ID mtDNA Whole Genome Panel for implementation in forensic casework.

Forensic casework

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions.

Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions ≥50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs ≥4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..

Journal Article

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

AmpSeqR: an R package for amplicon deep sequencing data analysis.

Amplicon sequencing (AmpSeq) is a methodology that targets specific genomic regions of interest for polymerase chain reaction (PCR) amplification so that they can be sequenced to a high depth of coverage. Amplicons are typically chosen to be highly polymorphic, usually with several highly informative, high frequency single nucleotide polymorphisms (SNPs) segregating in an amplicon of 100-200 base pair (bp). This allows high sensitivity detection and quantification of the frequency of each sequence within each sample making it suitable for applications such as low frequency somatic mosaicism detection or minor clone detection in mixed samples. AmpSeq is being increasingly applied to both biological and medical studies, in applications such as cancer, infectious diseases and brain mosaicism studies. Current bioinformatics pipelines for AmpSeq data processing lack downstream analysis, have difficulty distinguishing between true sequences and PCR sequencing errors and artifacts, and often require bioinformatic expertise. We present a new R package: AmpSeqR, designed for the processing of deep short-read amplicon sequencing data, with a focus on infectious diseases. The pipeline integrates several existing R packages combining them with newly developed functions to perform optimal filtering of reads to remove noise and improve the accuracy of the detected sequences data, permitting detection of very low frequency clones in mixed samples. The package provides useful functions including data pre-processing, amplicon sequence variants (ASVs) estimation, data post-processing, data visualization, and automatically generates a comprehensive Rmarkdown report that contains all essential results facilitating easy inclusion into reports and publications. AmpSeqR is publicly available at https://github.com/bahlolab/AmpSeqR.

High-Throughput Nucleotide Sequencing

Whole Genome Methylation Sequencing via Enzymatic Conversion (EM-seq): Protocol, Data Processing, and Analysis.

Whole genome bisulfite sequencing (WGBS) has been the gold standard technique for base resolution analysis of DNA methylation for the last 15 years. It has been, however, associated with technical biases, which lead to overall overestimation of global and regional methylation values, and significant artifacts in extreme cytosine-rich DNA sequence contexts. Enzymatic conversion of cytosine is the newest approach, set to replace entirely the use of the damaging bisulfite conversion of DNA. The EM-seq technique utilizes TET2, T4-BGT, and APOBEC in a two-step conversion process, where the modified cytosines are first protected by oxidation and glucosylation, followed by deamination of all unmodified cytosines to uracil. As a result, EM-seq is degradation-free and bias-free, requires low DNA input, and produces high library yields with longer reads, little batch variation, less duplication, uniform genomic coverage, accurate methylation over a larger number of captured CpGs, and no sequence-specific artifacts.

DNA Methylation

Beyond blacklists: a critical assessment of exclusion set generation strategies and alternative approaches.

MOTIVATION: Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Alternatively, "sponge" or decoy sequences have been proposed to reduce alignment artifacts. RESULTS: We examined the widely used Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to sensitivity to input data, aligner choice, and read length. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. We additionally investigated the effect of the T2T-CHM13 genome assembly on improving biological signals. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets, and recommend the use of the T2T-CHM13 assembly or, for the hg38 genome assembly, "sponge" sequences as an alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Software

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches.

Short-read sequencing data can be affected by alignment artifacts in certain genomic regions. Removing reads overlapping these exclusion regions, previously known as Blacklists, help to potentially improve biological signal. Tools like the widely used Blacklist software facilitate this process, but their algorithmic details and parameter choices are not always clearly documented, affecting reproducibility and biological relevance. We examined the Blacklist software and found that pre-generated exclusion sets were difficult to reproduce due to variability in input data, aligner choice, and read length. We also identified and addressed a coding issue that led to over-annotation of high-signal regions. We further explored the use of "sponge" sequences-unassembled genomic regions such as satellite DNA, ribosomal DNA, and mitochondrial DNA-as an alternative approach. Aligning reads to a genome that includes sponge sequences reduced signal correlation in ChIP-seq data comparably to Blacklist-derived exclusion sets while preserving biological signal. Sponge-based alignment also had minimal impact on RNA-seq gene counts, suggesting broader applicability beyond chromatin profiling. These results highlight the limitations of fixed exclusion sets and suggest that sponge sequences offer a flexible, alignment-guided strategy for reducing artifacts and improving functional genomics analyses.

Journal Article

Phylogenomic Analyses Reveal that Panguiarchaeum Is a Clade of Genome-Reduced Asgard Archaea Within the Njordarchaeia.

The Asgard archaea are a diverse archaeal phylum important for our understanding of cellular evolution because they include the lineage that gave rise to eukaryotes. Recent phylogenomic work has focused on characterizing the diversity of Asgard archaea in an effort to identify the closest extant relatives of eukaryotes. However, resolving archaeal phylogeny is challenging, and the positions of 2 recently described lineages-Njordarchaeales and Panguiarchaeales-are uncertain, in ways that directly bear on hypotheses of early evolution. In initial phylogenetic analyses, these lineages branched either with Asgards or with the distantly related Korarchaeota, and it has been suggested that their genomes may be affected by metagenomic contamination. Resolving this debate is important because these clades include genome-reduced lineages that may help inform our understanding of the evolution of symbiosis within Asgard archaea. Here, we performed phylogenetic analyses revealing that the Njordarchaeales and Panguiarchaeales constitute the new class Njordarchaeia within Asgard archaea. We found no evidence of metagenomic contamination affecting phylogenetic analyses. Njordarchaeia exhibit hallmarks of adaptations to (hyper-)thermophilic lifestyles, including biased sequence compositions that can induce phylogenetic artifacts unless adequately modeled. Panguiarchaeum is metabolically distinct from its relatives, with reduced metabolic potential and various auxotrophies. Phylogenetic reconciliation recovers a complex common ancestor of Asgard archaea that encoded the Wood-Ljungdahl pathway. The subsequent loss of this pathway during the reductive evolution of Panguiarchaeum may have been associated with the switch to a symbiotic lifestyle, potentially based on H2-syntrophy. Thus, Panguiarchaeum may contain the first obligate symbionts within Asgard archaea besides the lineage leading to eukaryotes.

Phylogeny

Validation and Optimization of Breeding Strategy for miR-141/200c Knockout Mice to Eliminate Off-Target Gene Silencing using FLPo Deleter.

MicroRNAs (miRNAs) of the miR-200 family specifically miR-141 and miR-200c regulate neurogenesis, differentiation, and epithelial-mesenchymal transitions in development and several diseases including cancer and stroke. The STOCK Mirc13tm1Mtm /Mmjax mouse line, which targets the miR-141/200c cluster, was originally generated and described by Park et al. 2012 as a conditional "knockout-first" allele requiring a two-step breeding strategy: FLP recombination to excise lacZ/neo cassettes followed by Cre recombination to delete the floxed miRNA cluster (1). However, subsequent studies either bypassed this step and reported knockouts based on direct crosses with Cre mouse lines, leaving residual lacZ/neo sequences that may silence upstream elements or introduce transcriptional artifacts or rare studies used less efficient FLPe Deleter mice. Here, we present a detailed and refined strategy to conditional miR-141/200c knockouts mice using FLPo Deleter mice to efficiently eliminate lacZ/neo cassettes. Our approach not only confirmed complete deletion of miR-141 and miR-200c in various organs such olfactory bulbs and lungs where these miRNAs are robustly expressed using various approach such as genotyping qPCR validation and in situ hybridization but showed that without the use of FLPo deleter mice deletion of miR-141/200c cluster amy also lead to loss of several close proximity physiologically important genes such as ptpn6, phb2, atn1 and eno1. By restoring a clean floxed allele using FLPo deleter mice prior to Cre deletion, we establish a reliable and interpretable mouse model for dissecting the roles of the miR-141/200c cluster miRNA in various disease models.

Journal Article

Duplex-Indel: a Snakemake pipeline for somatic Indel calling in Tn5 transposase-based duplex sequencing data.

SUMMARY: Duplex-Indel is a novel Snakemake workflow for detecting somatic small insertions and deletions (Indels) from Tn5 transposase-based duplex sequencing data. Duplex-Indel enhances the accuracy of mutation calling at the single-molecule level by requiring consensus support from both DNA strands for each somatic Indel, minimizing confounding from technical artifacts. Duplex-Indel extends somatic mutation calling in Tn5 transposase-based duplex sequencing data to include Indels. We have demonstrated the accuracy and robustness of Duplex-Indel using cancer cell lines. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are available under the MIT license on GitHub at https://github.com/ealee-lab/duplex-indel and archived on Zenodo at https://doi.org/10.5281/zenodo.19228799.

Transposases

QCatch: a framework for quality control assessment and analysis of single-cell sequencing data.

MOTIVATION: Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstream results. While tools like alevin-fry and simpleaf (and augmented execution context for the alevin-fry), offer flexibility and computational efficiency to process single-cell data, this ecosystem will further benefit from a standardized QC reporting tailored for its outputs. RESULTS: We introduce QCatch, a Python-based command-line tool that generates comprehensive and interactive HTML QC reports designed specifically for single-cell quantification results. Taking the output directory of alevin-fry or simpleaf as the input, QCatch is able to perform essential processing steps, like cell calling, and generate detailed QC reports that contain informative visualizations and statistics, including unique molecular identifier (UMI) count distributions, sequencing saturation estimates, and splicing status information, for QC assurance. Built for seamless integration into downstream analysis workflows, QCatch exports the processed results in a richly-annotated H5AD format file, a widely used data format common among many downstream single-cell data analysis tools. AVAILABILITY AND IMPLEMENTATION: The source code and documentation of QCatch are available on GitHub at https://github.com/COMBINE-lab/QCatch. QCatch can be installed via both Bioconda and PyPI.

Single-Cell Analysis

Longitudinal characterization of mixed-genotype SARS-CoV-2 infections in a military cohort reveals compartmentalized viral populations.

UNLABELLED: Mixed-genotype severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections are a concern due to the potential generation of novel recombinants that give rise to new variants. To better understand intra-host viral dynamics, we analyzed specimens from 24 participants from the U.S. Military Health System's Epidemiology, Immunology, and Clinical Characteristics of Emerging Infectious Diseases with Pandemic Potential COVID-19 cohort with suspected mixed-genotype SARS-CoV-2 infections. From an initial 24 suspected cases, we confirmed 17 as genuine coinfections and graded them by evidence: 7 were "strong"; 4 were "moderate"; 6 were "weak"; and 7 were deemed unlikely to be true mixed-genotype infections. Access to swabs from multiple body sites across the course of infection allowed us to observe compartmentalization and shifts in variant dominance that would have been missed by a single-timepoint analysis, as well as one recombinant Omicron BA.1/BA.2 genome. By using an evidence-based bioinformatic framework to assess sequencing data from well-characterized clinical cases, we distinguished genuine coinfections from bioinformatic artifacts. Our findings emphasize the importance of both extensive specimen collection and careful bioinformatic approaches in ascertaining dual genotype infections. IMPORTANCE: Novel recombinants of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) arise from coinfections with different lineages, but mixed infections are not screened for despite risk to public health, and most surveillance relies on single swabs. We analyzed a longitudinal data set with specimens from multiple body sites, providing an opportunity to assess intra-host dynamics. To distinguish true coinfection from bioinformatic artifacts with confidence, we applied a framework that grades evidence for mixed genotypes by incorporating lineage and clade with manually validated variant calls. This allowed investigation beyond abundance levels of mixed genotypes within a single specimen, including observations of compartmentalization and a recombinant virus. This work enables further study of evolutionary, immunological, and clinical implications of mixed SARS-CoV-2 genotypes. Detecting dual-genotype infections and discriminating between true dual-genotype infection vs potential bioinformatics-based artifacts support public health and military readiness. These efforts provide evidence to bolster decision-making in molecular epidemiological studies to track transmission and for the choice of effective countermeasures.

SARS-CoV-2

The diagnostic potential of combined quantitative polymerase chain reaction and next-generation sequencing using the same primers for periprosthetic joint infection.

Next-generation sequencing (NGS) enables the detection of specific pathogens unidentifiable by conventional cultures, but its application in orthopedics remains inconsistent due to background contamination and irreproducible findings. This study evaluated the diagnostic performance of a novel workflow combining broad-range 16S rRNA gene quantitative PCR (qPCR) screening with downstream NGS, focusing on bacterial biomass thresholds. The qPCR assay demonstrated excellent intrarater reliability, with an intraclass correlation coefficient (ICC) of 0.961 (95% confidence interval, 0.881 to 0.997). Based on serially diluted positive controls, a quantitative threshold of 10⁵ CFU/mL was established as the minimum concentration required for the consistent detection of fastidious taxa, such as Escherichia coli. When evaluated against conventional cultures using 95 sonicate fluid and 276 pre/intraoperative tissue samples, the qPCR assay achieved a sensitivity of 80% and a specificity of 72%. Subsequent NGS sequencing of 26 clinical samples and 9 controls showed concordance in 4 of 6 culture-positive infected cases with NGS taxonomy, whereas the remaining discrepancies were likely attributable to culture-based phenotypic misidentification. Notably, among the qPCR-positive cases, three were culture-negative, including two hip prosthesis loosening cases exhibiting polymicrobial profiles, and one post-traumatic osteoarthritis case harboring low-level Staphylococcus. Crucially, this post-traumatic patient developed delayed periprosthetic joint infection (PJI) 2 years post-surgery, with cultures identifying Staphylococcus previously detected by the initial NGS analysis. Integrating qPCR screening with targeted NGS effectively refines pathogen identification, filters environmental artifacts, and overcomes the diagnostic limitations of culture-negative infections in orthopedic practice.IMPORTANCENext-generation sequencing (NGS) enables the detection of specific pathogens in clinical samples that are not identifiable by conventional methods. However, NGS applications in orthopedics have not been quantitatively evaluated, and findings have been inconsistent owing to contaminants and the presence of non-credible causative organisms. These factors primarily stem from the failure to evaluate low-biomass samples and the absence of proper controls, such as negative controls or mock community DNA samples. This study demonstrates that interpreting results from low-biomass samples requires careful consideration because NGS relies on relative bacterial abundances; distinguishing likely pathogens from contaminants is particularly challenging when bacterial loads are low. We demonstrated that combining NGS with quantitative PCR (qPCR) and applying a Cq cutoff can reduce false positives.

Humans

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

Improved spike-in normalization clarifies the relationship between active histone modifications and transcription.

Spike-in normalization enables quantitative analysis of chromatin immunoprecipitation sequencing (ChIP-seq) signal. Here we introduce a robust dual spike-in normalization approach for ChIP-seq (ChIP-wrangler), optimize parameters and verify its accuracy in quantifying changes in ChIP-seq signal and detecting technical artifacts. We use ChIP-wrangler to revisit recent claims that active histone marks depend on transcription. We show that acute depletion of RNA polymerase II (RNAPII) has a modest impact on H3K27ac levels, with only 6% of peaks significantly changing after RNAPII depletion, indicating that histone acetylation maintenance is not entirely dependent on ongoing transcription. Promoters and enhancers are differentially affected, with 82% of decreasing acetylation peaks located at promoter-distal elements with enhancer-related motifs. ChIP-wrangler provides increased rigor and 'guardrails' for successful spike-in normalization and, as applied here, refines the understanding of crosstalk between RNAPII activity and transcription-associated histone marks.

Histones