Search PubMedSearch

SEARCH · Search PubMed

Results for “Single-cell ATAC-seq”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

SUMMARY: Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors." These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac. We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC. Alevin-fry-atac is highly scalable and, when using 32 threads, is 2.8 times faster than Chromap (the second fastest approach) while using only 33% of the memory required by Chromap. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual color-enhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger. AVAILABILITY AND IMPLEMENTATION: Alevin-fry-atac is written in Rust and C++17, and is freely-available under a BSD 3-clause license. It is integrated into piscem (https://github.com/COMBINE-lab/piscem) and alevin-fry (https://github.com/COMBINE-lab/alevin-fry), and is also supported directly as part of simpleaf (https://github.com/COMBINE-lab/simpleaf).

Single-Cell Analysis

Alevin-fry-atac enables rapid and memory frugal mapping of single-cell ATAC-seq data using virtual colors for accurate genomic pseudoalignment.

Ultrafast mapping of short reads via lightweight mapping techniques such as pseudoalignment has significantly accelerated transcriptomic and metagenomic analyses, often with minimal accuracy loss compared to alignment-based methods. However, applying pseudoalignment to large genomic references, like chromosomes, is challenging due to their size and repetitive sequences. We introduce a new and modified pseudoalignment scheme that partitions each reference into "virtual colors…. These are essentially overlapping bins of fixed maximal extent on the reference sequences that are treated as distinct "colors" from the perspective of the pseudoalignment algorithm. We apply this modified pseudoalignment procedure to process and map single-cell ATAC-seq data in our new tool alevin-fry-atac . We compare alevin-fry-atac to both Chromap and Cell Ranger ATAC . Alevin-fry-atac is highly scalable and, when using 32 threads, is approximately 2.8 times faster than Chromap (the second fastest approach) while using approximately one third of the memory and mapping slightly more reads. The resulting peaks and clusters generated from alevin-fry-atac show high concordance with those obtained from both Chromap and the Cell Ranger ATAC pipeline, demonstrating that virtual colorenhanced pseudoalignment directly to the genome provides a fast, memory-frugal, and accurate alternative to existing approaches for single-cell ATAC-seq processing. The development of alevin-fry-atac brings single-cell ATAC-seq processing into a unified ecosystem with single-cell RNA-seq processing (via alevin-fry ) to work toward providing a truly open alternative to many of the varied capabilities of CellRanger . Furthermore, our modified pseudoalignment approach should be easily applicable and extendable to other genome-centric mapping-based tasks and modalities such as standard DNA-seq, DNase-seq, Chip-seq and Hi-C.

Journal Article

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

MACS3: A Peak-calling Platform for Bulk and Single-cell Regulatory Genomics.

Since the original publication of Model-based Analysis for ChIP-Seq (MACS), the software has been widely used to identify enriched genomic regions in ChIP-seq, ATAC-seq, CUT&RUN, DNase-seq, and related regulatory genomics assays. Over the years, MACS has evolved substantially, with MACS version 3 (MACS3) now serving as the actively maintained implementation. MACS3 preserves the core MACS framework for fragment pileup, dynamic local background noise, statistical enrichment testing, and peak refinement, while adding functionality needed for contemporary bulk and single-cell workflows. It supports conventional bulk peak calling, paired-end and fragment-based file formats, modular signal processing, direct analysis of single-cell ATAC-seq fragment files, barcode-restricted pseudobulk and cluster-level peak calling, specialized ATAC-seq and variant-calling modules, as well as command-line and programmatic interfaces. MACS3 is distributed through standard software channels and supported by continuous testing across operating systems, Python versions, and CPU architectures. Here we describe the architecture, current capabilities, and recommended use of MACS3, providing an updated reference for applying the MACS framework in contemporary bulk and single-cell regulatory genomics workflows. MACS3 is open-source software available at https://github.com/macs3-project/MACS.

Bioinformatics software

HBO1 functions as an epigenetic barrier to hepatocyte plasticity and reprogramming during liver injury.

Hepatocytes can reprogram into biliary epithelial cells (BECs) during liver injury, but the underlying epigenetic mechanisms remain poorly understood. Here, we define the chromatin dynamics of this process using single-cell ATAC-seq and identify YAP/TEAD activation as a key driver of chromatin remodeling. An in vivo CRISPR screen highlights the histone acetyltransferase HBO1 as a critical barrier to reprogramming. HBO1 is recruited by YAP to target loci, where it promotes histone H3 lysine 14 acetylation (H3K14ac) and engages the chromatin reader zinc-finger MYND-type containing 8 (ZMYND8) to suppress YAP/TEAD-driven transcription. Loss of HBO1 accelerates chromatin remodeling, enhances YAP binding, and enables a more complete hepatocyte-to-BEC transition. Our findings position HBO1 as an epigenetic brake that restrains YAP-mediated reprogramming, suggesting that targeting HBO1 may enhance hepatocyte plasticity for liver regeneration.

Hepatocytes

A latent activated olfactory stem cell state revealed by single-cell transcriptomic and epigenomic profiling.

The olfactory epithelium is one of the few regions of the nervous system that sustains neurogenesis throughout life. Its experimental accessibility makes it especially tractable for studying molecular mechanisms that drive neural regeneration in response to injury. In this study, we used single-cell sequencing to identify transcriptional and epigenetic processes involved in determining olfactory epithelial stem cell fate during injury-induced regeneration. By combining gene expression and accessible chromatin profiles of individual lineage-traced olfactory stem cells, we identified transcriptional heterogeneity among activated stem cells at a stage when cell fates are being specified. We further identified a subset of resting cells that appears poised for activation, characterized by accessible chromatin around silent genes prior to their expression in response to injury. These results provide evidence for a latent activated stem cell state in which a subset of quiescent olfactory epithelial stem cells are epigenetically primed to support injury-induced regeneration.

Animals

scMGCL: accurate and efficient integration representation of single-cell multi-omics data.

MOTIVATION: Single-cell multi-omics data integration is essential for understanding cellular states and disease mechanisms, yet integrating heterogeneous data modalities remains a challenge. We present scMGCL, a graph contrastive learning framework for robust integration of single-cell ATAC-seq and RNA-seq data. Our approach leverages self-supervised learning on cell-cell similarity graphs, in which each modality's graph structure serves as an augmentation for the other. This cross-modality contrastive paradigm enables the learning of biologically meaningful, shared representations while preserving modality-specific features. RESULTS: Benchmarking against state-of-the-art methods demonstrates that scMGCL outperforms others in cell-type clustering, label transfer accuracy, and preservation of marker-gene correlations. Additionally, scMGCL significantly improves computational efficiency, reducing runtime and memory usage. The method's effectiveness is further validated through extensive analyses of cell-type similarity and functional consistency, providing a powerful tool for multi-omics data exploration. AVAILABILITY AND IMPLEMENTATION: Code and datasets are released at https://github.com/zlCreator/scMGCL.

Single-Cell Analysis

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

OmnibusX: A unified platform for accessible multi-omics analysis.

OmnibusX is an integrated, privacy-centric platform that enables code-free multi-omics data analysis by bridging computational methodologies with user-friendly interfaces. Designed to overcome challenges posed by fragmented analytical tools and high computational barriers, OmnibusX consolidates workflows for diverse technologies - including bulk RNA-seq, single-cell RNA-seq, single-cell ATAC-seq, and spatial transcriptomics - into a single, cohesive application. The application integrates established open-source tools such as Scanpy, DESeq2, SciPy, and scikit-learn into transparent, reproducible pipelines, offering users control over analytical parameters. Additionally, OmnibusX features proprietary modules, including a highly accurate cell-type prediction engine and an interactive plotting editor for generating publication-quality visualizations. Available as a standalone desktop application and an enterprise edition for centralized server deployment, OmnibusX ensures all data processing is conducted locally, eliminating external data transfer and usage tracking. By lowering technical barriers and enhancing reproducibility, OmnibusX aims to accelerate biological discovery and foster robust, data-driven collaborations. A fully documented trial version is accessible at: https://omnibusx.com/apps.

Computational Biology

Single-cell mapping of regulatory DNA-protein interactions.

Gene expression is controlled by transcription factors (TFs), whose genome binding is shaped by chromatin accessibility and histone modifications, yet mapping these interactions, particularly those with weak affinity or a transient nature, in single cells remains technically challenging. To address this gap, we developed docking and deamination followed by sequencing (D&D-seq), a single-cell immuno-tethering technology for profiling DNA-protein interactions. D&D-seq couples an antibody-binding nanobody to a cytosine base editor, a combination that enables detection of weak or transient factor binding through targeted cytosine-to-uracil editing at protein-bound genomic sites. This approach is compatible with standard single-cell multi-omic workflows and therefore allows integrated analyses of gene regulation. Using assay for transposase-accessible chromatin using sequencing (ATAC-seq) and single-cell ATAC-seq (scATAC-seq), we assessed chromatin accessibility as a functional readout of TF activity, and by coupling D&D-seq with whole-genome sequencing, we captured CTCF binding in both active and inactive chromatin compartments.

Animals

Teratoma Formation and Genomic Profiling Using Multi-Omics Approaches.

Teratoma formation is the gold standard assay for evaluating the developmental pluripotency of human and mouse embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs). Following subcutaneous injection into immunodeficient mice, pluripotent stem cells spontaneously differentiate into derivatives representing all three embryonic germ layers-ectoderm, mesoderm, and endoderm. Beyond serving as a functional assay for pluripotency, teratomas provide a unique three-dimensional model system for studying early human development and lineage specification in vivo. This chapter describes comprehensive protocols for teratoma formation in immunodeficient mice, tissue processing for multiple downstream genomic applications, and multi-omics profiling approaches. We detail methods for embryonic stem cell culture, teratoma generation via subcutaneous injection, tissue dissection and processing for chromatin immunoprecipitation followed by sequencing (ChIP-Seq), RNA sequencing (RNA-Seq), single-cell multiome profiling combining chromatin accessibility (ATAC-Seq) and gene expression (scRNA-Seq), and histological analysis using hematoxylin and eosin (H&E) staining. Additionally, we provide bioinformatics workflows for analyzing the resulting genomic datasets to characterize the epigenetic and transcriptional landscapes of teratoma-derived tissues. These methods enable comprehensive molecular characterization of developmental processes and provide valuable resources for stem cell biologists studying pluripotency, differentiation, and early embryonic development.

Teratoma

Prdm15 deficiency perturbs hematopoietic stem and progenitor cell homeostasis.

The maintenance of homeostasis in hematopoietic stem and progenitor cells (HSPCs) is essential for the proper development of the entire hematopoietic system. However, the mechanisms underlying this regulatory equilibrium remain elusive. Here, we report that Prdm15 deficiency in HSPCs induces the accumulation of immature hematopoietic stem cells in mice. A series of transplantation assays shows that these cells display impaired reconstitution capacity and competitive fitness, which are associated with abnormal differentiation trajectories and transcriptional alterations identified by single-cell RNA sequencing. Mechanistically, integrated multi-omics analyses including ATAC-seq and CUT&Tag sequencing of HSPCs indicate that Prdm15 deficiency induces significant transcriptional and epigenetic alterations, particularly affecting the methyltransferase KMT2C and altering H3K4me1 and H3K27ac modifications at the promoters of hematopoietic developmental genes. Collectively, our findings establish PRDM15 as a critical epigenetic regulator of HSPCs, offering valuable insights into the molecular mechanisms underlying hematopoietic homeostasis.

Cell differentiation

PATTY corrects open-chromatin bias for improved bulk and single-cell CUT&Tag profiling.

Precise profiling of epigenomes is essential for better understanding chromatin biology and gene regulation. Cleavage Under Targets & Tagmentation (CUT&Tag) is an efficient epigenomic profiling technique that can be performed on a low number of cells and at the single-cell level. With its growing adoption, CUT&Tag datasets spanning diverse biological systems are rapidly accumulating in the field. CUT&Tag assays use the hyperactive transposase Tn5 for DNA tagmentation. Tn5's preference toward accessible chromatin alters CUT&Tag sequence read distributions in the genome and introduces open-chromatin bias that can confound downstream analysis, an issue more substantial in sparse single-cell data. We show that open-chromatin bias extensively exists in published CUT&Tag datasets, including those generated with recently optimized high-salt protocols. To address this challenge, we present PATTY (Propensity Analyzer for Tn5 Transposase Yielded bias), a comprehensive computational method that corrects open-chromatin bias in CUT&Tag data by leveraging accompanying ATAC-seq. By integrating transcriptomic and epigenomic data using machine learning and integrative modeling, we demonstrate that PATTY enables accurate and robust detection of occupancy sites for both active and repressive histone modifications, including H3K27ac, H3K27me3, and H3K9me3, with experimental validation. We further develop a single-cell CUT&Tag analysis framework built on PATTY and show improved cell clustering when using bias-corrected single-cell CUT&Tag data compared to using uncorrected data. Beyond CUT&Tag, PATTY sets a foundation for further development of bias correction methods for improving data analysis for all Tn5-based high-throughput assays.

Journal Article

A stem-like chromatin program in small-cell lung cancer is associated with poor outcomes after chemoimmunotherapy.

Small-cell lung cancer (SCLC) is an aggressive malignancy with substantial tumor heterogeneity and limited clinically actionable biomarkers beyond established features such as liver metastases. We profile tumor-intrinsic chromatin accessibility in a patient-derived xenograft biobank and identify three recurrent chromatin programs: neuroendocrine, marked by ASCL1/NEUROD1 activity; immunogenic, marked by IRF-associated activity; and stem-like, marked by TEAD/OCT activity. These programs are reproduced at the cohort level across bulk and single-cell transcriptomic datasets comprising more than 800 tumors, including 300 extensive-stage samples. In patients treated with chemoimmunotherapy, the stem-like program is associated with inferior survival, including a median overall survival of 7.41 months versus 15.9 and 12.6 months for immunogenic and neuroendocrine groups, respectively. This association remains significant after adjustment for liver metastases, brain metastases, and elevated lactate dehydrogenase. These findings support a high-risk stem-like SCLC chromatin program for prospective biomarker refinement and therapeutic investigation.

ATAC-seq

A Standardized Protocol for Generating iPSC-Derived Human Microglia for Functional Genomic Assays.

Human induced pluripotent stem cell (iPSC)-derived microglia (iMG) provide an in vitro experimental system for studying human microglial biology, neuroinflammation, and genetic risk mechanisms associated with neurological disease. This chapter describes a standardized, scalable, and reproducible protocol for the differentiation of human iPSCs into functional microglia-like cells, with particular emphasis on applications in transcriptional and epigenomic network analysis. The protocol supports high-viability floating iMG production, compatibility with pooled CRISPR perturbation approaches, and downstream multiomic profiling, including single-cell RNA sequencing, chromatin accessibility assays, and proteomics. Detailed procedures are provided for iPSC maintenance, hematopoietic progenitor cell generation, microglial maturation, functional genomics integration, and quality control.

Humans

A distinct effector B cell population drives autoantibody production in SARS-CoV-2 infection.

Autoantibodies (autoAbs) are linked to mortality and Long COVID, yet their cellular origins remain unclear. We analyzed the INCOV cohort and identified 12 age- and sex-matched participants with varying autoAb abundance and integrated single-cell RNA-seq and ATAC-seq data from B cells, plasma proteomics, proteome-wide autoAb profiling, clinical data, and in vitro assays. AutoAb abundance inversely correlated with neutralizing IgG and declined as infection resolved, paralleling the contraction of atypical memory B cells (AtMs). In vitro, AtMs preferentially differentiated into autoAb-producing antibody-secreting cells upon TLR7/8 stimulation. CD11c+ AtMs (double-negative 2, DN2s) in autoAb-high individuals exhibited increased TLR7 signaling, oxidative stress, and isotype switching, regulated by transcription factors T-bet and XBP1. Integrated genetic and genomic analyses showed that DN2s had the strongest enrichment for autoimmune trait heritability and inferred regulatory effects of autoimmune risk variants among B cell subsets. These findings identify DN2s as key precursors of autoAb-producing cells during SARS-CoV-2 infection.

B cell

Disentangling covariate effects on single-cell-resolved epigenomes with DeepDive.

Understanding the effects of individual biological factors from single-cell-resolved epigenomic data is hindered by multicollinearity, particularly in human cohorts. We introduce DeepDive, a deep-learning framework designed to systematically disentangle known and unknown sources of variation in single-nucleus ATAC-seq data. DeepDive accurately reconstructs chromatin accessibility, outperforms state-of-the-art methods with incomplete covariate information, and robustly recovers true biological signals from even highly entangled covariates, unlocking counterfactual, "what-if," analyses. Applying DeepDive to pancreatic islet cells, we perform counterfactual analyses to prioritize covariates associated with a type 2 diabetes-linked beta-cell subtype and nominate transcription regulators. DeepDive offers a powerful and unbiased tool for mechanistic discovery in complex human disease cohorts.

disentanglement

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160 myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals