Search PubMedSearch

PubMed · 40581357

dbscATAC: a resource of single-cell super-enhancers/enhancers and gene markers derived from scATAC-seq data.

Abstract

MOTIVATION: scATAC-seq enables high-resolution mapping of cis-regulatory elements. It has been widely applied to uncover cell-type-specific regulatory networks and complement scRNA-seq analysis in numerous studies. However, a large number of datasets generated by scATAC-seq remain underutilized due to limited exploration of super-enhancers/typical enhancers and gene markers. A comprehensive resource enabling cell-type-specific annotation of cis-regulatory elements and their dynamic enhancer-gene linkages remains an urgent unmet need for scATAC-seq. RESULTS: We present dbscATAC, a specialized single-cell database for annotating super-enhancers, gene markers, and enhancer-gene interactions derived from scATAC-seq data. Using improved machine learning algorithms, we identified 213 835 super-enhancers across 520 tissue/cell types from three species, as well as 347 484 gene markers, 13 470 526 enhancers, and 10 402 346 enhancer-gene interactions derived from 1 668 076 single cells spanning 1028 tissue/cell types in 13 species. An easy-to-use online platform with multiple analytic modules and hierarchical query options was developed for searching, browsing and visualizing single-cell super-enhancers, enhancers, and gene markers. dbscATAC provides a comprehensive resource to facilitate the exploration of enhancer landscapes, gene regulation, and cell-type-specific characteristics in single-cell epigenomics. AVAILABILITY AND IMPLEMENTATION: The database with all the super-enhancer/enhancer annotation data is available at http://singlecelldb.com/dbscATAC/index.php. And the source code of dbscATAC for prediction of SEs, enhancers, and gene markers are available at https://github.com/EvansGao/dbscATAC. The source code, tissue/cell type description, and data summary can be downloaded at DOI: 10.6084/m9.figshare.28706414.scATAC-seq, Database, Super-enhancers/enhancers, Gene markers.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yingmei Li, Shahid Ullah, Yumei Xian, Yazhou Sun, Zilong Zheng, Xiaoyu Ma, Ming Shi, Changlin Zhang, Tian Li, Leli Zeng, Jie Chen, Yubin Y B Deng, Fuxin Wei, Tianshun Gao. 2025-07-01. dbscATAC: a resource of single-cell super-enhancers/enhancers and gene markers derived from scATAC-seq data.. https://doi.org/10.1093/bioinformatics%2Fbtaf364

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Mechanisms and functional implications of long-range enhancer-dependent gene regulation.

Metazoan development relies on the coordinated establishment of diverse gene regulatory programs that drive the formation of specific cell types, tissues and organs. The temporal and spatial control of gene expression is achieved through the concerted activity of multiple classes of cis-regulatory elements encoded in the genome. Among these, enhancers enable the establishment of specific and precise gene expression patterns and control gene expression over long linear distances, a property often referred to as distance-independent regulatory activity. However, enhancer activity is, in fact, inversely correlated with linear genomic distance, and target gene expression and transcriptional precision decrease with increasing enhancer-promoter linear distances. Here, we highlight emerging insights into multiple mechanisms that enable enhancers to precisely and robustly activate gene expression across large genomic distances. Finally, we provide a more speculative perspective on the potential advantages that long-range regulation might confer during the establishment of developmental gene expression programs.

Enhancer Elements, Genetic

Hi-Enhancer: a two-stage framework for prediction and localization of enhancers based on Blending-KAN and Stacking-Auto models.

MOTIVATION: Gene expression plays a crucial role in cell function, and enhancers can regulate gene expression precisely. Therefore, accurate prediction of enhancers is particularly critical. However, existing prediction methods have low accuracy or rely on fixed multiple epigenetic signals, which may not always be available. RESULTS: We propose a two-stage framework that accurately predicts enhancers by flexibly combining multiple epigenetic signals. In the first stage, we designed a Blending-KAN model, which integrates the results of various base classifiers and employs Kolmogorov-Arnold Networks (KAN) as a meta-classifier to predict enhancers based on flexible combinations of multiple epigenetic signals. In the second stage, we developed a Stacking-Auto model, which extracted sequence features using DNABERT-2 and located the enhancers based on the Stacking strategy and AutoGluon framework. The accuracy of the Blending-KAN model reached 99.69 ± 0.11% when five epigenetic signals were used. In cross-cell line prediction, the accuracy was more significant than or equal to 93.72%. With Gaussian noise, it still maintains an accuracy of 98.74 ± 0.03%. In the second stage, the accuracy of the Stacking-Auto model is 80.50%, which is better than the existing 17 methods. The results show that our models can be flexibly used to predict and locate enhancers utilizing a combination of multiple epigenetic signals. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/emanlee/Hi-Enhancer and https://doi.org/10.6084/m9.figshare.29262158.v1.

Enhancer Elements, Genetic

Adversarial attack of sequence-free enhancer prediction identifies chromatin architecture.

MOTIVATION: The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements "enhance" specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. RESULTS: We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks [adversarial particle swarm optimization (APSO)] to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference. AVAILABILITY AND IMPLEMENTATION: All software and code for data downloading, processing, enhancer inference, eXplainable AI (XAI), and complete figure generation are publicly available on GitHub at https://github.com/EpiGenomicsCode/ChromEnhancer and Zenodo at https://doi.org/10.5281/zenodo.15652797.

Enhancer Elements, Genetic