Search PubMedSearch

PubMed · 42244099

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

Abstract

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhichao Xiao, Haibo Ji, Quan Zou, Yijie Ding, Liang Yu. 2026-06-01. Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.. https://doi.org/10.1093/bioinformatics%2Fbtag359

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Modular synthetic cross-kingdom promoters enable coordinated expression in Escherichia coli and Saccharomyces cerevisiae.

Synthetic biology and metabolic engineering increasingly demand predictable and interoperable gene expression across phylogenetically distant organisms, as the need for portable genetic systems and transferable metabolic pathways continues to grow. However, fundamental differences in promoter architecture and transcriptional logic across kingdoms remain a key bottleneck in developing universal expression platforms. Here, we designed a set of modular hybrid promoters that enable tunable and quantitatively consistent gene expression in both Escherichia coli and Saccharomyces cerevisiae. These promoters integrate bacterial -10/-35 motifs and Shine-Dalgarno sequences with minimal yeast TATA boxes and Kozak sequences to ensure transcriptional and translational compatibility. The promoter set supported weak, moderate, and strong expression with high relative consistency across species. Applied to the biosynthetic pathway for the valuable pigment prodeoxyviolacein, the hybrid promoters enabled coordinated production in both hosts. This work establishes a broadly compatible promoter architecture and provides a foundational toolkit for cross-kingdom, multi-host synthetic biology.

Promoter Regions, Genetic

Structural Features of DNA in TATA-Containing and TATA-Less Core Promoters of RNA Polymerase II Differ.

Nucleotide motifs in the core promoters of eukaryotic protein-coding genes transcribed by RNA polymerase II (Pol II) play an important role in the transcription process. We analyzed the role of an octanucleotide located in the TATA box position. Depending on whether this octanucleotide can form a complex with the TATA-binding protein (TBP), the promoter is classified as either TATA-containing or TATA-less. We analyzed the differences in the primary and spatial structures, as well as their dynamics, in TATA-containing and TATA-less promoters of mammals and plants. We divided the complete promoter sets of six organisms (H. sapiens, M. musculus, C. familiaris, A. thaliana, Z. mays, and H. vulgare) from the EPDnew database into TATA-containing and TATA-less fractions. The sizes of the TATA-containing promoter fractions are significantly smaller than those of the TATA-less fractions in all studied organisms, except in A. thaliana, where the sizes of both fractions are approximately equal. We characterized promoter architecture using variation profiles of various base-pair step parameters, minor-groove width, and the conformational dynamics of native DNA. The architectures of TATA-containing and TATA-less promoters differ significantly. The possible mechanistic influence of DNA structural features on the formation of the pre-initiation complex (PIC) in both types of promoters is discussed.

Promoter Regions, Genetic

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic