Search PubMedSearch

SEARCH · Search PubMed

Results for “CTCF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

36 records · Page 2Linked to original sources

Chrom-Sig: de-noising 1D genomic profiles by signal processing methods.

MOTIVATION: Modern genomic research is driven by next-generation sequencing experiments such as ChIP-seq, CUT&Tag, and CUT&RUN that generate coverage files for transcription factor binding, as well as ATAC-seq that yield coverage files for chromatin accessibility. Due to the inherent technical noise present in the experimental protocols, researchers need statistically rigorous and computationally efficient methods to extract true biological signal from a mixture of signal and noise. However, existing approaches are often computationally demanding or require input or spike-in controls. RESULTS: We developed Chrom-Sig, a Python package to quickly de-noise 1D genomic coverage tracks by computing the empirical null distribution without prior assumptions or experimental controls. When tested on 19 ChIP-seq, CUT&RUN, ATAC-seq, and snATAC-seq datasets, Chrom-Sig can effectively decompose the data into signal and noise components. Notably, Chrom-Sig performs de-noising and peak calling in 1-2 h using around 20 GB of memory. The de-noised signal corroborates with biologically meaningful results: CTCF CUT&RUN data retained a high percentage of peaks overlapping CTCF binding motifs, while ATAC-seq and RNA Polymerase II data were enriched in enhancers and promoters. We envision Chrom-Sig to be a versatile and general tool for current and future genomic technologies. AVAILABILITY AND IMPLEMENTATION: Chrom-Sig is publicly available on GitHub (https://github.com/minjikimlab/chromsig) and Zenodo (doi: 10.5281/zenodo.17488772) under the MIT licence.

Genomics

ERCC2 mutations alter the genomic distribution pattern of somatic mutations and are independently prognostic in bladder cancer.

Excision repair cross-complementation group 2 (ERCC2) encodes the DNA helicase xeroderma pigmentosum group D, which functions in transcription and nucleotide excision repair. Point mutations in ERCC2 are putative drivers in around 10% of bladder cancers (BLCAs) and a potential positive biomarker for cisplatin therapy response. Nevertheless, the prognostic significance directly attributed to ERCC2 mutations and its pathogenic role in genome instability remain poorly understood. We first demonstrated that mutant ERCC2 is an independent predictor of prognosis in BLCA. We then examined its impact on the somatic mutational landscape using a cohort of ERCC2 wild-type (n = 343) and mutant (n = 39) BLCA whole genomes. The genome-wide distribution of somatic mutations is significantly altered in ERCC2 mutants, including T[C>T]N enrichment, altered replication time correlations, and CTCF-cohesin binding site mutation hotspots. We leverage these alterations to develop a machine learning model for predicting pathogenic ERCC2 mutations, which may be useful to inform treatment of patients with BLCA.

Humans

Barrier effects on the kinetics of cohesin-mediated loop extrusion.

Chromosome organization mediated by structural maintenance of chromosome complexes is crucial in many organisms. Cohesin extrudes chromatin into loops that are thought to lengthen until it is obstructed by CTCF proteins. In complex cellular environments, the loop extrusion machinery may encounter other chromatin-binding proteins. How these proteins interfere with the cohesin-meditated extrusion process is largely unexplored, but recent experiments have shown that some proteins serve as physical barriers that block cohesin translocation. Other proteins containing a cohesin-interaction motif serve as chemical barriers to induce cohesin pausing through interactions with it. Here, we develop an analytically solvable approach for the loop extrusion model incorporating barriers to investigate the effect of the barrier on the passive extrusion process. To further quantify the impact of barriers, we calculate the mean looping time it takes for cohesin to translocate to form a stable loop before dissociation. Our finding reveals that the physical barrier can accelerate the loop formation, and the degree of acceleration is closely related to the impedance strength of the physical barrier. In particular, the synergy of the cohesin loading site and the physical barrier site accelerates loop formation more significantly. The proximity of the cohesin loading site to the barrier site facilitates the rapid formation of stable loops in long genomes, which implies loop extrusion and chromatin-binding proteins might shape functional genomic organization. Conversely, chemical barriers consistently impede loop formation, with increasing impedance strength of the chemical barrier leading to longer loop formation time. Our study contributes to a more comprehensive understanding of the complexity of the loop extrusion process, providing a new perspective on the potential mechanisms of gene regulation.

Cohesins

Single-cell mapping of regulatory DNA-protein interactions.

Gene expression is controlled by transcription factors (TFs), whose genome binding is shaped by chromatin accessibility and histone modifications, yet mapping these interactions, particularly those with weak affinity or a transient nature, in single cells remains technically challenging. To address this gap, we developed docking and deamination followed by sequencing (D&D-seq), a single-cell immuno-tethering technology for profiling DNA-protein interactions. D&D-seq couples an antibody-binding nanobody to a cytosine base editor, a combination that enables detection of weak or transient factor binding through targeted cytosine-to-uracil editing at protein-bound genomic sites. This approach is compatible with standard single-cell multi-omic workflows and therefore allows integrated analyses of gene regulation. Using assay for transposase-accessible chromatin using sequencing (ATAC-seq) and single-cell ATAC-seq (scATAC-seq), we assessed chromatin accessibility as a functional readout of TF activity, and by coupling D&D-seq with whole-genome sequencing, we captured CTCF binding in both active and inactive chromatin compartments.

Animals

Selective control of HLA-DRB1 and HLA-DQA1 transcription through DR/DQ super enhancer microelements.

MHC-II gene expression requires promoter-proximal and distal organizing elements. A super enhancer located between the HLA-DR and -DQ genes was dissected into five microdomains (regions A-E) to define its mechanism of action. Regions A and B were required for maximal expression of HLA-DQA1. Region A facilitated HLA-DQA1 expression and substituted for region B in its absence, as well as mediating longer-range interactions with distal CTCF sites. Region D modulated HLA-DRB1 through specific 3D chromatin interactions, whereas a local insulator element (XL9) broadly interacted with the SE. Through deletion of HLA-DRB1 proximal-promoter elements, HLA-DQ expression increased by co-opting interactions with region D, indicating that competition between microelements regulates the absolute levels of MHC-II genes. Together, these data define an additional layer of control for MHC-II genes encoded in one of the most polymorphic and disease-associated regions of the human genome.

CP: immunology

Genome-wide profiling of histone modifications and transcription factor binding at single-cell resolution by DeChIC-seq.

Mapping of protein-DNA interactions at single-cell resolution remains a central challenge in epigenomics, particularly for transcription factors (TFs), whose sparse binding limits reliable detection. Here, we establish DeChIC-seq (DNA Deaminase-based Chromatin Immuno-Conversion sequencing), a conversion-based strategy that uses a protein A-DddAtox fusion to directly record protein-DNA interactions by inducing localized C-to-U conversions near antibody-bound chromatin. Retaining genome-wide background sequence information without immunoprecipitation, DeChIC-seq enables profiling of histone modifications and sensitive detection of TF binding. Integration with single-cell whole-genome amplification extends DeChIC-seq to single-cell applications (scDeChIC-seq), enabling chromatin profiling of individual cells. Applied to mouse embryogenesis, scDeChIC-seq resolves lineage-specific chromatin states through profiling of H3K4me3, CTCF, and RAD21 and sensitively detects TF binding, including that of NR5A2, TFAP2C, and KLF5, from extremely limited blastomere inputs. This underscores its strong potential for detecting TF-binding sites in scarce biological samples. DeChIC-seq establishes a conversion-based framework for chromatin profiling that enables mechanistic dissection of TF-driven gene regulation across rare cells, developmental systems, and disease contexts.

Animals

Allele-specific chromatin architecture shapes imprinted domains and coordinates a distal enhancer and antisense transcription at the mouse Mest-Copg2 domain.

Genomic imprinting results in parent-of-origin-dependent gene expression, but how three-dimensional genome organization contributes to imprinted gene regulation remains unclear. Using Capture Hi-C in mouse cortex and primary cortical neurons, we identified parental allele-specific chromatin architectures across multiple imprinted domains. These architectures largely originate from imprinting control regions and correlate with DNA methylation-sensitive CTCF binding. Active and inactive alleles of imprinted genes show distinct promoter interaction profiles and differential engagement with distal regulatory elements in both contact frequency and the epigenetic state of distal regions. A CRISPR interference screen identified a distal enhancer that regulates Mest-Copg2 imprinted expression through allele-specific chromatin interactions. In neurons, this enhancer activates Copg2 on the maternal allele, whereas on the paternal allele it drives Mest isoforms transcribed antisense to Copg2 and contributes to Copg2 repression. In summary, we show that allele-specific chromatin architecture coordinates maternal enhancer activity and paternal antisense transcription to control imprinted expression in neurons.

Animals

STAG2 loss in Ewing sarcoma alters enhancer-promoter contacts dependent and independent of EWS::FLI1.

Cohesin complexes carrying STAG1 or STAG2 organize the genome into chromatin loops. STAG2 loss-of-function mutations promote metastasis in Ewing sarcoma, a pediatric cancer driven by the fusion transcription factor EWS::FLI1. We integrated transcriptomic data from patients and cellular models to identify a STAG2-dependent gene signature associated with worse prognosis. Subsequent genomic profiling and high-resolution chromatin interaction data from Capture Hi-C indicated that cohesin-STAG2 facilitates communication between EWS::FLI1-bound long GGAA repeats, presumably acting as neoenhancers, and their target promoters. Changes in CTCF-dependent chromatin contacts involving signature genes, unrelated to EWS::FLI1 binding, were also identified. STAG1 is unable to compensate for STAG2 loss and chromatin-bound cohesin is severely decreased, while levels of the processivity factor NIPBL remain unchanged, likely affecting DNA looping dynamics. These results illuminate how STAG2 loss modifies the chromatin interactome of Ewing sarcoma cells and provide a list of potential biomarkers and therapeutic targets.

Sarcoma, Ewing

Regulation of immune signal integration and memory by inflammation-induced chromosome conformation.

Three-dimensional (3D) genome conformation is central to gene expression regulation, yet our understanding of its contribution to rapid transcriptional responses, signal integration, and memory in immune cells is limited. Here, we study the molecular regulation of the inflammatory response in primary macrophages using integrated transcriptomic, epigenomic, and chromosome conformation data, including base pair-resolution Micro Capture-C. We demonstrate that interleukin-4 (IL-4) primes the inflammatory response in macrophages by stably rewiring 3D genome conformation, juxtaposing endotoxin-, interferon-gamma-, and dexamethasone-responsive enhancers to their cognate gene promoters. CRISPR-based perturbations of enhancer-promoter contacts or CCCTC-binding factor (CTCF) boundary elements show that IL-4-driven conformation changes are required for enhanced and synergistic endotoxin-induced transcriptional responses, as well as transcriptional memory following stimulus removal. Moreover, transcriptional memory mediated by changes in chromosome conformation can occur in the absence of changes in chromatin accessibility or histone modifications. Collectively, these findings demonstrate that rapid and memory transcriptional responses to immunological stimuli are encoded in the 3D genome.

Animals

dcHiChIP: a comprehensive Nextflow-based pipeline for multiscale analysis of chromatin architecture from HiChIP data.

MOTIVATION: Despite the growing use of HiChIP to investigate protein-directed chromatin architecture, a comprehensive and reproducible pipeline for analysing these datasets-from raw reads to multiscale 3D genome features-remains lacking. Existing tools often focus on isolated components, such as loop calling or matrix generation, but fall short in integrating structural annotation, functional enrichment, and spatial modeling within a unified framework. To address this gap, we developed dcHiChIP, a modular, scalable Nextflow-based workflow that streamlines the analysis of HiChIP data, enabling both routine processing and in-depth exploration of chromatin organization and regulatory interactions. RESULTS: dcHiChIP enables robust and reproducible analysis of HiChIP datasets across multiple scales of chromatin architecture. It accepts raw sequencing data as input and generates high-quality loop calls, domain annotations, and 3D genome models. It also performs functional annotation and motif enrichment analyses. Applied to benchmark CTCF HiChIP datasets, dcHiChIP identifies major chromatin architectural features such as TADs/CCDs, A/B compartments, and chromatin stripes, and offers efficient, end-to-end execution with support for batch processing and workflow resumability. AVAILABILITY: dcHiChIP is publicly available on GitHub at https://github.com/SFGLab/dcHiChIP, with documentation at https://sfglab.github.io/dcHiChIP/. The software version used in this study is archived at Zenodo: https://doi.org/10.5281/zenodo.22030542.

Chromatin

Extrusion fountains are hallmarks of chromosome organization emerging upon zygotic genome activation.

The initiation of gene expression during development, known as zygotic genome activation (ZGA), is accompanied by massive changes in chromosome organization. However, the earliest events of chromosome folding and their functional roles remain unclear. Using Hi-C on zebrafish embryos, we discovered that chromosome folding begins early in development with the formation of "fountains", a novel element of chromosome organization. Emerging preferentially at enhancers, fountains exhibit an initial accumulation of cohesin, which later redistributes to CTCF sites at TAD borders. Knockouts of pioneer transcription factors driving ZGA enhancers result in the specific loss of fountains, establishing a causal link between enhancer activation and fountain formation. Polymer simulations demonstrate that fountains may arise as sites of facilitated cohesin loading, requiring two-sided but desynchronized loop extrusion, potentially caused by cohesin collisions with obstacles or internal switching. Moreover, we detected similar fountain patterns at enhancers in mouse cells. Fountains disappear upon acute cohesin depletion, as well as during mitosis, and reappear with cohesin loading in early G1. Altogether, fountains represent the first known enhancer-specific elements of chromosome organization and constitute starting points for chromosome folding during development, likely through facilitated cohesin loading.

Journal Article

Architectural logic of the 3D genome: mechanisms of dysregulation and emerging cancer therapeutics.

The three-dimensional (3D) genome provides an essential layer of organization that shapes genome function in space and time. Chromatin compartments and topologically associating domains (TADs) arise from the interplay between intrinsic properties of chromatin and architectural factors, including cohesin and CTCF. Despite substantial progress in defining these structural features, whether 3D genome architecture plays a causal role in regulating processes such as transcription, DNA replication, and DNA repair, or instead reflects underlying regulatory activity, remains unresolved. Here, we use the distinction between chromatin-intrinsic features and architectural factors as a framework to evaluate evidence for causality in genome structure-function relationships. We extend this framework to cancer, where both intrinsic alterations (including noncoding mutations, structural variants, and changes in chromatin state) and architectural factor perturbations (such as mutations in architectural proteins and dysregulation of transcriptional machinery) disrupt genome organization and contribute to disease progression. These findings suggest that alterations in genome structure can, in some contexts, actively reshape oncogenic programs. A major limitation in applying 3D genome insights to cancer biology is the cost and complexity of omics assays. Recent advances in artificial intelligence (AI) and machine learning (ML) enable inference and prediction of 3D genome organization from sequence and epigenomic features, providing insight into the extent to which genome folding is encoded intrinsically versus dynamically regulated in architectural factors. This perspective provides a unified view of how genome structure is established, how it relates to function, and how its disruption contributes to tumorigenesis.

3D genome

Architectural transcription factors collectively shape nuclear radial positioning of chromatin contacts.

The measurement of three-dimensional genome folding in the nucleus, mostly through Hi-C methods, is expressed as contact frequencies between genomic segments, without anchoring to physical axes of the spherical nucleus. Here, we mapped the chromatin contacts along nuclear radial axis and built radial score by factoring in contact frequencies. The chromatin high-order structures exhibit rich diversity along radial axis. Furthermore, the proximal trans contacts retrieved by radial score reveal conserved active/inactive chromatin segregation across intra- and interchromosomal interactions. Ablation of CTCF proteins disrupts chromatin loops with mild changes to chromatin radial positioning. By acutely perturbing multiple transcription factor (TF) occupancy, chromatin loop dissolutions are often accompanied by radial dissociations between two anchors. Our work provides a genome architecture reference map adhering to nuclear physical axis and suggests that multiple architectural TFs collectively shape nuclear positioning of chromatin and their contacts, with contacts serving as forces on chromatin positioning as well.

Chromatin

Transcriptomics-based exploration of ubiquitination-related biomarkers and potential molecular mechanisms in laryngeal squamous cell carcinoma.

BACKGROUND: One of the most common and prevalent cancers is laryngeal squamous cell carcinoma (LSCC), which poses a great threat to the life and health of the patient. Nonetheless, it has been demonstrated that ubiquitination is crucial for the development and course of LSCC. Therefore, it is particularly important to identify biomarkers for ubiquitination-related genes (UbRGs) in LSCC. METHODS: Differentially expressed genes (DEGs) in the LSCC versus controls were obtained by differential expression analysis. Also, key modular genes associated with LSCC were obtained using weighted gene co-expression network analysis (WGCNA). Next, DEGs, key module genes, and UbRGs were taken to intersect to obtain candidate genes. And then machine algorithms were to screen potential biomarkers, further their diagnostic value were analyzed and validated. Then, therapeutic agents for biomarkers were predict. In addition, the regulatory networks of the biomarkers were mapped. The expression levels of biomarkers were detected in clinical samples using reverse transcription-quantitative PCR (RT-qPCR). RESULTS: A total of eight candidate genes were acquired by the overlap 1,911 DEGs, the key modular genes of WGCNA, and 1,393 UbRGs. A sum of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) were identified by two machine learning, then these four biomarkers were validated in GSE127165 and the expression trend was consistent with TCGA-LSCC, they were recorded as biomarkers. Moreover, the accuracy of the biomarkers in predicting clinical aspects of LSCC was confirmed by the receiver operating characteristic (ROC) curves. Subsequently, cancers such as malignant neoplasms, colorectal cancers, tumors, and primary malignant neoplasms were significantly associated with the biomarkers, which further suggests that these four biomarkers were strongly associated with cancer. Meanwhile, the drugs garcinol, cocaine, and triazolam, among others, used for LSCC treatment were predicted. Finally, transcription factors (TFs) (BRD4, MYC, AR, and CTCF) were predicted to regulate the biomarkers. RT-qPCR assays illustrated that the expression trends of KAT2B, LNX1 and NBEAL2 remained consistent with the dataset. CONCLUSION: The identification of four biomarkers (WDR54, KAT2B, NBEAL2 and LNX1) associated with UbRGs could ultimately serve as a predictive clinical diagnosis of LSCC and provide insight into the molecular mechanisms of LSCC.

Humans

UnionLoops: a workflow for calling chromatin loops across related Hi-C datasets with improved specificity, precision, and sensitivity.

Chromatin loop calling from chromatin interaction data often exhibits substantial variability across related samples. We present UnionLoops, a computational workflow for chromatin loop calling across multiple related samples. UnionLoops integrates information across datasets to determine positions and dataset-specificity of looping interactions. It constructs a unified candidate loop set, applies consistent filtering and aggregation, and evaluates loop support across samples. We demonstrate that UnionLoops increases sensitivity for detecting shared chromatin loops, reduces spurious sample-specific calls, and improves concordance with independent genomic features, including CTCF and cohesin occupancy. UnionLoops enables improved biological interpretation of chromatin loop organization and dynamics across related conditions.

Chromatin

PRMT5 regulates alternative splicing of TCF3 under hypoxia to promote EMT and invasion in breast cancer.

Tumor hypoxia induced alterations in the epigenetic landscape and alternative splicing influence cellular adaptations. PRMT5 is a type II protein arginine methyltransferase that regulates several tumorigenic events in many cancer types. However, the regulation of PRMT5 and its direct implication on aberrant alternative splicing under hypoxia remains unexplored. In this study, we observed hypoxia-induced upregulation of PRMT5 via the CTCF in human breast cancer cells. Further, PRMT5-mediated symmetric arginine dimethylation H4R3me2s and H3R8me2s directly regulated the alternative splicing of TCF3. Under hypoxia, PRMT5-mediated histone dimethylation at the intronic conserved region (ICR) present between TCF3 exon 18a and exon 18b recruits DNMT3A, resulting in DNA methylation. DNA methylation at the TCF3-ICR is recognized and bound by MeCP2 resulting in RNA-Pol II pausing, promoting the recruitment of the negative splicing factor PTBP1 to the splicing locus of TCF3 pre-mRNA. PTBP1 promotes the exclusion of exon 18a which results in the production of the pro-invasive TCF3-18B (E47) isoform which promotes EMT and invasion of breast cancer cells under hypoxia. Collectively, our results indicate PRMT5-mediated symmetric arginine dimethylation of histones regulates alternative splicing of TCF3 gene thereby enhancing EMT and invasion in breast cancer hypoxia.

Humans

Machine learning on multiple epigenetic features reveals H3K27Ac as a driver of gene expression prediction across patients with glioblastoma.

Epigenetic mechanisms play a crucial role in driving transcript expression and shaping the phenotypic plasticity of glioblastoma stem cells (GSCs), contributing to tumor heterogeneity and therapeutic resistance. These mechanisms dynamically regulate the expression of key oncogenic and stemness-associated genes, enabling GSCs to adapt to environmental cues and evade targeted therapies. Importantly, epigenetic reprogramming allows GSCs to transition between cellular states, including therapy-resistant mesenchymal-like phenotypes, underscoring the need for epigenetic-targeting strategies to disrupt these adaptive processes. Understanding these epigenetic drivers of gene expression provides a foundation for novel therapeutic interventions aimed at eradicating GSCs and improving glioblastoma outcomes. Using machine learning (ML), we employ cross-patient prediction of transcript expression in GSCs by combining epigenetic features from various sources, including ATAC-seq, CTCF ChIP-seq, RNAPII ChIP-seq, H3K27Ac ChIP-seq, and RNA-seq. We investigate different ML and deep learning (DL) models for this task and ultimately build our final pipeline using XGBoost. The model trained on one patient generalizes to other 11 patients with high performance. Notably, H3K27Ac alone from a single patient is sufficient to predict gene expression in all 11 patients. Furthermore, the distribution of H3K27Ac peaks across the genomes of all patients is remarkably similar. These findings suggest that GSCs share a common distributional pattern of enhancer activity characterized by H3K27Ac, which can be utilized to predict gene expression in GSCs across patients. In summary, while GSCs are known for their transcriptomic and phenotypic heterogeneity, we propose that they share a common epigenetic pattern of enhancer activation that defines their underlying transcriptomic expression pattern. This pattern can predict gene expression across patient samples, providing valuable insights into the biology of GSCs.

Glioblastoma

Adult bi-paternal offspring generated through direct modification of imprinted genes in mammals.

Imprinting abnormalities pose a significant challenge in applications involving embryonic stem cells, induced pluripotent stem cells, and animal cloning, with no universal correction method owing to their complexity and stochastic nature. In this study, we targeted these defects at their source-embryos from same-sex parents-aiming to establish a stable, maintainable imprinting pattern de novo in mammalian cells. Using bi-paternal mouse embryos, which exhibit severe imprinting defects and are typically non-viable, we introduced frameshift mutations, gene deletions, and regulatory edits at 20 key imprinted loci, ultimately achieving the development of fully adult animals, albeit with a relatively low survival rate. The findings provide strong evidence that imprinting abnormalities are a primary barrier to unisexual reproduction in mammals. Moreover, this approach can significantly improve developmental outcomes for embryonic stem cells and cloned animals, opening promising avenues for advancements in regenerative medicine.

Animals