Search PubMedSearch

SEARCH · Search PubMed

Results for “Binding Sites”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Identification of RBP binding sites using RNA deaminases.

RNA-binding proteins (RBPs) are critical regulators of gene expression and RNA processing. Identification of their binding sites has important implications for their physiological and disease-related functions. Crosslinking and immunoprecipitation, followed by sequencing (CLIP-seq) and its derivatives, are the most commonly used methods to identify RBP binding sites, but are laborious and require a large amount of starting material. Recent advancements harnessing RNA deaminases in fusion to any RBP of interest, allow for the profiling of RBP binding sites from low-input samples in simpler procedures. Among these efforts, we developed STAMP (Surveying Targets by APOBEC-Mediated Profiling), which efficiently detects RBP-RNA interactions. This chapter describes the detailed protocol for the STAMP method, including plasmid construction, delivery and sorting, library preparation and bioinformatic data analysis.

RNA-Binding Proteins

IFNL1 gene promoter single nucleotide polymorphism rs7247086 enhances transcription through a STAT-binding site.

A single nucleotide polymorphism (SNP) within the human interferon lambda 1 (IFN-L1, IFN-λ1) gene promoter, rs7247086 (C/T) has been reported to be associated with severe dengue and possibly with psoriasis and COVID-19. However, its functional nature is unknown. The present study was undertaken to examine the effect of rs7247086 on transcription, by utilizing promoter and enhancer-reporter assays. We see that the T allele completes a consensus signal transducer and activator of transcription (STAT)-binding site. While we did not find strong evidence to show that the STAT-binding site drove transcription from the IFNL1 gene promoter, we saw that it acts like an enhancer in reporter assays. The T allele of rs7247086 within the STAT-binding site significantly increased transcription of the reporter gene compared to the C allele when incorporated into enhancer-reporter constructs in both HEK293 and A549 cell lines. Mechanistically, we obtained evidence from electrophoretic mobility shift assays to show that the T allele binds to STAT proteins more strongly than the C allele. In a cohort of healthy individuals, we saw that the T allele carriers, specifically males but not females, had significantly increased secretion of IFN-λ1 from their peripheral blood mononuclear cells after stimulation. Lastly, rs7247086 significantly associated with psoriasis, only in males but not in females.

Humans

The proneural proteins Atonal and Scute regulate neural target genes through different E-box binding sites.

For a particular functional family of basic helix-loop-helix (bHLH) transcription factors, there is ample evidence that different factors regulate different target genes but little idea of how these different target genes are distinguished. We investigated the contribution of DNA binding site differences to the specificities of two functionally related proneural bHLH transcription factors required for the genesis of Drosophila sense organ precursors (Atonal and Scute). We show that the proneural target gene, Bearded, is regulated by both Scute and Atonal via distinct E-box consensus binding sites. By comparing with other Ato-dependent enhancer sequences, we define an Ato-specific binding consensus that differs from the previously defined Scute-specific E-box consensus, thereby defining distinct E(Ato) and E(Sc) sites. These E-box variants are crucial for function. First, tandem repeats of 20-bp sequences containing E(Ato) and E(Sc) sites are sufficient to confer Atonal- and Scute-specific expression patterns, respectively, on a reporter gene in vivo. Second, interchanging E(Ato) and E(Sc) sites within enhancers almost abolishes enhancer activity. While the latter finding shows that enhancer context is also important in defining how proneural proteins interact with these sites, it is clear that differential utilization of DNA binding sites underlies proneural protein specificity.

3' Flanking Region

TFinder: A Python Web Tool for Predicting Transcription Factor Binding Sites.

Transcription is a key cell process that consists of synthesizing several copies of RNA from a gene DNA sequence. This process is highly regulated and closely linked to the ability of transcription factors to bind specifically to DNA. TFinder is an easy-to-use Python web portal allowing the identification of Individual Motifs (IM) such as Transcription Factor Binding Sites (TFBS). Using the NCBI API, TFinder extracts either promoter or gene terminal regulatory regions, through a simple query of NCBI gene name or ID. It enables simultaneous analysis across five different species for an unlimited number of genes. TFinder searches for Individual Motifs in different formats, including IUPAC codes and JASPAR entries. Moreover, TFinder also allows de novo generations of a Position Weight Matrix (PWM) and the use of already established PWM. Finally, the data are provided in a tabular and a graph format showing the relevance and the P-value of the Individual Motifs found as well as their location relative to the Transcription Start Site (TSS) or the terminal region of the gene. The results are then sent by email to users facilitating the subsequent data analysis and sharing. TFinder is written in Python and freely available on GitHub under the MIT license: https://github.com/Jumitti/TFinder. It can be accessed as a web application implemented in Streamlit at https://tfinder-ipmc.streamlit.app. Resources are available on Streamlit "Resources" tab. TFINDER strength is that it relies on an all-in-one intuitive tool allowing users inexperienced with bioinformatics tools to retrieve gene regulatory regions sequences in multiple species and to search for individual motifs in a huge number of genes.

Transcription Factors

How negative sampling shapes the performance of transcription factor binding site prediction models.

MOTIVATION: Transcription factors (TFs) are key players in gene regulation and development, where they activate and repress gene expression through DNA binding. Predicting transcription factor binding sites (TFBSs) has long been an active area of research, with many deep learning methods developed to tackle this problem. These models are often trained on TF ChIP-seq data, which is generally seen as only providing positive samples. The choice of datasets and negative sampling techniques is a critical yet often overlooked aspect of this work. RESULTS: In this study, we investigate the impact of different negative sampling techniques on TFBS prediction performance. We create high-quality test datasets based on ChIP-seq and ATAC-seq data, where true negatives can be identified as positions that are accessible but not bound by the TF in question. We then train models using various negative sampling techniques, including genomic sampling, shuffling, dinucleotide shuffling, neighborhood sampling, and cell line specific sampling, simulating cases where matching ATAC-seq data is not available. Our results show that, generally, metrics calculated on training datasets give inflated performance scores. Of the tested techniques, genomic sampling of negatives based on similarity to the positives performed by far the best, although still not reaching the performance of baseline models trained on high-quality datasets. Models trained on dinucleotide shuffled negatives performed poorly, despite being a common practice in the field. Our findings highlight the importance of carefully selecting negative sampling techniques for TFBS prediction, as they can significantly impact model performance and the interpretation of results. AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/NatanTourne/TFBS-negatives (DOI: 10.5281/zenodo.18007567).

Binding Sites

Role of the CTCF binding site in Human T-Cell Leukemia Virus-1 pathogenesis.

During HTLV-1 infection, the virus integrates into the host cell genome as a provirus with a single CCCTC binding protein (CTCF) binding site (vCTCF-BS), which acts as an insulator between transcriptionally active and inactive regions. Previous studies have shown that the vCTCF-BS is important for maintenance of chromatin structure, regulation of viral expression, and DNA and histone methylation. Here, we show that the vCTCF-BS also regulates viral infection and pathogenesis in vivo in a humanized (Hu) mouse model of adult T-cell leukemia/lymphoma. Three cell lines were used to initiate infection of the Hu-mice, i) HTLV-1-WT which carries an intact HTLV-1 provirus genome, ii) HTLV-1-CTCF, which contains a provirus with a mutated vCTCF-BS which abolishes CTCF binding, and a stop codon immediately upstream of the mutated vCTCF-BS which deletes the last 23 amino acids of the p12 gene, and iii) HTLV-1-p12stop that contains the intact vCTCF-BS, but retains the same stop codon in p12 as in the HTLV-1-CTCF cell line. Hu-mice were infected with mitomycin-treated or irradiated HTLV-1 producing cell lines. There was a delay in pathogenicity when Hu-mice were infected with the HTLV-1-CTCF virus compared to mice infected with either HTLV-1-p12 stop or HTLV-1-WT virus. Proviral load (PVL), spleen weights, and CD4 T cell counts were significantly lower in HTLV-1-CTCF infected mice compared to HTLV-1-p12stop infected mice. Furthermore, we found a direct correlation between the PVL in peripheral blood and death of HTLV-1-CTCF infected mice. In cell lines, we found that the vCTCF-BS regulates Tax expression in a time-dependent manner. The scRNAseq analysis of splenocytes from infected mice suggests that the vCTCF-BS plays an important role in activation and expansion of T lymphocytes in vivo. Overall, these findings indicate that the vCTCF-BS regulates Tax expression, proviral load, and HTLV pathogenicity in vivo.

Human T-lymphotropic virus 1

Unplugging lateral fenestrations of NALCN reveals a hidden drug binding site within the pore region.

The sodium (Na+) leak channel (NALCN) is a member of the four-domain voltage-gated cation channel family that includes the prototypical voltage-gated sodium and calcium channels (NaVs and CaVs, respectively). Unlike NaVs and CaVs, which have four lateral fenestrations that serve as routes for lipophilic compounds to enter the central cavity to modulate channel function, NALCN has bulky residues (W311, L588, M1145, and Y1436) that block these openings. Structural data suggest that occluded fenestrations underlie the pharmacological resistance of NALCN, but functional evidence is lacking. To test this hypothesis, we unplugged the fenestrations of NALCN by substituting the four aforementioned residues with alanine (AAAA) and compared the effects of NaV, CaV, and NALCN blockers on both wild-type (WT) and AAAA channels. Most compounds behaved in a similar manner on both channels, but phenytoin and 2-aminoethoxydiphenyl borate (2-APB) elicited additional, distinct responses on AAAA channels. Further experiments using single alanine mutants revealed that phenytoin and 2-APB enter the inner cavity through distinct fenestrations, implying structural specificity to their modes of access. Using a combination of computational and functional approaches, we identified amino acid residues critical for 2-APB activity, supporting the existence of drug binding site(s) within the pore region. Intrigued by the activity of 2-APB and its analogues, we tested compounds containing the diphenylmethane/amine moiety on WT channels. We identified clinically used drugs that exhibited diverse activity, thus expanding the pharmacological toolbox for NALCN. While the low potencies of active compounds reiterate the pharmacological resistance of NALCN, our findings lay the foundation for rational drug design to develop NALCN modulators with refined properties.

Binding Sites

Natural variants of CsSHN1 orchestrate a temporal regulatory cascade driving fruit skin netting in cucumber.

Fruit skin netting (russeting, Rs) forms when epidermal microcracks are sealed by a suberized periderm, reducing marketability. We previously identified the Rs locus (CsSHN1), which encodes an AP2/ERF transcription factor, as a major determinant of cucumber skin netting, but how fruit growth is temporally coupled to periderm formation remains unclear. Here, we integrated population genomics, time-series multiomics, DNA affinity purification sequencing (DAP-seq), and transgenic assays to decode the CsSHN1-mediated regulatory network. Six functionally relevant CsSHN1 variants were identified across 325 cucumber accessions. Allele distribution and selective sweep analyses revealed breeding-driven selection for smooth fruit skin. Overexpression of a netted allele in a smooth background induced epidermal fissures, altered cell geometry, and increased fruit size, demonstrating a dosage-sensitive effect. Time-series transcriptomics and metabolomics of near-isogenic lines (NILs) defined 3 developmental phases of netting: early suppression of lignin and trehalose genes preceding cracks, growth-driven fissuring accompanied by cell-wall remodeling and defense activation, and maturation-stage cell-wall degradation with strong induction of ligno-suberin biosynthesis. Across the cucumber genome, DAP-seq identified approximately 8,000 in vitro CsSHN1 binding sites. These binding sites were significantly enriched for the GCC-box motif and included genes involved in cutin and suberin biosynthesis. Together, these results show that CsSHN1 orchestrates fruit skin netting through a growth-coupled temporal regulatory cascade, providing a mechanistic framework for manipulating fruit epidermal properties.

Cucumis sativus

Deciphering estrogen receptor alpha-driven transcription in human endometrial stromal cells via transcriptome, cistrome, and integration with chromatin landscape.

OBJECTIVE: To investigate estrogen receptor gene 1 (ESR1) and estrogen-driven transcription in human endometrial stromal cells. DESIGN: RNA sequencing (RNA-seq) and Cleavage Under Targets and Release Using Nuclease (Cut&Run) were performed on telomerase-immortalized human endometrial stromal cells with Clustered Regularly Interspaced Short Palindromic Repeats-mediated ESR1 activation. Hi-C-based chromatin architecture analysis (H3K27ac HiChIP) was conducted in primary endometrial stromal cells. SUBJECTS: Biopsies from two healthy, reproductive-aged volunteers with regular menstrual cycles and no history of gynecological malignancies. EXPOSURE: The ESR1-activated and control endometrial stromal cells were treated with estradiol (E2) or vehicle. Primary endometrial stromal cells were treated with vehicle or a decidualization cocktail. MAIN OUTCOME MEASURES: Differential gene expression analysis (RNA-seq) identified ligand-independent and -dependent ESR1 activity. Cut&Run profiled ESR1 genomic binding in ESR1-activated cells. H3K27ac HiChIP mapped hormone-induced changes in chromatin looping in primary cells. RESULTS: Among seven tested guide RNAs (gRNA), the ESR1-3 gRNA induced robust ESR1 activation and restored E2 responsiveness. Bulk RNA-seq revealed both ligand-dependent and -independent ESR1 transcriptional programs regulating inflammation, proliferation, and cancer-related pathways. Notably, 72% of differentially expressed genes overlapped with genes active in human endometrial tissue during the proliferative estrogen-dominant phase, supporting their physiological relevance. The Cut&Run-seq identified genome-wide ESR1 binding sites, with most binding sites located at distal regulatory elements. Integration of Cut&Run data with H3K27ac HiChIP chromatin loops linked distal ESR1 binding sites to gene promoters, including genes involved in decidualization (e.g., FOXO1) and endometrial cancer (e.g., ERRFI1, NRIP1, and EPAS1). Functional assays showed that ESR1 promotes cell viability and, in the presence of E2, enhances migration. CONCLUSION: The CRISPR-mediated ESR1 activation restores estrogen responsiveness in endometrial stromal cells. Combined transcriptomic, cistromic, and chromatin architecture analyses reveal ESR1's role in regulating decidualization and inflammation-related gene networks, with relevance to endometrial pathologies including endometrial cancer. This model serves as a powerful tool to study estrogen signaling in endometrial stromal cell biology and related pathologies.

Humans

Lineage-specific splicing regulation of MAPT gene in the primate brain.

Divergence of precursor messenger RNA (pre-mRNA) alternative splicing (AS) is widespread in mammals, including primates, but the underlying mechanisms and functional impact are poorly understood. Here, we modeled cassette exon inclusion in primate brains as a quantitative trait and identified 1,170 (∼3%) exons with lineage-specific splicing shifts under stabilizing selection. Among them, microtubule-associated protein tau (MAPT) exons 2 and 10 underwent anticorrelated, two-step evolutionary shifts in the catarrhine and hominoid lineages, leading to their present inclusion levels in humans. The developmental-stage-specific divergence of exon 10 splicing, whose dysregulation can cause frontotemporal lobar degeneration (FTLD), is mediated by divergent distal intronic MBNL-binding sites. Competitive binding of these sites by CRISPR-dCas13d/gRNAs effectively reduces exon 10 inclusion, potentially providing a therapeutically compatible approach to modulate tau isoform expression. Our data suggest adaptation of MAPT function and, more generally, a role for AS in the evolutionary expansion of the primate brain.

tau Proteins

Enhanced transgene expression from single-stranded AAV vectors in human cells in vitro and in murine hepatocytes in vivo.

We identified that distal 10 nucleotides in the D-sequence in AAV2 inverted terminal repeat (ITR) share partial sequence homology to 1/2 binding site of glucocorticoid receptor-binding element (GRE). Here, we describe that (1) purified GR binds to AAV2 D-sequence, and the D-sequence competes with GR binding to its cognate binding site; (2) dexamethasone-mediated activation of GR pathway significantly increases the transduction efficiency of AAV2 vectors in human cells; (3) human osteosarcoma cells, U2OS, which lack expression of GR, are poorly transduced by AAV2 vectors, but stable transfection with a GR expression plasmid restores vector-mediated transgene expression; (4) replacement of the distal 10 nucleotides in the D-sequence of the AAV2 ITR with a full-length GRE consensus sequence significantly enhances transgene expression in human cells in vitro and in murine hepatocytes in vivo; and (5) none of the ITRs in AAV1, AAV3, AAV4, AAV5, and AAV6 genomes contains the GRE 1/2 binding site, and insertion of a full-length GRE consensus sequence in the AAV6-ITR also significantly enhances transgene expression from AAV6 vectors, both in vitro and in vivo. These novel vectors, termed generation Y AAV vectors, which are serotype, transgene, or promoter agnostic, should be useful in human gene therapy.

AAV vectors

Optimization of Structure-Guided Development of Chemical Probes for the Pseudoknot RNA of the Frameshift Element in SARS-CoV-2.

Targeting the RNA genome of SARS-CoV-2 is a viable option for antiviral drug development. We explored three ligand binding sites of the core pseudoknot RNA of the SARS-CoV-2 frameshift element. We iteratively optimized ligands, based on improved affinities, targeting these binding sites and report on structural and dynamic properties of the three identified binding sites. Available experimental 3D structures of the pseudoknot element were compared to SAXS and NMR data to validate its dominant folding state in solution. In order to experimentally map in silico predicted binding sites, NMR assignments of the majority of nucleobases were achieved by segmental labeling of the pseudoknot RNA and isotope-filtered NMR experiments at 1.2 GHz, demonstrating the value of NMR spectroscopy to supplement modelling and docking data. Optimized ligands with enhanced affinity were shown to specifically inhibit frameshifting without affecting 0-frame translation in cell-free translation assays, establishing the frameshift element as target for drug-like ligands of low molecular weight.

SARS-CoV-2

Post-transcriptional regulation of Profilin-2 by microRNAs and RNA-binding proteins forms a critical regulatory node for early embryonic cell fate decisions.

Post-transcriptional control by RNA binding proteins (RBPs) and microRNAs play central roles in mRNA stability and translation, yet how RBPs and microRNAs coordinate in developmental time to regulate cell fate remains poorly understood. Here, we demonstrate that post-transcriptional regulation of the Profilin 2 (Pfn2) transcript is essential for differentiation of embryonic stem cells (ESCs) into the primary germ layer lineages. The Pfn2 3'untranslated region has both an Iron Regulatory Protein binding site (IRE) and a nearby binding site for ESC enriched microRNAs. Deletion of this microRNA site leads to increased PFN2 and reduced FGF signaling during pluripotency transition prior to germ layer formation. In contrast, deletion of the IRE leads to decreased PFN2, a Wnt signaling defect, reduced nuclear beta-catenin, and a subsequent block in mesendodermal lineages during early germ layer formation. We further find that loss of the IRE site results in a cell autonomous defect in Wnt signaling and mesendodermal differentiation. The IRE site acts to stabilize beta-catenin, as disruption of the site leads to reduced nuclear beta-catenin levels. Together, these findings reveal the Pfn2 microRNA-IRE regulatory axis as a critical post-transcriptional regulatory node governing the switch from pluripotency to somatic differentiation.

MicroRNAs

iNOME-seq: in vivo simultaneous genome-wide mapping of chromatin accessibility, nucleosome positioning, DNA-binding protein sites, and DNA methylation in Arabidopsis.

We present iNOMe-seq, a novel method for in vivo simultaneous profiling of chromatin accessibility, nucleosome occupancy, DNA-binding protein sites, and DNA methylation in living tissues. iNOMe-seq utilizes an m5C methyltransferase to mark accessible cytosines in a GpC context, bypassing nucleosome-restricted regions. Using Arabidopsis thaliana, we demonstrate that iNOMe-seq improves chromatin accessibility quantification compared to existing methods. Furthermore, it allows for the spatial and temporal analysis of chromatin dynamics, transcription factor binding, and DNA methylation, offering insight into the role of epigenetic components in transcriptional regulation across tissues and genetic variations in natural populations.

Arabidopsis

Measuring FOXO Activity by Using qPCR-Based Expression Analysis of FOXO Target Genes.

FOXO transcription factors belong to the forkhead protein family and are distinguished by their unique forkhead (FKH) DNA-binding domain. In the realm of mammals, four FOXO paralogs are recognized: FOXO1, FOXO3, FOXO4, and FOXO6. These paralogs are evolutionary counterparts of the daf-16 gene discovered in the nematode C. elegans. A key feature shared by these paralogs is a consensus binding site known as the DAF-16 family protein-binding site (DBE: 5'-TTGTTTAC-3'). The functional outcome of FOXO transcription factors primarily hinges on their affinity for these specific binding sites within the promoters of their target genes. Nevertheless, it is worth noting that many of these target genes exhibit tissue-specific expression patterns. Consequently, there is not a single FOXO target gene whose expression can reliably serve as a universal indicator of FOXO activity across all cell types and tissues or in response to all stimuli. In light of these considerations, we present a collection of target genes that, when collectively assessed, can accurately gauge FOXO activation. In this chapter, we outline a specific protocol for utilizing quantitative reverse transcription polymerase chain reaction (qRT-PCR) to measure the expression levels of these genes.

Forkhead Transcription Factors

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Histone H3K27ac spreads from enriched chromatin domains into neighboring regions upon loss of CTCF binding.

Acetylation of histone H3 at lysine 27 (H3K27ac) is enriched at enhancers and highly transcribed genes. Our previous study showed that an H3K27ac-enriched chromatin domain expanded into neighboring regions following the deletion of CTCF-binding motifs flanking the domain. In this study, we explored the spreading of H3K27ac on a genome-wide scale by analyzing its distribution around CTCF-binding sites in human K562 cells and examining changes upon CTCF loss. We found that a subset of CTCF-binding sites demarcates H3K27ac-enriched domains. Upon loss of CTCF binding, H3K27ac levels increased in most regions adjacent to these domains, indicating that H3K27ac can spread into neighboring chromatin. This spreading was accompanied by elevated transcription of nearby genes. Chromatin features, including histone modifications, CTCF-binding intensity, and CTCF-mediated chromatin interactions, were associated with the H3K27ac spreading. Notably, enhancers were more enriched within domains that exhibited H3K27ac spreading compared to those that did not, and the deletion of enhancers from the CTCF motif-deficient β-globin locus attenuated the spreading. These findings indicate that CTCF-binding sites serve as boundaries for H3K27ac-enriched domains and that, in the absence of CTCF binding, H3K27ac can spread into neighboring regions. H3K27ac spreading appears to be influenced by multiple chromatin features and to contribute to the transcriptional increase of nearby genes.

CTCF