Search PubMedSearch

SEARCH · Search PubMed

Results for “Transcription initiation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Alternative tandem transcription initiation links noncoding variants to human disease through translational control.

Alternative tandem transcription initiation is a pervasive mechanism of gene regulation, yet its genetic impact on human disease remains largely unknown. Here, we systematically quantify the genetic regulation of alternative tandem transcription initiation across 25,859 samples from 49 normal human tissues and 33 tumor tissues. We identify approximately 0.4 million genetic variants associated with alternative transcription initiation in 5295 genes, with 32% operating independently of gene expression. Moreover, we discover 2238 multi-tissue alternative tandem transcription initiation outliers enriched for rare deleterious promoter and 5' UTR variants, demonstrating that both common and rare variants modulate transcription initiation. Strikingly, 74% of disease variants that colocalize with genetic variants regulating alternative transcription initiation cannot be identified through expression quantitative trait loci. Transcriptome-wide association studies identify 614 disease susceptibility genes associated with alternative transcription initiation, including known cancer drivers such as MAFF and MLLT10. Functional validation uncovers OSGEP as a breast cancer risk gene, where the alternative allele lengthens the 5' UTR and reduces protein abundance through upstream open reading frame-mediated translation repression, and suppresses breast cancer cell proliferation. Our findings establish alternative transcription initiation as a major, underappreciated mechanism associating noncoding variation with disease, providing a critical resource for interpreting disease risk loci.

Humans

BRD2 bridges TFIID and MOF-H4K16ac-containing nucleosomes to promote transcriptional initiation.

Members of the bromodomain and extraterminal domain (BET) protein family play a central role in transcription by RNA polymerase II (RNA Pol II). Small-molecule inhibitors that block interaction between BET bromodomains and acetylated histones have been developed for disease therapeutics. However, the BET protein BRD4 does not require bromodomains to perform its major transcriptional elongation control, and mechanisms by which other BET proteins regulate RNA Pol II remain insufficiently understood. Addressing the disparity between pan-BET degraders and BRD4-specific depletion, we report that the BET protein BRD2 generally functions to promote transcriptional initiation in a bromodomain-dependent manner at both promoters and enhancers in human cell lines. We demonstrate that BRD2 bromodomains preferentially bind to histone H4 harboring MOF-mediated H4K16ac, while the BRD2 C-terminal domain facilitates recruitment of TFIID. Our studies provide mechanistic insight into distinct roles for BRD2 and BRD4 in transcriptional initiation and elongation control for proper regulation of gene expression.

Humans

Evolution of Transcription Factor-containing Superfamilies in Eukaryotes.

Regulation of gene expression helps determine various phenotypes in most cellular life forms. It is orchestrated at different levels and at the point of transcription initiation by transcription factors (TFs). TFs bind to DNA through domains that are evolutionarily related, by shared membership of the same superfamilies (TF-SFs), to those found in other nucleic acid binding and protein-binding functions (nTFs for non-TFs). Here we ask how TF DNA binding sequence families in eukaryotes have evolved in relation to their nTF relatives. TF numbers scale by power law with the total number of protein-coding genes differently in different clades, with fungi usually showing sub-linear powers whereas chordates show super-linear scaling. The LECA probably encoded a complex regulatory machinery with both TFs and nTFs, but with an excess of nTFs when compared to the relative distribution of TFs and nTFs in extant organisms. Losses drive the evolution of TFs and nTFs, with the possible exception of TFs in animals for some tree topologies. TFs are highly dynamic in evolution, showing higher gain and loss rates than nTFs in some TF-SFs though both are conserved to similar extents. Gains of TFs and nTFs are driven by the appearance of a large number of new sequence clusters in a small number of nodes, which determine the presence of as many as a third of extant TFs and nTFs as well as the relative presence of TFs and nTFs. Whereas nodes showing explosion of TF numbers belong to multicellular clades, those for nTFs lie among the fungi and the protists.

Transcription Factors

The ARK2N-CK2 complex initiates transcription-coupled repair through enhancing the interaction of CSB with lesion-stalled RNAPII.

Transcription is extremely important for cellular processes but can be hindered by RNA polymerase II (RNAPII) pausing and stalling. Cockayne syndrome protein B (CSB) promotes the progression of paused RNAPII or initiates transcription-coupled nucleotide excision repair (TC-NER) to remove stalled RNAPII. However, the specific mechanism by which CSB initiates TC-NER upon damage remains unclear. In this study, we identified the indispensable role of the ARK2N-CK2 complex in the CSB-mediated initiation of TC-NER. The ARK2N-CK2 complex is recruited to damage sites through CSB and then phosphorylates CSB. Phosphorylation of CSB enhances its binding to stalled RNAPII, prolonging the association of CSB with chromatin and promoting CSA-mediated ubiquitination of stalled RNAPII. Consistent with this finding, Ark2n-/- mice exhibit a phenotype resembling Cockayne syndrome. These findings shed light on the pivotal role of the ARK2N-CK2 complex in governing the fate of RNAPII through CSB, bridging a critical gap necessary for initiating TC-NER.

DNA Repair Enzymes

Predicting and comparing transcription start sites in single cell populations.

The advent of 5' single-cell RNA sequencing (scRNA-seq) technologies offers unique opportunities to identify and analyze transcription start sites (TSSs) at a single-cell resolution. These technologies have the potential to uncover the complexities of transcription initiation and alternative TSS usage across different cell types and conditions. Despite the emergence of computational methods designed to analyze 5' RNA sequencing data, current methods often lack comparative evaluations in single-cell contexts and are predominantly tailored for paired-end data, neglecting the potential of single-end data. This study introduces scTSS, a computational pipeline developed to bridge this gap by accommodating both paired-end and single-end 5' scRNA-seq data. scTSS enables joint analysis of multiple single-cell samples, starting with TSS cluster prediction and quantification, followed by differential TSS usage analysis. It employs a Binomial generalized linear mixed model to accurately and efficiently detect differential TSS usage. We demonstrate the utility of scTSS through its application in analyzing transcriptional initiation from single-cell data of two distinct diseases. The results illustrate scTSS's ability to discern alternative TSS usage between different cell types or biological conditions and to identify cell subpopulations characterized by unique TSS-level expression profiles.

Transcription Initiation Site

Insights into the regulation of the HOTAIR proximal promoter.

HOTAIR (HOX transcript antisense RNA) is a HOXC-cluster long intervening non-coding RNA (lincRNA) whose cancer relevance is tightly coupled to how its transcription is wired into hormone, hypoxia, inflammatory, and developmental signaling. HOTAIR is known to associate with cancer cell proliferation, motility, tumor invasion, and metastasis. The present mini-review focuses on the regulatory architecture and mechanistic complexity of HOTAIR transcriptional regulation, with emphasis on three organizing principles. First, we consider the impact of promoter choice between a canonical proximal promoter (P1), which supports the 2.2-2.4 kb transcript, and an alternative upstream promoter/TSS (P2), which contributes to context-dependent transcription initiation. Second, we examine the long-distance enhancer-promoter communication between HOTAIR distal enhancer and P1/P2. Third, we summarize the recent epigenetic and epi-transcriptomic mechanisms involved in HOTAIR transcript initiation and elongation. A combination of these events determines isoform-specific transcription to govern cell-type-, context-, and cancer specific modulation of HOTAIR expression that promotes tumor formation and cancer progression. Finally, the review proposes how large-scale RNA datasets, long-read sequencing, and isoform-specific studies can refine our understanding of this versatile lincRNA's regulation.

Humans

Single-nucleotide transcription start sites profiling via Nascent Strand-Specific RNA sequencing uncovers IFN-γ-induced promoter dynamics.

Transcriptional regulation is a highly dynamic process in which nascent RNAs provide the most immediate readout of transcriptional activity. Precise mapping of transcription start sites (TSSs) is therefore critical for understanding promoter architecture and gene regulation, yet remains technically challenging. Here, we introduce Nascent Strand-Specific RNA sequencing (NSS-seq), a robust and streamlined method for genome-wide profiling of the capped 5' ends of nascent RNAs. By directly capturing transcription initiation events, NSS-seq overcomes the temporal delay inherent to conventional RNA-seq and enables time-resolved interrogation of transcriptional dynamics. Applied to interferon-γ (IFN-γ)-stimulation, NSS-seq uncovers previously unrecognized IFN-γ-responsive genes and transient transcription factor activation patterns underlying interferon-mediated tumor-suppressive functions. Together, NSS-seq provides a cost-effective and technically accessible platform for dissecting promoter-level regulatory dynamics during cellular responses.

Promoter Regions, Genetic

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data.

Long noncoding RNAs (lncRNAs) play critical regulatory roles in human development and disease. Although there are over 100,000 samples with available RNA sequencing (RNA-seq) data, many lncRNAs have yet to be annotated. The conventional approach to identifying novel lncRNAs from RNA-seq data is to find transcripts without coding potential but this approach has a false discovery rate of 30-75%. Other existing methods either identify only multi-exon lncRNAs, missing single-exon lncRNAs, or require transcriptional initiation profiling data (such as H3K4me3 ChIP-seq data), which is unavailable for many samples with RNA-seq data. Because of these limitations, current methods cannot accurately identify novel lncRNAs from existing RNA-seq data. To address this problem, we have developed software, Flnc, to accurately identify both novel and annotated full-length lncRNAs, including single-exon lncRNAs, directly from RNA-seq data without requiring transcriptional initiation profiles. Flnc integrates machine learning models built by incorporating four types of features: transcript length, promoter signature, multiple exons, and genomic location. Flnc achieves state-of-the-art prediction power with an AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the conventional approach to over 85%. Flnc is available via GitHub platform.

RNA-seq

Transcription Start Regions in PTU-intergenic regions drive cell cycle-dependent transcriptional activation events in Leishmania donovani.

Leishmania displays an unconventional mode of transcription, with long clusters of genes being transcribed polycistronically from Transcription Start Regions (TSRs), being processed into monocistronic units prior to translation. It has long been believed that transcription is constitutive: failure to identify consensus sequences across TSRs (except a GT-rich motif supporting transcription in Trypanosoma brucei) and absence of canonical eukaryotic transcription factors led to the conclusion that regulation is primarily post-transcriptional, with epigenetics playing a role in triggering transcription initiation. This study stems from our previous findings identifying a few genes to be activated in a cell cycle-dependent manner. Using nuclear run-on assays to analyze nascent transcripts of two chromosomes, chromosomes 2 and 14, we find that while most genes are constitutively transcribed, a subset of genes gets activated at specific cell cycle stages. Reporter assays reveal that this transcriptional activation is driven by the regions immediately upstream of the genes. Sequence analyses of these TSRs lying in polycistronic intergenic regions (PIRs) uncovered a 10-mer GT-rich motif, in synchrony with earlier findings in T. brucei identifying a GT-rich motif at bidirectional TSRs. We also identify a second 25-mer motif at these TSRs, and deletion analyses find this motif to be critical for regulating gene expression. The findings of this study reveal that transcriptional events in these unicellular parasites are more complex than believed thus far: not all transcriptional events are constitutive, polycistronic transcription is not the only mode of transcription, and cis-acting sequence elements regulate at least some transcriptional events in these parasites.IMPORTANCEEndemic to 90 countries, Leishmania parasites cause a spectrum of diseases called Leishmaniases. No vaccines for human use are available to date, and the drugs currently used to treat the disease are expensive, have toxic side effects, and have complex administration regimens, with emerging drug resistance compounding problems. Researchers continue to investigate Leishmania cellular processes, with the hope of uncovering new therapeutic target sites. Gene regulation in these parasites is unusual, being modulated by various mechanisms, including epigenetic modifications, gene dosage, and post-transcriptional processing. Transcription is typically polycistronic and constitutive, initiating from Transcription Start Regions (TSRs) lying upstream of the first gene in the polycistronic transcription unit (PTU). The work presented here reveals that a subset of genes is transcribed monocistronically in a cell cycle-dependent manner from Transcription Start Regions lying in the PTU-intergenic regions (PIRs), underscoring the complexities of gene regulation in these parasites.

Leishmania donovani

PotatoRTD and TomatoRTD: Comprehensive Reference Transcript Datasets for Accurate Transcriptome Analysis and Isoform Discovery.

Transcriptome annotations provide essential information on transcript locations, sequences and structures, including transcription start, end sites and splice junctions. They underpin key biological analyses such as gene and transcript quantification, and the study of transcriptional and post-transcriptional regulation, including alternative transcription initiation, polyadenylation and splicing. Accurate characterisation of transcript isoforms is critical for understanding how gene expression relates to functional protein products. However, for many species-including Solanaceae crops such as potato and tomato-current annotations suffer from limited isoform coverage, with tens or hundreds of thousands of splice junctions and transcript isoforms missing. This undermines the completeness and accuracy of transcript-level analyses. Here, by generating Iso-seq and RNA-seq on a range of tissues and samples, we have produced transcriptome annotations for both potato and tomato with improved coverage, diversity, accurate splice junctions, and transcript start and end sites. We have also made these high-quality resources accessible through genome browsers. These enhanced annotations will enable more accurate transcriptome analyses, supporting higher-resolution and novel biological discoveries.

Solanum tuberosum

A single cluster of RNA Polymerase II molecules is stably associated with active genes.

In eukaryotic nuclei, transcription is associated with the clustering of RNA Polymerase II (RNAPII) molecules. The mechanisms underlying cluster formation, their interactions with genes, and their impact on transcriptional activity remain heavily debated. Here we take advantage of the naturally occurring increase in transcriptional activity during Zygotic Genome Activation (ZGA) in Drosophila melanogaster embryos to characterize the functional roles of RNAPII clusters in a developmental context. Using single-molecule tracking and lattice light-sheet microscopy, we find that RNAPII cluster formation depends on transcription initiation, and that cluster lifetimes depend on transcriptional activity when not constrained by interphase duration. We show that single clusters are stably associated with active gene loci during transcription and that cluster intensities are strongly correlated with transcriptional output. Collectively our data and simulations on cluster formation kinetics show that RNAPII clusters reflect local accumulations of transcriptionally engaged polymerases and do not form through higher-order mechanisms such as phase separation.

Journal Article

The HIV-1 Transcriptional Program: From Initiation to Elongation Control.

A large body of work in the last four decades has revealed the key pillars of HIV-1 transcription control at the initiation and elongation steps. Here, I provide a recount of this collective knowledge starting with the genomic elements (DNA and nascent TAR RNA stem-loop) and transcription factors (cellular and the viral transactivator Tat), and later transitioning to the assembly and regulation of transcription initiation and elongation complexes, and the role of chromatin structure. Compelling evidence support a core HIV-1 transcriptional program regulated by the sequential and concerted action of cellular transcription factors and Tat to promote initiation and sustain elongation, highlighting the efficiency of a small virus to take over its host to produce the high levels of transcription required for viral replication. I summarize new advances including the use of CRISPR-Cas9, genetic tools for acute factor depletion, and imaging to study transcriptional dynamics, bursting and the progression through the multiple phases of the transcriptional cycle. Finally, I describe current challenges to future major advances and discuss areas that deserve more attention to both bolster our basic knowledge of the core HIV-1 transcriptional program and open up new therapeutic opportunities.

HIV-1

Structural Characterization of Native RNA Polymerase II Transcription Complexes and Nucleosomes in Drosophila melanogaster.

Structural studies of eukaryotic RNA polymerase II (Pol II) transcription often rely on in vitro assembly, which may not fully represent native conditions. To investigate Pol II transcription in metazoan cells, we developed a method to isolate native transcription complexes from Drosophila melanogaster embryos using FLAG-tag affinity purification and Micrococcal Nuclease treatment. Cryo-EM and proteomics studies revealed diverse transcription complexes and nucleosomes, including a metazoan Rpb4/Rpb7 stalk-less Pol II elongation complex and a hexameric nucleosome lacking an H2A/H2B dimer. Notably, nucleosome is found only downstream of the nucleosome elongation complex, underscoring it as a major energy barrier and a time-consuming step during Pol II progression through chromatin. Proteomics identified co-purified factors involved in transcription initiation, elongation, and RNA modification. This study provides a framework for investigations of transcription in cells, paving the way for future studies of transient and minor complexes.

Animals

Dlx3 transcriptional regulation of osteoblast differentiation: temporal recruitment of Msx2, Dlx3, and Dlx5 homeodomain proteins to chromatin of the osteocalcin gene.

Genetic studies show that Msx2 and Dlx5 homeodomain (HD) proteins support skeletal development, but null mutation of the closely related Dlx3 gene results in early embryonic lethality. Here we find that expression of Dlx3 in the mouse embryo is associated with new bone formation and regulation of osteoblast differentiation. Dlx3 is expressed in osteoblasts, and overexpression of Dlx3 in osteoprogenitor cells promotes, while specific knock-down of Dlx3 by RNA interference inhibits, induction of osteogenic markers. We characterized gene regulation by Dlx3 in relation to that of Msx2 and Dlx5 during osteoblast differentiation. Chromatin immunoprecipitation assays revealed a molecular switch in HD protein association with the bone-specific osteocalcin (OC) gene. The transcriptionally repressed OC gene was occupied by Msx2 in proliferating osteoblasts, while Dlx3, Dlx5, and Runx2 were recruited postproliferatively to initiate transcription. Dlx5 occupancy increased over Dlx3 in mature osteoblasts at the mineralization stage of differentiation, coincident with increased RNA polymerase II occupancy. Dlx3 protein-DNA interactions stimulated OC promoter activity, while Dlx3-Runx2 protein-protein interaction reduced Runx2-mediated transcription. Deletion analysis showed that the Dlx3 interacting domain of Runx2 is from amino acids 376 to 432, which also include the transcriptionally active subnuclear targeting sequence (376 to 432). Thus, we provide cellular and molecular evidence for Dlx3 in regulating osteoprogenitor cell differentiation and for both positive and negative regulation of gene transcription. We propose that multiple HD proteins in osteoblasts constitute a regulatory network that mediates development of the bone phenotype through the sequential association of distinct HD proteins with promoter regulatory elements.

Amino Acid Sequence

Deep learning guided programmable design of Escherichia coli core promoters from sequence architecture to strength control.

Core promoters are essential regulatory elements that control transcription initiation, but accurately predicting and designing their strength remains challenging due to complex sequence-function relationships and the limited generalizability of existing AI-based approaches. To address this, we developed a modular platform integrating rational library design, predictive modelling, and generative optimization into a closed-loop workflow for end-to-end core promoter engineering. Conserved and spacer region of core promoters exert distinct effects on transcriptional strength, with the former driving large-scale variation and the latter enabling finer gradation. Based on this insight, Mutation-Barcoding-Reverse Sequencing approach was used and constructed a synthetic promoter library comprising 112 955 variants with minimal redundancy and a 16 226-fold expression range. A Transformer-based model trained on this dataset achieved a Pearson correlation of 0.87 with experimentally measured promoter strengths. When combined with a conditional diffusion model, the system enabled de novo generation of promoter sequences with defined strengths, achieving a design-to-measurement correlation of 0.95 and maintaining high accuracy (R = 0.93) across varied sequence contexts. The designed promoters consistently preserved their intended strength gradients, demonstrating robust plug-and-play functionality. This work establishes a scalable and extensible platform (www.yudenglab.com) for deep learning-guided programmable design of Escherichia coli core promoters, enabling precise transcriptional control.

Promoter Regions, Genetic

E2F1 induces a G0-G1 reentry transcriptional program without changing chromatin accessibility.

Quiescent cells actively repress cell-cycle genes via chromatin-based mechanisms to maintain a non-dividing state, yet remain poised to reenter upon stimulation. E2F1, a canonical activator of cell-cycle genes, is sufficient to induce reentry from quiescence, but how it overcomes chromatin-mediated repression remains unclear. Here, we show that inducible E2F1 expression triggers exit from quiescence and progression through the cycle without changes in chromatin accessibility, by harnessing regulatory elements with limited, pre-existing accessibility. Using time-resolved transcriptomics, we demonstrate that E2F1 induces an accelerated transcriptional program compared to serum. Unlike serum, which triggers broad chromatin remodeling, E2F1-induced activation occurs in a context of limited accessibility. ChIP-seq reveals that E2F1 directly binds target sites in quiescent cells to upregulate canonical genes. Biochemical reconstitution shows that E2F1 binds nucleosomes and accesses internal E2F sites within histone-wrapped DNA. These findings suggest that E2F1 can engage nucleosome-associated DNA and initiate transcription without major chromatin reorganization, redefining transcription factor-chromatin dynamics during cell fate transitions and establishing E2F1 as a potent regulator of cell-cycle reentry.

Journal Article

Bimodality in E. coli gene expression: Sources and robustness to genome-wide stresses.

Bacteria evolved genes whose single-cell distributions of expression levels are broad, or even bimodal. Evidence suggests that they might enhance phenotypic diversity for coping with fluctuating environments. We identified seven genes in E. coli with bimodal (low and high) single-cell expression levels under standard growth conditions and studied how their dynamics are modified by environmental and antibiotic stresses known to target gene expression. We found that all genes lose bimodality under some, but not under all, stresses. Also, bimodality can reemerge upon cells returning to standard conditions, which suggests that the genes can switch often between high and low expression rates. As such, these genes could become valuable components of future multi-stable synthetic circuits. Next, we proposed models of bimodal transcription dynamics with realistic parameter values, able to mimic the outcome of the perturbations studied. We explored several models' tunability and boundaries of parameter values, beyond which it shifts to unimodal dynamics. From the model results, we predict that bimodality is robust, and yet tunable, not only by RNA and protein degradation rates, but also by the fraction of time that promoters remain unavailable for new transcription events. Finally, we show evidence that, although the empirical expression levels are influenced by many factors, the bimodality emerges during transcription initiation, at the promoter regions and, thus, may be evolvable and adaptable.

Escherichia coli

Promoter identity shapes splicing outcomes and fidelity.

Gene expression is a complex process subject to regulation at multiple functionally interconnected levels. One prominent example is the crosstalk between transcription and splicing regulation. Past work has shown that transcription can influence splicing in multiple ways, but a systematic investigation of this complex interplay is lacking. Here we employ massively parallel reporter assays of large combinatorial promoter-splice site libraries to dissect how promoter identity and transcription dynamics affect alternative splicing in human cells. We find that promoter identity, rather than expression level, exerts strong and highly context-specific effects on cassette exon inclusion, exceeding the effect of pharmacological inhibitors of transcription initiation or elongation. Groups of exons display coordinated promoter-dependent splicing behavior, and we identified predictive sequence and structural features underlying this sensitivity. Promoter and gene architecture also shape isoform diversity by modulating cryptic splice site usage. These findings present promoters as central regulators of splicing outcomes and fidelity.

Humans