Search PubMedSearch

SEARCH · Search PubMed

Results for “sequence-specificity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

CRISPR/Cas- and Argonaute-Based In Vivo Nucleic-Acid Imaging Technologies: Strategies, Challenges, and Perspectives.

Live-cell monitoring of sequence-specific nucleic acids is essential to understanding genome organization, RNA regulation, and disease progression. Clustered regularly interspaced short palindromic repeat (CRISPR)/CRISPR-associated protein (Cas) and Argonaute (Ago) systems provide programmable, guide-directed recognition of DNA or RNA and are increasingly used as platforms for in vivo bioimaging. This review summarizes the structural and mechanistic features of representative CRISPR and Ago effectors and discusses design strategies for sensitive, specific, and multiplexed imaging of genomic loci, extrachromosomal DNA, and endogenous RNA in living cells. We compare the analytical performance and limitations of CRISPR- and Ago-based imaging, with particular emphasis on the major technical and biological challenges affecting their accuracy, applicability, and reliability. Finally, this review offers insights into developing high-resolution and user-friendly bioimaging platforms for fundamental biology and future translational applications.

CRISPR

Nucleosomes and IDRs suppress promiscuous GCN4 binding on minichromosomes.

Eukaryotic sequence-specific transcription factors (TFs) must find their cognate DNA targets hidden in genomic chromatin amid an excess of nonspecific sequences and degenerate motifs. Although static TF interactions with nucleosomal targets have been elucidated, how TFs efficiently search for cognate sites within native gene-sized chromatin domains has been unclear. Here we used purified Saccharomyces cerevisiae HIS3 minichromosomes and single-molecule imaging to compare association and dissociation kinetics of transcription activator GCN4 on chromatin and naked genomic DNA. GCN4 displays widespread and stable off-target binding on bare DNA because of entrapment by degenerate sites and interactions with nonspecific DNA of increasing length, indicative of one-dimensional (1D) diffusion. Nucleosome organization on the minichromosome reduces promiscuous GCN4 residence times by obstructing TF association and restricting 1D target search within nucleosome-free regions. Furthermore, the intrinsically disordered GCN4 activation domain independently enhances targeting efficiency and specificity by accelerating association-dissociation kinetics in vitro and in living cells. Altogether, both nucleosome organization and activation domains independently suppress promiscuous GCN4 binding, which, if unchecked, may cause aberrant cryptic transcription known to occur upon chromatin disruptions.

Nucleosomes

The CGG triplet repeat binding protein 1 counteracts R-loop induced transcription-replication stress.

The CGG triplet repeat binding protein 1 (CGGBP1) binds to CGG repeats and has several important cellular functions, but how this DNA sequence-specific binding factor affects transcription and replication processes is an open question. Here, we show that CGGBP1 binds human gene promoters containing short (<&#x2009;5) CGG-repeat tracts prone to R-loop formation. Loss of CGGBP1 leads to deregulated transcription, transcription-replication-conflicts (TRCs) and accumulation of Serine-5 phosphorylated RNA polymerase II (RNAPII), indicative of promoter-proximal stalling and a defect in transcription elongation. Consistently, an episomal CGG-repeat-containing model locus as well as endogenous genes show deregulated transcription, R-loop accumulation and increased RNAPII chromatin occupancy in CGGBP1-depleted cells. We identify the DEAD-box RNA:DNA helicases DDX41 and DHX15 as interaction partners specifically recruited by CGGBP1. Co-depletion experiments show that DDX41 and CGGBP1 work in the same pathway to unwind R-loops and avoid TRCs. Together, our work shows that short trinucleotide repeats are a source of genome-destabilizing secondary structures, and cells rely on specific DNA-binding factors to maintain proper transcription and replication coordination at short CGG repeats.

Humans

High prevalence of PRDM9-independent recombination hotspots in placental mammals.

In many mammals, recombination events are concentrated in hotspots directed by a sequence-specific DNA-binding protein named PRDM9. Intriguingly, PRDM9 has been lost several times in vertebrates, and notably among mammals, it has been pseudogenized in the ancestor of canids. In the absence of PRDM9, recombination hotspots tend to occur in promoter-like features such as CpG islands. It has thus been proposed that one role of PRDM9 could be to direct recombination away from PRDM9-independent hotspots. However, the ability of PRDM9 to direct recombination hotspots has been assessed in only a handful of species, and a clear picture of how much recombination occurs outside of PRDM9-directed hotspots in mammals is still lacking. In this study, we derived an estimator of past recombination activity based on signatures of GC-biased gene conversion in substitution patterns. We quantified recombination activity in PRDM9-independent hotspots in 52 species of boreoeutherian mammals. We observe a wide range of recombination rates at these loci: several species (such as mice, humans, some felids, or cetaceans) show a deficit of recombination, while a majority of mammals display a clear peak of recombination. Our results demonstrate that PRDM9-directed and PRDM9-independent hotspots can coexist in mammals and that their coexistence appears to be the rule rather than the exception. Additionally, we show that the location of PRDM9-independent hotspots is relatively more stable than that of PRDM9-directed hotspots, but that PRDM9-independent hotspots nevertheless evolve slowly in concert with DNA hypomethylation.

Animals

Characterization of the integration protein of bacteriophage lambda as a site-specific DNA-binding protein.

The Int protein specified by bacteriophage lambda is required for the recombination event that integrates the viral DNA into the host genome at its specific attachment site. Using a DNA-binding assay, we have partially purified the Int protein and studied some of the features of its binding specificity and regulation. The DNA-binding activity is attributed to Int protein because the activity is eliminated by a nonsense mutation or a deletion in the int gene, and is rendered thermolabile by temperature-sensitive mutations in the int gene. The DNA-binding activity is specific for DNA carrying an appropriate attachment site, suggesting that Int protein directs the sequence-specific recognition essential for integrative recombination. The specific DNA-binding activity is also missing after infection by phage carrying mutations in the cII and cIII regulatory genes of lambda. This finding corroborates the conclusion from other types of experiments that regulation of the int and cI genes by cII/cIII provides for coordinate regulation of both major events of the lysogenic response, establishment of repression and insertion of viral DNA.

Carrier Proteins

Variants in the interferon regulatory factor 5 gene confer genetic risk for systemic lupus erythematosus in a Han Chinese population.

BACKGROUND: Interferon regulatory factor 5 (IRF5), integral to interferon signaling pathways, has been identified as a susceptibility locus for systemic lupus erythematosus (SLE). Nevertheless, the relationship between IRF5 variants and SLE risk within the Han Chinese demographic remains inadequately characterized. MATERIALS AND METHODS: Genotyping of two functional single nucleotide variants (SNVs) in IRF5 was conducted in 167 individuals with SLE and 246 healthy controls utilizing sequence-specific primer polymerase chain reaction (PCR-SSP). Chi-square and Fisher's exact tests were employed to assess associations. RESULTS: The rs10954213 variant demonstrated a significant association with SLE susceptibility under the recessive model (GG vs. AG+AA, OR = 2.20, 95% CI: 1.30-3.75, p&#x2009;=&#x2009;0.003, adjusted p [pc]&#x2009;=&#x2009;0.030) and homozygous model (GG vs. AA, OR = 2.43, 95% CI: 1.36-4.42, p&#x2009;=&#x2009;0.003, pc = 0.032). Similarly, the rs2004640 variant was associated with an increased risk of SLE across allelic (T vs. G, OR = 1.66, 95% CI: 1.22-2.26, p&#x2009;=&#x2009;0.001, pc = 0.011), dominant (TG+TT vs. GG, OR = 1.77, 95% CI: 1.19-2.63, p&#x2009;=&#x2009;0.005, pc = 0.047), and homozygous models (TT vs. GG, OR = 3.72, 95% CI: 1.58-8.78, p&#x2009;=&#x2009;0.002, pc = 0.016). Haplotype analysis identified protective haplotype HT1 (A/G, OR = 0.54, 95% CI: 0.41-0.73, p&#x2009;<&#x2009;0.001) and risk haplotype HT4 (G/T, OR = 2.51, 95% CI: 1.42-4.42, p&#x2009;=&#x2009;0.001). CONCLUSIONS: These findings indicate that IRF5 gene variants substantially modulate susceptibility to SLE in the Han Chinese population. They hold potential as biomarkers for evaluating SLE risk and offer valuable perspectives into disease pathogenesis.

Adult

TripLexicon: prediction and analysis of gene regulatory RNA-DNA interactions.

MOTIVATION: Non-coding RNA (ncRNA) plays a crucial role in gene regulation, including by forming sequence-specific RNA-DNA interactions at gene regulatory elements. One form of interaction takes place via the formation of RNA:DNA:DNA triple helices (triplexes). Accurate computational prediction of triplex formation from nucleotide sequences is an important tool in ncRNA research but remains somewhat inaccessible and complex. To address this, we created TripLexicon, a web-based interface for accessing and analyzing predicted gene regulatory RNA-DNA interactions in human and mouse. RESULTS: Predicted interactions can be accessed from RNA-, DNA-, and region-centric perspectives. For each RNA transcript, visualizations at genome and nucleotide resolution are available, providing insight into target genes and regions, as well as putative functional domains of the transcript. Predicted target genes can immediately be subjected to ontology and pathway enrichment analysis, providing rapid insight into potential functions mediated by the RNA-DNA interactions of the queried transcript. DNA and region queries are designed to identify potentially important ncRNA interactors at sites of interest. AVAILABILITY AND IMPLEMENTATION: TripLexicon is accessible at https://triplexicon.uni-frankfurt.de. This website is free and open to all users and there is no login requirement. All data and code is uploaded to Zenodo: https://zenodo.org/records/17143608 and the code for the webserver is available on Github: https://github.com/SchulzLab/TripLexicon.

Software

Target-Site Selection by Transcription Factors: Roles of DNA, Chromatin, and Cofactor-Mediated Regulation.

Transcription factors (TFs) are sequence-specific DNA-binding proteins that regulate gene-expression programs and cell fate. The ability of a defined combination of four TFs to reprogram differentiated cells into induced pluripotent stem cells illustrates the powerful role of TFs in determining cellular identity. However, TFs usually recognize short and degenerate DNA motifs of approximately 6-12 base pairs, generating thousands to millions of potential motif matches in mammalian genomes. In living cells, TFs occupy only a restricted subset of these sites, indicating that motif presence alone is insufficient for functional target selection. Several layers of regulation contribute to this selective occupancy, including DNA methylation, nucleosome organization, histone modifications, chromatin remodeling, TF oligomerization, TF availability and localization, and cofactors that regulate DNA-binding domains. This review outlines how DNA/chromatin features and TF-centered mechanisms contribute to target-site selection. The principal aim is to highlight DNA-binding domain-directed cofactor regulation as an underappreciated mechanism that modulates TF-DNA binding and may help explain selective genomic occupancy.

Target-site selection

CRISPR-mediated proximity labeling unveils ABHD14B as a host factor to regulate HBV cccDNA transcriptional activity.

BACKGROUND: The long-term goal of chronic hepatitis B research is a functional cure (HBsAg seroclearance). Although currently used nucleos(t)ide analogs can efficiently inhibit viral replication, they do not reduce viral RNAs or proteins produced from covalently closed circular DNA (cccDNA), and rarely achieve a functional cure. To overcome this situation, revealing the mode of the existence of cccDNA is required, including identifying the interreacting proteins with cccDNA. Here, we aimed to identify novel proteins that interact with cccDNA. METHODS: Using an in vitro HBV infection model and a sequence-specific proximity labelling method consisting of dead Cas9 and biotin identification (BioID2), we comprehensively determined proteins that possibly interact with cccDNA. After identifying the candidate proteins, the HBV RNA transcription levels were examined by knocking out the associated genes. RESULTS: We identified ABHD14B as a protein that interacts with cccDNA and inhibits HBV RNA transcription from cccDNA. ABHD14B decreases the acetylation levels of histone proteins that control the transcription levels of HBV RNA in cccDNA. Moreover, ABHD14B interacts with TFII-I, which binds directly to cccDNA in a sequence-dependent manner. These results suggest that the host protein, ABHD14B, is recruited to cccDNA via the TFII-I protein, inhibiting HBV RNA transcription from cccDNA by deacetylating cccDNA histones. CONCLUSIONS: ABHD14B was newly identified as a suppressor of HBV RNA transcription from cccDNA, which may improve our understanding of the mode of existence of cccDNA, providing a basis for development of a functional cure.

Humans

Write and Read: Harnessing Synthetic DNA Modifications for Nanopore Sequencing.

An exciting feature of nanopore sequencing is its ability to record multi-omic information on the same sequenced DNA molecule. Well-trained models allow the detection of nucleotide-specific molecular signatures through changes in ionic current as DNA molecules translocate through the nanopore. Thus, naturally occurring DNA modifications, such as DNA methylation and hydroxymethylation, may be recorded simultaneously with the genetic sequence. Additional genomic information, such as chromatin state or the locations of bound transcription factors, may also be recorded if their locations are chemically encoded into the DNA. Here, we present a versatile "write-and-read" framework, where chemo-enzymatic DNA labeling with unnatural synthetic tags results in predictable electrical fingerprints in nanopore sequencing. As a proof-of-concept, we explore a DNA glucosylation approach that selectively modifies 5-hydroxymethylcytosine (5hmC) with glucose or glucose-azide adducts. We demonstrate that these modifications generate distinct and reproducible electrical shifts, enabling the direct detection of chemically altered nucleotides. We further demonstrate that enzymatic alkylation, such as the enzymatic transfer of azide residues to the N6 position of adenines, also produces characteristic nanopore signal shifts relative to the native adenine and 6-methyladenine. Beyond direct nucleotide detection, this approach introduces new possibilities for bio-orthogonal DNA labeling, enabling an extended alphabet of sequence-specific detectable moieties. The future use of programmable chemical modifications for simultaneous analysis of multiple omics features on individual molecules opens new avenues for genetic research and discovery.

5-hydroxymethylcytosine (5hmC)

Foxh1 is a locus-specific PRC2 recruiter governing germ layer silencing.

Polycomb Repressive Complex 2 (PRC2) establishes H3K27me3 marks to shape spatiotemporal gene expression during embryogenesis. While its dysregulation is linked to developmental disorders, cancer, and aging, the mechanisms guiding PRC2 to specific genomic loci remain a subject of ongoing debate. A prevailing model proposes that PRC2 recruitment occurs via its intrinsic affinity for chromatin rather than through sequence-specific transcription factors. Here, we provide evidence that the maternally deposited pioneer transcription factor Foxh1 plays a critical role in directing PRC2 to specific genomic loci during zygotic genome activation in Xenopus. Foxh1 is a critical transcription factor mediating Nodal signaling, but it also plays an earlier role by pre-binding enhancers prior to signaling activation. This pre-binding is essential for forming enhanceosome complexes that trigger mesendodermal gene expression and drive gastrulation, in cooperation with other maternal transcription factors. Using maternal Foxh1-null embryos, we demonstrate that Foxh1 directly recruits Ezh2, the catalytic subunit of PRC2, to Foxh1-bound loci. Loss of Foxh1 impairs Ezh2 recruitment, leading to a global reduction in H3K27me3. These findings support a dual-function model in which Foxh1 not only activates endodermal gene expression in endoderm, but also recruits PRC2 to silence the same genes in ectoderm. This dual activity of Foxh1 allows the spatially coordinated epigenetic states of the endodermal gene regulatory program during early embryogenesis.

CRISPR/Cas9

Quantitative analysis of DNA-GATA1 binding alterations linked to hematopoietic disorders.

GATA1 is a crucial transcription factor involved in hematopoiesis and mutations in this gene are linked to severe hematological disorders, including anemia, thrombocytopenia, Down syndrome-related transient abnormal myelopoiesis (DS-TAM), and myeloid leukemia of Down syndrome (ML-DS). Despite significant clinical interest in the molecular level characterization of GATA1 mutations, a comprehensive understanding of their impact on DNA binding is limited. Efforts to conduct detailed studies on full-length recombinant GATA1 have faced significant technical challenges, while alternative approaches are limited by low throughput or qualitative nature. Here, we introduce a native holdup (nHU) assay designed to systematically quantify DNA-protein interactions and is suitable for studying the impact of transcription factor mutations on DNA binding affinity. First, using the erythroid-specific ATP2B4 promoter as a model, we demonstrate that nHU can capture sequence-specific interactions and detect even subtle differences in DNA binding affinities. Then, we quantitatively characterize the impact of pathological mutations on DNA binding affinities in the context of full-length human GATA1. Our findings reveal that the GATA1s isoform, lacking the N-terminal transactivation domain (N-TAD), binds to DNA with increased affinity, while the R307C mutation reduces binding to the ATP2B4 erythroid promoter. In harmony with these observations, GATA1s exhibits increased functional activity, while the R307C mutation results in decreased activity. This study demonstrates the power of the nHU assay for studying DNA interactions of transcription factor variants and providing insight into the molecular mechanism of related diseases.

GATA1 Transcription Factor

Pentatricopeptide repeat protein targeting CUG repeat RNA ameliorates RNA toxicity in a myotonic dystrophy type 1 mouse model.

Myotonic dystrophy type 1 (DM1) is an autosomal dominant multisystemic disorder caused by the expansion of a CTG-triplet repeat in the 3' untranslated region of the dystrophia myotonica protein kinase (DMPK) gene. It results in the transcription of toxic RNAs that contain expanded CUG repeats (CUGexp). Splicing factors, such as muscleblind-like 1 (MBNL1), are sequestered by CUGexp, thereby disrupting the normal splicing program that is essential for various cellular functions. Pentatricopeptide repeat (PPR) proteins, originally found in plants, regulate RNA in organelles by binding in a sequence-specific manner. Here, we designed PPR proteins that specifically bind to the hexamer of CUG repeat RNAs (CUG-PPRs) and showed that CUG-PPR1 could ameliorate RNA toxicity induced by CUGexp in cell models of DM1. A single systemic recombinant adeno-associated virus (AAV9) vector-mediated gene delivery of CUG-PPR1 demonstrated long-term therapeutic effects on myotonia and restored splicing activity in a mouse model of DM1. These results highlight the potential of PPR molecules to target pathogenic RNA sequences in DM1 and potentially other RNA-mediated disorders.

Animals

Cytomegalovirus strain differentiation by DNA restriction analysis.

The heterogeneity of CMV DNA obtained from standard strains and new isolates, including a vaccination strain (Towne 125), was investigated. The cleavage patterns produced by the restriction endonucleases Eco RI and Bam 1 revealed stable strain specificities of CMV. On the other hand, a remarkable homology of sequence-specific CMV DNA fragmentation was demonstrated. A CMV subtyping relevant to clinical questions seems to be improbable.

Antigens, Viral

An updated compendium and reevaluation of the evidence for nuclear transcription factor occupancy over the mitochondrial genome.

In most eukaryotes, mitochondrial organelles contain their own genome, usually circular, which is the remnant of the genome of the ancestral bacterial endosymbiont that gave rise to modern mitochondria. Mitochondrial genomes are dramatically reduced in their gene content due to the process of endosymbiotic gene transfer to the nucleus; as a result most mitochondrial proteins are encoded in the nucleus and imported into mitochondria. This includes the components of the dedicated mitochondrial transcription and replication systems and regulatory factors, which are entirely distinct from the information processing systems in the nucleus. However, since the 1990s several nuclear transcription factors have been reported to act in mitochondria, and previously we identified 8 human and 3 mouse transcription factors (TFs) with strong localized enrichment over the mitochondrial genome using ChIP-seq (Chromatin Immunoprecipitation) datasets from the second phase of the ENCODE (Encyclopedia of DNA Elements) Project Consortium. Here, we analyze the greatly expanded in the intervening decade ENCODE compendium of TF ChIP-seq datasets (a total of 6,153 ChIP experiments for 942 proteins, of which 763 are sequence-specific TFs) combined with interpretative deep learning models of TF occupancy to create a comprehensive compendium of nuclear TFs that show evidence of association with the mitochondrial genome. We find some evidence for chrM occupancy for 50 nuclear TFs and two other proteins, with bZIP TFs emerging as most likely to be playing a role in mitochondria. However, we also observe that in cases where the same TF has been assayed with multiple antibodies and ChIP protocols, evidence for its chrM occupancy is not always reproducible. In the light of these findings, we discuss the evidential criteria for establishing chrM occupancy and reevaluate the overall compendium of putative mitochondrial-acting nuclear TFs.

Genome, Mitochondrial

OligoSeq: Rapid nanopore-sequencing of single-stranded oligonucleotides.

Nanopore-based DNA sequencing technology has achieved remarkable success in sequencing increasingly long DNA strands (e.g., over a million nucleotides long) for genomics research and biotechnology applications. However, the same level of progress has not been achieved for DNA oligonucleotides (usually &#x2264; 300 nucleotides long). Oligonucleotides play a crucial role in genome engineering efforts through oligo library generation and in DNA data storage, where they are used to encode computer information, such as binary (digital) data in DNA libraries. To enable these applications, accurate sequencing of oligonucleotides in a way that allows to assess for sequence variability, quality and length is essential. But sequencing solutions for oligonucleotides - particularly DNA primers for PCR, oligo DNA libraries used for mutagenesis or cDNA libraries used in gene expression analysis - remain inadequate. To address this gap, OligoSeq is presented as an innovative approach that integrates two complementary techniques: AmpliSeq (based on PCR) and RevSeq (based on reverse complementation with sequence-specific or random primers) to facilitate sequencing of single-stranded oligonucleotides using reference sequence anchor matches of more than &#x2265; 90% identity spanning from about 70% to 10% with AmpliSeq or RevSeq with random nonamers, respectively, and resolving the final reference sequence based on the most likely candidate from basecall frequencies, regardless of length and double-stranding method. OligoSeq can be integrated with nanopore sequencing technology pipelines and can be used as a reference for other sequencing platforms requiring double-stranded adapters, offering a practical and scalable alternative for standard quality control in single-stranded oligonucleotide synthesis. The use of nanopore technology, compatible with the double-stranding methods showcased, is shown to be the most cost-effective method for resolving original DNA sequences of different length and quality, and to assess its sequence variability, compared to other methods such as Illumina, PacBio or HPLC/MS.

Sequence Analysis, DNA

Virus-induced gene silencing as a tool for functional genomics in weeds: Challenges and future directions.

Virus-induced gene silencing (VIGS) has evolved from a conceptual demonstration of antiviral defense into a pivotal reverse-genetics platform for plant functional genomics. By exploiting engineered DNA- or RNA-based viral vectors, VIGS enables rapid, sequence-specific transcript knockdown through RNA-mediated degradation of target transcripts. Recent refinements in vector design, inoculation strategies, and viral species selection, such as TRV, BSMV, and FoMV, have expanded its application to previously recalcitrant plants, including major crops and emerging weed models. In weeds, functional genomics remains particularly challenging due to high genetic variability, limited genomic resources, and incompatibility with conventional viral vectors and transformation systems. In this context, VIGS provides a tractable approach to investigate genes associated with herbicide resistance, metabolic adaptation, and stress tolerance. Beyond weed biology, its application to studies of immune signaling, hormonal crosstalk, and secondary metabolism highlights VIGS as a versatile biotechnology for elucidating gene function and supporting next-generation strategies in plant improvement and integrated pest management.

Journal Article

Toxin-Antitoxin Systems of Staphylococcus aureus.

Toxin-antitoxin (TA) systems are small genetic elements found in the majority of prokaryotes. They encode toxin proteins that interfere with vital cellular functions and are counteracted by antitoxins. Dependent on the chemical nature of the antitoxins (protein or RNA) and how they control the activity of the toxin, TA systems are currently divided into six different types. Genes comprising the TA types I, II and III have been identified in Staphylococcus aureus. MazF, the toxin of the mazEF locus is a sequence-specific RNase that cleaves a number of transcripts, including those encoding pathogenicity factors. Two yefM-yoeB paralogs represent two independent, but auto-regulated TA systems that give rise to ribosome-dependent RNases. In addition, omega/epsilon/zeta constitutes a tripartite TA system that supposedly plays a role in the stabilization of resistance factors. The SprA1/SprA1AS and SprF1/SprG1 systems are post-transcriptionally regulated by RNA antitoxins and encode small membrane damaging proteins. TA systems controlled by interaction between toxin protein and antitoxin RNA have been identified in S. aureus in silico, but not yet experimentally proven. A closer inspection of possible links between TA systems and S. aureus pathophysiology will reveal, if these genetic loci may represent druggable targets. The modification of a staphylococcal TA toxin to a cyclopeptide antibiotic highlights the potential of TA systems as rather untapped sources of drug discovery.

Antitoxins