Search PubMedSearch

SEARCH · Search PubMed

Results for “Motif”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The ETS domain transcription factor Elk-1 contains a novel class of repression domain.

The ETS domain transcription factor Elk-1 serves as an integration point for different mitogen-activated protein (MAP) kinase pathways. Phosphorylation of Elk-1 by MAP kinases triggers its activation. However, while the activation process is well understood, its downregulation-inactivation is less well characterized. The ETS DNA-binding domain plays a role in the downregulation of Elk-dependent promoter activity following mitogenic activation by recruiting the mSin3A-HDAC complex. Here we have identified a novel evolutionarily conserved repression domain in Elk-1, termed the R motif, which serves to reduce the basal transcriptional activity of Elk-1 and dampen its response to mitogenic signals. This domain is highly potent and portable and can repress transcription in trans. The R motif is related to the CRD1 repression domain in p300 and can functionally replace this domain and confer p21(waf1/cip1) inducibility on p300. However, the R motif acts in a context-dependent manner and is not p21(waf1/cip1) responsive in Elk-1. Thus, the Elk-1 R motif and the p300 CRD1 motif represent a new class of repression domains that are regulated in a context-dependent manner.

Amino Acid Motifs

Distinct HLA Associations for Antibody Multireactivity With Citrulline-Containing Type II Collagen Epitopes Versus More Limited Antibody Reactivity With Citrulline-Containing IgG Epitopes in Rheumatoid Arthritis.

OBJECTIVE: Anticitrullinated protein antibodies (ACPAs) in rheumatoid arthritis (RA) can be promiscuous, with cross-reactive binding to many antigens containing short motifs, or private with little cross-reactivity. Also, ACPA reactivity patterns differ among patients with RA, including for motif-containing epitopes in important self-antigens like collagen and IgG (bound by RA-associated rheumatoid factors [RFs]), with limited understanding of the underlying mechanism. The objective of this study was to determine if HLA alleles associate with ACPA reactivity patterns. METHODS: For 100 ACPA+RF+ participants with RA, serum IgG binding was quantified by enzyme-linked immunosorbent assay to 10 citrulline-containing peptides derived from Type II collagen and IgG1 (nine with motifs), and HLA loci were genotyped. Also, antibody and serum multireactivity were evaluated. HLA alleles present differentially in RA participants with high versus low IgG binding to specific peptides, as well as with multireactivity versus limited reactivity were identified by Fisher's exact test. RESULTS: Serum IgG multireactivity for citrulline-glycine motif-containing collagen peptides was high, at least partially due to promiscuous antibodies. HLA-DQA1*01:02 was present in more participants with anticitrullinated collagen antibodies and multireactive sera. In contrast, serum multireactivity was low for IgG1-derived peptides due at least in part to more private antibodies. Shared epitope-containing HLA-DRB1*04:01 was present more frequently in participants with RA-associated RFs irrespective of the citrulline-serine motif and less frequently in participants with anticitrullinated collagen antibodies. Several HLA alleles associated with specific antibody reactivities. CONCLUSION: Different HLA alleles may contribute to the different reactivity patterns of promiscuous anticitrullinated collagen antibodies and more private RA-associated RFs.

Humans

Oriented binding of transcription factors to nucleosomes remodels chromatin at human promoters.

Transcription factors (TFs) can access nucleosomes via five distinct modes: gyre-spanning, periodic-binding, dyad-binding, and end-binding modes as well as an oriented binding mode, where the TF binding motif shows orientational preference relative to the nucleosome. Here, we report the first structure of an oriented TF:nucleosome complex, where two ELF2 proteins bind to a double motif located at superhelical location +4, unwinding four helical turns of DNA from the nucleosome. We further show that unlike previously described pioneer factors, ELF2 is able to occupy all of its unmethylated, high-affinity double motifs in vivo. Motifs of ELF2 and another oriented nucleosome binder, YY1, are highly enriched downstream of transcription start sites (TSSs) of highly expressed genes, with the motifs oriented in such a way that the TSS becomes accessible upon TF binding. Our results suggest that oriented binding may be generally important for high transcriptional activity.

Nucleosomes

Regulation and function of the HPV16 CircE7 RNA.

High-risk human papillomaviruses (HPV), including HPV16, produce circular RNA that encompasses the E7 oncogene (circE7). CircE7 can be detected in HPV16-positive cells and tumors, is preferentially localized to the cytoplasm, is N6-methyladenosine (m6A)-modified, and can be translated into the E7 oncoprotein. Here, we explored the regulation and function of circE7. Mutation of m6A motifs flanking the backsplice junction revealed a single m6A motif to be essential for circE7 formation. Mutation of this m6A motif promoted linear splicing of the E6*I splice site (226^409), suggesting that linear and circular E7 splicing are inversely regulated. Additionally, mutation of an IRES-like motif in circE7 significantly decreased E7 protein expression, without having significant effects on circE7 RNA levels. Knockdown of YTHDC1, but not other m6A-binding proteins, decreased both circE7 RNA and protein expression. BaseScope ISH was used to confirm the expression of circE7 in head and neck squamous cell carcinoma cell lines and tumors. Using both qRT-PCR and BaseScope ISH, we found that serum and amino acid starvation significantly increased circE7. Finally, we generated an HPV16 genome with mutations in the circE7 m6A motif (Mut2). Stable transduction of primary keratinocytes with Mut2 confirmed the loss of circE7 and increased expression of E6*I. The Mut2 HPV16 genome exhibited significantly decreased viral replication but an increased ability to transform primary keratinocytes. Our studies reveal that the precise regulation of circE7 and E6*I by m6A is critical for the ability of HPV16 to infect and transform keratinocytes.IMPORTANCEHigh-risk human papillomaviruses (HPVs), such as HPV16, must carefully control how much E6 and E7 proteins they make. This study shows that HPV16 toggles a single chemical tag on the viral RNA (an m⁶A mark) to control the production of early region RNAs, including a circular RNA called circE7. The same site coordinately regulates splicing of the E6*I isoform. CircE7 uses m⁶A-binding proteins to control its production and a specific sequence to promote its translation. It is present in HPV-positive cancers and can respond to nutrient starvation. Regulation of circE7 through this m6A site also impacted viral replication and transformation capacity, indicating that this regulatory mechanism is critical for HPV biology.

RNA splicing

Transcription Start Regions in PTU-intergenic regions drive cell cycle-dependent transcriptional activation events in Leishmania donovani.

Leishmania displays an unconventional mode of transcription, with long clusters of genes being transcribed polycistronically from Transcription Start Regions (TSRs), being processed into monocistronic units prior to translation. It has long been believed that transcription is constitutive: failure to identify consensus sequences across TSRs (except a GT-rich motif supporting transcription in Trypanosoma brucei) and absence of canonical eukaryotic transcription factors led to the conclusion that regulation is primarily post-transcriptional, with epigenetics playing a role in triggering transcription initiation. This study stems from our previous findings identifying a few genes to be activated in a cell cycle-dependent manner. Using nuclear run-on assays to analyze nascent transcripts of two chromosomes, chromosomes 2 and 14, we find that while most genes are constitutively transcribed, a subset of genes gets activated at specific cell cycle stages. Reporter assays reveal that this transcriptional activation is driven by the regions immediately upstream of the genes. Sequence analyses of these TSRs lying in polycistronic intergenic regions (PIRs) uncovered a 10-mer GT-rich motif, in synchrony with earlier findings in T. brucei identifying a GT-rich motif at bidirectional TSRs. We also identify a second 25-mer motif at these TSRs, and deletion analyses find this motif to be critical for regulating gene expression. The findings of this study reveal that transcriptional events in these unicellular parasites are more complex than believed thus far: not all transcriptional events are constitutive, polycistronic transcription is not the only mode of transcription, and cis-acting sequence elements regulate at least some transcriptional events in these parasites.IMPORTANCEEndemic to 90 countries, Leishmania parasites cause a spectrum of diseases called Leishmaniases. No vaccines for human use are available to date, and the drugs currently used to treat the disease are expensive, have toxic side effects, and have complex administration regimens, with emerging drug resistance compounding problems. Researchers continue to investigate Leishmania cellular processes, with the hope of uncovering new therapeutic target sites. Gene regulation in these parasites is unusual, being modulated by various mechanisms, including epigenetic modifications, gene dosage, and post-transcriptional processing. Transcription is typically polycistronic and constitutive, initiating from Transcription Start Regions (TSRs) lying upstream of the first gene in the polycistronic transcription unit (PTU). The work presented here reveals that a subset of genes is transcribed monocistronically in a cell cycle-dependent manner from Transcription Start Regions lying in the PTU-intergenic regions (PIRs), underscoring the complexities of gene regulation in these parasites.

Leishmania donovani

Crystal structures of Parechovirus A1 3Dpol reveal a mechanism of conformational stabilization in +ssRNA virus RNA-dependent RNA polymerase.

Parechovirus A1 (PeV A1) 3Dpol is an RNA-dependent RNA polymerase responsible for replication of the virus genome. We solved crystal structures of PeV A1 3Dpol structure in complex with GTP and in apo-state at 1.8-2.0 Å resolutions. In the 3Dpol-GTP complex, the conformation of the conserved motif B loop was stabilized by zinc ion coordination by cysteine residues. Apo-state structures of PeV A1 3Dpol showed significant conformational flexibility in the motif B loop, in the absence of zinc. While one of the conformational states of apo-3Dpol was similar to the 3Dpol-GTP complex structure, the alternative apo-3Dpol conformation showed a 4.3 Å movement of the motif B loop out of the active site cavity relative to the complex of 3Dpol with GTP. We propose that PeV A1 3Dpol activity is regulated by conformational stabilization of the motif B loop by zinc coordination.

Crystal structure

Genome-wide identification and expression profiling of HSD3B and SDR42E1 genes in the Pacific oyster (Crassostrea gigas): potential associations with gonadal development.

Sex steroids are lipid-soluble signaling molecules that regulate sex differentiation, reproductive development and physiological homeostasis in animals. 3β-Hydroxysteroid dehydrogenase/Δ5-Δ4 isomerase (3β-HSD) is a key steroidogenic enzyme, whereas SDR42E1, an extended short-chain dehydrogenase/reductase, has been implicated in sterol- and steroid-related metabolism. However, the composition, evolutionary relationships and expression patterns of the HSD3B- and SDR42E1-related genes in bivalve gonadal development remain poorly characterized. In this study, five PF01073-containing genes, comprising three CgHsd3b and two CgSdr42e1 genes, were identified in the Pacific oyster Crassostrea gigas. Phylogenetic analysis separated the proteins into HSD3B-related and SDR42E1-related groups, and gene-structure and motif analyses indicated subfamily-level divergence. All five proteins retained the SDR domain but differed in exon-intron structure and motif composition. Each contained the extended-SDR TGxxGxxG motif, whereas exact classical [ST]GxxxGxG and NNAG motifs were absent. Tyr- and Lys-equivalent residues were conserved, while the HSD3B1 Ser-equivalent position contained Thr in two C. gigas proteins and Ser in one. These features support their classification as extended-SDR proteins but do not establish enzymatic activity or substrate specificity. The three CgHsd3b genes were dispersed on one chromosome, whereas CgSdr42e1-1 and CgSdr42e1-2 were adjacent on another chromosome, suggesting a possible local duplication event for the CgSdr42e1 pair. Public RNA-seq data showed distinct tissue- and gonadal-stage expression patterns, with several genes displaying gonad-biased or female-stage-associated expression. Independent RT-qPCR profiling of the representative genes CgHsd3b-3 and CgSdr42e1-1 detected stage-dependent expression, although tissue rankings differed from those in the public RNA-seq datasets. These differences may reflect the use of independent biological samples, tissue composition, normalization procedures, and platform-specific measurements. Because enzymatic assays, metabolite measurements, cellular localization, and functional perturbation were not performed, the results identify candidate genes whose expression is associated with gonadal development rather than demonstrating regulatory roles. This study provides a comparative framework for future functional investigation of sterol- and steroid-related metabolism in bivalves.

Animals

TFinder: A Python Web Tool for Predicting Transcription Factor Binding Sites.

Transcription is a key cell process that consists of synthesizing several copies of RNA from a gene DNA sequence. This process is highly regulated and closely linked to the ability of transcription factors to bind specifically to DNA. TFinder is an easy-to-use Python web portal allowing the identification of Individual Motifs (IM) such as Transcription Factor Binding Sites (TFBS). Using the NCBI API, TFinder extracts either promoter or gene terminal regulatory regions, through a simple query of NCBI gene name or ID. It enables simultaneous analysis across five different species for an unlimited number of genes. TFinder searches for Individual Motifs in different formats, including IUPAC codes and JASPAR entries. Moreover, TFinder also allows de novo generations of a Position Weight Matrix (PWM) and the use of already established PWM. Finally, the data are provided in a tabular and a graph format showing the relevance and the P-value of the Individual Motifs found as well as their location relative to the Transcription Start Site (TSS) or the terminal region of the gene. The results are then sent by email to users facilitating the subsequent data analysis and sharing. TFinder is written in Python and freely available on GitHub under the MIT license: https://github.com/Jumitti/TFinder. It can be accessed as a web application implemented in Streamlit at https://tfinder-ipmc.streamlit.app. Resources are available on Streamlit "Resources" tab. TFINDER strength is that it relies on an all-in-one intuitive tool allowing users inexperienced with bioinformatics tools to retrieve gene regulatory regions sequences in multiple species and to search for individual motifs in a huge number of genes.

Transcription Factors

Role of RNA G-Quadruplexes in the Japanese Encephalitis Virus Genome and Their Recognition as Prospective Antiviral Targets.

G-quadruplexes (GQs) have been primarily studied in the context of cancer and neurodegenerative pathologies. However, recent research has shifted focus to their existence and functional roles in viral genomes, revealing GQ-regulated key pathways in various human pathogenic viruses. While GQ structures have been reported in the genomes of emerging and re-emerging viruses, RNA viruses have been understudied compared to DNA viruses, including notable examples such as human immunodeficiency virus-1, hepatitis C virus, Ebola virus, Nipah virus, Zika virus, and SARS-CoV-2. The flavivirus family, comprising the Japanese encephalitis virus (JEV), poses a significant global threat due to recurring outbreaks yet lacks approved antivirals. In this study, we identified and characterized eight putative G-quadruplex-forming motifs within essential genes involved in genome replication, assembly, and internalization in the host cell, conserved across different JEV isolates. The formation and stability of these motifs were validated through a multitude of biophysical and cell-based assays. The interaction and binding affinity of these motifs with the known GQ-binding ligand BRACO-19 were supported by biophysical assays, confirming the capability of these motifs to form GQ structures. Notably, BRACO-19 also exerted antiviral properties through reduction of viral replication and infectious virus titers as well as inhibition of viral protein expression, as evaluated by the cell-based assays. This comprehensive molecular characterization of G-quadruplex structures within the JEV genome highlights their potential as promising antiviral targets for intervention strategies against JEV infection through GQ-specific ligands.

G-Quadruplexes

Population-scale disease-associated tandem repeat analysis reveals locus and ancestry-specific insights.

Tandem repeat (TR) expansions, including short TRs (motifs ≤6 bp) and variable number TRs (motifs >6 bp), underlie many monogenic disorders, with variable length and sequence influencing pathogenicity, penetrance, severity, and onset. Accurate genotype-phenotype correlation and disease prevalence estimation require characterization beyond repeat length. Here we present a population-scale analysis of 66 disease-associated TR loci using long-read assemblies from 2530 diverse haplotypes from 1265 unaffected donors. Integrating repeat length, motif composition, local ancestry, linkage disequilibrium, and phylogenetic analyses, we reveal extensive locus-, population-, and allele-specific variation shaping disease risk. Up to 8.5% of individuals carry expansions above established pathogenic thresholds, many containing interrupting motifs or sequence structures that attenuate pathogenicity. After excluding alleles from loci with uncertain disease association, non-pathogenic interrupted expansions, and carrier states inconsistent with inheritance patterns, ~4% carried expansions predicted to confer disease risk, largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses uncover population-specific TR architectures contributing to epidemiological disparities in repeat expansion disorders. Phylogenetic analyses identify conserved ancestral alleles and loci with recent instability. We describe variable linkage disequilibrium patterns and recombination signatures around specific disease-associated TR loci. Our findings emphasize integrating sequence, ancestry, and evolutionary context to understand the complex landscape of disease-associated TRs.

Humans

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

An interpretable deep learning framework uncovers features governing CRISPR-Cas9 genome-editing efficiency.

MOTIVATION: CRISPR-Cas9 genome-editing efficiency is strongly influenced by the sequence composition and positional context of single-guide RNAs (sgRNAs). Although numerous deep learning-based models have been developed to predict Cas9 efficiency from sgRNA sequences, most operate as black boxes, offering limited insight into the sequence determinants underlying Cas9 activity. In addition, previous studies often overlook how the positional context of sequence motifs within sgRNAs influences their effects on Cas9 binding or cleavage. RESULTS: We introduce DeepCC9, an interpretable machine learning framework that combines explicit sequence feature extraction with a residual block-based deep architecture to improve interpretability and identify composition- and position-based motifs governing Cas9 genome-editing efficiency. We applied this method to multiple Cas9 variant datasets, achieving superior predictive performance compared with existing methods while enabling direct interpretation of sequence motifs and their positional effects. Our analysis uncovered 74 sequence motifs enriched or depleted at specific positions within sgRNAs and strongly associated with Cas9 efficiency, providing mechanistic insight into sequence features that influence guide performance. Together, these results establish DeepCC9 as a generalizable and interpretable framework for modeling sequence-function relationships and advancing the understanding of the sequence determinants underlying CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The authors have implemented their algorithm in the Python programming language (version 3.X), which is accessible using (https://zenodo.org/records/20073890).

Deep Learning

Functional genomic analysis of non-canonical DNA regulatory elements of the aryl hydrocarbon receptor.

The aryl hydrocarbon receptor (AHR) is a ligand-dependent transcription factor activated by environmental toxicants like halogenated and polycyclic aromatic hydrocarbons, which then binds to DNA and regulates gene expression. AHR is implicated in numerous physiological processes, including liver and immune function, cell cycle control, oncogenesis, and metabolism. Traditionally, AHR binds a consensus DNA sequence (GCGTG), the xenobiotic response element (XRE), recruits coregulators, and modulates gene expression. Yet, recent evidence suggests AHR can also regulate gene expression via a non-consensus sequence (GGGA), termed the non-consensus XRE (NC-XRE). The prevalence and functional significance of NC-XRE motifs in the genome have remained unclear. While ChIP and reporter studies hinted at AHR-NC-XRE interactions, direct evidence for transcriptional regulation in a native context was lacking. In this study, we analyzed AHR binding to NC-XRE sequences genome-wide in mouse liver, integrating ChIP-seq and RNA-seq data to identify candidate AHR target genes containing NC-XRE motifs in their regulatory regions. We found NC-XRE motifs in 82% of AHR-bound DNA, significantly enriched compared to random regions, and present in promoters and enhancers of AHR targets. Functional genomics on the Serpine1 gene revealed that deleting NC-XRE motifs reduced TCDD-induced Serpine1 upregulation, demonstrating direct regulation. These findings provide the first direct evidence for AHR-mediated regulation via NC-XRE in a natural genomic context, advancing our understanding of AHR-bound DNA and its impact on gene expression and physiological relevance.

Journal Article

GHT-SELEX demonstrates unexpectedly high intrinsic sequence specificity and complex DNA binding of many human transcription factors.

A long-standing challenge in human regulatory genomics is that transcription factor (TF) DNA-binding motifs are short and degenerate, while the genome is large. Motif scans therefore produce many false-positive binding site predictions. By surveying 179 TFs across 25 families using >1,500 cyclic in vitro selection experiments with fragmented, naked, and unmodified genomic DNA - a method we term GHT-SELEX (Genomic HT-SELEX) - we find that many human TFs possess much higher sequence specificity than anticipated. Moreover, genomic binding regions from GHT-SELEX are often surprisingly similar to those obtained in vivo (i.e. ChIP-seq peaks). We find that comparable specificity can also be obtained from motif scans, but performance is highly dependent on derivation and use of the motifs, including accounting for multiple local matches in the scans. We also observe alternative engagement of multiple DNA-binding domains within the same protein: long C2H2 zinc finger proteins often utilize modular DNA recognition, engaging different subsets of their DNA binding domain (DBD) arrays to recognize multiple types of distinct target sites, frequently evolving via internal duplication and divergence of one or more DBDs. Thus, contrary to conventional wisdom, it is common for TFs to possess sufficient intrinsic specificity to independently delineate cellular targets.

C2H2

Lineage structure and penicillin-binding protein variability in clinical Streptococcus pneumoniae isolates from Southwest China exhibiting reduced susceptibility to penicillin.

BACKGROUND: Reduced susceptibility to penicillin in Streptococcus pneumoniae is mediated primarily by alterations in penicillin-binding proteins (PBPs) and often coexists with multidrug resistance within successful lineages. The region-specific genomic characterization of clinically relevant pneumococci with reduced penicillin susceptibility in Southwest China remains limited. METHODS: We performed whole-genome sequencing of 204 clinical S. pneumoniae isolates collected from five institutions in Southwest China (2018-2022) that met our operational screening definition of reduced susceptibility to penicillin (PEN MIC ≥0.12 μg/mL). Molecular serotypes, MLST types, and Global Pneumococcal Sequence Clusters (GPSCs) were assigned; virulence and antimicrobial resistance determinants were profiled; and a core genome phylogeny was reconstructed with international contextualization through the use of PubMLST genomes meeting the same MIC criterion. Amino acid variability in PBP1a/PBP2b/PBP2x was quantified using TIGR4 numbering, and highly variable noncatalytic residues located within 15 Å of catalytic motifs were prioritized via structure-guided screening. RESULTS: The isolates showed a high burden of resistance to non-β-lactam antibiotics (erythromycin, 98.5%; tetracycline, 82.8%; trimethoprim-sulfamethoxazole, 64.7%), while fluoroquinolone susceptibility was largely preserved (≥97%), and vancomycin/linezolid resistance was not detected. Twenty-seven serotypes were identified, among which 19F (23.5%) and 19A (14.2%) were dominant, and the estimated PCV13 coverage was 69.6%. GPSC1 was the dominant lineage (36.8%), and the lineage composition among our isolates differed markedly from those in the PubMLST-USA and PubMLST-Thailand subsets. Virulence and resistance gene carriage differed markedly between GPSC1 and non-GPSC1 isolates, with enrichment of pilus operons, mef(A)/msr(D), and folA/folP in GPSC1. PBP variations were clustered in transpeptidase domains and motif-adjacent regions while essential catalytic residues were conserved; with the structure-guided filter, 12, 11, and 11 motif-proximal noncatalytic candidate sites were prioritized in PBP1a, PBP2b, and PBP2x, respectively. CONCLUSION: Clinical S. pneumoniae isolates with reduced penicillin susceptibility collected in Southwest China demonstrated resistance and accessory gene profiles that were strongly structured by a GPSC-defined lineage background. Our site-resolved, structure-guided PBP analysis provides a regional PBP variability landscape and a compact set of recurrent motif-proximal candidate substitutions to support surveillance and downstream functional validation.

Streptococcus pneumoniae

An expanded realm of anti-CRISPR-associated proteins and regulatory mechanisms.

Many bacteriophages encode anti-CRISPR (Acr) proteins that inhibit bacterial CRISPR-Cas immune systems. Rapid acr gene expression upon phage entry enables CRISPR-Cas neutralization but can impact phage fitness if unregulated. Therefore, Acr production is often controlled by distinct families of co-encoded anti-CRISPR-associated (Aca) proteins, which are usually helix-turn-helix (HTH) regulators that bind DNA within acr-aca operon promoters. Previously, we demonstrated that the Aca2 family additionally represses Acr production translationally by binding structured RNA motifs within the 5' untranslated region (UTR) of the acr-aca mRNA. Here, through systematic bioinformatic analyses, we provide evidence of structured RNA motifs in the 5' UTRs of operons encoding members of other Aca families and show that Aca1 also specifically binds its cognate RNA motif. Additionally, many Aca proteins are predicted to regulate not only their own but also adjacent operons with potential anti-defence genes. Indeed, we show that Aca14, newly identified in this study, represses two predicted anti-defence operons. Aca14 is a ribbon-helix-helix domain protein, revealing regulatory diversity beyond the canonical HTH Aca family members. Collectively, our findings expand our understanding of acr regulation in mobile genetic elements and reveal novel mechanisms by which phages fine-tune anti-defence gene expression.

5' Untranslated Regions

TRIM63 Overexpression in FISH-Negative MiTF Family Altered Renal Cell Carcinoma (MiTF RCC).

TFE3 and TFEB break-apart fluorescent in situ hybridization (FISH) assays are the "gold standard" for diagnostic confirmation of microphthalmia-associated transcription factor (MiTF) family-altered renal cell carcinoma (MiTF RCC), which includes TFE3-rearranged RCC and TFEB-altered RCC. However, FISH assays, for multiple reasons, may lead to equivocal or false-negative results, especially in cryptic fusions resulting from intrachromosomal inversions involving 5' partner genes, such as non-POU domain-containing octamer-binding protein (NONO); GRIPI-associated protein 1 (GRIPAP1); RNA-binding motif protein, X chromosome (RBMX); and RNA-binding motif protein 10 (RBM10). When FISH results are negative in cases with strong morphological suspicion of the listed tumor entities, pathologists may recommend targeted RT-PCR or panel-based RNA fusion sequencing for diagnostic confirmation. Our recent RNA in situ hybridization (RNA ISH)-based study demonstrated RNA expression of the tripartite motif containing 63 (TRIM63) to be highly enriched in TFE3-rearranged RCC and TFEB-altered RCC, including 2 FISH false-negative RCC cases harboring RBM10::TFE3 fusion. Based on these observations, we hypothesized that TRIM63 positivity could aid in diagnosing cases that are negative by conventional FISH assay but remain morphologically suspicious, representing an unmet clinical need in this area. We collected 20 RCC cases with morphological suspicion (with equivocal/indeterminate immunohistochemistry panel) of MiTF RCC, which were TRIM63 positive, negative/equivocal for TFE3/TFEB gene rearrangement by FISH, and underwent next-generation sequencing (NGS). On NGS correlation, 14 of 20 (70%) FISH-negative TRIM63-positive tumors harbored an MiTF gene rearrangement. In the remaining 6 cases, we were unable to fully ascertain the MiTF rearrangement status due to the inherent limitation of the NGS panel utilized. The cases with MiTF gene rearrangement include TFE3 rearrangement in 60% (12/20) and TFEB low-level copy gains (with an additional missense mutation in 1 case) in 10% (2/20) of samples. RBM10:TFE3 fusion was seen in 67% (8/12) of TFE3-rearranged RCC in this cohort. TRIM63 RNA ISH assay could aid in identifying cases that harbor TFE3 or TFEB rearrangement associated with false-negative or equivocal TFE3/TFEB FISH results, especially those involving gene fusions with a paracentric Xp11 inversion. Overall, employment of TRIM63 RNA ISH coupled with TFE3/TFEB FISH assays and follow-up genomic interrogation enhanced diagnostic accuracy for patients with MiTF RCC.

Carcinoma, Renal Cell

Single nucleotide polymorphism-based validation of exonic splicing enhancers.

Because deleterious alleles arising from mutation are filtered by natural selection, mutations that create such alleles will be underrepresented in the set of common genetic variation existing in a population at any given time. Here, we describe an approach based on this idea called VERIFY (variant elimination reinforces functionality), which can be used to assess the extent of natural selection acting on an oligonucleotide motif or set of motifs predicted to have biological activity. As an application of this approach, we analyzed a set of 238 hexanucleotides previously predicted to have exonic splicing enhancer (ESE) activity in human exons using the relative enhancer and silencer classification by unanimous enrichment (RESCUE)-ESE method. Aligning the single nucleotide polymorphisms (SNPs) from the public human SNP database to the chimpanzee genome allowed inference of the direction of the mutations that created present-day SNPs. Analyzing the set of SNPs that overlap RESCUE-ESE hexamers, we conclude that nearly one-fifth of the mutations that disrupt predicted ESEs have been eliminated by natural selection (odds ratio = 0.82 +/- 0.05). This selection is strongest for the predicted ESEs that are located near splice sites. Our results demonstrate a novel approach for quantifying the extent of natural selection acting on candidate functional motifs and also suggest certain features of mutations/SNPs, such as proximity to the splice site and disruption or alteration of predicted ESEs, that should be useful in identifying variants that might cause a biological phenotype.

Alleles