Search PubMedSearch

SEARCH · Search PubMed

Results for “Motif”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software

Motif-Centered Analyses Reveal Universal and Tissue-Specific Mutagenic Mechanisms Operating in the Human Body.

Somatic mutations are inevitable in human genomes and can lead to tumorigenesis, yet baseline mutagenesis in non-cancerous normal cells remain poorly understood. Here, we analyzed the mutation profiles of 11,949 normal samples across 25 tissues obtained from whole-genome and whole-exome sequencing datasets. We applied stringent statistical hypothesis for detecting enrichment and enrichment-adjusted Minimal Estimate of Mutation Load (MEML) in trinucleotide motifs preferred by known mutagenic processes. We found several cancer-associated mutational motifs in cancer-free tissues. Samples enriched with C→T mutations in nCg motif associated with clock-like spontaneous meCpG deamination were detected across all tissues. We revealed another clock-like motif, T→C substitutions in aTn motif associated with exposure to small epoxides and other SN2 electrophiles, in several tissues. Donors with several non-cancerous diseases showed significantly higher, age-independent, and concordant accumulation of aTn and nCg motifs compared to healthy donors. Motifs associated with chemical exposures showed sporadic, tissue and disease-specific mutagenesis. APOBEC-induced C→T and C→G mutations in tCw motif were enriched in bladder, lung, small intestine, liver, and breast with preference for APOBEC3A-like mutagenesis in most. Together, our analyses elucidated several ongoing mutagenic processes in normal human tissues and provided a robust analytical framework for identifying mutagenic sources from somatic mutation catalogues.

Journal Article

Immunopeptidomics-driven MHC class II peptide-binding motif discovery for 2 common canine DR alleles.

Despite the central role of major histocompatibility complex (MHC) class II in adaptive immunity, peptide-binding motifs have yet to be characterized for any canine MHC class II allele. Here, we report the first immunopeptidomics-derived binding motifs for DLA-DRB1*015:01 (DLA-DR15) and DLA-DRB1*012:01 (DLA-DR12), 2 alleles overrepresented in breeds predisposed to immune-mediated diseases. Because dogs co-express DLA-DR and DLA-DQ, the MHC class II Ab clone YKIX334.2 was validated to be DLA-DR-specific, enabling allele-selective immunoaffinity purification of DLA-DR molecules from homozygous DLA-DR15 and DLA-DR12 donor spleens. Mass spectrometry and GibbsCluster motif deconvolution of 838 DLA-DR15-associated and 644 DLA-DR12-associated peptides eluted from their respective peptide-binding grooves revealed distinct allele-specific binding motifs, with characterization of anchor residue preferences, peptide-length distributions, cross-species comparisons with human and murine MHC class II motifs, and source protein composition of the eluted self-peptidome. To evaluate the translational utility of these motifs, recombinant DLA-DR15 and DLA-DR12 molecules were used to screen rabies virus glycoprotein and nucleoprotein peptide libraries via fluorescence-based peptide competition assays, identifying high-affinity candidate binders for both alleles. Spearman rank correlation between immunopeptidomics-derived position-specific scoring matrix scores and peptide competition assay rankings demonstrated modest associations, consistent with these approaches capturing complementary dimensions of peptide-MHC class II interaction. Ultimately, these findings establish what we believe is the first allele-specific peptide-binding motif framework for canine MHC class II, providing a foundation for DLA-allele-informed CD4+ T-cell epitope discovery studies and Ag-specific immune response characterization in the dog.

Animals

Motif-centered analyses reveal universal and tissue-specific mutagenic mechanisms operating in the human body.

Somatic mutations are inevitable in human genomes and can lead to cancer initiation and tumor progression. Although many mutagenic processes have been linked to cancer, their activities in normal tissues before malignant transformation remain poorly characterized. Here, we analyzed the mutation profiles of 10,625 normal samples across 25 tissues obtained from whole-genome and whole-exome sequencing datasets. We applied stringent statistical hypothesis for detecting enrichment and enrichment-adjusted Minimal Estimate of Mutation Load in trinucleotide motifs preferred by known mutagenic processes. We found several cancer-associated mutational motifs in cancer-free tissues. Samples enriched with C→T mutations in nCg motif associated with clock-like spontaneous meCpG deamination were detected across all tissues. We also identified a second clock-like motif, T→C substitutions in aTn motif associated with exposure to small epoxides and other SN2 electrophiles, in several tissues. Motifs associated with other environmental and chemical mutagens showed sporadic and tissue-specific mutagenesis. APOBEC-induced C→T and C→G mutations in tCw motif were enriched in bladder, lung, small intestine, liver, and breast with preference for APOBEC3A-like mutagenesis in most tissues. Together, our analyses elucidated several cancer-associated mutagenic processes in normal tissues and provided a robust analytical framework for quantifying mutagenic activities from somatic mutation catalogs.

Humans

Characterising the motif composition and allele length distribution of ZFHX3 GGC repeat expansions in amyotrophic lateral sclerosis.

A pathogenic GGC repeat expansion in zinc finger homeobox 3 (ZFHX3), encoding a pure polyglycine (polyG) tract, causes spinocerebellar ataxia type 4 (SCA4). Intermediate expansions of other SCA loci have been implicated in amyotrophic lateral sclerosis (ALS), while repeat motif composition is recognised to influence pathogenicity in neurodegenerative diseases. Given the genetic pleiotropy between ALS and SCA, we evaluated whether ZFHX3 GGC expansions are associated with ALS and characterised repeat motif composition. ZFHX3 GGC repeat sizes were genotyped using ExpansionHunter in short-read whole-genome sequencing data from ALS cases and healthy controls of European ancestry. Repeat sizes were visually inspected using REViewer, and motif configurations were manually derived from a subset. Receiver operating characteristic analysis and Youden's J statistic identified a candidate repeat size threshold. Logistic regression tested associations of repeat length and motif composition with ALS, while regression models assessed clinical phenotypes. Across 5785 ALS cases and 7982 controls, no association was observed between ZFHX3 expansions and ALS risk. Longer alleles showed a nominal association with later disease onset, however this did not remain significant after Bonferroni correction. Among 802 ALS cases and 800 controls, 50 distinct motif compositions were identified, including 11 encoding pure polyG tracts characteristic of pathogenic SCA4 expansions; none were associated with ALS. Although no association with ALS was observed, this study established the dynamic nature of ZFHX3 repeat motif composition and configuration. Variation within and between repeat sizes, including pure polyG repeats, supports consideration of motif composition alongside allele length when evaluating neurodegenerative disease risk.

Journal Article

RETRACTED: Investigation of the effect of UV-B light on Arabidopsis MYB4 (AtMYB4) transcription factor stability and detection of a putative MYB4-binding motif in the promoter proximal region of AtMYB4.

Here, we have investigated the possible effect of UV-B light on the folding/unfolding properties and stability of Arabidopsis thaliana MYB4 (AtMYB4) transcription factor in vitro by using biophysical approaches. Urea-induced equilibrium unfolding analyses have shown relatively higher stability of the wild-type recombinant AtMYB4 protein than the N-terminal deletion forms after UV-B exposure. However, as compared to wild-type form, AtMYB4Δ2 protein, lacking both the two N-terminal MYB domains, showed appreciable alteration in the secondary structure following UV-B exposure. UV-B irradiated AtMYB4Δ2 also displayed higher propensity of aggregation in light scattering experiments, indicating importance of the N-terminal modules in regulating the stability of AtMYB4 under UV-B stress. DNA binding assays have indicated specific binding activity of AtMYB4 to a putative MYB4 binding motif located about 212 bp upstream relative to transcription start site of AtMYB4 gene promoter, while relatively weak DNA binding activity was detected for another putative MYB4 motif located at -908 bp in AtMYB4 promoter. Gel shift and fluorescence anisotropy studies have shown increased binding affinity of UV-B exposed AtMYB4 to the promoter proximal MYB4 motif. ChIP assay has revealed binding of AtMYB4 to the promoter proximal (-212 position) MYB4 motif (ACCAAAC) in vivo. Docking experiments further revealed mechanistic detail of AtMYB4 interaction with the putative binding motifs. Overall, our results have indicated that the N-terminal 62-116 amino acid residues constituting the second MYB domain plays an important role in maintaining the stability of the C-terminal region and the overall stability of the protein, while a promoter proximal MYB-motif in AtMYB4 promoter may involve in the regulation of its own expression under UV-B light.

Arabidopsis

Characterization of a Ku-binding motif in the C-terminal region of RAG2.

We applied an unsupervised interactome analysis with the RAG2 C-terminal region (R2CT) in v-abl pro-B cells undergoing V(D)J recombination. Mass-spectrometry analyses showed that Ku70 and Ku80 were among the top 10 hits. To further strengthen these observations, we performed Proximity Ligation Assay (PLA) and characterize the existence of a GFP-R2CT-Ku complex formation in cellulo. The interaction of several partners with Ku70/80 (Ku) through Ku-binding motifs (KBMs) in their sequences governs their enrolment in NHEJ repair complexes. Through sequence analysis, we identified a KBM within R2CT (R-KBM, amino acids 589-527). We confirmed by calorimetry a specific micromolar interaction between this RAG2 region and Ku70/80/DNA complex. The RAG2 motif KBM can be subdivided in two conserved parts that have no interaction individually. AlphaFold2 prediction coupled with molecular dynamic simulations indicate that the C-terminal part of the RAG2 motif interacts with Ku80 on the same site than the NHEJ factor XLF. These in silico analyses indicated that the N-terminal part of the RAG2 motif interacts with DNA adjacent to Ku with a major role of the K503 residue in agreement with disruption of the interaction observed with the K503E mutant. This study further extends the large ensemble of proteins recruited at DSBs by KBM motifs and substantiates the model of a tight coupling between DNA breakage and repair during V(D)J recombination, mediated by the Ku-RAG2 C-terminus interaction.

Ku Autoantigen

Heat-responsive ONSEN long terminal repeats integrate heat shock factor motifs, DNA methylation and natural sequence variation in Arabidopsis.

ONSEN is a heat-activated Ty1/copia retrotransposon in Arabidopsis thaliana controlled by heat shock factors (HSFs) and epigenetic silencing. Heat shock element (HSE)-like sequences in ONSEN long terminal repeats (LTRs) contribute to heat responsiveness, but relationships among sequence architecture, basal DNA methylation and natural variation remain unclear. We combined transcription-factor motif prediction, transposable-element comparisons, methylome and RNA sequencing (RNA-seq) data, and Arabidopsis genome assemblies. In silico disruption of five HSE cores eliminated HSF-family motif compatibility in the selected design and all 5119 exact-guanine-cytosine (GC) alternatives. Across 16 curated Columbia-0 terminal windows, ONSEN contained 33-49 non-redundant HSF motif-coordinate placements per 800 bp window and was strongly enriched relative to 1930 non-ONSEN transposable elements across score thresholds and continuous metrics. Direct comparison with 779 non-ONSEN LTR retrotransposons showed selectively elevated basal CHH methylation (where H = A, C or T) at ONSEN termini. Genome-wide RNA-seq analysis revealed broad heat-responsive gene and transposable-element changes, including strong ONSEN induction, whereas candidate-window analysis distinguished ONSEN from most HSF-rich non-ONSEN outliers. ONSEN-like variants across eight accessions generally retained HSF-compatible motifs while altering predicted DNA binding with one finger-family motif composition. Together, these findings define ONSEN terminal regions as HSF-rich regulatory sequences that retain heat-responsive potential within a methylated chromatin context and identify candidates for functional analysis.

DNA Methylation

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Discovering human transcription factor physical interactions with genetic variants, novel DNA motifs, and repetitive elements using enhanced yeast one-hybrid assays.

Identifying transcription factor (TF) binding to noncoding variants, uncharacterized DNA motifs, and repetitive genomic elements has been technically and computationally challenging. Current experimental methods, such as chromatin immunoprecipitation, generally test one TF at a time, and computational motif algorithms often lead to false-positive and -negative predictions. To address these limitations, we developed an experimental approach based on enhanced yeast one-hybrid assays. The first variation of this approach interrogates the binding of >1000 human TFs to repetitive DNA elements, while the second evaluates TF binding to single nucleotide variants, short insertions and deletions (indels), and novel DNA motifs. Using this approach, we detected the binding of 75 TFs, including several nuclear hormone receptors and ETS factors, to the highly repetitive Alu elements. Further, we identified cancer-associated changes in TF binding, including gain of interactions involving ETS TFs and loss of interactions involving KLF TFs to different mutations in the TERT promoter, and gain of a MYB interaction with an 18-bp indel in the TAL1 superenhancer. Additionally, we identified TFs that bind to three uncharacterized DNA motifs identified in DNase footprinting assays. We anticipate that these enhanced yeast one-hybrid approaches will expand our capabilities to study genetic variation and undercharacterized genomic regions.

Algorithms

A Sequence Motif Enables Widespread Use of Non-Canonical Redox Cofactors in Natural Enzymes.

Non-canonical redox cofactors (NRCs) are promising alternatives to nicotinamide adenine dinucleotide (phosphate) (NAD(P)+) for biomanufacturing due to low cost and exquisite electron delivery control, yet their adoption is limited by the scarcity of compatible enzymes. Here, we screened the aldehyde dehydrogenase (ALDH) protein family and identified a conserved RH/QxxR sequence motif that enables widespread NRC activity among natural enzymes. Bos taurus ALDH3a1 and Pseudanabaena biceps ALDH exhibit unprecedented turnover with nicotinamide mononucleotide (NMN+), with kcat values matching or exceeding that of NAD+ and surpassing most engineered NRC-active enzymes by 10 to 105-fold, based on the relative NRC to native activity. Structural and dynamic analyses reveal this motif reinforces cofactor positioning and pre-organizes the active site without dependence on the adenosine monophosphate moiety of NAD+. When introduced into diverse ALDH scaffolds, the RH/QxxR motif enhances NMN+ activity up to 60-fold. In addition to NMN+, this motif also supports activity across multiple non-nucleotide, simple synthetic NRCs such as 1-(2-carbamoylmethyl)nicotinamide (AmNA+). These findings elucidate Nature's solution to the engineering challenge of obtaining NRC-active enzymes and offers a blueprint to mine latent evolutionary plasticity in natural enzymes that serve as superior engineering starting points.

Active site pre-organization

Disruption of GxxxG motifs in pATOM36 impairs biogenesis of the mitochondrial protein translocase of the outer membrane in Trypanosoma brucei.

Mitochondrial biogenesis requires efficient import of cytosolically produced proteins and correct segregation of the mitochondrial genome during cytokinesis. In Trypanosoma brucei, a parasitic protozoan with a single mitochondrion harboring a single-unit mitochondrial genome, protein import across the outer membrane is mediated by the ATOM complex. An important, yet poorly understood role is played by the integral membrane protein pATOM36 of the outer mitochondrial membrane, which is essential for both ATOM complex assembly and mitochondrial DNA segregation. Here, we combined in vivo functional mutational analysis and structural modeling to investigate the function of pATOM36. AlphaFold3-based models predict five highly tilted helices forming a funnel-shaped cavity open toward the cytoplasm, reminiscent of membrane protein insertases. In the model, the protein is sealed towards the mitochondrial intermembrane space by tight helix packing, with conserved GxxxG motifs potentially facilitating these helix-helix interactions. Progressive replacement of these glycines by isoleucines does not affect protein production or correct localization but leads to defective ATOM complex biogenesis and arrest of growth, while mitochondrial DNA segregation is largely unaffected. Based on the predicted structure, these effects can be rationalized by hydrophobic bulking that interferes with associated electrostatic interactions. This hypothesis is supported by experimental mutational analysis of the respective electrostatic interactions in the presence of native GxxxG motifs. Together, our data support the hypothesis that pATOM36 functions as an outer mitochondrial insertase and arose by convergent evolution. The GxxxG motifs, also found in unrelated yeast and human outer membrane insertases, are crucial for protein activity.

Trypanosoma brucei brucei

Harnessing deep learning for proteome-scale detection of amyloid signaling motifs.

MOTIVATION: Amyloid signaling sequences adopt the cross-β fold that is capable of self-replication in the templating process. Propagation of the amyloid fold from the receptor to the effector protein is used for signal transduction in the immune response pathways in animals, fungi, and bacteria. So far, a dozen of families of amyloid signaling motifs (ASMs) have been classified. Unfortunately, due to the wide variety of ASMs it is difficult to identify them in large protein databases available, which limits the possibility of conducting experimental studies. To date, various deep learning (DL) models have been applied across a range of protein-related tasks, including domain family classification and the prediction of protein structure and protein-protein interactions. RESULTS: In this study, we develop tailor-made bidirectional LSTM and BERT-based architectures to model ASM, and compare their performance against a state-of-the-art machine learning grammatical model. Our research is focused on developing a discriminative model of generalized ASMs, capable of detecting ASMs in large datasets. The DL-based models are trained on a diverse set of motif families and a global negative set, and used to identify ASMs from remotely related families. We analyze how both models represent the data and demonstrate that the DL-based approaches effectively detect ASMs, including novel motifs, even at the genome scale. AVAILABILITY AND IMPLEMENTATION: The models are provided as a Python package, asmscan-bilstm, and a Docker image at https://github.com/chrispysz/asmscan-proteinbert-run. The source code can be accessed at https://github.com/jakub-galazka/asmscan-bilstm and https://github.com/chrispysz/asmscan-proteinbert. Data and results are at https://github.com/wdyrka-pwr/ASMscan.

Deep Learning

Disruption of a six-nucleotide miRNA motif improves PKD1 dosage and ameliorates polycystic kidney disease.

Disrupting microRNA interactions to restore protein expression from haploinsufficient genes offers a promising precision-therapy strategy for monogenic disorders. PKD1 heterozygosity underlies autosomal dominant polycystic kidney disease (ADPKD), a disorder affecting nearly 12 million people worldwide, where reduced PKD1 dosage drives progressive cyst formation and kidney failure. We previously identified a 55-bp cis-repressive element in the PKD1 3'UTR. Here, we define a six-nucleotide miR-17 seed match within this element that is sufficient to reproduce PKD1 repression. In vivo base substitution of this motif stabilizes Pkd1 messenger RNA and increases polycystin-1 (PC1) protein levels, producing a robust reduction in cyst growth and preservation of kidney function in mouse models. To therapeutically recapitulate this effect, we developed a steric-blocking oligonucleotide that occludes the motif, stabilizes PKD1 transcript levels, increases PC1 expression, and mitigates cyst-pathogenic events in both murine and patient-derived ADPKD cells. Together, these findings establish a minimal, targetable cis-regulatory motif and provide proof of concept for oligonucleotide-mediated PKD1 derepression, while offering a potentially generalizable strategy to restore other haploinsufficient genes.

Animals

Ribosome dynamics at the conserved PGP motif governs 2A peptide-bond-skipping efficiency.

Viral 2A oligopeptides drive an unusual ribosome recoding event in which peptide-bond formation fails at a conserved PG↓P motif, producing two discrete proteins without canonical termination. Despite decades of study, the molecular basis of 2A-mediated peptide-bond skipping remains poorly understood. Here, we combine quantitative 2A reporters with high-resolution ribosome profiling to interrogate ribosome dynamics at the core 2A sequences. We identify a pausing event at the terminal proline codon of the PGP motif that functions as a kinetic decision point: ribosome dwell time at this site inversely correlates with skipping efficiency. Increasing nascent chain flexibility by inserting linkers immediately upstream of the 2A sequence reduces ribosome occupancy at the terminal proline codon and enhances peptide-bond skipping. Strikingly, amino acid repeats positioned distally upstream also modulate 2A activity, indicating long-range coupling between nascent chain properties outside of the ribosome and the peptidyl transferase center inside the ribosome. In particular, hydrophobic residues potently suppress skipping, an effect that can be rescued by extending flexible segments within the peptide exit tunnel. Together, our findings support a model in which nascent chain features-beyond the core 2A motif-dynamically tune ribosomal recoding efficiency through co-translational feedback into the catalytic center.

Ribosomes

Dense RNA motif modifications enable robust in vivo prime editing and enhance efficiencies of diverse editing systems.

Prime editing holds promise for therapeutic applications. However, viral delivery of the prime editor presents challenges for clinical translation due to concerns regarding long-term expression. Meanwhile, systemic delivery using non-viral vectors has been limited by low efficiency, the need for repeated injections and reliance on doses that exceed clinically translatable levels. Here we develop engineered prime editing guide RNAs (pegRNAs) with densely modified RNA motifs and demonstrate their application for efficient in vivo prime editing. By systemically delivering the prime editor in RNA format via a single injection of lipid nanoparticles, we achieved nearly 70% editing efficiency in the bulk mouse liver, indicating successful editing of the majority of hepatocytes. Notably, a single injection at a clinically translatable lipid nanoparticle dose was sufficient to suppress target protein expression in vivo, resulting in a near 80-fold increase in editing efficiency compared with conventional end-modified pegRNAs. Furthermore, incorporating densely modified RNA motifs, including the widely used MS2 motif, proved broadly applicable across various RNA sequences and split RNA-guided genome editing platforms, resulting in up to an 11-fold increase in base editing efficiency. These findings present a generalizable approach for enhancing the therapeutic potential of prime editing and expanding the utility of RNA-based therapeutics.

Journal Article