Search PubMedSearch

PubMed · 42561589

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

Abstract

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Explore related subjects

Keep this discovery

BibTeXRIS

Takumi Suto, Md Harun-Or-Roshid, Hiroyuki Kurata. 2026-08-01. Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.. https://doi.org/10.1016/j.cmpb.2026.109581

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Whole-transcriptome RNA sequencing and ceRNA network analyses provide novel insights into the antibacterial immune response of Hippocampus abdominalis against Vibrio harveyi.

Long non-coding RNAs (lncRNAs) stand as newly-arisen molecular types that exert regulatory effects, able to operate as competitive endogenous RNAs (ceRNAs) to engage microRNAs (miRNAs) in interaction, resulting in the recovery of target mRNA expression and activity. Increasing evidences indicate that the ceRNA network affects various biological processes in mammals, including development, cellular differentiation, metabolism, immune response, and disease pathogenesis. In teleost fish, the lncRNA-miRNA-mRNA regulatory networks have been reported occasionally. However, up to now, the roles of lncRNAs in the big-belly seahorse (Hippocampus abdominalis) remains unclear. In this study, we reported for the first time, via whole-transcriptome RNA sequencing, the lncRNA mediated ceRNA regulatory network in Vibrio harveyi-infected H. abdominalis. A total of 4197 differentially expressed mRNAs (DE-mRNAs), 1317 DE-lncRNAs, and 183 DE-miRNAs were identified. Furthermore, the crosstalk between miRNAs and lncRNAs as well as between miRNAs and mRNAs was inferred based on the negative correlations between miRNAs and their target lncRNAs/mRNAs. A core immune associated lncRNA-miRNA-mRNA putative regulatory network was thus constructed, comprising 211 lncRNA-miRNA and 224 mRNA-miRNA pairs. In conclusion, our findings provide an integrative overview of the ceRNA regulatory networks on the underlying immune responses to V. harveyi infection in the big-belly seahorse, and offer a solid theoretical foundation for the comparative immunological research of teleost fish.

Animals

Circular RNAs in amyotrophic lateral sclerosis.

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder characterized by the progressive loss of motor neurons, with most cases lacking a clear genetic basis. Emerging evidence highlights the involvement of non-coding RNAs, particularly circular RNAs (circRNAs), in disease onset and progression. Here, we investigated circRNAs implicated in ALS and related motor neuron diseases (MNDs). Here, we provide a general overview of circular RNA metabolism and cellular functions. We then present our systematic literature review that identified ALS-associated circRNAs, followed by in silico analyses of 15 circular RNA candidates that were selected based on the most compelling data regarding ALS. Our results revealed that several circular RNAs regulate ALS-related genes, such as unfolded protein response, oxidative stress, cell cycle regulation, and apoptosis. Protein-RNA interaction analysis further showed that ALS-related circRNAs can sponge 20 RNA-binding proteins. Additionally, molecular docking analysis demonstrated that ALS-associated FUS variants significantly alter its binding affinity to circular RNAs. RNA-seq data from ALS patients confirmed significant alterations in the expression of host genes of ALS-related circRNAs and hub proteins in ALS-affected CNS tissues. Collectively, our findings identify circRNAs as potential key contributors to ALS pathogenesis.

Amyotrophic Lateral Sclerosis

A contextual activity score (CAS) for inferring ADAR-associated transcriptional activity across RNA-seq, single-cell, and spatial transcriptomics.

BACKGROUND AND OBJECTIVE: Adenosine-to-inosine RNA editing, catalyzed by Adenosine Deaminases Acting on RNA (ADARs), is a widespread modification involved in neural function, immune regulation, and cancer. The Alu Editing Index (AEI) is the standard metric to estimate ADAR activity but requires raw sequencing reads and is poorly suited for single-cell and spatial transcriptomic data. This study aimed to develop an alternative framework for inferring ADAR-associated transcriptional activity from gene expression data across diverse transcriptomic technologies. METHODS: We developed the Contextual Activity Score (CAS), a framework based on transcriptional signatures from ADAR perturbation experiments. Context-specific signatures were generated for human neurons, mouse neurons, and cancer models to infer ADAR1 and ADAR2 activity. CAS was computed from normalized gene expression matrices using regulon-based enrichment analysis. Performance was evaluated by comparing with the Alu Editing Index across bulk RNA sequencing datasets, simulated sequencing depths, and library preparation protocols. RESULTS: CAS showed strong concordance with the Alu Editing Index across multiple datasets, while remaining robust to reduced sequencing depth and different library protocols. Unlike the Alu Editing Index, CAS can be applied to single-cell and spatial transcriptomic data and enables the independent assessment of ADAR2 activity. In cancer and neuronal contexts, CAS captured biologically meaningful variations in ADAR-associated transcriptional activity at sample, cell-type, and spatial levels. CONCLUSION: CAS provides a scalable approach applicable across multiple RNA-seq protocols for estimating ADAR-associated transcriptional activity using gene expression data. This method, implemented in an open-source R package for broad adoption, expands the ability to study ADAR-associated transcriptional activity across transcriptomic modalities where direct editing quantification is challenging, such as single-cell and spatial transcriptomics.

Adenosine Deaminase