Search PubMedSearch

SEARCH · Search PubMed

Results for “Sequence Analysis, DNA”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

1,671 records · Page 8Linked to original sources

Temporal shifts in gyrA mutation types and sublineage replacement in ST11 Salmonella enterica serovar Enteritidis over a decade (2014-2023): A genomic epidemiological study in Guangxi, China.

The overuse or abuse of antibiotics drives the global health threat of antimicrobial resistance. Although bans on certain veterinary antibiotics, such as colistin, have proven effective, the impact of fluoroquinolone stewardship on the evolution of the foodborne pathogen Salmonella enterica serovar Enteritidis (S. Enteritidis) remains unclear. Here, we conducted a decade-long (2014-2023) retrospective longitudinal genomic epidemiological analysis of 441 ST11 S. Enteritidis isolates from Guangxi, China, alongside a global reference dataset of 4297 genomes. Our aim was to elucidate the effect of real-world antibiotic stewardship on the shift of gyrA point mutations and lineage distribution. Surveillance identified three global epidemic clade sublineages (GEC-L2, L3, L4), with the multidrug-resistant GEC-L4 (i.e., GC-c or MMC2), characterized by the gyrA mutation with amino acid substitution D87Y, being domestically dominant (70.07%, 309/441). Following China's 2016 ban on the veterinary use of critical fluoroquinolones, the proportion of the highly resistant GEC-L4 sublineage decreased continuously (from 86.84% in 2017 to 56.00% in 2023), while the less resistant GEC-L3 sublineage (i.e., GC-b or MMC1), mainly characterized by gyrA D87G, increased simultaneously (from 13.16% to 44.00%). This phenomenon might be attributed to the fact that the GEC-L4 sublineage exhibited a higher fitness cost compared with the GEC-L3 sublineage, as confirmed by the competition assay. A Random Forest Model validated that the gyrA mutation with amino acid substitution D87Y was the paramount feature for these sublineages' identification. In contrast, global data showed a continuous increase in gyrA mutations (from 8.63% in 2006 to 68.85% in 2024), primarily D87Y (from 1.44% to 31.15%) and D87N (from 4.32% to 22.95%), correlating with rising average fluoroquinolone consumption. This study provides direct genomic evidence that national-level antibiotic stewardship can drive the replacement of highly resistant sublineages with moderately resistant ones. These findings offer crucial scientific evidence for evaluating the impact of antibiotic management policies and inform strategies for the rational use of antimicrobials.

China

Tracing the evolution and diversity of human parvovirus B19 across human history.

Human parvovirus B19 (B19V) is an ubiquitously spread, exclusively human pathogen, mainly posing risks to children, as well as pregnant and immunocompromised individuals. Despite evidence of B19V infection of human populations as far back as 7,000 years, the evolutionary history of B19V remains poorly understood. In this study, we present B19V genomic data from the remains of 53 globally distributed individuals spanning more than 8,000 years, including 7 children. Our findings suggest that the most recent common ancestor of all present B19V lineages existed around 12,000 years ago, at the end of the last Ice Age. Additionally, we identified an extinct Eurasian clade that participated in the recombination event that led to the emergence of B19V genotype 2 (GT-2). We date this event to ∼3,200-1,800 BP, potentially in the greater Mediterranean area. Our study shows aspects of how ancient parvovirus variants arose, disseminated, and impacted human health through time.

ancient DNA

In vitro evaluation of sacituzumab govitecan in non-small cell lung cancer with actionable genomic alterations.

PURPOSE: The TROP2-directed antibody-drug conjugate sacituzumab govitecan (SG) has shown substantial therapeutic benefit in several malignancies; however, preclinical evidence supporting its activity in non-small cell lung cancer (NSCLC) is rare. MATERIALS AND METHODS: We evaluated 16 NSCLC cell lines harboring actionable genomic alterations for TROP2 expression and treated them with SG or its unconjugated payload, SN-38, for 3 days to determine cytotoxic effects. Apoptosis and DNA damage signaling were assessed using flow cytometry and western blot. SG internalization and lysosomal trafficking were visualized by confocal microscopy. RESULTS: SG had greater cytotoxic potency than SN-38, across all NSCLC cell lines, independent of genomic subtype or TROP2 expression level. Cell lines that were sensitive to SN-38 showed enhanced vulnerability to SG (P < 0.0001). Higher SLFN11 expression, a recognized determinant of SN-38 responsiveness, correlated with lower SG IC50 values. Both SG and SN-38 triggered apoptotic and DNA damage responses within 6-48 h, with SG inducing stronger activation of these pathways than SN-38. SG was efficiently taken up in CUTO17 and SNU-3173 adenocarcinoma cells, with more than 60% of the conjugate internalized within 3 h and subsequently localized to lysosomes. CONCLUSION: Our study provides in vitro evidence supporting the potential activity of SG in NSCLC with actionable genomic alterations. The efficacy of SG closely paralleled intrinsic sensitivity to the SN-38 payload, suggesting that DNA-damage responses, rather than oncogenic drivers, predominantly contribute to SG activity.

Actionable genomic alterations

A dual-dimensional CRISPR toolkit enables one-step high-efficiency multiplex genome editing in Komagataella phaffii.

Against the backdrop of green biomanufacturing, engineering methanol-utilizing Komagataella phaffii (K. phaffii) represents an effective strategy to expand the one carbon (C1) product profile and speed up the industrialization of C1-based bioeconomy. To address the technical challenges of low efficiency and cumbersome experimental procedures for multiplex gene editing and precise large-fragment integration during the reconstruction of complex metabolic pathways in K. phaffii, this study established a CRISPR toolkit - Efficient Multi-Gene Editing System 3.0 (EMGES 3.0) - which enabled one-step large-fragment integration coupled with multiplex gene knockout. EMGES 3.0 was constructed through the synergistic optimization of a repair-engineered chassis and an episomal CRISPR vector. For chassis engineering, five DNA repair modules: &#x394;lig4 (DNA Ligase IV, non-homologous end joining end ligation), ppMRE11(The endogenous MRE11 gene from Pichia pastoris) overexpression (The Meiotic Recombination 11, DNA double-strand break end resection), &#x394;rad9 (Radiation-Sensitive 9, DNA damage checkpoint regulation), &#x394;mph1 (Mutator Phenotype Helicase 1, improvement of homologous recombinant strand extension), and PapRecT-PaSSB co-expression (stabilization of recombination intermediates) were integrated to generate the highly recombinogenic strain Y09. For vector engineering, cenARS was replaced by panARS and the endogenous promoter PGAP was employed to drive the double hammerhead ribozyme-single guide RNA-hepatitis delta virus ribozyme (double HH-sgRNA-HDV: dHgH)-mediated sgRNA expression, yielding the optimized vector Nov_pGAP_panARS_pLAT1_Cas9. These two features on K. phaffii together enhanced the EMGES 3.0 to a higher standard of transformation rate and editing efficiency. According to our results, EMGES 3.0 achieved dual-functional gene knockout efficiencies between 76.6% and 100%. For insertion of medium-long fragments (>4.5&#x202f;kb), the efficiency achieved 93.3%. In addition, the one-step integration of ultra-long fragments (>16&#x202f;kb) achieved 14.8%, which was reported for the first time. Furthermore, the efficiency of simultaneous long-fragment integration at three neutral loci reached 38.4% (>15&#x202f;kb). We applied the system for one-step production of free fatty acids (FFAs, yield: 5.82 &#x223c; 7.30&#x202f;mg/L/OD600) and resveratrol (yield: 1.14 &#x223c; 1.28&#x202f;mg/L) using methanol as the sole carbon source. EMGES 3.0 provides a robust technical foundation for complex compounds biosynthesis and high-yield industrial strains, while also advancing K. phaffii as an industrial synthetic biology chassis for efficient C1 utilization.

CRISPR-Cas Systems

Biomarker Analysis from Patients with Metastatic PDAC Treated with TGF&#x3b2; Antibody NIS793 plus Abraxane + Gemcitabine versus Abraxane + Gemcitabine Alone in a Phase II, Open-Label, Randomized Study.

PURPOSE: Transforming growth factor &#x3b2; (TGF&#x3b2;) plays a dual role in cancer, acting as a tumor suppressor early in the disease but promoting progression and immune evasion when dysregulated. In pancreatic ductal adenocarcinoma (PDAC), TGF&#x3b2;-driven desmoplasia fosters chemoresistance and immunosuppression, limiting therapeutic efficacy. NIS793, a fully human mAb targeting TGF&#x3b2;, demonstrated antifibrotic and immunomodulatory activity in preclinical models and early-phase trials. PATIENTS AND METHODS: We conducted a randomized, open-label, phase II study in treatment-na&#xef;ve patients with metastatic PDAC (mPDAC) to evaluate NIS793 &#xb1; spartalizumab (anti-PD-1) combined with nab-paclitaxel (or Abraxane)/gemcitabine (ABRA/GEM) versus ABRA/GEM alone. The primary endpoint was progression-free survival (PFS); secondary endpoints included overall survival (OS), safety, pharmacokinetics, and biomarker analyses. Exploratory assessments included paired tumor RNA sequencing, cell-free DNA profiling, and plasma proteomics. RESULTS: NIS793 demonstrated target engagement and suppression of TGF&#x3b2; signaling, confirmed by transcriptomic and proteomic analyses. Stromal remodeling was evident, with significant downregulation of cancer-associated fibroblast markers (Acta2, Fap) and collagen-related signatures. Despite proof of mechanism, clinical efficacy was not observed: Median PFS and OS were comparable or numerically worse in the NIS793 arm versus control (HR for OS in NIS793 + ABRA/GEM vs. ABRA/GEM: 1.32; 95% confidence interval, 0.84-2.07). The safety profile was manageable, with no unexpected toxicities. Biomarker data revealed increased expression of neutrophil-related genes after treatment, suggesting potential induction of tumor-promoting inflammation. CONCLUSIONS: NIS793 effectively inhibited TGF&#x3b2; signaling and led to stromal remodeling but failed to improve outcomes in mPDAC. These findings highlight the complexity of TGF&#x3b2; biology and caution against its blockade in combination with chemotherapy for PDAC. Future strategies should consider context-dependent effects of TGF&#x3b2; inhibition.

Humans

Diagnostic and clinical utility of exome sequencing and chromosomal microarray in children with GDD/iD: a meta-analysis.

BACKGROUND: Global developmental delay/intellectual disability (GDD/ID) is among the most common neurodevelopmental disorders, with up to half of cases are attributed to genetic factors. Chromosome microarray (CMA) has traditionally been the primary genetic test for idiopathic GDD/ID. However, whole exome sequencing (WES) and whole genome sequencing (WGS) have recently emerged, substantially increasing diagnostic yields in these populations. METHODS: We conducted a comprehensive literature search of PubMed, Scopus, EMBASE, and the Cochrane Library from inception to April 29, 2025. Studies reporting the diagnostic utility of these tests in children with GDD/ID were included and analyzed. RESULTS: A total of 102 studies, comprising 55,752 children, were reviewed. The pooled diagnostic yield of WES was 0.37 (95% CI: 0.33-0.41; I2 = 93%), significantly higher than that of CMA at 0.19 (95% CI: 0.16-0.21; I2 = 95%). Subgroup analyses showed that WES yielded significantly higher diagnostic rates than CMA in both same-sample comparisons (OR = 2.27, 95% CI: 1.08-4.78) and different-sample comparisons (OR = 1.65, 95% CI: 1.15-2.37). Only one study evaluated WGS, reporting a diagnostic yield of 0.27. Meta-regression revealed a significant association between CMA diagnostic yield and the proportion of male participants (p&#x2009;<&#x2009;0.01), but not with WES. No significant difference in diagnostic utility was observed between isolated GDD/ID and GDD/ID with comorbidities. CONCLUSION: In children with unexplained GDD/ID, WES demonstrates superior diagnostic and clinical utility compared to CMA. Incorporating WES as a first-line investigation in the diagnostic evaluation of GDD/ID may be warranted.

Humans

Complete genome sequence of the Anaplasma phagocytophilum clinical isolate NCH-1.

Anaplasma phagocytophilum is an obligate intracellular gram-negative bacterium and etiologic agent of human granulocytic anaplasmosis. A. phagocytophilum genomic sequencing has historically been performed via short-read platforms. Our optimized bacterial isolation protocol combined with Nanopore sequencing produced a single, closed 1,481,805 bp circular A. phagocytophilum strain NCH-1 chromosome.

Anaplasma phagocytophilum

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Comprehensive analysis of mRNA-microRNA-lncRNA expression profiles in post-traumatic elbow heterotopic ossification using RNA sequencing and experimental validation.

BACKGROUND: This study aimed to profile the molecular signatures of post-traumatic elbow heterotopic ossification (HO) to identify key regulators and potential therapeutic targets. METHODS: Total RNA from post-traumatic elbow HO tissues (n=4) and normal bone tissues (n=6) was subjected to high-throughput sequencing to identify differentially expressed mRNAs (DEGs), microRNAs (DEMs), and lncRNAs (DELs). Bioinformatics analyses included Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment, protein-protein interaction network construction, and transcription factor (TF)-microRNA-mRNA network analysis. The expression trends of four most upregulated and four most downregulated DEGs were validated by real-time quantitative reverse transcription polymerase chain reaction (qRT-PCR). RESULTS: We identified 2,138 DEGs, 40 DEMs, and 905 DELs. DEGs were significantly enriched in biological process "bone mineralization," cellular component "plasma membrane," molecular function "integrin binding," and pathways including PI3K-Akt, NF-&#x3ba;B, JAK-STAT, and TNF signaling pathways. Hub genes with high connectivity included MMP9, IL6, MMP3, CTSK, and BGLAP. Integrated network analysis highlighted the transcription factor JUN and key microRNAs (hsa-miR-124-3p, hsa-miR-548c-3p, and hsa-miR-135b). The qRT-PCR results confirmed the expression trends of selected DEGs. CONCLUSIONS: This study, for the first time, profiled the differentially expressed mRNAs, microRNAs, and lncRNAs in post-traumatic elbow HO using high-throughput RNA sequencing. These findings provide valuable insights into the molecular mechanisms of HO following elbow trauma. The identified hub genes (MMP9, IL6, MMP3, CTSK, and BGLAP), key TF (JUN), and key microRNAs (hsa-miR-124-3p, hsa-miR-548c-3p, and hsa-miR-135b) may serve as potential therapeutic targets for preventing and treating post-traumatic elbow HO.

Humans

Draft genome sequence of Enterococcus casseliflavus strain MBBL_MP4 isolated from healthy bovine milk.

We report the draft genome sequence of Enterococcus casseliflavus MBBL_MP4, recovered from healthy bovine milk. The 3.45-Mbp genome assembly comprises 27 contigs and indicates low pathogenic potential, with no acquired antimicrobial resistance or known virulence genes. This genome provides a valuable resource for the genomic characterization of bovine-associated E. casseliflavus.

Enterococcus casseliflavus

Single-cell RNA sequencing provides further insights into the immunostimulatory action of freeze-dried Lactiplantibacillus plantarum on Penaeus vannamei shrimp.

Immunostimulation through dietary interventions opened new avenues in developing disease control and prevention tools for shrimp aquaculture. We have previously shown that feeding with freeze-dried Lactiplantibacillus plantarum (LAB) increased disease resistance of Penaeus vannamei against both Vibrio parahaemolyticus and white spot syndrome virus (WSSV) based on bulk RNA sequencing of shrimp gills. This tissue participates in ion transport and serves as a first line of defense against environmental stressors and pathogenic infections. However, characterization of their cell composition and functions remains limited. Here, we implemented a single-cell RNA sequencing approach to further gather insights into how feeding with freeze-dried LAB modulates host immunity which may not be evident with bulk RNA sequencing approach. A total of five clusters with unique transcriptional signatures were identified, corresponding to pillar cells, septal cells, and sessile hemocytes. Pseudo-bulk analyses at global- and cluster-levels showed differential expression of genes related to host immunity and metabolism. We further revealed how overall transcriptomic changes are not exclusively caused by gene expression changes but may also be driven by cell population dynamics. This study highlighted how single-cell RNA sequencing approach may shed light on the mechanisms of action of immunostimulants which may be masked in bulk transcriptome analyses.

Animals

Evaluation of one-step amplicon-based targeted enrichment for SARS-CoV-2 whole-genome sequencing using the Midnight amplicon scheme.

Genomic surveillance proved invaluable during the COVID-19 pandemic for tracking SARS-CoV-2 variants and guiding outbreak responses, underscoring the ongoing need to reduce whole-genome sequencing (WGS) costs and improve workflow efficiency to ensure accessibility in resource limited settings. Here, we evaluated a one-step reverse transcription polymerase chain reaction (RT-PCR) approach using the Midnight V2 primer scheme for targeted amplification of the SARS-CoV-2 genome, assessed its compatibility with Illumina sequencing, and compared its performance to a well-established two-step method. Initially, we determined optimal RT-PCR reaction conditions using the Midnight V2 primer panel for the one-step RT-PCR kit and scaled reaction volumes for both RT-PCR and library preparation. Clinical specimens (n&#x202f;=&#x202f;53) that had undergone routine WGS for surveillance purposes using the established two-step RT-PCR method were compared using the one-step RT-PCR assay. For samples with genome completeness greater than 70%, both methods gave comparable results with similar sequence coverage and 100% concordance for lineage assignment. Further investigation revealed a higher percentage of reads aligning to the SARS-CoV-2 genome with a greater depth of coverage using the one-step method compared to the two-step method. Finally, analysis of scaled one-step and library reaction volumes revealed significant cost savings for samples undergoing WGS. Overall, the results presented here verify the accuracy and reproducibility of one-step targeted amplification and offer an efficient and cost-effective workflow for routine SARS-CoV-2 genomic surveillance.

Humans

Integrated exome and mitochondrial genome sequencing reveals the genetic landscape of primary mitochondrial diseases: findings from a large Tunisian cohort.

Primary mitochondrial diseases are a heterogeneous group of neurometabolic disorders recognized as the most common metabolic genetic diseases. They manifest at any age, affecting any tissue or organ, especially those with high energy demands, and are caused by pathogenic variants in both mitochondrial and nuclear genomes. Here, we aimed to describe the genetic spectrum of a Tunisian pediatric cohort with suspected mitochondrial diseases. We recruited 47 unrelated families who underwent exome sequencing as a first-tier test followed by whole mitochondrial genome sequencing for unsolved cases. Dedicated bioinformatic pipelines and prediction tools were used to determine the potential disease-causing variants. Sanger sequencing confirmed the presence and segregation within parents. For the newly identified variants, structural modeling was conducted to study the impact of these variants on protein structure and motions. Dual genome sequencing yielded a molecular diagnosis in 33/47 families (70%) and 18/47 (38%) showed disease-causing variants in genes encoding mitochondrial proteins. Among them, four families disclosed novel variants in FASTKD2, SERAC1 and GATB, which were supported by in-depth in silico and structural analyses demonstrating their deleterious effect. The remaining families (32%, 15/47) disclosed other metabolic and neurological disorders. An exome-first strategy delivers a high diagnostic yield in Tunisia, where consanguinity remains high and simultaneously captures mitochondrial and non-mitochondrial etiologies. Mitochondrial sequencing remains indispensable in the case of an inconclusive exome. Thus, our data expand the clinical and genetic spectrum of primary mitochondrial diseases in Tunisia, an underrepresented and admixed population.

Humans

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

Complete genome sequence of multidrug-resistant Salmonella enterica subsp. enterica serovar Enteritidis SD191 isolated from chicken liver, harboring a novel imipenem resistance mechanism.

We present the complete genome sequence of Salmonella enterica subsp. enterica serovar Enteritidis SD191 isolated from Gallus gallus liver in China, harboring plasmid pSE191. The genome reveals multiple antibiotic resistance mechanisms and phenotypic imipenem resistance without canonical genes.

antibiotic resistance

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine