Search PubMedSearch

SEARCH · Search PubMed

Results for “RNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

RNA Sequencing Protocols for Short-Read Sequencing.

RNA sequencing (RNA-seq) methodologies allow the discovery of novel variants and transcripts. These comprise three general steps: (1) capture of RNA species of interest, (2) conversion of RNA to complementary DNA (cDNA), and (3) modification of cDNA to fit the sequencing platform. Here we describe four different library preparation protocols for short-read sequencing: cDNA synthesis with poly(A) selection, library preparation with ribosomal depletion, and cDNA synthesis with SMART® (Switching Mechanism at 5' end of RNA Template) technology for low and Pico inputs.

Gene Library

ModiCal: A Targeted Calibration Workflow for Site-Specific m5C Validation by Nanopore Direct RNA Sequencing.

Accurate identification of RNA 5-methylcytidine (m5C) at the single-nucleotide resolution remains a central challenge in nanopore direct RNA sequencing (DRS). Current global scanning and modification-aware basecalling methods enable transcriptome-wide profiling but often yield high false-positive rates and lack site-specific accuracy. To address this, we repurposed ModiDeC, originally a de novo multimodification classifier, into a targeted, high-precision validation tool for RNA modification sites with prior biochemical knowledge. This was implemented through a three-step calibration workflow that alternates between biochemical and computational modules using the well-characterized m5C2278 site in 25S rRNA as a starting point. Baseline training uses short synthetic RNAs carrying either a methylated or unmodified C2278 as ground truth, followed by IVT-derived calibration and validation in methyltransferase knockout yeast. The baseline model accurately detected the bona fide m5C2278 site but initially produced off-target predictions. Iterative retraining with unmodified IVT signals progressively reduced and ultimately eliminated false positives while maintaining a strong signal at the bona fide site. The final model retained enzyme-dependent detection in wild-type versus knockout yeast and, when explicitly targeted, was also able to detect the second rRNA site, C2870, which remained invisible in the initial analysis. Application to native human prerRNA processing intermediates further resolved two distinct m5C deposition regimes on 28S rRNA, while generalization to dengue virus genomic RNA confirmed that the same calibration logic transfers across diverse RNA contexts. Together, this study establishes a reproducible and transferable framework that integrates biochemical validation with iterative neural network refinement, providing a route toward reliable site-specific m5C confirmation by nanopore direct RNA sequencing.

RNA Methylation

Targeted reflex RNA sequencing for enhanced variant classification on exome and genome sequencing improves patient outcomes.

RNA sequencing (RNA-seq) has been utilized to provide functional evidence regarding the impact of splicing variants. This study explores the utility of targeted reflex RNA-seq to inform classification of predicted splicing variants identified through clinical exome sequencing (ES) and genome sequencing (GS). A retrospective analysis was conducted on consecutive ES/GS cases completed at a single center in which targeted reflex RNA-seq was performed following identification of eligible variants. There were 131 cases (4.1%) that had at least one RNA-seq eligible variant reported, with eight of these cases having two unique eligible variants. Of the 139 eligible variants, 125 were classified as variants of uncertain significance (VUS). Sixty-four cases had targeted reflex RNA-seq completed with 27 cases having at least one variant reclassified (42.2%). After reclassification, 23 cases had positive results, and two cases had a likely diagnosis of an autosomal recessive condition. Clinical outcomes data regarding positive RNA-seq cases showed that 71% (10/14) had clinical management changes and 43% (6/14) had treatment changes. Incorporation of targeted reflex RNA-seq analysis into the diagnostic pipeline of rare diseases enhances variant classification and resolves uncertainty regarding predicted splice variants, leading to an estimated 1.6% increase in diagnostic yield of clinical ES/GS.

Journal Article

New insights on Plasmodium gene expression from direct RNA sequencing.

Oxford Nanopore Technology (ONT) direct RNA sequencing enables the sequencing of native RNA molecules without cDNA conversion. The long-read approach captures full-length reads spanning entire genes and has transformed the study of gene expression in Plasmodium parasites by enabling analysis of untranslated regions, isoforms, and alternative splicing. In addition, ONT provides unique insights into non-coding RNAs, RNA modifications, and polyadenylated tail dynamics, which are expanding our understanding of post-transcriptional regulation in Plasmodium, including processes beyond translational repression in gametocytes and sporozoites. Here, we discuss the past and future applications of direct RNA sequencing in Plasmodium research and highlight its advantages, limitations, and future prospects.

Oxford Nanopore Technology

From transcriptomic profiling to precision oncology: a bibliometric analysis of RNA sequencing in acute myeloid leukemia.

BACKGROUND: RNA sequencing (RNA-seq) has become an important tool for investigating the molecular heterogeneity of acute myeloid leukemia (AML); however, the global development and thematic evolution of this field remain inadequately characterized. OBJECTIVE: To map the global landscape of AML RNA-seq research and identify major knowledge domains, emerging themes, and temporal changes in research priorities. METHODS: Publications indexed in the Web of Science Core Collection and Scopus between January 1, 2007, and August 18, 2025, were retrieved. After database filtering, merging, and deduplication, 3,460 articles and reviews were included. CiteSpace, VOSviewer, the bibliometrix R package, and Microsoft Excel were used to analyze publication trends, collaboration networks, co-citation structures, keyword evolution, and citation bursts. RESULTS: Publication output increased steadily, accelerating after 2014. China contributed the largest number of publications (n = 547, 15.8%), whereas the United States had the highest total citation count. Major publication outlets spanned hematology, oncology, genomics, and molecular biology. Co-citation analysis identified prominent themes involving next-generation sequencing, gene mutations, KMT2A rearrangements, epigenetic dysregulation, leukemia-initiating cells, drug resistance, biomarkers, T-cell biology, and single-cell sequencing. Earlier literature emphasized sequencing technologies, gene expression profiling, and molecular alterations, whereas recent publications show increasing representation of cellular heterogeneity, single-cell transcriptomics, drug resistance, biomarker applications, immune-related research, and computational interpretation. CONCLUSION: While molecular characterization remains foundational, AML RNA-seq research has broadened to encompass increasingly prominent cellular, functional, computational, and translational dimensions. This study provides a structured overview of the field; nevertheless, bibliometric prominence should not be interpreted as direct evidence of clinical utility.

RNA sequencing

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5 ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91 µg of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8 million filtered reads per sample, with 63±7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9 million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans

Characterization of METTL3/14-mediated m6A modification in human transcriptome using Nanopore direct RNA sequencing.

Post-transcriptional RNA modifications modulate diverse aspects of RNA metabolism. N6-methyladenosine (m6A), one of the most abundant internal RNA modifications, is deposited by the core methyltransferase complex, METTL3 and METTL14. Oxford Nanopore Technologies (ONT) platform permits direct, single RNA molecule sequencing while preserving native modifications. However, without rigorous benchmarking, the accuracy and reproducibility of modification detection remain uncertain. Here, we leveraged ONT to comprehensively profile bona fide m6A modifications in cellular RNAs at single-nucleotide resolution by integrating two direct RNA sequencing chemistries (RNA002 and RNA004) with the m6Anet and Dorado modification-detection models. We independently depleted METTL3 and METTL14 in human cells and rigorously validated modification calls through several assays and independent orthogonal methods (GLORI and miCLIP). We find that Dorado detected a higher number of m6A events and enabled simultaneous detection of other RNA modifications (5-methylcytosine, pseudouridine, and inosine). Pairing Dorado with an in vitro transcribed, unmodified control under stringent filtering, we provide compelling evidence supporting a global reduction in m6A sites and stoichiometry within coding sequences and across genes, particularly in highly modified genes and sites, and at consensus DRACH motifs. We report a differential and complex regulation of modified transcripts, accompanied by a global reduction in poly(A) tail length. Notably, METTL3 and METTL14 depletion produced distinct transcript-specific effects, supporting non-redundant roles within the m6A writer complex. Together, our study illustrates a notable advancement of ONT capabilities and establishes a robust transcriptome-wide framework for RNA modification detection, thereby laying the groundwork for exploring the contribution of METTL3/METTL14 to cellular functions and disease.

Humans

Enzymes in high-throughput RNA sequencing: Applications and challenges.

High-throughput RNA sequencing provides genome-wide information on the dynamics of RNA in each cell and how the dynamics responds to environmental changes. Next-generation sequencing by the Illumina platform currently provides the highest information output as compared to other platforms. A key component of next generation sequencing of each RNA is the successful end-to-end reverse-transcription into a cDNA strand. This can be highly challenging given the propensity of each RNA to adopt ordered structures and to contain post-transcriptional modifications. While many reverse transcriptase (RT) enzymes have been developed over the years to maximize read-through of an RNA, their processivity and efficiency varies, raising the question of how to select the RT for the experiment at hand. Here, we use tRNA as a model for genome-wide sequencing, as tRNA has a stable secondary and tertiary structure and has a high density and wide variety of post-transcriptional modifications, presenting one of the most challenging problems of sequencing RNA. We compare the efficiency of end-to-end cDNA synthesis of tRNA among several recent RT enzymes and provide a general sequencing workflow that is applicable to most of these enzymes.

High-Throughput Nucleotide Sequencing

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

RNA sequencing offers new diagnostic opportunities in neurodevelopmental disorders: A systematic review.

PURPOSE: Transcriptomics by way of RNA sequencing (RNAseq) has emerged as a means to increase the diagnostic yield in genetic conditions. In this systematic review, we focus on the contribution of transcriptomics to improve the diagnostic yield in neurodevelopmental disorders. METHODS: We performed a systematic literature search in PubMed until January 2024, including articles describing diagnostic RNAseq on at least 1 individual with a primary neurodevelopmental phenotype. We extracted data on cohort size, phenotype, sample tissue, previously used diagnostic methods, added diagnostic yield of RNAseq, the use of control samples, and technical aspects of the RNA sequencing methodology. RESULTS: A total of 17 articles were eligible for inclusion in the systematic review. We found an average added diagnostic yield of 15.5% through RNA sequencing for individuals with neurodevelopmental disorders. There is heterogeneity in the tissue type, reported quality measures, and the computational pipeline. CONCLUSION: The significantly increased diagnostic yield demonstrates the value of this novel tool in the diagnostic setting of neurodevelopmental disorders. Our results offer an overview of common methodologies for RNAseq and allow us to formulate recommendations for genetic labs and clinicians when implementing RNAseq as a diagnostic tool. Lastly, we provide recommendations for future publications to increase transparency and reproducibility.

Humans

Single-cell RNA sequencing provides further insights into the immunostimulatory action of freeze-dried Lactiplantibacillus plantarum on Penaeus vannamei shrimp.

Immunostimulation through dietary interventions opened new avenues in developing disease control and prevention tools for shrimp aquaculture. We have previously shown that feeding with freeze-dried Lactiplantibacillus plantarum (LAB) increased disease resistance of Penaeus vannamei against both Vibrio parahaemolyticus and white spot syndrome virus (WSSV) based on bulk RNA sequencing of shrimp gills. This tissue participates in ion transport and serves as a first line of defense against environmental stressors and pathogenic infections. However, characterization of their cell composition and functions remains limited. Here, we implemented a single-cell RNA sequencing approach to further gather insights into how feeding with freeze-dried LAB modulates host immunity which may not be evident with bulk RNA sequencing approach. A total of five clusters with unique transcriptional signatures were identified, corresponding to pillar cells, septal cells, and sessile hemocytes. Pseudo-bulk analyses at global- and cluster-levels showed differential expression of genes related to host immunity and metabolism. We further revealed how overall transcriptomic changes are not exclusively caused by gene expression changes but may also be driven by cell population dynamics. This study highlighted how single-cell RNA sequencing approach may shed light on the mechanisms of action of immunostimulants which may be masked in bulk transcriptome analyses.

Animals

Model-directed generation of artificial CRISPR-Cas13a guide RNA sequences improves nucleic acid detection.

CRISPR guide RNA sequences deriving exactly from natural sequences may not perform optimally in every application. Here we implement and evaluate algorithms for designing maximally fit, artificial CRISPR-Cas13a guides with multiple mismatches to natural sequences that are tailored for diagnostic applications. These guides offer more sensitive detection of diverse pathogens and discrimination of pathogen variants compared with guides derived directly from natural sequences and illuminate design principles that broaden Cas13a targeting.

CRISPR-Cas Systems

Heat Inactivation of Nipah Virus for Downstream Single-Cell RNA Sequencing Does Not Interfere with Sample Quality.

Single-cell RNA sequencing (scRNA-seq) technologies are instrumental to improving our understanding of virus-host interactions in cell culture infection studies and complex biological systems because they allow separating the transcriptional signatures of infected versus non-infected bystander cells. A drawback of using biosafety level (BSL) 4 pathogens is that protocols are typically developed without consideration of virus inactivation during the procedure. To ensure complete inactivation of virus-containing samples for downstream analyses, an adaptation of the workflow is needed. Focusing on a commercially available microfluidic partitioning scRNA-seq platform to prepare samples for scRNA-seq, we tested various chemical and physical components of the platform for their ability to inactivate Nipah virus (NiV), a BSL-4 pathogen that belongs to the group of nonsegmented negative-sense RNA viruses. The only step of the standard protocol that led to NiV inactivation was a 5 min incubation at 85 °C. To comply with the more stringent biosafety requirements for BSL-4-derived samples, we included an additional heat step after cDNA synthesis. This step alone was sufficient to inactivate NiV-containing samples, adding to the necessary inactivation redundancy. Importantly, the additional heat step did not affect sample quality or downstream scRNA-seq results.

Nipah Virus

Integrating RNA sequencing with deep learning-based metabolic toxicity prediction: A new perspective on screening prioritized liquid crystal monomers.

Nearly 99 % of liquid crystal monomers (LCMs) toxicological data remains gaps, especially to aquatic organisms. Herein, this study proposes a rapid and high-throughput screening method for identifying priority LCMs in natural water. Using six fluorinated LCMs (LCMsF) with significant enrichment characteristics in zebrafish as examples, RNA sequencing revealed that LCMsF-induced metabolic disturbances are predominant, including 28 Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway abnormalities attributed to 498 differentially expressed genes. Notably, the intricate sequencing process resulted in the inability to rapid identify additional 857 LCMsF that may induce metabolic disturbances. To address this, LCMsT-MTP, a predictive deep learning model based on RNA sequencing, was developed. This model integrates a comprehensive representation of LCMsF structures and metabolic toxicity target sequences. LCMsT-MTP improves upon traditional methods that are limited to single targets and mechanisms by facilitating the simultaneous identification of 21 metabolic toxicities induced by LCMsF. In addition, the LCMsT-MTP model was further applied to non-fluorinated LCMs (LCMsNone F) that satisfy the applicability domains test. Accordingly, a metabolic toxicity priority list of LCMs was proposed, with ∼95 % of LCMs classified as high or medium risk. Priority list validation by molecular dynamics confirmed that the interactions of LCMsF/LCMsNone F and metabolic toxicity targets in representative KEGG pathways were distinct.

Animals

Consistently processed RNA sequencing data from 50 sources enriched for pediatric data.

Larger cohorts improve the power of tumor gene expression analysis, but the signal is muddied if datasets are processed using different methods or have inaccurate metadata. Here we present five compendia containing consistently processed gene expression data derived from 16,446 diverse RNA sequencing datasets. To create the compendia, we obtained access to RNA sequence data from repositories containing public data as well as clinical partners with access to non-published data. We then assessed the quality, quantified gene expression, harmonized clinical metadata, and released the expression values and metadata without access restrictions. These datasets have been used for diverse projects ranging from identifying similarities between tumor types to assessing how well cell lines recapitulate tumors. They have also been used for n-of-1 analysis to identify genes with unusual expression patterns in a single sample and to infer molecular diagnosis. The comparison to new data is enabled by our dockerized, freely available pipeline. The compendia have been cited in at least 20 publications.

Humans

scnanoseq: an nf-core pipeline for Oxford Nanopore single-cell RNA-sequencing.

MOTIVATION: Recent advancements in long-read single-cell RNA sequencing (scRNA-seq) have facilitated the quantification of full-length transcripts and isoforms at the single-cell level. Historically, long-read data would need to be complemented with short-read single-cell data in order to overcome the higher sequencing errors to correctly identify cellular barcodes and unique molecular identifiers. Improvements in Oxford Nanopore sequencing, and development of novel computational methods have removed this requirement. Though these methods now exist, the limited availability of modular and portable workflows remains a challenge. RESULTS: Here, we present, nf-core/scnanoseq, a secondary analysis pipeline for long-read single-cell and single-nuclei RNA that delivers gene and transcript-level quantification. The scnanoseq pipeline is implemented using Nextflow and is built upon the nf-core framework, enabling portability across computational environments, scalability and reproducibility of results across pipeline runs. The nf-core/scnanoseq workflow follows best practices for analyzing single-cell and single-nuclei data, performing barcode detection and correction, genome and transcriptome read alignment, unique molecular identifier deduplication, gene and transcript quantification, and extensive quality control reporting. AVAILABILITY AND IMPLEMENTATION: The source code, and detailed documentation are freely available at https://github.com/nf-core/scnanoseq and https://nf-co.re/scnanoseq under the MIT License. Documentation for the version of nf-core/scnanoseq used for this paper, including default parameters and descriptions of output files are available at https://nf-co.re/scnanoseq/1.1.0.

Single-Cell Analysis

RNA sequencing resolves a novel noncanonical splice-region variant in PHKA2 causing glycogen storage disease type IX α2: a case report.

BACKGROUND: Glycogen storage disease type IX α2 (GSD IX α2) is an X-linked hepatic glycogenosis caused by pathogenic variants in PHKA2. Noncanonical splice-region variants located outside the invariant GT/AG dinucleotides pose significant interpretive challenges, as in silico predictions alone are often insufficient for definitive classification. CASE DESCRIPTION: We report a 2.9-year-old boy presenting with short stature, hepatomegaly, markedly elevated aminotransferases, fasting hypoglycemia with ketonuria, hypercholesterolemia, coagulation parameter abnormalities (decreased fibrinogen and prolonged thrombin time), and histological evidence of early hepatic fibrosis as demonstrated by Masson's trichrome staining (portal fibrosis and perisinusoidal fibrosis). Whole-exome sequencing (WES) identified a hemizygous, previously unreported PHKA2 variant [NM_000292.3:c.2517+5G>T, genomic location (GRCh38): NC_000023.11: g.18907895G>T], initially classified as a variant of uncertain significance (VUS) under American College of Medical Genetics and Genomics (ACMG) criteria. RNA sequencing of peripheral blood leukocytes demonstrated predominant exon 22 skipping in 94.2% of informative junction reads, predicting a frameshift and premature termination codon [p.(Gly788Profs*74)] with predicted loss of the C-terminal CBL 2 subdomain. Incorporating this transcript-level evidence, the variant was reclassified as pathogenic (PVS1 + PM2_Supporting + PP4). Following dietary management with uncooked cornstarch supplementation, the patient showed progressive biochemical improvement over a 2.2-year follow-up. CONCLUSIONS: This case expands the mutational spectrum of PHKA2 and demonstrates that RNA sequencing of accessible tissues is a practical and diagnostically informative strategy for resolving noncanonical splice-region variants in pediatric hepatic GSD. Early hepatic fibrosis detected by histological examination before age 3 years underscores the importance of longitudinal hepatic surveillance in GSD IX α2.

Glycogen storage disease type IX α2 (GSD IX

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles