Search PubMedSearch

SEARCH · Search PubMed

Results for “Long-read RNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Biallelic VPS41 Variants in Autosomal Recessive Spinocerebellar Ataxia 29 Resolved by Long-Read Sequencing and RNA Analysis.

BACKGROUND: Biallelic variants in VPS41, encoding a subunit of the HOPS complex, cause autosomal recessive spinocerebellar ataxia 29 (SCAR29), a rare neurodevelopmental disorder with an incompletely defined phenotypic and molecular spectrum. METHODS: We investigated a 24-year-old man with cerebellar ataxia, hypotonia, and intellectual disability. Exome sequencing identified four candidate VPS41 variants. Because maternal DNA was unavailable, long-read genome sequencing was performed to determine allelic configuration, followed by RNA and protein analyses. RESULTS: In addition to typical SCAR29 features, the patient showed previously unreported findings, including swan-neck deformities and pes cavus. Long-read genome sequencing demonstrated that two VPS41 variants were in trans. RNA analysis revealed distinct splicing consequences: one allele produced an out-of-frame transcript predicted to undergo nonsense-mediated decay, whereas the other generated an in-frame exon-skipped transcript. These complementary defects reduced VPS41 expression at both transcript and protein levels, supporting pathogenicity and variant reclassification. CONCLUSION: Our findings expand the phenotypic spectrum of VPS41-related disease and highlight the value of long-read allelic resolution in clarifying pathogenic mechanisms in rare genetic disorders.

Humans

Isoform-Level Analysis Reveals Reproducible Early Changes in Transcript Usage During Human Vaccine Responses.

Vaccine-induced transcriptional responses have been extensively characterized at the gene level, but whether vaccination also alters transcript isoform usage remains largely unexplored. Here, we reanalyzed longitudinal whole-blood RNA-seq data from a discovery cohort of mRNA COVID-19 vaccine recipients using the IsoformSwitchAnalyzeR framework and validated the findings in an independent cohort. Key findings were validated by full-length RNA long-read sequencing and extended to four additional vaccine cohorts covering distinct platforms and pathogens. mRNA vaccination induced a rapid and transient wave of differential transcript usage, peaking at 24 h post-vaccination with 131 isoforms significantly altered across 107 genes, before largely resolving by Day 14. Isoform switching events were reproducible across independent cohorts and confirmed by full-length RNA long-read sequencing. Structural annotation of switching transcripts, including RMI2, WARS1, and NT5C3A, revealed changes affecting predicted protein domains and signal peptides. Notably, highly concordant isoform switching patterns were observed across MVA-based SARS-CoV-2, influenza, and Ebola vaccine cohorts and showed dose-dependent modulation. Overall, differential transcript isoform usage is a rapid and transient feature of the early human immune response to vaccination that was observed across multiple vaccine platforms. These findings reveal an underappreciated layer of transcriptional regulation that complements conventional gene-level analyses and warrants integration into future vaccine immunogenicity studies.

Humans

NextLongIso: a comprehensive Nextflow pipeline for multi-dimensional long-read RNA-seq analysis.

SUMMARY: Long-read RNA sequencing technologies, including Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT), enable direct characterization of full-length transcripts and transcriptome complexity. However, analysis of long-read RNA-seq data remains fragmented across multiple tools, limiting the ability to obtain a unified view of transcript structure, expression, and regulatory variation in long-read transcriptomes. We present NextLongIso, a scalable and reproducible Nextflow pipeline that enables coordinated analysis of multiple layers of transcript regulation. Rather than focusing solely on transcript reconstruction, NextLongIso integrates transcript discovery with downstream regulatory analyses to jointly characterize alternative splicing, isoform switching, transcript boundary dynamics (including alternative promoters and polyadenylation), and transposable element-associated transcription from both PacBio and ONT datasets. By eliminating complex cross-tool data harmonization, this unified framework facilitates the transition from transcript identification to functional interpretation of transcriptomic variation. AVAILABILITY AND IMPLEMENTATION: NextLongIso is implemented in Nextflow and is freely available at github: https://github.com/YidanSunResearchLab/nf-LongIso.git and Zenodo: https://doi.org/10.5281/zenodo.21049837.

Software

Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.

RNA splicing shapes neuronal identity and disease risk, yet current maps lack the developmental resolution and depth to resolve this complexity. Here, we integrate deep long-read RNA sequencing and proteomics in induced pluripotent stem cell-derived cortical neurons to generate a high-resolution proteogenomic atlas of human neuron development. We identify 182,371 mRNA isoforms (over half previously unknown) and provide direct peptide evidence for the translation of hundreds of novel protein-coding sequences. Population genetics demonstrates that variants affecting novel exons and splice sites are under negative selection, underscoring the potential significance of these isoforms. During neuronal maturation, we observe that autism risk genes undergo dynamic isoform switching, including microexon inclusion and intron retention, that remodel key protein domains and regulatory regions. Furthermore, we uncover widespread, long-range coordination between alternative transcript processing events, including transcription start sites, exon splicing, and polyadenylation. Finally, our atlas enables variant reinterpretation in autism, highlighting the value of an isoform-centric view for interpreting pathogenic variation in neurodevelopment.

Humans

Insights into the regulation of the HOTAIR proximal promoter.

HOTAIR (HOX transcript antisense RNA) is a HOXC-cluster long intervening non-coding RNA (lincRNA) whose cancer relevance is tightly coupled to how its transcription is wired into hormone, hypoxia, inflammatory, and developmental signaling. HOTAIR is known to associate with cancer cell proliferation, motility, tumor invasion, and metastasis. The present mini-review focuses on the regulatory architecture and mechanistic complexity of HOTAIR transcriptional regulation, with emphasis on three organizing principles. First, we consider the impact of promoter choice between a canonical proximal promoter (P1), which supports the 2.2-2.4 kb transcript, and an alternative upstream promoter/TSS (P2), which contributes to context-dependent transcription initiation. Second, we examine the long-distance enhancer-promoter communication between HOTAIR distal enhancer and P1/P2. Third, we summarize the recent epigenetic and epi-transcriptomic mechanisms involved in HOTAIR transcript initiation and elongation. A combination of these events determines isoform-specific transcription to govern cell-type-, context-, and cancer specific modulation of HOTAIR expression that promotes tumor formation and cancer progression. Finally, the review proposes how large-scale RNA datasets, long-read sequencing, and isoform-specific studies can refine our understanding of this versatile lincRNA's regulation.

Humans

Functional chimeric mRNAs encode proteins in mammalian immunity.

Individual mammalian mRNAs and proteins are typically believed to originate from single genomic loci, with isoform diversity arising through cis-splicing of pre-mRNA. Whether mRNA from distant genes can undergo trans-splicing to generate functionally relevant chimeric transcripts has remained unclear. Here we develop a pipeline combining long-read direct RNA sequencing with non-targeted and targeted validation to identify chimeric transcripts in macrophages. Chromatin conformation capture studies reveal that inflammation induces interchromosomal DNA interactions, positioning parent genes proximally to facilitate the formation of chimeric mRNA. Notably, we identify a protein-coding chimeric mRNA representing a fusion between the pore-forming protein gasdermin D (GSDMD)1,2 and a C-terminal domain translated out of frame from Tmem106a (Gsdmd-Tmem106a) in mice. We show that inflammasome priming upregulates Gsdmd-Tmem106a, with the protein localizing to the plasma membrane. After activation of the inflammasome, GSDMD-TMEM106A directly interacts with canonical GSDMD N termini to accelerate and enhance pore formation and IL-1β release. Finally, we show that GSDMD-TMEM106A balances host defence and immunopathology in vivo: its loss protects against lethal sepsis but compromises antibacterial defence, whereas overexpression enhances host protection while increasing sepsis lethality. We establish that protein-coding chimeric mRNAs formed by regulated transcript fusion events are operative during inflammation and immunity.

Journal Article

Beyond the gene: isoform diversity as a key contributor to human brain disorders.

The human brain exhibits exceptional transcriptomic complexity, with alternative splicing, promoter usage, and polyadenylation generating extensive transcript-isoform diversity. Isoform dysregulation is increasingly implicated in neurodevelopmental and psychiatric disorders (NPDs), yet the landscape, function, and genetic regulation of brain isoforms remain poorly understood due to limitations of short-read RNA sequencing. Advances in long-read sequencing (LR-seq) enable scalable full-length transcriptome profiling with single-cell and spatial resolution across developmental stages. Here, we review recent progress in isoform discovery, quantification, functional annotation, and genetic regulation, highlighting emerging links to human neurodevelopment and disease. LR-seq studies have uncovered tens of thousands of previously unannotated brain isoforms, with neuronal maturation characterized by increased exon inclusion and progressive 3' untranslated region (3' UTR) lengthening. Isoform-resolved genetic mapping outperforms gene-level analyses for NPD gene discovery and mechanistic interpretation. We argue that a shift from gene-centric to isoform-centric frameworks is essential to fully capture regulatory complexity in human neurogenetics. Together, these advances establish isoform diversity as a fundamental yet underappreciated axis of brain gene regulation and a key entry point for dissecting NPD biology.

Humans

New insights on Plasmodium gene expression from direct RNA sequencing.

Oxford Nanopore Technology (ONT) direct RNA sequencing enables the sequencing of native RNA molecules without cDNA conversion. The long-read approach captures full-length reads spanning entire genes and has transformed the study of gene expression in Plasmodium parasites by enabling analysis of untranslated regions, isoforms, and alternative splicing. In addition, ONT provides unique insights into non-coding RNAs, RNA modifications, and polyadenylated tail dynamics, which are expanding our understanding of post-transcriptional regulation in Plasmodium, including processes beyond translational repression in gametocytes and sporozoites. Here, we discuss the past and future applications of direct RNA sequencing in Plasmodium research and highlight its advantages, limitations, and future prospects.

Oxford Nanopore Technology

Elevated intron retention implicates neuroinflammation in brains of individuals with alcohol use disorder.

Intron retention, a form of alternative RNA splicing, can occur as part of normal gene regulation or result from disruption of the splicing machinery. Retained introns can potentially form double-stranded RNA, activating innate immune sensors and inflammation. This mechanism has been implicated in cancer but has not been studied in neuropsychiatric diseases like alcohol use disorder. We systematically analysed transcriptome-wide intron retention events in post-mortem brain tissue from 142 individuals (66 with alcohol use disorder and 76 controls), encompassing 320 region-specific samples from the superior frontal cortex, nucleus accumbens, central nucleus and basolateral amygdala. Analyses were adjusted for demographic, technical and biological covariates. Validation was performed in alcohol-preferring (P) rats using long-read sequencing. In complementary experiments, immunofluorescent staining was used to detect double-stranded RNA in rat brain tissue, while single-cell RNA-sequencing was performed to test activation of double-stranded RNA-sensing pathways in human brains. Brains from individuals with alcohol use disorder showed significantly higher total intron retention compared with controls, independent of age, with females showing greater increases than males. A total of 368 introns were positively associated with alcohol use disorder, and these introns were significantly longer and had weaker splice acceptor sites compared with non-associated introns. Genes harbouring these intron retention events were enriched in Purkinje neurons, visual cortex neurons and oligodendrocytes. Computational predictions indicated these long introns could form duplex RNA structures. Increased double-stranded RNA was confirmed experimentally in multiple brain regions of alcohol-consuming rats, where it co-localized primarily with neuronal nuclei and dendrites. In individuals with alcohol use disorder, we found that multiple pathways including double-stranded RNA responses, neuroinflammation, interferon and NF-κB signalling, adaptive immunity and apoptosis were activated. In addition, NeuN-positive neuronal counts significantly decreased in both the prefrontal and visual cortices. Furthermore, single-cell analysis demonstrated upregulation of TICAM1, the target of double-stranded RNA sensor TLR3, in oligodendrocytes, as well as widespread activation of downstream inflammatory pathways across glial and neuronal cell types. These findings provide the first evidence that chronic alcohol consumption promotes an overall increase of intron retention in the brain and is associated with the presence of double-stranded RNA. Furthermore, the double-stranded RNA may contribute to neuronal loss and brain pathology by activating a neuroinflammatory response.

alcohol use disorder

Full-length single-cell spatial transcriptomics reveals spatial and cell-type-specific transcript isoforms in the primate brain.

The primate brain exhibits complex RNA alternative splicing heterogeneity crucial for functional complexity, yet systematic spatial isoform characterization has been lacking. We developed Fullscope-seq, a full-length single-molecule large field-of-view spatial transcriptomics sequencing method at single-cell resolution, based on programmed concatenation cDNA for multiple long-read sequencing platforms. Applying Fullscope-seq to the macaque brain, we uncovered thousands of genes exhibiting differential transcript usage (DTU) across cortical layers, cell types and brain regions. Fullscope-seq resolved hundreds of major isoform switches across distinct brain regions and identified DTUs between superficial and deep cortical layers. Cortical layer-specific DTUs showed cell-composition dependence, whereas regional DTUs were regulated according to both cellular composition and spatial contexts. These isoform variations showed substantial enrichment for neuropsychiatric disorder-associated genes and were conserved across platforms and species. Our study establishes a scalable framework for spatial isoform analysis and provides a resource for understanding transcriptomic diversity in complex tissues.

Animals

Dysregulation of U12-Type Splicing in Lupus Neutrophils.

OBJECTIVE: Neutrophil dysfunction is a hallmark of systemic lupus erythematosus (SLE), but its molecular basis remains unclear. This study explores transcriptional and posttranscriptional changes in low-density granulocytes (LDGs), a proinflammatory neutrophil subset expanded in SLE, focusing on NADPH oxidase (Nox) function and minor intron splicing. METHODS: LDGs and normal-density granulocytes (NDGs) were isolated from patients with SLE and healthy controls (HCs). CYBA (p22phox) expression was evaluated at transcript and protein levels. Nox activity was measured using luminol assays. Bulk RNA sequencing (RNA-seq) and rMATS software were used to assess alternative splicing, particularly of U12-type intron-containing genes. RESULTS: CYBA expression was reduced in SLE LDGs (n = 11) compared to SLE and HC NDGs (n = 6), with levels resembling those in chronic granulomatous disease neutrophils. SLE LDGs exhibited impaired Nox activity (n = 7 SLE, n = 12 HC). CYBA is a U12 intron-containing gene, and transcriptomic analysis revealed broad down-regulation of this gene class in SLE LDGs, suggesting minor spliceosome dysfunction. rMATS analysis showed increased U12-type intron retention and widespread splicing defects-including exon skipping and mutually exclusive exon use-in genes such as GBP5, MAEA, and STX10. These abnormalities were validated in an independent long-read RNA-seq data set from SLE peripheral blood mononuclear cells. Importantly, splicing disruptions correlated with disease activity and autoantibody profiles. CONCLUSION: Impaired U12-dependent splicing may contribute to neutrophil dysfunction in SLE, potentially via defective oxidative burst and altered immune regulation. These findings highlight the minor spliceosome as a novel player in lupus pathogenesis.

Humans

Cancer-associated fusion transcripts: mechanisms, functional roles, and clinical implications.

Fusion transcripts are hybrid RNA molecules generated through genomic rearrangements or RNA-level fusion mechanisms. They represent important molecular features of many cancers and can function as oncogenic drivers, diagnostic biomarkers, prognostic indicators, and therapeutic targets. Since the discovery of the BCR::ABL1 fusion in chronic myeloid leukemia, numerous cancer-associated fusion transcripts have been identified across hematologic malignancies and solid tumors. These fusion events encompass diverse biological mechanisms, including constitutively active kinases, aberrant transcription factors, epigenetic regulators, and non-coding fusion RNAs. This review summarizes current knowledge of the mechanisms underlying fusion transcript formation, including genomic rearrangement-dependent and rearrangement-independent processes, as well as fusion circular RNAs. The functional roles of fusion transcripts in cancer biology and their clinical relevance as diagnostic, prognostic, and predictive biomarkers are discussed. In addition, recent advances in fusion transcript detection and characterization are reviewed, including next-generation sequencing, long-read sequencing, single-cell approaches, artificial intelligence-assisted computational methods, and CRISPR/Cas9-mediated strategies for functional modeling and functional validation of fusion transcripts. Despite the rapid expansion of fusion transcript catalogs, the biological and clinical significance of most identified fusion events remains incompletely understood. Future progress will depend on integrating advanced sequencing technologies, artificial intelligence-assisted computational prioritization, and systematic functional validation to distinguish clinically actionable fusion transcripts from biologically neutral events. Such multidisciplinary approaches will be essential for translating fusion transcript research into precision oncology and improving cancer diagnosis, patient stratification, and targeted therapy.

Humans

Single-cell multi-omics dissects transcript isoform and immune repertoire dynamics in human immunosenescence.

Immunosenescence, a major hallmark of systemic aging, refers to the progressive functional decline of the immune system. This decline not only compromises host defense and immunological memory but also fuels chronic inflammation and tissue degeneration (collectively known as inflammaging). While single-cell RNA sequencing (scRNA-seq) has revealed transcriptomic alterations associated with immune aging, analyses restricted to transcript abundance fail to capture deeper regulatory layers, such as transcript isoform diversity and the remodeling of immune receptor repertoires. To address this limitation, we present a human peripheral immune single-cell multi-omics atlas that integrates gene expression, transcript isoform diversity, and immune receptor repertoires. By combining single-cell full-length transcriptome sequencing (scCycloneSEQ), short-read scRNA-seq, and single-cell immune receptor sequencing (scTCR/BCR-seq), we systematically profiled peripheral blood mononuclear cells (PBMCs) from healthy donors aged 30-40 and 60-70 years. Our analyses uncovered extensive age-related remodeling of immune cell composition, functional states, and TCR/BCR diversity. Notably, we found that CD4+ effector memory T cells exhibited widespread differential isoform usage (DIU), 3'UTR length variation, and a marked reshaping of cytotoxic T lymphocyte (CTL) clonotypes-all of which were closely associated with aging-related inflammation and cellular senescence. This multi-omics atlas delineates key molecular features of immunosenescence and provides a high-resolution resource for deciphering the regulatory architecture underlying immune aging.

TCR/BCR

Paralogous evolution of the ITS2 region in Xiphophorus.

Ribosomal ITS2 is widely used in phylogenetic studies, yet its multigene organization and potential paralogy can obscure true species relationships. This proof-of-concept study investigates whether ITS2 sequences derived from long-read genomic data in multiple Xiphophorus species primarily reflect orthologous history or are shaped by ancient and local duplications. Phylogenetic analyses reveal two major, reciprocally mirroring ITS2 clades that represent long-standing paralogous rDNA lineages rather than simple allelic variants. The two paralogons show strong asymmetry in copy retention and loss for the majority of the species analyzed in this study. Exceptionally some other species are confined to one paralogon group and exhibit alternating ITS2 variants consistent with persistent ancestral polymorphism. A striking copy number imbalance in X. variatus, combined with its phylogenetic incongruence relative to the established species tree, is best explained by historical rDNA introgression followed by biased concerted evolution that nearly erased one paralogous copy. Despite incomplete homogenization, heterogeneous evolutionary rates, and occasional long-branch artifacts, the recovered paralog-specific topologies largely recapitulate the accepted Xiphophorus species phylogeny, indicating that ITS2 retains a robust organismal signal while also recording episodes of introgression and differential paralog evolution. These results demonstrate that explicit recognition of ITS2 paralogs can both improve phylogenetic interpretation and open avenues for future sequence-structure-based analyses of rDNA evolution and genus-level systematics in Xiphophorus.

Gene duplication

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Long-read sequencing reveals widespread novel splicing and neojunction-derived neoantigens in nasopharyngeal carcinoma.

The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; ∼44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.

Humans

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus