Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Long-read RNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article↗

Identification of Novel Wraparound Transcripts in JC Polyomavirus.

JC polyomavirus (JCPyV) is a ubiquitous pathogen that causes progressive multifocal leukoencephalopathy (PML). Although a recent study using next-generation sequencing (NGS) provided detailed transcriptome atlases for polyomaviruses (PyVs) such as BK polyomavirus and simian virus 40, the transcriptome of JCPyV remains poorly characterized. Here, we conducted a comprehensive analysis using both short-read and long-read NGS technologies to construct a transcriptome atlas of JCPyV. RNA extracted from IMR-32 and HEK293 cells transfected with the circular JCPyV genome was analyzed, leading to the identification of 39 previously uncharacterized viral transcripts in addition to 12 known ones. Among the novel transcripts, we identified wraparound transcripts, conserved across PyVs, which are generated through continuous, multicyclic transcription of the circular viral genome. These included both late transcripts containing leader-to-leader repeated sequences and SuperT transcripts with multiple LxCxE motifs. Notably, wraparound transcripts, including SuperT transcripts, were also detected in brain tissues from PML patients. Collectively, this study significantly expands our understanding of the JCPyV transcriptome, revealing the expression of wraparound transcripts in PML lesions. These findings provide valuable insights into the molecular basis of JCPyV gene expression and PML pathogenesis, potentially facilitating the development of effective countermeasures against PML.

JC Virus↗

Long-read sequencing to interrogate strain-level variation among adherent-invasive Escherichia coli isolated from human intestinal tissue.

Adherent-invasive Escherichia coli (AIEC) is a pathovar linked to inflammatory bowel diseases (IBD), especially Crohn's disease, and colorectal cancer. AIEC are genetically diverse, and in the absence of a universal molecular signature, are defined by in vitro functional attributes. The relative ability of difference AIEC strains to colonize, persist, and induce inflammation in an IBD-susceptible host is unresolved. To evaluate strain-level variation among tissue-associated E. coli in the intestines, we develop a long-read sequencing approach to identify AIEC by strain that excludes host DNA. We use this approach to distinguish genetically similar strains and assess their fitness in colonizing the intestine. Here we have assembled complete genomes using long-read nanopore sequencing for a model AIEC strain, NC101, and seven strains isolated from the intestinal mucosa of Crohn's disease and non-Crohn's tissues. We show these strains can colonize the intestine of IBD susceptible mice and induce inflammatory cytokines from cultured macrophages. We demonstrate that these strains can be quantified and distinguished in the presence of 99.5% mammalian DNA and from within a fecal population. Analysis of global genomic structure and specific sequence variation within the ribosomal RNA operon provides a framework for efficiently tracking strain-level variation of closely-related E. coli and likely other commensal/pathogenic bacteria impacting intestinal inflammation in experimental settings and IBD patients.

Animals↗

How advances in machine learning drive early detection and risk prediction of early-onset colorectal cancer.

Early-onset colorectal cancer (EOCRC), defined as colorectal cancer diagnosed before age 50, is rising across high- and middle-income settings whilst organised screening stays anchored to older age thresholds. Blood-based liquid biopsy, combined with machine learning, is the most plausible route to early detection in this group because it does not depend on bowel preparation, endoscopy capacity, or adherence to stool-based testing. The gap is structural: incidence climbs fastest in the population below the age at which any guideline-endorsed modality is offered. The analytical challenge is that early-stage tumour-derived signals in plasma are low in abundance and distributed across heterogeneous molecular layers: circulating tumour DNA mutations, aberrant methylation, cfDNA fragmentomics, and small non-coding RNA. Machine learning converts these into a single calibrated probability. This review examines where artificial intelligence (AI)-driven liquid biopsy genuinely adds diagnostic value in EOCRC, distinguishes components in which learned models are decorative from those in which they are mechanistically necessary, and identifies the validation deficit separating research cohorts from deployable clinical tools. It summarises the first-generation tools used clinically for early detection and post-treatment monitoring, then considers analytes from exosome-bound microRNAs to long-read whole-genome sequencing of circulating plasma DNA, which reads cytosine modification natively, resolves methylation and fragmentation on single molecules, and characterises structural events short reads cannot anchor. Any analyte can feed a learned model, but more diverse input yields better discrimination. The central argument is that approved, guideline-included blood tests were validated in populations aged 45 and above, and their performance in younger patients cannot be assumed.

cfDNA fragmentomics↗

Long-read direct infrared sequencing of crude PCR products for prediction of resistance to HIV-1 reverse transcriptase and protease inhibitors.

Patients infected with human immunodeficiency virus type 1 (HIV-1) are being treated with a number of different combinations of antiretroviral compounds that target the essential viral enzymes reverse transcriptase and protease. Different sets of HIV-1 mutations that confer drug resistance have been well defined; they allow reasonable prediction of the drug sensitivity pattern from analysis of the HIV-1 genotype in vivo. Since periodical monitoring of genotypic resistance is expected to improve clinical management in a large number of infected patients, practical and cost-effective methods are highly desirable to set at least medium-scale sequencing in clinical diagnostic settings. We present a complete protocol for direct sequencing of HIV-1 reverse transcriptase and protease-coding regions. Features making the system amenable to routine clinical use include: 1. Highly robust presequencing steps (plasma RNA extraction, reverse transcription, and nested PCR); 2. Direct use of the crude unpurified PCR product as the sequencing template; and 3. Use of infrared-labeled sequencing primers consistently allowing long reads, thus obviating the need for sequencing of both DNA strands.

Base Sequence↗

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans↗

Dysregulation of U12-Type Splicing in Lupus Neutrophils.

OBJECTIVE: Neutrophil dysfunction is a hallmark of systemic lupus erythematosus (SLE), but its molecular basis remains unclear. This study explores transcriptional and posttranscriptional changes in low-density granulocytes (LDGs), a proinflammatory neutrophil subset expanded in SLE, focusing on NADPH oxidase (Nox) function and minor intron splicing. METHODS: LDGs and normal-density granulocytes (NDGs) were isolated from patients with SLE and healthy controls (HCs). CYBA (p22phox) expression was evaluated at transcript and protein levels. Nox activity was measured using luminol assays. Bulk RNA sequencing (RNA-seq) and rMATS software were used to assess alternative splicing, particularly of U12-type intron-containing genes. RESULTS: CYBA expression was reduced in SLE LDGs (n = 11) compared to SLE and HC NDGs (n = 6), with levels resembling those in chronic granulomatous disease neutrophils. SLE LDGs exhibited impaired Nox activity (n = 7 SLE, n = 12 HC). CYBA is a U12 intron-containing gene, and transcriptomic analysis revealed broad down-regulation of this gene class in SLE LDGs, suggesting minor spliceosome dysfunction. rMATS analysis showed increased U12-type intron retention and widespread splicing defects-including exon skipping and mutually exclusive exon use-in genes such as GBP5, MAEA, and STX10. These abnormalities were validated in an independent long-read RNA-seq data set from SLE peripheral blood mononuclear cells. Importantly, splicing disruptions correlated with disease activity and autoantibody profiles. CONCLUSION: Impaired U12-dependent splicing may contribute to neutrophil dysfunction in SLE, potentially via defective oxidative burst and altered immune regulation. These findings highlight the minor spliceosome as a novel player in lupus pathogenesis.

Humans↗

Cancer-associated fusion transcripts: mechanisms, functional roles, and clinical implications.

Fusion transcripts are hybrid RNA molecules generated through genomic rearrangements or RNA-level fusion mechanisms. They represent important molecular features of many cancers and can function as oncogenic drivers, diagnostic biomarkers, prognostic indicators, and therapeutic targets. Since the discovery of the BCR::ABL1 fusion in chronic myeloid leukemia, numerous cancer-associated fusion transcripts have been identified across hematologic malignancies and solid tumors. These fusion events encompass diverse biological mechanisms, including constitutively active kinases, aberrant transcription factors, epigenetic regulators, and non-coding fusion RNAs. This review summarizes current knowledge of the mechanisms underlying fusion transcript formation, including genomic rearrangement-dependent and rearrangement-independent processes, as well as fusion circular RNAs. The functional roles of fusion transcripts in cancer biology and their clinical relevance as diagnostic, prognostic, and predictive biomarkers are discussed. In addition, recent advances in fusion transcript detection and characterization are reviewed, including next-generation sequencing, long-read sequencing, single-cell approaches, artificial intelligence-assisted computational methods, and CRISPR/Cas9-mediated strategies for functional modeling and functional validation of fusion transcripts. Despite the rapid expansion of fusion transcript catalogs, the biological and clinical significance of most identified fusion events remains incompletely understood. Future progress will depend on integrating advanced sequencing technologies, artificial intelligence-assisted computational prioritization, and systematic functional validation to distinguish clinically actionable fusion transcripts from biologically neutral events. Such multidisciplinary approaches will be essential for translating fusion transcript research into precision oncology and improving cancer diagnosis, patient stratification, and targeted therapy.

Humans↗

Single-cell multi-omics dissects transcript isoform and immune repertoire dynamics in human immunosenescence.

Immunosenescence, a major hallmark of systemic aging, refers to the progressive functional decline of the immune system. This decline not only compromises host defense and immunological memory but also fuels chronic inflammation and tissue degeneration (collectively known as inflammaging). While single-cell RNA sequencing (scRNA-seq) has revealed transcriptomic alterations associated with immune aging, analyses restricted to transcript abundance fail to capture deeper regulatory layers, such as transcript isoform diversity and the remodeling of immune receptor repertoires. To address this limitation, we present a human peripheral immune single-cell multi-omics atlas that integrates gene expression, transcript isoform diversity, and immune receptor repertoires. By combining single-cell full-length transcriptome sequencing (scCycloneSEQ), short-read scRNA-seq, and single-cell immune receptor sequencing (scTCR/BCR-seq), we systematically profiled peripheral blood mononuclear cells (PBMCs) from healthy donors aged 30-40 and 60-70 years. Our analyses uncovered extensive age-related remodeling of immune cell composition, functional states, and TCR/BCR diversity. Notably, we found that CD4+ effector memory T cells exhibited widespread differential isoform usage (DIU), 3'UTR length variation, and a marked reshaping of cytotoxic T lymphocyte (CTL) clonotypes-all of which were closely associated with aging-related inflammation and cellular senescence. This multi-omics atlas delineates key molecular features of immunosenescence and provides a high-resolution resource for deciphering the regulatory architecture underlying immune aging.

TCR/BCR↗

Enhanced CRISPR-Cas3-mediated genome editing using circularized crRNAs.

Type I-E CRISPR-Cas3 represents a genome-editing technology in which large deletions averaging several kilobases are introduced in target regions. However, its genome-editing efficiency varies considerably across targets and cell types, making it difficult to achieve consistent results. Here, we investigated the efficacy and stability of circularized CRISPR RNAs (ccrRNAs) to enhance CRISPR-Cas3-mediated genome editing in human cells. Using in vitro single-strand DNA cleavage assays, we demonstrated that ccrRNA induces Cascade complex formation. Significant genome-editing activity targeting the EMX1 and B2M genes was observed in cellular assays using K562 cells. Long-read sequencing identified large-scale deletion mutations at the target loci and no detectable off-target effects using ccrRNA. Furthermore, ccrRNAs exhibited extended intracellular stability compared with that for linear crRNAs, resulting in an enhanced editing efficiency. These results demonstrate that ccrRNAs enable stable, efficient, and highly specific genome editing and support the broader application of the long-range deletion system.

CRISPR-Cas3↗

Paralogous evolution of the ITS2 region in Xiphophorus.

Ribosomal ITS2 is widely used in phylogenetic studies, yet its multigene organization and potential paralogy can obscure true species relationships. This proof-of-concept study investigates whether ITS2 sequences derived from long-read genomic data in multiple Xiphophorus species primarily reflect orthologous history or are shaped by ancient and local duplications. Phylogenetic analyses reveal two major, reciprocally mirroring ITS2 clades that represent long-standing paralogous rDNA lineages rather than simple allelic variants. The two paralogons show strong asymmetry in copy retention and loss for the majority of the species analyzed in this study. Exceptionally some other species are confined to one paralogon group and exhibit alternating ITS2 variants consistent with persistent ancestral polymorphism. A striking copy number imbalance in X. variatus, combined with its phylogenetic incongruence relative to the established species tree, is best explained by historical rDNA introgression followed by biased concerted evolution that nearly erased one paralogous copy. Despite incomplete homogenization, heterogeneous evolutionary rates, and occasional long-branch artifacts, the recovered paralog-specific topologies largely recapitulate the accepted Xiphophorus species phylogeny, indicating that ITS2 retains a robust organismal signal while also recording episodes of introgression and differential paralog evolution. These results demonstrate that explicit recognition of ITS2 paralogs can both improve phylogenetic interpretation and open avenues for future sequence-structure-based analyses of rDNA evolution and genus-level systematics in Xiphophorus.

Gene duplication↗

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans↗

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms↗

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals↗

Long-read sequencing reveals widespread novel splicing and neojunction-derived neoantigens in nasopharyngeal carcinoma.

The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; ∼44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.

Humans↗

A transposable element insertion in AUX/IAA16 disrupts splicing and causes auxin resistance in Bassia scoparia.

A dicamba-resistant population of kochia (Bassia scoparia) identified in Colorado, USA in 2012 was used to generate a synthetic mapping population that segregated for dicamba resistance. Linkage mapping associating dicamba injury with genotype derived from restriction-site-associated DNA sequencing identified a single locus in the kochia genome associated with resistance on chromosome 4. A mutant version of Auxin/Indole-3-Acetic Acid 16 (AUX/IAA16; a gene previously implicated in dicamba resistance in kochia) was found near the middle of this locus in resistant plants. Long-read sequencing of dicamba-resistant plants identified a recently inserted long-terminal repeat (LTR) retrotransposon TRIM element near the beginning of the second exon of AUX/IAA16, leading to disruption of normal splicing and a mutated degron domain. Stable transgenic lines of Arabidopsis thaliana ectopically expressing the mutant and wild-type alleles of AUX/IAA16 were developed. Arabidopsis thaliana plants expressing the mutant AUX/IAA16 allele grew shorter roots on control media. However, transgenic root growth was less inhibited on media containing either dicamba (5 μM) or IAA (0.5 μM) when compared with non-transgenic plants or those expressing the wild-type allele of AUX/IAA16. In vitro assays indicate reduced binding affinity and more rapid dissociation of the mutant AUX/IAA16 with TIR1 in the presence of several auxins, and protein modeling suggests the substitution of the glycine residue in the degron domain of AUX/IAA16 is especially important for resistance. A fitness cost associated with the mutant allele of AUX/IAA16 has implications for resistance evolution and management of kochia populations with this resistance mechanism.

Indoleacetic Acids↗

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus↗

Characterization and analysis of the full-length transcriptome of Frankliniella occidentalis (Thysanoptera: Thripidae).

BACKGROUND: Frankliniella occidentalis, an insect belonging to the order Thysanoptera, causes severe damage to agricultural and horticultural crops, resulting in significant economic losses worldwide. The development of molecular and sequencing technologies has helped elucidate the molecular mechanisms regulating its growth and development as well as its damaging activity. However, much remains to be explored. To further investigate the molecular complexity of this species, we sequenced the full-length transcriptome of mixed samples obtained from specimens at all developmental stages. RESULTS: Of all transcripts, 89.04% matched with the reference genome; additionally, 29,750 alternative splicing events, 2,342 genes with poly(A) sites, and 153 candidate fusion transcript events were identified, and 4,235 long noncoding RNAs were discovered. CONCLUSIONS: This is the first full-length transcriptome of F. occidentalis reported to date. This study greatly contributes to the understanding of the molecular complexity and diversity of this insect, providing a basis to develop specific molecular targets as well as resources for gene function studies in other insects.

Animals↗