Search PubMedSearch

SEARCH · Search PubMed

Results for “Long-read transcriptomics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Population-scale detection of methylation outliers from long-read genome sequencing.

BACKGROUND: Aberrant DNA methylation can mediate the functional effects of rare genetic variation and contribute to imprinting disorders, repeat expansion diseases, and other pathogenic regulatory mechanisms. Long-read sequencing technologies now enable genome-wide detection of CpG methylation alongside genetic variation from a single assay. However, methods for systematic identification and interpretation of methylation outliers from long-read sequencing data remain limited. METHODS: We developed METAFORA, a computational workflow for detecting methylation outlier regions from PacBio and Oxford Nanopore long-read sequencing data. METAFORA constructs population-level methylation references, segments the genome into correlated CpG blocks, infers technical and biological sources of variation through hidden factor estimation, models uncertainty due to variable depth sequencing, and computes covariate-adjusted methylation outlier scores for individual samples. We applied METAFORA across large long-read sequencing cohorts and integrated methylation outliers with multi-omic data. METAFORA is implemented as a snakemake workflow available at https://github.com/tjense25/METAFORA. RESULTS: METAFORA identified methylation outlier regions associated with rare structural variants, tandem repeat expansions, and imprinting abnormalities. We found outlier regions were enriched for molecular outliers across transcriptomic and chromatin accessibility datasets, supporting their functional relevance in gene regulation. In a representative case, METAFORA identified an imprinting defect affecting the GNAS locus associated with an STX16 deletion. CONCLUSIONS: METAFORA enables scalable detection and interpretation of methylation outliers from long-read sequencing data and provides a framework for integrating epigenetic outliers with genomic and multi-omic analyses. These approaches may improve interpretation of rare regulatory variation and support discovery of clinically relevant epigenetic abnormalities in genomic medicine.

DNA methylation

A novel allele of Sh1 facilitates the development of waxy-sweet corn from waxy corn.

Waxy corn and sweet corn represent 2 major classes of fresh-eating corn, each with distinct sensory attributes and nutritional compositions. Developing a new variety that combines both waxy and sweet traits would address rising consumer demand and expand new market potential. From a fast neutron-mutagenized population of the waxy corn inbred line HB522, we isolated a novel mutant, designated as wx-sweet, whose kernels simultaneously exhibit waxy and sweet characteristics at the milk-filling stage. Through bulked segregant analysis combined with fine mapping, we mapped the causal locus to SHRUNKEN1 (Sh1) on chromosome 9, which was confirmed by an allelism test with a characterized Mu-insertion allele of Sh1. A 7,227-bp Copia-type long terminal repeat retrotransposon insertion was identified in exon 2 of Sh1 in the wx-sweet mutant by long-read sequencing. Consistently, the novel sh1 allele significantly reduced sucrose synthase activity. Genetic and physiological analyses demonstrate that sh1 and wx1 act synergistically to fine-tune carbohydrate metabolism in the endosperm. Integrated transcriptomic and metabolomic profiling uncover extensive transcriptional reprogramming and redirected metabolic flux, leading to substantial accumulation of sucrose and a range of oligosaccharides. These metabolic shifts underlie the unique simultaneous dual waxy-sweet texture in fresh-eating wx-sweet kernels. In summary, our work not only provides valuable genetic resources for breeding next-generation fresh-eating corn but also, for the first time, elucidates the molecular mechanism by which the sh1 and wx1 mutations cooperatively shape the waxy-sweet endosperm phenotype.

Zea mays

The First Highly Contiguous Genome Assembly for the Western Bluebird (Sialia mexicana).

The western bluebird (Sialia mexicana) is a secondary cavity-nesting thrush that has experienced historical population declines, local extirpations, and more recent recoveries associated with nest box programs. Despite these regional successes, recent eBird estimates suggest continued range-wide declines and substantial geographic variation in population trajectories, making this species a useful system for future studies of demographic change, connectivity, and conservation genomics. However, genomic resources for western bluebirds remain limited, and no reference genome currently exists for any species in the genus Sialia. Here, we present the first high-quality de novo reference genome for S. mexicana. Using PacBio HiFi long-read sequencing from an adult female, we generated a highly contiguous, phased 1.3 Gb nuclear assembly with a contig N50 of 24.8 Mb and high BUSCO completeness of 98.3%. We annotated the nuclear genome using transcriptomic and protein evidence, identifying 16,656 protein-coding genes and 26,060 transcripts/protein isoforms. We also assembled a complete ∼16 kb mitochondrial genome from Illumina short-read data. This reference genome provides a foundational resource for future studies of population structure, genetic diversity, connectivity, demographic history, and adaptation in western bluebirds and related taxa.

Animals

Insights into the regulation of the HOTAIR proximal promoter.

HOTAIR (HOX transcript antisense RNA) is a HOXC-cluster long intervening non-coding RNA (lincRNA) whose cancer relevance is tightly coupled to how its transcription is wired into hormone, hypoxia, inflammatory, and developmental signaling. HOTAIR is known to associate with cancer cell proliferation, motility, tumor invasion, and metastasis. The present mini-review focuses on the regulatory architecture and mechanistic complexity of HOTAIR transcriptional regulation, with emphasis on three organizing principles. First, we consider the impact of promoter choice between a canonical proximal promoter (P1), which supports the 2.2-2.4 kb transcript, and an alternative upstream promoter/TSS (P2), which contributes to context-dependent transcription initiation. Second, we examine the long-distance enhancer-promoter communication between HOTAIR distal enhancer and P1/P2. Third, we summarize the recent epigenetic and epi-transcriptomic mechanisms involved in HOTAIR transcript initiation and elongation. A combination of these events determines isoform-specific transcription to govern cell-type-, context-, and cancer specific modulation of HOTAIR expression that promotes tumor formation and cancer progression. Finally, the review proposes how large-scale RNA datasets, long-read sequencing, and isoform-specific studies can refine our understanding of this versatile lincRNA's regulation.

Humans

Dysregulation of U12-Type Splicing in Lupus Neutrophils.

OBJECTIVE: Neutrophil dysfunction is a hallmark of systemic lupus erythematosus (SLE), but its molecular basis remains unclear. This study explores transcriptional and posttranscriptional changes in low-density granulocytes (LDGs), a proinflammatory neutrophil subset expanded in SLE, focusing on NADPH oxidase (Nox) function and minor intron splicing. METHODS: LDGs and normal-density granulocytes (NDGs) were isolated from patients with SLE and healthy controls (HCs). CYBA (p22phox) expression was evaluated at transcript and protein levels. Nox activity was measured using luminol assays. Bulk RNA sequencing (RNA-seq) and rMATS software were used to assess alternative splicing, particularly of U12-type intron-containing genes. RESULTS: CYBA expression was reduced in SLE LDGs (n = 11) compared to SLE and HC NDGs (n = 6), with levels resembling those in chronic granulomatous disease neutrophils. SLE LDGs exhibited impaired Nox activity (n = 7 SLE, n = 12 HC). CYBA is a U12 intron-containing gene, and transcriptomic analysis revealed broad down-regulation of this gene class in SLE LDGs, suggesting minor spliceosome dysfunction. rMATS analysis showed increased U12-type intron retention and widespread splicing defects-including exon skipping and mutually exclusive exon use-in genes such as GBP5, MAEA, and STX10. These abnormalities were validated in an independent long-read RNA-seq data set from SLE peripheral blood mononuclear cells. Importantly, splicing disruptions correlated with disease activity and autoantibody profiles. CONCLUSION: Impaired U12-dependent splicing may contribute to neutrophil dysfunction in SLE, potentially via defective oxidative burst and altered immune regulation. These findings highlight the minor spliceosome as a novel player in lupus pathogenesis.

Humans

Unraveling a Diagnostic Enigma: A TECPR2 Case Solved Through Multi-Omic Genomics.

TECPR2 is a key regulator of autophagy, encoded by the TECPR2 gene. Pathogenic variants in this gene have been linked to a rare hereditary sensory and autonomic neuropathy with intellectual disability (HSAN9). We report a teenage female with a syndromic intellectual disability disorder associated with neuromuscular abnormalities. Multi-omics analysis including genomics, transcriptomics, and proteomics, together with muscle biopsy from the affected individual, were used in this clinical case. Through trio exome sequencing we identified two heterozygous variants in the TECPR2 gene, NM_014844.4: c.480G>A; p.(Gln160=) and c.2846C>A; p.(Ala949Glu). Both were classified as variants of uncertain significance due to the lack of supporting evidence for pathogenicity. Subsequent long-read sequencing phased the variants and confirmed they were in trans. Additional functional studies using RNAseq and proteomics analyses verified the pathogenicity of the variants. This case study demonstrated the value of a multi-omics assisted analysis, which complemented the traditional phenotype-first approach in reaching a definitive clinical diagnosis.

Humans

Elevated intron retention implicates neuroinflammation in brains of individuals with alcohol use disorder.

Intron retention, a form of alternative RNA splicing, can occur as part of normal gene regulation or result from disruption of the splicing machinery. Retained introns can potentially form double-stranded RNA, activating innate immune sensors and inflammation. This mechanism has been implicated in cancer but has not been studied in neuropsychiatric diseases like alcohol use disorder. We systematically analysed transcriptome-wide intron retention events in post-mortem brain tissue from 142 individuals (66 with alcohol use disorder and 76 controls), encompassing 320 region-specific samples from the superior frontal cortex, nucleus accumbens, central nucleus and basolateral amygdala. Analyses were adjusted for demographic, technical and biological covariates. Validation was performed in alcohol-preferring (P) rats using long-read sequencing. In complementary experiments, immunofluorescent staining was used to detect double-stranded RNA in rat brain tissue, while single-cell RNA-sequencing was performed to test activation of double-stranded RNA-sensing pathways in human brains. Brains from individuals with alcohol use disorder showed significantly higher total intron retention compared with controls, independent of age, with females showing greater increases than males. A total of 368 introns were positively associated with alcohol use disorder, and these introns were significantly longer and had weaker splice acceptor sites compared with non-associated introns. Genes harbouring these intron retention events were enriched in Purkinje neurons, visual cortex neurons and oligodendrocytes. Computational predictions indicated these long introns could form duplex RNA structures. Increased double-stranded RNA was confirmed experimentally in multiple brain regions of alcohol-consuming rats, where it co-localized primarily with neuronal nuclei and dendrites. In individuals with alcohol use disorder, we found that multiple pathways including double-stranded RNA responses, neuroinflammation, interferon and NF-κB signalling, adaptive immunity and apoptosis were activated. In addition, NeuN-positive neuronal counts significantly decreased in both the prefrontal and visual cortices. Furthermore, single-cell analysis demonstrated upregulation of TICAM1, the target of double-stranded RNA sensor TLR3, in oligodendrocytes, as well as widespread activation of downstream inflammatory pathways across glial and neuronal cell types. These findings provide the first evidence that chronic alcohol consumption promotes an overall increase of intron retention in the brain and is associated with the presence of double-stranded RNA. Furthermore, the double-stranded RNA may contribute to neuronal loss and brain pathology by activating a neuroinflammatory response.

alcohol use disorder

Estimating protein isoform abundances with [Formula: see text].

A single gene can encode multiple versions of a protein, dubbed isoforms, with varying functionality. Cellular control of isoform abundances is critical for multiple aspects of biology and is only partially regulated by transcript levels. While long-read sequencing facilitates transcript quantification, quantifying the resulting protein isoforms on a large scale is a major challenge, complicating biological interpretation of transcript alterations. Standard "bottom up" mass spectrometry can assess only short portions of isoforms called peptides, and these peptides often map onto more than one isoform. We introduce [Formula: see text] (Protein isoform Abundance Quantification), a Bayesian method that leverages multiomic information from the peptidome and transcriptome to provide accurate estimates of isoform abundance even when peptide mapping is ambiguous. [Formula: see text] offers several advantages over existing methods in a unified framework. It provides uncertainty quantification, integrates multiomic information for improved accuracy, and provides a rigorous framework for hypothesis testing. Extensive simulations show that [Formula: see text] consistently outperforms competing methods in detecting differentially abundant protein isoforms and estimating their abundances. We use [Formula: see text] to investigate differences in isoform abundance levels between people with schizophrenia and control subjects, confirming a long-held hypothesis that levels of the C4A isoform of Complement Component 4 are increased in schizophrenia while C4B is not. These results demonstrate that [Formula: see text] can identify significant variations in isoform abundance levels not previously possible.

Protein Isoforms

scnanoseq: an nf-core pipeline for Oxford Nanopore single-cell RNA-sequencing.

MOTIVATION: Recent advancements in long-read single-cell RNA sequencing (scRNA-seq) have facilitated the quantification of full-length transcripts and isoforms at the single-cell level. Historically, long-read data would need to be complemented with short-read single-cell data in order to overcome the higher sequencing errors to correctly identify cellular barcodes and unique molecular identifiers. Improvements in Oxford Nanopore sequencing, and development of novel computational methods have removed this requirement. Though these methods now exist, the limited availability of modular and portable workflows remains a challenge. RESULTS: Here, we present, nf-core/scnanoseq, a secondary analysis pipeline for long-read single-cell and single-nuclei RNA that delivers gene and transcript-level quantification. The scnanoseq pipeline is implemented using Nextflow and is built upon the nf-core framework, enabling portability across computational environments, scalability and reproducibility of results across pipeline runs. The nf-core/scnanoseq workflow follows best practices for analyzing single-cell and single-nuclei data, performing barcode detection and correction, genome and transcriptome read alignment, unique molecular identifier deduplication, gene and transcript quantification, and extensive quality control reporting. AVAILABILITY AND IMPLEMENTATION: The source code, and detailed documentation are freely available at https://github.com/nf-core/scnanoseq and https://nf-co.re/scnanoseq under the MIT License. Documentation for the version of nf-core/scnanoseq used for this paper, including default parameters and descriptions of output files are available at https://nf-co.re/scnanoseq/1.1.0.

Single-Cell Analysis

Multiomics approaches to cardiovascular disease: technological innovations and clinical translation.

Cardiovascular diseases (CVDs) remain the leading cause of global morbidity and mortality, reflecting a persistent gap between clinical phenotyping and the molecular mechanisms that govern disease initiation, progression, and interindividual variability. Recent advances in emerging technologies have fundamentally reshaped cardiovascular physiology by enabling high-resolution, cross-layer profiling of the heart and vasculature across genomic, epigenomic, transcriptomic, proteomic, metabolomic, lipidomic, glycomic, and fluxomic layers, increasingly at single-cell and spatial resolution. These approaches reveal CVD as a coordinated, multilayered process driven by dynamic interactions among cell types, regulatory programs, and metabolic states, rather than isolated gene-level defects. In this review, we synthesize how emerging multiomic, computational, and functional genomic technologies are redefining the study of cardiovascular disease across molecular, cellular, and tissue levels. We highlight recent innovations in single-cell and spatial atlases, long-read sequencing, proteomics and metabolomics, integrative data modeling, and functional omics approaches, including genome-scale perturbation screens and single-cell perturbation frameworks. These platforms enable mechanistic dissection of regulatory circuits, distinguish primary disease drivers from secondary adaptations, and directly assess therapeutic reversibility, advancing the field beyond associative biomarker discovery toward mechanism-guided target prioritization. We further discuss key methodological and translational challenges accompanying high-dimensional cardiovascular data, including preanalytical variability, control selection, temporal misalignment across molecular layers, population diversity, and reference bias. By integrating technological innovation with computational rigor and functional validation, this review frames emerging omics-enabled strategies as a unified, physiologically grounded framework for translating molecular insight into clinically meaningful cardiovascular phenotypes and advancing precision cardiovascular medicine.

Humans

De Novo Assembly of the Trypanosoma congolense Genome Reveals an Organization Influenced by Antigenic Variation but Distinct from Trypanosoma brucei.

Antigenic variation allows pathogens to evade mammalian adaptive immunity through the continuous change in exposed antigens. In African trypanosomes, antigenic variation involves changes in expressed Variant Surface Glycoproteins (VSGs). Understanding of VSG expression control and change amongst African trypanosomes is most advanced in Trypanosoma brucei. In the important animal trypanosome, Trypanosoma congolense, incomplete genome assembly has held back understanding of the mechanics of antigenic variation. Here, we have used long-read DNA sequencing and Hi-C DNA interaction analysis to provide a telomere-to-telomere assembly of the T. congolense genome. This assembly reveals a genome comprising 12 diploid chromosomes, one tetraploid chromosome, and more than 100 small chromosomes. With this assembly we reveal several features of VSG organization and expression that differ from T. brucei. The majority of the T. congolense VSG archive, estimated at ∼1,500 genes, localizes to subtelomeres in 12 of the 13 large chromosomes, but these loci are notably smaller than are found in T. brucei. Furthermore, transcriptome analysis suggests expression of VSGs across the T. congolense subtelomeres, which are not separated within the nucleus from non-VSG chromosome regions, suggesting that there is no dedicated VSG expression site. Strikingly, one chromosome contains approximately 40% of the VSG archive and is largely transcriptionally silent, potentially acting as the major reservoir of new VSG variants. Finally, we show that VSG expression can be detected from multiple small chromosomes. In summary, the new genome assembly provides a platform for understanding a potentially unusual operation of VSG expression and switching in T. congolense.

Trypanosoma congolense

Promises and pitfalls of long-read sequencing for resolving microbial complexity.

Long-read sequencing (LRS) has driven a transition in microbial genomics, overcoming the assembly fragmentation inherent to short-read sequencing. This review elucidates the impact of LRS across isolate genomics, metagenomics, and multi-omics domains. By spanning extensive repetitive regions, LRS facilitates the reconstruction of circular chromosomes and precisely resolves mobile genetic elements (MGEs). In metagenomics, LRS enables strain-level resolution, the recovery of circular metagenome-assembled genomes, and the precise localization of MGEs within host replicons. Furthermore, the single-molecule, amplification-free properties of LRS provide enhanced resolution of native epigenetic modifications and full-length transcriptomes. Despite these advancements, widespread implementation remains constrained by multidimensional challenges, including stringent high-molecular-weight DNA requirements, depth deficits, and computational overhead. Nevertheless, LRS is increasingly becoming the method of choice for isolate genomics and metagenomics. As detection technologies and algorithms progress, LRS will further improve our ability to decipher the structural and functional diversity of microbial ecosystems.

Metagenomics

Isoform-Level Analysis Reveals Reproducible Early Changes in Transcript Usage During Human Vaccine Responses.

Vaccine-induced transcriptional responses have been extensively characterized at the gene level, but whether vaccination also alters transcript isoform usage remains largely unexplored. Here, we reanalyzed longitudinal whole-blood RNA-seq data from a discovery cohort of mRNA COVID-19 vaccine recipients using the IsoformSwitchAnalyzeR framework and validated the findings in an independent cohort. Key findings were validated by full-length RNA long-read sequencing and extended to four additional vaccine cohorts covering distinct platforms and pathogens. mRNA vaccination induced a rapid and transient wave of differential transcript usage, peaking at 24 h post-vaccination with 131 isoforms significantly altered across 107 genes, before largely resolving by Day 14. Isoform switching events were reproducible across independent cohorts and confirmed by full-length RNA long-read sequencing. Structural annotation of switching transcripts, including RMI2, WARS1, and NT5C3A, revealed changes affecting predicted protein domains and signal peptides. Notably, highly concordant isoform switching patterns were observed across MVA-based SARS-CoV-2, influenza, and Ebola vaccine cohorts and showed dose-dependent modulation. Overall, differential transcript isoform usage is a rapid and transient feature of the early human immune response to vaccination that was observed across multiple vaccine platforms. These findings reveal an underappreciated layer of transcriptional regulation that complements conventional gene-level analyses and warrants integration into future vaccine immunogenicity studies.

Humans

Single-cell multi-omics dissects transcript isoform and immune repertoire dynamics in human immunosenescence.

Immunosenescence, a major hallmark of systemic aging, refers to the progressive functional decline of the immune system. This decline not only compromises host defense and immunological memory but also fuels chronic inflammation and tissue degeneration (collectively known as inflammaging). While single-cell RNA sequencing (scRNA-seq) has revealed transcriptomic alterations associated with immune aging, analyses restricted to transcript abundance fail to capture deeper regulatory layers, such as transcript isoform diversity and the remodeling of immune receptor repertoires. To address this limitation, we present a human peripheral immune single-cell multi-omics atlas that integrates gene expression, transcript isoform diversity, and immune receptor repertoires. By combining single-cell full-length transcriptome sequencing (scCycloneSEQ), short-read scRNA-seq, and single-cell immune receptor sequencing (scTCR/BCR-seq), we systematically profiled peripheral blood mononuclear cells (PBMCs) from healthy donors aged 30-40 and 60-70 years. Our analyses uncovered extensive age-related remodeling of immune cell composition, functional states, and TCR/BCR diversity. Notably, we found that CD4+ effector memory T cells exhibited widespread differential isoform usage (DIU), 3'UTR length variation, and a marked reshaping of cytotoxic T lymphocyte (CTL) clonotypes-all of which were closely associated with aging-related inflammation and cellular senescence. This multi-omics atlas delineates key molecular features of immunosenescence and provides a high-resolution resource for deciphering the regulatory architecture underlying immune aging.

TCR/BCR

Chromosome-level genome assembly with telomeric repeats at scaffold ends for Rhabdosargus sarba.

Rhabdosargus sarba, the goldlined seabream, is a euryhaline marine fish of great aquaculture potential. Genome sequencing and assembly of R. sarba was carried utilizing a multi-platform sequencing strategy that included long-read sequencing (PacBio HiFi), short-read sequencing (Illumina), and chromatin interaction mapping (Hi-C). The final genome assembly size after scaffolding was 764.59 Mb in 31 scaffolds with an N50 length of 33.98 Mb. Repeat profiling of primary assembly showed that 28.71% of the genome comprises of repeat elements. Gene prediction utilising the evidence from ab initio prediction and transcriptome data revealed 26,913 protein encoding genes and functional annotation and pathway analysis showed their participation in 332 pathways. This genome is an excellent resource for future research on genetic improvement and molecular breeding programmes for R. sarba.

Animals

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43 Mb, with a Scaffold N50 length of 19.05 Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75 Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals