Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Long-read sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus↗

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant↗

New insights on Plasmodium gene expression from direct RNA sequencing.

Oxford Nanopore Technology (ONT) direct RNA sequencing enables the sequencing of native RNA molecules without cDNA conversion. The long-read approach captures full-length reads spanning entire genes and has transformed the study of gene expression in Plasmodium parasites by enabling analysis of untranslated regions, isoforms, and alternative splicing. In addition, ONT provides unique insights into non-coding RNAs, RNA modifications, and polyadenylated tail dynamics, which are expanding our understanding of post-transcriptional regulation in Plasmodium, including processes beyond translational repression in gametocytes and sporozoites. Here, we discuss the past and future applications of direct RNA sequencing in Plasmodium research and highlight its advantages, limitations, and future prospects.

Oxford Nanopore Technology↗

MHASS: Microbiome HiFi Amplicon Sequencing Simulator.

SUMMARY: Microbiome HiFi Amplicon Sequence Simulator (MHASS) creates realistic synthetic PacBio HiFi amplicon sequencing datasets for microbiome studies, by integrating genome-aware abundance modeling, realistic dual-barcoding strategies, and empirically derived pass-number distributions from actual sequencing runs. MHASS generates datasets tailored for rigorous benchmarking and validation of long-read microbiome analysis workflows, including ASV clustering and taxonomic assignment. AVAILABILITY AND IMPLEMENTATION: Implemented in Python with automated dependency management, the source code for MHASS is freely available at https://github.com/rhowardstone/MHASS along with installation instructions. Our code is also published on Zenodo at https://doi.org/10.5281/zenodo.17486364. The data underlying this article are available on GitHub at https://github.com/rhowardstone/MHASS_evaluation/.

Software↗

Complete genome sequence of the Anaplasma phagocytophilum clinical isolate NCH-1.

Anaplasma phagocytophilum is an obligate intracellular gram-negative bacterium and etiologic agent of human granulocytic anaplasmosis. A. phagocytophilum genomic sequencing has historically been performed via short-read platforms. Our optimized bacterial isolation protocol combined with Nanopore sequencing produced a single, closed 1,481,805 bp circular A. phagocytophilum strain NCH-1 chromosome.

Anaplasma phagocytophilum↗

A Step-by-Step Guide to Sequencing and Assembly of Complete Bacterial Genomes Using the Oxford Nanopore MinION.

The Oxford Nanopore (ONT) MinION enables sequencing of longer DNA/RNA fragments compared to other sequencers, such as Illumina, etc. This nanopore method provides distinct advantages for generating complete genome assemblies from microorganisms. Specifically, the R9.4 flow cells used for MinION sequencing have much lower error rates compared with earlier versions of the ONT platform. Coupled with base calling using Dorado software, higher-quality long reads can now be generated for complete bacterial genome assembly. In this chapter, we describe a detailed MinION method to assemble a complete genome from a microorganism, polish the final assembly, and evaluate the genome quality using various software tools. Because of the low cost for MinION sequencing, this platform could be an asset for virtually any laboratory interested in generating complete genomes from microorganisms.

Genome, Bacterial↗

Dynamic Centromeres Under Epigenetic Constraint.

Centromeres are essential chromosomal loci specified epigenetically by CENP-A chromatin, yet they undergo rapid sequence turnover, structural remodeling, and occasional repositioning. In this review, we integrate recent advances enabled by long-read genome assemblies and high-resolution chromatin mapping to synthesize current understanding of centromere organization across taxa. We examine how satellite repeats, transposable elements, molecular drive, and meiotic conflict generate extreme centromere diversity. We further explore how DNA methylation and H3K9me3 heterochromatin constrain CENP-A positioning, stabilize centromeric domains, and shape boundary dynamics during centromere drift, duplication, and de novo formation. Together, these perspectives show how centromeres accommodate evolutionary change while preserving the stringent requirements of faithful chromosome segregation.

Journal Article↗

Population-scale disease-associated tandem repeat analysis reveals locus and ancestry-specific insights.

Tandem repeat (TR) expansions, including short TRs (motifs ≤6 bp) and variable number TRs (motifs >6 bp), underlie many monogenic disorders, with variable length and sequence influencing pathogenicity, penetrance, severity, and onset. Accurate genotype-phenotype correlation and disease prevalence estimation require characterization beyond repeat length. Here we present a population-scale analysis of 66 disease-associated TR loci using long-read assemblies from 2530 diverse haplotypes from 1265 unaffected donors. Integrating repeat length, motif composition, local ancestry, linkage disequilibrium, and phylogenetic analyses, we reveal extensive locus-, population-, and allele-specific variation shaping disease risk. Up to 8.5% of individuals carry expansions above established pathogenic thresholds, many containing interrupting motifs or sequence structures that attenuate pathogenicity. After excluding alleles from loci with uncertain disease association, non-pathogenic interrupted expansions, and carrier states inconsistent with inheritance patterns, ~4% carried expansions predicted to confer disease risk, largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses uncover population-specific TR architectures contributing to epidemiological disparities in repeat expansion disorders. Phylogenetic analyses identify conserved ancestral alleles and loci with recent instability. We describe variable linkage disequilibrium patterns and recombination signatures around specific disease-associated TR loci. Our findings emphasize integrating sequence, ancestry, and evolutionary context to understand the complex landscape of disease-associated TRs.

Humans↗

Dogme: a nextflow pipeline for reprocessing nanopore RNA and DNA modifications.

MOTIVATION: Oxford Nanopore (ONT) sequencing allows for the direct detection of RNA and DNA modifications from unamplified nucleic acids, which is a significant advantage over other platforms. However, the rapid updates to ONT basecalling models and the evolving landscape of computational tools for modification detection bring about challenges for reproducible and standardized analyses. To address these challenges, we developed Dogme to automate basecalling, alignment, modification detection, and transcript quantification. Dogme automates the reprocessing of ONT POD5 files by integrating basecalling using Dorado, read mapping using minimap2 and subsequent analysis steps such as running modkit. The pipeline supports three major types of sequencing data-direct RNA (dRNA), complementary DNA (cDNA), and genomic DNA (gDNA). Dogme facilitates detection of diverse RNA modifications supported by Dorado such as N6-methyladenosine (m6A), 5-methylcytosine (m5C), inosine, pseudouridine, 2'-O-methylation (Nm) and DNA methylation, while concurrently quantifying full-length transcript isoforms LR-Kallisto for transcript quantification for dRNA and cDNA. RESULTS: We applied Dogme to three separate mouse C2C12 myoblast replicates using direct RNA sequencing on MinION flow cells. We detected 96 603 m6A, 43 476 m5C, 8829 inosine, 10 055 pseudouridine, and 30 320 Nm sites in three biological replicates. The pipeline produced reproducible modification profiles and transcript expression levels across replicates, demonstrating its utility for integrative long-read transcriptomic and epigenomic analyses. AVAILABILITY AND IMPLEMENTATION: Dogme is implemented in Nextflow and is freely available under the MIT license at https://github.com/mortazavilab/dogme, with documentation provided for installation and usage.

RNA↗

Global diversity and evolution of Salmonella enterica serovar Panama: a genomic epidemiology study.

BACKGROUND: Non-typhoidal Salmonella is a globally important bacterial pathogen, typically associated with foodborne gastrointestinal infection. Some non-typhoidal Salmonella serovars can also colonise typically sterile sites in people to cause invasive non-typhoidal Salmonella disease. Salmonella enterica serovar Panama is responsible for a substantial number of cases of human bloodstream infection, but despite its global dissemination, numerous outbreaks, and a reported association with invasive non-typhoidal Salmonella disease, S enterica serovar Panama (S Panama) is understudied. We aimed to describe the genomic epidemiology and evolutionary history of S Panama to provide a vital baseline of understanding for this globally important serovar. METHODS: In this genomic epidemiology study, we analysed S Panama genomes derived from historical collections, national surveillance datasets, and publicly available epidemiological and whole-genome sequencing data which span the years 1931-2019. Maximum likelihood and Bayesian phylodynamic approaches were used to investigate population structure and evolutionary history and to infer geotemporal dissemination. A combination of different bioinformatic approaches with short-read and long-read data were used to characterise geographical and clade-specific trends in antimicrobial resistance (AMR) and genetic markers for invasiveness. FINDINGS: We analysed 836 S Panama genomes, of which 559 (67%) were sequenced as part of this study. The collection represents all inhabited continents and includes isolates collected between 1931 and 2019. We identified the presence of four geographically linked S Panama clades (C1 [ie, the Latin America and the Caribbean clade; n=338], C2 [ie, the European clade; n=124], C3 [ie, the Martinique clade; n=131], and C4 [ie, the Asia and Oceania clade; n=104]) and regional trends in AMR profiles. Most isolates (715 [86%] of 836) were pan-susceptible to antibiotics and belonged to clades circulating in Latin America and the Caribbean (64%, n=458). Most antibiotic-resistant isolates in our collection (113 [93%] of 121) fell within clades C4 (ie, the Asia and Oceania clade) and C2 (ie, the European clade), the latter of which had the highest invasiveness index values based on the conservation of 196 extraintestinal predictor genes. INTERPRETATION: This first large-scale phylogenetic analysis of S Panama has revealed important information about the population structure, AMR, global ecology, and genetic markers of invasiveness of the identified genomic subtypes. Our findings provide an important baseline for understanding S Panama infection. The presence of multidrug-resistant clades with elevated invasiveness index values should be monitored through ongoing surveillance, as such clades could pose an increased public health risk. FUNDING: UK Research and Innovation Global Challenges Research Fund and Biotechnology and Biological Sciences Research Council, UK Medical Research Council, Wellcome Trust, John Lennon Memorial Scholarship, Institut Pasteur, Santé publique France, Fondation Le Roch-Les Mousquetaires, Investissement d'Avenir Programme, and Australian National Health and Medical Research Council.

Humans↗

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals↗

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software↗

Dysregulation of U12-Type Splicing in Lupus Neutrophils.

OBJECTIVE: Neutrophil dysfunction is a hallmark of systemic lupus erythematosus (SLE), but its molecular basis remains unclear. This study explores transcriptional and posttranscriptional changes in low-density granulocytes (LDGs), a proinflammatory neutrophil subset expanded in SLE, focusing on NADPH oxidase (Nox) function and minor intron splicing. METHODS: LDGs and normal-density granulocytes (NDGs) were isolated from patients with SLE and healthy controls (HCs). CYBA (p22phox) expression was evaluated at transcript and protein levels. Nox activity was measured using luminol assays. Bulk RNA sequencing (RNA-seq) and rMATS software were used to assess alternative splicing, particularly of U12-type intron-containing genes. RESULTS: CYBA expression was reduced in SLE LDGs (n = 11) compared to SLE and HC NDGs (n = 6), with levels resembling those in chronic granulomatous disease neutrophils. SLE LDGs exhibited impaired Nox activity (n = 7 SLE, n = 12 HC). CYBA is a U12 intron-containing gene, and transcriptomic analysis revealed broad down-regulation of this gene class in SLE LDGs, suggesting minor spliceosome dysfunction. rMATS analysis showed increased U12-type intron retention and widespread splicing defects-including exon skipping and mutually exclusive exon use-in genes such as GBP5, MAEA, and STX10. These abnormalities were validated in an independent long-read RNA-seq data set from SLE peripheral blood mononuclear cells. Importantly, splicing disruptions correlated with disease activity and autoantibody profiles. CONCLUSION: Impaired U12-dependent splicing may contribute to neutrophil dysfunction in SLE, potentially via defective oxidative burst and altered immune regulation. These findings highlight the minor spliceosome as a novel player in lupus pathogenesis.

Humans↗

Optimizing a culture-enriched hybrid metagenomics pipeline to assess the AMR footprint of livestock manure in anaerobic digestate.

The role of environmental samples from livestock production systems, including manure and anaerobic digestate, as reservoirs of antimicrobial resistance genes (ARGs) is likely underestimated because conventional metagenomic approaches can overlook low-abundance ARGs and often lack the resolution to associate these genes with their microbial hosts and co-localized mobile genetic elements (MGEs). We evaluated whether culture-enriched metagenomics (CEMG), with and without antibiotic selection, enhances ARG detection in anaerobic digestate and improves the resolution of ARG-MGE-host associations using hybrid short- and long-read metagenomic assembly. CEMG increased ARG recovery; mean ARG abundance rose from 15.4 counts per million (CPM) in metagenomic fresh digestate (FD) to 124 CPM in CEMG without antibiotics and 160 CPM in antibiotic-selective CEMG. In FD, only 9 unique ARGs were detected, whereas CEMG recovered 112, including ARGs of clinical importance, such as glycopeptide resistance, beta-lactamase genes, and the cfr 23S rRNA methyltransferase conferring cross-resistance to multiple antibiotic classes. Antibiotic selection induced targeted, class-specific shifts in ARG profiles, with ARGs associated with tetracycline resistance consistently enriched across treatments. Hybrid metagenomic assembly resolved the genomic context of 784 ARGs, of which 59.3% were co-localized with at least one class of MGEs, predominantly plasmids and integrative conjugative elements/integrative mobilizable elements. Biocide and metal resistance genes frequently co-occurred with ARGs on the same contigs. Together, these findings demonstrate that antibiotic-selective culture enrichment enhances resistome surveillance by improving detection of low-abundance ARGs, while hybrid assembly provides critical genomic context for assessing their mobility and host associations.IMPORTANCELivestock manure and its byproducts, such as anaerobic digestate, are recognized as important environmental reservoirs of antimicrobial resistance genes (ARGs) and resistant bacteria, yet current metagenomic approaches may underestimate this risk by failing to detect low-abundance but clinically relevant ARGs. Here, we show that integrating culture enrichment with hybrid metagenomics improves ARG recovery and reveals ARG co-localization with mobile genetic elements and putative bacterial hosts. This approach captures a cultivable and condition-responsive fraction of the resistome that is not readily accessible through direct metagenomic sequencing alone, providing a more informative framework for environmental AMR surveillance.

anaerobic digestion↗

Single-cell multi-omics dissects transcript isoform and immune repertoire dynamics in human immunosenescence.

Immunosenescence, a major hallmark of systemic aging, refers to the progressive functional decline of the immune system. This decline not only compromises host defense and immunological memory but also fuels chronic inflammation and tissue degeneration (collectively known as inflammaging). While single-cell RNA sequencing (scRNA-seq) has revealed transcriptomic alterations associated with immune aging, analyses restricted to transcript abundance fail to capture deeper regulatory layers, such as transcript isoform diversity and the remodeling of immune receptor repertoires. To address this limitation, we present a human peripheral immune single-cell multi-omics atlas that integrates gene expression, transcript isoform diversity, and immune receptor repertoires. By combining single-cell full-length transcriptome sequencing (scCycloneSEQ), short-read scRNA-seq, and single-cell immune receptor sequencing (scTCR/BCR-seq), we systematically profiled peripheral blood mononuclear cells (PBMCs) from healthy donors aged 30-40 and 60-70 years. Our analyses uncovered extensive age-related remodeling of immune cell composition, functional states, and TCR/BCR diversity. Notably, we found that CD4+ effector memory T cells exhibited widespread differential isoform usage (DIU), 3'UTR length variation, and a marked reshaping of cytotoxic T lymphocyte (CTL) clonotypes-all of which were closely associated with aging-related inflammation and cellular senescence. This multi-omics atlas delineates key molecular features of immunosenescence and provides a high-resolution resource for deciphering the regulatory architecture underlying immune aging.

TCR/BCR↗

Characterisation of Trichuris incognita n sp in Côte d'Ivoire: a morphological, genomic, and genome-wide association with drug sensitivity study.

BACKGROUND: Trichuriasis is a neglected tropical disease that affects up to 500 million individuals and can cause considerable morbidity. For decades, trichuriasis was thought to be caused by one species of whipworm, Trichuris trichiura. The aim of this study was to investigate the origin of differences in response rates to the best available anthelmintic treatment for trichuriasis-a combination of albendazole and ivermectin-in Côte d'Ivoire by analysing the parasite population. METHODS: In this morphological, genomic, and genome-wide association study (GWAS) with drug sensitivity we used long-read and short-read sequencing approaches and assembled a high-quality reference genome of Trichuris incognita n sp isolated in a primary interventional study conducted in the Lagunes district of Côte d'Ivoire. Children aged 6-12 years were screened between July 14, 2022, and July 31, 2022; children positive for T trichiura on duplicate Kato-Katz smears and with infection intensity of 200 eggs per gram or more were eligible and treated first with albendazole (400 mg) and ivermectin (200 μg/kg) then with oxantel pamoate (20 mg/kg). We constructed a species tree of the Trichuris genus using 12 434 orthologous groups. We sequenced individual worms, which were used to confirm the phylogenetic placement and investigate patterns of adaptation through comparative genomic analyses. Finally, we conducted a GWAS to compare albendazole-ivermectin sensitive worms to drug non-sensitive worms. FINDINGS: 670 children were screened, of whom 243 were enrolled and from whom 271 worms were isolated after the first treatment and 827 worms after the second treatment. Sufficient DNA was recovered from 747 worms of which 721 were suitable for further bioinformatic analysis; of these, 179 were albendazole-ivermectin sensitive worms and 542 were drug non-sensitive worms. We present and characterise a new, human-infecting Trichuris species named T incognita n sp, which is morphologically indistinguishable from T trichiura, but forms a distinct phylogenetic clade, closer to Trichuris suis than to the canonical human-infective T trichiura. Comparative genomic analysis of genes suspected to confer resistance to either albendazole or ivermectin in helminths revealed a high number of β-tubulin orthologs, present in the whole population of T incognita n sp, compared with the canonical T trichiura species, but these genes were not associated with a resistant phenotype. The GWAS did not provide conclusive evidence of adaptation to drug pressure within the same species. INTERPRETATION: Our results demonstrate that trichuriasis can be caused by multiple whipworm species, and that differences in response rates might result from species responding differently to drug treatment, rather than from the intraspecies establishment of resistance. This discovery, coupled with the high tolerability of T incognita n sp to albendazole-ivermectin, marks a substantial shift in how we understand and approach whipworm infections. FUNDING: European Research Council.

Trichuris↗

Community-driven updates for comprehensive long-read metagenomics and enhanced binning in nf-core/mag v5.

SUMMARY: nf-core/mag is a reproducible Nextflow pipeline for best-practice metagenomic de novo assembly and binning within the nf-core framework. Here we present a major update that adds support for long-read-only assembly and bin refinement, includes five new binning tools, expands taxonomic classification to viruses and eukaryotes, and improves bin quality evaluation with new tools and latest databases. Through sustained community-driven development spanning seven years and four primary curator teams, nf-core/mag remains actively developed as an open source workflow for metagenomic analysis, benefiting from contributions from across the broader metagenomics, nf-core, and Nextflow ecosystem. AVAILABILITY AND IMPLEMENTATION: The source code of nf-core/mag v5 is available on GitHub (https://github.com/nf-core/mag) under the open source MIT license, with v5.5.0 source code archived on Zenodo (https://zenodo.org/records/21735731). Documentation is viewable on the nf-core website (https://nf-co.re/mag).

Metagenomics↗

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals↗