Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Long-read assemblers”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Community-driven updates for comprehensive long-read metagenomics and enhanced binning in nf-core/mag v5.

SUMMARY: nf-core/mag is a reproducible Nextflow pipeline for best-practice metagenomic de novo assembly and binning within the nf-core framework. Here we present a major update that adds support for long-read-only assembly and bin refinement, includes five new binning tools, expands taxonomic classification to viruses and eukaryotes, and improves bin quality evaluation with new tools and latest databases. Through sustained community-driven development spanning seven years and four primary curator teams, nf-core/mag remains actively developed as an open source workflow for metagenomic analysis, benefiting from contributions from across the broader metagenomics, nf-core, and Nextflow ecosystem. AVAILABILITY AND IMPLEMENTATION: The source code of nf-core/mag v5 is available on GitHub (https://github.com/nf-core/mag) under the open source MIT license, with v5.5.0 source code archived on Zenodo (https://zenodo.org/records/21735731). Documentation is viewable on the nf-core website (https://nf-co.re/mag).

Metagenomics↗

Chromosome-Level Reference Genome of the Desert Night Lizard Xantusia vigilis.

We present a reference-quality genome assembly for the desert night lizard (Xantusia vigilis). The night lizards (Xantusiidae) are a family of small-bodied lizards found in North America (Xantusia), Central America (Lepidophyma), and Cuba (Cricosaura). The night lizard family has an independent evolutionary history of at least 80 million years from its sister taxa within Scincoidea. The Xantusiids have several unique ecological, behavioral and evolutionary characteristics. For instance, the family contains the only squamate species that form diploid, unisexual, parthenogenic lineages. In addition, most night lizards are viviparous and form stable kin groups that are maintained over multiple years, an unusual life history strategy among lizards. Combining PacBio long-read sequencing, Hi-C, and RNAseq data we developed a reference-quality genome for the desert night lizard, X. vigilis. We assembled a complete mitochondrion and ~ 2.2 Gb nuclear genome, with 20 scaffolds that correlate in size to the X. vigilis karyotype. In addition, we found that X. vigilis chromosome 1 aligns with gene content of both of macrochromosome 1 and microchromosome 9 from a genome assembly of a species in the sister family Cordylidae (Hemicordylus capensis).

Xantusia↗

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii↗

Chromosome-level genome assembly of Ampulex clypecomplana Chen & Li (Hymenoptera: Ampulicidae).

Ampulex clypecomplana Chen & Li, 2010 (Hymenoptera: Ampulicidae) is an important predatory insect in Hymenoptera. However, molecular information about this predatory insect is currently limited. In this study, we employed ONT long-read sequencing, MGI-SEQ short-read sequencing, Hi-C sequencing and transcriptomic data to assemble the high-quality genome of A. clypecomplana. The genome assembly length was 338.43 Mb, with a Scaffold N50 length of 19.05 Mb. Our BUSCO analysis further confirmed the gene coverage completeness of the genome assembly to be 99.2%. Phylogenetic analysis indicated that A. clypecomplana appeared approximately 132 million years ago. We annotated 110.75 Mb of repetitive sequences, accounting for 32.72% of the entire genome. In A. clypecomplana, we identified 180 gene expansions and 1029 genes that underwent contraction or loss. The high-quality genome of A. clypecomplana provides a valuable genetic resource for future research in evolution, molecular biology, and applied studies.

Animals↗

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, Hi-C, and karyotyping. These data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic variants. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Humans↗

Development and extensive sequencing of a broadly-consented Genome in a Bottle matched tumor-normal pair.

The Genome in a Bottle Consortium (GIAB), hosted by the National Institute of Standards and Technology (NIST), is developing new matched tumor-normal samples, the first to be explicitly consented for public dissemination of genomic data and cell lines. Here, we describe a comprehensive genomic dataset from the first individual, HG008, including DNA from an adherent, epithelial-like pancreatic ductal adenocarcinoma (PDAC) tumor cell line and matched normal cells from duodenal and pancreatic tissues. Data for the tumor-normal matched samples comes from seventeen distinct state-of-the-art whole genome measurement technologies, including high depth short and long-read bulk whole genome sequencing (WGS), single cell WGS, and Hi-C, and karyotyping. In future publications, these data will be used by the GIAB Consortium to develop matched tumor-normal benchmarks for somatic variant detection. We expect these data to facilitate innovation for whole genome measurement technologies, de novo assembly of tumor and normal genomes, and bioinformatic tools to identify small and structural somatic mutations. This first-of-its-kind broadly consented open-access resource will facilitate further understanding of sequencing methods used for cancer biology.

Journal Article↗

Genome evolution and long-term demographic history in true crocodiles.

Reference-quality genomes remain scarce for true crocodiles (Crocodylus), limiting comparative analyses of genome evolution and demographic history. Here, we generated and analyzed 2 long-read genomes, 1 for Crocodylus intermedius and 1 for C. niloticus, to investigate genome architecture, coalescent effective population size (Ne), and patterns of molecular evolution across crocodilians. Comparative analyses revealed broadly similar repeat landscapes in both species and extensive macro-synteny with Alligator sinensis, indicating strong structural conservation across crocodilian genomes. Using phased diploid assemblies and MSMC2, we reconstructed historical Ne trajectories and found marked differences between species. Crocodylus intermedius exhibited persistently low Ne throughout most of the late Quaternary, with a pronounced decline during the Late Pleistocene-early Holocene transition. In contrast, C. niloticus showed substantially larger Ne over comparable time intervals. Genome-wide codon-based analyses identified significant heterogeneity in dN/dS (ω) among crocodilian lineages. Crocodylus niloticus showed the lowest genome-wide ω, whereas elevated values in C. intermedius and other lineages were consistent with reduced long-term efficacy of purifying selection under smaller historical population sizes. Branch-site tests identified candidate genes under positive selection in both focal species, with functional categories related to ion transport, endocrine regulation, and cellular signaling. Together, these results provide genomic resources for Crocodylus and support an association between long-term demographic history and genome-wide patterns of molecular evolution across crocodilians.

Animals↗

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals↗

Nallo: a Nextflow pipeline for comprehensive human long-read genome analysis.

MOTIVATION: Long-read sequencing (LRS) is increasingly used for human medical research and clinical diagnostics due to its capacity to generate complete genome information. However, there is a lack of robust and easy-to-use pipelines for comprehensive LRS data analysis. RESULTS: Here we present Nallo, a Nextflow pipeline for analysis of PacBio and Oxford Nanopore data, with additional support for rare disease research projects. The pipeline detects a wide range of genetic variants, performs genome assembly, and reports CpG methylation. It also enables annotation and ranking of variants based on their predicted functional consequences. AVAILABILITY AND IMPLEMENTATION: Nallo is available from GitHub: https://github.com/genomic-medicine-sweden/nallo.

Humans↗

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

Ribosomal DNA and RNA (rDNA and rRNA) sequences are usually discarded from sequencing analyses. But with hundreds of copies of rDNA genes it is unknown whether they possess sequence variations that form different types of ribosomes that affect human physiology and disease. Here, we developed an algorithm for variant-calling between paralog genes (termed RGA) and compared rDNA variations found in short- and long-read sequencing data from the 1,000 Genomes Project (1KGP) and Genome In A Bottle (GIAB). We additionally developed a novel protocol for long-read sequencing full-length rRNA (RIBO-RT) from actively translating ribosomes. Our analyses identified hundreds of rDNA variants, most of which, surprisingly, are short insertion-deletions (indels) and dozens of highly abundant rRNA variants that are incorporated into translationally active ribosomes. To visualize variant ribosomes at the single cell level, we developed an in-situ rRNA sequencing method (SWITCH-seq) which revealed that variants are co-expressed within individual cells. Strikingly, by analyzing rDNA, we found that variants assemble into distinct ribosome subtypes. We discovered that these subtypes acquire different rRNA structures by successfully employing dimethyl sulfate (DMS) probing of full length rRNA. With this atlas we investigated rRNA variation changes across human tissues and cancer types. This revealed tissue-specific rRNA subtype expression in endoderm/ectoderm-derived tissues. In cancer, low abundant rRNA variants can become highly expressed, which suggests the presence of cancer-specific ribosomes. Together, this study identifies and comprehensively characterizes the diversity of ribosomes at the level of rRNA variants which is dominated by indel variants, their chromosomal location and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Journal Article↗

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5&#xa0;Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis↗

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans↗

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics↗

Characterisation of Trichuris incognita n sp in C&#xf4;te d'Ivoire: a morphological, genomic, and genome-wide association with drug sensitivity study.

BACKGROUND: Trichuriasis is a neglected tropical disease that affects up to 500 million individuals and can cause considerable morbidity. For decades, trichuriasis was thought to be caused by one species of whipworm, Trichuris trichiura. The aim of this study was to investigate the origin of differences in response rates to the best available anthelmintic treatment for trichuriasis-a combination of albendazole and ivermectin-in C&#xf4;te d'Ivoire by analysing the parasite population. METHODS: In this morphological, genomic, and genome-wide association study (GWAS) with drug sensitivity we used long-read and short-read sequencing approaches and assembled a high-quality reference genome of Trichuris incognita n sp isolated in a primary interventional study conducted in the Lagunes district of C&#xf4;te d'Ivoire. Children aged 6-12 years were screened between July 14, 2022, and July 31, 2022; children positive for T trichiura on duplicate Kato-Katz smears and with infection intensity of 200 eggs per gram or more were eligible and treated first with albendazole (400 mg) and ivermectin (200 &#x3bc;g/kg) then with oxantel pamoate (20 mg/kg). We constructed a species tree of the Trichuris genus using 12&#x2009;434 orthologous groups. We sequenced individual worms, which were used to confirm the phylogenetic placement and investigate patterns of adaptation through comparative genomic analyses. Finally, we conducted a GWAS to compare albendazole-ivermectin sensitive worms to drug non-sensitive worms. FINDINGS: 670 children were screened, of whom 243 were enrolled and from whom 271 worms were isolated after the first treatment and 827 worms after the second treatment. Sufficient DNA was recovered from 747 worms of which 721 were suitable for further bioinformatic analysis; of these, 179 were albendazole-ivermectin sensitive worms and 542 were drug non-sensitive worms. We present and characterise a new, human-infecting Trichuris species named T incognita n sp, which is morphologically indistinguishable from T trichiura, but forms a distinct phylogenetic clade, closer to Trichuris suis than to the canonical human-infective T trichiura. Comparative genomic analysis of genes suspected to confer resistance to either albendazole or ivermectin in helminths revealed a high number of &#x3b2;-tubulin orthologs, present in the whole population of T incognita n sp, compared with the canonical T trichiura species, but these genes were not associated with a resistant phenotype. The GWAS did not provide conclusive evidence of adaptation to drug pressure within the same species. INTERPRETATION: Our results demonstrate that trichuriasis can be caused by multiple whipworm species, and that differences in response rates might result from species responding differently to drug treatment, rather than from the intraspecies establishment of resistance. This discovery, coupled with the high tolerability of T incognita n sp to albendazole-ivermectin, marks a substantial shift in how we understand and approach whipworm infections. FUNDING: European Research Council.

Trichuris↗

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software↗

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames↗

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow&#x2122; (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software↗

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 &#xd7; homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering &#x223c;18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals↗