Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Short-read sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Assembling genomes of non-model plants: A case study with evolutionary insights from Ranunculus (Ranunculaceae).

Whereas genome sequencing and assembly technologies are improving, cost can still be prohibitive for plant species with large, complex genomes. As a consequence, genomics work on some taxa in evolutionarily pivotal positions in the vascular plant tree of life has been hampered. The species-rich genus Ranunculus (Ranunculaceae) is an important angiosperm group for the study of polyploidy, apomixis, and reticulate evolution. However, neither mitochondrial nor high-quality nuclear genome sequences are available. This limits phylogenomic, functional, and taxonomic analyses thus far. Here, we tested Illumina short-read, Oxford Nanopore Technology (ONT) and PacBio (HiFi) long-read, and hybrid-read assembly strategies. We sequenced the diploid progenitor species R. cassubicifolius (R. auricomus species complex) and selected the best assemblies in terms of completeness, contiguity, and quality scores. We first assembled the plastome (156 kbp, 85 genes) and mitogenome (1.18 Mbp, 40 genes) sequences using Illumina and Illumina-PacBio-hybrid strategies, respectively. We also present an updated plastome and the first mitogenome phylogeny of Ranunculaceae, including studies of gene loss (e.g., infA, ycf15, or rps) with evolutionary implications. For the nuclear genome sequence, we favored a PacBio-based assembly polished three times with filtered short reads and subsequently scaffolded into eight pseudochromosomes by chromatin conformation data (Hi-C). We obtained a haploid genome sequence of 2.69 Gbp, with 94.1% complete BUSCO genes found and 35 482 annotated genes, and inferred ancient gene duplications compared to existing Ranunculales genomes. The genomic information presented here will enable advanced evolutionary-functional analyses for the species complex, but also for the genus and beyond Ranunculaceae.

Ranunculus↗

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans↗

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software↗

Integrative Long-Read Multi-Omics of a Patient With GPI Deficiency: A Molecular Case Study of a Candidate Dual-Effect GPI Variant.

The molecular determinants of phenotypic severity in red cell enzymopathies are often obscured by the disconnect between coding sequence variants and their regulatory landscapes. Here we present a single-patient molecular case study that uses an integrative multi-omic approach-combining short-read WGS, PacBio HiFi long-read sequencing, native CpG methylation profiling, and Iso-Seq full-length transcriptomics-to characterize a severe, transfusion-dependent hemolytic anaemia. We identified a compound heterozygous state in the glucose-6-phosphate isomerase (GPI) gene, with no wild-type allele present. One allele (Haplotype 1) carried a missense variant (p.His191Arg); the other (Haplotype 2) carried a distinct missense variant, c.1414C>T (p.Arg472Cys), previously reported as biochemically unstable. Long-read phasing placed the two variants in trans. Allele-resolved transcript counts showed a directionally consistent but statistically non-significant trend toward higher expression of Haplotype 2 across two Iso-Seq replicates. Notably, the c.1414C>T transition abolishes a local CpG dinucleotide; in a small number of haplotype-2 reads spanning this position, the corresponding cytosine on the wild-type/Haplotype-1 background was methylated. We did not measure GPI protein abundance, enzymatic activity, or stability in this patient, and we do not establish that methylation at this site regulates GPI transcription. On the basis of these correlative observations in a single patient, we propose-as a hypothesis for future testing-that a coding variant might simultaneously perturb protein stability and disrupt a local epigenetic mark, and we outline the experiments required to test whether such a dual effect contributes to disease. This case illustrates the value of integrative long-read multi-omics for generating mechanistic hypotheses about variants of uncertain significance, while underscoring that causal claims require dedicated functional validation.

Humans↗

Optimized Amplicon Strategy for Long-Read Sequencing of the Chikungunya Virus Genome.

Chikungunya virus (CHIKV) is a positive-sense RNA alphavirus transmitted to humans primarily by Aedes aegypti and Aedes albopictus mosquitoes. Its global circulation and significant public health impact underscore the need to better understand the molecular mechanisms driving CHIKV pathogenesis and transmission. Although robust molecular biology methods exist for CHIKV genome sequencing, a major limitation for surveillance and research is the inability to determine whether two nucleotide variations co-occur within the same viral genome when they are separated beyond the span of typical short-read designs. Here, we describe an optimized approach for processing CHIKV RNA samples that generates large amplicons suitable for long-read nanopore sequencing. This protocol enables amplification of the complete CHIKV genome in only two or three amplicons and facilitates detection of co-occurring nucleotide variations across 4-7.5 kb within the same molecule, thereby simplifying sequencing workflows and improving resolution in studies of viral evolution.

Chikungunya virus↗

Global diversity and evolution of Salmonella enterica serovar Panama: a genomic epidemiology study.

BACKGROUND: Non-typhoidal Salmonella is a globally important bacterial pathogen, typically associated with foodborne gastrointestinal infection. Some non-typhoidal Salmonella serovars can also colonise typically sterile sites in people to cause invasive non-typhoidal Salmonella disease. Salmonella enterica serovar Panama is responsible for a substantial number of cases of human bloodstream infection, but despite its global dissemination, numerous outbreaks, and a reported association with invasive non-typhoidal Salmonella disease, S enterica serovar Panama (S Panama) is understudied. We aimed to describe the genomic epidemiology and evolutionary history of S Panama to provide a vital baseline of understanding for this globally important serovar. METHODS: In this genomic epidemiology study, we analysed S Panama genomes derived from historical collections, national surveillance datasets, and publicly available epidemiological and whole-genome sequencing data which span the years 1931-2019. Maximum likelihood and Bayesian phylodynamic approaches were used to investigate population structure and evolutionary history and to infer geotemporal dissemination. A combination of different bioinformatic approaches with short-read and long-read data were used to characterise geographical and clade-specific trends in antimicrobial resistance (AMR) and genetic markers for invasiveness. FINDINGS: We analysed 836 S Panama genomes, of which 559 (67%) were sequenced as part of this study. The collection represents all inhabited continents and includes isolates collected between 1931 and 2019. We identified the presence of four geographically linked S Panama clades (C1 [ie, the Latin America and the Caribbean clade; n=338], C2 [ie, the European clade; n=124], C3 [ie, the Martinique clade; n=131], and C4 [ie, the Asia and Oceania clade; n=104]) and regional trends in AMR profiles. Most isolates (715 [86%] of 836) were pan-susceptible to antibiotics and belonged to clades circulating in Latin America and the Caribbean (64%, n=458). Most antibiotic-resistant isolates in our collection (113 [93%] of 121) fell within clades C4 (ie, the Asia and Oceania clade) and C2 (ie, the European clade), the latter of which had the highest invasiveness index values based on the conservation of 196 extraintestinal predictor genes. INTERPRETATION: This first large-scale phylogenetic analysis of S Panama has revealed important information about the population structure, AMR, global ecology, and genetic markers of invasiveness of the identified genomic subtypes. Our findings provide an important baseline for understanding S Panama infection. The presence of multidrug-resistant clades with elevated invasiveness index values should be monitored through ongoing surveillance, as such clades could pose an increased public health risk. FUNDING: UK Research and Innovation Global Challenges Research Fund and Biotechnology and Biological Sciences Research Council, UK Medical Research Council, Wellcome Trust, John Lennon Memorial Scholarship, Institut Pasteur, Santé publique France, Fondation Le Roch-Les Mousquetaires, Investissement d'Avenir Programme, and Australian National Health and Medical Research Council.

Humans↗

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans↗

Long-read sequencing reveals widespread novel splicing and neojunction-derived neoantigens in nasopharyngeal carcinoma.

The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; &#x223c;44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.

Humans↗

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames↗

Temporal shifts in K-locus composition and expansion of dual-carbapenemase-producing ST11-KL62 Klebsiella pneumoniae: a retrospective genomic surveillance study.

OBJECTIVES: To characterize longitudinal changes in carbapenem-resistant Klebsiella pneumoniae (CRKP) and investigate the recent increase in dual-carbapenemase-producing ST11-KL62 isolates. METHODS: We retrospectively analysed 1,239 non-duplicate CRKP isolates recovered at a tertiary hospital in China during 2018-2025. Antimicrobial susceptibility testing, whole-genome sequencing, K-locus and resistance/virulence gene profiling, core-genome single-nucleotide polymorphism analysis, reference-guided plasmid comparison, conjugation and stability assays, and a murine lethality model were used. RESULTS: ST11 accounted for 936/1,239 isolates (75.5%). KL47 declined from 39/151 (25.8%) in 2018-2019 to 45/755 (6.0%) in 2024-2025, whereas KL62 increased from 3/151 (2.0%) to 147/755 (19.5%). Among 148 ST11-KL62 isolates, 13/148 (8.8%) co-harboured blaKPC-2 and blaNDM-1, of which 12/13 (92.3%) met the study's molecular definition of hypervirulent CRKP. Pairwise single-nucleotide polymorphism distances among local ST11-KL62 isolates ranged from 0 to 43 (median, 14), suggesting that clonal expansion may have contributed to their increase. Complete genome analysis of ZD872 located blaKPC-2, blaNDM-1, and major virulence-associated genes on distinct plasmids; related plasmid backbones were predicted in other isolates using short-read comparisons. ZD872 exhibited a hypervirulent phenotype in the murine model. CONCLUSIONS: The ST11 CRKP population underwent temporal shifts in K-locus composition, including expansion of a closely related ST11-KL62 subset carrying dual carbapenemases and hypervirulence-associated markers. These findings support integrating longitudinal genomic surveillance with local transmission analysis.

Carbapenem-resistant Klebsiella pneumoniae↗

The First Highly Contiguous Genome Assembly for the Western Bluebird (Sialia mexicana).

The western bluebird (Sialia mexicana) is a secondary cavity-nesting thrush that has experienced historical population declines, local extirpations, and more recent recoveries associated with nest box programs. Despite these regional successes, recent eBird estimates suggest continued range-wide declines and substantial geographic variation in population trajectories, making this species a useful system for future studies of demographic change, connectivity, and conservation genomics. However, genomic resources for western bluebirds remain limited, and no reference genome currently exists for any species in the genus Sialia. Here, we present the first high-quality de novo reference genome for S. mexicana. Using PacBio HiFi long-read sequencing from an adult female, we generated a highly contiguous, phased 1.3&#x2005;Gb nuclear assembly with a contig N50 of 24.8&#x2005;Mb and high BUSCO completeness of 98.3%. We annotated the nuclear genome using transcriptomic and protein evidence, identifying 16,656 protein-coding genes and 26,060 transcripts/protein isoforms. We also assembled a complete &#x223c;16&#x2005;kb mitochondrial genome from Illumina short-read data. This reference genome provides a foundational resource for future studies of population structure, genetic diversity, connectivity, demographic history, and adaptation in western bluebirds and related taxa.

Animals↗

StrainMake: reproducible hybrid metagenomics with MAG recovery and strain-level resolution.

SUMMARY: Metagenomic workflows involve complex multi-step analyses, from quality control and assembly to binning, annotation, and strain-level profiling. Few existing metagenomic pipelines achieve the combination of flexibility, reproducibility, and hybrid assembly support within a unified workflow. We present StrainMake, a Snakemake-based workflow for de novo metagenomic analysis from short, long, or hybrid sequencing data. StrainMake integrates widely used tools across all major steps-quality control, assembly, binning, dereplication, taxonomic and functional annotation-while also providing non-redundant gene catalogues, community-scale metabolic models, and strain-level microdiversity metrics. The modular design enables the use of alternative tools, scalable execution on HPC systems, and full reproducibility through Snakemake and Conda. RESULTS: Applied to the CAMI II strain-madness dataset, StrainMake produced high-quality assemblies and metagenome-assembled genomes (MAGs), while enabling strain-resolved comparisons across samples. Hybrid assemblies improved contiguity, whereas short-read assemblies offered faster runtimes, illustrating the workflow's benchmarking capacity. AVAILABILITY AND IMPLEMENTATION: StrainMake is open source and available at https://github.com/UMMISCO/strainmake, together with comprehensive documentation. Generated data are deposited in Zenodo (doi: 10.5281/zenodo.16950162).

Metagenomics↗

Global emergence and transmission dynamics of carbapenemase-producing Citrobacter freundii sequence type 22 high-risk international clone: a retrospective, genomic, epidemiological study.

BACKGROUND: Carbapenemase-producing Citrobacter (CPC) species have recently been recognised as emerging pathogens associated with nosocomial infections in humans. The increased rate of Citrobacter freundii infections is a public health concern and there is a paucity of genomic data regarding its global transmission dynamics. We aimed to characterise the genetic features of CPC species, and their associated carbapenemase-encoding plasmids, obtained from hospitalised patients in China and from publicly available global data, with a particular focus on high-risk clones. METHODS: This was a retrospective, genomic epidemiological study of CPC species obtained from a tertiary hospital in Zhejiang Province, China, from March 5, 2013, to March 5, 2023. We used antimicrobial susceptibility testing, short-read and long-read whole-genome sequencing, phylogenomic analysis, and plasmid structure analysis. A global dataset of complete plasmid sequences encoding blaKPC, blaNDM, and blaIMP was constructed from the National Center for Biotechnology Information (NCBI) RefSeq database to provide insights into their diversity and distribution. All carbapenemase-producing Citrobacter freundii genomes from the NCBI GenBank database were incorporated in the comparative genomic analyses. Bayesian phylogeographical analysis and growth rate assays were carried out to characterise the high-risk C freundii sequence type (ST) 22 clone. FINDINGS: 1724 Citrobacter species isolates were collected from diverse clinical specimens, with 48 identified as CPC species. Citrobacter koseri (22 [46%] of 48) and C freundii (20 [42%]) were the predominant CPC species. Comparative analysis found C freundii carried significantly higher median numbers of plasmid replicons (5&#xb7;0 [IQR 3&#xb7;3-6&#xb7;0] vs 2&#xb7;0 [2&#xb7;0-3&#xb7;0]; p<0&#xb7;0001) and acquired antimicrobial resistance genes (12&#xb7;0 [7&#xb7;3-15&#xb7;8] vs 3&#xb7;0 [3&#xb7;0-5&#xb7;3]; p<0&#xb7;0001) than did C koseri. Molecular characterisation identified Inc-type plasmids, In823::Kl.pn.I3/In1589-like/In837-like integrons, Tn6296/Tn125/Tn5060 transposons, and insertion sequences (eg, IS26, IS3000, IS5, ISAba125, ISCR1), collectively facilitating the dissemination of carbapenemase genes. Global analysis of 3126 carbapenemase-encoding plasmids found epidemic plasmids with broad host ranges and global diversity. Phylogenetic investigation of predominant carbapenemase-encoding plasmids showed their persistence across geographical regions, temporal spans, and Enterobacterales species, exhibiting high genetic similarity to our clinical plasmids. A phylogenetic tree of 726 global carbapenemase-producing C freundii genomes showed that ST22 (227 [31&#xb7;3%]) represents the predominant multidrug-resistant clone across community, health-care, and environmental niches. Transmission across continents contributes to the global predominance of the ST22 clone, which carries a high load of resistance genes (median 15&#xb7;0 [IQR 11&#xb7;0-17&#xb7;0] vs 12&#xb7;0 [3&#xb7;0-16&#xb7;0]; p<0&#xb7;0001) and enhanced plasmid maintenance capacity (median replicons 5&#xb7;0 [IQR 4&#xb7;0-7&#xb7;0] vs 4&#xb7;0 [3&#xb7;0-6&#xb7;0]; p<0&#xb7;0001) relative to non-ST22 clones. INTERPRETATION: Our study provides evidence to suggest that Citrobacter species are emerging carriers of carbapenem-resistance genes. These findings provide insight into the population structure of CPC species and highlight C freundii ST22 as a prominent high-risk international clone. FUNDING: National Natural Science Foundation of China, National Health Commission Scientific Research Fund-Zhejiang Provincial Major Health Science and Technology Plan Project, Zhejiang Province Natural Science Foundation Project, Outstanding Youth Foundation of Jiangsu Province of China, the Priority Academic Program Development of Jiangsu Higher Education Institutions, and Postgraduate Research and Practice Innovation Program of Jiangsu Province.

Citrobacter freundii↗

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis↗

Ancient climate changes and relaxed selection shape cave colonization in North American cavefishes.

Extreme environments serve as natural laboratories for studying evolutionary processes, with caves offering replicated instances of independent colonizations. The timing, mode and genetic underpinnings underlying cave-obligate organismal evolution remain enigmatic. We integrate phylogenomics, fossils, palaeoclimatic modelling and newly sequenced genomes to elucidate the evolutionary history and adaptive processes of cave colonization in the study group, the North American Amblyopsidae fishes. Amblyopsid fishes present a unique system for investigating cave evolution, encompassing surface, facultative cave-dwelling and cave-obligate (troglomorphic) species. Using 1105 exon markers and total-evidence dating, we reconstructed a robust phylogeny that supports the nested position of eyed, facultative cave-dwelling species within blind cavefishes. We identified three independent cave colonizations, dated to the Early Miocene (18.5 Ma), Late Miocene (10.0 Ma) and Pliocene (3.0 Ma). Evolutionary model testing supported a climate-relict hypothesis, suggesting that global cooling trends since the Early-Middle Eocene may have influenced cave colonization. Comparative genomic analyses of 487 candidate genes revealed both relaxed and intensified selection on troglomorphy-related loci. We found more loci under relaxed selection, supporting neutral mutation as a significant mechanism in cave-obligate evolution. Our findings provide empirical support for climate-driven cave colonization and offer insights into the complex interplay of selective pressures in extreme environments.

Animals↗

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6&#x2009;Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia↗

Genome size estimation from long read overlaps.

MOTIVATION: Accurate genome size estimation is an important component of genomic analyses such as assembly and coverage calculation, though existing tools are primarily optimized for short-read data. RESULTS: We present LRGE, a novel tool that uses read-to-read overlap information to estimate genome size in a reference-free manner. LRGE calculates per-read genome size estimates by analysing the expected number of overlaps for each read, considering read lengths and a minimum overlap threshold. The final size is taken as the median of these estimates, ensuring robustness to outliers such as reads with no overlaps. Additionally, LRGE provides an expected confidence range for the estimate. We validate LRGE on a large, diverse bacterial dataset and confirm it generalizes to eukaryotic datasets. On bacterial genomes, LRGE outperforms k-mer-based methods in both accuracy and computational efficiency and produces genome size estimates comparable to those from assembly-based approaches, like Raven, while using significantly less computational resources. AVAILABILITY AND IMPLEMENTATION: Our method, LRGE (Long Read-based Genome size Estimation from overlaps), is implemented in Rust and is available as a precompiled binary for most architectures, a Bioconda package, a prebuilt container image, and a crates.io package as a binary (lrge) or library (liblrge). The source code is available at https://github.com/mbhall88/lrge and an archive at https://doi.org/10.5281/zenodo.17183812 under an MIT license.

Genome Size↗