Search PubMedSearch

SEARCH · Search PubMed

Results for “Pangenome graph”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans

Pairwise graph edit distance characterizes the impact of the construction method on pangenome graphs.

MOTIVATION: Pangenome variation graphs are an increasingly used tool to perform genome analysis, aiming to replace a linear reference in a wide variety of genomic analyses. The construction of a variation graph from a collection of chromosome-size genome sequences is a difficult task that is generally addressed using a number of heuristics. The question that arises is to what extent the construction method influences the resulting graph, and the characterization of variability. RESULTS: We aim to characterize the differences between variation graphs derived from the same set of genomes with a metric which expresses and pinpoint differences. We designed a pairwise variation graph comparison algorithm, which establishes an edit distance between variation graphs, threading the genomes through both graphs. We applied our method to pangenome graphs built from yeast and human chromosome collections, and demonstrate that our method effectively characterizes discordances between pangenome graph construction methods and scales to real datasets. AVAILABILITY AND IMPLEMENTATION: pancat compare is published as free Rust software under the AGPL3.0 open source license. Source code and documentation are available at https://github.com/dubssieg/rs-pancat-compare. Snapshot available on Software Heritage at swh:1:dir:61acda8ba3dac1709ed60530147d3871831be629.

Algorithms

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software

Scalable computation of ultrabubbles in pangenomes by orienting bidirected graphs.

MOTIVATION: Pangenome graphs are increasingly used in bioinformatics, ranging from environmental surveillance and crop improvement to the construction of population-scale human pangenomes. As these graphs grow in size, methods that scale efficiently become essential. A central task in pangenome analysis is the discovery of variation structures. In directed graphs, the most widely studied such structures, superbubbles, can be identified in linear time. Their canonical generalization to bidirected graphs, ultrabubbles, more accurately models DNA reverse complementarity. However, existing ultrabubble algorithms are quadratic in the worst case. RESULTS: We show that all ultrabubbles in a bidirected graph containing at least one tip or one cutvertex-a common property of pangenome graphs-can be computed in linear time. Our key contribution is a new linear-time orientation algorithm that transforms such a bidirected graph into a directed graph of the same size, in practice. Orientation conflicts are resolved by introducing auxiliary source or sink vertices. We prove that ultrabubbles in the original bidirected graph correspond to weak superbubbles in the resulting directed graph, enabling the use of existing linear-time algorithms. Our approach achieves speedups of up to 25× over the ultrabubble implementation in vg, and of >200× over BubbleGun, enabling scalable pangenome analyses. For example, on the v2.0 pangenome graph constructed by the Human Pangenome Reference Consortium from 232 individuals, after reading the input, our method completes in under 3 min, while vg requires >1 hour, and four times more RAM. AVAILABILITY AND IMPLEMENTATION: Our method is implemented in the BubbleFinder tool github.com/algbio/BubbleFinder, via the new ultrabubbles subcommand.

Algorithms

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article

Pangenomic analyses in the cultivated grapevine confirm high genomic collinearity and extensive dispensable gene content likely involved in adaptation.

Pangenomes have now been developed for several horticultural crops, yet the extent to which genome diversity in sequence and organization contribute to plant adaptation and major agronomic traits remains poorly understood. Here, we assembled the genomes of 9 cultivated grapevine varieties and compared the genomes of 15 cultivated grapevine varieties for variation in gene and TE content. We found that genomic collinearity is highly conserved among varieties. We still observed substantial variation across genomes. Notably, we identified across varieties 55,662 orthologous genes, of which 55.3% appears to be dispensable. Dispensable genes are enriched for functions related to adaptation to biotic and abiotic constraints, suggesting that they may play a role in adaptation. Comparing our results with a recently published study, we found substantial differences with ∼12.6% of the genes we classified as core genes being classified as dispensable genes in this other study. We then constructed a pangenome graph and used it to performed genome-wide association studies for 3 important traits in grapevine production, which allowed us to include large structural variants as markers in the analyses. We identified 32 loci that we did not detect when we used the PN40024 genome as a reference, 20 of which are newly reported associations. Overall, our results indicates that despite recent advances in characterizing plant pangenomes, current gene classification into core and dispensable gene categories should be taken with caution. They also highlight the value of incorporating structural variants into GWAS, to better characterize the genetic architecture of agronomic traits.

Vitis

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis

Pouria Salehi Nowbandegani.

Dr. Laura Zahn asked Dr. Pouria Salehi Nowbandegani about their study, "Defining and cataloging variants in pangenome graphs," and how they came to study this aspect of genomics.

Humans

ntSynt-viz: Visualizing synteny patterns across multiple genomes.

With the explosion of chromosome-scale genome assemblies being generated in recent years, there is vast potential for comparative genomics analyses through detecting multi-genome synteny. While existing tools can detect synteny blocks between multiple genomes, their text-based outputs make it challenging to intuitively explore large-scale synteny patterns. Interpretable, information-rich and easy-to-use synteny visualization tools are imperative to enable important biological insights from the synteny block data output by the aforementioned utilities. Here, we present ntSynt-viz, a command-line tool for automated sorting, normalization and plotting of multi-genome synteny blocks. We show how ntSynt-viz provides clearer and more easily interpretable chromosome painting ribbon plots compared to the state-of-the-art tools NGenomeSyn and plotsr when evaluating synteny between 14 human genomes, and compared to NGenomeSyn when comparing 9 hoverfly genomes. As plotsr is limited to comparing genomes with equal chromosome numbers, it was not applicable to the hoverfly dataset. Furthermore, we demonstrate how ntSynt-viz can also be applied to visualize syntenic patterns encoded in pangenome graphs, using a Minigraph-Cactus graph built from 16 Drosophila genomes. We expect that ntSynt-viz will provide crucial insights into large-scale synteny patterns between divergent genomes, thereby advancing research into key evolutionary questions.

Synteny

A pangenome framework uncovers the role of deletions in repeated evolution of cave-derived traits.

Structural variants (SVs) are increasingly recognized as key contributors to adaptive evolution, yet they remain underexplored compared with single-nucleotide variation. To understand how large-scale genomic changes shape repeated evolution, we leveraged multiple levels of sequence data across the powerful evolutionary model system of the Mexican tetra fish (Astyanax mexicanus). We constructed one of the first pangenome graphs from a naturally evolving vertebrate, enabling comprehensive discovery of SVs among 120 fish from 11 populations. We discover substantial amounts of structural variation and explore the roles of genomic biases and selection in shaping the distribution of these variants. More than 2400 high-confidence cave-specific deletions are enriched in biological pathways involved in vision, metabolism, and behavior and cluster nonrandomly in quantitative trait loci linked to cavefish traits. Additionally, 67 genes harbor unique deletions between independent cavefish lineages. These reused genes show evidence of population-specific selection (99% contain selective sweeps compared with 8%-15% in genes lacking SVs), indicating that deletions likely rose in frequency through repeated positive selection rather than drift. Together, these results reveal that recurrent deletion events have repeatedly contributed to the evolution of cave-adapted phenotypes and highlight deletions as underexplored contributors of adaptive evolution in extreme environments.

Animals

Recurrent structural variation and recent turnover at the 17q21.31 locus in humans and great apes.

The 17q21.31 locus in humans harbors several complex structural haplotypes including a ~970kb inversion. Different inversion haplotypes have been associated with susceptibility to microdeletions causing Koolen-de Vries syndrome and variation in fecundity and recombination rates. Here, using 210 haplotype-resolved human genome assemblies and pangenome graph-based approaches we characterize 11 distinct structural haplotypes, several of which have not been previously described. Extending our analyses to a set of haplotype-resolved great-ape genomes, we characterize the structure of an independent inversion in chimpanzees which extends an additional 650kb, encompasses 5 additional genes, and is ~2 million years younger than the human inversion. We further determine that gorillas exhibit an independent duplication of the KANSL1 gene which may predispose them to Koolen-de Vries syndrome causing microdeletions. Using short read sequencing data we characterize 17q21.31 haplotype diversity worldwide in ~5174 individuals from 107 populations finding increased frequencies of KANSL1 duplication-containing haplotypes in both European and South Asian populations as well as 8 double recombination events between inverted and non-inverted haplotypes ranging in size from 20-180kb. Finally, using 626 ancient Eurasian human genomes we show the frequency of haplotypes containing KANSL1 duplications has increased ~6-fold over the past 12 thousand years in Europe. Together, our results highlight the dynamics, complexity, and recurrent, independent evolution of a medically relevant locus across humans and great apes.

Journal Article

National genomic projects in Asia and Africa: a review.

National genome projects (NGPs) are increasingly shaping precision medicine by improving representation of population-specific genetic diversity. This review compiles findings from NGPs across Asia and Africa, regions that remain underrepresented in global genomic databases despite their extensive demographic and genetic diversity. A total of 53 studies from 24 countries were identified to understand (1) the genomic approach utilized, (2) novel findings that have emerged, and (3) strategies for improving research in these regions. The NGPs implement population-based variome databases (20 NGPs), linear reference genome assemblies (8 NGPs), and graph-based pangenome assemblies (1 NGP). Novel variants ranged between 0.28% (China) and 19.6% (Iran), whereas rare variants accounted for up to 88.9% of the detected variants in the Chinese population. Each NGP documents its country's evolutionary and migration history, which impacts disease frequency and pharmacogenomic variants. Clinically, NGPs revealed strong population stratification in disease-associated and pharmacogenomic variants. For example, the GJB2 rs72474224 hearing-loss variant ranged from 13% in Vietnam and 12% in Hong Kong to 0.0894% in Turkey, while the VKORC1 rs9923231 pharmacogenomic variant reached 89.2% in Taiwan but was 20%-25% in European-related Russian subpopulations. These findings demonstrate that clinically relevant allele frequencies, pathogenicity assessments, and drug-response markers differ substantially across ancestries. This review highlights ongoing efforts and strategies to enhance the representativeness of genomic data through NGPs in Asia and Africa. We also suggest future directions for national projects, including integrating family-based studies, multi-omic data, and standardized pipelines to accelerate discovery and support the equitable implementation of precision medicine.

Humans

A graph-based approach for the visualisation and analysis of bacterial pangenomes.

BACKGROUND: The advent of low cost, high throughput DNA sequencing has led to the availability of thousands of complete genome sequences for a wide variety of bacterial species. Examining and interpreting genetic variation on this scale represents a significant challenge to existing methods of data analysis and visualisation. RESULTS: Starting with the output of standard pangenome analysis tools, we describe the generation and analysis of interactive, 3D network graphs to explore the structure of bacterial populations, the distribution of genes across a population, and the syntenic order in which those genes occur, in the new open-source network analysis platform, Graphia. Both the analysis and the visualisation are scalable to datasets of thousands of genome sequences. CONCLUSIONS: We anticipate that the approaches presented here will be of great utility to the microbial research community, allowing faster, more intuitive, and flexible interaction with pangenome datasets, thereby enhancing interpretation of these complex data.

Bacteria

Exploiting uniqueness: seed-chain-extend alignment on elastic founder graphs.

SUMMARY: Sequence-to-graph alignment is a central challenge of computational pangenomics. To overcome the theoretical hardness of the problem, state-of-the-art tools use seed-and-extend or seed-chain-extend heuristics to alignment. We implement a complete seed-chain-extend alignment workflow based on indexable elastic founder graphs (iEFGs) that support linear-time exact searches unlike general graphs. We show how to construct iEFGs, find high-quality seeds, chain, and extend them at the scale of a telomere-to-telomere assembled human chromosome. AVAILABILITY AND IMPLEMENTATION: Our sequence-to-graph alignment tool and the scripts to replicate our experiments are available in https://github.com/algbio/SRFAligner.

Software

Theseus: fast and optimal affine-gap sequence-to-graph alignment.

MOTIVATION: Sequence-to-graph alignment is a central problem in bioinformatics, with applications in multiple sequence alignment (MSA) and pangenome analysis, among others. However, current algorithms for optimal affine-gap alignment impose high memory and computational requirements, limiting their scalability to aligning long sequences to complex graphs. Practical solutions partially address this problem using heuristic strategies that ultimately trade off optimality for speed. RESULTS: This work presents Theseus, a novel, fast, and optimal affine-gap sequence-to-graph alignment algorithm. Theseus leverages similarities between genomic sequences to accelerate the alignment computation and reduces the overall memory requirements without compromising optimality. To that end, Theseus processes only a subset of the dynamic programming cells, using a sparse-data strategy that enables efficient sequence-to-graph alignment. Moreover, our algorithm supports optimal affine-gap alignment on arbitrary directed graphs, including those with cycles. We evaluate Theseus on two key problems: MSA and pangenome read mapping. For MSA, we compare it against SPOA, abPOA, and POASTA. Theseus is 1.6× to 17.6× faster than POASTA, and 7.3× faster, on average, than SPOA, both optimal aligners. Compared with abPOA, Theseus ensures optimality and scales to the largest problems. For pangenome read mapping, we benchmark Theseus against the alignment stage of the mapping tool vg map, along with the alignment kernels of SPOA, abPOA, and POASTA. Theseus outperforms the other methods, showing a 1.9× to 16.9× speedup on short reads. Moreover, Theseus is 1.5× to 36.3× faster than vg when aligning against synthetic cyclic graphs. AVAILABILITY AND IMPLEMENTATION: Theseus code and documentation are publicly available at https://github.com/albertjimenezbl/theseus-lib.

Algorithms

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals