Search PubMedSearch

SEARCH · Search PubMed

Results for “coverage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Animals

Blended Length Genome Sequencing (blend-seq): Combining Short Reads with Low-Coverage Long Reads to Maximize Variant Discovery.

We introduce blend-seq, a workflow for combining data from traditional short-read sequencing pipelines with low-coverage long reads, to improve variant discovery for single samples without the full cost of high-coverage long reads. We demonstrate that with only 4x long-read coverage augmenting 30x short reads, we can improve SNP discovery across the genome, exceeding performance beyond even high-coverage short reads (60x). For genotype-agnostic discovery of structural variants, we see a threefold improvement in recall while maintaining precision by using the low-coverage long reads on their own, and show how we can improve genotyping accuracy by adding in the short-read data. In addition, we demonstrate how the long reads can better phase these variants, incorporating long-context information in the genome to substantially outperform phasing with short reads alone. Our experiments highlight the complementary nature of short- and long-read technologies: the former contributing higher depth for genotyping and the latter better resolution of larger events or those in difficult regions.

cost optimization

Global disparities in COVID-19 vaccine coverage associated with trajectories of SARS-CoV-2 adaptation.

BACKGROUND: Vaccination serves as an effective intervention for health promotion and disease prevention across the socioecological systems and has played an important role during the COVID-19 pandemic. However, global disparities in vaccine coverage have increased uncertainty about the trajectories of viral adaptation, and the potential interplay between SARS-CoV-2 adaptation and vaccine rollout warrants further quantification. METHODS: Using over 13 million SARS-CoV-2 genomes across 86 countries from March 2020 to September 2022, we analyzed nonlinear associations between SARS-CoV-2 adaptation and vaccination coverage, considering public health and social measures, international travel, and infection dynamics, before and after the emergence of Omicron. Additionally, we examined the relationship between SARS-CoV-2 adaptation and COVID-19 mortality. RESULTS: During the pre-Omicron period, we found positive associations between nonsynonymous to synonymous divergence (dN/dS) ratios in the S1 subunit and medium levels of adjusted vaccine coverage (effect size: 0.96 [95% CI 0.47, 1.45]), while the association became insignificant at high levels (effect size: -1.89 [95% CI -4.20, 0.43]). However, no significant associations were found when Omicron dominated, possibly due to the immune escape ability of Omicron variants and the complex immune landscape shaped by mass hybrid immunity. Moreover, we observed evidence of dynamic interdependence and positive correlations between COVID-19 mortality and SARS-CoV-2 adaptation, with COVID-19 mortality interpreted as a proxy for uncontrolled viral spread. CONCLUSIONS: Our findings suggest a complex nonlinear relationship between vaccine-induced immunity and SARS-CoV-2 adaptation, with high vaccine coverage potentially linked to lower positive selection. We also observed directional coupling between COVID-19 mortality and SARS-CoV-2 adaptation. This may have implications for fair and fast vaccination in pandemic preparedness and response. CLINICAL TRIAL NUMBER: Not applicable.

Humans

Limited Impact of Column Chemistry and Length on Proteome Coverage Under High-Speed DIA.

The evolution of mass spectrometry (MS)-based proteomics has been driven by continuous technological advances in sample preparation, liquid-phase separations, instrumentation, and data acquisition. Chromatographic performance has been recognized as a contributing factor to identification depth, particularly on earlier-generation MS platforms. Recent advances in MS sampling speed and sensitivity now raise the question of how strongly chromatographic quality continues to determine overall proteome coverage. We investigate how column chemistry and length influence proteome coverage and chromatographic selectivity under modern data-independent acquisition conditions, and whether traditional optimization priorities still apply. Spanning a matrix of experiments with five distinct stationary phases, including C18 chemistries, C8, and Phenyl-Hexyl, across eight column lengths (40-140 mm), we evaluate protein identification performance using data-independent acquisition on the Orbitrap Astral mass spectrometer. Despite differences in stationary-phase chemistry and column length, we observed remarkably convergent proteome coverage metrics. All C18 and C8 phases consistently achieved over 150,000 precursor- and approximately 9000 protein group identifications, regardless of column length variations. While retention fingerprints persisted across chemistries, these chromatographic differences did not translate into meaningful variations in proteome coverage under high-speed acquisition conditions at 200 Hz. Within the range of modern sub-2 μm reversed-phase materials tested, identification depth showed limited dependence on column chemistry and length, suggesting that for state-of-the-art stationary phases, method development priorities may increasingly favor operational robustness, throughput, and reproducibility over traditional separation optimization.

Proteome

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal

MetaGLIMPSE: Meta-imputation of low-coverage sequencing data for modern and ancient genomes.

The advent of efficient and accurate imputation for low-coverage sequencing offers an unbiased alternative to SNP array imputation, increasing the accuracy of rare variant imputation across all populations. Since imputation accuracy generally increases with larger reference panels and closer ancestry match between target and reference samples, leveraging imputation from multiple reference panels improves imputation accuracy; however, individual reference panel genotypes are often privacy protected. Meta-imputation bypasses individual-level data by combining single-panel imputed genotypes through estimating panel- and marker-specific weights. We present a meta-imputation method, MetaGLIMPSE, that combines estimates from multiple reference panels for low-coverage sequencing imputation. Across all our scenarios, for both modern and ancient DNA samples, MetaGLIMPSE consistently outperforms the best single-panel imputation for coverages of 0.1×-8× and across all minor-allele frequencies, equaling the combined panel imputation for some parameters. Finally, MetaGLIMPSE is computationally efficient, meta-imputing 500 whole genomes in 16% of the time of GLIMPSE2.

Humans

esloco: simulation-based estimation of local coverage in long-read DNA sequencing.

SUMMARY: Long-read DNA sequencing is increasingly applied for whole-genome studies, yet experimental planning often lacks reliable estimates of target region coverage, leading to costly and time-consuming pilot studies and replicates. We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read sequencing experiments, including scenarios with unknown target regions (e.g. viral integration, CRISPR-Cas9) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of sequencing depth and read length distribution, esloco enables informed predictions of local sequencing outcomes. Benchmarking across a 45-gene panel demonstrated close agreement with empirical data, underscoring the framework's reliability. AVAILABILITY AND IMPLEMENTATION: esloco is a Python package available on PyPI (https://pypi.org/project/esloco/), GitHub (https://github.com/aweich/esloco), and Zenodo (https://doi.org/10.5281/zenodo.17776161).

Sequence Analysis, DNA

Large-scale simulation of coverage and error rate tradeoffs for cancer detection in cell-free DNA whole-genome sequencing.

MOTIVATION: Cell-free DNA (cfDNA) whole-genome sequencing (WGS) is a promising approach for detecting cancer recurrence. It enables cancer detection by identifying all tumor-derived cfDNA (ctDNA) molecules carrying somatic single nucleotide variants (sSNVs). While ideally, a sequencing platform should be highly accurate for reliable ctDNA detection, in reality, all sequencing platforms introduce sequencing errors that generate false positives indistinguishable from true SNVs. Understanding how sequencing parameters influence ctDNA detection sensitivity at low tumor fractions (TFs) in cfDNA samples is essential for guiding sequencing strategies in clinical contexts. To model cfDNA sequencing for tumor detection, which contains asymmetric noise and multiple interacting parameters, analytical modeling is intractable, motivating large-scale parallelized simulation. RESULTS: We developed a simulation framework to generate in silico cfDNA data across 10 cancer types. In total, 480 million cfDNA samples were simulated from tumor WGS profiles. Overall, the lowest detectable TF differs substantially between cancer types under identical sequencing conditions due to variations in mutational load. For cancers with high mutational load, 3× coverage with low-error techniques reliably detects TFs below 0.1%. In contrast, cancers with low mutational load require at least six-fold higher coverage to achieve comparable detection thresholds. Increasing sequencing quality scores from Q30 to Q55 at 30× coverage further enhances sensitivity, enabling detection of TFs as low as 1 × 10-5. This study provides a comprehensive framework for optimizing sequencing parameters, offering valuable guidance for tailoring future technology development for specific cancer types and clinical applications. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/UMCUGenetics/cfdetect/tree/main.

Whole Genome Sequencing

Genome-resolved assessment of archaeal diversity in full-scale anaerobic digesters reveals variability in mcrA primer coverage.

AIMS: Methanogenic archaea are key players in anaerobic digestion, driving methane production in biogas reactors. This study aimed to assess the diversity of methanogenic archaea in full-scale anaerobic digesters using genome-resolved metagenomics and to systematically evaluate the taxonomic coverage of commonly used mcrA-targeted qPCR primer sets against this genomic framework. METHODS AND RESULTS: Methanogenic diversity was assessed using 113 dereplicated archaeal metagenome-assembled genomes (MAGs) recovered from 109 full-scale anaerobic digesters treating diverse substrates. Genome-resolved analyses revealed a diverse archaeal community spanning multiple phyla, dominated by Halobacteriota and Methanobacteriota, with additional representatives from Methanobacteriota_B, Thermoplasmatota, and Thermoproteota. The presence of the mcrA gene was identified in a subset 55 MAGs, which were subsequently used as the genomic framework to evaluate six commonly used mcrA qPCR primer sets in silico. This subset clustered into nine phylogenetic groups and formed the basis for the primer coverage analysis. The evaluation revealed marked differences in taxonomic coverage among primer sets. Most primers preferentially detected Methanobacteriales and Methanosarcinales, while underrepresenting or excluding other methanogenic lineages, including H₂-dependent methylotrophic Methanomassiliicoccaceae. CONCLUSIONS: Commonly used mcrA primer sets differ substantially in their ability to capture methanogenic diversity, with some showing broad representation of reactor-associated methanogens and others exhibiting strong lineage-specific biases. Genome-resolved metagenomics provides an effective framework for benchmarking primer performance and supports the selection and improvement of molecular tools for more accurate monitoring of anaerobic digestion systems.

Archaea

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans

4CMenB vaccine coverage of invasive serogroup B meningococci collected in Belgium between 2016 and 2022.

Neisseria meningitidis infections can cause life-threatening meningitis and septicemia. In Europe, serogroup B (MenB) is the leading cause of invasive meningococcal disease (IMD), particularly in young children. Genomic surveillance of circulating MenB strains through whole genome sequencing (WGS) provides a powerful tool to assess the potential impact of vaccination strategies, including the 4CMenB vaccine, which is available for infants from 2 months of age. Here, we present a retrospective WGS-based analysis of clinical MenB IMD cases (n = 311) recovered in Belgium from 2016 to 2022 by the Belgian National Reference Center. High-quality WGS data were obtained for 281 of these strains, demonstrating high genetic diversity of the antigen targets included in the 4-component meningococcal serogroup B vaccine 4CMenB (fHbp, PorA, NHBA and NadA) and at the 4CMenB Antigen Sequence Types (BAST) level. Novel antigen combinations, not yet assigned a BAST ID, were detected in 23.5% of isolates. Vaccine coverage was predicted using the Genetic Meningococcal Antigen Typing System (gMATS) and the Meningococcal Deduced Vaccine Antigen Reactivity (MenDeVAR) index. Of the 281 strains, 79.5% (lower limit-upper limit: 68.0-91.5%) were predicted to be covered by the vaccine by gMATS, and 80.7% (lower limit-upper limit: 66.5-95.4%) by MenDeVAR. No evidence of variation in vaccine coverage was found throughout the study period nor between different age groups, demonstrating the broad applicability of 4CMenB. This study highlights the benefits of a pathogen surveillance program and the need for experimental characterization of continuously evolving antigenic subvariants of Neisseria meningitidis.

Humans

Self-Report Health Screening Tools in Female Athletes: A Systematic Review of Domain Coverage, Validation, and Use Across Participation Levels.

BACKGROUND: Female athlete health encompasses multiple interconnected domains; however, the self-report screening tools used to assess these domains have not been comprehensively synthesised. OBJECTIVE: To systematically identify self-report health screening tools used to assess female athlete health, map domain coverage, determine validation reporting, and describe application across participation levels. METHODS: This systematic review was pre-registered with PROSPERO ( CRD420251056910 ) and conducted in accordance with PRISMA guidelines. Four databases (PubMed, MEDLINE, SPORTDiscus and Web of Science) were searched from inception to January 2026 using female health and screening-related terms. Methodological quality was appraised using Joanna Briggs Institute and National Institutes of Health tools, and findings were synthesised descriptively. Eligible, peer-reviewed studies reported the use, development or validation of self-report health screening tools assessing one or more domains relevant to female health applied in female athlete populations, spanning recreational through elite participation levels. All sports and activities were included. The search was restricted to English language with no date limits. RESULTS: In total, 360 studies (1990-2026) representing 134,506 female participants spanning recreational to elite sport and 273 screening tools were included. Mental health (n = 77, 34.1%), disordered eating (n = 33, 14.6%) and body image (n = 30, 13.3%) predominated. Domains related to female health, including menstrual health, pelvic floor health, pregnancy/postpartum and breast health were comparatively underrepresented. Most studies reported tools were used for risk identification (n = 323, 80.3%). Validation reporting was inconsistent, with half (n = 180, 50%) reporting use of at least one validated tool. Tool use was concentrated in professional and elite sport, with limited inclusion of recreational, masters and disability athlete cohorts. Health literacy constructs were explicitly assessed in 12.5% of studies (n = 45). CONCLUSIONS: Health screening in female athlete populations remains fragmented and uneven in domain coverage, with inconsistent validation reporting. Development of integrated, multi-domain and contextually inclusive screening frameworks is warranted.

Journal Article

Long-read, high-coverage reference genome of the nymphalid butterfly Catonephele acontius (Nymphalidae: Biblidinae).

Catonephele acontius (Nymphalidae:Biblidinae:Epicalinii) is a butterfly species with a wide distribution across the Neotropics including the Amazon. Here, we present a long-read high-coverage reference genome for this species to serve as a genomic resource for future studies on Biblidinae butterflies, a group that is the subject of ongoing studies of seasonal adaptation under climate change. We used PacBio HiFi and IsoSeq reads to generate a highly contiguous and well-annotated reference genome. Five libraries were constructed, 4 using RNA from different tissues and 1 using high molecular weight (HMW) DNA from a wild-caught female. The DNA was sequenced using PacBio HiFi technology, and the RNA was sequenced using long read PacBio IsoSeq technology. About 20 Gb of raw HiFi data were generated and assembled to an initial size of 520.7 Mb (39 × homozygous coverage) in 90 contigs. The assembly was then polished and decontaminated into 40 contigs with an N50 of 19.927 Mb (BUSCO completeness: 99.0%; duplication: 0.5%; fragmentation: 0.7%; and missing: 0.3%). Final assembly size was 519.2 Mb. Repeats were annotated, showing that the genome consisted of 40.4% transposable elements. IsoSeq transcriptome data from antennae, leg, ovary, and digestive tissue was then used to structurally and functionally annotate gene models for the softmasked genome, uncovering ∼18,500 genes, with 70% of them given functional annotation. This reference assembly joins many published genomes in the Nymphalidae family but represents one of the first high-quality genomes from the Biblidinae subfamily. It provides a valuable resource to study the evolution of plastic and seasonal traits and will help investigate the genetic processes that may influence these species' responses to rapid climate change.

Animals

Absolute copy number aware CNV calling of sub-megabase segments in ultra-low coverage single-cell DNA sequencing data.

Recent advances in ultra-low coverage whole-genome sequencing (WGS) of single cells have enabled detailed analysis of copy number variation at a throughput approaching that of single-cell RNA sequencing. However, downstream computational methods have not seen comparable advances and are largely adaptations of deep sequencing methodology with reduced precision. Here, we present ASCENT, a computational method built to take full advantage of modern direct tagmentation-based WGS at ultra-low depth. Using joint segmentation with high-resolution bins, we accurately detect small segments, achieving accurate copy number profiles even at 100 000 reads per cell. ASCENT implements true absolute copy state inference for single cells, based on statistical modeling of coverage rather than comparison to a reference, while taking variable segment copy state into account. Further, ASCENT implements per-segment copy-neutral loss of heterozygosity (LOH) calling without the need for non-tumor or bulk WGS reference. When applied to a pediatric B-ALL sample, ASCENT finds copy-neutral LOH in a small segment and a minor subclone defined by breakpoints missed in bulk WGS. Thus, by applying appropriate computational methods, single-cell WGS provides clear advantages over bulk, even at a relatively low cell number and sequencing depth.

DNA Copy Number Variations

Meta-CD: a metagenomic sequencing coverage and depth calculator for target species.

Metagenomic Coverage and Depth Calculator (Meta-CD) is a convenient, biologist-friendly tool for determining coverage and depth to enhance taxonomic detection, functional profiling, and metagenome-assembled genome (MAG) recovery in metagenomics. It supports experimental design and post-sequencing analysis, modeling how genome size, relative abundance, sequencing depth, and DNA quantity influence detection of target species.

metagenomics

Development of a low-coverage whole genome sequencing screen for apomixis using a diverse set of Malus germplasm.

In the past decade, plant biologists have made several major discoveries pertaining to the genetic basis of apomixis (clonal propagation by seed) that have shown promise in preserving high-value hybrid rice and sorghum genotypes. This progress was made possible by foundational gene discovery efforts in model species and natural apomicts, but pleiotropic obstacles still limit its broad agricultural adoption, especially in eudicots. Thus, it follows that investigations of novel apomicts should lead to the development of new molecular tools for plant breeding. The two most common ways to identify clonal seed production are flow-cytometry seed screens and genome sequencing to compare the DNA sequences of the maternal parent and progeny, traditionally using low-throughput markers. While flow-cytometry has been the dominant method for more than two decades, it provides indirect information on the genetics of a resulting embryo and can be ineffective in certain species. Here we developed a method using short-read whole-genome sequencing at moderately low coverage (averaging 3X and 6X) to screen diverse Malus genotypes maintained in a USDA germplasm collection for clonal seed production. In total, we sequenced 55 genotypes, 1,216 of their embryos, and identified 17 previously undescribed apomictic genotypes. Several more were detected with the flow cytometry seed screen, which helped resolve certain types of reproduction and sources of noise in low-coverage datasets. This low-pass screening-by-sequencing method is a relatively low-cost, rapid method for detecting apomictic genotypes in diverse plant germplasm and when used thoughtfully in conjunction with flow cytometry, provides a new way to visualize the genetic outcomes of sexual and asexual reproduction in plants.

Apomixis

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article