Search PubMedSearch

SEARCH · Search PubMed

Results for “genome assembly error”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Revisiting the genome assembly of Lupinus species reveals differential diploidization after a shared whole-genome duplication.

Accurate genome assemblies are essential for comparative genomics, yet Hi-C-guided scaffolding can introduce structural errors that misrepresent chromosome architecture and bias evolutionary inferences. Here, we identified pervasive scaffolding errors-including artificial fusions, internal inversions, and incomplete contig mounting-in 2 previously published Lupinus genomes (L. cosentinii and L. digitatus) using a segmentation method based on long terminal repeat (LTR) retrotransposon density. We reassembled both genomes, producing chromosome-level references of 472.7 Mb (16 chromosomes) and 427.2 Mb (21 chromosomes), with BUSCO completeness >98.5%. Synteny validation and reapplication of LTR profiling confirmed that all prior errors were resolved. Using these corrected genomes together with 4 additional Lupinus species and 2 outgroup legumes, we investigated postpolyploid evolution. Synonymous substitution rate (Ks) analysis revealed a genus-specific whole-genome duplication (WGD) event (Ks = 0.17) shared by all 6 Lupinus species. The proportion of WGD-derived genes varied markedly, from 60% in L. digitatus to only 36% in L. mutabilis, indicating differential diploidization. While all species retained a core set of WGD duplicates enriched in cytoskeleton organization, ion transport, and defense responses, each exhibited lineage-specific functional trajectories: cell wall modification in L. cosentinii and L. digitatus, nitrogen metabolism in L. albus and L. angustifolius, flower development in L. luteus, and stress/lipid metabolism in L. mutabilis. Our corrected assemblies provide optimal references for Lupinus comparative genomics, and our findings demonstrate that a shared WGD event can lead to both conserved and highly divergent postpolyploid fates, likely underpinning adaptive diversification within the genus.

Lupinus

The complete telomere-to-telomere sequence of a mouse Y chromosome.

The mouse Y chromosome is essential for male reproduction, yet the GRCm39 reference contains 25 gaps, particularly in repetitive and complex regions. Here, we assembled a telomere-to-telomere Y chromosome (mT2T Y) of 95.21 Mb from a C57BL/6 mouse incorporating parental genomes. This assembly fills all gaps, corrects structural errors, and adds over 8.70 Mb of previously unassembled sequence to the reference genome. We annotated 142 previously unidentified genes, identified Y specific satellite arrays, and mapped homologous recombination loci in the pseudoautosomal region (PAR). Analysis of X Y homologous gene expression revealed a Y chromosome dosage compensation mechanism. By combining mT2T Y with T2T mhaESC, we completed the T2T assembly of all C57BL/6 chromosomes, designated T2T mhaESC+Y, providing a complete C57BL/6 reference genome.

Animals

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9 Gb and represented 23 pairs of chromosomes with an NG50 of 150 Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans

Autocycler: long-read consensus assembly for bacterial genomes.

MOTIVATION: Long-read sequencing enables complete bacterial genome assemblies, but individual assemblers are imperfect and often produce sequence-level and structural errors. Consensus assembly using Trycycler can improve accuracy, but its lack of automation limits scalability. There is a need for an automated method to generate high-quality consensus bacterial genome assemblies from long-read data. RESULTS: We present Autocycler, a command-line tool for generating accurate bacterial genome assemblies by combining multiple alternative long-read assemblies of the same genome. Without requiring user input, Autocycler builds a compacted De Bruijn graph from the input assemblies, clusters and filters contigs, trims overlaps, and resolves consensus sequences by selecting the most common variant at each locus. It also supports manual curation when desired, allowing users to refine assemblies in challenging or important cases. In our evaluation using Oxford Nanopore Technologies reads from five bacterial isolates, Autocycler outperformed individual assemblers, automated pipelines, and other consensus tools, producing assemblies with lower error rates and improved structural accuracy. AVAILABILITY AND IMPLEMENTATION: Autocycler is implemented in Rust, open-source, and freely available at github.com/rrwick/Autocycler. It runs on Linux and macOS and is extensively documented.

Genome, Bacterial

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

A Step-by-Step Guide to Sequencing and Assembly of Complete Bacterial Genomes Using the Oxford Nanopore MinION.

The Oxford Nanopore (ONT) MinION enables sequencing of longer DNA/RNA fragments compared to other sequencers, such as Illumina, etc. This nanopore method provides distinct advantages for generating complete genome assemblies from microorganisms. Specifically, the R9.4 flow cells used for MinION sequencing have much lower error rates compared with earlier versions of the ONT platform. Coupled with base calling using Dorado software, higher-quality long reads can now be generated for complete bacterial genome assembly. In this chapter, we describe a detailed MinION method to assemble a complete genome from a microorganism, polish the final assembly, and evaluate the genome quality using various software tools. Because of the low cost for MinION sequencing, this platform could be an asset for virtually any laboratory interested in generating complete genomes from microorganisms.

Genome, Bacterial

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

De novo haplotype-resolved genome assembly of the endemic kiwifruit Actinidia hubeiensis.

The genus Actinidia, which encompasses the widely cultivated kiwifruit, is characterized by its rich species diversity. Wild Actinidia species serve as invaluable germplasm reservoirs for crop improvement. As an important kiwifruit species, Actinidia hubeiensis represents a unique taxonomic group endemic to Hubei Province, contributing valuable genetic diversity to the genus Actinidia. Here, we present a haplotype-resolved genome assembly for A. hubeiensis. The two haplotype assemblies (Hap1 and Hap2) spanned 658.03 Mb (N50 = 23.16 Mb) and 597.19 Mb (N50 = 20.89 Mb), encoding 35,741 and 36,647 high-confidence protein-coding genes, respectively. Based on comprehensive assessments, both haplotypes demonstrated high completeness (BUSCO completeness > 99%), excellent continuity (LAI up to 21.67), low base-error rates (QV > 40), and nearly complete read mapping rates (> 98%). This genome assembly provides crucial genomic resources for the genus, enriching our understanding of kiwifruit biodiversity and offering new insights into the genetic background and evolutionary characteristics of this distinctive species.

Actinidia

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Chromosome-level genome assembly of the small-sized Taihang donkey (Equus asinus).

China harbors a rich diversity of donkey breeds, with small-sized donkeys (<110&#x2009;cm) representing a largely underexplored group. Here, we present the first high-quality, chromosome-level genome assembly of a small-sized donkey, generated using PacBio HiFi sequencing (286.7&#x2009;Gb), Hi-C scaffolding (240.47&#x2009;Gb), and annotated with RNA-seq data. The final assembly has a total length of 2.7&#x2009;Gb and comprises 32 chromosomes (including both X and Y chromosomes), in which five chromosomes were fully assembled without gaps. It possesses a scaffold N50 of 106.70&#x2009;Mb and 84 contigs (contig N50&#x2009;=&#x2009;63.60&#x2009;Mb), and captures 99.2% of BUSCO genes. The assembly achieved a consensus quality value (QV) of 77.44, corresponding to an extremely low base-level error rate, indicating exceptional nucleotide accuracy. This high-quality genome provides a valuable resource for investigating genetic variation, adaptive evolution, and domestication processes in small-sized donkeys, and will facilitate the conservation and sustainable utilization of rich donkey genetic resources in China.

Animals

Reference genome bias in light of species-specific chromosomal reorganization and translocations.

BACKGROUND: Whole-genome sequencing efforts, have during the past decade, unveiled the central role of genomic rearrangements-such as chromosomal inversions-in evolutionary processes, including local adaptation in a wide range of taxa. However, employment of reference genomes from distantly or even closely related species for mapping and the subsequent variant calling can lead to errors and/or biases in the datasets generated for downstream analyses. RESULTS: Here, we capitalize on the recently generated chromosome-anchored genome assemblies for Arctic cod (Arctogadus glacialis), polar cod (Boreogadus saida), and Atlantic cod (Gadus morhua) to evaluate the extent and consequences of reference bias on population sequencing datasets (approx. 15-20&#x2009;&#xd7;&#x2009;coverage) for both Arctic cod and polar cod. Our findings demonstrate that the choice of reference genome impacts the mapping statistics, including mapping depth and mapping quality, as well as core population genetic estimates, such as heterozygosity levels, nucleotide diversity (&#x3c0;), and cross-species genetic divergence (DXY). Furthermore, using a more distantly related reference genome can lead to inaccurate detection and characterization of chromosomal inversions, i.e., in terms of size (length) and location (position), due to inter-chromosomal reorganizations between species. Additionally, we observe that some of the verified species-specific inversions are split across multiple genomic regions when mapped against a heterospecific reference. CONCLUSIONS: Inaccurate identification of chromosomal rearrangements as well as biased population genetic measures could potentially lead to erroneous interpretation of species-specific genomic diversity, impede the resolution of local adaptation, and thus, impact predictions of their genomic potential to respond to climatic and other environmental perturbations.

Animals

Functional integration of the bacteriophage T4 DNA replication complex: The multiple roles of the ssDNA binding protein (gp32).

Single-stranded DNA binding protein (gp32) serves as the central regulatory component of the multi-subunit T4 bacteriophage DNA replication system by coordinating the system's three functional sub-assemblies, resulting in phage DNA synthesis in T4-infected E. coli cells at the high speeds (~1,000 nts s-1) and the high fidelity (< 1 error per 107 nts) required for genomic function within this cellular eco-system. Gp32 proteins continuously bind to, slide as cooperatively-linked clusters on, and un-bind from transiently exposed single-stranded (ss) DNA templates to carry out their coordinating functions, as well as to protect genomic sequences from nuclease activity and block the formation of interfering secondary structures. The N-terminal domains (NTDs) of gp32 mediate cooperative interactions within ssb clusters, but the roles of the disordered C-terminal domains (CTD) in the nucleation of gp32-ssDNA filaments at ss-dsDNA junctions are less well understood. We here present microsecond-resolved single-molecule F&#xf6;rster resonance energy transfer studies of the initial steps of gp32 assembly on short oligo-deoxythymidine lattices of varying lattice length and polarity near model ss-dsDNA junctions. These data are analyzed to define the molecular steps and related free energy surfaces involved in initiating gp32 cluster formation, which show that the nucleation mechanisms and regulatory interactions driven by gp32 proteins at ss-dsDNA junctions are significantly directed by lattice polarity. We propose a model for the role of the CTDs in orienting gp32 monomers at lattice positions close to ss-dsDNA junctions that suggests how intrinsically disordered CTD domains might facilitate and control non-base-sequence-specific binding in both the nucleation and the dissociation of the gp32-ssDNA filaments involved in phage DNA replication and related processes.

Journal Article

Preserving centromere identity: right amounts of&#xa0;CENP-A&#xa0;at the right place and time.

Four decades ago, the discovery of centromere protein-A (CENP-A) marked a pivotal breakthrough in chromosome biology, revealing the epigenetic foundation of centromere identity. CENP-A, a histone H3 variant, directs the formation of the microtubule-binding kinetochore complex, designating the chromosomal site for its assembly and underpins the accurate partitioning of genetic material during cell division. Errors in cell division can give rise to DNA instability and aneuploidy, implicated in human diseases such as cancer. Therefore, discovering the underlying pathways and mechanisms responsible for the formation, regulation and maintenance of the centromere is important to our understanding of genome stability, epigenetic inheritance, and in providing the knowledge to help generate possible treatments and therapeutics. Here, we review various molecular pathways and mechanisms implicated in maintaining centromere identity and highlight some of the key outstanding questions with a focus on the human centromere.

Humans