Search PubMedSearch

SEARCH · Search PubMed

Results for “long read”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Integrative Genotyping and Analysis of Canine Structural Variation Using Long-read and Short-read Data.

Structural variation makes an important contribution to canine evolution and phenotypic differences. Although recent advances in long-read sequencing have enabled the generation of multiple canine genome assemblies, most prior analyses of structural variation have relied on short-read sequencing. To offer a more complete assessment of structural variation in canines, we performed an integrative analysis of structural variants present in 12 canine samples with available long-read and short-read sequencing data along with genome assemblies. Use of long-reads permits the discovery of heterozygous variation that is absent in existing haploid assembly representations while offering a marked increase in the ability to identify insertion variants relative to short-read approaches. Examination of the size spectrum of structural variants shows that dimorphic LINE-1 and SINE variants account for over 45% of all deletions and identified 1,410 LINE-1s with intact open reading frames that show presence-absence dimorphism. Using a graph-based approach, we genotype newly discovered structural variants in an existing collection of 1,879 resequenced dogs and wolves, generating a variant catalog containing a 56.5% increase in the number of deletions and 705% increase in the number of insertions previously found in the analyzed samples. Examination of allele frequencies across admixture components present across breed clades identified 283 structural variants evolving with a signature of selection.

Animals

Nallo: a Nextflow pipeline for comprehensive human long-read genome analysis.

MOTIVATION: Long-read sequencing (LRS) is increasingly used for human medical research and clinical diagnostics due to its capacity to generate complete genome information. However, there is a lack of robust and easy-to-use pipelines for comprehensive LRS data analysis. RESULTS: Here we present Nallo, a Nextflow pipeline for analysis of PacBio and Oxford Nanopore data, with additional support for rare disease research projects. The pipeline detects a wide range of genetic variants, performs genome assembly, and reports CpG methylation. It also enables annotation and ranking of variants based on their predicted functional consequences. AVAILABILITY AND IMPLEMENTATION: Nallo is available from GitHub: https://github.com/genomic-medicine-sweden/nallo.

Humans

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Long-read sequencing reveals a hidden Alu-mediated splice defect in CPLANE1, causing orofaciodigital syndrome type VI.

Orofaciodigital syndrome type VI (OFD VI) is a recessive ciliopathy characterized by excessive polydactyly, molar tooth sign, cleft lip, and developmental delay, caused by pathogenic variants in CPLANE1. Here, we present a patient with OFD VI that remained genetically unexplained after routine genetic testing, including short-read whole genome sequencing (WGS). Using long-read sequencing, we found two biallelic splice-site variants in CPLANE1, c.8633-4_8633-3del, and an Alu element insertion close to an exon-intron boundary. Transcript analysis showed that each variant independently resulted in exon skipping, and quantitative expression studies revealed reduced total CPLANE1 mRNA levels in patient-derived fibroblasts. Based on these findings, we were able to re-classify the c.8633-4_8633-3del variant from a variant of uncertain significance (VUS) to likely pathogenic. The identification of an Alu element insertion missed by short-read WGS highlights the added diagnostic value of long-read sequencing in uncovering cryptic, transposable element-associated pathogenic variants.

Journal Article

Promises and pitfalls of long-read sequencing for resolving microbial complexity.

Long-read sequencing (LRS) has driven a transition in microbial genomics, overcoming the assembly fragmentation inherent to short-read sequencing. This review elucidates the impact of LRS across isolate genomics, metagenomics, and multi-omics domains. By spanning extensive repetitive regions, LRS facilitates the reconstruction of circular chromosomes and precisely resolves mobile genetic elements (MGEs). In metagenomics, LRS enables strain-level resolution, the recovery of circular metagenome-assembled genomes, and the precise localization of MGEs within host replicons. Furthermore, the single-molecule, amplification-free properties of LRS provide enhanced resolution of native epigenetic modifications and full-length transcriptomes. Despite these advancements, widespread implementation remains constrained by multidimensional challenges, including stringent high-molecular-weight DNA requirements, depth deficits, and computational overhead. Nevertheless, LRS is increasingly becoming the method of choice for isolate genomics and metagenomics. As detection technologies and algorithms progress, LRS will further improve our ability to decipher the structural and functional diversity of microbial ecosystems.

Metagenomics

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

Allele Level Sequencing of Killer Cell Immunoglobulin-Like Receptor Genes Using Oxford Nanopore Long Read Sequencing.

The human Killer cell Immunoglobulin-like Receptor (KIR) genes, found on chromosome 19, encode for cell surface protein receptors that, through interaction with their ligand, modulate the action of Natural Killer (NK) cells and some subsets of T lymphocytes. KIR genes exhibit extensive variation through variable gene content, copy number, and allele polymorphism. The combination of KIR genes and their ligands is implicated in various clinical settings including haematopoietic stem cell and solid organ transplant, and infectious disease progression. KIR gene content has been used in the selection of optimal stem cell donors with haplotype variations in recipient and donor giving differential clinical outcomes. With the introduction of massively parallel clonal next generation sequencing and single molecule long read third generation sequencing, allele level determination of KIR genotypes has become feasible. We describe a method for amplicon-based long read sequencing on the Oxford Nanopore Technologies platform that provides largely unambiguous allele level typing of KIR genes. The method was validated using DNA extracted from 48 10th International Histocompatibility Workshop (IHWS) cell lines with previously published allele level KIR genotypes and 176 Western Australian samples previously tested for the presence or absence of KIR genes. Our long-read sequencing method was able to accurately determine KIR alleles with an overall concordance of 97%-99% with the published data. Importantly, phasing ambiguity caused by the inability to phase heterozygous base positions over long stretches of gene sequence was resolved in several samples. Thus, our long read PCR sequencing strategy can be used to determine KIR genotypes at allele resolution level.

Humans

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Biallelic VPS41 Variants in Autosomal Recessive Spinocerebellar Ataxia 29 Resolved by Long-Read Sequencing and RNA Analysis.

BACKGROUND: Biallelic variants in VPS41, encoding a subunit of the HOPS complex, cause autosomal recessive spinocerebellar ataxia 29 (SCAR29), a rare neurodevelopmental disorder with an incompletely defined phenotypic and molecular spectrum. METHODS: We investigated a 24-year-old man with cerebellar ataxia, hypotonia, and intellectual disability. Exome sequencing identified four candidate VPS41 variants. Because maternal DNA was unavailable, long-read genome sequencing was performed to determine allelic configuration, followed by RNA and protein analyses. RESULTS: In addition to typical SCAR29 features, the patient showed previously unreported findings, including swan-neck deformities and pes cavus. Long-read genome sequencing demonstrated that two VPS41 variants were in trans. RNA analysis revealed distinct splicing consequences: one allele produced an out-of-frame transcript predicted to undergo nonsense-mediated decay, whereas the other generated an in-frame exon-skipped transcript. These complementary defects reduced VPS41 expression at both transcript and protein levels, supporting pathogenicity and variant reclassification. CONCLUSION: Our findings expand the phenotypic spectrum of VPS41-related disease and highlight the value of long-read allelic resolution in clarifying pathogenic mechanisms in rare genetic disorders.

Humans

Targeted long-read genomic and epigenomic profiling enhances timely comprehensive variant discovery in hypotonia and muscle weakness.

BACKGROUND: Identifying the genetic basis of hypotonia and muscle weakness is critical for patient management and family counseling. However, diagnosis is often hindered by diverse genomic alterations, including repeat expansions, structural variants (SVs), and methylation defects. Standard-of-care testing, largely based on short-read sequencing, is limited in its ability to detect this heterogeneous variation landscape, leaving many patients undiagnosed or requiring lengthy sequential testing. Long-read sequencing represents a promising solution. However, its application as a first-tier diagnostic assay for hypotonia remains unexplored. METHODS: We retrospectively analyzed 227 patients with hypotonia to assess diagnostic yield, time-to-diagnosis, and costs associated with standard-of-care testing. A long-read whole-genome sequencing (LR-WGS) workflow with targeted analysis of hypotonia-associated genes was developed to detect and prioritize pathogenic SNVs, SVs, and CNVs, repeat expansions, and methylation changes at key disease loci. The workflow was validated in a reference-positive cohort with known diagnoses (n = 15) and applied to an unsolved cohort (n = 14). Variant interpretation followed ACMG guidelines and was confirmed with orthogonal methods. RESULTS: Standard-of-care testing achieved a diagnostic yield of 42% with an average time-to-diagnosis of 68.7 days; however, 30% of diagnosed patients experienced significant delays (average 169 days) due to sequential testing. The LR-WGS based approach identified all known pathogenic variants in the positive cohort, including SMN1 deletions, methylation defects at 15q11.2/Prader-Willi locus, FMR1 repeat expansions, and sequence and copy-number variants in > 100 genes underlying myopathies and muscular dystrophies. The targeted long-read pipeline reduced prioritized variant calls by 97.9-99.9% and, in the unsolved cohort, yielded one definitive diagnosis (de novo COL6A3 deletion) and one possible diagnosis (aberrant methylation and copy number at POMK), for an additional 14% yield. Among patients diagnosed after sequential testing (n = 29), LR-WGS is expected to reduce time-to-diagnosis by ~ 85% and decrease cumulative diagnostic delays, with projected healthcare cost savings of $396,000-439,000. Across the entire 227 patient cohort, LR-WGS is anticipated to reduce testing costs by 6.5%, yielding an average savings of $105 per patient. CONCLUSIONS: LR-WGS enables comprehensive discovery of genomic and epigenomic variants in hypotonia and muscle weakness, improving diagnostic yield, shortening diagnostic timelines, and reducing costs compared with current standard-of-care testing.

Humans

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere

nf-core/pacvar: a pipeline for analyzing long-read PacBio whole genome and repeat expansion sequencing data.

MOTIVATION: Pacific Biosciences (PacBio) single-molecule, long-read sequencing enables whole genome annotation and the characterization of 20 complex repetitive repeat regions, especially relevant to neurodegenerative diseases, through their PureTarget panel. Long-read whole-genome sequencing (WGS) also allows for the detection of structural variants that would be difficult to detect with traditional short-read sequencing. However, the raw unaligned Binary Alignment Map data need to be processed before analysis. There is a need for an intuitive comprehensive bioinformatic pipeline that can analyze these data. RESULTS: We present nf-core/pacvar, a comprehensive pipeline for analyzing both PacBio single-molecule PureTarget and WGS data that demultiplexes and parallelizes pre-processing, variant calling and repeat characterization. nf-core/pacvar is compatible with little configuration and has few dependencies. This pipeline enables rapid end-to-end, parallel processing of PacBio single-molecule whole genome and targeted repeat expansion sequencing. AVAILABILITY AND IMPLEMENTATION: nf-core/pacvar is available on nf-core website (https://nf-co.re/pacvar/) and on github (https://github.com/nf-core/pacvar) under MIT License (DOI: 10.5281/zenodo.14813048).

Software

ALPINE: a scalable pipeline for comprehensive classification of gene-editing outcomes from long-read amplicon sequencing.

SUMMARY: CRISPR genome editing has enabled precise genetic modification for gene and cell therapies, but edits often produce heterogeneous on-target outcomes, including homology-directed repair (HDR) knock-ins, DNA repair template integrations, and structural variants. Existing tools are frequently limited to short reads or lack viral vector-specific integration categories needed for therapeutic development. Here, we present ALPINE (Amplicon Long-read Pipeline for INtegration Evaluation), a scalable and reproducible pipeline for classifying and quantifying gene-editing outcomes from long-read amplicon sequencing supporting both PacBio HiFi and Oxford Nanopore platforms. ALPINE classifies reads into 10+ categories, including DNA repair vector integration subtypes, and performs variant calling near the gene-edited site with batch, multi-sample reporting. Uniquely, ALPINE can distinguish between cells treated with multiple DNA repair vectors and identify distinct molecular features, such as inverted terminal repeats (ITRs), enabling comprehensive characterization of complex gene editing outcomes. Dual-target benchmarking on simulated datasets demonstrated high accuracy for transgene integration events. Independent validation on public crosslinked-HDR dataset confirmed ALPINE's integration detection capabilities, and application to edited T cell samples demonstrated comprehensive gene-editing outcome profiling. AVAILABILITY: ALPINE is available under MIT license at https://github.com/Maggi-Chen/ALPINE and https://doi.org/10.5281/zenodo.20272510. All analysis scripts and visualization code used in this manuscript are available at https://github.com/Maggi-Chen/ALPINE-manuscript-analysis. Simulated datasets are deposited at Zenodo (https://doi.org/10.5281/zenodo.20260865). Public dataset PRJNA913199 is available through NCBI SRA.

Gene Editing

Toward the clinical application of long-read sequencing in repeat-expansion disorders.

Repeat-expansion disorders (REDs) are a mechanistically and clinically well-defined subgroup of rare diseases caused by the expansion of short tandem repeats (STRs). These expansions can exceed several kilobases and show complex features, such as noncanonical secondary structures, somatic instability, repeat interruptions and allele-specific methylation. These characteristics are highly relevant for understanding disease mechanisms, clinical variability, prognosis and potentially therapeutic decision-making, but cannot be fully resolved using traditional diagnostic methods or short-read sequencing technologies. By contrast, long-read sequencing (LRS) enables accurate investigation of STR complexity in a single assay, facilitates the discovery of new pathogenic repeat expansions and drives advances in diagnostics, clinical and basic research, which may allow for better patient stratification in future clinical trials. This Perspective discusses recent LRS-driven discoveries, methodological and bioinformatic advances, and emerging diagnostic applications to illustrate the potential of LRS in reshaping both research and clinical practice.

Humans

First clinical diagnosis of FAME3 via commercial Long-Read sequencing reveals mosaic repeat expansion in MARCHF6 gene.

Familial Adult Myoclonic Epilepsy type 3 (FAME3) is a rare autosomal dominant disorder characterized by cortical tremor and epilepsy, caused by a noncoding pentanucleotide repeat expansion (TTTTA/TTTCA)n in the MARCHF6 gene. Conventional genetic testing often fails to detect this expansion due to its repetitive structure and intronic location. We evaluated a 61-year-old woman with refractory myoclonic and generalized tonic-clonic seizures, whose prior genetic testing-including exome and genome sequencing-was non-diagnostic. Using PacBio HiFi long-read whole-genome sequencing and the tandem repeat genotyping tool TRGT, we identified a pathogenic MARCHF6 intronic expansion. The proband harbored one allele with 15 TTTTA repeats and a second allele with a compound expansion of 661 TTTTA and 12 TTTCA repeats. Three affected relatives shared similarly expanded alleles, but with increasing repeat size in the latter generations. Importantly, analysis using TRGT-instability revealed repeat mosaicism in all affected individuals, reflected by variability in motif counts across individual sequencing reads. This somatic heterogeneity may contribute to the phenotypic penetrance, variable expressivity and pleiotropism seen in FAME3 disease expression. To our knowledge, this is the first clinical diagnosis of FAME3 using a commercially available long-read sequencing platform, underscoring its diagnostic utility in resolving complex repeat expansion disorders and uncovering biologically relevant mosaicism.

Humans

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing.

BACKGROUND: Allan-Herndon-Dudley syndrome (AHDS) is an X-linked disorder caused by pathogenic variants in the SLC16A2 gene. Although most reported variants are found in protein-coding regions or adjacent junctions, structural variations (SVs) within non-coding regions have not been previously reported. METHODS: We investigated two male siblings with severe neurodevelopmental disorders and spasticity, who had remained undiagnosed for over a decade and were negative from exome sequencing, utilizing long-read HiFi genome sequencing. We conducted a comprehensive analysis including short-tandem repeats (STRs) and SVs to identify the genetic cause in this familial case. RESULTS: While coding variant and STR analyses yielded negative results, SV analysis revealed a novel hemizygous deletion in intron 1 of the SLC16A2 gene (chrX:74,460,691 - 74,463,566; 2,876 bp), inherited from their carrier mother and shared by the siblings. Determination of the breakpoints indicates that the deletion probably resulted from Alu/Alu-mediated rearrangements between homologous AluY pairs. The deleted region is predicted to include multiple transcription factor binding sites, such as Stat2, Zic1, Zic2, and FOXD3, which are crucial for the neurodevelopmental process, as well as a regulatory element including an eQTL (rs1263181) that is implicated in the tissue-specific regulation of SLC16A2 expression, notably in skeletal muscle and thyroid tissues. CONCLUSIONS: This report, to our knowledge, is the first to describe a non-coding deletion associated with AHDS, demonstrating the potential utility of long-read sequencing for undiagnosed patients. Although interpreting variants in non-coding regions remains challenging, our study highlights this region as a high priority for future investigation and functional studies.

Humans