Search PubMedSearch

SEARCH · Search PubMed

Results for “Short-read sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Short-read sequencing cannot determine whether co-occurring variants within a cancer gene lie on the same allele (cis) or opposing alleles (trans), a distinction with direct therapeutic consequences: trans configurations confirm biallelic tumor suppressor inactivation, whereas cis configurations generate compound oncogenic alleles with enhanced activity. Among 768 patients with prostate, breast, or ovarian cancers, we used mutational signatures to nominate cryptic genomic instability cases lacking a causative biallelic event on short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%) through methylation detection, long insertion resolution, and structural variant characterization, confirming trans inactivation in every resolved tumor suppressor case. Analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, 78.7% of which exceeded the 500 bp short-read phasing limit, and long-read phasing revealed recurrent compound cis alleles in NOTCH1, PIK3CA, PDGFRB, and KIT. Haplotype phasing addresses an overlooked gap in cancer variant interpretation and warrants integration into precision oncology.

Journal Article

Sequencing approaches in hereditary cancer testing: strengths, limitations and future directions.

Over the past three decades, Hereditary Cancer Testing (HCT) has evolved from single gene assays into multigene panel testing (MGPT), which allows for the screening of all known hereditary cancer genes in a single assay. MGPT is currently the standard approach for clinical HCT. However, with decreasing sequencing costs and increased instrument throughput, the scalability of exome sequencing (ES) and genome sequencing (GS) for HCT indications is becoming more viable. These methods provide broader insights into the coding exons and/or the entire genome, respectively. ES/GS data can also be reanalyzed to identify variants in novel genes that were not characterized at the time of initial testing, or to support research efforts aimed at uncovering additional associations between germline variants and cancer predisposition. Additionally, the emerging use of long-read sequencing (LRS) is noteworthy, enabling improved variant detection compared to short-read sequencing, especially for complex/structural variants and variation in difficult-to-sequence or paralogous regions in genes such as PMS2. This has the potential to increase the accuracy of HCT, reduce the turnaround time, find previously unidentifiable cancer risk variants, and ultimately increase the diagnostic yield. This article provides a comprehensive summary of the sequencing approaches used in HCT, discussing their strengths and limitations. We also highlight the added value of complementing DNA-only testing with RNA and tumor sequencing. Furthermore, we explore LRS-based approaches and discuss opportunities for their implementation in routine genetic testing for hereditary cancer.

Humans

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans

Toward the clinical application of long-read sequencing in repeat-expansion disorders.

Repeat-expansion disorders (REDs) are a mechanistically and clinically well-defined subgroup of rare diseases caused by the expansion of short tandem repeats (STRs). These expansions can exceed several kilobases and show complex features, such as noncanonical secondary structures, somatic instability, repeat interruptions and allele-specific methylation. These characteristics are highly relevant for understanding disease mechanisms, clinical variability, prognosis and potentially therapeutic decision-making, but cannot be fully resolved using traditional diagnostic methods or short-read sequencing technologies. By contrast, long-read sequencing (LRS) enables accurate investigation of STR complexity in a single assay, facilitates the discovery of new pathogenic repeat expansions and drives advances in diagnostics, clinical and basic research, which may allow for better patient stratification in future clinical trials. This Perspective discusses recent LRS-driven discoveries, methodological and bioinformatic advances, and emerging diagnostic applications to illustrate the potential of LRS in reshaping both research and clinical practice.

Humans

Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions.

The architecture and function of biological systems are inherently three-dimensional, yet most existing spatial transcriptomic technologies remain restricted to thin tissue sections, limiting their capacity to resolve cellular organization and microenvironments within intact tissue volumes. To address this limitation, we developed volumetric DNA microscopy, a scalable, optics-free approach for spatial transcriptome profiling directly within intact biological specimens. The method encodes spatial information into DNA molecules that form a dense intermolecular network in situ, enabling the reconstruction of three-dimensional spatial relationships through short-read sequencing and computational analysis. Here we detail the complete workflow including in situ cDNA synthesis, spatial encoding through DNA nanoball formation, dual-scale proximity bridging between neighboring nanoballs and spatial reconstruction via geodesic spectral embedding. Sequencing libraries can be generated within 7-8 d by a competent graduate-level molecular biologist, followed by standardized downstream computational analysis. Because the workflow requires only routine molecular biology reagents and a benchtop sequencer, volumetric DNA microscopy provides a versatile platform for exploring genetic and morphological features in intact tissues.

Spatial Transcriptomics

Megamimivirus double-stranded DNA linear genomes flanked by highly diverse terminal inverted repeats.

UNLABELLED: Giant viruses have fundamentally expanded our understanding of virology by challenging the conventional boundaries of both virion size and genome complexity. However, the scarcity of isolates has left many of their unique biological features unexplored. Here, we report the isolation and characterization of four new giant virus species belonging to the subfamily Megamimivirinae, sampled from distinct environments across China. Among these, Megavirus daqingense is the first giant virus isolated from an oil reservoir; it exhibits virion stability under high salinity, chloroform exposure, and elevated temperatures, suggesting fitness adaptations to subsurface conditions. Using a hybrid sequencing approach that integrates short- and long-read technologies, we assembled complete linear genomes for all four isolates, each flanked by long terminal inverted repeats (TIRs). Comparative genomic and synteny analyses identified 29 distinct TIRs from 46 megamimivirus genomes. Gene content within these TIRs was highly diverse, with no orthologous proteins conserved across all repeats. Furthermore, TIR genes experienced weaker purifying selection than those in non-TIR regions (i.e., the genomic regions excluding the TIRs), consistent with their role as drivers of genome plasticity. Notably, we discovered for the first time that identical tRNA genes are shared between TIRs and non-TIR regions of eukaryotic viruses. Collectively, our work provides insights into the structural and evolutionary complexity of megamimiviruses, revealing TIRs as reservoirs of genetic diversity and hotspots for gene transfer, thereby playing a pivotal role in shaping the dynamic architecture of giant virus genomes. IMPORTANCE: Terminal inverted repeats (TIRs) are critical structural elements at the termini of linear genomes essential for fundamental processes such as recombination, replication, and integration across diverse organisms. However, the inherent limitations of short-read sequencing technologies have left the complete structure, diversity, and evolutionary significance of long TIRs in giant viruses unexplored. In this study, we leverage hybrid sequencing and comparative genomic analyses to unveil the complexity of TIRs across the subfamily Megamimivirinae. We demonstrate that TIRs are dynamic genomic hotspots characterized by remarkable gene diversity and unexpected conservation of specific tRNA genes. These findings establish TIRs as key drivers of genome plasticity, serving as hotspots for horizontal gene transfer and genetic innovation. By resolving the long-hidden terminal structures of megamimivirus genomes, this work provides a foundational framework for understanding how TIRs shape the evolution of giant viruses and, more broadly, advances our understanding of genome architecture in large DNA viruses.

Megavirus

Single-cell multi-omics dissects transcript isoform and immune repertoire dynamics in human immunosenescence.

Immunosenescence, a major hallmark of systemic aging, refers to the progressive functional decline of the immune system. This decline not only compromises host defense and immunological memory but also fuels chronic inflammation and tissue degeneration (collectively known as inflammaging). While single-cell RNA sequencing (scRNA-seq) has revealed transcriptomic alterations associated with immune aging, analyses restricted to transcript abundance fail to capture deeper regulatory layers, such as transcript isoform diversity and the remodeling of immune receptor repertoires. To address this limitation, we present a human peripheral immune single-cell multi-omics atlas that integrates gene expression, transcript isoform diversity, and immune receptor repertoires. By combining single-cell full-length transcriptome sequencing (scCycloneSEQ), short-read scRNA-seq, and single-cell immune receptor sequencing (scTCR/BCR-seq), we systematically profiled peripheral blood mononuclear cells (PBMCs) from healthy donors aged 30-40 and 60-70 years. Our analyses uncovered extensive age-related remodeling of immune cell composition, functional states, and TCR/BCR diversity. Notably, we found that CD4+ effector memory T cells exhibited widespread differential isoform usage (DIU), 3'UTR length variation, and a marked reshaping of cytotoxic T lymphocyte (CTL) clonotypes-all of which were closely associated with aging-related inflammation and cellular senescence. This multi-omics atlas delineates key molecular features of immunosenescence and provides a high-resolution resource for deciphering the regulatory architecture underlying immune aging.

TCR/BCR

Structural genome variation drives adaptation of the xylose-fermenting yeast Scheffersomyces stipitis to lignocellulosic hydrolysates.

Second-generation (2G) bioethanol from lignocellulosic feedstocks is a sustainable alternative to fossil fuels. However, its production is constrained by the poor performance of industrial microbes in hydrolysates that are generated during biomass pretreatment. Scheffersomyces stipitis is a native xylose fermenting yeast and a promising platform for 2G bioethanol production, and adaptive evolution under hydrolysate stress has yielded strains with enhanced performance. However, the chromosomal basis of this adaptation is unknown. Here, we demonstrate that chromosome scale structural variation, rather than point mutations, underlies the improved phenotype of the evolved strains. By integrating long- and short-read genome sequencing, we identify two major chromosomal rearrangements in the top performing isolate: a reciprocal translocation between chromosomes 1 and 2 that disrupts the NUDIX hydrolase gene YSA1, and the formation of a mitotically stable 175 kb minichromosome derived from chromosome 5. Functional analyses show that disruption of YSA1 enhances xylose utilisation and ethanol yield, while the minichromosome contributes to improved performance in hydrolysate conditions. These findings provide direct evidence that balanced rearrangements and minichromosome formation can be selected during prolonged stress and can generate adaptive phenotypes. Taken together, our study establishes genome reorganisation as a key driver of adaptation in S. stipitis.

Xylose

Development and evaluation of a multiplex PCR-based dual-platform targeted sequencing framework for precise differentiation of lumpy skin disease virus.

BACKGROUND: Lumpy skin disease virus (LSDV) shares over 96% genomic identity with goatpox and sheeppox viruses, presenting severe diagnostic challenges due to cross-reactivity. METHODS: To address this bottleneck, we established a targeted sequencing framework integrating multiplex PCR with short-read and long-read platforms. By sequentially screening target pathogens, identifying low-homology genes, and designing short and gradient long-fragment primer pools, we evaluated these dual-platform panels using highly homologous poxvirus samples. RESULTS: The short-read panel stably detected target viruses at inputs as low as 5.26 ×101 copies/μL. Under strict alignment criteria, LSDV mapping rates reached 42.91%, suppressing non-target signals to 3.05%. The Nanopore-Targeted Sequencing (NTS) long-amplicon strategy successfully eliminated homologous interference. By applying length-dependent diagnostic thresholds (≥ 100 reads for short amplicons; ≥ 50 reads for long amplicons), precise species-level identification was achieved, maintaining near-zero cross-reads (0-5) in ultra-long regions. Crucially, the field-deployable NTS workflow enabled complete detection in approximately 4 h. CONCLUSION: This complementary strategy seamlessly meets both laboratory demands for high-sensitivity enrichment and frontline requirements for rapid typing, providing a reliable tool for LSDV surveillance, mutation tracking, and outbreak control.

Capripoxvirus differentiation

Complete genome sequence of the Anaplasma phagocytophilum clinical isolate NCH-1.

Anaplasma phagocytophilum is an obligate intracellular gram-negative bacterium and etiologic agent of human granulocytic anaplasmosis. A. phagocytophilum genomic sequencing has historically been performed via short-read platforms. Our optimized bacterial isolation protocol combined with Nanopore sequencing produced a single, closed 1,481,805 bp circular A. phagocytophilum strain NCH-1 chromosome.

Anaplasma phagocytophilum

MetaStrainer: accurate reconstruction of bacterial strain genotypes from short-read metagenomic samples.

MOTIVATION: Metagenomics provides broad insights from microbial communities, but more biological relevant phenotypes are attributed to subtle changes at the strain-level rather than species. Despite development of several tools using different algorithms, resolving individual strains from short-read pair-end sequencing data remains challenging. RESULTS: Here we present MetaStrainer, a tool capable of reconstructing strain genotypes from metagenomic data. Compared with existing approaches, MetaStrainer substantially increases genotype accuracy, correctly identifies the number of strains, and accurately estimates their relative abundances. Accuracy of reconstructed genotypes is robust to choice of mapping reference. AVAILABILITY: MetaStrainer is implemented in Python 3. Source code and instructions are available on GitHub at www.github.com/lbobay/MetaStrainer and on Zenodo: 10.5281/zenodo.17872331.

Metagenomics

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Genomic Diversity and Extended-Spectrum β-Lactamase Gene Contexts of Community Resident-Carried Escherichia coli in Ecuador.

Community carriage of extended-spectrum β-lactamase (ESBL)-producing Escherichia coli represents an important reservoir of antimicrobial resistance. However, the genomic diversity and population structure of ESBL-producing E. coli circulating in community settings remain poorly characterized. This study aimed to characterize ESBL-producing E. coli isolated from fecal samples of residents in Ecuador, with an emphasis on the diversity and genomic context of ESBL genes. ESBL-producing E. coli was isolated from fecal samples obtained from 55 residents using MacConkey agar supplemented with cefotaxime. Whole-genome sequencing of the isolates was performed using a hybrid approach combining long- and short-read platforms. Plasmids and β-lactamase genes were identified using DFAST and PlasmidFinder. Bacterial identification and antimicrobial susceptibility testing were conducted by MALDI-TOF MS and the broth microdilution method, respectively. ESBL-producing E. coli were isolated from 35 of 55 fecal samples (63.6%). Complete circular genomes were obtained from 31 isolates. All isolates harbored bla CTX-M genes, predominantly belonging to the bla CTX-M-1 group, whereas 65.7% carried bla TEM, mainly bla TEM-1, and related variants. Although β-lactamase genes were predominantly plasmid-borne, chromosomal integration was detected in 40% of the isolates. Notably, 87.5% of the isolates harbored IncF plasmids with multiple replicons. Conserved IS26-flanked transposons carrying bla CTX-M and bla TEM were frequently identified in the plasmids. Phylogenetic analysis revealed substantial genomic diversity across seven phylogroups, together with closely related isolates detected within and between households. These findings provide high-resolution genomic insights into the ESBL determinants circulating in community residents and reveal region-specific patterns of ESBL genomic diversity.

CTX-M β-lactamases

Exploring differences across pangenome-graph representations using Escherichia coli O157:H7 as a model.

Pangenome graphs are increasingly used to represent population-scale bacterial diversity, yet construction methods span fundamentally different representation paradigms whose outputs and sensitivities to assembly quality remain poorly quantified. We systematically reviewed microbial pangenome graph tools and benchmarked seven representative methods spanning gene-cluster, compacted coloured de Bruijn graph, one hybrid approach and one multiple sequence alignment method. Using a repeat-rich Escherichia coli O157:H7 dataset with complete genomes and matched short-read data, we constructed graphs from identical inputs and observed orders-of-magnitude differences in graph size and fragmentation, indicating that global topology is driven by representation strategy. Varying completeness composition revealed that assembly fragmentation is a first-order determinant of graph structure: gene-cluster graphs contracted as draft assemblies replaced complete genomes, whereas compacted coloured de Bruijn graphs expanded, with distinct degree-prevalence fingerprints across tools. In contrast, the multiple sequence alignment method could not be evaluated across fragmented inputs because it did not run reliably on draft-assembly datasets. Computational cost mirrored these shifts and depended strongly on completeness composition, including a pronounced runtime penalty for one compacted coloured de Bruijn graph implementation on all-draft inputs. Finally, analysis of Shiga toxin loci showed that pangenome-level reconciliation by gene-cluster-based tools does not reliably correct assembly artefacts at challenging multi-copy genes and that performance varies by locus. Together, these findings show that pangenome graphs are representation-dependent models of bacterial diversity, and that, in this repeat-rich O157:H7 benchmark dataset, assembly completeness is a primary determinant of their topology, scalability, and locus-level accuracy.

Escherichia coli O157

nf-core/pacsomatic: a scalable somatic analytic pipeline using PacBio HiFi data.

MOTIVATION: Pacific Biosciences (PacBio) HiFi long-read sequencing enables robust characterization of complex genomic regions, repetitive elements, and structural variants (SVs) that are often inaccessible to short-read technologies. To fully leverage HiFi reads to advance cancer genomics and epigenetics, researchers require an end-to-end, scalable and optimized bioinformatics workflow. The nf-core framework meets this need by providing rigorously tested, community-curated pipelines that ensure reproducibility, transparency, and broad compatibility across computational environments. RESULTS: We present nf-core/pacsomatic, an automated Nextflow DSL2 pipeline designed for comprehensive paired tumor-normal somatic analysis using PacBio HiFi data. The workflow includes steps for read alignments against reference genome, somatic SNV/indel, SV, and CNV calling, CpG methylation profiling and differential methylation region (DMR) detection. Additional downstream modules support functional annotation, mutational signature analysis, tumor purity and ploidy estimation, and homologous recombination deficiency (HRD) assessment. Utilizing nf-core's modular design and containerized execution, nf-core/pacsomatic provides a stable framework for the reproducible discovery of biological insights. AVAILABILITY: nf-core/pacsomatic is available under the MIT License at nf-core (https://nf-co.re/pacsomatic) and github (https://github.com/nf-core/pacsomatic).

Software

Long-read sequencing reveals widespread novel splicing and neojunction-derived neoantigens in nasopharyngeal carcinoma.

The widespread transcriptomic diversity driven by alternative splicing (AS) contributes to all hallmarks of cancer and represents a critical source of neoantigens for personalized immunotherapy. However, unlike other major malignancies, the full repertoire of AS in nasopharyngeal carcinoma (NPC) remains underexplored. Here, we employ long-read sequencing (LR-seq) to generate a high-resolution, isoform-level transcriptomic atlas from a cohort of 14 NPC tumor samples and four immortalized nasopharyngeal epithelial cell lines. We identify a substantial number of full-length novel transcripts (22,687; ∼44.38%), which reveal diverse splicing patterns and previously unannotated splicing events. By integrating short-read RNA-seq data to quantify isoform expression, we discover a subset of novel transcripts that are differentially expressed between tumor samples and immortalized nasopharyngeal epithelial cell lines. Furthermore, LR-seq enables precise identification of chimeric readthrough fusion transcripts, such as CLDN15-FIS1 and FOXRED2-TXN2 Finally, we develop a computational framework, tumor-specific splicing neoantigen detection (TS-SNAD), to predict neoantigens originating from novel exon-exon junctions (neojunctions) in tumor-specific novel transcripts. Using this framework, we identify neojunction-derived neoantigens and experimentally validate the immunogenicity of selected HLA-B*40:01-restricted neoantigens. These neojunction-derived peptides constitute a new class of noncanonical neoantigens with significant potential for developing personalized cancer vaccines for NPC.

Humans