Search PubMedSearch

SEARCH · Search PubMed

Results for “Haplotype phasing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Unraveling the complex genetic landscape of OTOF-related hearing loss: a deep dive into cryptic variants and haplotype phasing.

BACKGROUND: Pathogenic variants in OTOF are a major cause of auditory synaptopathy. However, challenges remain in interpreting OTOF variants, including difficulties in confirming haplotype phasing using traditional short-read sequencing (SRS) due to the large gene size, the potential incomplete penetrance of certain variants, and difficulties in assessing variants at non-canonical splice sites. This study aims to revisit the genetic landscape of OTOF variants in a Taiwanese non-syndromic auditory neuropathy spectrum disorder (ANSD) cohort using a combination of sequencing technologies, predictive tools, and experimental validations. METHODS: We performed SRS to analyze OTOF variants in 65 unrelated Taiwanese patients diagnosed with non-syndromic ANSD, complemented by long-read sequencing (LRS) for haplotype phasing. A prediction-to-validation pipeline was implemented to assess the pathogenicity of cryptic variants using SpliceAI software and minigene assays. RESULTS: Biallelic pathogenic OTOF variants were identified in 33 patients (50.8%), while monoallelic variants were found in five patients. Three novel variants, c.3864G > A (p.Ala1288 =), c.4501G > A (p.Ala1501Thr), and c.5813 + 2T > C, were detected. The pathogenicity of two non-canonical mis-splicing variants, c.3894 + 5G > C and c.3864G > A (p.Ala1288 =), was confirmed by minigene assays. LRS-based haplotype phasing revealed that the common missense variant c.5098G > C (p.Glu1700Gln) and the novel variant c.5975A > G (p.Lys1992Arg) are in cis and form a founder pathogenic allele in the Taiwanese population. CONCLUSIONS: Our study highlights the genetic heterogeneity of DFNB9 and emphasizes the importance of population-specific variant interpretation. The integration of advanced sequencing technologies, predictive algorithms, and functional validation assays will improve the accuracy of molecular diagnosis and inform personalized treatment strategies for individuals with DFNB9.

Humans

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Short-read sequencing cannot determine whether co-occurring variants within a cancer gene lie on the same allele (cis) or opposing alleles (trans), a distinction with direct therapeutic consequences: trans configurations confirm biallelic tumor suppressor inactivation, whereas cis configurations generate compound oncogenic alleles with enhanced activity. Among 768 patients with prostate, breast, or ovarian cancers, we used mutational signatures to nominate cryptic genomic instability cases lacking a causative biallelic event on short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%) through methylation detection, long insertion resolution, and structural variant characterization, confirming trans inactivation in every resolved tumor suppressor case. Analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, 78.7% of which exceeded the 500 bp short-read phasing limit, and long-read phasing revealed recurrent compound cis alleles in NOTCH1, PIK3CA, PDGFRB, and KIT. Haplotype phasing addresses an overlooked gap in cancer variant interpretation and warrants integration into precision oncology.

Journal Article

Long-Read Haplotype Phasing Resolves Allelic Configuration as a Missing Layer of Precision Oncology.

Conventional short-read sequencing cannot determine whether co-occurring variants within a cancer gene reside on the same allele (cis) or on opposing alleles (trans), a distinction with direct biological and therapeutic consequences. Trans configurations confirm biallelic tumor suppressor inactivation and inform therapy selection, while cis configurations generate compound oncogenic alleles with enhanced activity. We analyzed 768 patients with prostate, breast, or ovarian cancers in the PROBLEM cohort, using mutational signatures to nominate cryptic genomic instability cases where the causative biallelic event was not apparent from short-read sequencing. Long-read nanopore sequencing resolved 32 of 46 cryptic cases (69.6%), leveraging its unique advantages in direct methylation detection, long insertion resolution, and complex structural variant characterization, confirming trans biallelic inactivation in all resolved tumor suppressor cases. Systematic analysis of 4,496 MiOncoSeq samples identified 17,519 multi-hit gene pairs, of which 78.7% exceeded the 500 bp short-read phasing limit. Long-read phasing further revealed recurrent compound cis oncogenic alleles in NOTCH1, PIK3CA, PDGFRB, and KIT with functionally synergistic activity. Haplotype phasing resolves a systematically overlooked gap in cancer variant interpretation and warrants broader integration into precision oncology workflows.

Journal Article

Haplotype-specific expression of a terpene synthase underlies linalool variation in the grapevine cultivar Riesling.

Grapevine cultivars vary widely in monoterpenoid content, yet the genetic and regulatory mechanisms underlying this variation remain poorly characterized beyond highly aromatic Muscat types. We profiled free volatiles and monoterpenoid glycosides in a Riesling × Cabernet Sauvignon F1 mapping population, revealing extensive variation and transgressive segregation consistent with multigenic control. QTL mapping identified 70 significant loci associated with 48 volatile compounds and monoterpene glycosides, including two major QTLs explaining 33.6% and 33.4% of phenotypic variance in (3S)-linalool accumulation. Integration of haplotype-resolved transcriptomics with metabolite data, enabled by a chromosome-scale diploid Riesling genome assembly, resolved a (3S)-linalool/nerolidol synthase cluster on chromosome 10 and identified VviTPS54 as the strongest candidate underlying linalool variation. VviTPS54 exhibited haplotype-specific expression strongly correlated with (3S)-linalool accumulation across genotypes, while no QTL was detected at the 1-deoxy-D-xylulose-5-phosphate synthase 1 (VviDXS1) locus previously identified in Muscat cultivars. In addition, VviDXS1 expression was not correlated with terpene levels, indicating that regulatory variation within terpene synthase clusters, rather than methylerythritol phosphate (MEP) pathway flux, drives monoterpenoid composition in this population. These results establish regulatory variation of terpene synthases as a key mechanism underlying monoterpenoid diversity in grapevine and demonstrate that resolving such variation requires haplotype-phased genome assemblies coupled with haplotype-resolved transcriptomics to detect allele-specific expression differences at complex, heterozygous loci.

Grapevine

A personalized multi-platform assessment of somatic mosaicism in the human frontal cortex.

Somatic mutations in individual cells create genomic mosaicism, influencing genetic disorders and cancers. While clonal mutations in cancers are well-studied, rarer somatic variants in normal tissues remain poorly characterized. This study systematically evaluates detection methods using a personalized donor-specific assembly (DSA) from a neurotypical individual's dorsolateral prefrontal cortex assessed with Oxford Nanopore, NovaSeq, linked-read sequencing, Cas9-targeted long-read sequencing (TEnCATS), and single-neuron MALBAC amplification. The haplotype-resolved DSA improved cross-platform analysis, dramatically increasing phasing rates. Germline SNVs, structural variations (SVs), and transposable elements (TEs) were recalled with 99.4%-99.7% accuracy in bulk tissue, and phased haplotype analysis reduced false positives by 15.4%-75.1% for putative somatic candidates. Long-read single-neuron sequencing detected nine somatic SV candidates, demonstrating enhanced sensitivity for rare variants, while TEnCATS identified eight low-frequency somatic TE candidates. These findings highlight advanced methodologies for precise somatic variant detection, critical for understanding mosaicism's role in health and disease.

Multi-platform Sequencing

The evolution of separate sexes in waterhemp is associated with surprising chromosomal diversity and complexity.

The evolution of separate sexes is hypothesized to occur through distinct pathways involving few large-effect or many small-effect alleles. However, we lack empirical evidence for how these different genetic architectures shape the transition from quantitative variation in sex expression to distinct male and female phenotypes. To explore these processes, we leveraged the recent transition of Amaranthus tuberculatus to dioecy within a predominantly monoecious genus, along with a sex-phenotyped population genomic dataset, and six newly generated chromosome-level haplotype phased assemblies. We identify a ~3 Mb region strongly associated with sex through complementary SNP genotype and sequence-depth-based analyses. Comparative genomics of these proto-sex chromosomes within the species and across the Amaranthus genus demonstrates remarkable variability in their structure and genic content, including numerous polymorphic inversions. No such inversion underlies the extended linkage we observe associated with sex determination. Instead, we identify a complex presence/absence polymorphism reflecting substantial Y-haplotype variation-structured by ancestry, geography, and habitat-but only partially explaining phenotyped sex. Just over 10% of sexed individuals show phenotype-genotype mismatch in the sex-linked region, and along with observation of leakiness in the phenotypic expression of sex, suggest additional modifiers of sex and dynamic gene content within and between the proto-X and Y. Together, this work reveals a complex genetic architecture of sex determination in A. tuberculatus characterized by the maintenance of substantial haplotype diversity, and variation in the expression of sex.

Haplotypes

Rhesus blood group haplotype determination by nanopore sequencing and adaptive sampling enables the precise determination of complex allele combinations that could not be accurately determined by standard methods.

BACKGROUND: Patients with chronic transfusion needs such as those with sickle cell disease face a high risk of developing antibodies against high-prevalence antigens in the RH blood group system, complicating transfusion therapy and potentially necessitating stem cell transplantation. Molecular characterization of the RH system is hindered by hybrid alleles and high sequence homology between RHD and RHCE, limiting the effectiveness of conventional short-read sequencing. STUDY DESIGN AND METHODS: We analyzed 11 control and 20 patient samples, some of which could not be reliably genotyped by standard methods. RESULTS: Nanopore sequencing with adaptive sampling enables targeted, amplification-free long-read sequencing of the RH locus, resolving homologous and complex hybrid structures and enabling complete haplotype phasing for all samples, including samples that could not be accurately determined by standard methods like serology and short-read sequencing. Four new alleles were identified and for 13 out of 20 patients the results led to a change in the transfusion regimen. DISCUSSION: These findings show that nanopore sequencing with adaptive sampling allows unambiguous genotyping of the RH system, improves detection of complex variants, and supports better-matched transfusion strategies for chronically transfused patients.

Rh-Hr Blood-Group System

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 × depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Individualized antisense oligonucleotides for SCN2A-related developmental epileptic encephalopathy.

SCN2A variants are among the most common genetic causes of developmental and epileptic encephalopathies (DEEs), which can present with uncontrolled seizures at birth and account for 1-2% of all epileptic encephalopathies. A substantial fraction of causal variants are gain-of-function or mixed-function variants associated with increased channel open probability or greater sodium current flux. Here two parallel n = 1 clinical studies were conducted in two patients (9-year-old and 14-year-old boys) with SCN2A-related DEE. Individualized allele-selective antisense oligonucleotides (ASOs) were designed to target heterozygous intronic single-nucleotide polymorphisms (SNPs) for decreased expression of mutant SCN2A transcript while preserving the wild-type copy. Primary endpoints included quantitative change from baseline in seizure frequency and neurodevelopment, including motor scores. Efficacy measures were also individualized to each patient's phenotype, including refractory seizures, developmental delay, autism spectrum disorder, choreoathetosis and gastrointestinal dysfunction. Patients experienced a reduction in seizure frequency (26% and 90% in the two patients, respectively), decreased use of concomitant medications and improvement in neurodevelopmental skills. Both ASOs were well tolerated, with no ASO-related serious adverse events. Continued long-term follow-up of these preliminary positive safety and efficacy findings is needed to confirm the disease-modifying potential of these ASOs. Haplotype phasing in a separate cohort of infants with SCN2A-related disorder (SCN2A-RD), diagnosed by rapid whole-genome sequencing, identified 16% of patients with compatible SNPs. These data provide a pathway from n = 1 to n of more patients with SCN2A-RD and other monogenic disorders. ClinicalTrials.gov registration: NCT06314490 .

Adolescent

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Tubers are wide-spread across angiosperms with some species such as Solanum tuberosum L. (potato) serving as a vital crop for human consumption. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of Flowering Time, that functions as tuberigen, the equivalent of florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of Atlantic, a tetraploid cultivar. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129,218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Journal Article

An allelic resolution gene atlas for tetraploid potato provides insights into tuberization and stress resilience.

Tubers are modified underground stems that enable asexual, clonal reproduction and serve as a mechanism for overwintering and avoidance of herbivory. Potato (Solanum tuberosum L.) is cultivated for its tubers, which serve as a major crop. Genes responsible for tuber initiation and disease resistance have been characterized in potato including StSP6A, a homolog of flowering time, that functions as a tuberigen, the equivalent of a florigen. To elucidate additional molecular and genetic mechanisms underlying potato biology including tuber initiation, tuber development, and stress responses, we generated a developmental and abiotic/biotic-stress gene expression atlas from 34 tissues and treatments of the tetraploid potato cultivar, Atlantic. Using the haplotype-phased tetraploid Atlantic genome assembly and expression abundances of 129 218 genes, we constructed gene coexpression modules that represent networks associated with distinct developmental stages as well as stress responses. Functional annotations were given to modules and used to identify genes involved in tuberization and stress resilience. Structural variation from a pan-genomic analysis across four cultivated potato genome assemblies as well as domestication and wild introgression data allowed for deeper insights into the modules to identify key genes involved in tuberization and stress responses. This study underscores the importance of transcriptional regulation in tuberization and provides a comprehensive framework for future research on potato development and improvement.

Solanum tuberosum

Haplotype-aware long-read error correction.

Error correction of long reads is an important initial step in genome assembly workflows. For organisms with ploidy greater than one, it is important to preserve haplotype-specific variation during read correction. This challenge has driven the development of several haplotype-aware correction methods. However, existing methods are based on either ad-hoc heuristics or deep learning approaches. In this paper, we introduce a rigorous formulation for this problem. Our approach builds on the minimum error correction framework used in reference-based haplotype phasing. We prove that the proposed formulation for error correction of reads in de novo context, i.e., without using a reference genome, is NP-hard. To make our exact algorithm scale to large datasets, we introduce practical heuristics. Experiments using PacBio HiFi sequencing datasets from human and plant genomes show that our approach achieves accuracy comparable to state-of-the-art methods. Implementation: https://github.com/at-cg/HALE .

Clustering

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9 Gb and represented 23 pairs of chromosomes with an NG50 of 150 Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans

Biallelic VPS41 Variants in Autosomal Recessive Spinocerebellar Ataxia 29 Resolved by Long-Read Sequencing and RNA Analysis.

BACKGROUND: Biallelic variants in VPS41, encoding a subunit of the HOPS complex, cause autosomal recessive spinocerebellar ataxia 29 (SCAR29), a rare neurodevelopmental disorder with an incompletely defined phenotypic and molecular spectrum. METHODS: We investigated a 24-year-old man with cerebellar ataxia, hypotonia, and intellectual disability. Exome sequencing identified four candidate VPS41 variants. Because maternal DNA was unavailable, long-read genome sequencing was performed to determine allelic configuration, followed by RNA and protein analyses. RESULTS: In addition to typical SCAR29 features, the patient showed previously unreported findings, including swan-neck deformities and pes cavus. Long-read genome sequencing demonstrated that two VPS41 variants were in trans. RNA analysis revealed distinct splicing consequences: one allele produced an out-of-frame transcript predicted to undergo nonsense-mediated decay, whereas the other generated an in-frame exon-skipped transcript. These complementary defects reduced VPS41 expression at both transcript and protein levels, supporting pathogenicity and variant reclassification. CONCLUSION: Our findings expand the phenotypic spectrum of VPS41-related disease and highlight the value of long-read allelic resolution in clarifying pathogenic mechanisms in rare genetic disorders.

Humans

Haplotype-resolved reconstruction and functional interrogation of cancer karyotypes.

Complex karyotype changes are widespread in cancer genomes. A major gap in cancer genome characterization is the resolution of rearranged chromosomes with chromosome-length continuity. Here, we describe a two-tiered approach to determine the segmental composition of rearranged chromosomes with haplotype resolution. First, we present refLinker, a bioinformatic method for robust determination of chromosomal haplotypes using cancer Hi-C data. By contrast with existing methods, refLinker is insensitive to the presence of large-scale DNA deletions, duplications, and high-level amplification in cancer genomes. Second, we demonstrate a computational strategy to determine the segmental structure of rearranged chromosomes using haplotype-specific Hi-C contacts. We apply these methods to breast cancer genomes and provide direct evidence for long-range transcriptional changes associated with rearrangements of the inactive X chromosome. Together, these results highlight refLinker's broad utility for studying the functional consequences of chromosomal rearrangements.

Humans

A novel method for across-chromosome phasing without relative data.

MOTIVATION: Across-chromosome phasing identifies which haplotypes of different chromosomes come from the same parent. This differs from within-chromosome phasing, which uses linkage disequilibrium patterns to determine which alleles were co-inherited within each chromosome but does not match haplotypes across different chromosomes. While across-chromosome phasing can be conducted using genotypes from parents or close relatives, current methods perform poorly for samples of unrelated individuals. Here, we introduce a novel approach for across-chromosome phasing that employs a window-based SNP-similarity metric, eliminating the need for data from close relatives or detection of identical-by-descent haplotypes. RESULTS: Using UK Biobank offspring with both parents genotyped as a gold standard, we evaluated the performance of our method by phasing the offspring without using parental data. In genomic data with no within-chromosome phase errors, our algorithm achieved a mean across-chromosome phasing accuracy of 95%, with 53% of individuals phased perfectly. When data was pre-phased computationally using a standard within-chromosome phasing algorithm, mean accuracy for across-chromosome phasing dropped to 83.1%. Thus, our method is limited primarily by the accuracy of within-chromosome phasing accuracy and can approach near-perfect across-chromosome phasing accuracy as within-chromosome phasing accuracy improves. AVAILABILITY AND IMPLEMENTATION: The implementation was executed within a multi-node computational environment of University of Colorado Boulder Research Computing (Blanca Cluster: https://www.colorado.edu/rc/resources/blanca), employing parallelization techniques in the C programming language. The source code has been made publicly accessible online at https://github.com/emmanuelsapin/AcrossChromosomesPhasing, thereby facilitating reproducibility of the results for researchers with authorized access to the UK Biobank dataset.

Algorithms

Haplotype-resolved 3D genome maps reveal RNAPII-mediated allelic regulation in hybrid rice.

To understand how the two parental genomes coordinate transcription in hybrids, chromatin architecture must be resolved at the haplotype level. Here, using phased Bridge-Linker Hi-C, we reconstructed a haplotype-resolved three-dimensional (3D) genome of the elite hybrid rice (Oryza sativa) line Shanyou 63 (SY63). We identified extensive allele-specific chromatin conformations. Furthermore, we generated allele-resolved RNAPII ChIA-PET maps and phased transcriptomes to explore how chromatin interactions contribute to allelic regulation. Although maternal and paternal homologs share broadly similar chromatin features, we detected widespread haplotype-biased RNAPII binding and chromatin looping at high resolution. These allele-specific RNAPII-mediated contacts were significantly associated with biased expression. Stronger RNAPII binding on one haplotype promoted the formation of long-range regulatory loops with distal genes, thereby contributing to allele-biased transcription at a subset of loci, even when promoter-proximal RNAPII occupancy was comparable between alleles. These results demonstrate that subtle differences in RNAPII engagement and 3D regulatory wiring between parental haplotypes can reshape transcriptional output in hybrids, providing new insights into the mechanisms underlying the allelic regulation of gene expression.

Allele-specific chromatin interactions

Haplotype of multiple polymorphisms resolved by enzymatic amplification of single DNA molecules.

We have developed a reliable method for the direct resolution of haplotypes or linkage phase from individuals who are multiply heterozygous in a given genomic region. The method is based on single-molecule dilution (SMD) of genomic template and amplification via biphasic polymerase chain reaction (booster PCR). We have verified the feasibility of the SMD method for a highly polymorphic region within the beta-globin cluster by analysis of triply heterozygous individuals of known haplotype. This approach should be useful in many studies in population or evolutionary genetics and in a variety of clinical settings.

Base Sequence