Search PubMedSearch

SEARCH · Search PubMed

Results for “Variant analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Refining the genetic diagnostic puzzle: A case report on a Chinese ARPKD patient with a reciprocal balanced translocation and c.2507 T > C (p.V836A) in PKHD1.

INTRODUCTION: Autosomal recessive polycystic kidney disease (ARPKD) ranks among the most severe chronic kidney diseases (CKD). Its primary cause is variants in the Polycystic Kidney and Hepatic Disease 1 gene (PKHD1). The clinical spectrum of ARPKD varies widely, ranging from mild late-onset symptoms to severe perinatal mortality. However, achieving an early genetic diagnosis in ARPKD patients before clinical symptoms appear proves challenging. CASE PRESENTATION: This case is a 4-year-old boy who experienced a convulsion characterized by a generalized tonic attack lasting approximately 3-5 minutes and later sought treatment to our hospital. However, routine abdominal ultrasound examination accidentally detected that he had diffuse liver lesions, splenomegaly, and bilateral renal enlargement with renal pelvis dilation. Given the uncertainty regarding the underlying cause of the patient's structural abnormalities and convulsions, karyotyping, whole exome sequencing (WES), structural variant analysis (SV analysis) of whole genome sequencing (WGS) were recommended. The result of SV analysis revealed that he has an RBT impacting PKHD1 and the precise location of breakpoints was confirmed through Long-Range Polymerase Chain Reaction (LR-PCR). However, WES did not screen out pathogenic variants initially, the WES data was reviewed subsequently based on SV analysis results. CONCLUSION: We identified an infrequent variant combination, c.2507T>C (p.V836A) in PKHD1 and an RBT with broken PKHD1, which extends the genetic spectrum of ARPKD, and provide a basis for further genetic counselling to the family.

Humans

PRISM: privacy-preserving rare disease analysis using fully homomorphic encryption.

MOTIVATION: Rare diseases affect millions of people worldwide, yet their genomic foundations remain poorly understood due to limited patient data and strict privacy regulations, such as the General Data Protection Regulation (GDPR) (https://gdpr.eu/tag/gdpr/) in March 2025. These restrictions can hinder the collaborative analysis of genomic data necessary for uncovering disease-causing variants. RESULTS: We present PRISM, a novel privacy-preserving framework based on fully homomorphic encryption (FHE) that facilitates rare disease variant analysis across multiple institutions without exposing sensitive genomic information. To address the challenges of centralized trust, PRISM is built upon a Threshold FHE scheme. This approach decentralizes key management across participating institutions and ensures no single entity can unilaterally decrypt sensitive data. Our method filters disease-causing variants under recessive, dominant, and de novo inheritance models entirely on encrypted data. We propose two algorithmic variants: a multiplication-intensive (MUL-IN) approach and an addition-intensive (ADD-IN) approach. The ADD-IN algorithms minimize the number of costly multiplication operations, enabling up to a 17× improvement in runtime for recessive/dominant filtering and 22× for de novo filtering, compared to MUL-IN methods. While ADD-IN produces larger ciphertexts, efficient parallelization via SIMD and multithreading allows it to handle millions of variants in reasonable time. To the best of our knowledge, this is the first study that utilizes FHE for privacy-preserving rare disease analysis across multiple inheritance models, demonstrating its practicality and scalability in a single-cloud setting. AVAILABILITY AND IMPLEMENTATION: The source code and the data used in this work can be found in https://github.com/mdppml/PRISM.git.

Computer Security

Cardiovascular prognostic impact of missense vs nonmissense lamin A/C variants: A systematic review and meta-analysis.

BACKGROUND: Variants in the LMNA gene, responsible for laminopathies, are associated with severe cardiovascular outcomes, including arrhythmias and heart failure (HF). However, the differential prognostic impact of missense vs nonmissense variants remains unclear. OBJECTIVE: The primary end point of this systematic review and meta-analysis was to compare the cardiovascular outcome defined as combined malignant ventricular arrhythmias (MVAs) and HF among patients with missense vs nonmissense variants in the LMNA gene. Secondary outcomes included a comparison of MVAs and HF-related events analyzed separately. METHODS: A systematic search of PubMed, Ovid MEDLINE, and Cochrane Library was conducted according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Meta-analyses were performed using fixed or random effects models, depending on heterogeneity. PROSPERO identifier: CRD42024584721. RESULTS: 12 studies comprising 1818 participants were included. Of these, 969 had missense variants, and 849 had nonmissense variants. The nonmissense group showed a significantly higher rate of cardiovascular events (30.5% vs 21.3%; odds ratio [OR] 2.22; P < .001). MVAs were more frequent in nonmissense carriers (25.5% vs 18.9%; OR 2.37; P < .001). Although limited by the small number of studies (n = 5) and single-study bias, the incidence of HF-related severe events seemed similar between the groups (18.2% vs 23.9%; OR 0.956; P = .801). CONCLUSION: Nonmissense LMNA variants are associated with worse cardiovascular outcomes, particularly arrhythmic events, whereas HF-related events seem comparable between nonmissense and missense variants.

Humans

Genetic architecture of postpartum psychosis: from common to rare genetic variation.

Postpartum psychosis is a severe psychiatric condition marked by the abrupt onset of psychosis, mania, or psychotic depression following childbirth. Despite evidence for a strong genetic basis, the roles of common and rare genetic variation remain poorly understood. Leveraging data from Swedish national registers and genomic data from the All of Us Research Program, we estimated family-based heritability at 55% and whole-genome sequencing-based heritability at 46%. Rare coding variant analysis identified HMGCR as a gene in which rare damaging variants confer risk for postpartum psychosis (FDR&#x2009;<&#x2009;0.05). Analyses of 240,009 participants from the All of Us Research Program and 58,990 participants from the Mount Sinai BioMe Biobank identified significant associations linking deleterious rare variants in HMGCR to vascular dementia and mental disorder, not otherwise specified, supporting the gene's broader psychiatric relevance. Additionally, among the top 200 genes ranked by association statistics, 17% of bipolar disorder, 21% of schizophrenia, and 16-25% of multiple autoimmune disorders exhibit a possible association with postpartum psychosis. These findings reveal unique genetic contributions and shared pathways, providing a foundation for understanding pathophysiology and advancing therapeutic strategies.

Humans

The Data Distillery: A Graph Framework for Semantic Integration and Querying of Biomedical Data.

The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.

Journal Article

Whole-genome sequencing implicates rare, low-frequency and structural non-coding variation at the SCN5A locus in Brugada syndrome.

Brugada syndrome (BrS) is an inherited cardiac condition characterized by a hallmark ECG pattern and an increased risk of sudden cardiac death. Central to the aetiology of BrS, the SCN5A region harbours both common non-coding risk variants and rare coding variants that are causative in approximately 20% of patients. However, rare non-coding genetic variation in this region remains largely unexplored. Here, we used whole-genome sequencing (WGS) of 752 European-ancestry BrS cases and 1,827 ancestry-matched controls to identify BrS-associated rare non-coding genetic variation at the SCN5A locus. Sliding-window and cis-regulatory element (CRE)-based rare-variant aggregate testing implicated three conserved CREs, including a dense aggregation of case singleton variants within a 178 bp enhancer in intron 17 of SCN5A which replicated in an independent BrS cohort. Prioritised BrS-associated rare and low-frequency non-coding variants within these elements were predicted to alter cardiac transcription factor motifs, and altered CRE activity in hiPSC-CM luciferase assays or were associated with BrS-relevant ECG endophenotypes in the UK Biobank. Single-variant analysis across the region identified a Bonferroni-significant five-fold case-enriched low-frequency variant within a known CRE in intron 1 of SCN5A, which replicated, was associated with slower cardiac conduction in the UK Biobank and accounted for part of the BrS GWAS signal at this locus. Structural variant analyses identified a 10.5 kb deletion upstream of SCN5A in a BrS case that encompassed a cardiac CRE and reduced sodium current density in a hiPSC-CM model, as well as a 6 kb BrS-enriched retrotransposon insertion in SCN5A that appeared to underlie part of the GWAS signal in this region. Together, these findings implicate rare and low-frequency non-coding variation at the SCN5A locus in BrS susceptibility and demonstrate the value of targeted WGS analysis of key disease loci.

Journal Article

HAP-SAMPLE2: data-based resampling for association studies with admixture.

MOTIVATION: HAP-SAMPLE2 extends the functionality of the original HAP-SAMPLE tool for simulating genotype-phenotype data, now with features to handle population admixture and rare variant analysis. It allows users to define parameters such as disease prevalence and allele effect sizes for both common and rare variant simulations. RESULTS: HAP-SAMPLE2 provides an efficient means for simulating complex datasets, suitable for large-scale projects like the 1000 Genomes Project. Its capabilities for population admixture allow users to create admixed populations or preserve substructures while introducing novel variation through artificial recombination. Additionally, the tool supports burden testing for rare variants using fixed and Madsen-Browning weighting schemes. AVAILABILITY AND IMPLEMENTATION: The software, along with a detailed vignette, is available on GitHub: https://github.com/M3dical/HAPSAMPLE2.

Software

Population-level genomic surveillance of human norovirus using wastewater-based whole-genome sequencing.

Wastewater-based surveillance has garnered increasing attention as a valuable approach for capturing community-level infection dynamics that are often difficult to detect through clinical reporting systems alone. In this study, we analyzed human norovirus genotype distributions and whole-genome-level variations in wastewater samples collected in Gwangju, Korea. These results were interpreted in conjunction with a documented foodborne outbreak to evaluate the epidemiological relevance of wastewater-based monitoring. Human norovirus concentrations were quantified using TaqMan Array Card-based RT-qPCR, and whole-genome next-generation sequencing (NGS) was performed to obtain viral read counts and reads per kilobase per million filtered reads values. Overall, strong correlations were observed between RT-qPCR-based concentrations and NGS-derived metrics. Genotype dynamics varied among wastewater treatment plants, reflecting differences in catchment size and local population characteristics. In particular, the relative abundance of GII.17[P17] increased during epidemiological week 50, temporally coinciding with a documented local foodborne outbreak. Variant analysis revealed that wastewater samples exhibited mixed nucleotide patterns, with multiple alleles coexisting at varying relative frequencies rather than fixed substitutions. Notably, some nonsynonymous variants detected in clinical samples were also observed in wastewater samples collected surrounding the outbreak period. Together, these findings demonstrate that wastewater-based whole-genome surveillance can capture both genotype-level shifts and nucleotide-level dynamics at the population scale, highlighting its potential as a complementary tool for monitoring community-level norovirus circulation and outbreak-associated genotype dynamics.IMPORTANCEWastewater-based surveillance is increasingly recognized as a promising approach for capturing community-level infection dynamics that are often missed by clinical surveillance. In this study, we applied whole-genome sequencing to wastewater samples collected in Gwangju, South Korea, to comprehensively characterize human norovirus genotype distributions and genetic variation. Distinct genotype patterns were observed across wastewater treatment plants, reflecting differences in catchment population size and local characteristics. Notably, an increase in the GII.17[P17] genotype detected in wastewater coincided with a foodborne outbreak investigated in Gwangju, demonstrating the potential of wastewater surveillance to reflect ongoing community transmission and emerging outbreak-associated genotypes. In addition, wastewater samples contained diverse and coexisting genetic variants, capturing population-level viral diversity and evolutionary dynamics that are not readily detected through clinical surveillance alone. These findings highlight the value of wastewater-based whole-genome surveillance for monitoring community-level viral circulation and support its integration as a complementary strategy to existing clinical surveillance systems.

genotype dynamics

Genetic analysis of three patients from two unrelated Chinese families with autosomal recessive spastic ataxia of Charlevoix-Saguenay.

Autosomal recessive spastic ataxia of Charlevoix-Saguenay (ARSACS) is a rare early-onset neurodegenerative disorder characterized by progressive cerebellar ataxia, spasticity, and sensorimotor peripheral neuropathy. This disorder is caused by homozygous or compound heterozygous variants in the sacsin (SACS) gene on chromosome 13q12.12. Three patients with ARSACS from two unrelated Chinese families were recruited for this study. Patient #1 was an 18-year-old male who had been walking unstably for 12 years. Patient #2, the younger sister of Patient #1, was a 5-year-old girl who had been walking unstably for 2 years. Patient #3 was a 19-year-old female who had been walking unstably and a tendency to fall for 17 years. For Patient #1, whole-exome sequencing (WES) identified a hemizygous variant c.8310_8313delAGAT (p.Asp2771fs4*) in SACS (NM_014363.6), with the father being heterozygous, the mother wild-type, and Patient #2 hemizygous, as verified by Sanger sequencing. Additional copy number variant analysis of the WES data indicated that Patient #1 had a heterozygous gross deletion of chr13q12.12 (chr13:23,808,732&#x2009;-&#x2009;24,890,322). Low-coverage whole-genome sequencing results revealed that Patient #2 carried a chr13q12.12 deletion (chr13:23,520,000-24,940,000). Together with Sanger sequencing results, this gross deletion was speculated to have been inherited from the mother, further explaining the hemizygous state of c.8310_8313delAGAT (p.Asp2771fs4*) in Patients #1 and #2. Through WES, Patient #3 was identified as having suspected compound heterozygous variants of c.2881&#xa0;C&#x2009;>&#x2009;T (p.Arg961*) and c.6409&#xa0;C&#x2009;>&#x2009;T (p.Gln2137*), inherited from the father and mother, respectively, as confirmed by Sanger sequencing. This study identified three variants in SACS. The c.8310_8313delAGAT (p.Asp2771fs4*) is novel, whereas c.2881&#xa0;C&#x2009;>&#x2009;T (p.Arg961*) and c.6409&#xa0;C&#x2009;>&#x2009;T (p.Gln2137*) have been reported previously. Moreover, this study highlights the growing trend that ARSACS has become increasingly prevalent worldwide rather than being localized to a specific region or race. As an increasing number of patients with ARSACS are diagnosed, the genetic spectrum of ARSACS will gradually broaden, providing an accurate genetic basis for prenatal diagnosis of mothers in the years ahead, if possible.

Adolescent

Mitigating pH-induced instability in deruxtecan-based ADCs: an onboard-mixing icIEF approach for robust charge heterogeneity characterization.

Accurate charge variant analysis of antibody-drug conjugates (ADCs) is essential for understanding product heterogeneity and ensuring quality control. However, Deruxtecan (DXd)-based ADCs present a unique analytical challenge due to the intrinsic instability of the payload, where the lactone ring readily undergoes hydrolysis under alkaline conditions, resulting in time-dependent shifts in charge distribution during imaged capillary isoelectric focusing (icIEF). In this study, we describe the development of an onboard-mixing icIEF method designed to minimize pH-induced degradation during sample preparation. By separating ADC samples from carrier ampholytes (CAs) prior to injection and enabling real-time mixing within the instrument, this approach effectively suppresses premature lactone ring opening and stabilizes charge variant profiles. Comparative studies between conventional premixing and onboard-mixing approach demonstrated that the latter significantly enhances reproducibility, particularly for acidic variants that are highly sensitive to structural conversion. Comprehensive method validation confirmed excellent precision, linearity, and sensitivity, with consistent performance across run-to-run and intra-day analyses. The results underscore the importance of controlling microenvironmental pH exposure in the analysis of chemically instable ADCs. The proposed onboard-mixing strategy provides a robust and efficient solution for icIEF-based characterization, reducing analytical artifacts while simplifying method development. This approach is broadly applicable to ADCs and other biotherapeutics containing pH-sensitive functional groups.

Hydrogen-Ion Concentration

Mechanism of age-related accumulation of mtDNA mutations in human blood.

Accumulation of mutant mitochondrial DNA (mtDNA) heteroplasmy is among the strongest signatures of ageing1. Here we investigated the underlying mechanism by calling mtDNA sequence, mtDNA abundance and mtDNA heteroplasmic variants in human blood using whole-genome sequences from approximately 750,000 individuals. We observed that mtDNA single-nucleotide variants (mtSNVs) accumulate sharply at age 60 years, occur at low levels of heteroplasmy, exhibit little evidence of positive selection and are likely to be predominantly neutral. The mutational spectrum of mtSNVs does not reflect oxidative lesions, as is commonly invoked, but is more consistent with mtDNA replication errors. To understand why mtSNVs become detectable with age, we performed a genome-wide association study for heteroplasmic mtSNV burden, identifying germline variants near TERT, TCL1A and SMC4, all of which have been linked to clonal haematopoiesis (CH)2. Rare-variant analysis also showed that high mtSNV burden is associated with mutations in numerous CH driver genes. These genetic associations persisted&#xa0;even after exclusion of individuals with known CH driver mutations. Our results support a model in which 'cryptic' mtDNA mutations initially arise randomly as replication errors but are undetectable in bulk. They then become apparent only through age-related expansion of cellular clones in blood. We propose that the high copy number and mutation rate of mtDNA make it a sensitive blood-based marker of somatic mosaicism due to CH. Our work mechanistically unifies three prominent signatures of ageing: common germline variants in TERT, CH and observed accrual of&#xa0;mtDNA mutations.

Humans

Optimized methods for the targeted surveillance of extended-spectrum beta-lactamase-producing Escherichia coli in human stool.

Understanding transmission pathways of important opportunistic, drug-resistant pathogens, such as extended-spectrum beta-lactamase (ESBL)-producing Escherichia coli, is essential to implementing targeted prevention strategies to interrupt transmission and reduce the number of infections. To link transmission of ESBL-producing E. coli (ESBL-EC) between two sources, single-nucleotide resolution of E. coli strains, as well as E. coli diversity within and between samples, is required. However, the microbiological methods to best track these pathogens are unclear. Here, we compared different steps in the microbiological workflow to determine the impact different pre-enrichment broths, pre-enrichment incubation times, selection in pre-enrichment, selective plating, and DNA extraction methods had on recovering ESBL-EC from human stool samples, with the aim to acquire high-quality DNA for sequencing and genomic epidemiology. We demonstrate that using a 4-h pre-enrichment in Buffered Peptone Water, plating on cefotaxime-supplemented MacConkey agar and extracting DNA using Lucigen MasterPure DNA Purification kit improves the recovery of ESBL-EC from human stool and produced high-quality DNA for whole-genome sequencing. We conclude that our optimized workflow can be applied for single-nucleotide variant analysis of an ESBL-EC from stool.IMPORTANCEDrug-resistant infections are increasingly difficult to treat with antibiotics. Preventing infections is thus highly beneficial. To do this, we need to understand how drug-resistant bacteria spread to take action to stop infection and transmission. This requires us to accurately trace these bacteria between different sources. In this study, we compared different laboratory methods to see which worked best for detecting extended-spectrum beta-lactamase (ESBL)-producing E. coli, a common cause of urinary tract or bloodstream infections, from human stool samples. We found that enriching stool in a nutrient broth for 4 h, then plating the bacterial suspension on antibiotic-selective MacConkey agar, and finally extracting DNA from the bacteria using a specific DNA purification kit resulted in improved recovery of ESBL E. coli and high-quality DNA. Sequencing multiple isolates from stool allowed us to distinguish unambiguously and at high resolution between different variants of ESBL E. coli present in stool.

Humans

Association of FOXC1 Duplications With Juvenile Open-Angle Glaucoma.

IMPORTANCE: While FOXC1 single-nucleotide variants and deletions are well-established causes of Axenfeld-Rieger syndrome, few FOXC1 duplications have been reported. This study investigated families with duplications encompassing the FOXC1 gene to refine the associated phenotypic spectrum and contribution to glaucoma. OBJECTIVE: To investigate the prevalence and phenotype of FOXC1 duplications in 2 large glaucoma registries. DESIGN, SETTING, AND PARTICIPANTS: This retrospective observational genetic cohort study included participants recruited from the Australian & New Zealand Registry of Advanced Glaucoma (ANZRAG) and the Massachusetts Eye and Ear (MEE) cohort from 2008 through 2025. Participants with glaucoma, and available relatives, underwent genomic testing to identify duplications encompassing FOXC1 using exome sequencing and genotyping arrays (ANZRAG) or whole-genome sequencing (MEE). Data analyses were conducted from 2022 through 2025. MAIN OUTCOMES AND MEASURES: Prevalence of FOXC1 duplications, age at glaucoma onset, and phenotype, including ocular and systemic features. RESULTS: Twenty individuals from 10 families (50% female and 50% male; 70% self-described as broadly European [Australian/British, British, English/German, English/Polish, European, or Scottish], 25% as Asian [Chinese or Filipino], and 5% as Latin American [Salvadoran]) were identified with FOXC1 duplications. All genetically tested individuals were diagnosed with glaucoma, demonstrating high penetrance. Seventeen individuals were referred with juvenile open-angle glaucoma (JOAG), 1 with primary open-angle glaucoma, 1 with primary congenital glaucoma, and 1 with anterior segment dysgenesis. The diagnosis of 4 individuals from 1 family with ectropion uveae was revised to anterior segment dysgenesis. Systemic features were reported for 2 participants (10.5%), including subtle dental findings and mild facial dysmorphism. Duplications encompassing FOXC1 were among the most common monogenic contributors to JOAG. In the ANZRAG group, they accounted for 13.5% (95% CI, 6.7%-25.3%) of JOAG probands with a genetic diagnosis, second to MYOC (53.8%; 95% CI, 40.5%-66.7%). In the MEE group, FOXC1 duplications accounted for 9.5% (95% CI, 2.7%-28.9%) of JOAG probands with a genetic diagnosis. CONCLUSIONS AND RELEVANCE: These findings suggest FOXC1 duplications are an underrecognized, highly penetrant, but variably expressive, genetic variation associated with JOAG. Findings for the relatively modest number of individuals in the retrospective study were associated with wide confidence intervals. This limitation is often inherent to studies of JOAG, a rare condition for which individual genetic variants account for only a subset of cases. Despite this, the findings highlight the genetic heterogeneity of JOAG and support the potential importance of considering routine genetic copy-number variant analysis for individuals with JOAG.

Humans

A Novel PTPN2 Isoform Differentially Regulates Immune Response.

Genome-wide association studies implicate the PTPN2 gene locus (18p11.21) in risk for several autoimmune diseases, including inflammatory bowel disease. Through genetic fine mapping, we identified the single-nucleotide polymorphism rs80262450 in the PTPN2 gene as the putative causal variant. Analysis of GTEx tissue samples and genetically engineered myeloid cell lines carrying risk and nonrisk alleles of rs80262450 demonstrated increased expression of the PTPN2 splice isoform 4 (PTPN2.4), suggesting that the rs80262450 enhances disease susceptibility by favoring production of PTPN2.4. Furthermore, we found that PTPN2.4 contains a nuclear export sequence (NES) that leads to its retention in the cytoplasm. Differential localization of PTPN2.4 isoform results in a distinct protein binding profile revealed by mass-spectrometry analysis, and its overexpression increased TNF-&#x3b1;. PTPN2.4 knockdown reduced pro-inflammatory cytokines in human macrophages. Mutations within the NES motif abolished the unique localization and function of PTPN2.4. Lastly, increased expression of PTPN2.4 was found in Crohn's disease tissues, demonstrating its involvement in the disease. Together, we identified the pathogenic isoform PTPN2.4 as a novel driver of intestinal inflammation and a potential target to attenuate inflammation in IBD.

Humans

ONT-only genome assembly of a Korean male individual using a semen sample.

BACKGROUND: Long-read sequencing has enabled the generation of high-quality human genome assemblies, but many previous assemblies were based on blood-derived DNA and often relied on limited data types from a single sequencing strategy. OBJECTIVE: This study aimed to generate high-quality phased genome assemblies of a Korean individual using multiple independent long-read datasets produced from a single sequencing platform and to evaluate their utility for chromosome-scale assembly and variant detection. METHODS: Genomic DNA was extracted from a semen sample of a Korean male. Long-read, ultra-long-read, and chromatin conformation capture sequencing data were generated using Oxford Nanopore Technologies. These datasets were integrated to construct phased genome assemblies, followed by correction of noticeable phasing errors and assessment of assembly continuity, chromosomal representation, telomeric repeat recovery, and variant detection performance. RESULTS: The final phased assemblies spanned approximately 2.9&#xa0;Gb and represented 23 pairs of chromosomes with an NG50 of 150&#xa0;Mb. Telomeric repeats were detected at 36 and 37 of the 48 chromosomal ends in the two assemblies, indicating high end-to-end completeness. In addition, we successfully identified structural variants, including small variants. These results demonstrate that combining multiple Oxford Nanopore data types can produce highly continuous and informative phased human genome assemblies. CONCLUSIONS: We generated high-quality phased genome assemblies of a Korean individual using Oxford Nanopore long-read sequencing data derived from semen DNA. This publicly available genome resource will support broader applications of long-read sequencing in human genomics and variant analysis.

Humans

Genetic diversity of clinical Mycobacterium bovis BCG isolates from an immunocompromised patient with BCG infection.

BACKGROUND: The Bacillus Calmette-Gu&#xe9;rin (BCG) vaccine is widely administered to prevent severe tuberculosis but can cause serious adverse events, including disseminated BCGosis, in immunocompromised individuals. However, studies investigating the in vivo genetic adaptation and microevolution of this live-attenuated vaccine during prolonged infection remain limited. METHODS: Two clinical Mycobacterium bovis BCG isolates (BCG01 and BCG02) and a lot-matched vaccine strain (VAC) underwent whole-genome sequencing. Phenotypic drug susceptibility testing was performed on the clinical isolates. Genomic relatedness was assessed using SNP-distance clustering and maximum-likelihood phylogeny against global reference strains. Comparative variant analysis was performed to identify mutations specific to BCG01 and BCG02 relative to VAC, and genotypic drug resistance was assessed using TB-Profiler. RESULTS: Phylogenomic analyses and SNP distance confirmed that both clinical isolates were derived from the BCG Tokyo 172 vaccine strain. BCG02 exhibited twice the mutational burden of BCG01, acquiring mutations in genes associated with cell wall biosynthesis (mas, ppsA), regulatory adaptation (pknK, dnaA), and surface antigens (pecA, PE/PPE). Crucially, whereas BCG01 remained susceptible to first-line drugs, BCG02 acquired a canonical rpoB Ser450Leu mutation (100% frequency) and a heteroresistant inhA Ile194Thr mutation (13% frequency), resulting in multidrug-resistant (MDR) BCGosis. CONCLUSION: Although structurally stable, the BCG Tokyo 172 vaccine strain can undergo rapid, clinically significant microevolution and clonal selection within immunocompromised hosts. The in vivo acquisition of multidrug resistance underscores the critical need for pre-vaccination immune screening and comprehensive laboratory monitoring of BCG-associated adverse events.

BCGosis

EV DNA from pancreatic cancer patient-derived cells harbors molecular, coding, non-coding signatures and mutational hotspots.

DNA packaged into cancer cell-derived EV is not well appreciated. Here, we uncovered signatures of EV DNA secreted by pancreatic cancer cells. The cancer cells and non-cancer counterparts exhibit distinct low vs. high molecular weight (LMW vs. HMW) EV DNA fragments distribution, respectively. Genome sequencing and Single Nucleotide Variants analysis revealed that 95% of reads and 94% of SNVs map to noncoding regions of the genome. Given that ~1% of the human genome represents coding regions, the 5% mapping rate to coding regions suggests a non-random enrichment of certain coding regions and mutations. The LMW DNA fragments not only set cancer cells apart, but also harbor cancer specific enrichment of unique coding regions, the top nine being FAM135B, COL22A1, TSNARE1, KCNK9, ZFAT, JRK, MROH5, GSDMD, and MIR3667HG. Additionally, the cancer cells' LMW DNA fragments exhibit dense centromeric mapping more strikingly on chromosomes 3, 7, 9, 10, 11, 13, 17, and 20. Mutational profiling turned up close to 200 mutations specific for the cancer cells. Altogether, our analyses suggest that centromeric regions might hold clues to EV DNA content from pancreatic cancer, the molecular, mutational signatures thereof, and rationalizes the need for a new approach to DNA biomarker research.

Humans

Utilization of long-read sequencing for the detection of structural rearrangements with AgileStructure.

MOTIVATION: Changes in genome organisation contribute to genetic disease when they disrupt gene function or regulation. Structural rearrangements may interrupt coding sequence or alter expression through promoter loss or gain, chromatin changes, copy-number variation, or disruption of short-range regulatory elements. Although short-read sequencing excels at detecting small variants, it performs poorly at resolving breakpoints of large rearrangements, especially in repetitive or low-complexity regions. Long-read sequencing overcomes these limitations, but analytical tools have not kept pace, making accurate identification and annotation of large structural variants challenging. RESULTS: We developed AgileStructure, a desktop application for locating and annotating large&#x2011;scale genomic rearrangements using aligned long&#x2011;read data. The software enables user&#x2011;guided exploration of breakpoint&#x2011;spanning reads, supporting accurate interpretation of complex events and filling a key gap in current structural variant analysis workflows. AVAILABILITY AND IMPLEMENTATION: Source code, binaries, user guide, and example aligned read data, are available on GitHub: https://github.com/msjimc/AgileStructure. An archived version is also available on Zenodo at https://doi.org/10.5281/zenodo.18610110.

Software