Search PubMedSearch

SEARCH · Search PubMed

Results for “variant masking”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

11 recordsLinked to original sources

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

Clinically relevant pseudoexons of the GALNS gene and their antisense-based correction.

BACKGROUND: Biallelic pathogenic variants in the GALNS gene lead to Mucopolysaccharidosis Type IVA (MPS IVA), a rare lysosomal storage disorder. GALNS encodes the enzyme N-acetylgalactosamine-6-sulfatase, whose deficiency causes accumulation of glycosaminoglycans and leads to a broad spectrum of clinical manifestations primarily affecting the osteoarticular system. Several studies have shown that, in 10%-15% of patients with the biochemical phenotype of MPS IVA, standard molecular genetic testing fails to identify one or both causative variants in the GALNS gene. METHODS: We performed an in-depth investigation of GALNS' splicing, with a special focus on deep-intronic mutations that lead to activation of pseudoexons (PEs). Using bioinformatic tools, we analyzed all deep-intronic variants in GALNS available in public databases and subjected the most relevant ones to in vitro analyses using minigenes. RESULTS: We characterized eight PE-activating variants, one of which (c.121-210C > T) represents a recurrent pathogenic variant which has long been hidden behind the mask of a polymorphic variant. In addition, we demonstrate that GALNS' splicing can produce a diverse range of mRNA isoforms containing so-called wild-type PEs, which are present at low levels as part of non-productive splicing, and weak canonical exons which are prone to skipping. We show that PE-activating variants cluster within wild-type PEs, highlighting the need for closer scrutiny of these regions during genetic testing. Finally, we applied modified U7 small nuclear RNAs and circular RNAs to efficiently block the identified PEs and pave the way for personalized antisense-based therapy for MPS IVA patients. CONCLUSION: The results of this study expand the understanding of GALNS gene splicing, indicating hotspots for splicing mutations. The presented data not only help to increase the diagnostic yield for MPS IVA but also unveil new therapeutic approaches for a number of MPS IVA patients.

Humans

STX1B variant-specific synaptic dysfunction is associated with network hyperexcitability in human iPSC-derived neurons.

BACKGROUND: Variants in STX1B/syntaxin-1B are linked to a spectrum of fever-associated epilepsy syndromes. While studies in murine models have provided mechanistic insights, their relevance to human disease in a heterozygous context may be limited. METHODS: We investigated two pathogenic STX1B variants using isolated single neurons and neuronal network cultures derived from patient-specific induced pluripotent stem cells. These carried either a de novo p.G226R variant, associated with severe developmental epilepsy, or an InDel variant (p.K45delinsRCMIE/p.L46M) linked to a transient familial seizure syndrome. Synaptic function and network excitability were assessed using patch-clamp and multi-electrode array recordings, alongside morphological and transcriptomic profiling. FINDINGS: G226R exhibited both gain- and loss-of-function characteristics, with increased miniature excitatory postsynaptic current frequency in networks but not in autapses, and synaptic failure during sustained high-frequency stimulation. For the InDel variant, the predicted loss-of-function phenotype based on reduced syntaxin-1B levels was not detectable at the single-cell level, likely masked by compensatory synaptic upregulation. At the network level, however, both variants were associated with neuronal hyperexcitability, characterised by more frequent and prolonged bursting activity, with a much stronger phenotype in G226R-containing networks. Transcriptomic profiling revealed a differential dysregulation of synaptic and other neuronal genes. INTERPRETATION: The divergence between morphological, electrophysiological and transcriptomic findings suggests that compensatory mechanisms may contribute to network hyperexcitability. Initially engaged to maintain homoeostasis, they may ultimately contribute to a pathological network state. The graded severity of network alterations across STX1B variants correlates with the clinical phenotypes. FUNDING: BMBF (Treat ION-01GM2210A, SNAREopathies-01EW1809A), 2023 FEBS Summer Fellowship, Fortüne programme (2610-0-0), EKFS college precise.net, Open Access Publishing Fund of University of Tübingen.

Humans

Equity in genome sequencing for rare disease diagnosis: a cross-sectional analysis of data from the UK 100,000 Genomes Project.

BACKGROUND: Genome sequencing has improved rare disease diagnosis and is now part of routine clinical care in the National Health Service in England. Automated prioritisation pipelines narrow millions of variants per patient to a small subset for clinical review, a process that relies on allele frequency resources that do not fully represent human genetic diversity. We assessed ancestry-related differences in variant prioritisation and diagnostic outcomes in patients from the UK 100,000 Genomes Project. METHODS: We analysed 29,405 rare disease probands with genome sequencing and linked clinical outcomes data. We used multivariable regression to assess ancestry-related differences in the number of variants prioritised for clinical review, the proportion of prioritised variants that were recorded as diagnostic, and diagnostic yield. We also evaluated the use of ancestry-stratified allele frequency filters derived from an independent, diverse UK cohort (n = 33,724). FINDINGS: Compared with the European ancestry group, the East African group had nearly three times more variants prioritised for clinical review (IRR 2.77, 95% CI 2.33-3.29). Other non-European groups also had significantly higher counts. Diagnostic yield was similar across ancestry groups after adjustment (LRT p = 0.1650). Prioritised variants were less likely to be recorded as diagnostic in East African (OR 0.32, 95% CI 0.22-0.46), West African (0.47, 0.39-0.57), South Asian (0.65, 0.58-0.73), and Middle Eastern (0.68, 0.54-0.86) groups. Applying ancestry-stratified allele-frequency filters removed 3.1% of prioritised variants overall-24.3% in the East African group-without loss of diagnostic sensitivity, including 29.5% of recorded VUS in this group. INTERPRETATION: Differences in the likelihood of prioritised variants being recorded as diagnostic partly reflect limitations of current allele frequency resources, which use broad population groupings that mask within-group diversity. Increased representation of diverse ancestries in reference databases and better estimation of ancestry-appropriate allele frequencies will help reduce inefficiencies and improve equity in variant prioritisation for rare disease diagnosis. FUNDING: The UK Department of Health and Social Care and the EU's Horizon 2020 Research and Innovation Programme.

Humans

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics

Artificial intelligence-assisted detection and optical differentiation of colorectal lesions in Lynch syndrome surveillance (CADLY2): a multicentre, open-label, randomised controlled superiority trial.

BACKGROUND: Artificial intelligence (AI)-based computer-aided detection (CADe) systems improve adenoma detection in average-risk colorectal cancer screening. Meanwhile, evidence in Lynch syndrome surveillance is sparse and inconsistent. We assessed the effect of CADe on adenoma detection during Lynch syndrome surveillance. Computer-aided optical diagnosis (CADx) performance for optical differentiation of colorectal lesions was evaluated as a secondary aim. METHODS: CADLY2 was an international, multicentre, open-label, randomised controlled superiority trial at nine specialised hereditary cancer surveillance centres in Belgium, Germany, the Netherlands, and Spain. Adults aged 18 years or older with genetically confirmed Lynch syndrome scheduled for surveillance colonoscopy were randomly assigned (1:1) to high-definition white-light (HD-WL) colonoscopy alone or to HD-WL colonoscopy with computer-aided assistance from CAD EYE (Fujifilm, Tokyo, Japan). CAD EYE was used for CADe during withdrawal and for CADx after lesion detection. Randomisation was done centrally through a secure web-based system using Pocock's minimisation algorithm with a stochastic component and was stratified by centre, sex, previous colorectal cancer, underlying pathogenic variant, and interval since previous colonoscopy. Allocation concealment was ensured through the centralised web-based system. Patients were masked to group allocation until the start of withdrawal in procedures with mild sedation, or until completion of the procedure in procedures with propofol-based sedation. Endoscopists were not masked. The primary outcome was adenoma detection rate, defined as the proportion of patients with at least one histopathologically confirmed adenoma, analysed in the full analysis set (defined as all randomly allocated patients with available data for the primary outcome). The diagnostic performance of the CADx system was evaluated as a secondary outcome. The safety analysis set comprised all randomly allocated patients who underwent a study colonoscopy. This study is registered with the German Clinical Trials Register, DRKS00030695, and is completed. FINDINGS: Between May 9, 2023, and Oct 30, 2025, 757 patients were randomly allocated to HD-WL colonoscopy (377 patients) or to AI-assisted colonoscopy (380 patients); 733 patients were included in the full analysis set (369 HD-WL and 364 AI-assisted). The median age was 49 years (IQR 38-59) in the HD-WL group and 50 years (38-59) in the AI-assisted group; 213 (58%) were female and 156 (42%) male in the HD-WL group, and 207 (57%) were female and 157 (43%) male in the AI-assisted group. The adenoma detection rate was 30·9% (114 of 369 patients) with HD-WL versus 33·8% (123 of 364 patients) with CADe assistance (odds ratio 1·14 [95% CI 0·83-1·57], p=0·41). For CADx differentiation of neoplastic versus non-neoplastic lesions in the paired lesion-level analysis, with histopathology as the reference standard and sessile serrated lesions and traditional serrated adenomas classified as non-neoplastic, CADx sensitivity was 85·9% (95% CI 82·0-89·1) and specificity was 91·4% (89·4-93·0). Three adverse events occurred in the AI-assisted group: two mild post-polypectomy bleedings and one serious pulmonary embolism or deep venous thrombosis unrelated to the procedure. No adverse events occurred in the HD-WL group. INTERPRETATION: CADe-assisted colonoscopy did not show the absolute improvement in adenoma detection rate that was assumed in the prespecified sample-size calculation. CADx did not clearly improve lesion differentiation beyond expert optical diagnosis in expert Lynch syndrome surveillance settings. FUNDING: Third-party research funding of the National Center for Hereditary Tumor Syndromes, University Hospital Bonn.

Humans

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans

Long-read sequencing resolves complex CYP21A2 variants and identifies 2+0 carriers in 21-hydroxylase deficiency.

The complex CYP21A2 variants arising from high homology with its pseudogene CYP21A1P challenge the diagnosis of 21-hydroxylase deficiency (21-OHD). This study systematically evaluated long-read sequencing (LRS) for identifying complex structural variants of the CYP21A2 gene in 21-OHD in comparison with conventional molecular diagnostic methods, including multiplex ligation-dependent probe amplification (MLPA), CNVplex, and SNaPshot. Twenty patients with suspected 21-OHD and defined CYP21A2 structural variants identified via initial MLPA screening were enrolled. Variants were further analyzed using CNVplex and SNaPshot, then all samples underwent LRS for comprehensive variant detection, breakpoint mapping, and haplotype resolution. LRS overcame key limitations of conventional methods. It reliably identified a novel large-fragment deletion and defined its boundaries. Notably, LRS identified "2+0" carriers, where deletions masked by duplications cause false-negatives with standard techniques. Moreover, LRS accurately distinguished CYP21A1P/CYP21A2_CH-4 and CH-9 chimera subtypes which were indistinguishable by the combined conventional assays. Furthermore, LRS enabled the precise identification and characterization of TNXA/TNXB chimeric deletions. These are frequently misclassified as CYP21A1P/CYP21A2 chimeras by conventional methods but are critical for diagnosing associated conditions such as CAH-X syndrome. LRS provides a superior, integrated solution for the molecular diagnosis of 21-OHD, offering precise structural variant characterization, accurate carrier detection, and reliable breakpoint mapping. Its application enhances diagnostic accuracy, supports advanced genetic counseling, and paves the way for genotype-informed clinical management.

Journal Article

Beyond Bulk: Cell-Type-Resolved Epigenomics as the Path Forward in Alzheimer's Disease Research.

Alzheimer's disease (AD) is a complex neurodegenerative disorder in which most risk variants are noncoding and are enriched at gene regulatory regions, implicating epigenetic mechanisms as central mediators of disease pathogenesis. For most of the history of AD epigenetics research, bulk tissue analysis has dominated, obscuring the fundamentally distinct epigenomic landscapes of individual brain cell types and masking cell-type-specific contributions to disease. Advances in single-cell and single-nucleus sequencing, fluorescence-activated nuclei sorting and multiplexed epigenomic platforms have transformed this landscape, enabling cell-type-resolved profiling of chromatin accessibility, DNA methylation, histone modifications and transcription across the major neuronal, glial and neurovascular populations of the human brain. Here, we review these advances, structured around the argument that cell-type resolution is not a methodological refinement but a conceptual necessity. We describe the distinct epigenomic programs disrupted in neurons, microglia, astrocytes, oligodendrocytes and neurovascular cells in AD, highlighting how each cell type responds to pathology. We discuss the discovery of epigenomic erosion, the progressive loss of cell-type-specific epigenomic identity across virtually all brain cell populations as AD advances, as a unifying disease mechanism linking chromatin dysregulation to cognitive decline. Finally, we identify critical gaps in current knowledge, including the near-complete absence of cell-type-resolved histone modification and DNA methylation data for most brain cell types, the underrepresentation of rare populations in standard preparations and the untapped potential of metabolic acylation marks as indicators of the epigenome-metabolism interface in neurodegeneration.

Humans

Robust pleiotropy-decomposed polygenic scores identify distinct contributions to elevated coronary artery disease polygenic risk.

BACKGROUND: Polygenic risk score (PRS) have proved to offer robust risk prediction for coronary artery disease (CAD). However, the global CAD PRS summarizes the joint effects of all the markers in the genome, masking potential genetic heterogeneity that may be important for disease interpretation and targeted interventions. METHODS: Using summary-level data, we identified 43 significant CAD-related traits based on genetic correlations, and further classified them into eight pleiotropy clusters based on their biological functions. We then partitioned the genome into 2,353 near-independent regions. Variants in each region were assigned to the trait most genetically similar to CAD, and then were labeled with the corresponding pleiotropy cluster. We grouped variants without labels into a ninth, non-specific cluster. The Pleiotropy Decomposed (PD) PRSs for each of the nine clusters were calculated using variants assigned to each cluster for 407,903 samples of European ancestry from the UK Biobank (UKBB). RESULTS: We decomposed the CAD PRS into nine PD-PRSs and further stratified individuals with high CAD-PRS into nine subgroups. Each PD-PRS accounted for a higher proportion of the global CAD-PRS within its corresponding subgroup than in the remaining subjects with high CAD-PRS (e.g., 25.2% (0.07) vs. 10.06% (0.07) for lipids-PD-PRS). Additionally, these subgroups showed distinct clinical features. For example, in the lipids-related subgroup, lipoprotein(a) and LDL-cholesterol levels were 67.5% and 18.3% higher, respectively, compared to the remaining high-risk individuals. Furthermore, significant interactions were observed between blood pressure and BP PD-PRS, and between current smoking and respiratory system PD-PRS. CONCLUSION: Our findings suggest that PD-PRSs may reveal substantial genetic and phenotypic heterogeneity among individuals with high CAD-PRS. The unique PD-PRS compositions of each individual can highlight the relative importance of different pleiotropic regions.

Humans

Predicting dynamic expression patterns in budding yeast with a fungal DNA language model.

Predicting gene expression from DNA sequence remains challenging due to complex regulatory codes. We introduce a masked DNA language model pretrained on 165 fungal genomes closely related to budding yeast that captures conserved regulatory grammar. Fine-tuning the LM on yeast RNA-seq data-including high-resolution transcriptional regulator induction time courses generated in this study-yielded Shorkie, a model that substantially improves gene expression prediction compared to baselines trained without self-supervision. Shorkie identified canonical transcription factor (TF) binding motifs and tracked their usage across induction experiments. Furthermore, Shorkie accurately predicted variant effects, outperforming leading sequence-to-expression models in cis-eQTL classification and achieving high concordance with massively parallel reporter assays. Interpretability analyses revealed Shorkie's ability to resolve promoter dynamics, splicing signals, and temporal changes in regulatory motif usage. This framework demonstrates that evolutionary-scale pretraining combined with transfer learning substantially improves our ability to decode gene regulation from sequence, providing insights into noncoding variants and regulatory networks.

Journal Article