Search PubMedSearch

SEARCH · Search PubMed

Results for “single nucleotide variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software

β1- and β2-adrenergic Receptor Haplotypes Regulate Therapeutic Responses to Placebo and the Biased Ligand β-blocker Bucindolol.

BACKGROUND: ADRB1 and ADRB2, encoding cardiac myocyte &#x3b2;1- and &#x3b2;2-adrenergic receptors (ARs) that mediate pathologic myocardial remodeling in response to chronically increased signaling, contain N-terminus haplotype variants capable of influencing agonist- or biased ligand-induced receptor internalization that uncouples canonical signaling and initiates EGFR/ERK1/2 cardioprotection. METHODS: In two heart failure (HF) clinical trial genetic substudies we investigated effects of internalizing vs. internalization-resistant ADRB1/ADRB2 haplotypes on clinical or biomarker responses to the biased ligand &#x3b2;-blocker bucindolol vs. placebo or vs. the nonbiased &#x3b2;1-antagonist metoprolol, and in haplotyped isolated human heart preparations we measured ERK1/2 activation in response to these same interventions. RESULTS: In subjects with &#x2265;3 internalizing ADRB1+ADRB2 haplotypes (6.7% subcohort) placebo treatment was associated with fewer clinical events compared to subjects with internalization-resistant haplotypes (Odds Ratio (OR) 0.28, 95% CI (0.10, 0.82)). In contrast, placebo treatment in subjects with &#x2265;3 internalization-resistant haplotypes (70% subcohort) was associated with more clinical events in comparison to subjects with internalizing haplotype counterparts (OR 1.64 (1.46, 1.84)). Bucindolol treatment was equal to placebo in the &#x2265;3 internalizing subcohort, but was superior to placebo in the internalization-resistant subcohort (bucindolol vs. placebo OR 0.49 (0.41, 0.58)). In subjects with all 4 haplotypes internalization-resistant (25% subcohort), bucindolol vs. placebo reduced time to first event rates by 62.3&#xb1;17.5% (P <0.01, 1.68&#xb1;0.34 fold > the all-haplotypes parent population and additive to 1.92&#xb1;0.58 fold when the ADRB1 haplotype contained Arg389 rather than Gly389). The same bucindolol vs. placebo pattern was observed for NT-proBNP or norepinephrine reduction vs. metoprolol. In these comparisons ADRB2 and ADRB1 haplotypes behaved similarly, and although the haplotypes differed in frequency between Black and non-Black subjects, within haplotypes there were no by-race differences in therapeutic effects. Bucindolol but not metoprolol activated ERK1/2 signaling in isolated ventricular preparations with &#x2265;3 internalization-resistant haplotypes. CONCLUSIONS: 1) Both &#x3b2;1- and &#x3b2;2-AR haplotypes regulate therapeutic responses in HF; internalizing species confer protection against clinical events in placebo-treated subjects, while in internalization-resistant haplotypes the biased ligand &#x3b2;-blocker bucindolol but not the non-biased ligand metoprolol is associated with favorable effects. 2) The biased ligand cardioprotective effect may be related to internalization-dependent or -independent ERK1/2 activation.

Beta Adrenergic Receptors

MutBERT: probabilistic genome representation improves genomics foundation models.

MOTIVATION: Understanding the genomic foundation of human diversity and disease requires models that effectively capture sequence variation, such as single nucleotide polymorphisms (SNPs). While recent genomic foundation models have scaled to larger datasets and multi-species inputs, they often fail to account for the sparsity and redundancy inherent in human population data, such as those in the 1000 Genomes Project. SNPs are rare in humans, and current masked language models (MLMs) trained directly on whole-genome sequences may struggle to efficiently learn these variations. Additionally, training on the entire dataset without prioritizing regions of genetic variation results in inefficiencies and negligible gains in performance. RESULTS: We present MutBERT, a probabilistic genome-based masked language model that efficiently utilizes SNP information from population-scale genomic data. By representing the entire genome as a probabilistic distribution over observed allele frequencies, MutBERT focuses on informative genomic variations while maintaining computational efficiency. We evaluated MutBERT against DNABERT-2, various versions of Nucleotide Transformer, and modified versions of MutBERT across multiple downstream prediction tasks. MutBERT consistently ranked as one of the top-performing models, demonstrating that this novel representation strategy enables better utilization of biobank-scale genomic data in building pretrained genomic foundation models. AVAILABILITY AND IMPLEMENTATION: https://github.com/ai4nucleome/mutBERT.

Humans

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

Chloroplast Haplotype Analysis Reveals High Genetic Similarity Among Central Asian Prunus Species.

Genetic variation in four wild Prunus taxa (P. fruticosa, P. erythrocarpa, P. verrucosa and P. griffithii var. tianshanica) was investigated for the first time using six chloroplast DNA regions (matK, r rpl16, ycf1_1, ycf1_2, ndhF and trnH-psbA) analysed through CAPS-based SNP detection. The results revealed weak chloroplast differentiation among P. erythrocarpa, P. verrucosa and P. griffithii var. tianshanica. However, chloroplast variation exhibited a strong geographic signal across the studied populations. The observed chloroplast variation primarily reflected geographic structuring rather than clear differentiation among these closely related taxa. In contrast, P. fruticosa showed distinct chloroplast haplotypes not shared with the other taxa. These findings demonstrate that the developed chloroplast CAPS marker system is effective for detecting chloroplast haplotype variation but has limited discriminatory power among closely related wild Prunus taxa. Further studies using nuclear markers and genome-wide approaches will be required to better resolve their genetic relationships and evolutionary history.

Haplotypes

Parallel genetic adaptation amid a background of changing effective population sizes in divergent yellow perch (Perca flavescens) populations.

Aquatic ecosystems are highly dynamic environments vulnerable to natural and anthropogenic disturbances. High-economic-value fisheries are one of many ecosystem services affected by these disturbances, and it is critical to accurately characterize the genetic diversity and effective population sizes of valuable fish stocks through time. We used genome-wide data to reconstruct the demographic histories of economically important yellow perch (Perca flavescens) populations. In two isolated and genetically divergent populations, we provide independent evidence for simultaneous increases in effective population sizes over both historic and contemporary time scales including negative genome-wide estimates of Tajima's D, 3.1 times more single nucleotide polymorphisms than adjacent populations, and contemporary effective population sizes that have increased 10- and 47-fold from their minimum, respectively. The excess of segregating sites and negative Tajima's D values probably arose from mutations accompanying historic population expansions with insufficient time for purifying selection, whereas linkage disequilibrium-based estimates of Ne also suggest contemporary increases that may have been driven by reduced fishing pressure or environmental remediation. We also identified parallel, genetic adaptation to reduced visual clarity in the same two habitats. These results suggest that the synchrony of key ecological and evolutionary processes can drive parallel demographic and evolutionary trajectories across independent populations.

Animals

Genomic diversity, inbreeding, and selection signatures in duroc, landrace, and yorkshire pigs from a long-term closed breeding system.

Duroc (DD), Landrace (LL), and Yorkshire (YY) are among the most widely used commercial pig breeds, having undergone intense long-term selection within closed breeding systems. This study presents a comprehensive genomic analysis of genetic diversity, inbreeding patterns, and selection signatures in DD, LL, and YY populations that have been subject to close breeding for over 15 years. Genomic and pedigree data were available for 1,088 animals (DD&#x2009;=&#x2009;348, LL&#x2009;=&#x2009;276, YY&#x2009;=&#x2009;464), genotyped using the GenoBaits&#xae; Porcine 100&#xa0;K SNP panel. Principal component analysis and genetic diversity metrics revealed distinct population structures among the three breeds. Pairwise genetic differentiation supported this pattern, with DD showing the greatest divergence from LL (0.34&#x2009;&#xb1;&#x2009;0.24) and YY (0.33&#x2009;&#xb1;&#x2009;0.24), while LL and YY were more closely related (FST&#x2009;=&#x2009;0.22&#x2009;&#xb1;&#x2009;0.19). Linkage disequilibrium (LD) analysis further confirmed these differences, as DD exhibited the highest average r&#xb2; (0.34), followed by LL (0.28) and YY (0.25). Within-breed genetic diversity metrics, including observed heterozygosity (HO: 0.37 in DD, 0.39 in LL, 0.38 in YY), expected heterozygosity (HE: 0.36 in DD, 0.37 in LL, 0.38 in YY), and minor allele frequency (MAF: 0.27 in DD, 0.28 in LL, 0.29 in YY), indicated greater genetic variability in LL and YY compared to DD. Runs of homozygosity (ROH) analyses revealed different patterns of autozygosity, with DD exhibiting more long ROH indicative of recent inbreeding, while YY harbored a higher number of short ROH, suggestive of more ancient demographic events. ROH-based inbreeding coefficients (FROH) consistently exceeded pedigree-based estimates (FPED) across all breeds, highlighting the presence of recent or unrecorded inbreeding that pedigree data may not fully capture. According to Generation Proxy Selection Mapping (GPSM), 17, 1, and 12 significant SNPs were detected in DD, LL, and YY, respectively. Functional annotation of ROH islands and GPSM-significant loci revealed both breed-specific and overlapping QTLs related to traits such as growth, reproduction, and carcass. In general, the findings of this study contribute to a deeper understanding of the genomic consequences of long-term closed breeding and provide reference information to support consideration of breeding strategies that balance continued selection for productivity with the maintenance of genetic diversity in modern commercial pig populations.

Animals

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans

Heterogeneous effects of genetic variants and traits associated with fasting insulin on cardiometabolic outcomes.

Elevated fasting insulin levels (FI), indicative of altered insulin secretion and sensitivity, may precede type 2 diabetes (T2D) and cardiovascular disease onset. In this study, we group FI-associated genetic variants based on their genetic and phenotypic similarities and identify seven clusters with distinct mechanisms contributing to elevated FI levels. Clusters fall into two types: "non-diabetogenic hyperinsulinemia," where clusters are not associated with increased T2D risk, and "diabetogenic hyperinsulinemia," where T2D associations are driven by body fat distribution, liver function, circulating lipids, or inflammation. In over 1.1 million multi-ancestry individuals, we demonstrated that diabetogenic hyperinsulinemia cluster-specific polygenic scores exhibit varying risks for cardiovascular conditions, including coronary artery disease, myocardial infarction (MI), and stroke. Notably, the visceral adiposity cluster shows sex-specific effects for MI risk in males without T2D. This study underscores processes that decouple elevated FI levels from T2D and cardiovascular risk, offering new avenues for investigating process-specific pathways of disease.

Humans

Alternative tandem transcription initiation links noncoding variants to human disease through translational control.

Alternative tandem transcription initiation is a pervasive mechanism of gene regulation, yet its genetic impact on human disease remains largely unknown. Here, we systematically quantify the genetic regulation of alternative tandem transcription initiation across 25,859 samples from 49 normal human&#xa0;tissues and 33 tumor tissues. We identify approximately 0.4 million genetic variants associated with alternative transcription initiation in 5295 genes, with 32% operating independently of gene expression. Moreover, we discover 2238 multi-tissue alternative tandem transcription initiation outliers enriched for rare deleterious promoter and 5' UTR variants, demonstrating that both common and rare variants modulate transcription initiation. Strikingly, 74% of disease variants that colocalize with genetic variants regulating alternative transcription initiation cannot be identified through expression quantitative trait loci. Transcriptome-wide association studies identify 614 disease susceptibility genes associated with alternative transcription initiation, including known cancer drivers such as MAFF and MLLT10. Functional validation uncovers OSGEP as a breast cancer risk gene, where the alternative allele lengthens the 5' UTR and reduces protein abundance through upstream open reading frame-mediated translation repression, and suppresses breast cancer cell proliferation. Our findings establish alternative transcription initiation as a major, underappreciated mechanism associating noncoding variation with disease, providing a critical resource for interpreting disease risk loci.

Humans

The impact of common and rare genetic variants on bradyarrhythmia development.

To broaden our understanding of bradyarrhythmias and conduction disease, we performed common variant genome-wide association analyses in up to 1.3&#x2009;million individuals and rare variant burden testing in 460,000 individuals for sinus node dysfunction (SND), distal conduction disease (DCD) and pacemaker (PM) implantation. We identified 13, 31 and 21 common variant loci for SND, DCD and PM, respectively. Four well-known loci (SCN5A/SCN10A, CCDC141, TBX20 and CAMK2D) were shared for SND and DCD, while others were more specific for SND or DCD. SND and DCD showed a moderate genetic correlation (rg&#x2009;=&#x2009;0.63). Cardiomyocyte-expressed genes were enriched for contributions to DCD heritability. Rare-variant analyses implicated LMNA for all bradyarrhythmia phenotypes, SMAD6 and SCN5A for DCD and TTN, MYBPC3 and SCN5A for PM. These results show that variation in multiple genetic pathways (for example, ion channel function, cardiac developmental programs, sarcomeric structure and cellular homeostasis) appear critical to the development of bradyarrhythmias.

Humans

Genomic diversity and selection signatures in Asian Zebu Cattle: insights into adaptation and genetic erosion.

Indigenous cattle breeds in Asia are highly adapted to their local environments providing essential commodities such as meat, milk and draught power while also playing a key role in traditional ceremonies, and sports. Despite ongoing efforts to characterize and conserve these breeds, the increasing trend of indiscriminate crossbreeding of Zebu cattle with high-yielding taurine breeds, threatens their genetic diversity. This study investigates the population structure, inbreeding levels, effective population size, gene flow and identification of selection footprints of Asian Zebu (Bos indicus) cattle. Using an Axiom 60&#xa0;K SNP chip, we analyzed genotypes from 1303 cattle across 36 populations in nine countries, including seven taurine outgroups and 29 Zebu populations from Bangladesh, Cambodia, India, Myanmar, Pakistan, and Sri Lanka. Zebu populations demonstrated moderate genetic diversity, with heterozygosity levels averaging 0.356, inbreeding coefficients ranging from 0.026 to 0.074 and genetic differentiation (FST) varied between 0.01 and 0.11. Breed clusters aligned closely with their geographic locations except for Achai (Pakistan) and Baru Harak (Sri Lanka) breeds that appeared in both Zebu and taurine clusters indicating evidence of taurine admixture. Genomic analyses identified regions under selection using extended haplotype homozygosity (EHH) and fixation index (FST) methods. Candidate genes associated with key biological functions related to environmental responsiveness, including heat tolerance (HSP90AA1), immunity (RIPK3), metabolism and fertility (REC8, CLIC4, TSSK4), were identified, reflecting adaptive traits critical for Zebu survival and utility across diverse environments. These findings provide valuable insights for conservation and management strategies aimed at preserving the unique genetic diversity of Asian Bos indicus breeds.

Animals

Identifying causal genetic variants for high-altitude adaptation through blood eQTL analysis in plateau populations.

A substantial number of genetic variants have been associated with high-altitude adaptation (HAA), yet most of them are located in non-coding genomic regions, leaving their specific functions and underlying mechanisms largely unknown. In this study, we analyze whole-genome and transcriptome sequencing data from a self-established cohort comprising 61 native highlanders (NHs) and 164 acclimatized newcomers (ANs), identifying 6,586 cis- and 34,203 trans-expression quantitative trait loci (eQTLs), along with 130 cell type-specific eQTLs. By further combining these data with a large East Asia (~30% Tibetan) genome-wide association study (GWAS) cohort, we employ colocalization and causal inference analyses to prioritize 85 cis-eQTLs associated with HAA and identify several novel candidate causal genes, including EXOC8, which is experimentally confirmed to regulate erythroid differentiation. Additionally, network analysis of these causal genes uncovers multiple regulatory pathways, mainly involving energy metabolism, autophagy, ubiquitination and inflammation. Our study offers a comprehensive eQTL map and reveals causal chains of "variant-gene-phenotype" for HAA-related traits, which provides new insights into potential regulatory mechanisms and targets for prevention and treatment of altitude sickness.

Quantitative Trait Loci

Genome-to-genome analysis reveals associations between human and mycobacterial genetic variation in tuberculosis patients from Tanzania.

The risk and prognosis of tuberculosis (TB) are influenced by a complex interplay between human and bacterial genetic factors. While previous genomic studies have largely examined human and bacterial genomes separately, we adopted an integrated approach to uncover host-pathogen interactions. We leveraged paired human and Mycobacterium tuberculosis (M.tb) genomic data from 1000 adult TB patients from Tanzania and used a "genome-to-genome" approach to search for associations between human and M.tb genetic variants and to identify interacting genetic loci. Our analyses revealed two significant host-pathogen genetic associations. The first significant association (p&#x2009;=&#x2009;4.7e-11) links a human intronic variant in PRDM15 (rs12151990), a gene involved in apoptosis regulation, with an M.tb variant in Rv2348c (I101M), which encodes a T cell-stimulating antigen. The second significant association (p&#x2009;=&#x2009;6.3e-11) connects a human intergenic variant near TIMM21 and FBXO15 (rs75769176) - also associated with TB severity (p&#x2009;=&#x2009;0.04) - with an M.tb variant in FixA (T67M). While FBXO15 is involved in the regulation of antigen processing and TIMM21 affects mitochondrial function, FixA's role remains undefined due to limited functional characterization. Additionally, we observed that a group of M.tb T cell epitope variants were significantly associated with HLA-DRB1 variation, suggesting that, despite their rarity, certain epitopes may still be subjected to immune selective pressure. Together, these findings identify previously unknown sites of genomic conflicts between humans and M.tb, advancing our understanding of how this pathogen evades selection pressure and persist in human populations.

Humans

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A&#xa0;key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

Genetic variation influences food-sharing sociability in honey bees.

Individual variation in sociability is a central feature of every society. This includes honey bees, with some individuals well connected and sociable, and others at the periphery of their colony's social network. However, the genetic and molecular bases of sociability are poorly understood. Trophallaxis-a behavior involving sharing liquid with nutritional and signaling properties-comprises a social interaction and a proxy for sociability in honey bee colonies: more sociable bees engage in more trophallaxis. Here, we identify genetic and molecular mechanisms of trophallaxis-based sociability by combining genome sequencing, brain transcriptomics, and automated behavioral tracking. A genome-wide association study (GWAS) identified 18 single nucleotide polymorphisms (SNPs) associated with variation in sociability. Several SNPs were localized to genes previously associated with sociability in other species, including in the context of human autism, suggesting shared molecular mechanisms of sociability. Variation in sociability also was linked to differential brain gene expression, particularly genes associated with neural signaling and development. Using comparative genomic and transcriptomic approaches, we also detected evidence for divergent mechanisms underpinning sociability across species, including those related to reward sensitivity and encounter probability. These results highlight both potential evolutionary conservation of the molecular roots of sociability and points of divergence.

Animals

Leveraging functional annotations to map rare variants associated with Alzheimer disease with gruyere.

Increased availability of whole-genome sequencing (WGS) has facilitated the study of rare variants (RVs) in complex diseases. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most do not fully leverage the availability of variant-level functional annotations. We propose genome-wide rare variant enrichment evaluation (gruyere), an empirical Bayesian framework that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimer's Disease Sequencing Project to identify Alzheimer disease (AD)-associated genes and annotations. Growing evidence suggests that the disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not examined rare non-coding effects that incorporate such cell-type-specific information. To address this gap, we (1) define per-gene non-coding RV test sets using predicted enhancer and promoter regions in microglia and other brain cell types (oligodendrocytes, astrocytes, and neurons) and (2) include cell-type-specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 13 significant genetic associations not detected by other RV methods, four of which remain significant in omnibus tests. We find that deep-learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a robust framework incorporating functional annotations, coding RVs, and cell-type-associated non-coding RVs to perform genome-wide association tests, uncovering AD-relevant genes and annotations.

Alzheimer Disease

Novel genetic determinants contribute to hearing loss in a central European cohort with enlarged vestibular aqueduct.

BACKGROUND: The enlarged vestibular aqueduct (EVA) is the most commonly detected inner ear malformation. Biallelic pathogenic variants in the SLC26A4 gene, coding for the anion exchanger pendrin, are frequently involved in determining Pendred syndrome and nonsyndromic autosomal&#xa0;recessive hearing loss DFNB4 in EVA patients. In Caucasian cohorts, the genetic determinants of EVA remain unknown in approximately 50% of cases. We have recruited a cohort of 32 Austrian patients with hearing loss and EVA to define the prevalence and type of pathogenic sequence alterations in SLC26A4 and discover novel EVA-associated genes. METHODS: Sanger sequencing, single nucleotide polymorphism (SNP) assays, copy number variation (CNV) testing, and Exome Sequencing (ES) were employed for gene analysis. Cell-based functional and molecular assays were used to discriminate between gene variants with and without impact on protein function. RESULTS: SLC26A4 biallelic variants were detected in 5/32 patients (16%) and monoallelic variants in 5/32 patients (16%). The pathogenicity of the uncharacterized SLC26A4 protein variants was assigned or excluded based on their ion transport function and cellular abundance. The monoallelic or biallelic Caucasian EVA haplotype was detected in 7/32 (22%) patients, but its pathogenicity could not be confirmed. X-linked pathogenic variants in POU3F4 (2/32, 6%) and biallelic pathogenic variants in GJB2 (2/32, 6%) were also found. No CNV of SLC26A4 and STRC genes was detected. ES of eleven undiagnosed patients with bilateral EVA detected rare sequence variants in six EVA-unrelated genes (monoallelic variants in SCD5, REST, EDNRB, TJP2, TMC1, and two variants in CDH23) in five patients (5/11, 45%). Cell-based assays showed that the TJP2 variant leads to a mislocalized protein product forming dimers with the wild-type, supporting autosomal dominant pathogenicity. The genetic causes of hearing loss and EVA remained unidentified in (14/32) 44% of patients. CONCLUSIONS: The present investigation confirms the role of SLC26A4 in determining hearing loss with EVA, identifies novel genes in this pathophysiological context, highlights the importance of functional testing to exclude or assign pathogenicity of a given gene variant, proposes a possible diagnostic workflow, suggests a novel pathomechanism of disease for TJP2, and highlights voids of knowledge that deserve further investigation.

Humans