Search PubMedSearch

SEARCH · Search PubMed

Results for “gnomAD”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Estimation of carrier frequencies of autosomal and X-linked recessive genetic conditions based on gnomAD v4.0 data in different ancestries.

PURPOSE: Monogenic rare diseases contribute significantly to infant deaths and pediatric hospitalizations and cause burden to the patients and their families. The American College of Medical Genetics and Genomics recommended in 2021 that carrier screening of autosomal recessive and X-linked conditions with a carrier frequency of ≥1/200 and a severe or moderate phenotype should be offered when planning or during pregnancy. In November 2023 gnomAD v4.0 was released. It contains in total 807,162 individuals, being nearly 5× larger than previous versions, which have been used to estimate gene carrier frequencies (GCF). METHODS: We utilized gnomAD v4.0 (GRCh38) to calculate the GCFs for available genetic ancestry groups for variants having pathogenic or likely pathogenic classification (>80% of submissions) in ClinVar. We calculated GCF separately for exomes and genomes, combined data, and at-risk couple frequencies (ACF) per genetic ancestry group. RESULTS: In total, 324 genes had a GCF ≥1/200 in at least 1 ancestry subgroup. The number of genes with GCF ≥1/200 varied greatly between subgroups. ACFs were more similar, Ashkenazi Jewish having the highest ACF of 6.11%. CONCLUSION: Improved understanding of carrier risks and updated carrier screening content would allow patients to make more informed reproductive decisions.

Humans

Clinically Relevant Pharmacogenomic Variant Frequencies in Kazakh, Russian, and Uzbek Population Groups Residing in Kazakhstan.

Central Asian populations remain underrepresented in pharmacogenomic research, limiting the availability of population-specific data for genotype-informed prescribing and precision medicine. This study analyzed clinically relevant pharmacogenomic variant frequencies in Kazakh, Russian, and Uzbek population groups residing in Kazakhstan using genome-wide genotype data from 1301 individuals: Kazakh (n = 1111), Russian (n = 156), and Uzbek (n = 34). ClinPGx, a PharmGKB-based clinical annotation framework that prioritizes variant-drug associations according to levels of evidence, was used to select variants with evidence levels 1A, 1B, and 2A. In total, 112 directly genotyped variants were retained for population-specific allele and genotype frequency analysis. All 112 variants were queried against the gnomAD v4.1 genome and exome reference datasets. Of these, matching allele-frequency data for the predefined reported allele were available in at least one of the two gnomAD datasets for 103 variants, whereas for 9 variants the VEP-based query did not return a matching gnomAD frequency for that allele. Frequencies were reported for the same predefined reported allele across all groups, and differences between the study groups were assessed using 95% confidence intervals, Fisher's exact tests, and false discovery rate correction. Genotype counts and the proportions of individuals carrying at least one copy of the reported allele were also summarized for all selected variants. Several pharmacogenomic variants showed population-specific frequency patterns, including NUDT15 rs116855232, SLCO1B1 rs4149056, VKORC1 rs9934438, and UGT1A1 rs10929302. Comparison with gnomAD showed that the observed frequencies were variant-specific and could not be consistently approximated by a single broad genetic ancestry group. Reference-based population structure analysis provided additional ancestry context and supported separate reporting by population group. The study did not evaluate clinical outcomes or make individual prescribing recommendations, and the small Uzbek sample size limits the precision of frequency estimates for this group, particularly for rare variants. Overall, this study provides a clinically prioritized pharmacogenomic frequency resource for underrepresented population groups in Kazakhstan and supports broader Central Asian representation in pharmacogenomic implementation research.

Central Asia

Twelve Japanese patients with POLG-related disorders: Population-specific genetic differences of POLG variants in Japan and Europe.

BACKGROUND: POLG encodes mitochondrial DNA (mtDNA) polymerase γ. Pathogenic POLG variants cause mitochondrial diseases, including progressive external ophthalmoplegia. POLG-related disorders are relatively common in Europe, possibly because of the high prevalence of carriers in the general population, but remain rare in Japan for unclear reasons. METHODS: We performed long-range PCR on mtDNA from skeletal muscle and/or peripheral blood from 3146 patients with suspected mitochondrial disease between 1993 and 2021. We selected 167 individuals with clinical features suggestive of POLG-related disorders for POLG gene analysis; all lacked pathogenic mtDNA point mutations, and most had multiple mtDNA deletions and/or a family history of mitochondrial disease. RESULTS: Among the 167 patients (median age: 52 years, range: 0-83 years, 11% pediatric cases), we identified 12 Japanese patients with POLG-related disorders and six POLG variants, including one novel variant. The six variants were p.Y955C, p.R943H, p.T599I, p.M299L, p.Y1210* (c.3626_3629dupGATA), and the novel variant p.F377S (c.1130T>C). Neither these six variants nor the 10 previously reported cases from Japan included the POLG variants that are more frequent in Europe. We also analyzed three population databases: two whole-genome sequencing databases covering 61,000 and 9850 Japanese individuals, respectively, and one global population database (gnomAD) covering 730,000 individuals worldwide. POLG variants that are more frequent in Europe were not detected in the Japanese databases or among East Asian individuals in gnomAD. CONCLUSIONS: Our findings suggest population-specific genetic differences in POLG between Japanese and European populations, explaining the lower frequency of POLG-related disorders in Japan.

CPEO

Whole-exome characterization of host genetic variation in HIV-associated genes across the high-prevalence Mizo population, Northeast India.

BACKGROUND: The Mizoram state of Northeast India has one of the highest HIV prevalence rates in Asia, yet the host genetic factors influencing HIV susceptibility in this Tibeto-Burman population remain uncharacterised. METHODS: We performed whole-exome sequencing using Illumina NovaSeq 6000, mean coverage 100X on 76 HIV-negative Mizo individuals. Variants were called using GATK HaplotypeCaller v4.3 against GRCh38p14, annotated with ANNOVAR, and filtered using hard-quality thresholds (QD&#xa0;&#x2265;&#xa0;2, SOR&#xa0;&#x2264;&#xa0;3, MQ&#xa0;&#x2265;&#xa0;40, DP&#xa0;&#x2265;&#xa0;10, GQ&#xa0;&#x2265;&#xa0;20). The allele frequencies were compared against gnomAD v2.1.1 population databases. Hardy-Weinberg equilibrium was assessed using the Wigginton exact test with Bonferroni correction. RESULTS: Post-quality filtering resulted in 12,011 sample-variants across 2,821 unique positions from 36 HIV-associated loci (33 protein-coding genes, 2 chemokine ligands, and 3 lncRNA targets). Of these, 784 observations (51 unique positions) were high-impact nonsynonymous or loss-of-function variants. ADAR rs2229857 (p.K384R, NM_015840) was the most frequently observed variant (Mizo carrier frequency&#xa0;=&#xa0;0.895; 95% CI: 0.806-0.946). CXCR1 rs16858808 (p.R335C) showed the greatest population enrichment (Mizo carrier frequency&#xa0;=&#xa0;0.197; 95% CI: 0.123-0.300; 7.65-fold carrier-frequency enrichment versus gnomAD South Asian; CADD&#xa0;=&#xa0;15.60). Sixteen of 20 tested variants deviated from Hardy-Weinberg equilibrium after Bonferroni correction (p&#xa0;<&#xa0;0.0025), predominantly showing excess homozygosity consistent with the endogamous Mizo population. The protective variant CCR5-&#x394;32 was absent in all the 76 individuals tested. CONCLUSION: This first whole-exome characterization of HIV host genes in the Mizo population identifies CXCR1 rs16858808 as the most population-enriched functional variant and reveals a pervasive endogamy signature. These findings provide a population-specific genetic framework for future HIV susceptibility studies and ART pharmacogenomics research.

Humans

Optimizing gene panels for equitable reproductive carrier screening: The Goldilocks approach.

PURPOSE: Professional organizations recommend pan-ancestry carrier screening for autosomal recessive and X-linked conditions. Advances in DNA sequencing have allowed the analysis of hundreds of genes; however, the optimal number of genes for carrier screening remains unclear. The American College of Medical Genetics and Genomics (ACMG) has proposed a tiered approach recommending screening for 113 genes. METHODS: We analyzed ClinVar and gnomAD v4.1.0, for genes associated with serious autosomal recessive and X-linked conditions and modeled screening performance across panels of varying compositions and sizes in diverse genetic ancestries. We also reevaluated the ACMG gene list using the updated gnomAD data. RESULTS: We identified potential inconsistencies in the ACMG gene lists, particularly in the carrier test performance (defined as a positive yield) for underrepresented genetic ancestry groups. Modeling of the population data for 1310 genes revealed that the screening of 152, 248, 531, and 725 genes achieved 90%, 95%, 99%, and 99.7% positive yields, respectively, in couples. Real-world data from the screening of more than 60,000 couples were used to validate the model. CONCLUSION: Our methodology optimizes the gene content of carrier screening panels for diverse ancestry groups, provides a mechanism for continually updating guidelines, ensures consistency with genomic population data, and improves equity across populations.

Humans

Tandem splice acceptor sites: Profiling their relevance to human disease.

PURPOSE: Interpretation of variation, particularly the creation or disruption of tandem splice acceptor sites (NAGNnAG variants), challenges genomic medicine practice. METHODS: We analyzed the creation and disruption of dinucleotide AG sites within &#xb1;30 bases of natural splice-acceptor sites in the GRCh37 human reference genome. These results were compared with variant data from the ClinVar and gnomAD databases, as well as with data from 779 National Institutes of Health Undiagnosed Diseases Program study participants. Using RNA sequencing, we assessed the splicing at NAGNnAG variants for 107 of the Undiagnosed Diseases Program participants and compared the empirical data with SpliceAI predictions. RESULTS: Creation or disruption of NAGNnAG sites within 30 bases of the natural splice acceptor are enriched in ClinVar compared with gnomAD; however, such variants in the 2 databases are rarely differentiated by SpliceAI scores. Empirical evaluation via RNA sequencing analysis supported novel acceptor site usage from -21 to +30; splice-altering variants did not predominate in a specific region or have SpliceAI scores invariantly, suggesting increased spliceogenicity. CONCLUSION: NAGNnAG variants within 30 bp of the natural splice acceptor have a high probability of clinical relevance and are poorly contextualized for clinical utility. Their interpretation benefits from empirical evaluation via RNA analysis.

Humans

Exploring penetrance of clinically relevant variants in over 800,000 humans from the Genome Aggregation Database.

Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.

Humans

T-rex: standardized analysis of germline variants in whole-exome sequencing trios.

Whole-exome sequencing (WES) enables the identification of rare germline variants contributing to pediatric diseases. Trio-based sequencing, comparing affected children with their parents, is particularly effective for rare disease genetics. However, WES data analysis requires bioinformatics expertise, varies across institutions, and is often incompatible with clinical workflows. We developed T-Rex (Trio Rare variant analysis of EXomes), a cross-platform desktop application that enables the standardized and local analysis of WES germline Trio data without the need for programming knowledge. T-Rex integrates state-of-the-art tools for alignment, dual-variant calling (GATK HaplotypeCaller&#x2009;+&#x2009;VarScan2), annotation (SNPEff/SNPSift), rare-variant filtering based on population frequencies (gnomAD), and family-based statistical testing, including the Transmission Disequilibrium Test with multiple-testing correction. Benchmarking of the dual-caller strategy on the Genome in a Bottle Ashkenazim Trio demonstrates high precision (99.2%) while maintaining robust sensitivity (91.1%). User testing (n&#x2009;=&#x2009;13) confirmed quick learning across clinicians and researchers. Application to a cohort of n&#x2009;=&#x2009;121 pediatric cancer Trio datasets, filtering for rare protein-coding variants (MAF&#x2009;&#x2264;&#x2009;0.1% in gnomAD v4.1), validated all assessable previously reported pathogenic variants. Overall, T-Rex enables clinicians to robustly analyze WES Trio data in compliance with data protection regulations without requiring additional software licenses. As one of the first platforms for comprehensive WES Trio analysis that requires no programming expertise while providing reproducible, end-to-end workflows for clinical genomics, T-Rex facilitates collaborative research between clinics and reduces reliance on external providers.

Humans

Somatic likelihood tiering: an interpretable post-calling triage protocol for tumor-only whole-exome variant review.

Tumor-only whole-exome sequencing (WES) is used when matched normal tissue is unavailable, but one sample can produce thousands of variants. Somatic likelihood tiering (SLT) is an interpretable post-calling protocol that ranks Mutect2 calls into four review-priority tiers using population-frequency, germline-quality, cancer-knowledge, PureCN posterior, and clonal-hematopoiesis evidence. Layer 2 distinguishes common, rare-callable, and unevaluable gnomAD states; missing or unmatchable gnomAD evidence is not positive rarity evidence. On the SEQC2 HCC1395 benchmark, the callability-aware SLT-A row contained 101 calls, 78 truth variants, 77.2% PPV (95% Wilson confidence interval 68.1%-84.3%), and a Number Needed to Review (NNR) of 1.29 (1.19-1.47). The conservative SLT-C catchment retained 352 of 455 truth variants (77.4%, 73.3%-81.0%) and all tiers together retained 430 of 455 truth variants. SNV performance is the primary calibration frame: SLT-C retained 341 of 439 SNV truth variants, whereas indel results were exploratory because only 16 truth indels were available. Clinical cohorts are reported as recall and concordance versus partially dependent matched-normal Mutect2 references, not independent clinical sensitivity. Patient-level bootstrap intervals were principal: HdM-BLCA-1 SLT-A recall was 18.2% (14.0%-23.5%), and LUAD-TW SLT-A recall was 49.1% (26.6%-63.3%) among 32 evaluable patients. The HdM-BLCA-1 median SLT-A queue remained 1277 variants per patient, so SLT reduces first-pass candidate counts but does not measure review time or eliminate FFPE candidate-count burden. SLT provides an auditable tumor-only WES review queue, not a substitute for matched-normal sequencing, independent orthogonal validation, or definitive somatic classification.

Humans

Inference of elevated mutation rates and variant effects using 700k exomes.

Genomic sequencing is now widely accessible for genetic diagnostics and is emerging as a component of newborn screening. This technological development generates the need to characterize incoming mutations, create comprehensive datasets of genes causing rare Mendelian disorders, and identify pathogenic variants. Large-scale exome sequencing datasets such as Genome Aggregation Database (gnomAD) have been assembled to help address these challenges. The recent release of gnomAD (v4; n = 730,947) uncovers millions of rare coding variants, many of which have arisen more than once by independent recurrent mutations in the rapidly growing recent human population. Here, we use newly developed theoretical understanding of sampling properties of rare variants to estimate key population genetics parameters of practical importance to human genetics such as demography history, mutation rate, and selection. Solely relying on population data, our method Population Inferred Estimates of Selection (PIES) identifies novel genes with loss-of-function mutational hotspots likely due to selection in spermatogonia. PIES efficiently estimates selection coefficients for heterozygous loss-of-function variants. Combining population genetics inference with variant effect predictors, PIES predicts pathogenic missense mutations and improves variant prioritization for genetic diagnostics and newborn screening.

Journal Article

Genetic Ancestry and Carrier Variant Frequency Enrichment in a Colombian Andean Population: Insights From the Eje Cafetero.

Colombia is one of the most genetically diverse populations in Latin America, and its demographic process has promoted the persistence and local enrichment of deleterious alleles, increasing the frequency of autosomal recessive disorders, particularly in semi-isolated Andean populations such as the Eje Cafetero. However, exome-based reference data from this region remain scarce, limiting ancestry-aware variant interpretation and carrier screening strategies. We aimed to characterize the ancestry proportions of this population using exome data, and to estimate the carrier frequency and distribution of pathogenic and likely pathogenic (P/LP) variants in clinically relevant recessive genes. We conducted a cross-sectional study with whole-exome sequencing (WES) in 316 unrelated individuals from the Colombian Eje Cafetero. P/LP variants were evaluated in 454 genes associated with autosomal recessive disorders. The global ancestry proportions were estimated using a validated panel of 250 exome-compatible ancestry-informative markers. Carrier frequencies were compared against Non-Finnish Europeans (NFE) and Admixed Americans (AMX) from gnomAD v4. The cohort showed predominant European ancestry (mean 51%), followed by Native American (36%) and African (13%) components. We identified 151 carriers of 89 distinct pathogenic variants across autosomal recessive genes. The most frequent variants were SERPINA1 c.863A>T (5.5%), CFTR c.1210-11T>G (3.5%), and PYGM c.1094C>T (1.5%). Also, recurrent variants were significantly enriched compared with both NFE and AMX populations, supporting regional founder effects. This study represents one of the most comprehensive exome-based genetic characterizations of the Colombian Eje Cafetero, revealing ancestry-specific enrichment of clinically relevant autosomal recessive variants driven by founder effects.

Female

Investigating genetic susceptibility to concussion through rare variants in ion channel and neurotransmission genes.

Some individuals appear more susceptible to concussion or mild traumatic brain injury (mTBI) and the severity, range, and the persistence of post-concussion symptoms vary considerably between affected individuals. Genetic factors are likely to contribute to this variability. Symptomatic overlap of post-concussion syndrome with neurological conditions such as familial hemiplegic migraine (FHM) caused by rare pathogenic variants in ion channel and synapse protein genes, with high sensitivity to head trauma for some patients, suggests that variation in similar pathways may influence concussion susceptibility and recovery. To investigate this hypothesis, we performed whole exome sequencing in 93 unrelated individuals who had sustained a single or multiple concussions and examined rare protein-altering variants in FHM genes, other neuronal ion channel and transporter genes, and genes involved in neurotransmission. We identified 62 different rare missense variants across 24 genes in 59 participants (63%), with 26 individuals carrying 2 or more variants. The prevalence of specific likely damaging rare variants in the 16 ion channel-related genes that were identified was approximately fivefold higher than that observed from gnomAD population controls (Odds Ratio = 5.44, 95% CI [4.13,7.18], P < 0.0001). Notably, voltage-gated calcium and sodium channel genes, including SCN9A, together with neurotransmission-related genes such as SNCAIP, harboured multiple potentially deleterious variants. These findings suggest that rare deleterious variants in genes involved in ion homeostasis and neurotransmission may contribute to an individual's susceptibility to concussion or more severe post-concussion symptoms. This study provides a foundation for future genetic and functional investigations aimed at improving our understanding of concussion susceptibility and outcomes. Further validation in larger cohorts and mechanistic studies is warranted to determine their utility as biomarkers of concussion risk and prognosis.

Humans

Androgens mediate sexual dimorphism in Pilarowski-Bjornsson syndrome.

Sex-specific penetrance in autosomal-dominant Mendelian conditions is largely understudied. The neurodevelopmental disorder Pilarowski-Bjornsson syndrome (PILBOS) was initially described in females. Here, we describe the clinical and genetic characteristics of the largest PILBOS cohort to date, showing that both sexes can exhibit PILBOS features, although males are overrepresented. A mouse model carrying a human-derived Chd1 missense variant (Chd1R616Q/+) displays female-restricted phenotypes, including growth deficiency, anxiety, and hypotonia. Orchiectomy unmasks a growth-deficiency phenotype in male Chd1R616Q/+ mice, while testosterone rescues the phenotype in females, implicating androgens in phenotype modulation. In the gnomAD and UK Biobank databases, rare missense variants in CHD1 are overrepresented in males, supporting a male-protective effect. We identify 33 additional highly constrained autosomal genes with missense variant overrepresentation in males. Our results support androgen-regulated sexual dimorphism in PILBOS and open avenues toward understanding the mechanistic basis of sexual dimorphism in other autosomal Mendelian disorders.

Male

CanVar-UK: A collaborative platform for germline interpretation in cancer susceptibility genes.

Germline variants in cancer susceptibility genes (CSGs) are typically inherited rather than arising de novo. Hence, wide cascade testing of families across geographies is common, meaning consistency in variant classification is particularly critical. Variant interpretation requires collation of variant-level data from diverse sources, as well as assembly of comprehensive clinical data, often necessitating sharing of information between genomic testing centers. Here, we describe CanVar-UK, a freely accessible web platform bespoke designed to support interpretation of germline CSG variants. CanVar-UK contains variant-level data for over 1.1 million single-nucleotide variants (SNVs), comprising all possible coding SNVs in 116 established CSGs. The data sources with which variants are annotated include in silico scores from 11 clinically relevant tools, population allele frequencies from gnomAD v4.1, case counts from multiple cohorts, including National Health Service (NHS) clinical laboratory testing, variant-level readouts from 47 selected functional and splicing datasets across 19 CSGs, genetic epidemiology studies, and live linkage to existing consensus classifications in the ClinVar database. The diagnostic discussion forum is only available to registered diagnostic scientist users. Through this, a variant-tagged email message can be dispatched in real time across the diagnostic forum community of >1,500 users, with all exchanges and classifications captured and stored in the platform. Already widely used by NHS diagnostic clinical scientists in the UK, CanVar-UK has a rapidly growing international diagnostic user base (>800 UK and >600 non-UK registered users). Survey of the NHS diagnostic user community illustrates the wide-ranging utility of CanVar-UK within their clinical workflows for interpretation of germline CSG variants.

Journal Article

A recurrent CCDC82 frameshift variant associated with syndromic neurodevelopmental disorder in a consanguineous Pakistani family.

BACKGROUND: Intellectual disabilities (IDs) are part of neurodevelopmental disorders (NDDs) and are genetically heterogeneous conditions characterized by impairments in cognition, learning, and adaptive functioning. Despite advances in gene discovery, many individuals, particularly those from understudied populations, remain without a molecular diagnosis. Recent reports implicate CCDC82 (HGNC: 26282) as an autosomal recessive ID gene, although the phenotypic spectrum and biological context remain incompletely defined. METHODS: Exome sequencing (ES) was performed in a consanguineous Pakistani family (PKMR06A) with four affected individuals presenting with moderate to severe ID. Variant segregation was confirmed by Sanger sequencing. In silico analyses, including pathogenicity prediction, protein structural modeling, and domain intolerance assessment, were used to evaluate the functional consequences of the identified variant. Spatiotemporal gene expression patterns were examined using bulk and single-cell human brain transcriptomic datasets. RESULTS: Clinically, affected individuals of family PKMR06A presented with early childhood global developmental delay, speech delay, hypotonia, gait abnormalities, spasticity, and mild facial dysmorphism. Genetic screening revealed a recurrent rare homozygous frameshift variant in CCDC82 (NM_024725.4): c.373del; p.(Asp125Ilefs*6), segregating with disease in all available affected individuals of the family. The identified c.373del variant was absent from the gnomAD database and was classified as pathogenic (PVS1, PM2, and PP1) based on ACMG/AMP criteria. The c.373del variant is predicted to introduce a premature termination codon, p.(Asp125Ilefs*6), leading to deletion of essential coiled-coil domains from the encoded protein, supporting a loss-of-function mechanism. In silico, transcriptomic analyses demonstrated preferential CCDC82 expression during prenatal human brain development, providing developmental context for the neurodevelopmental phenotype associated with the identified truncating variant. CONCLUSIONS: This study expands the mutational landscape of CCDC82 and provides additional clinical and molecular evidence supporting its role in autosomal recessive NDD. The findings reinforce the importance of CCDC82 in human neurodevelopment and highlight the value of genomic investigation in underrepresented populations.

Autosomal recessive

Novel splice site variants in GBA1 are associated with Gaucher disease and genotype-phenotype correlations.

BACKGROUND: Variants in GBA1 are associated with neurodegenerative disease. This study aimed to explore pathogenic GBA1 variants. METHODS: Four patients with progressive myoclonic epilepsy (PME) and extremely low &#x3b2;-glucosidase levels were recruited. Whole-exome sequencing and long-range PCR were performed to identify GBA1 variants. Bioinformatic analyses were used to predict the impact of the identified variants. A literature review was performed to explore the genotype-phenotype correlations. GBA1 expression data across different brain regions and developmental stages were analyzed using the BrainSpan database. RT-PCR was performed to verify the splicing effects. RESULTS: Compound heterozygous GBA1 variants were identified in four patients. Five distinct variants were detected, including two novel splice site variants (c.308-2A>G and c.762-2A>C) and three previously reported variants. All identified variants were rare or absent in gnomAD. Splice site variants c.308-2A>G and c.762-2A>C were predicted to cause aberrant splicing. Minigene-based splicing assays coupled with RT-PCR and Sanger sequencing confirmed that both variants cause complete exon skipping (exon 4 and exon 7, respectively). All patients presented with PME onset in childhood/adolescence, intellectual regression, low &#x3b2;-glucosidase, and diffuse brain atrophy and were subsequently diagnosed with Gaucher disease type 3. GBA1 expression in the brain showed two distinct peaks: one in infancy and another after five years of age. The onset age of PME aligned with the second GBA1 expression peak (after five years of age). CONCLUSION: This study identified compound heterozygous GBA1 variants, including two novel candidate pathogenic splice site variants, in Gaucher disease type 3 patients, expanding the known mutational spectrum.

Humans

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense

A dataset of estimated heterozygous individual and carrier couple frequencies for pan-ancestry carrier screening.

The data described in this publication supported the development and evaluation of pan-ancestry reproductive carrier screening panels for autosomal recessive (AR) and X-linked (XL) conditions. Raw data included combined sets of DNA variants in 1,350 AR/XL genes obtained from the ClinVar and gnomAD databases. The dataset enabled calculations of positive yield for individuals and couples across both ancestry-specific and pan-ancestry, optimised "Goldilocks"-ranked gene panels, addressing population-specific variations in the frequencies of heterozygous individuals and carrier couples. The positive yield analysis offered a performance metric for carrier screening panels, facilitating the modeling of screening performance for panels of varying sizes and composition and providing resources for optimizing panel content to ensure equity across underrepresented genetic ancestries The dataset can support ongoing research into the equitable application of carrier screening and offers significant reuse potential for refining population genetic screening practices, validating computational models, and developing frameworks to update carrier screening panels in alignment with evolving genomic data, including in underrepresented and minority populations.

Carrier screening