Search PubMedSearch

SEARCH · Search PubMed

Results for “coding variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Pathogenic XPO1 variants cause a dominant neurodevelopmental disorder.

PURPOSE: XPO1 functions in key cellular processes, including nucleo-cytoplasmic export and mitosis. The gene is deleted in a subset of patients with the 2p15p16.1 microdeletion syndrome; however, no monogenic XPO1-related disorder has been described to date. METHODS: We collected clinical data of individuals with de novo XPO1 variants through online matchmaking. We used Drosophila to study XPO1 function in development and habituation learning. RESULTS: A total of 22 individuals met the criteria to be included in the main study cohort. Of these, half have putative loss-of-function variants, and half have coding variants (10 missense and 1 in-frame deletion variant). We found an overlapping phenotype, consistent with a monogenic neurodevelopmental disorder. We demonstrate XPO1 functions in development by ubiquitous and neuron-specific knockdown in Drosophila. GABAergic neuron specific knockdown flies demonstrated impaired habituation. CONCLUSION: Our results establish XPO1 as a novel dominant monogenic neurodevelopmental disorder gene and demonstrate a central role for XPO1 in development.

Exportin 1 Protein

The contribution of common and rare genetic variation to emotional and behavioural symptoms in childhood and adolescence.

Genetic factors influence vulnerability to common mental health conditions, but their role in early-life mental health remains understudied. We analysed genotype array (n&#x2009;=&#x2009;4709-6687) and exome sequence data (n&#x2009;=&#x2009;4500-5424) from the Millennium Cohort Study (MCS) and Avon Longitudinal Study of Parents and Children (ALSPAC) to assess the contribution of common variants and rare deleterious coding variants to internalising and externalising symptoms across development. In longitudinal analysis spanning ages 5-17 years, we identified several associations between common genetic variation, indexed by polygenic indices (PGIs), and both symptom domains that generally remained stable across development. Effect sizes were modest, with the largest estimates observed for PGIs for attention deficit hyperactivity disorder (ADHD) and externalising behaviour with externalising symptoms (&#x3b2;&#x2009;=&#x2009;0.13-0.18; p-adj<3.5&#xd7;10&#x207b;29). Evidence for direct genetic effects was strongest for externalising symptoms, including for associations with the ADHD and externalising behaviour PGIs. Concordant results were observed in the Born in Bradford cohort. A higher exome-wide burden of deleterious rare variants was associated with increased externalising and internalising symptoms (&#x3b2;&#x2009;=&#x2009;0.04-0.06, p-adj<0.03); within-family models indicated direct genetic effects on externalising in MCS (&#x3b2;&#x2009;=&#x2009;0.07; p&#x2009;<&#x2009;0.05, p-adj>0.05) and on internalising symptoms in ALSPAC (&#x3b2;&#x2009;=&#x2009;0.12, p-adj<0.02). Common and rare genetic variants contributed independently, jointly explaining 2% of the variance in internalising and 5-7% in externalising symptoms. This study shows that early-life mental health is influenced by both common and rare genetic variation, with several associations explained by direct genetic effects.

Journal Article

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

Expanding the Genomic Spectrum of NHLRC2-Associated FINCA Disease: Integrated Bioinformatic Characterization of a Novel Deep Intronic Variant Predicted to Activate a Pseudoexon.

NHLRC2-associated FINCA disease is an ultra-rare autosomal recessive multisystem disorder caused by biallelic pathogenic variants in NHLRC2. Its mutational spectrum and genotype-phenotype correlations remain incompletely defined, and the contribution of non-coding variants is poorly understood. Here, we report a male infant with a severe FINCA-like phenotype, including early-onset hemolytic anemia, pulmonary involvement, neurodevelopmental impairment, growth failure, recurrent infections, and fatal progression at 8.5 months. Whole-genome sequencing identified a compound heterozygous NHLRC2 genotype comprising the previously reported pathogenic missense variant c.442G>T (p.Asp148Tyr) and a novel deep intronic variant, c.331+6863A>G. Segregation analysis confirmed inheritance from different parents. Integrated genomic and splicing analysis predicted that c.331+6863A>G creates a strong cryptic donor splice site and supports pseudoexon inclusion. Reconstruction of the predicted aberrant transcript indicated premature termination and potential susceptibility to nonsense-mediated mRNA decay. To our knowledge, this is the first reported deep intronic NHLRC2 variant predicted to activate pseudoexon inclusion. Although experimental validation was unavailable, convergent clinical, segregation, population, and computational evidence supports c.331+6863A>G as the most plausible second disease-associated allele. This case expands the genomic spectrum of NHLRC2-associated FINCA disease and highlights the diagnostic value of phenotype-driven whole-genome sequencing.

Humans

Whole-genome sequencing implicates rare, low-frequency and structural non-coding variation at the SCN5A locus in Brugada syndrome.

Brugada syndrome (BrS) is an inherited cardiac condition characterized by a hallmark ECG pattern and an increased risk of sudden cardiac death. Central to the aetiology of BrS, the SCN5A region harbours both common non-coding risk variants and rare coding variants that are causative in approximately 20% of patients. However, rare non-coding genetic variation in this region remains largely unexplored. Here, we used whole-genome sequencing (WGS) of 752 European-ancestry BrS cases and 1,827 ancestry-matched controls to identify BrS-associated rare non-coding genetic variation at the SCN5A locus. Sliding-window and cis-regulatory element (CRE)-based rare-variant aggregate testing implicated three conserved CREs, including a dense aggregation of case singleton variants within a 178 bp enhancer in intron 17 of SCN5A which replicated in an independent BrS cohort. Prioritised BrS-associated rare and low-frequency non-coding variants within these elements were predicted to alter cardiac transcription factor motifs, and altered CRE activity in hiPSC-CM luciferase assays or were associated with BrS-relevant ECG endophenotypes in the UK Biobank. Single-variant analysis across the region identified a Bonferroni-significant five-fold case-enriched low-frequency variant within a known CRE in intron 1 of SCN5A, which replicated, was associated with slower cardiac conduction in the UK Biobank and accounted for part of the BrS GWAS signal at this locus. Structural variant analyses identified a 10.5 kb deletion upstream of SCN5A in a BrS case that encompassed a cardiac CRE and reduced sodium current density in a hiPSC-CM model, as well as a 6 kb BrS-enriched retrotransposon insertion in SCN5A that appeared to underlie part of the GWAS signal in this region. Together, these findings implicate rare and low-frequency non-coding variation at the SCN5A locus in BrS susceptibility and demonstrate the value of targeted WGS analysis of key disease loci.

Journal Article

Protective TMEM106B-rs3173615 delays age at onset in GRN mutation carriers.

One of the major causative genes involved in Frontotemporal dementia (FTD) is Granulin (GRN), encoding for Progranulin (PGRN). GRN mutation carriers show a substantial heterogeneity with high variability in age at onset and pathological presentation, even within the same family or identical mutations, suggesting the presence of additional genetic factors. Single nucleotide polymorphisms in the Transmembrane protein 106B (TMEM106B) locus were identified as a genetic risk-associated factor for FTD. The top variant identified was the non-coding rs1990622, with the major allele (T) associated with an increased risk to develop FTD, while subjects with the minor allele (C) were less likely to develop disease, suggesting a protective effect. In this study, we investigate in a large Italian cohort of GRN mutation carriers, how the coding variant TMEM106B-rs3173615, in linkage disequilibrium with rs1990622, modulates age at onset, survival, and PGRN levels, including, up to date, the highest sample size of homozygous protective allele carriers. Genetic screening for TMEM106B-rs3173615 was performed on a total of 187 GRN mutation carriers, comprising 131 FTD patients and 56 pre-symptomatic subjects. Individuals with the protective genotype (GG) had a risk of FTD onset reduced by 80%, with a median age at onset of 77 years compared to a median age at onset of 63 years for individuals without the protective genotype. TMEM106B-rs3173615 acts as a genetic modifier of age at onset in the presence of GRN mutations and could be considered in clinical practice to optimize risk stratification for FTD.

Humans

Creating an atlas of variant effects to resolve variants of uncertain significance and guide cardiovascular medicine.

Cardiovascular diseases are leading global causes of death and disability, often presenting as interrelated phenotypes of atherosclerotic vascular disease, heart failure and arrhythmias. Cardiovascular diseases arise from interactions between environmental factors and predisposing genotypes and include common Mendelian lipid disorders, cardiomyopathies and arrhythmia syndromes. The identification of a pathogenic variant through genetic testing can inform disease diagnosis, risk prediction, treatment and family screening. However, a major roadblock in genomic medicine is that for many variants, especially missense variants, we lack sufficient evidence to enable a definitive classification, and therefore these variants are deemed as 'variants of uncertain significance'. In this Review, we describe how multiplexed assays of variant effects can enable the functional assessment of nearly all coding variants in a target sequence, potentially offering a proactive approach to identifying the functional significance of gene variants that are observed later in a patient. We discuss validation, including the role of in silico variant effect predictors, and how multiplexed experimental methods are informing cardiovascular disease biology and ultimately resolving the problem of variants of uncertain significance at scale.

Humans

Pharmacogenomic diversity in Amazonian Indigenous populations: implications for Berlin-Frankfurt-M&#xfc;nster acute lymphoblastic leukemia therapy.

PURPOSE: This study aimed to characterize pharmacogenomic variation in genes involved in the metabolism and transport of drugs used in Berlin-Frankfurt-M&#xfc;nster-based therapy in Amazonian Indigenous individuals and to compare allele frequencies with major continental populations. METHODS/PATIENTS: Whole-exome sequencing data previously generated from 64 healthy Indigenous individuals from 12 Amazonian ethnic groups were analyzed. A total of 120 genes associated with drugs used in Berlin-Frankfurt-M&#xfc;nster protocols were selected. Variants were annotated and filtered using bioinformatic quality-control criteria, and allele frequencies were compared with African, Admixed American, East Asian, European, and South Asian populations from the 1000 Genomes Project. Multidimensional scaling was used to assess population-level genetic similarity. RESULTS: After quality control, 648 variants were identified. Twenty-eight variants were observed exclusively in the Indigenous study population, including four nonsynonymous coding variants with moderate predicted impact. Significant allele-frequency differences were observed for ADA rs11555566, CBR3 rs881711, and CYP2B6 rs3745274; rs881711 and rs3745274 differed from all five reference populations. Multidimensional scaling showed a distinct Indigenous pharmacogenomic profile, with greater similarity to the Admixed American population. CONCLUSIONS: Amazonian Indigenous populations exhibit substantial pharmacogenomic diversity in genes relevant to Berlin-Frankfurt-M&#xfc;nster-based therapy. These findings identify candidate variants for functional and clinical validation and reinforce the importance of including underrepresented populations in pharmacogenomic research.

Acute lymphoblastic leukemia

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

Molecular characterization of CD36 deficiency in blood donors of Middle Eastern and African origin reveals transcript-level defects beyond genomic variants.

BACKGROUND: The increasing diversity of blood donor populations has created new challenges for transfusion services worldwide. The identification of donors lacking relevant high-prevalence antigens is becoming increasingly important to ensure compatible blood products for alloimmunized patients and to support the development of rare donor registries. CD36 (ISBT 045) is a glycoprotein expressed on platelets, monocytes, and erythroid precursor cells. CD36 deficiency has been reported across multiple populations and is of relevance due to its association with anti-CD36 isoantibodies, which may cause platelet transfusion refractoriness and fetal/neonatal alloimmune thrombocytopenia. STUDY DESIGN AND METHODS: We analyzed CD36 expression in 1250 blood donors of diverse ancestry using flow cytometry. CD36-negative samples underwent molecular characterization using Sanger sequencing and next-generation sequencing of genomic DNA, complemented by cDNA analysis and cloning to investigate transcript-level alterations. RESULTS: We identified CD36 deficiency in 27 donors (2.16%). Genomic sequencing revealed 18 distinct coding variants, including three novel variants, most in the heterozygous state. In one CD36-negative donor, cDNA analysis demonstrated a 52-bp deletion in exon four and complete skipping of exon 9, despite the absence of splice-site variants in genomic DNA. Cloning confirmed coexistence of aberrant and wild-type transcripts in this individual. CONCLUSION: Our findings demonstrate that CD36 deficiency can arise from transcript-level defects in the absence of detectable coding or splice-site variants. These results indicate that genomic sequencing alone may be insufficient to fully resolve CD36-negative phenotypes and highlight the importance of integrating transcriptomic approaches to improve molecular diagnostics and transfusion support in increasingly diverse donor populations.

CD36 deficiency

Exome sequencing reveals neurodevelopmental genes in simplex consanguineous Iranian families with syndromic autism.

BACKGROUND AND OBJECTIVE: Autosomal recessive genetic disorders pose significant health challenges in regions where consanguineous marriages are prevalent. The utilization of exome sequencing as a frequently employed methodology has enabled a clear delineation of diagnostic efficacy and mode of inheritance within multiplex consanguineous families. However, these aspects remain less elucidated within simplex families. METHODS: In this study involving 12 unrelated simplex Iranian families presenting syndromic autism, we conducted singleton exome sequencing. The identified genetic variants were validated using Sanger sequencing, and for the missense variants in FOXG1 and DMD, 3D protein structure modeling was carried out to substantiate their pathogenicity. To examine the expression patterns of the candidate genes in the fetal brain, adult brain, and muscle, RT-qPCR was employed. RESULTS: In four families, we detected an autosomal dominant gene (FOXG1), an autosomal recessive gene (CHKB), and two X-linked autism genes (IQSEC2 and DMD), indicating diverse inheritance patterns. In the remaining eight families, we were unable to identify any disease-associated genes. As a result, our variant detection rate stood at 33.3% (4/12), surpassing rates reported in similar studies of smaller cohorts. Among the four newly identified coding variants, three are de novo (heterozygous variant p.Trp546Ter in IQSEC2, heterozygous variant p.Ala188Glu in FOXG1, and hemizygous variant p.Leu211Met in DMD), while the homozygous variant p.Glu128Ter in CHKB was inherited from both healthy heterozygous parents. 3D protein structure modeling was carried out for the missense variants in FOXG1 and DMD, which predicted steric hindrance and spatial inhibition, respectively, supporting the pathogenicity of these human mutants. Additionally, the nonsense variant in CHKB is anticipated to influence its dimerization - crucial for choline kinase function - and the nonsense variant in IQSEC2 is predicted to eliminate three functional domains. Consequently, these distinct variants found in four unrelated individuals with autism are likely indicative of loss-of-function mutations. CONCLUSIONS: In our two syndromic autism families, we discovered variants in two muscular dystrophy genes, DMD and CHKB. Given that DMD and CHKB are recognized for their participation in the non-cognitive manifestations of muscular dystrophy, it indicates that some genes transcend the boundary of apparently unrelated clinical categories, thereby establishing a novel connection between ASD and muscular dystrophy. Our findings also shed light on the complex inheritance patterns observed in Iranian consanguineous simplex families and emphasize the connection between autism spectrum disorder and muscular dystrophy. This underscores a likely genetic convergence between neurodevelopmental and neuromuscular disorders.

Humans

Glycemic and bodyweight effects of GIPR coding variation reflect differences in surface expression and intrinsic functional impairment.

The glucose-dependent insulinotropic polypeptide receptor (GIPR) is a major therapeutic target in type 2 diabetes and obesity. Missense variation in GIPR could confer phenotypic effects through alterations to constitutive activity or functional responses to GIP or pharmacological agonists. In this study, we aimed to provide a deep understanding of the molecular mechanisms that underpin the cellular and physiological impacts of GIPR coding variation by studying 30 GIPR coding variants in cellular models and pancreatic islets. Many variants showed impaired GIP-induced cyclic adenosine monophosphate responses, and population-based association analysis highlighted that these loss-of-function variants decrease body mass index but increase glycemia. In many cases, reduced function was partly driven by reduced expression at the cell surface due to impaired stability and redirection toward proteasomal degradation. Molecular dynamics simulations suggest distinct variant-induced perturbations in inter- and intrahelical interactions, which interfere with receptor stability. This study highlights the mechanisms and consequences of GIPR coding variation, which may have implications for the therapeutic targeting of this receptor.

Humans

At least three alternatively spliced mRNAs encoding two alpha subunits of the Go GTP-binding protein can be expressed in a single tissue.

Hybridization blot (Northern) analysis of mRNA coding for alpha subunits of the Go signal-transducing protein detects three bands at 5.7, 4.2, and 3.2 kilobases (kb). We showed previously that the largest is a splice variant coding for the type 2 form of the polypeptide (alpha o2) and the two smaller RNAs react with a probe specific for the seventh of the eight exons that code for the type 1 form (alpha o1). In the present work we demonstrate that the 3.2- and 4.2-kb mRNAs also result from alternative splicing, the splice site being located 31 nucleotides downstream from the termination codon of the open reading frame, and that therefore the alpha o mRNA is made up of at least nine exons. All three alpha o mRNAs are expressed in both heart and brain, more in the latter than the former, as well as in the hamster insulin-secreting tumor (HIT) cell from which the cDNAs encoding the splice variants had been cloned. In contrast, in lung and testis we found only the 5.7-kb alpha o2 mRNA. The same analysis was unable to detect alpha o-specific sequences in either kidney, pancreas (whole), spleen, or liver, while at the same time detecting strong bands for alpha s mRNA. A comparison of the nucleotide sequences of the 5'- and 3'-untranslated regions of the hamster cDNAs cloned here indicated that previously cloned alpha o cDNAs all belong to the same alpha o1A slice subclass derived from 3.2-kb mRNA. The comparison also revealed that the sequences of the untranslated regions are highly conserved among three species (rat, hamster, and brain). Their 3' tails are 99.1% (HIT versus bovine, 200 known bases) and 99.7% (HIT versus rat, 229 bases) identical, and their 5' leader sequences are 92.7% (HIT versus bovine, 165 known bases) and 90.7% (HIT versus rat, 670 bases) identical. This indicates that untranslated regions of mRNAs need not exhibit high degrees of species variation.

Animals

Ovine trophoblast protein-one: evidence for possible glycosylation.

1. The polymerase chain reaction has been used to amplify specifically the cDNA coding for the secreted form of ovine trophoblast protein-one from a preparation of total cellular RNA extracted from sheep embryos removed from ewes 16 days after mating. 2. Cloning and sequencing of the amplified cDNA revealed two new sequence variants: SPW49 having 93% similarity with deduced amino acid sequences from published cDNA data, and SPW27 a variant coding for a deleted form of ovine trophoblast protein-one. 3. The gene for ovine trophoblast protein-one is intronless. 4. This study provides further evidence for the existence of an ovine trophoblast protein-one gene family. 5. Both variants contain a potential N-glycosylation site not apparent in published sequences for ovine trophoblast protein-one.

Amino Acid Sequence

Genetic Determinants of Pulmonary Artery Size in over 50,000 Subjects with and without COPD.

RATIONALE: Pulmonary artery (PA) enlargement is a non-invasive imaging biomarker associated with pulmonary hypertension and mortality in COPD; however, its genetic determinants remain incompletely understood. OBJECTIVES: To characterize the genetic architecture of PA size across COPD-enriched and population-based cohorts. METHODS: We performed genome-wide association analyses of PA diameter using whole-genome sequencing in COPDGene (n=9,418) and ECLIPSE (n=1,859), and imputed-genotype data from the UK Biobank (n=37,073). We replicated lead variants in the Framingham Heart Study (FHS; n=3,289), incorporated all four studies into a joint meta-analysis, and identified independent signals through conditional analyses. Candidate effector genes were prioritized using coding variant annotation, colocalization, and integrative regulatory evidence. MEASUREMENTS AND MAIN RESULTS: We identified 44 independent genome-wide significant PA diameter signals within 39 loci, including 8 variants replicated in FHS, novel associations near FRMD4B, SLC20A2, BORCS7-ASMT, and KCNRG, and 5 signals in conditional analysis including multiple signals at ANO1. Genetic effects were concordant across imaging modalities and cohorts of differing COPD burden. Effector-gene prioritization nominated ABCC8, PDGFD, HMCN1, CCNE1, and TBX20, implicating pathways in vascular remodeling, developmental regulation, smooth muscle and endothelial function, ion-channel signaling, and extracellular matrix organization. Colocalization with pulse pressure GWAS demonstrated substantial shared causal variation between pulmonary and systemic vascular biology. CONCLUSIONS: In this largest genetic study of pulmonary vascular imaging to date, PA diameter exhibits a polygenic architecture consistent across imaging modalities and cohorts of differing COPD burden. The prioritized effector genes bridge rare-variant pulmonary hypertension biology with common-variant systemic vascular biology.

Pulmonary artery diameter

Inference of elevated mutation rates and variant effects using 700k exomes.

Genomic sequencing is now widely accessible for genetic diagnostics and is emerging as a component of newborn screening. This technological development generates the need to characterize incoming mutations, create comprehensive datasets of genes causing rare Mendelian disorders, and identify pathogenic variants. Large-scale exome sequencing datasets such as Genome Aggregation Database (gnomAD) have been assembled to help address these challenges. The recent release of gnomAD (v4; n = 730,947) uncovers millions of rare coding variants, many of which have arisen more than once by independent recurrent mutations in the rapidly growing recent human population. Here, we use newly developed theoretical understanding of sampling properties of rare variants to estimate key population genetics parameters of practical importance to human genetics such as demography history, mutation rate, and selection. Solely relying on population data, our method Population Inferred Estimates of Selection (PIES) identifies novel genes with loss-of-function mutational hotspots likely due to selection in spermatogonia. PIES efficiently estimates selection coefficients for heterozygous loss-of-function variants. Combining population genetics inference with variant effect predictors, PIES predicts pathogenic missense mutations and improves variant prioritization for genetic diagnostics and newborn screening.

Journal Article

The Germline SH2B3rs111340708 Splicing Variant Drives Intron Retention and Protein Instability by Impacting Clinical Outcomes in Core Binding Factor AML.

The SH2B3 gene, also known as LNK, encodes an adaptor protein that negatively regulates key hematopoietic signaling pathways, including JAK-STAT, MAPK, and PI3K/AKT, thereby maintaining hematopoietic homeostasis. SH2B3 interacts with major signaling regulators such as JAK2, MPL, FLT3, and KIT. Loss-of-function alterations have been reported in several hematologic malignancies, supporting its role as a leukemia predisposition gene. We previously identified a germline start-loss mutation (c.3G&#xa0;>&#xa0;A) in SH2B3 in a family with early-onset myeloproliferative neoplasm, demonstrating that this variant causes SH2B3 haploinsufficiency. In the present study, next-generation sequencing of 149 de novo AML patients identified a frequent intronic polymorphism (rs111340708), located within intron 6 (IVS6) of SH2B3. Although this variant has a reported minor allele frequency (MAF) of approximately 12% in European populations, it was enriched in our AML cohort, reaching 34.2% in Core Binding Factor leukemias (CBFLs). The presence of the rs111340708 variant was associated with inferior overall survival, whereas no significant association with progression-free survival was observed. Functional analyses demonstrated that this polymorphism promotes aberrant IVS6 intron retention in AML cells, resulting in reduced abundance of correctly spliced SH2B3 transcripts and predicted generation of truncated peptides and/or nonsense-mediated decay. Consistently, immunoblot analyses of AML patient samples and hematologic cell lines revealed heterogeneous SH2B3 protein expression, including additional SH2B3-immunoreactive species in variant carriers, together with reduced levels of the canonical SH2B3 protein. Collectively, these findings identify a common germline splicing polymorphism as a novel mechanism contributing to SH2B3 functional impairment in AML and highlight the potential relevance of non-coding variants in leukemia pathogenesis, with possible implications for risk stratification and future therapeutic strategies.

Humans