Search PubMedSearch

SEARCH · Search PubMed

Results for “intron retention”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Elevated intron retention implicates neuroinflammation in brains of individuals with alcohol use disorder.

Intron retention, a form of alternative RNA splicing, can occur as part of normal gene regulation or result from disruption of the splicing machinery. Retained introns can potentially form double-stranded RNA, activating innate immune sensors and inflammation. This mechanism has been implicated in cancer but has not been studied in neuropsychiatric diseases like alcohol use disorder. We systematically analysed transcriptome-wide intron retention events in post-mortem brain tissue from 142 individuals (66 with alcohol use disorder and 76 controls), encompassing 320 region-specific samples from the superior frontal cortex, nucleus accumbens, central nucleus and basolateral amygdala. Analyses were adjusted for demographic, technical and biological covariates. Validation was performed in alcohol-preferring (P) rats using long-read sequencing. In complementary experiments, immunofluorescent staining was used to detect double-stranded RNA in rat brain tissue, while single-cell RNA-sequencing was performed to test activation of double-stranded RNA-sensing pathways in human brains. Brains from individuals with alcohol use disorder showed significantly higher total intron retention compared with controls, independent of age, with females showing greater increases than males. A total of 368 introns were positively associated with alcohol use disorder, and these introns were significantly longer and had weaker splice acceptor sites compared with non-associated introns. Genes harbouring these intron retention events were enriched in Purkinje neurons, visual cortex neurons and oligodendrocytes. Computational predictions indicated these long introns could form duplex RNA structures. Increased double-stranded RNA was confirmed experimentally in multiple brain regions of alcohol-consuming rats, where it co-localized primarily with neuronal nuclei and dendrites. In individuals with alcohol use disorder, we found that multiple pathways including double-stranded RNA responses, neuroinflammation, interferon and NF-κB signalling, adaptive immunity and apoptosis were activated. In addition, NeuN-positive neuronal counts significantly decreased in both the prefrontal and visual cortices. Furthermore, single-cell analysis demonstrated upregulation of TICAM1, the target of double-stranded RNA sensor TLR3, in oligodendrocytes, as well as widespread activation of downstream inflammatory pathways across glial and neuronal cell types. These findings provide the first evidence that chronic alcohol consumption promotes an overall increase of intron retention in the brain and is associated with the presence of double-stranded RNA. Furthermore, the double-stranded RNA may contribute to neuronal loss and brain pathology by activating a neuroinflammatory response.

alcohol use disorder

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning

PRMT5-mediated intron retention triggers innate and adaptive immunity against cancer.

PRMT5 is expressed at high levels in many cancers, where it regulates diverse cellular pathways that contribute to oncogenesis. Here, we have defined a new role for PRMT5 in regulating and coordinating the interplay between the innate and adaptive immune response. This occurs, in part, through the influence of PRMT5 and E2F1 on RNA splicing and the presence of retained introns (RIs). We found that RIs have a propensity to form double-stranded RNAs that contribute to the innate response. Furthermore, many RIs contain non-canonical open-reading frames (ncORFs), which can be translated and then processed into small peptides that assemble with the MHC class I complex. Significantly, RI-derived peptides are highly immunogenic and, as a murine cancer vaccine, carrying a string of antigenic RI peptides, delayed tumour growth and enhanced survival. RIs are present in human tumour cells, and we identified T lymphocytes in human cancer patients, with antigen specificity for RI-derived peptides, that killed human tumour cells in vitro. Regulating intron retention thus offers a new therapeutic approach to enhance tumour immunogenicity.

Animals

The Germline SH2B3rs111340708 Splicing Variant Drives Intron Retention and Protein Instability by Impacting Clinical Outcomes in Core Binding Factor AML.

The SH2B3 gene, also known as LNK, encodes an adaptor protein that negatively regulates key hematopoietic signaling pathways, including JAK-STAT, MAPK, and PI3K/AKT, thereby maintaining hematopoietic homeostasis. SH2B3 interacts with major signaling regulators such as JAK2, MPL, FLT3, and KIT. Loss-of-function alterations have been reported in several hematologic malignancies, supporting its role as a leukemia predisposition gene. We previously identified a germline start-loss mutation (c.3G > A) in SH2B3 in a family with early-onset myeloproliferative neoplasm, demonstrating that this variant causes SH2B3 haploinsufficiency. In the present study, next-generation sequencing of 149 de novo AML patients identified a frequent intronic polymorphism (rs111340708), located within intron 6 (IVS6) of SH2B3. Although this variant has a reported minor allele frequency (MAF) of approximately 12% in European populations, it was enriched in our AML cohort, reaching 34.2% in Core Binding Factor leukemias (CBFLs). The presence of the rs111340708 variant was associated with inferior overall survival, whereas no significant association with progression-free survival was observed. Functional analyses demonstrated that this polymorphism promotes aberrant IVS6 intron retention in AML cells, resulting in reduced abundance of correctly spliced SH2B3 transcripts and predicted generation of truncated peptides and/or nonsense-mediated decay. Consistently, immunoblot analyses of AML patient samples and hematologic cell lines revealed heterogeneous SH2B3 protein expression, including additional SH2B3-immunoreactive species in variant carriers, together with reduced levels of the canonical SH2B3 protein. Collectively, these findings identify a common germline splicing polymorphism as a novel mechanism contributing to SH2B3 functional impairment in AML and highlight the potential relevance of non-coding variants in leukemia pathogenesis, with possible implications for risk stratification and future therapeutic strategies.

Humans

Gene expression is stable despite widespread cis and trans regulatory divergence in Saccharomyces yeasts.

Regulatory evolution can alter phenotypes, but cis- and trans-regulatory mechanisms may also diverge extensively while total transcript abundance remains stable. Comparisons of parental expression with allele-specific expression in F1 hybrids provide a framework for separating cis- and trans-regulatory effects because both parental alleles are measured in a shared trans-regulatory environment. Here, we analyzed RNA sequencing data from Saccharomyces cerevisiae, Saccharomyces paradoxus, and their F1 hybrid. Among the 4,164 genes with sufficient allele-specific support for strict classification, 2,134 (51.2%) showed detectable cis and/or trans regulatory divergence. However, hybrid expression remained largely conserved, with 81.5% of genes not significantly different from either parent. Compensatory cis-trans divergence predominated over reinforcing divergence; cross-replicate estimation reduced the apparent magnitude of this excess, but opposite-sign effects remained predominant in all 20 non-overlapping replicate comparisons. To connect gene expression to genome sequence, we analyzed the strongly cis-diverged locus LYS2 and found species differences in promoter architecture, including an S. cerevisiae-specific AT-rich insertion, altered spacing among candidate regulatory features, and a promoter-proximal TATA-like element unique to S. cerevisiae. Sequence-based nucleosome prediction suggests that these differences create a broader promoter-proximal nucleosome-depleted region in S. cerevisiae than in S. paradoxus. We also quantified allele-resolved intron retention and found that allele-resolved intron retention was broadly conserved, with only rare locus-specific hybrid-associated shifts. Together, these results show that regulatory divergence is widespread but often buffered in the hybrid, whereas intron-retention divergence is comparatively limited.

Saccharomyces

Beyond in silico prediction: multi-omics to identify a pathogenic deep intronic HNRNPK variant in Au-Kline syndrome.

Pathogenic variants in HNRNPK are associated with autosomal dominant Au-Kline syndrome (AKS, Au-Kline-Okamoto syndrome, OMIM #616580). This syndrome is characterized by developmental delay and intellectual disability, hypotonia, and distinctive facial features. Despite the use of whole-genome sequencing (WGS) as a powerful diagnostic tool, we nearly dismissed a novel intronic variant (NM_031263.4(HNRNPK):c.214-55 T > A) affecting HNRNPK splicing and function. Although commonly used bioinformatic splice prediction tools, including SpliceAI and PDIVAS, yielded inconclusive results, Face2Gene analysis indicated a high phenotypic similarity to AKS. Characteristic facial features described by Choufani et al. [1] supported the clinical diagnosis of AKS. Subsequent functional studies demonstrated aberrant splicing with intron retention, and DNA methylation profiling revealed a positive HNRNPK-specific episignature. These insights and the de novo status support an evaluation as likely pathogenic. This case report supports the relevance of facial analysis and comprehensive variant validation strategies, particularly for deep intronic variants with ambiguous in silico splicing predictions.

Journal Article

The periphery of nuclear speckles defines a spatially and temporally regulated compartment of long-lived intron-retained RNAs that resolves during mitosis.

RNA localization adds a fundamental layer to gene expression by determining when and where translation-ready mRNAs become available, yet how this timing is coordinated with nuclear architecture and cell-cycle progression remains unclear. Here we identify a subnuclear RNA niche at the nuclear speckle periphery that couples intron retention to cell-cycle-timed RNA release. Using compartment-resolved transcriptional inhibition, sequence-based deep learning and single-molecule and super-resolution RNA imaging in human pluripotent stem cells, we define a class of nuclear RNAs with long-lived retained introns that persist for hours and are enriched in transcripts encoding regulators of genome maintenance and mitosis, including centromere and kinetochore assembly, DNA repair and telomere maintenance. Long-lived retained introns exhibit elevated GC content, predicted structural stability and enrichment for nuclear speckle-associated RNA-binding proteins. In interphase, these RNAs localize to a distinct nuclear speckle-peripheral RNA niche in a spatial arrangement conserved across cell types. During mitotic remodelling, they undergo coordinated, kinase-dependent splicing and are released into the cytoplasm of early G1 daughter cells. Together, these findings link cis-encoded intronic features, subnuclear organization and mitotic remodelling to temporal control of RNA fate.

Mitosis

Dysregulation of U12-Type Splicing in Lupus Neutrophils.

OBJECTIVE: Neutrophil dysfunction is a hallmark of systemic lupus erythematosus (SLE), but its molecular basis remains unclear. This study explores transcriptional and posttranscriptional changes in low-density granulocytes (LDGs), a proinflammatory neutrophil subset expanded in SLE, focusing on NADPH oxidase (Nox) function and minor intron splicing. METHODS: LDGs and normal-density granulocytes (NDGs) were isolated from patients with SLE and healthy controls (HCs). CYBA (p22phox) expression was evaluated at transcript and protein levels. Nox activity was measured using luminol assays. Bulk RNA sequencing (RNA-seq) and rMATS software were used to assess alternative splicing, particularly of U12-type intron-containing genes. RESULTS: CYBA expression was reduced in SLE LDGs (n = 11) compared to SLE and HC NDGs (n = 6), with levels resembling those in chronic granulomatous disease neutrophils. SLE LDGs exhibited impaired Nox activity (n = 7 SLE, n = 12 HC). CYBA is a U12 intron-containing gene, and transcriptomic analysis revealed broad down-regulation of this gene class in SLE LDGs, suggesting minor spliceosome dysfunction. rMATS analysis showed increased U12-type intron retention and widespread splicing defects-including exon skipping and mutually exclusive exon use-in genes such as GBP5, MAEA, and STX10. These abnormalities were validated in an independent long-read RNA-seq data set from SLE peripheral blood mononuclear cells. Importantly, splicing disruptions correlated with disease activity and autoantibody profiles. CONCLUSION: Impaired U12-dependent splicing may contribute to neutrophil dysfunction in SLE, potentially via defective oxidative burst and altered immune regulation. These findings highlight the minor spliceosome as a novel player in lupus pathogenesis.

Humans

RNU4ATAC-opathy: Clinical, molecular, and transcriptomic insights from a large cohort.

PURPOSE: We aim to better define the genotype and phenotype spectrum of RNU4ATAC-opathy, demonstrate the utility of RNA sequencing (RNA-seq) for variant classification, and highlight the challenges in detecting variants in this noncoding gene. METHODS: Sixty individuals with molecularly confirmed RNU4ATAC-opathy were recruited from multiple clinical and research centers internationally. RNA-seq was available for 7 affected individuals. RESULTS: We report the clinical and molecular findings of 60 individuals, including 42 not previously described, and 33 distinct RNU4ATAC variants, 13 of which are novel. Core features in this cohort-present in most individuals assessed and varying in severity-include microcephaly, short stature, skeletal anomalies, developmental delay, cerebral anomalies, skin conditions, and immune deficiency. Additional findings, such as diabetes, holoprosencephaly, and the absence of various core features in some individuals, highlight the broad phenotypic spectrum. All individuals who underwent RNA-seq showed a consistent pattern of minor intron retention. In 6 individuals, RNA-seq enabled the reclassification of variants of uncertain significance as likely pathogenic. Although RNU4ATAC variants are generally covered by clinical exomes, they are often overlooked in analysis because of their noncoding nature. CONCLUSION: This study highlights the variability of phenotypes and genotypes associated with RNU4ATAC-opathy. Laboratories should ensure RNU4ATAC and other noncoding genes are appropriately assessed by their analysis pipelines.

Lowry-Wood syndrome

Alternative transcription of the mouse Gh gene identifies an immune-associated transcript with species-specific structural divergence.

Growth hormone (GH) in mice is primarily expressed in the anterior pituitary, although Gh expression has been reported in extrapituitary tissues, including immune organs. However, the structure of immune-associated Gh transcripts remains poorly characterized. To determine whether splenic Gh transcripts differ from pituitary Gh mRNA, 5'- and 3'-rapid amplification of cDNA ends (RACE) analyses were performed. While 3' RACE showed a shared polyadenylation site, 5' RACE identified a novel exon located approximately 2 kb upstream of the conventional exon 1, generating a transcript (spl-Gh mRNA) with a distinct first exon but shared downstream exons with pituitary Gh mRNA (pit-Gh mRNA). RT-PCR analysis revealed that spl-Gh mRNA is predominantly expressed in immune tissues such as spleen and bone marrow, and its distribution did not correlate with Pit-1 mRNA expression. Quantitative RT-PCR further demonstrated that spl-Gh mRNA was expressed at levels comparable to those of pit-Gh mRNA in the mouse spleen, indicating that spl-Gh is one of the major Gh transcript forms in this tissue. Sequence analysis indicated that spl-Gh mRNA is predicted to retain coding potential for a GH protein. Comparative genomic analyses further demonstrated that genomic features associated with the spl-Gh transcriptional unit are conserved only in a subset of closely related Mus species. In contrast, although a spl-Gh-related transcript was detected in rat spleen, no properly spliced mouse-like transcript was identified under the present experimental conditions. The detected transcript exhibited intron retention and an in-frame stop codon, suggesting that it is unlikely to produce a functional GH protein. These findings identify a distinct immune-associated Gh transcript generated through alternative transcription of the mouse Gh gene and suggest that immune-associated Gh transcriptional mechanisms have undergone species-specific divergence among rodents. Together, these findings reveal previously unrecognized complexity in Gh gene regulation and highlight species-specific differences in immune-associated Gh transcripts.

Animals

Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.

RNA splicing shapes neuronal identity and disease risk, yet current maps lack the developmental resolution and depth to resolve this complexity. Here, we integrate deep long-read RNA sequencing and proteomics in induced pluripotent stem cell-derived cortical neurons to generate a high-resolution proteogenomic atlas of human neuron development. We identify 182,371 mRNA isoforms (over half previously unknown) and provide direct peptide evidence for the translation of hundreds of novel protein-coding sequences. Population genetics demonstrates that variants affecting novel exons and splice sites are under negative selection, underscoring the potential significance of these isoforms. During neuronal maturation, we observe that autism risk genes undergo dynamic isoform switching, including microexon inclusion and intron retention, that remodel key protein domains and regulatory regions. Furthermore, we uncover widespread, long-range coordination between alternative transcript processing events, including transcription start sites, exon splicing, and polyadenylation. Finally, our atlas enables variant reinterpretation in autism, highlighting the value of an isoform-centric view for interpreting pathogenic variation in neurodevelopment.

Humans

A high resolution A-to-I editing map in the mouse identifies editing events controlled by pre-mRNA splicing.

Pre-mRNA-splicing and adenosine to inosine (A-to-I) RNA-editing occur mostly cotranscriptionally. During A-to-I editing, a genomically encoded adenosine is deaminated to inosine by adenosine deaminases acting on RNA (ADARs). Editing-competent stems are frequently formed between exons and introns. Consistently, studies using reporter assays have shown that splicing efficiency can affect editing levels. Here, we use Nascent-seq and identify ∼90,000 novel A-to-I editing events in the mouse brain transcriptome. Most novel sites are located in intronic regions. Unlike previously assumed, we show that both ADAR (ADAR1) and ADARB1 (ADAR2) can edit repeat elements and regular transcripts to the same extent. We find that inhibition of splicing primarily increases editing levels at hundreds of sites, suggesting that reduced splicing efficiency extends the exposure of intronic and exonic sequences to ADAR enzymes. Lack of splicing factors NOVA1 or NOVA2 changes global editing levels, demonstrating that alternative splicing factors can modulate RNA editing. Finally, we show that intron retention rates correlate with editing levels across different brain tissues. We therefore demonstrate that splicing efficiency is a major factor controlling tissue-specific differences in editing levels.

Adenosine Deaminase

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

Single-Cell Splicing Isoform Atlas of the Adult Human Heart and Heart Failure.

BACKGROUND: Alternative splicing plays crucial roles in normal heart development and cardiac disease by influencing protein-coding sequences, functional domains, and molecular networks. However, a detailed characterization of the human heart isoform landscape remains incomplete. METHODS: Leveraging long-read single-nucleus RNA sequencing and computational analysis, we dissected full-length isoform heterogeneities, expression patterns, and usage shifts across cell types, cell states, and cardiac conditions of the adult left ventricle. We applied in silico approaches to assess the functional relevance of identified isoforms; validated isoform compositions of representative cardiac genes using reverse transcription quantitative polymerase chain reaction and targeted amplicon sequencing; and developed a web server for interactive navigation of our results. RESULTS: The data revealed that isoform heterogeneity is widespread in the cardiac cellular system, serving as a posttranscriptional buffer mechanism that calibrates the molecule reservoirs in human hearts. In healthy left ventricles, ≈30% of cell type-specific genes were polyform, using multiple isoforms tailored to cell type-specific programs. Among ubiquitously expressed genes, >300 showed differential isoform usage with cell type specificity in normal hearts. Comparisons of cardiomyocytes across conditions uncovered 379 genes with marked isoform usage shifts, most of which are predicted to change protein coding outcomes through direct changes in protein coding sequences and switches between intron retention and non-protein-coding biotypes. In contrast, cell state-specific programs tend to operate on monoform genes associated with changes among cell states. In addition, our data revealed heart failure-associated differential isoform usage events in stromal and immune cell types in the cardiac microenvironment. CONCLUSIONS: We present a comprehensive atlas of splicing isoforms in the normal adult heart and heart failure through long-read single-nucleus RNA sequencing and computational analyses. The results suggest crucial roles of isoforms in buffering core cellular programs and contributing to disease-associated cell states. The full-length details of these cell-specific isoforms serve as an important reference for downstream translational and mechanistic studies and are available on our online data portal at https://github.com/gaolabtools/heart-isoform-atlas.

Humans

Novel compound heterozygous POR variants in a neonate with Antley-Bixler syndrome and 46,XY DSD: a case report and literature review.

BACKGROUND: Cytochrome P450 oxidoreductase deficiency (PORD) is an ultra-rare autosomal recessive disorder within the congenital adrenal hyperplasia (CAH) spectrum, characterized by a broad clinical spectrum involving steroidogenesis defects, genital anomalies, and skeletal abnormalities. CASE PRESENTATION: We report a phenotypically female neonate with a 46,XY karyotype whose postnatal diagnostic evaluation was initiated after newborn screening revealed elevated 17-hydroxyprogesterone (17-OHP) concentration. The patient presented with mild hypertelorism, mild nasal hypoplasia, and low-set bilateral ears, along with female external genitalia consistent with disorder of sex development (DSD) and anal atresia. Radiological evaluation revealed femoral bowing and subsequent fracture. The craniofacial and skeletal abnormalities were consistent with the features of Antley-Bixler syndrome (ABS). Endocrine evaluation revealed elevated progesterone, markedly reduced testosterone, and secondary hyperaldosteronism. Genetic analysis identified three novel variants in the POR gene (NM_001395413.1): the patient harbored a paternal c.1187_1195dup (p.Pro396_Glu398dup) variant and two maternally inherited variants in cis, c.1447G>A (p.Gly483Ser) and c.1806 + 4_1806 + 28del. Protein structural modeling predicted that the p.Pro396_Glu398dup and p.Gly483Ser may disrupt the flavin adenine dinucleotide (FAD)-binding domain. RNA sequencing (RNA-seq) confirmed that the intronic variant c.1806 + 4_1806 + 28del caused aberrant splicing, resulting in partial intron retention and predicted impairment of the nicotinamide adenine dinucleotide phosphate (NADPH)-binding domain. According to American College of Medical Genetics and Genomics (ACMG) guidelines and incorporating functional evidence, c.1187_1195dup and c.1806 + 4_1806 + 28del were reclassified as likely pathogenic (LP), whereas c.1447G>A remained a variant of uncertain significance (VUS). CONCLUSIONS: This study describes a neonate with PORD caused by three novel POR variants and expands the known clinical spectrum of PORD by identifying rare manifestations including anal atresia and hearing loss. RNA-seq provided valuable functional evidence for variant interpretation and facilitated accurate molecular diagnosis. These findings highlight the importance of integrating genetic phasing, transcript-level functional analysis, and comprehensive clinical evaluation for precise diagnosis and counseling in rare endocrine disorders.

Humans

Gene expression is stable despite widespread cis and trans regulatory divergence in Saccharomyces yeasts.

Regulatory evolution can alter phenotypes, but cis- and trans-regulatory mechanisms may also diverge extensively while total transcript abundance remains stable. Comparisons of parental expression with allele-specific expression in F1 hybrids provide a framework for separating cis- and trans-regulatory effects because both parental alleles are measured in a shared trans-regulatory environment. Here, we analyzed RNA sequencing data from Saccharomyces cerevisiae, Saccharomyces paradoxus, and their F1 hybrid. Regulatory divergence was widespread, with 61.3% of tested orthologs showing significant divergence in at least one cis or trans component. However, hybrid expression remained largely conserved, with 81.6% of genes not significantly different from either parent. Compensatory cis-trans divergence predominated over reinforcing divergence, consistent with widespread buffering of transcript abundance. To connect genome-wide patterns to mechanism, we analyzed the strongly cis-diverged locus LYS2 and found species differences in promoter architecture, including an S. cerevisiae-specific AT-rich insertion, altered spacing among candidate regulatory features, and a promoter-proximal TATA-like element unique to S. cerevisiae. Sequence-based nucleosome prediction suggests that these differences create a broader promoter-proximal nucleosome-depleted region in S. cerevisiae than in S. paradoxus. We also quantified allele-resolved intron retention and found that splicing was broadly conserved, with only rare locus-specific hybrid-associated shifts. Together, these results show that regulatory divergence is widespread but often buffered in the hybrid, whereas post-transcriptional divergence is comparatively limited.

Gene expression

Somatic and germinal mosaicism of a canonical splicing variant causing limb-girdle muscular dystrophy type 1B.

Limb-girdle muscular dystrophy type 1B is one of several muscular dystrophies caused by pathogenic variants in the LMNA gene. In this study, we investigated the clinical, pathological, and genetic findings of an LGMD1B family. Genetic sequencing identified the proband and her younger brother both carried the canonical splicing c.513 + 1G > A variant in the LMNA gene. The variant was absent in the proband's mother, and a certain percentage of the LMNA variant was identified in the venous blood, urine, and semen sample of the proband's father by pyrophosphate sequencing. Further cDNA analysis demonstrated that the canonical splicing c.513 + 1G > A variant in intron 2 induced retention of the first 45 bp of intron 2, resulting in an in-frame insertion of 15 amino acids. Our study directly confirmed the presence of somatic and germinal mosaicism in the LGMD1B family and the pathogenicity of the canonical splicing variant in the LMNA gene.

Humans

Dual Aberrant Splicing Caused by an Apparently Missense CHD7 Variant, c.5273A>G (p.Asp1758Gly), in CHARGE Syndrome.

CHARGE syndrome is a rare congenital disorder primarily attributed to heterozygous pathogenic variants of the CHD7 gene. Most pathogenic CHD7 variants are loss-of-function (LoF) variants, whereas the interpretation of missense variants remains challenging in the absence of functional evidence for their pathogenicity. We report a female infant presenting with clinical features characteristic of CHARGE syndrome. Targeted sequencing identified a heterozygous CHD7 variant (NM_017780.4:c.5273A>G), initially annotated as a missense substitution p.Asp1758Gly. This variant has been previously reported and registered with conflicting pathogenicity classifications; however, its transcript-level consequences remain unclear. Long-PCR-based RNA sequencing of total RNA from peripheral blood mononuclear cells revealed two aberrant splicing patterns associated with the variant: a predominant transcript carrying a 28-bp deletion due to cryptic donor splice-site activation, and a minor transcript with partial intron 24 retention. Both transcripts were predicted to result in premature termination codons. These findings demonstrate that c.5273A>G functions as a LoF variant through dual aberrant splicing rather than a simple missense substitution. This case underscores the importance of RNA-level splicing analysis for the accurate interpretation and classification of CHD7 missense variants.

CHD7