Search PubMedSearch

SEARCH · Search PubMed

Results for “intron position”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

Calcium sensors and their interacting protein kinases: genomics of the Arabidopsis and rice CBL-CIPK signaling networks.

Calcium signals mediate a multitude of plant responses to external stimuli and regulate a wide range of physiological processes. Calcium-binding proteins, like calcineurin B-like (CBL) proteins, represent important relays in plant calcium signaling. These proteins form a complex network with their target kinases being the CBL-interacting protein kinases (CIPKs). Here, we present a comparative genomics analysis of the full complement of CBLs and CIPKs in Arabidopsis and rice (Oryza sativa). We confirm the expression and transcript composition of the 10 CBLs and 25 CIPKs encoded in the Arabidopsis genome. Our identification of 10 CBLs and 30 CIPKs from rice indicates a similar complexity of this signaling network in both species. An analysis of the genomic evolution suggests that the extant number of gene family members largely results from segmental duplications. A phylogenetic comparison of protein sequences and intron positions indicates an early diversification of separate branches within both gene families. These branches may represent proteins with different functions. Protein interaction analyses and expression studies of closely related family members suggest that even recently duplicated representatives may fulfill different functions. This work provides a basis for a defined further functional dissection of this important plant-specific signaling system.

Amino Acid Sequence

Elevated intron retention implicates neuroinflammation in brains of individuals with alcohol use disorder.

Intron retention, a form of alternative RNA splicing, can occur as part of normal gene regulation or result from disruption of the splicing machinery. Retained introns can potentially form double-stranded RNA, activating innate immune sensors and inflammation. This mechanism has been implicated in cancer but has not been studied in neuropsychiatric diseases like alcohol use disorder. We systematically analysed transcriptome-wide intron retention events in post-mortem brain tissue from 142 individuals (66 with alcohol use disorder and 76 controls), encompassing 320 region-specific samples from the superior frontal cortex, nucleus accumbens, central nucleus and basolateral amygdala. Analyses were adjusted for demographic, technical and biological covariates. Validation was performed in alcohol-preferring (P) rats using long-read sequencing. In complementary experiments, immunofluorescent staining was used to detect double-stranded RNA in rat brain tissue, while single-cell RNA-sequencing was performed to test activation of double-stranded RNA-sensing pathways in human brains. Brains from individuals with alcohol use disorder showed significantly higher total intron retention compared with controls, independent of age, with females showing greater increases than males. A total of 368 introns were positively associated with alcohol use disorder, and these introns were significantly longer and had weaker splice acceptor sites compared with non-associated introns. Genes harbouring these intron retention events were enriched in Purkinje neurons, visual cortex neurons and oligodendrocytes. Computational predictions indicated these long introns could form duplex RNA structures. Increased double-stranded RNA was confirmed experimentally in multiple brain regions of alcohol-consuming rats, where it co-localized primarily with neuronal nuclei and dendrites. In individuals with alcohol use disorder, we found that multiple pathways including double-stranded RNA responses, neuroinflammation, interferon and NF-κB signalling, adaptive immunity and apoptosis were activated. In addition, NeuN-positive neuronal counts significantly decreased in both the prefrontal and visual cortices. Furthermore, single-cell analysis demonstrated upregulation of TICAM1, the target of double-stranded RNA sensor TLR3, in oligodendrocytes, as well as widespread activation of downstream inflammatory pathways across glial and neuronal cell types. These findings provide the first evidence that chronic alcohol consumption promotes an overall increase of intron retention in the brain and is associated with the presence of double-stranded RNA. Furthermore, the double-stranded RNA may contribute to neuronal loss and brain pathology by activating a neuroinflammatory response.

alcohol use disorder

Evolutionary history and recombination in the mitochondrial carrier SLC25 superfamily analyzed by similarities in the exon and transmembrane α-helix sequences.

Mitochondrial carriers (MCs), which constitute a superfamily also called the solute carrier family 25 (SLC25), are characterized by conserved signature motif sequences and a six-transmembrane α-helical transporter domain. They transport a wide variety of substrates ranging from protons, inorganic ions, citric acid cycle intermediates, and amino acids to nucleotides and cofactors. The superfamily members can be divided into subfamilies, each with a distinct substrate specificity. In an attempt to understand how different subfamilies have evolved, we analyzed the protein sequences of the exons (with conserved boundaries) and the six transmembrane α-helices of MCs from highly diverged organisms. The results show that some MC subfamilies have all exons and transmembrane α-helices most similar to a closely related subfamily, which is consistent with a scenario of gene duplication and mutational divergence from a last common ancestor. However, several MC subfamilies appear to be mosaics of exons and transmembrane α-helices most similar to different and distant subfamilies, which in some cases could be explained by recombination between the superfamily genes during evolution. It seems that this latter mechanism could have played a role in the formation of new subfamilies with different substrate specificities by the combination of MC transporter domain segments that had been optimized previously for binding specific portions of the substrates. This study presents novel evolutionary relationships between MC subfamilies and may provide clues for how protein superfamilies have expanded and how to investigate their evolution.

Evolution, Molecular

Comparison of cloned rabbit and mouse beta-globin genes showing strong evolutionary divergence of two homologous pairs of introns.

Cloned beta-globin genes of both mouse and rabbit each contain a large and a small intervening sequence (intron) of about equal length at precisely the same positions relative to the coding sequence. The homologous introns show some sequence similarity, particularly at the junctions with the coding sequence. They most probably arose from a common ancestral sequence and diverged substantially during evolution.

Animals

Beyond in silico prediction: multi-omics to identify a pathogenic deep intronic HNRNPK variant in Au-Kline syndrome.

Pathogenic variants in HNRNPK are associated with autosomal dominant Au-Kline syndrome (AKS, Au-Kline-Okamoto syndrome, OMIM #616580). This syndrome is characterized by developmental delay and intellectual disability, hypotonia, and distinctive facial features. Despite the use of whole-genome sequencing (WGS) as a powerful diagnostic tool, we nearly dismissed a novel intronic variant (NM_031263.4(HNRNPK):c.214-55 T > A) affecting HNRNPK splicing and function. Although commonly used bioinformatic splice prediction tools, including SpliceAI and PDIVAS, yielded inconclusive results, Face2Gene analysis indicated a high phenotypic similarity to AKS. Characteristic facial features described by Choufani et al. [1] supported the clinical diagnosis of AKS. Subsequent functional studies demonstrated aberrant splicing with intron retention, and DNA methylation profiling revealed a positive HNRNPK-specific episignature. These insights and the de novo status support an evaluation as likely pathogenic. This case report supports the relevance of facial analysis and comprehensive variant validation strategies, particularly for deep intronic variants with ambiguous in silico splicing predictions.

Journal Article

Characterization of group I introns in generating circular RNAs as vaccines.

Circular RNAs are an increasingly important class of RNA molecules that can be engineered as RNA vaccines and therapeutics. Here, we screened eight different group I introns for their ability to circularize and delineated different features that are important for their function. First, we identified the Scytalidium dimidiatum group I intron as causing minimal innate immune activation inside cells, underscoring its potential to serve as an effective RNA vaccine without triggering unwanted reactogenicity. Additionally, mechanistic RNA structure analysis was used to identify the P9 domain as important for circularization, showing that swapping sequences can restore pairing to improve the circularization of poor circularizers. We also determined the diversity of sequence requirements for the exon 1 and exon 2 (E1 and E2) domains of different group I introns and engineered a S1 tag within the domains for positive purification of circular RNAs. In addition, this flexibility in E1 and E2 enables substitution with less immunostimulatory sequences to enhance protein production. Our work deepens the understanding of the properties of group I introns, expands the panel of introns that can be used, and improves the manufacturing process to generate circular RNAs for vaccines and therapeutics.

RNA, Circular

A novel splice-altering TNC variant (c.5247A > T, p.Gly1749Gly) in an Chinese family with autosomal dominant non-syndromic hearing loss.

BACKGROUND: This study aims to analyze the pathogenic gene in a Chinese family with non-syndromic hearing loss and identify a novel mutation site in the TNC gene. METHODS: A five-generation Chinese family from Anhui Province, presenting with autosomal dominant non-syndromic hearing loss, was recruited for this study. By analyzing the family history, conducting clinical examinations, and performing genetic analysis, we have thoroughly investigated potential pathogenic factors in this family. The peripheral blood samples were obtained from 20 family members, and the pathogenic genes were identified through whole exome sequencing. Subsequently, the mutation of gene locus was confirmed using Sanger sequencing. The conservation of TNC mutation sites was assessed using Clustal Omega software. We utilized functional prediction software including dbscSNV_AdaBoost, dbscSNV_RandomForest, NNSplice, NetGene2, and Mutation Taster to accurately predict the pathogenicity of these mutations. Furthermore, exon deletions were validated through RT-PCR analysis. RESULTS: The family exhibited autosomal dominant, progressive, post-lingual, non-syndromic hearing loss. A novel synonymous variant (c.5247A > T, p.Gly1749Gly) in TNC was identified in affected members. This variant is situated at the exon-intron junction boundary towards the end of exon 18. Notably, glycine residue at position 1749 is highly conserved across various species. Bioinformatics analysis indicates that this synonymous mutation leads to the disruption of the 5' end donor splicing site in the 18th intron of the TNC gene. Meanwhile, verification experiments have demonstrated that this synonymous mutation disrupts the splicing process of exon 18, leading to complete exon 18 skipping and direct splicing between exons 17 and 19. CONCLUSION: This novel splice-altering variant (c.5247A > T, p.Gly1749Gly) in exon 18 of the TNC gene disrupts normal gene splicing and causes hearing loss among HBD families.

Adult

Ovalbumin gene: evidence for a leader sequence in mRNA and DNA sequences at the exon-intron boundaries.

Selected regions of cloned EcoRI fragments of the chicken ovalbumin gene have been sequenced. The positions where the sequences coding for ovalbumin mRNA (ov-mRNA) are interrupted in the genome have been determined, and a previously unreported interruption in the DNA sequences coding for the 5' nontranslated region of the messenger has been discovered. Because directly repeated sequences are found at exon-intron boundaries, the nucleotide sequence alone cannot define unique excision-ligation points for the processing of a possible ov-mRNA precursor. However, the sequences in these boundary regions share common features; this leads to the proposal that there are, in fact, unique excision-ligation points common to all boundaries.

Animals

Structural basis of substrate recognition by human tRNA splicing endonuclease TSEN.

Heterotetrameric human transfer RNA (tRNA) splicing endonuclease TSEN catalyzes intron excision from precursor tRNAs (pre-tRNAs), utilizing two composite active sites. Mutations in TSEN and its associated RNA kinase CLP1 are linked to the neurodegenerative disease pontocerebellar hypoplasia (PCH). Despite the essential function of TSEN, the three-dimensional assembly of TSEN-CLP1, the mechanism of substrate recognition, and the structural consequences of disease mutations are not understood in molecular detail. Here, we present single-particle cryogenic electron microscopy reconstructions of human TSEN with intron-containing pre-tRNAs. TSEN recognizes the body of pre-tRNAs and pre-positions the 3' splice site for cleavage by an intricate protein-RNA interaction network. TSEN subunits exhibit large unstructured regions flexibly tethering CLP1. Disease mutations localize far from the substrate-binding interface and destabilize TSEN. Our work delineates molecular principles of pre-tRNA recognition and cleavage by human TSEN and rationalizes mutations associated with PCH.

Introns

Aberrant DNA methylation of genes regulating CD4+ T cell HIV-1 reservoir in women with HIV.

BACKGROUND: The HIV-1 reservoir in CD4+ T cells (HRCD4) pose a major challenge to curing HIV, with many of its mechanisms still unclear. HIV-1 DNA integration and immune responses may alter the host's epigenetic landscape, potentially silencing HIV-1 replication. METHODS: This study used bisulphite capture DNA methylation sequencing in CD4+ T cells from the blood of 427 virally suppressed women with HIV to identify differentially methylated sites and regions associated with HRCD4. RESULTS: The average total HRCD4 size was 1409 copies per million cells, with most proviruses defective and only a small proportion intact. The study identified 245 differentially methylated CpG sites and 85 regions linked to HRCD4 size, with 52% of significant sites in intronic regions. Genes associated with HRCD4 were involved in viral replication, HIV-1 latency and cell growth and apoptosis. HRCD4 size was inversely related to DNA methylation of interferon signalling genes and positively associated with methylation at known HIV-1 integration sites. HRCD4-associated genes were enriched on the pathways related to immune defence, transcription repression and host-virus interactions. CONCLUSIONS: These findings suggest that HIV-1 reservoir is linked to aberrant DNA methylation in CD4+ T cells, offering new insights into epigenetic mechanisms of HIV-1 latency and potential molecular targets for eradication strategies. KEY POINTS: Study involved 427 women with HIV. Identified 245 aberrant DNA methylation sites and 85 methylation regions in CD4+ T cells linked to the HIV-1 reservoir. Highlighted genes are involved in viral replication, immune defence, and host genome integration. Findings suggest potential molecular targets for eradication strategies.

Humans

Positional grammar of transcription factor binding partitions developmental and stress-response regulation in plants.

Understanding how transcription factor binding site (TFBS) position influences gene regulation remains a fundamental challenge in plants. Here, we integrate conserved multiDAP TFBS maps for 244 transcription factors (TFs) with single-nucleus chromatin accessibility, cell type-resolved gene expression, and hormone-response datasets across Brassicaceae species to determine how TFBS position relates to regulatory function. Although conserved TFBSs are enriched near transcription start sites (TSSs), TSS-proximal accessibility poorly predicts cell type-specific expression. Instead, cell type-specific expression correlates best with conserved TFBSs embedded in cell type-restricted chromatin, with TF family-specific distributions across distal promoters and introns. In contrast, TSS-proximal TFBSs in broadly accessible chromatin are associated with rapid transcriptional responses to abiotic and biotic stress hormones. Coding sequence TFBSs mark a distinct regulatory context in which the same DNA sequence encodes both amino acid sequence and TF motifs, including evidence that CDS-localized ABR1 binding may contribute to repression during hormone response. Finally, distal upstream regions contain conserved multi-family TF clusters with enhancer-like features overlapping rare cell type-specific accessible chromatin and enriched near genes controlling embryonic, meristematic, and hormone-dependent developmental patterning. Together, these results support a positional grammar in which TFBS position and chromatin context jointly partition developmental, stress-responsive, and repressive regulatory output in plants.

Transcription Factors

Genome-wide identification and expression profiling of HSD3B and SDR42E1 genes in the Pacific oyster (Crassostrea gigas): potential associations with gonadal development.

Sex steroids are lipid-soluble signaling molecules that regulate sex differentiation, reproductive development and physiological homeostasis in animals. 3β-Hydroxysteroid dehydrogenase/Δ5-Δ4 isomerase (3β-HSD) is a key steroidogenic enzyme, whereas SDR42E1, an extended short-chain dehydrogenase/reductase, has been implicated in sterol- and steroid-related metabolism. However, the composition, evolutionary relationships and expression patterns of the HSD3B- and SDR42E1-related genes in bivalve gonadal development remain poorly characterized. In this study, five PF01073-containing genes, comprising three CgHsd3b and two CgSdr42e1 genes, were identified in the Pacific oyster Crassostrea gigas. Phylogenetic analysis separated the proteins into HSD3B-related and SDR42E1-related groups, and gene-structure and motif analyses indicated subfamily-level divergence. All five proteins retained the SDR domain but differed in exon-intron structure and motif composition. Each contained the extended-SDR TGxxGxxG motif, whereas exact classical [ST]GxxxGxG and NNAG motifs were absent. Tyr- and Lys-equivalent residues were conserved, while the HSD3B1 Ser-equivalent position contained Thr in two C. gigas proteins and Ser in one. These features support their classification as extended-SDR proteins but do not establish enzymatic activity or substrate specificity. The three CgHsd3b genes were dispersed on one chromosome, whereas CgSdr42e1-1 and CgSdr42e1-2 were adjacent on another chromosome, suggesting a possible local duplication event for the CgSdr42e1 pair. Public RNA-seq data showed distinct tissue- and gonadal-stage expression patterns, with several genes displaying gonad-biased or female-stage-associated expression. Independent RT-qPCR profiling of the representative genes CgHsd3b-3 and CgSdr42e1-1 detected stage-dependent expression, although tissue rankings differed from those in the public RNA-seq datasets. These differences may reflect the use of independent biological samples, tissue composition, normalization procedures, and platform-specific measurements. Because enzymatic assays, metabolite measurements, cellular localization, and functional perturbation were not performed, the results identify candidate genes whose expression is associated with gonadal development rather than demonstrating regulatory roles. This study provides a comparative framework for future functional investigation of sterol- and steroid-related metabolism in bivalves.

Animals

A germline PDGFRB splice site variant associated with infantile myofibromatosis and resistance to imatinib.

PURPOSE: Infantile myofibromatosis is characterized by the development of myofibroblastic tumors in young children. In most cases, the disease is caused by somatic gain-of-function variants in platelet-derived growth factor (PDGF) receptor beta (PDGFRB). Here, we reported a novel germline intronic PDGFRB variant, c.2905-8G>A, in 6 unrelated infants with multifocal myofibromatosis and their relatives. METHODS: We performed constitutional and tumor DNA and RNA sequencing to identify novel variants, which were subsequently characterized in cellular assays. RESULTS: All patients had multiple skin nodules, 4 had bone lesions, and 2 had aggressive disease with bowel obstruction. The c.2905-8G>A substitution creates an alternative acceptor splice site in intron 21, inserting 2 codons in the PDGFRB transcript. Functional studies revealed that the splice change induced a partial loss of function, contrasting with previously described variants. In 4 tumor samples, we identified a second somatic hit at position Asp850 in PDGFRB exon 18, triggering constitutive receptor activation and resistance to imatinib. In addition to vinblastine and methotrexate, 2 patients received imatinib without objective response. One of them switched to dasatinib with concomitant improvement. CONCLUSION: This splice-site PDGFRB variant favors the development of myofibroma, featuring an acquired oncogenic variant in the same gene and resistance to targeted therapy.

Humans

Characterization of non-crossover recombination spectrum by single-microspore sequencing in maize and rice.

Meiotic DNA double-strand breaks (DSB) are crucial for chromosome recombination. The repair of DSB gives two outcomes: crossover (CO) and non-crossover (NCO). CO involves the bidirectional exchange between homologous chromosomes, whereas NCO refers to the unidirectional transfer of chromosome fragments. NCO can be categorized into NCO with gene conversion and NCO without gene conversion. Due to technological constraints, previous studies have focused more on CO than on NCO. In this study, we isolated single microspores from meiotic tetrads of maize (Zea mays) and rice (Oryza sativa) and conducted deep single-microspore genome sequencing to characterize NCO gene conversion (NCO-GC). Under highly stringent conditions, 101 CO and 902 NCO-GC tracts were identified in four maize tetrads, while 173 CO and 279 NCO-GC tracts were identified in six rice tetrads. In both maize and rice, NCO-GC was more prone to occur in the upstream and downstream of genes, as well as the introns. It also had a significant distribution in transposon regions. A common A-rich motif was enriched in the NCO-GC tracts of maize and rice. GC-biased gene conversion (gBGC) likely contributed to the bimodality of the GC content at the third codon position (GC3), and we discovered a significant proportional relationship between the number of DSBs and the GC content. These findings provide evidence that NCO-GC exhibits a distinct pattern compared with CO and may play an important role in gene and genome evolution.

Oryza