Search PubMedSearch

SEARCH · Search PubMed

Results for “gene features”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-&#x3b1;/NF-&#x3ba;B, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans

Identification of potential biomarkers and mechanisms for keloid disorder based on comprehensive bioinformatics analysis and machine learning algorithms.

BACKGROUND: Keloid disorder (KD) encompasses a spectrum of fibroproliferative dermal conditions, the pathogenesis remains complex and incompletely understood. This study sought to identify biomarkers and potential therapeutic targets for KD through an integrative bioinformatics approach and machine learning analysis of RNA sequencing data. METHODS: RNA sequencing was performed on skin tissue samples from 13 patients with KD and 14 healthy controls. Using weighted gene co-expression network analysis and differential expression analysis revealed differentially expressed key module genes, and the CytoHubba plugin identified candidate genes. Subsequently analyzed using least absolute shrinkage and selection operator (LASSO) and support vector machine recursive feature elimination (SVM-RFE) methods to pinpoint feature genes associated with KD. Following this, biomarkers were determined through expression level validation, enrichment analysis, and immune infiltration analysis. RESULTS: A total of 420 differentially expressed key module genes were identified, and the top 10 genes with DMNC values were selected as candidate genes. Five feature genes were selected through LASSO and SVM-RFE, with NID2, MFAP2, COL8A1, and P4HA3 showing significant expression differences between KD and control samples, along with consistent expression patterns across datasets, identified as potential biomarkers. These four biomarkers were proved to possess high diagnostic potential, and they were found to exhibit significant positive correlations with one another. Functional enrichment analysis indicated that the primary KEGG pathways associated with these biomarkers included "steroid hormone biosynthesis" and "cytokine-cytokine receptor interaction." Moreover, immune infiltration analysis revealed that the four biomarkers were negatively correlated with type 17 T helper cells and positively correlated with 15 immune cell types, including activated B cells and central memory CD4 T cells. CONCLUSION: In conclusion, NID2, MFAP2, COL8A1, and P4HA3 were identified as key biomarkers for KD, offering new avenues for more targeted and effective diagnostic and therapeutic strategies for managing this condition.

Humans

Reanalysis of BRCA1/2 negative high risk ovarian cancer patients reveals novel germline risk loci and insights into missing heritability.

While up to 25% of ovarian cancer (OVCA) cases are thought to be due to inherited factors, the majority of genetic risk remains unexplained. To address this gap, we sought to identify previously undescribed OVCA risk variants through the whole exome sequencing (WES) and candidate gene analysis of 48 women with ovarian cancer and selected for high risk of genetic inheritance, yet negative for any known pathogenic variants in either BRCA1 or BRCA2. In silico SNP analysis was employed to identify suspect variants followed by validation using Sanger DNA sequencing. We identified five pathogenic variants in our sample, four of which are in two genes featured on current multi-gene panels; (RAD51D, ATM). In addition, we found a pathogenic FANCM variant (R1931*) which has been recently implicated in familial breast cancer risk. Numerous rare and predicted to be damaging variants of unknown significance were detected in genes on current commercial testing panels, most prominently in ATM (n = 6) and PALB2 (n = 5). The BRCA2 variant p.K3326*, resulting in a 93 amino acid truncation, was overrepresented in our sample (odds ratio = 4.95, p = 0.01) and coexisted in the germline of these women with other deleterious variants, suggesting a possible role as a modifier of genetic penetrance. Furthermore, we detected loss of function variants in non-panel genes involved in OVCA relevant pathways; DNA repair and cell cycle control, including CHEK1, TP53I3, REC8, HMMR, RAD52, RAD1, POLK, POLQ, and MCM4. In summary, our study implicates novel risk loci as well as highlights the clinical utility for retesting BRCA1/2 negative OVCA patients by genomic sequencing and analysis of genes in relevant pathways.

Adult

LymphGen-Sig: Integrating Genetic and Transcriptional States to Predict Therapeutic Response in Diffuse Large B-Cell Lymphoma.

PURPOSE: Genetic classification may advance precision medicine in diffuse large B-cell lymphoma (DLBCL), but existing tools like LymphGen (LG) are limited by complexity and incomplete classification and do not incorporate nongenetic features that affect disease biology and therapeutic outcomes. To address these limitations, we developed LG-sig (LGsig), a gene expression-based platform that classifies all DLBCLs and harmonizes both genetic and nongenetic dimensions of the disease. METHODS: LGsig was built on the distinct subtype-specific gene expression signature of each LG class using paired genomic and transcriptomic data (National Cancer Institute/British Columbia Cancer Agency; N = 764). Model development was restricted to DLBCLs classified into MYD88L265P&#xa0;and&#xa0;CD79B&#xa0;mutations (MCD), BCL6&#xa0;translocation and&#xa0;NOTCH2&#xa0;mutations (BN2), EZH2&#xa0;mutations and&#xa0;BCL2&#xa0;translocation (EZB), or SGK1&#xa0;and&#xa0;TET2&#xa0;mutations (ST2). Gene features were selected by differential gene expression, with 294 genes being optimal for classification using a nearest shrunken centroid classifier. LGsig classifications were designated as MCDsig, BN2sig, ST2sig, and EZBsig. The final model was applied to RNAseq from archival samples from the POLARIX trial (N = 678) to assess outcomes after polatuzumab vedotin-R-CHP (pola-R-CHP) or rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) for each LGsig subtype. RESULTS: LGsig accurately identified LG subtypes using transcriptional data alone and extended assignments to all previously LG-unclassified cases. Importantly, LG-unclassified DLBCLs reassigned by LGsig mirrored the transcriptional and clinical features of their corresponding LG counterparts, supporting their reclassification. In addition, LGsig reassigned LG A53 DLBCLs, characterized by aneuploidy and TP53 alterations, into more biologically and therapeutically relevant LGsig clusters. Finally, LGsig improved the performance of LG as a biomarker in the POLARIX study, by identifying distinct DLBCL subtypes exhibiting a survival benefit with pola-R-CHP over R-CHOP in both LG-classified and LG-unclassified cases. CONCLUSION: LGsig expands molecular classification beyond current genetic classifiers in DLBCL by integrating both genetic and transcriptional dimensions of the disease to better inform subtype-specific therapeutic strategies.

Journal Article

Clinical features and ALDH5A1 gene findings in 13 Chinese cases with succinic semialdehyde dehydrogenase deficiency.

BACKGROUND AND AIMS: To investigate the clinical features, ALDH5A1 gene variations, treatment, and prognosis of patients with succinic semialdehyde dehydrogenase (SSADH) deficiency. MATERIALS AND METHODS: This retrospective study evaluated the findings in 13 Chinese patients with SSADH deficiency admitted to the Pediatric Department of Peking University First Hospital from September 2013 to September 2023. RESULTS: Thirteen patients (seven male and six female patients; two sibling sisters) had the symptoms aged from 1 month to 1 year. Their urine 4-hydroxybutyrate acid levels were elevated and were accompanied by mildly increased serum lactate levels. Brain magnetic resonance imaging (MRI) showed symmetric abnormal signals in both sides of the globus pallidus and other areas. All 13 patients had psychomotor retardation, with seven showing epileptic seizures. Among the 18 variants of the ALDH5A1 gene identified in these 13 patients, six were previously reported, while 12 were novel variants. Among the 12 novel variants, three (c.85_116del, c.206_222dup, c.762C&#x2009;>&#x2009;G) were pathogenic variants; five (c.427delA, c.515G&#x2009;>&#x2009;A, c.637C&#x2009;>&#x2009;T, c.755G&#x2009;>&#x2009;T, c.1274T&#x2009;>&#x2009;C) were likely pathogenic; and the remaining four (c.454G&#x2009;>&#x2009;C, c.479C&#x2009;>&#x2009;T, c.1480G&#x2009;>&#x2009;A, c.1501G&#x2009;>&#x2009;C) were variants of uncertain significance. The patients received drugs such as L-carnitine, vigabatrin, and taurine, along with symptomatic treatment. Their urine 4-hydroxybutyric acid levels showed variable degrees of reduction. CONCLUSIONS: A cohort of 13 cases with early-onset SSADH deficiency was analyzed. Onset of symptoms occurred from 1 month to 1 year of age. Twelve novel variants of the ALDH5A1 gene were identified.

Child, Preschool

Genome-wide identification of CHY zinc finger and RING finger (CHYR) genes in pepper and functional characterization of CaCHYR5 in response to Phytophthora capsici infection.

CHY zinc finger and RING finger (CHYR) proteins play crucial roles in the growth and development, as well as stress response. To date, no systematic or comprehensive analysis of the CHYR gene family has been performed in pepper (Capsicum annuum L.). In this study, we identified 8 CaCHYR genes (CaCHYR1-CaCHYR8), which were classified into 3 groups based on phylogenetic relationships. CaCHYR members within the same group exhibited similar distributions of conserved motifs and exon-intron structures. Chromosomal localization analysis showed that 8 CaCHYR genes were unevenly distributed on 6 chromosomes. Segmental duplication, rather than tandem duplication, was found to be the major contributor to the expansion of this gene family. CaCHYR genes feature a variety of cis-elements involved in developmental processes, phytohormone responses, and stress adaptation. Expression analysis based on RNA-seq data revealed that CaCHYR genes exhibited distinct spatial expression patterns across different tissues and in response to Phytophthora capsici infection (PCI), and quantitative real-time PCR (qRT-PCR) further confirmed that three of them (CaCHYR2, CaCHYR3, and CaCHYR5) exhibited altered expression under PCI. Furthermore, transient overexpression of CaCHYR5 in pepper leaves increased susceptibility to PCI, suggesting its potential negative regulatory role in pepper defense against P. capsici. Collectively, these findings reveal the expression patterns and regulatory functions of pepper CHYR genes in growth and development, laying a groundwork for breeding pepper cultivars tolerant to PCI.

Phytophthora capsici infection (PCI)

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans

Immune Cell-Stratified Regulatory Contexts Associated With BMI-Related Multi-System Disease Risk: A Cell-Stratified Mendelian Randomization Study Using Single-Cell eQTL Data.

AIMS: Body mass index (BMI) is associated with multisystem disease risk, but the immune cell-specific regulatory contexts underlying BMI-related genetic associations with disease outcomes remain unclear. METHODS: We applied a cell-stratified Mendelian randomization framework integrating European-ancestry BMI GWAS data, GWAS datasets for 33 disease outcomes across five disease systems, single-cell cis-eQTL data from 28 peripheral blood immune cell types, and dynamic CD4+ T cell eQTL data. SuSiE-based colocalization was used to identify BMI-associated loci sharing causal variants with immune-cell gene expression. These variants were used as cell-stratified instruments for Mendelian randomization. RESULTS: Across 28 immune cell types, 1326 colocalized variants regulating 1426 genes were identified. In primary MR analyses, genetically predicted BMI showed Bonferroni-significant associations with 26 disease outcomes. Cell-stratified analyses identified 87 Bonferroni-significant associations across 17 disease outcomes. Cardiovascular diseases showed the broadest cell-stratified associations, followed by respiratory and metabolic diseases. CD4+ T cell regulatory contexts contributed one of the largest shares of prioritized associations, and BMI-related effects varied across CD4+ T cell activation states. Cross-disease prioritization highlighted recurrent immune feature genes, including TRAF3 and FGFR1. CONCLUSION: These findings prioritize CD4+ T cell regulatory contexts as potential immunogenetic links between BMI and multi-system disease risk, while requiring further validation in diverse populations and mechanistic models.

Humans

The opioid receptor-ligand network in human cancers: pan-cancer multi-omics profiling and translational implications.

BACKGROUND: Opioid receptor-ligand signalling has been implicated in tumour biology and perioperative outcomes; however, its pan-cancer molecular landscape and clinical relevance remain incompletely defined. METHODS: We performed a pan-cancer multi-omics analysis of eight predefined opioid receptor-ligand genes across 33 tumour types from The Cancer Genome Atlas. Analyses included gene expression analysis using the linear models for microarray data (limma) package, genomic alterations, DNA methylation, regulatory network inference, pathway activity estimation using gene set variation analysis, and survival modelling. Multivariable Cox regression models were adjusted for age, sex, and tumour stage. RESULTS: Opioid receptor-ligand genes exhibited heterogeneous and generally low-to-moderate expression across tumour types. Genomic and epigenetic alterations were tumour-specific and variably associated with gene expression. Selected genes showed associations with overall survival in a tumour-dependent manner; however, these associations were attenuated after adjustment for clinical covariates and were accompanied by wide confidence intervals in some cohorts. Pathway analyses suggested associations with broader biological programmes, including epithelial-mesenchymal transition and immune-related pathways. Regulatory analyses identified candidate transcription factors and miRNAs, although these findings are exploratory. CONCLUSIONS: This pan-cancer analysis provides a systematic overview of opioid receptor-ligand gene features across human cancers. The observed associations are context-dependent and should be interpreted as hypothesis-generating. Further mechanistic and prospective studies are required to determine the clinical relevance of opioid signalling in cancer and perioperative settings.

Humans

YIPF&#x3b1;1A expression is regulated by multilayered molecular mechanisms.

Yip domain family (YIPF) proteins are five-pass transmembrane proteins that localize primarily to the Golgi apparatus. These proteins assemble into higher-order complexes with each &#x3b1;-subunit pairing specifically with a &#x3b2;-subunit to form a dimer which then assemble into complexes with two to four dimers. Notably, &#x3b2;-subunit expression depends on the corresponding &#x3b1;-subunit partner, and conventional transient overexpression of &#x3b1;-subunits has been extremely inefficient, hindering deeper analysis of YIPF complexes. To identify the cause of poor exogenous expression, we examined YIPF gene features and found two properties correlated with low expression: (i) rare-codon enrichment in the CDS and (ii) extended 3' UTRs. Experimental analyses focusing on YIPF&#x3b1;1A revealed that rare-codon enrichment suppresses expression mainly at the mRNA level, consistent with translation-coupled mRNA decay, whereas inclusion of the native 3'&#xa0;UTR enhances expression by increasing mRNA abundance. Deletion mapping further showed that a proximal 3' UTR segment (51-150) is necessary and sufficient for mRNA stabilization, thereby elevating both mRNA and protein levels. Conversely, a distal 3' UTR fragment (1116-2230) increased mRNA but not protein levels, suggesting translational repression resulting in a reduced protein-to-mRNA ratio. Together, these findings explain the discrepancy between endogenous and exogenous YIPF&#x3b1;1A expression and propose a multilayered regulatory model in which rare codons decrease mRNA, the proximal 3' UTR stabilizes mRNA, and the distal 3' UTR reduces translation. Impact statement Our work advances YIPF biology and identifies post&#x2011;transcriptional mechanisms governing multi&#x2011;pass membrane proteins. We show rare&#x2011;codon and 3' UTR&#x2011;based control of trafficking proteins-an area largely unexplored-and introduce a new paradigm for membrane&#x2011;traffic regulation that will guide future studies of complex assembly, localization, and homeostasis.

3' Untranslated Regions

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500&#xa0;m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed >&#x2009;99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071&#x1d40; (=&#x2009;ATCC 10145&#x1d40;), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33&#xa0;Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8&#xa0;kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~&#x2009;22&#xa0;kb, ~&#x2009;17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family

Clinical and genetic features of Ph-negative myeloproliferative neoplasms with dual-driver gene positivity.

OBJECTIVES: To investigate the clinical laboratory characteristics and gene mutation features of dual-driver gene positivity in patients with Philadelphia chromosome-negative myeloproliferative neoplasm (Ph-negative MPN). METHODS: We conducted a retrospective analysis of clinical data and genetic test results from 203 newly diagnosed patients with Ph-negative MPN. Of these, 194 had single-driver gene positivity and 9 had dual-driver gene positivity. High-throughput sequencing was used to detect mutations in JAK2, CALR, and MPL. Clinical characteristics and gene mutation profiles were compared between the two patient groups. RESULTS: The incidence of dual-driver gene positivity was 4.4% (9/203), with the most common combinations being JAK2 with CALR (4 patients) and JAK2 with MPL (4 patients). Compared with the single-driver group, the dual-driver group had a significantly higher risk of bleeding [4.1% (8/194) vs. 33.3% (3/9), P&#x2009;=&#x2009;0.008] and a higher proportion of uncommon mutations [3.6% (7/194) vs. 33.3% (3/9), P&#x2009;=&#x2009;0.006]. No statistically significant differences were observed between the two groups regarding age, thrombosis incidence, splenomegaly, or routine blood test indicators. During follow-up, 1 patient in the dual-driver group died from cerebrovascular disease. No leukaemia transformation or disease-related deaths occurred among the remaining patients. DISCUSSION: The increased bleeding risk in dual-driver patients may be related to a higher proportion of CALR mutations, elevated platelet counts, and higher variant allele frequencies, though these findings require validation in larger cohorts due to the small sample size. The higher prevalence of uncommon mutations suggests a more complex mutational landscape in this subgroup. CONCLUSION: Patients with Ph-negative MPN and dual-driver gene positivity may have a higher risk of bleeding and a more complex gene mutation profile.

Humans

Histone H3K27ac spreads from enriched chromatin domains into neighboring regions upon loss of CTCF binding.

Acetylation of histone H3 at lysine 27 (H3K27ac) is enriched at enhancers and highly transcribed genes. Our previous study showed that an H3K27ac-enriched chromatin domain expanded into neighboring regions following the deletion of CTCF-binding motifs flanking the domain. In this study, we explored the spreading of H3K27ac on a genome-wide scale by analyzing its distribution around CTCF-binding sites in human K562 cells and examining changes upon CTCF loss. We found that a subset of CTCF-binding sites demarcates H3K27ac-enriched domains. Upon loss of CTCF binding, H3K27ac levels increased in most regions adjacent to these domains, indicating that H3K27ac can spread into neighboring chromatin. This spreading was accompanied by elevated transcription of nearby genes. Chromatin features, including histone modifications, CTCF-binding intensity, and CTCF-mediated chromatin interactions, were associated with the H3K27ac spreading. Notably, enhancers were more enriched within domains that exhibited H3K27ac spreading compared to those that did not, and the deletion of enhancers from the CTCF motif-deficient &#x3b2;-globin locus attenuated the spreading. These findings indicate that CTCF-binding sites serve as boundaries for H3K27ac-enriched domains and that, in the absence of CTCF binding, H3K27ac can spread into neighboring regions. H3K27ac spreading appears to be influenced by multiple chromatin features and to contribute to the transcriptional increase of nearby genes.

CTCF

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite.

Nucleotide recoding RNA sequencing methods (NR-seq; TimeLapse-seq, SLAM-seq, TUC-seq, etc.) are powerful approaches for assaying transcript population dynamics. In addition, these methods have been extended to probe a host of regulated steps in the RNA life cycle. Current bioinformatic tools significantly constrain analyses of NR-seq data. To address this limitation, we developed EZbakR (https://github.com/isaacvock/EZbakR), an R package to facilitate a more comprehensive set of NR-seq analyses, and fastq2EZbakR (https://github.com/isaacvock/fastq2EZbakR), a Snakemake pipeline for flexible preprocessing of NR-seq datasets, collectively referred to as the EZbakR suite. Together, these tools generalize many aspects of the NR-seq analysis workflow. The fastq2EZbakR pipeline can assign reads to a diverse set of genomic features (e.g., genes, exons, splice junctions), and EZbakR can perform analyses on any combination of these features. EZbakR extends standard NR-seq mutational modeling to support multi-label analyses (e.g., s4U and s6G dual labeling), and implements an improved hierarchical model to better account for transcript-to-transcript variance in metabolic label incorporation. EZbakR also generalizes dynamical systems modeling of NR-seq data to support analyses of premature mRNA processing and flow between subcellular compartments. Finally, EZbakR implements flexible and well-powered comparative analyses of all estimated parameters via design matrix-specified generalized linear modeling. The EZbakR suite will thus allow researchers to make full, effective use of NR-seq data.

Software

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

TCRspec: A Recognition Interface-Informed Multimodal Method for TCR-pMHC Specificity Prediction.

Specific recognition between T-cell receptors (TCRs) and peptide-major histocompatibility complexes (pMHCs) is central to adaptive immunity, yet accurate prediction of TCR-pMHC specificity remains challenging. Existing models mainly rely on sequence features or isolated molecular structures, limiting their ability to capture interface-level determinants within the ternary recognition complex. Here, we constructed the multimodal TCR-pMHC ternary complex (MM-TCR) data set, integrating paired TCR-pMHC sequences, V/J gene annotations, and modeled TCR-pMHC complex structures refined by short molecular dynamics-based relaxation. Based on MM-TCR, we developed TCRspec, an interpretable multimodal framework combining sequence embeddings, gene-usage features, and complex-level structural representations. Under a stringent CD-HIT TCR-cluster-disjoint split, TCRspec achieved an average AUROC of 0.896 and AUPRC of 0.882 across seven antigen-specific test data sets, outperforming representative baseline models. Cross-validation and ablation analyses confirmed the contribution of ternary complex structural information and MD-refined structures. In independent OOD peptide-TCR systems, TCRspec retained discriminative performance and identified model-inferred peptide positions associated with TCR recognition, providing a structure-informed framework for TCR specificity prediction.

Receptors, Antigen, T-Cell