Search PubMedSearch

SEARCH · Search PubMed

Results for “spatial transcriptomics data”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

jsPCA: fast, scalable, and interpretable identification of spatial domains and variable genes across multi-slice and multi-sample spatial transcriptomics data.

MOTIVATION: Spatial transcriptomics technologies record genome-wide measurements of gene expression with high spatial resolution. These technologies generate large and high-dimensional datasets requiring efficient automated methods for their analysis. We introduce joint spatial PCA (jsPCA), a novel, fast, scalable and interpretable method for the automatic identification of spatial domains and variable genes in multi-slice and multi-sample spatial transcriptomics data. RESULTS: jsPCA relies on a simple mathematical formulation of a spatial covariance defined as the product of the gene expression covariance with the spatial autocorrelation. The principal components of this spatial covariance yield a biologically meaningful low-dimensional representation. From this representation, spatial domains are derived by simple clustering and spatially variable genes are identified directly from the principal component coefficients. A joint representation of multiple slices and samples without spatial alignment is obtained by computing common principal components via joint diagonalization. By leveraging data sparsity and non-convex manifold optimization, jsPCA leads to computing time in the order of seconds to minutes, substantially outperforming state-of-the-art approaches. We benchmarked jsPCA against 10 state-of-the-art methods on two reference databases. Our approach demonstrated excellent performance, comparable or better than state-of-the-art methods, while being much faster, interpretable, and scalable to very large datasets.

Journal Article

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

Representation learning for multi-modal spatially resolved transcriptomics data.

MOTIVATION: Spatial transcriptomics enables in-depth molecular characterization of samples on a morphology and RNA level while preserving spatial location. Integrating the resulting multi-modal data is an unsolved problem, and developing new solutions in precision medicine depends on improved methodologies. RESULTS: We introduce AESTETIK, a convolutional deep learning model that jointly integrates spatial, transcriptomics, and morphology information to learn accurate spot representations. AESTETIK yielded substantially improved cluster assignments on widely adopted technology platforms (e.g. 10x Genomics™, NanoString™) across multiple datasets. We achieved performance enhancement on structured tissues (e.g. brain) with a 21% increase in median ARI over previous state-of-the-art methods. Notably, AESTETIK also demonstrated superior performance on cancer tissues with heterogeneous cell populations, showing a 2-fold increase in breast cancer, 79% in melanoma, and 21% in liver cancer. We expect that these advances will enable a multi-modal understanding of key biological processes. AVAILABILITY AND IMPLEMENTATION: AESTETIK is implemented in Python 3 and is available as open source software at http://www.github.com/ratschlab/aestetik. The Snakemake pipeline for reproducing the results is available at http://www.github.com/ratschlab/st-rep.

Spatial Transcriptomics

Spatial mutual nearest neighbors for spatial transcriptomics data.

MOTIVATION: Mutual nearest neighbors (MNN) is a widely used computational tool to perform batch correction for single-cell RNA-sequencing data. However, in applications such as spatial transcriptomics, it fails to take into account the 2D spatial information. RESULTS: Here, we present spatialMNN, an algorithm that integrates multiple spatial transcriptomic samples and identifies spatial domains. Our approach begins by building a k-nearest neighbors (kNN) graph based on the spatial coordinates, prunes noisy edges, and identifies niches to act as anchor points for each sample. Next, we construct a MNN graph across the samples to identify similar niches. Finally, the spatialMNN graph can be partitioned using existing algorithms, such as the Louvain algorithm to predict spatial domains across the tissue samples. We demonstrate the performance of spatialMNN using large datasets, including one with N = 31 10x Genomics Visium samples. We also evaluate the computing performance of spatialMNN to other popular spatial clustering methods. AVAILABILITY AND IMPLEMENTATION: Our software package is available on GitHub (https://github.com/Pixel-Dream/spatialMNN). The code is available on Zenodo (https://doi.org/10.5281/zenodo.15073963).

Algorithms

SpaceBar enables clone tracing in spatial transcriptomic data.

We report a cellular barcoding strategy, SpaceBar, that enables simultaneous clone tracing and spatial transcriptomics profiling. Our approach uses a library of 96 synthetic barcode sequences that can be robustly detected by imaging based spatial transcriptomics (seqFISH), delivered such that each cell is labeled with a combination of barcodes. We used these barcodes to label melanoma cells in a tumor xenograft model and profiled both clone identity and spatial gene expression in situ. We developed a gene scoring metric that quantifies how strongly gene expression is driven by intrinsic cellular cues or extrinsic environmental signals. Our framework distinguishes between clonal dynamics and environmentally-driven transcriptional regulation in complex tissue contexts.

Journal Article

DPAS-Graph: adaptive spatial-feature relation learning for spatial RNA-to-protein prediction and virtual protein profiling.

Paired spatial multi-omics provides a supervised basis for learning RNA-protein correspondence in situ, but predicting protein abundance from spatial transcriptomic data alone remains challenging across tissue contexts and protein panels. Here, we present DPAS-Graph, an adaptive relation-learning framework for spatial RNA-to-protein prediction. Rather than directly merging spatial proximity and transcriptomic similarity as fixed graph priors, DPAS-Graph represents them as two relation channels on a shared edge support and updates their contributions during representation learning for protein prediction. Its Niche-Coupled Field Encoder combines layer-wise edge-relation modeling, intra-branch relation refinement, and cross-branch residual correction to learn spot representations for protein abundance prediction. In a leave-one-dataset-out benchmark across seven paired spatial multi-omics datasets, DPAS-Graph achieved lower aggregate prediction errors and improved spot-level agreement of protein expression profiles, with gains mainly reflected in error-based metrics and PCC-Spot. Spatial autocorrelation and protein-derived domain agreement analyses were further used to characterize the spatial behavior of the predicted protein maps. When applied to external RNA-only spatial sections, DPAS-Graph generated qualitatively interpretable marker-level virtual protein maps, illustrating its use as a complementary tool for protein-level interpretation of transcriptomics-only spatial data.

RNA

Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model.

MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.

Algorithms

Learning directed acyclic graphs for ligands and receptors based on spatially resolved transcriptomic data of ovarian cancer.

To unravel the mechanism of immune activation and suppression within tumors, a critical step is to identify transcriptional signals governing cell-cell communication between tumor and immune/stromal cells in the tumor microenvironment. Central to this communication are interactions between secreted ligands and cell-surface receptors, creating a highly connected signaling network among cells. Recent advancements in in situ-omics profiling, particularly spatial transcriptomic (ST) technology, provide unique opportunities to directly characterize ligand-receptor signaling networks that power cell-cell communication. In this paper, we propose a novel statistical method, LRnetST, to characterize the ligand-receptor interaction networks between adjacent tumor and immune/stroma cells based on ST data. LRnetST utilizes a directed acyclic graph model with a novel approach to handle the zero-inflated distributions of ST data. It also leverages existing ligand-receptor regulation databases as prior information, and employs a bootstrap aggregation strategy to achieve robust network estimation. Application of LRnetST to ST data of high-grade serous ovarian tumor samples revealed both common and distinct ligand-receptor regulations across different tumors. Some of these interactions were validated through both a MERFISH dataset and a CosMx SMI dataset of independent ovarian tumor samples. These results cast light on biological processes relating to the communication between tumor and immune/stromal cells in ovarian tumors. An open-source R package of LRnetST is available on GitHub at https://github.com/jie108/LRnetST.

Humans

A contextual activity score (CAS) for inferring ADAR-associated transcriptional activity across RNA-seq, single-cell, and spatial transcriptomics.

BACKGROUND AND OBJECTIVE: Adenosine-to-inosine RNA editing, catalyzed by Adenosine Deaminases Acting on RNA (ADARs), is a widespread modification involved in neural function, immune regulation, and cancer. The Alu Editing Index (AEI) is the standard metric to estimate ADAR activity but requires raw sequencing reads and is poorly suited for single-cell and spatial transcriptomic data. This study aimed to develop an alternative framework for inferring ADAR-associated transcriptional activity from gene expression data across diverse transcriptomic technologies. METHODS: We developed the Contextual Activity Score (CAS), a framework based on transcriptional signatures from ADAR perturbation experiments. Context-specific signatures were generated for human neurons, mouse neurons, and cancer models to infer ADAR1 and ADAR2 activity. CAS was computed from normalized gene expression matrices using regulon-based enrichment analysis. Performance was evaluated by comparing with the Alu Editing Index across bulk RNA sequencing datasets, simulated sequencing depths, and library preparation protocols. RESULTS: CAS showed strong concordance with the Alu Editing Index across multiple datasets, while remaining robust to reduced sequencing depth and different library protocols. Unlike the Alu Editing Index, CAS can be applied to single-cell and spatial transcriptomic data and enables the independent assessment of ADAR2 activity. In cancer and neuronal contexts, CAS captured biologically meaningful variations in ADAR-associated transcriptional activity at sample, cell-type, and spatial levels. CONCLUSION: CAS provides a scalable approach applicable across multiple RNA-seq protocols for estimating ADAR-associated transcriptional activity using gene expression data. This method, implemented in an open-source R package for broad adoption, expands the ability to study ADAR-associated transcriptional activity across transcriptomic modalities where direct editing quantification is challenging, such as single-cell and spatial transcriptomics.

Adenosine Deaminase

Integrated histopathology, spatial and single cell transcriptomics resolve cellular drivers of early and late alveolar damage in COVID-19.

The most common cause of death due to COVID-19 remains respiratory failure. Yet, our understanding of the precise cellular and molecular changes underlying lung alveolar damage is limited. Here, we integrate single cell transcriptomic data of COVID-19 and donor lung tissue with spatial transcriptomic data stratifying histopathological stages of diffuse alveolar damage. We identify changes in cellular composition across progressive damage, including waves of molecularly distinct macrophages and depletion of epithelial and endothelial populations. Predicted markers of pathological states identify immunoregulatory signatures, including IFN-alpha and metallothionein signatures in early damage, and fibrosis-related collagens in late damage. Furthermore, we predict a fibrinolytic shutdown via endothelial upregulation of SERPINE1/PAI-1. Cell-cell interaction analysis revealed macrophage-derived SPP1/osteopontin signalling as a key regulator during early steps of alveolar damage. These results provide a comprehensive, spatially resolved atlas of alveolar damage progression in COVID-19, highlighting the cellular mechanisms underlying pro-inflammatory and pro-fibrotic pathways in severe disease.

COVID-19

SCMO: a deep learning model integrating the single-cell resolution TME ecosystem and multi-omics for survival prediction in CRC patients.

BACKGROUND: Colorectal cancer (CRC) remains a leading cause of global cancer mortality, highlighting the need for precise survival prediction to guide clinical decisions. Although tissue-level multi-omics is widely utilized for survival prediction, its limited resolution cannot capture tumor heterogeneity. Single-cell RNA sequencing (scRNA-seq) enables dissection of the tumor microenvironment (TME) at cellular resolution, supporting personalized prognostic assessment. METHODS: We collected 213 CRC scRNA-seq samples and established a CRC-specific TME atlas comprising 339,060 cells. Using this atlas as a reference, we deconvolved bulk RNA-seq data from TCGA-CRC cohort with the EcoTyper algorithm to reconstruct TME features. Clinical, genomic, and transcriptomic data were obtained from the Xena platform; microbial data were sourced from the BIC database. We integrated TME and multi-omics features through a self-normalizing neural network to construct a deep learning model (single-cell resolution TME ecosystem with multi-omics data [SCMO]) for survival prediction. To enhance interpretability, we utilized the Integrated Gradients algorithm and spatial transcriptomic data to analyze multi-omics and TME features. We performed anticancer drug screening with tumor necrosis factor receptor-associated protein 1 (TRAP1), a critical feature according to the Integrated Gradients algorithm, as a potential target. RESULTS: We identified 13 survival-related TME features from the CRC-specific atlas: 12 cell states and one multi-cellular ecosystem. SCMO, which combined TME and multi-omics features, improved survival prediction and outperformed existing methods, achieving a concordance index of 0.762. The SCMO demonstrated robust performance for long-term predictions, achieving areas under the curve (AUCs) of 0.752, 0.772, and 0.869 for 1-, 3-, and 5-year predictions in the training set, with corresponding test set AUCs of 0.639, 0.756, and 0.772. TME features from the SCMO model revealed that ecosystem density increased with CRC malignancy. Multi-omics features included TRAP1 as a potential drug target. Drug screening identified saikosaponin A as a novel TRAP1 inhibitor, and its anticancer activity was validated in vitro. We developed SCMO-Lite, a simplified model incorporating 12 high-attribution-weight multi-omics features, which demonstrated robust risk stratification. CONCLUSIONS: SCMO combines analytical precision with biological interpretability, offering novel insights for oncology survival prediction.

Humans

Enhancing pan-cancer spatial transcriptomics at single-cell resolution with stPainter.

Subcellular spatial transcriptomics can resolve tissue architecture at cellular scale, but sparse gene panels and limited detection sensitivity constrain downstream analysis. Existing enhancement methods often require tissue-matched single-cell RNA sequencing (scRNA-seq) references and dataset-specific retraining. Here we show that stPainter, a conditional generative model pretrained on a pan-cancer scRNA-seq atlas, can enhance spatial transcriptomics data without matched references or retraining. Using a latent diffusion architecture guided by Stochastic Differential Equations (SDE), stPainter reconstructs expanded expression profiles from sparse measurements and produces latent representations for clustering and cell-state analysis. When we apply stPainter upon 6 spatial transcriptomics datasets of different cancer types, we demonstrate that our model empowers downstream biological analyses, including fine-grained subpopulation clustering and pathway enrichment. Comparison with spatially resolved proteomics (CODEX) provided independent support for regional agreement between imputed cellular compositions and protein-level tissue organization. These results establish stPainter as a scalable approach for analyzing tumor microenvironments without auxiliary sequencing data.

Spatial Transcriptomics

CAGNet: a structure-aware clustering-alternated graph network for cell-cell interaction inference in spatial transcriptomics.

MOTIVATION: Understanding cell-cell interactions (CCIs) in spatial transcriptomics is crucial for uncovering the spatial organization and functional heterogeneity of tissues. However, existing graph-based models typically rely on static clustering or fixed adjacency structures, which limits their ability to capture dynamic cellular relationships. RESULTS: We propose CAGNet, a two-stage framework for CCI inference from spatial transcriptomics data. In Stage 1, a Graph Attention Network encoder with joint feature and graph reconstruction learns structure-aware node embeddings from spatial gene expression profiles. In Stage 2, an alternating optimization mechanism iteratively updates cluster centers via KL-guided soft assignment and refines node embeddings through spatial graph reconstruction, establishing a closed-loop between representation learning and clustering. Experiments on three 10x Genomics Visium datasets demonstrate that CAGNet consistently outperforms six CCI inference baselines across ACC, AUC, AP, Precision, Recall, and F1. CAGNet also achieves the highest Adjusted Rand Index on all three datasets against six spatial domain identification methods, confirming that the learned embeddings capture biologically relevant spatial organization. Information-theoretic analysis further shows that CAGNet retains the highest mutual information between input features and learned embeddings among all compared methods. Ablation studies and 5-fold cross-validation confirm the contribution of each component and the reproducibility of the results. AVAILABILITY: The proposed method is implemented in the CAGNet package available at http://github.com/mahan1233333-maker/CAGNet .

Spatial Transcriptomics

Malignant epithelial states drive immune dysfunction in ampulla of Vater carcinoma.

BACKGROUND: Ampulla of Vater (AoV) carcinoma is a rare malignancy arising at the junction of intestinal and pancreatobiliary epithelium. Its heterogeneous clinical behavior and histological diversity have hindered therapeutic advances, and the cellular basis of this heterogeneity remains unclear. We aimed to construct a single-cell transcriptomic atlas of AoV carcinoma, with a focus on identifying epithelial subtypes and their interactions with the tumor microenvironment (TME). METHODS: We performed single-cell RNA sequencing on eight primary AoV tumors and four matched normal tissues. Comprehensive clustering and transcriptomic analyses identified cell-type composition, epithelial heterogeneity, and tumor-immune interactions. Findings were validated using deconvolution of bulk RNA-seq data from 62 AoV carcinoma patients. Results Malignant epithelial cells were categorized into four distinct subtypes: Int-Wnt, PB-KRAS, Int-Hypoxia, and Cycling stage. PB-KRAS cells exhibited stem-like transcriptional programs and high genomic instability. Deconvolution analysis of bulk RNA-seq data from the independent AoV cohort revealed that enrichment of the PB-KRAS subtype correlated with tumor recurrence and poor survival. Our immune profiling analysis discovered a significant association between PB-KRAS subtype and GZMK+ CD8+ T cells, which are in a pre-dysfunctional state, alongside SPP1+ macrophages exhibiting immunosuppressive traits. Spatial transcriptome data further supports the immunosuppressive natures of TME around PB-KRAS subtype malignant epithelial cells in AoV carcinoma. CONCLUSIONS: Our study presents a single-cell atlas of AoV carcinoma, highlighting the molecular diversity of malignant epithelium and its association with the immune microenvironment. The PB-KRAS subtype emerges as a stem-like, immunosuppressive tumor state associated with poor prognosis, providing insights for future therapeutic targeting.

Ampulla of Vater carcinoma

Integrated single-cell and spatial transcriptomic analyses reveal malignant epithelial glycolytic heterogeneity and spatial niche remodeling during colorectal cancer progression.

Colorectal cancer (CRC) progression is shaped by metabolic reprogramming and complex interactions within the tumor microenvironment. However, the cellular heterogeneity, spatial organization, and clinical relevance of glycolytic activity in CRC remain incompletely understood. In this study, we integrated single-cell RNA sequencing, bulk transcriptomics, and spatial transcriptomics data to systematically characterize glycolytic heterogeneity in CRC. Glycolytic activity was quantified using five independent scoring methods, consistently showing that epithelial cells exhibited the highest glycolytic activity across the two single-cell cohorts. Stratification of CopyKAT-verified aneuploid malignant epithelial cells into high-glycolysis (HG) and low-glycolysis (LG) subgroups by glycolysis scores revealed that HG cells exhibited higher stemness scores and chromosomal copy number variations. Cell-cell communication analysis revealed that, compared with LG cells, HG cells exhibited increased interaction frequency and strength with immune and stromal populations, indicating enhanced malignant epithelial-microenvironment crosstalk. Spatial transcriptomics analyses further revealed that glycolytic activity varied across normal colorectal tissue, primary CRC, and colorectal liver metastases, accompanied by progressive remodeling of epithelial-associated spatial niches and MIF-mediated intercellular communication. Bulk transcriptomic analysis identified a glycolysis-related prognostic signature with robust predictive performance, which served as an independent prognostic factor for overall survival in CRC cohorts. Collectively, these findings indicate that glycolytic heterogeneity is a key feature of CRC malignant epithelial cells and is closely associated with tumor progression, microenvironmental remodeling, and clinical outcomes.

Humans

Deep FLASH-seq profiling of purified canine sensory neurons uncovers species-specific signatures relevant to pain and itch.

Naturally occurring pain and itch disorders in the domestic dog represent an important and underexploited opportunity for translational sensory neuroscience. These conditions largely mirror human disease, highlighting the need for detailed comparative understanding of canine somatosensory neurobiology. Here, we present a single-cell transcriptomic characterisation of the canine dorsal root ganglion (DRG), providing molecular insights into sensory neuron diversity in a species of direct veterinary and biomedical relevance. We develop a novel mechanical dissociation and fluorescence-activated cell sorting strategy enabling purification of intact whole neurons from adult canine DRG, followed by deep, full-length RNA sequencing using FLASH-seq. This approach yields high-quality transcriptional profiles with molecular depth analogous to deep neuronal profiling in human DRG, enabling resolution of neuronal identities and subtype-specific gene programs. Using these data, we identify canine sensory neuron clusters conforming to conserved principles of DRG molecular organization observed across species, including peptidergic and noncanonical peptidergic nociceptors, low-threshold mechanoreceptors, proprioceptors, and thermosensory populations. Cross-species comparisons with human and mouse DRG datasets reveal broad conservation of pain- and itch-relevant pathways and therapeutic targets, alongside biologically meaningful divergence. We further identify species-specific differences in subtype-restricted expression of the pharmacologically relevant receptors IL31RA and SSTR2 , which we validate using in situ hybridization and contextualize with human spatial transcriptomic data. Finally, we provide evidence that domestication-associated genes are nonrandomly enriched in specific sensory neurons, suggesting that evolutionary history may have shaped somatosensory function. These data represent a resource for comparative sensory neuroscience and inform translational interpretation of pain and itch therapeutics across species.

Animals

Dissecting Sex-Specific Pathology in K18-hACE2 Transgenic Mice Infected With Different SARS-CoV-2 Variants.

Sex-biased differences in COVID-19 outcomes in relation to individual SARS-CoV-2 variants are not well understood. In this study, lungs and nasal cavities of age-matched female and male K18-hACE2 transgenic mice were collected for dissecting sex-specific differences in pathology after infection of SARS-CoV-2 614 G, Delta, or Omicron variant. Overall, Delta infection induced the most severe inflammation and pathology in nasal cavity and lung followed by the 614 G, then Omicron variant. Sex differences in host responses to SARS-CoV-2 infection were variant-specific. Delta-infected males showed increased pulmonary infiltration of CD163+ "M2" macrophages, Ly6G+ neutrophils, and NKR-P1C + NK cells during early onset of infection, and elevated lung inflammatory cytokines such as IL-10, IL-6, and IP-10 than Delta-infected females. Conversely, females had increased lung CD4 + T cell recruitment after Omicron infection and significantly elevated lung MCP-1 secretion after Delta infection than males. Lung spatial transcriptomics data revealed that Delta-infected females had enriched gene pathways related to humoral immune response and interferon signaling, while males had enriched pathways associated with extracellular matrix production, chemokine signaling, and cell chemotaxis. Taken together, this study highlights the complex infection dynamics with respect to individual SARS-CoV-2 variants and underscores the importance of sex as a confounding factor for COVID-19 pathology.

Animals

Dietary Polyphenol Acteoside-Related Molecular Signatures in Clear Cell Renal Cell Carcinoma: Multi-Omics Profiling and Functional Validation of IMPDH1.

Clear cell renal cell carcinoma (ccRCC) is characterized by substantial metabolic and molecular heterogeneity, but the disease-relevant programs associated with acteoside, a dietary polyphenol, remain poorly understood. We integrated predicted acteoside targets with bulk, single-cell, and spatial transcriptomic data from ccRCC and combined molecular subtyping with cross-cohort machine-learning analysis. Acteoside-related signatures were preferentially enriched in malignant compartments and increased with tumor grade and stage. Consensus clustering identified two molecular subtypes with distinct biological and clinical features. C1 was associated with immune activation, metabolic activity, and more favorable survival, whereas C2 showed greater genomic instability, reduced renal epithelial differentiation, and poorer outcomes. We further benchmarked multiple machine-learning strategies and established a 10-gene prognostic model that retained predictive performance across independent cohorts, with IMPDH1 emerging as the strongest risk-associated feature. Functional experiments confirmed the biological relevance of IMPDH1: its knockdown suppressed ccRCC cell proliferation, DNA synthesis, colony formation, and migration, whereas overexpression produced the opposite effects. Together, these findings indicate that acteoside-related molecular signatures capture clinically relevant heterogeneity in ccRCC and provide a framework for linking dietary-polyphenol-related molecular space with tumor biology. The identification and functional validation of IMPDH1 further highlight its potential importance in ccRCC progression.

IMPDH1