Search PubMedSearch

SEARCH · Search PubMed

Results for “deconvolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

OmicsTweezer: A distribution-independent cell deconvolution model for multi-omics Data.

Cell deconvolution estimates cell type proportions from bulk omics data, enabling insights into tissue microenvironments and disease. However, practical applications are often hindered by batch effects between bulk data and referenced single-cell data, a challenge that is frequently overlooked. To address this discrepancy, we developed OmicsTweezer, a distribution-independent cell deconvolution model. By integrating optimal transport with deep learning, OmicsTweezer aligns simulated and real data in a shared latent space, effectively mitigating data shifts and inter-omics distribution differences. OmicsTweezer is versatile, capable of deconvolving bulk RNA-seq, bulk proteomics, and spatial transcriptomics. Extensive evaluations on simulated and real-world datasets demonstrate its robustness and accuracy. Furthermore, applications in prostate and colon cancer showcase OmicsTweezer's ability to identify biologically meaningful cell types. As a unified deconvolution framework for multi-omics data, OmicsTweezer offers an efficient and powerful tool for studying disease microenvironments.

Humans

Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model.

MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.

Algorithms

Continuous DNA Methylation Deconvolution-Based Surrogate for B-Cell Differentiation State in CLL.

Chronic Lymphocytic Leukemia (CLL) is clinically divided into IGHV mutated (M-CLL) and IGHV unmutated (U-CLL) subtypes, which are thought to arise from distinct cells of origin along the B-cell differentiation pathway. We measured genome-scale DNA methylation in purified CLL samples ( n = 89) and utilized reference-based cell deconvolution techniques to develop a continuous metric of epigenetic similarity across a B-naive-like to B-memory-like scale (B-Index). B-Index accurately classifies CLL into clinical subtypes (98.8%), has a stronger epigenetic signal than IGHV gene percent identity, and demonstrates additional epigenetic signal within the M-CLL subgroup. We demonstrate that U-CLL is epigenetically more similar to B-memory than B-naive cells and reconcile previous reports of a B-naive-like epigenetic signal. The B-memory-like program of U-CLL is enriched for binding sites of transcription factors related to the germinal center activation pathway. Our findings provide epigenetic evidence for discerning CLL mechanisms of initiation and cell of origin. We also identified an epigenetic signal associated with tumor burden, which may have some relation to viral infections such as Epstein-Barr-Virus. Our cell-type deconvolution-based approach to developing a continuous metric for CLL epigenetic differentiation state can be applied to other tumors with multiple subtypes across differentiation stages.

B-memory-like

Cell cycle-dependent protein dynamics in budding yeast resolved by deconvolution of bulk proteomics.

The cell division cycle is characterised by oscillatory dynamics in regulatory mechanisms and biosynthesis, coordinated with genome replication and segregation. To understand these dynamics, quantitative cell cycle-dependent protein concentration data are essential. Unfortunately, accurately resolving cell cycle-dependent protein dynamics is challenging because single-cell proteomics is currently infeasible and bulk proteomics requires - inherently imperfect - cell synchronisation. Here, we developed a computational method to deconvolve cell cycle-dependent protein concentration dynamics and applied it to new budding yeast bulk proteome data. Key to this method was a yeast population model, parameterised with experimental cell cycle progression and volume growth data, for quantifying the desynchronisation in sampled populations. We performed deconvolution on 3272 proteins, using cross-validation to determine regularisation parameters, and identified 539 proteins with cell cycle-dependent dynamics. Many of these dynamics were consistent with known yeast biology and dynamic proteins were enriched for several metabolic process, extending previous observations and supporting the emerging picture of metabolic activity as varying substantially over cell cycle phases. We consider the generated cell cycle-resolved budding yeast proteome data a key resource.

Journal Article

Sparse deconvolution of cell type medleys in spatial transcriptomics.

Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling However, current deconvolution models aim to estimate the proportions of distinct cell types in each spatial transcriptomics spot by integrating reference single-cell data. These models often assume strong overlap between the reference and spatial datasets, neglecting biology-grounded constraints such as sparsity and cell-type variations, as well as technical sparsity. As a result, these methods rely on over-permissive algorithms that ignore given constraints leading to inaccurate predictions, particularly in heterogeneous or unmatched datasets. We introduce Weight-Induced Sparse Regression (WISpR), a machine learning algorithm that integrates spot-specific hyperparameters and sparsity-driven modeling. Unlike conventional approaches that neglect biology-grounded constraints, WISpR accurately predicts cell-type distributions while preserving biological coherence, i.e., spatially and functionally consistent cell-type localization, even in unmatched datasets. Benchmarking against five alternative methods across ten datasets, WISpR consistently outperformed competitors and predicted cellular landscapes in both normal and cancerous tissues. By leveraging sparse cell-type arrangements, WISpR provides biologically informed, high-resolution cellular maps. Its ability to decode tissue organization in both healthy and diseased states highlights WISpR's practical utility for spatial transcriptomics, particularly in challenging settings involving noise, sparsity, or reference mismatches.

Humans

Palaeoproteomic Deconvolution of Physical and Genetic Collagen Mixtures.

Species identification in palaeoproteomics relies on genome-derived protein sequences which are often poor-quality, and lacks tools to cope with multi-species samples. Here, we address both challenges through the analysis of "physical and genetic mixtures". Species that are absent from our database are considered a "genetic mixture", i.e. a patchwork of peptides from closely related species. Inversely, various overlapping peptide stretches allow us to resolve complex "physical mixtures". This is benchmarked by analysing physical mixtures of modern bone fragments, including genetic mixtures. We illustrate the impact of our approach via a rapid and high-throughput analysis of >2500 bone fragments, revealing the Eemian-era faunal environment around Scladina Cave, including the first Palaeoloxodon antiquus identified at this site.

bioarchaeology

Deconvoluting clonal and cellular architecture in IDH-mutant acute myeloid leukemia.

Isocitrate dehydrogenase 1/2 (IDH) mutations are early initiating events in acute myeloid leukemia (AML). The complex clonal architecture and cellular heterogeneity in IDH-mutant AML underlies the heterogeneous clinical presentation and outcomes. Integrating single-cell genotyping and transcriptomics, we demonstrate a stem-like and inflammatory phenotype of IDH-mutant AML and identify clone-specific programs associated with NPM1, NRAS, and SRSF2 co-mutations. Furthermore, these clones had distinct responses to treatment with combination IDH inhibitors and chemotherapy, including elimination, reconstitution of myeloid differentiation, or retention within progenitor populations. At relapse after IDH inhibitor monotherapy, we identify upregulated stemness, inflammation, mitochondrial metabolism, and anti-apoptotic factors, as well as downregulated major histocompatibility complex (MHC) class II antigen presentation. At the pre-leukemic stage, we observe upregulation of IDH2-associated pathways, including inflammation. We deliver a detailed phenotyping of IDH-mutant AML and a framework for dissecting contributions of recurrently mutated genes in AML at diagnosis and following therapy, with implications for precision medicine.

Leukemia, Myeloid, Acute

StrainR2 accurately deconvolutes strain-level abundances in synthetic microbial communities.

MOTIVATION: Synthetic microbial communities offer an opportunity to conduct reductionist research in tractable model systems. However, deriving abundances of highly related strains within these communities is currently unreliable. 16S rRNA gene sequencing does not resolve abundance at the strain level and other methods such as quantitative polymerase chain reaction (qPCR) scale poorly and are resource prohibitive for complex communities. We present StrainR2, which utilizes shotgun metagenomic sequencing to provide high accuracy strain-level abundances for all members of a synthetic community, provided their genomes. RESULTS: Both in silico, and using sequencing data derived from gnotobiotic mice colonized with a synthetic fecal microbiota, StrainR2 resolves strain abundances with greater accuracy and efficiency than other tools utilizing shotgun metagenomic sequencing reads. We demonstrate that StrainR2's accuracy is comparable to that of qPCR on a subset of strains resolved using absolute quantification. AVAILABILITY AND IMPLEMENTATION: Software is available at GitHub and implemented in C, R, and Bash. Software is supported on Linux and MacOS, with packages available on Bioconda or as a Docker container. The source code at the time of publication is also available on figshare at the doi: 10.6084/m9.figshare.29420780.

Mice

transfactor: transcription factor activity estimation via probabilistic gene expression deconvolution.

Gene expression is a primary modality being studied to differentiate between biological cells. Contemporary single-cell studies simultaneously measure genome-wide transcription levels for thousands of individual cells in a single experiment. While the characterization of cell population differences has often occurred through differential gene expression analysis, tiny effect sizes become statistically significant when thousands of cells are available for each population, compromising biological interpretation. Moreover, these large studies have spurred the development of methods to infer gene regulatory networks (GRNs) directly from the data, and GRN databases are becoming more comprehensive. In this work, we propose a statistical model for gene expression measures and an inference method that leverage GRNs to deconvolve transcription factor (TF) activity from gene expression, by probabilistically assigning mRNA molecules to TFs. This shifts the paradigm from investigating gene expression differences to regulatory differences at the level of TF activity, aiding interpretation and allowing prioritization of a limited number of TFs responsible for significant contributions to the observed gene expression differences. The inferred TF activities result in intuitive prioritization of TFs in terms of the (difference in) estimated number of molecules they produce, in contrast to other widely used methods relying on arbitrary enrichment scores. Our model allows the incorporation of prior information on the regulatory potential between each TF and target gene and is able to deal with both repressing and activating interactions. We compare our approach to other TF activity estimation methods using two simulation experiments and two case studies. Single-cell RNA-sequencing; TF activity; bioinformatics; GRN.

Transcription Factors

ChromaFactor: Deconvolution of single-molecule chromatin organization with non-negative matrix factorization.

The investigation of chromatin organization in single cells holds great promise for identifying causal relationships between genome structure and function. However, analysis of single-molecule data is hampered by extreme yet inherent heterogeneity, making it challenging to determine the contributions of individual chromatin fibers to bulk trends. To address this challenge, we propose ChromaFactor, a novel computational approach based on non-negative matrix factorization that deconvolves single-molecule chromatin organization datasets into their most salient primary components. ChromaFactor provides the ability to identify trends accounting for the maximum variance in the dataset while simultaneously describing the contribution of individual molecules to each component. Applying our approach to two single-molecule imaging datasets across different genomic scales, we find that these primary components demonstrate significant correlation with key functional phenotypes, including active transcription, enhancer-promoter distance, and genomic compartment. Also, we find that some bulk trends exist at the single-cell level, but only in a small fraction of cells, suggesting that critical changes in genome organization may be driven by specific rare subpopulations rather than occurring uniformly across all cells. ChromaFactor offers a robust tool for understanding the complex interplay between chromatin structure and function on individual DNA molecules, pinpointing which subpopulations drive functional changes and fostering new insights into cellular heterogeneity and its implications for bulk genomic phenomena.

Animals

Deconvolution of evolutionary architecture unmasks a high-risk, subclonal-rich subtype in treatment-naive small cell lung cancer.

BACKGROUND: Intratumoral heterogeneity (ITH) drives therapeutic resistance in small cell lung cancer (SCLC). However, conventional single-sample analysis has limited horizontal, cross-patient comparisons, leaving the overarching evolutionary architecture in treatment-naive tumors poorly understood. This study aims to deconvolve these architectures to identify clinically relevant evolutionary subtypes. METHODS: We analyzed whole-exome sequencing data from 41 treatment-naive SCLC patients. To overcome the cross-patient comparability bottleneck, we developed a novel probabilistic framework using a refined Gaussian Mixture Model (GMM). This standardized subclonal structures into four hierarchical strata, enabling the identification of evolutionary subtypes via unsupervised clustering. To address the scarcity of SCLC public data, prognostic concordance was robustly explored in The Cancer Genome Atlas (TCGA) lung squamous cell carcinoma (LUSC) based on shared smoking etiology, with lung adenocarcinoma (LUAD) serving as a negative control. RESULTS: The cohort robustly segregated into "Clonal-dominant" (Group 1, n=28) and "Subclonal-rich" (Group 2, n=13) subtypes. Group 1 evolution was primarily driven by tobacco signatures (SBS4). Conversely, Group 2 exhibited late-stage acquisition of a DNA mismatch repair deficiency (MMRd) signature (SBS15), fueling trace subclonal diversification. Clinically, Group 2 demonstrated a significantly lower objective response rate (ORR) to platinum-based regimens (25.0% vs. 81.3%, P=0.02). Furthermore, the Subclonal-rich architecture independently predicted inferior overall survival (OS) [adjusted hazard ratio (adj. HR) =2.93, P=0.02], driven predominantly by limited-stage disease. Cross-cancer analysis validated this histology-dependent, high-heterogeneity adverse pattern in early-stage LUSC but not in LUAD. CONCLUSIONS: This hypothesis-generating study demonstrates that a "Subclonal-rich" architecture, driven by acquired MMRd, identifies high-risk, chemo-resistant SCLC. Our GMM approach suggests that pre-existing heterogeneity may serve as a potential, histology-dependent prognostic marker that warrants prospective validation for tailoring future therapeutic regimens.

Gaussian Mixture Model (GMM)

GBMdeconvoluteR accurately infers proportions of neoplastic and immune cell populations from bulk glioblastoma transcriptomics data.

BACKGROUND: Characterizing and quantifying cell types within glioblastoma (GBM) tumors at scale will facilitate a better understanding of the association between the cellular landscape and tumor phenotypes or clinical correlates. We aimed to develop a tool that deconvolutes immune and neoplastic cells within the GBM tumor microenvironment from bulk RNA sequencing data. METHODS: We developed an IDH wild-type (IDHwt) GBM-specific single immune cell reference consisting of B cells, T-cells, NK-cells, microglia, tumor associated macrophages, monocytes, mast and DC cells. We used this alongside an existing neoplastic single cell-type reference for astrocyte-like, oligodendrocyte- and neuronal progenitor-like and mesenchymal GBM cancer cells to create both marker and gene signature matrix-based deconvolution tools. We applied single-cell resolution imaging mass cytometry (IMC) to ten IDHwt GBM samples, five paired primary and recurrent tumors, to determine which deconvolution approach performed best. RESULTS: Marker-based deconvolution using GBM-tissue specific markers was most accurate for both immune cells and cancer cells, so we packaged this approach as GBMdeconvoluteR. We applied GBMdeconvoluteR to bulk GBM RNAseq data from The Cancer Genome Atlas and recapitulated recent findings from multi-omics single cell studies with regards associations between mesenchymal GBM cancer cells and both lymphoid and myeloid cells. Furthermore, we expanded upon this to show that these associations are stronger in patients with worse prognosis. CONCLUSIONS: GBMdeconvoluteR accurately quantifies immune and neoplastic cell proportions in IDHwt GBM bulk RNA sequencing data and is accessible here: https://gbmdeconvoluter.leeds.ac.uk.

Humans

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans

Admission whole-blood transcriptomic characterization of a neutrophil-predominant systemic immune response in patients with acute traumatic brain injury.

BACKGROUND: Acute traumatic brain injury (TBI) is accompanied by systemic immune responses, but their whole-blood transcriptomic features at hospital arrival remain incompletely characterized. We aimed to characterize these features in patients with acute TBI compared with healthy controls. METHODS: In this single-center prospective observational study, we performed whole-blood RNA sequencing on hospital-arrival samples from 42 patients with acute TBI and 21 healthy controls. Analyses included differential expression (limma-voom; FDR < 0.05, |log2FC| > 0.7), functional enrichment, Ingenuity Pathway Analysis, CIBERSORTx LM22 deconvolution, and per-sample neutrophil degranulation signature scoring. RESULTS: Differential expression analysis identified 996 upregulated and 863 downregulated genes, with marked upregulation of inflammation-, innate immunity-, and neutrophil-related genes including DUSP1, HMGB2, MMP9, and S100A8. Canonical pathways with positive IPA z-scores included Neutrophil degranulation, Neutrophil Extracellular Trap Signaling Pathway, and Toll-like Receptor Signaling; upstream regulators included TNF, IL1B, IFNG, and STAT3. Deconvolution identified 7 of 22 differing subsets (q < 0.05), with relatively higher myeloid and lower lymphoid fractions in TBI. The Neutrophil degranulation signature score correlated with Injury Severity Score within TBI (Spearman &#x3c1; = +0.55; q < 0.001). CONCLUSIONS: Admission whole-blood transcriptomics characterized a neutrophil-predominant systemic transcriptional response in patients with acute TBI. This response was also evident among patients without major extracranial injury and was associated with total ISS. However, because the study lacked an appropriately matched non-TBI trauma comparator, the findings should be interpreted as a descriptive characterization of a systemic injury response accompanying TBI and do not establish a TBI-specific molecular signature or mechanism.

gene expression

Epigenetic Liquid Biopsy Enables Universal Mutation-Agnostic Molecular Surveillance for High-Risk Neuroblastoma.

PURPOSE: Liquid biopsy monitoring in pediatric solid tumors is limited by low mutational burden and lack of trackable genomic drivers. We sought to develop a mutation-agnostic, methylation-based liquid biopsy framework enabling universal molecular surveillance of high-risk neuroblastoma. EXPERIMENTAL DESIGN: Using whole-genome Oxford Nanopore Technologies sequencing of high-risk neuroblastoma tumors, we compared tumor-derived methylation profiles with a comprehensive atlas of normal human cell types and identified 72 neuroblastoma-specific differentially methylated regions (meNBL) that were reliably detectable in cell-free DNA (cfDNA). Marker robustness and specificity were validated using independent neuroblastoma methylation datasets and assessed against methylation profiles from other cancer types. We established neuroblastoma as a distinct methylation entity within the reference atlas by integrating a panel of 25 meNBLs, enabling quantitative estimation of tumor-derived cfDNA. Assay performance was evaluated across diagnostic, remission, relapse, and healthy control samples and compared with mutation-based and copy number-based approaches. RESULTS: Neuroblastoma-derived cfDNA was consistently detected at diagnosis and relapse but was absent in healthy controls and during confirmed remission. Methylation-based deconvolution demonstrated high specificity, with no detectable background signal in controls, and improved performance relative to copy number-based tumor fraction estimation. Longitudinal profiling enabled early molecular detection of relapse and reliable disease monitoring. CONCLUSIONS: We establish a robust, mutation-independent methylation-based liquid biopsy strategy for neuroblastoma that enables accurate, quantitative disease monitoring across all high-risk patients, including those lacking trackable genomic alterations. This approach supports the clinical translation of methylation-based cfDNA deconvolution as a broadly applicable platform for pediatric precision oncology.

Humans

Epigenetic and immunological alterations in umbilical cord blood of overweight/obese women with gestational diabetes mellitus: insights into DNA methylation signatures and immune cell dysregulation.

BACKGROUND: Gestational diabetes mellitus (GDM) is a common pregnancy complication associated with adverse maternal and neonatal outcomes. Epigenetic modifications may reflect intrauterine metabolic exposure and contribute to immune and metabolic alterations. This study aimed to explore DNA methylation profiles in umbilical cord blood from overweight and obese women with and without GDM. METHODS: Umbilical cord blood samples from 30 overweight/obese pregnant women (with and without GDM) were analyzed using the Illumina 850&#xa0;K methylation array to identify differentially methylated positions (DMPs) and regions (DMRs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to assess the functional relevance of methylation changes. Immune cell composition was estimated using deconvolution analysis and further examined in an independent single-cell RNA sequencing (scRNA-seq) cohort. Lasso regression was applied to identify CpG sites associated with GDM status and construct a preliminary methylation-based classification model. RESULTS: A total of 23,331 hypermethylated and 29,501 hypomethylated DMPs were identified between women with and without GDM, with hypomethylation predominating. Enrichment analyses indicated associations with neurodevelopmental pathways, metabolic processes, immune regulation, and epigenetic modification. Immune deconvolution analysis suggested reduced proportions of CD4+ T cells (p&#x2009;<&#x2009;0.05) and a trend toward decreased NK cells in the GDM group, alongside increased CD8+ T cells and neutrophils. Seven CpG sites were selected for model construction and demonstrated strong discriminatory performance within this cohort. CONCLUSION: This exploratory study identifies distinct cord blood DNA methylation patterns associated with GDM in overweight/obese pregnancies. The findings suggest potential links between epigenetic alterations and immune cell composition in GDM-exposed offspring. The identified CpG signature warrants further validation in larger, prospective cohorts to determine its clinical applicability.

Humans

MOADE: a multimodal autoencoder for dissociating bulk multi-omics data.

In single cell biology, the complexity of tissues may hinder lineage cell mapping or tumor microenvironment decomposition, requiring digital dissociation of bulk tissues. Many deconvolution methods focus on transcriptomic assay, not easily applicable to other omics due to ambiguous cell markers and reference-to-target difference. Here, we present MOADE, a multimodal autoencoder pipeline linking multi-dimensional features to jointly predict personalized multi-omic profiles and cellular compositions, using pseudo-bulk data constructed by internal non-transcriptomic reference and external scRNA-seq data. MOADE is evaluated through rigorous simulation experiments and real multi-omic data from multiple tissue types, outperforming nine deconvolution pipelines with superior generalizability and fidelity.

Humans

Malignant epithelial states drive immune dysfunction in ampulla of Vater carcinoma.

BACKGROUND: Ampulla of Vater (AoV) carcinoma is a rare malignancy arising at the junction of intestinal and pancreatobiliary epithelium. Its heterogeneous clinical behavior and histological diversity have hindered therapeutic advances, and the cellular basis of this heterogeneity remains unclear. We aimed to construct a single-cell transcriptomic atlas of AoV carcinoma, with a focus on identifying epithelial subtypes and their interactions with the tumor microenvironment (TME). METHODS: We performed single-cell RNA sequencing on eight primary AoV tumors and four matched normal tissues. Comprehensive clustering and transcriptomic analyses identified cell-type composition, epithelial heterogeneity, and tumor-immune interactions. Findings were validated using deconvolution of bulk RNA-seq data from 62 AoV carcinoma patients. Results Malignant epithelial cells were categorized into four distinct subtypes: Int-Wnt, PB-KRAS, Int-Hypoxia, and Cycling stage. PB-KRAS cells exhibited stem-like transcriptional programs and high genomic instability. Deconvolution analysis of bulk RNA-seq data from the independent AoV cohort revealed that enrichment of the PB-KRAS subtype correlated with tumor recurrence and poor survival. Our immune profiling analysis discovered a significant association between PB-KRAS subtype and GZMK+ CD8+ T cells, which are in a pre-dysfunctional state, alongside SPP1+ macrophages exhibiting immunosuppressive traits. Spatial transcriptome data further supports the immunosuppressive natures of TME around PB-KRAS subtype malignant epithelial cells in AoV carcinoma. CONCLUSIONS: Our study presents a single-cell atlas of AoV carcinoma, highlighting the molecular diversity of malignant epithelium and its association with the immune microenvironment. The PB-KRAS subtype emerges as a stem-like, immunosuppressive tumor state associated with poor prognosis, providing insights for future therapeutic targeting.

Ampulla of Vater carcinoma