Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “single cell RNA sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

In Vivo CRISPR Interference Screen Reveals Long Noncoding RNA Portfolio Crucial for Cutaneous Squamous Cell Carcinoma Tumor Growth.

Cutaneous squamous cell carcinoma (cSCC) accounts for 20% of all skin cancer mortality globally, making it the second-highest subtype of skin cancer. The high prevalence of cSCC in humans highlights the need to uncover alternative actors and mechanisms influencing skin cancer development. Significant advances have been made to better understand some key factors in cSCC growth. However, little is known about the role of noncoding RNAs, particularly of a specific subclass termed long noncoding RNA (lncRNA). By performing pseudobulk analysis of single-cell sequencing data from normal and cSCC human skin tissues, we determined a global portfolio of lncRNAs specifically expressed in keratinocyte subpopulations. Integration of CRISPR interference screens in vitro and the xenograft model identified several lncRNAs impacting the growth of cSCC cancer lines both in vitro and in vivo. Among these, we further validated LINC00704 and LINC01116 as proliferation-regulating lncRNAs in cSCC lines and potential biomarkers of cSCC growth. Taken together, our study provides a comprehensive signature of lncRNAs with roles in regulating cSCC growth.

RNA, Long Noncoding↗

Immature Neutrophil Programs Associate With Burn Mortality and Extend Across Critical Illnesses.

Severe burns provoke a systemic "genomic storm," yet cell states associated with divergent outcomes remain unclear. We profiled blood cells by single-cell RNA-Sequencing (73 014 cells) from adult patients with burn injuries within postburn day 17 (n = 4) and healthy donors (n = 5), integrated data with bulk signatures of burn size, inhalation injury, and mortality, and evaluated clinical associations in the American Burn Association National Burn Repository. Burn was associated with emergency hematopoiesis marked by expansion of hematopoietic stem/progenitor-like cells, immature neutrophils, and plasmablast/plasma cell states, alongside depletion of naïve CD4+/CD8+ T cells and dendritic cells. Larger burns (>20% TBSA) showed enrichment of humoral transcriptional programs, including plasmablast/plasma cell activation and suppression of cytotoxic CD8+ T-cell states. In multivariable models, inhalation injury was a stronger predictor of death (adjusted odds ratio [OR] 1.9) than burn size (adjusted OR 1.1) and shared greater overlap with the most perturbed single cells in non-survivors; 55% of co-perturbed cells were neutrophils, implicating granulocyte dysregulation as a common lethal axis. We identified a neutrophil-specific 5-gene panel (OLFM4, RETN, LCN2, ARG1, and BTNL3) that discriminated survivors vs non-survivors after burns (area under the curve [AUC] > 0.9) and generalized to trauma (n = 158; AUC 0.81) and intensive care unit COVID-19 (n = 103; AUC 0.75), providing information orthogonal to conventional biomarkers and severity scores. Cytomorphology corroborated transcriptomic immaturity, with ~2-fold higher band neutrophils and larger neutrophil size in a fatal case. Computational drug-reversal analysis highlighted galectin-1 inhibition as a candidate modulator of mortality-associated neutrophil programs. Together, our findings suggest that immature neutrophils represent a shared immune feature across severe burns and other forms of critical illness.

Humans↗

Spatial mutual nearest neighbors for spatial transcriptomics data.

MOTIVATION: Mutual nearest neighbors (MNN) is a widely used computational tool to perform batch correction for single-cell RNA-sequencing data. However, in applications such as spatial transcriptomics, it fails to take into account the 2D spatial information. RESULTS: Here, we present spatialMNN, an algorithm that integrates multiple spatial transcriptomic samples and identifies spatial domains. Our approach begins by building a k-nearest neighbors (kNN) graph based on the spatial coordinates, prunes noisy edges, and identifies niches to act as anchor points for each sample. Next, we construct a MNN graph across the samples to identify similar niches. Finally, the spatialMNN graph can be partitioned using existing algorithms, such as the Louvain algorithm to predict spatial domains across the tissue samples. We demonstrate the performance of spatialMNN using large datasets, including one with N = 31 10x Genomics Visium samples. We also evaluate the computing performance of spatialMNN to other popular spatial clustering methods. AVAILABILITY AND IMPLEMENTATION: Our software package is available on GitHub (https://github.com/Pixel-Dream/spatialMNN). The code is available on Zenodo (https://doi.org/10.5281/zenodo.15073963).

Algorithms↗

From transcriptomic profiling to precision oncology: a bibliometric analysis of RNA sequencing in acute myeloid leukemia.

BACKGROUND: RNA sequencing (RNA-seq) has become an important tool for investigating the molecular heterogeneity of acute myeloid leukemia (AML); however, the global development and thematic evolution of this field remain inadequately characterized. OBJECTIVE: To map the global landscape of AML RNA-seq research and identify major knowledge domains, emerging themes, and temporal changes in research priorities. METHODS: Publications indexed in the Web of Science Core Collection and Scopus between January 1, 2007, and August 18, 2025, were retrieved. After database filtering, merging, and deduplication, 3,460 articles and reviews were included. CiteSpace, VOSviewer, the bibliometrix R package, and Microsoft Excel were used to analyze publication trends, collaboration networks, co-citation structures, keyword evolution, and citation bursts. RESULTS: Publication output increased steadily, accelerating after 2014. China contributed the largest number of publications (n = 547, 15.8%), whereas the United States had the highest total citation count. Major publication outlets spanned hematology, oncology, genomics, and molecular biology. Co-citation analysis identified prominent themes involving next-generation sequencing, gene mutations, KMT2A rearrangements, epigenetic dysregulation, leukemia-initiating cells, drug resistance, biomarkers, T-cell biology, and single-cell sequencing. Earlier literature emphasized sequencing technologies, gene expression profiling, and molecular alterations, whereas recent publications show increasing representation of cellular heterogeneity, single-cell transcriptomics, drug resistance, biomarker applications, immune-related research, and computational interpretation. CONCLUSION: While molecular characterization remains foundational, AML RNA-seq research has broadened to encompass increasingly prominent cellular, functional, computational, and translational dimensions. This study provides a structured overview of the field; nevertheless, bibliometric prominence should not be interpreted as direct evidence of clinical utility.

RNA sequencing↗

A pan-cancer single-cell atlas uncovers the role of sex hormones and chromosomes in sex-divergent reprogramming of the tumor microenvironment.

BACKGROUND: Sex bias is pervasive in tumors; however, how sex chromosomes and hormone-responsive signaling shape the tumor microenvironment (TME) remains insufficiently characterized. Considering the critical impact of the TME on tumor progression and response to immunotherapy, a pan-cancer investigation of sex-specific and cancer-context-dependent TME features is warranted. METHOD: Based on stringent inclusion criteria, we constructed a high-resolution pan-cancer single-cell sequencing atlas by integrating 31 publicly available single-cell RNA-seq datasets, comprising a total of 1,831,436 cells by integrating 468 samples from eight types of non-sex-specific solid tumors (282 males and 186 females). After correcting for batch effects, we identified major and minor cellular subsets. Multiple computational approaches were applied to investigate sex-associated differences in cellular composition, gene expression, pathway activity, malignant cell states and intercellular communication. RESULTS: We systematically compared sex-specific TME features across eight common solid malignancies. Male-biased CD8+ T cell exhaustion emerged as a recurrent but non-uniform feature, with its magnitude varying across cancer types and being modified by tissue-specific contexts. This pattern was associated with androgen-response signature scores and expression-based loss of the Y chromosome (LOY) scores. M2-like macrophage polarization showed a more cancer-type-dependent pattern; although female-biased enrichment was observed in selected malignancies, it did not represent a uniform pan-cancer feature. Expression-based X chromosome inactivation (XCI)/XCI escape-related programs, estrogen-response signature scores and stromal components, including fibroblasts and endothelial cells, were associated with macrophage and immune-regulatory states in specific tumor contexts. Tumor cells of male origin displayed higher genomic instability and more aggressive phenotypes, with androgen-response signatures and LOY contributing to the development of a male biased malignant state. Furthermore, expression-based LOY scores in malignant cells were associated with CD8+ T cell exhaustion based on transcriptomic proxies. CONCLUSION: Our study uncovers extensive but heterogeneous sex-specific differences in the TME across multiple cancer types. We propose a regulatory framework linking sex chromosomes, hormone-responsive signaling and TME interactions, which is consistent with recurrent male-biased CD8⁺ T cell exhaustion and context-dependent M2-like macrophage polarization. Importantly, the magnitude and, in some cancers, the direction of these sex-biased features are modified by tissue-specific contexts. These findings underscore the need to include sex chromosome and hormone status as essential biological variables in studies of the tumor microenvironment and the design of immunotherapies.

Tumor Microenvironment↗

Elevated intron retention implicates neuroinflammation in brains of individuals with alcohol use disorder.

Intron retention, a form of alternative RNA splicing, can occur as part of normal gene regulation or result from disruption of the splicing machinery. Retained introns can potentially form double-stranded RNA, activating innate immune sensors and inflammation. This mechanism has been implicated in cancer but has not been studied in neuropsychiatric diseases like alcohol use disorder. We systematically analysed transcriptome-wide intron retention events in post-mortem brain tissue from 142 individuals (66 with alcohol use disorder and 76 controls), encompassing 320 region-specific samples from the superior frontal cortex, nucleus accumbens, central nucleus and basolateral amygdala. Analyses were adjusted for demographic, technical and biological covariates. Validation was performed in alcohol-preferring (P) rats using long-read sequencing. In complementary experiments, immunofluorescent staining was used to detect double-stranded RNA in rat brain tissue, while single-cell RNA-sequencing was performed to test activation of double-stranded RNA-sensing pathways in human brains. Brains from individuals with alcohol use disorder showed significantly higher total intron retention compared with controls, independent of age, with females showing greater increases than males. A total of 368 introns were positively associated with alcohol use disorder, and these introns were significantly longer and had weaker splice acceptor sites compared with non-associated introns. Genes harbouring these intron retention events were enriched in Purkinje neurons, visual cortex neurons and oligodendrocytes. Computational predictions indicated these long introns could form duplex RNA structures. Increased double-stranded RNA was confirmed experimentally in multiple brain regions of alcohol-consuming rats, where it co-localized primarily with neuronal nuclei and dendrites. In individuals with alcohol use disorder, we found that multiple pathways including double-stranded RNA responses, neuroinflammation, interferon and NF-κB signalling, adaptive immunity and apoptosis were activated. In addition, NeuN-positive neuronal counts significantly decreased in both the prefrontal and visual cortices. Furthermore, single-cell analysis demonstrated upregulation of TICAM1, the target of double-stranded RNA sensor TLR3, in oligodendrocytes, as well as widespread activation of downstream inflammatory pathways across glial and neuronal cell types. These findings provide the first evidence that chronic alcohol consumption promotes an overall increase of intron retention in the brain and is associated with the presence of double-stranded RNA. Furthermore, the double-stranded RNA may contribute to neuronal loss and brain pathology by activating a neuroinflammatory response.

alcohol use disorder↗

Protocol of electron microscope in situ nucleic acid hybridization for the exclusive detection of double-stranded DNA sequences in cells containing large amounts of homologous single-stranded DNA and RNA sequences: application to adenovirus type 5 infected HeLa cells.

In order to gain a further insight into the relationships of the complex process of replication of adenovirus genomes to the substructures which occur in the nuclei of adenovirus type 5 (Ad5) infected HeLa cells, we have visualized directly, at the electron microscopic level, viral double-stranded DNA (dsDNA) in late infected nuclei by the use of a post-embedding in situ hybridization technique with a biotinylated specific DNA probe. The procedure is based on the removal of single-stranded (ss) nucleic acids by S1 nuclease. The highest levels of signal density for viral dsDNA were detected over the fibrils of the large, centrally located viral genome storage site and over the viral nucleoids of both clustered and isolated viruses. Lower but significant signals were observed over the fibrillo-granular network of the peripheral replicative zones, where both transcription and replication of viral DNA occur. On the other hand, the labeling of the enclosed viral ssDNA accumulation sites, also involved in viral replication but not transcription, was negligible, which suggests that, in the latter, the newly synthesized viral dsDNA immediately extends into the adjacent peripheral replicative zone to be transcribed and/or replicated.

Adenoviruses, Human↗

REACTOR: REgulon Activity analysis and Comparison Tool for single-cell transcriptOmics Research.

SUMMARY: We introduce REACTOR, a computational tool designed to detect differential activity of transcriptional regulators and their target genes (regulons) in single-cell RNA-sequencing data. It expands the currently available framework for regulon analysis by introducing a robust statistical test to detect differential regulon activity between conditions, such as disease versus control, with multiple replicates. By contrasting different conditions, REACTOR enables identification of key condition- and cell type-specific regulons. To demonstrate the use of REACTOR, we illustrate its performance in a publicly available COVID-19 dataset. AVAILABILITY: REACTOR R-package together with an implementation vignette are available at https://www.github.com/elolab/REACTOR.

Regulon↗

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans↗

CROPseq-multi: a universal solution for multiplexed perturbation in high-content pooled CRISPR screens.

Forward genetic screens seek to dissect complex biological systems by systematically perturbing genetic elements and observing the resulting phenotypes. While standard screening methodologies introduce individual perturbations, multiplexing perturbations improves the performance of single-target screens and enables combinatorial screens for the study of genetic interactions. Current tools for multiplexing perturbations are limited by technical challenges and do not offer compatibility across diverse screening methodologies, including enrichment, single-cell sequencing, and optical pooled screens. Here, we report the development of CROPseq-multi (CSM), a CROPseq1-inspired lentiviral system to multiplex Streptococcus pyogenes (Sp) Cas9-based perturbations with versatile readout compatibility and high performance for both perturbation and barcode identification. CSM has equivalent per-guide activity to CROPseq and low lentiviral recombination frequencies. Dual-guide CSM libraries are constructed in a single, facile molecular cloning step that facilitates the use of unique molecular identifiers. CSM is compatible with enrichment screening methodologies, single-cell RNA-sequencing readouts, and optical pooled screens. For optical pooled screens, an optimized and multiplexed in situ detection protocol improves barcode counts 10-fold (for mRNA detection), enables detection of recombination events, and reduces the number of sequencing cycles required for decoding by 3-fold relative to CROPseq. CROPseq-multi-v2 (CSMv2) adds compatibility for detection methods based on T7 RNA polymerase in vitro transcription2-5. CSM provides a single system for CRISPR screens that is compatible with individual and combinatorial perturbations, diverse SpCas9-based perturbation technologies, and multiple high-content, single-cell phenotypic readouts.

CRISPR Cas9↗

TCGA-based identification of prognostic biomarkers and candidate traditional Chinese medicine compounds in papillary thyroid carcinoma: An observational study.

This study aimed to identify prognostic genes associated with papillary thyroid carcinoma (PTC) and explore candidate traditional Chinese medicine (TCM) compounds using integrated bioinformatics and molecular docking. In this observational study, PTC gene expression profiles and clinical data were obtained from The Cancer Genome Atlas. Differentially expressed genes were screened using differential-expression sequencing (DESeq2), followed by protein-protein interaction network analysis to identify hub genes. Their expression, diagnostic value, immune relevance, prognostic significance, protein-level validation, and single-cell distribution were assessed using gene expression profiling interactive analysis, receiver operating characteristic analysis, immune infiltration analysis, Kaplan-Meier survival analysis, the human protein atlas, and single-cell RNA-sequencing data. Candidate TCM compounds were predicted using symptom mapping (SymMap) and the TCM Systems Pharmacology Database and Analysis Platform, and molecular docking was performed to evaluate potential ligand-target interactions. Five hub genes, colony-stimulating factor 2, apolipoprotein E, fibronectin 1 (FN1), collagen type I alpha 1 chain (COL1A1), and intercellular adhesion molecule 1, were identified and found to be significantly upregulated in PTC tissues, with diagnostic value in receiver operating characteristic analysis. Immune infiltration analysis showed associations with macrophages, dendritic cells, and T helper 1 cells, whereas single-cell analysis demonstrated heterogeneous expression across immune and stromal cell populations, including fibroblasts. Higher FN1 and COL1A1 expression was associated with poorer outcomes. Immunohistochemistry supported the expression patterns, while single-cell analysis provided exploratory cell-type-level context for the cellular distribution of selected genes. Ginseng and Smilax glabra were predicted as common candidate TCMs, and docking suggested favorable binding between their active compounds and selected hub targets. Colony-stimulating factor 2, apolipoprotein E, FN1, COL1A1, and intercellular adhesion molecule 1 may be biologically relevant hub genes in PTC, while FN1 and COL1A1 may have prognostic value. Predicted TCM compounds provide preliminary computational evidence for possible compound-target interactions, requiring experimental and clinical validation.

Female↗

Immune-Like Malignant Epithelial Programs Shape Tumor-Immune Interactions and Inform Prognostic Stratification in Lung Adenocarcinoma.

Lung adenocarcinoma (LUAD) is characterized by marked cellular heterogeneity, yet how malignant epithelial states contribute to immune regulation and clinical outcomes remains incompletely defined. We integrated single-cell RNA-sequencing data to map the cellular landscape of LUAD and identify malignant epithelial cells based on inferred copy-number alterations. Epithelial states were further examined through trajectory inference, transcription factor analysis, and cell-cell communication profiling. Single-cell-derived genes were subsequently integrated with TCGA and independent GEO cohorts to construct and validate a machine learning-based prognostic signature. Malignant epithelial cells displayed distinct functional programs, including an immune-like state associated with genomic instability, immune-related transcriptional activity, tumor-immune communication, and patient outcomes. The resulting immune-like malignant epithelial cell signature (IMEC-Sig) consistently stratified survival across multiple cohorts. Low IMEC-Sig scores were accompanied by greater immune infiltration, higher immune checkpoint expression, and increased immunophenoscore, whereas high scores were linked to a comparatively immunosuppressive phenotype. Pan-cancer analyses further identified KRT8 as a gene associated with unfavorable prognosis, and functional experiments showed that KRT8 silencing suppressed proliferation, migration, invasion, and colony formation in LUAD cells. Together, these findings connect malignant epithelial heterogeneity with the immune context and clinical outcomes, support IMEC-Sig as a biologically informed prognostic tool, and nominate KRT8 as a potential therapeutic target in LUAD.

Humans↗

A machine learning-derived and functionally validated circadian rhythm signature predicts clinical outcomes and in silico drug sensitivity in colorectal cancer.

BACKGROUND: Colorectal cancer (CRC) displays considerable heterogeneity in clinical outcomes, highlighting the need for reliable prognostic biomarkers. While the aberrant expression of circadian rhythm-related genes has been implicated in cancer pathogenesis, its comprehensive role in CRC progression and predicted therapeutic vulnerabilities remains inadequately characterized. METHODS: Bulk and single-cell RNA-sequencing data were integrated from multiple CRC cohorts. A circadian rhythm signature (CRS) was developed through machine learning algorithms and validated for prognostic value. Comprehensive analyses of tumor microenvironment, genomic alterations, and drug sensitivity were performed. Furthermore, the biological function of the core gene, BHLHE40, was validated in CRC cell lines through CCK-8, EdU, and wound healing assays. RESULTS: Single-cell analysis demonstrated an elevated expression signature of circadian rhythm-related genes in dendritic cells. The optimized CRS, comprising 14 circadian rhythm-related genes, successfully categorized patients into high- and low-risk groups. Patients with a high CRS showed markedly poorer overall survival and computationally inferred immunosuppressive features, including reduced CD8+ T cell infiltration and increased M2 macrophage polarization. Genomic analysis revealed enhanced mutation burden in TP53 and alterations in RTK-RAS/WNT pathways. Notably, in vitro assays confirmed that BHLHE40 is significantly overexpressed in CRC cells. Knockdown of BHLHE40 markedly inhibited tumor cell proliferation and migration. Drug sensitivity profiling identified bexarotene and SMER-3 as potential therapeutic options for high-CRS patients. A nomogram integrating CRS with clinical parameters demonstrated superior predictive accuracy for 1-, 3-, and 5-year survival. CONCLUSIONS: The CRS represents a promising prognostic biomarker that reflects tumor immune status and genomic features, providing valuable insights for personalized treatment strategies in CRC.

Circadian rhythm↗

High-throughput single-cell proteomics and transcriptomics from same cells with a nanoliter-scale, spin-transfer approach.

Single-cell multiomic platforms provide a comprehensive snapshot of cellular states and cell types by offering critical insights into the spatiotemporal regulation of biomolecular networks at a systems level, thereby defining the basis of multicellularity. Here, we introduce nanoSPINS, an advanced platform that enables high-throughput profiling and integrative analysis of the transcriptome and proteome from the same single cells using RNA sequencing and isobaric labeling LC-MS-based proteomics, respectively. NanoSPINS can efficiently transfer mRNA-containing droplets across two microarrays via a centrifugation-based approach, while proteins are retained on the initial platform. Benchmarking of nanoSPINS on two cell lines demonstrates its ability to generate global proteomic and transcriptomic profiles that align well with previously established methodologies/platforms. The incorporation of isobaric TMTpro labeling into this single-cell multiomics platform significantly enhances the throughput of single-cell proteomic analyses. Through the high-throughput quantification of the proteome and transcriptome, nanoSPINS not only facilitates the identification of molecular features at both mRNA and protein level but also provides larger sample sizes for improved statistical power in clustering and differential abundance. Given the broad applicability of single-cell multiomics in biological research and clinical settings, we believe nanoSPINS represents a powerful platform for the characterization of heterogeneous cell populations.

Single-Cell Analysis↗

Ligand-based directed differentiation to produce granulosa-like cells expressing steroidogenic enzyme genes.

The ovarian granulosa cells are responsible for producing hormones and supporting oocytes through maturation and meiotic resumption. There is a need to generate granulosa-like cells (GLCs) from human induced pluripotent stem cells (hiPSCs) to better model human gonadal development and to test the effects of exogenous or pharmaceutical compounds on the ovary. Here we report a rapid ligand-based protocol for differentiating hiPSCs into cells that express markers of the transient developmental lineages and steroidogenic pathway genes. Single-cell RNA-sequencing (scRNA-seq) analysis identified canonical granulosa cell genes were expressed in a subset of cells and identified new genes of interest that were significantly associated with computationally modeled pseudotime. HSD17B1 was expressed in resulting GLCs but at low levels, suggesting an immature granulosa cell phenotype. The GLCs were produced using a simple culture method that could be augmented for granulosa cell functions such as sustaining oocyte growth. Producing GLCs through protocols such as this one is a first step toward designing large-scale ovarian endocrinology assays and developing personalized cell-based fertility and hormone restoration technologies in the future. This rapid protocol produced cells that express steroidogenic enzyme genes etoc blurb. Kubo and colleagues present a 5-day rapid protocol to generate immature granulosa-like cells from hiPSCs. Cells differentiated with inhibition of DKK1, a WNT signaling target gene, expressed gonadal ridge markers and FOXL2 transcripts and protein. Additionally, steroidogenic enzyme genes were expressed. A small population of differentiated cells were identified as expressing early-stage granulosa cell genes by single-cell RNA-seq.

Female↗

Integrative analysis of single-cell sequencing identifies CD8+ TIM3+ CD101+ T cell-associated genes as prognostic biomarkers in breast cancer.

BACKGROUND: Breast cancer is a prevalent and deadly malignancy that significantly impacts women's quality of life and imposes financial burdens. Despite therapeutic advancements, tumour heterogeneity and frequent relapses remain major challenges. Accordingly, this study aimed to characterize immune features associated with CD8+ TIM3+ CD101+ T cells and develop a prognostic signature for breast cancer. METHODS: This study integrated single-cell and bulk transcriptomic datasets to characterize CD8+ TIM3+ CD101+ T cell (CCT)-related immune features and construct a prognostic signature in breast cancer. Single-cell RNA-seq data were sourced from the Gene Expression Omnibus (GEO) repository, and bulk transcriptomic data were from The Cancer Genome Atlas (TCGA) and GEO databases. Analytical methods included pseudo-time trajectory reconstruction (Monocle2), intercellular signalling analysis (CellChat), functional enrichment (ClusterProfiler), and immune profiling (ssGSEA). Prognostic modeling was conducted using least absolute shrinkage and selection operator (LASSO) Cox regression, with validation via Kaplan-Meier and time-dependent receiver operating characteristic (ROC) analyses. RESULTS: Single-cell analysis identified 17 clusters spanning seven cell types, including T cells, myeloid cells, and epithelial cells. T-cell sub-clustering revealed four subtypes. Pseudotime analysis suggested a potential state-transition relationship between CD8+ CD101- TIM3+ and CD8+ CD101+ TIM3+ T-cell states. A total of 121 differentially expressed genes were enriched in vital biological processes. An 11-gene prognostic model showed strong predictive power across cohorts. Single-cell T-cell reclustering identified a CD8+ CD101+ TIM3+ T-cell subpopulation, which was primarily characterized by the expression of markers such as CD101 and HAVCR2/TIM3. CONCLUSIONS: This study maps cellular heterogeneity and molecular networks in breast cancer, offering insights for targeted therapy and improved prognosis.

Breast invasive carcinoma↗

shinyDeepGxP: a user-friendly R shiny app for predicting surface protein abundance from scRNA-seq expression using deep learning in blood cells.

MOTIVATION: Understanding accurate immune cell heterogeneity and function in single-cell datasets requires access to protein-level information, which is often unavailable due to experimental limitations. RESULTS: We present shinyDeepGxP, an interactive web application featuring our deep learning model, DeepGxP, for predicting surface protein abundance from single-cell RNA-sequencing (scRNA-seq) data. This platform makes DeepGxP accessible to researchers without programming skills. Users can upload scRNA-seq count matrices and use "Predict Protein" to predict the abundance of 224 biologically relevant surface proteins. shinyDeepGxP provides visualizations to help identify distinct cell populations based on predicted protein profiles. Moreover, users can choose "Explore Model" to reveal key RNA predictors and their associated biological pathways for each protein. Overall, shinyDeepGxP is a user-friendly, freely available web tool that provides protein-level detail for RNA-only single-cell datasets, enabling multimodal discovery without additional experiments. AVAILABILITY AND IMPLEMENTATION: shinyDeepGxP can be launched on https://shiny.crc.pitt.edu/deepgxp/.

Journal Article↗

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans↗