Search PubMedSearch

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

CeLLTra: aligning cell names with gene expression via a pathway-informed transformer.

MOTIVATION: Single-cell RNA sequencing (scRNA-Seq) technology enables detailed exploration of gene expression at the individual cell level, crucial for annotating cell types and understanding cellular diversity. Traditional methods for cell type annotation often rely on marker genes and manual labeling, posing challenges due to low data quality and incomplete reference datasets. RESULTS: We developed CeLLTra, a novel contrastive learning framework that leverages a Transformer-based model integrating biological pathway information to group genes into super tokens, effectively capturing comprehensive gene expression from scRNA-Seq data. By combining this pathway-informed Transformer with a pretrained domain-specific language model, CeLLTra accurately aligns cell-type annotations with gene expression profiles. Evaluations on a large-scale human scRNA-Seq dataset showed that CeLLTra significantly outperformed state-of-the-art methods in supervised and zero-shot cell-type prediction. Additionally, CeLLTra generalized well to external datasets, improving clustering performance and enabling better characterization of cancerous cell states in tumor-infiltrating myeloid cells from non-small cell lung cancer patients. AVAILABILITY AND IMPLEMENTATION: CeLLTra is freely available on GitHub (https://github.com/WJZheng-group/CeLLTra) and Zenodo (https://doi.org/10.5281/zenodo.17666735). The datasets underlying this article are the following: GSE201333 and GSE127465. All these datasets are publicly available and can be freely accessed on the Gene Expression Omnibus repository.

Humans

Differential cell signaling testing for cell-cell communication inference from single-cell data by dominoSignal.

MOTIVATION: Algorithms for ligand-receptor network inference have emerged as commonly used tools to estimate cell-cell communication from reference single-cell data. Many studies employ these algorithms to compare signaling between conditions and lack methods to statistically identify signals that are significantly different. We previously developed the cell communication inference algorithm Domino, which considers ligand and receptor gene expression in association with downstream transcription factor activity scoring. We developed the dominoSignal software to innovate upon Domino and extend its functionality to test statistically differential cellular signaling. RESULTS: This new functionality includes the compilation of active signals as linkages from multiple subjects in a single-cell data set and testing condition-dependent signaling linkage. The software is applicable for analysis of single-cell data sets with multiple subjects as biological replicates as well as with bootstrapped replicates from data sets with few or pooled subjects. We use simulation studies to benchmark the number of subjects in compared groups and cells within an annotated cell type sufficient to accurately identify differential linkages. We demonstrate the application of the Differential Cell Signaling Test (DCST) in the dominoSignal software to investigate consequences of cancer cell phenotypes and immunotherapy on cell-cell communication in tumor microenvironments. These applications in cancer studies demonstrate the ability of differential cell signaling analysis to infer changes to cell communication networks from therapeutic or experimental perturbations, which is broadly applicable across biological systems. AVAILABILITY: dominoSignal is available through Bioconductor at https://www.bioconductor.org/packages/release/bioc/html/dominoSignal.html.

Cell Communication

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

Quantitative molecular cartography of emergency myelopoiesis reveals conserved modules of hematopoietic activation.

Hematopoietic stem and progenitor cells (HSPCs) respond to infections, inflammation, and regenerative challenges using emergency myelopoiesis (EM) pathways to amplify myeloid cell production. However, it remains unclear how various EM inducers regulate HSPCs using shared or distinct molecular mechanisms. Here, we generate a comprehensive and generalizable cell annotation method (HemaScribe) and a refined quantitative model of hematopoietic differentiation (HemaScape) using single-cell RNA sequencing (scRNA-seq) of murine HSPCs, which we apply to a broad range of EM modalities. We uncover multiple strategies for enhancing myelopoiesis that act at different levels of the HSPC hierarchy and are associated with both unique and shared transcriptional response modules. In particular, we identify a myeloid progenitor-based EM activation module across diverse inflammatory challenges that is conserved in humans and informs outcomes in adult and pediatric acute myeloid leukemia. Our work illuminates fundamental regulatory mechanisms in hematopoietic regeneration that have direct translational applications in disease contexts.

Animals

Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope images.

MOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc.

Microscopy, Fluorescence

Putative function and prognostic molecular marker of mast cells in colorectal cancer.

BACKGROUND: The increased demand for markers for colorectal cancer (CRC) highlights the importance of investigating immune cells involved in CRC progression. This study aims to dissect the mast cells in CRC, characterize the role of mast cells in CRC development, coordinate molecular communication between mast cells and malignant cells, and construct and validate a prognostic classification model based on mast cell markers. METHODS: Single-cell transcriptome data of CRC patients were extracted from GSE146771 for cell classification and annotation. The malignant cells were identified by copykat and the communication between mast cells and malignant cells was analyzed by CellChat. Least absolute shrinkage and selection operator (LASSO) regression analysis and Cox regression analysis of mast cell markers were performed in the TCGA-COAD cohort to construct a prognostic classification model. qRT-PCR was performed to detect the mRNA expression of the molecules in the classification model in P815 and MC-9 cells. The co-culture experiment of MC38 and P815 cells were performed in 12-well transwell dish. Wound healing assay and Transwell assay were performed to detect cell migration and invasion. RESULTS: 10,186 high-quality cells in GSE146771 were annotated to 9 cell types. Six markers in mast cells (HDC, GATA2, ASAH1, BTBD19, TIMP1, FAM110A) were selected to construct a classification model. The high-risk score defined showed high infiltration of immunosuppressive cells, including endothelial cells, CAFs, Tregs and high angiogenesis and epithelial-mesenchymal transition (EMT) activities. In the model, HDC were abnormally low expressed in P815 cells, while BTBD19, FAM110A, GATA2, ASAH1 and TIMP1 showed excessive expression in P815 cells. Knockdown of GATA2 in the co-culture system of P815 and MC38 cells blocked cell migration and invasion. CONCLUSION: This study identified the cell types within CRC, elaborated the cellular functions of mast cells in CRC development and their molecular communication to coordinate malignant cells, and highlighted the molecular components and biological features that constitute promising prognostic classification model.

Mast Cells

VDJ-Insights: simplifying the annotation of genomic immunoglobulin and T cell receptor regions.

MOTIVATION: Accurate annotation of germline immunoglobulin (IG) and T cell receptor (TCR) loci is critical for understanding adaptive immunity. RESULTS: VDJ-Insights provides a user-friendly software package for characterizing these complex immune regions. In addition, it assesses gene segment functionality, identifies recombination signal sequences, and annotates complementarity-determining regions 1 and 2. VDJ-Insights achieved over 99% concordance with curated annotations from multiple species, outperforming existing annotation tools. When applied to 95 haplotypes from the Human Pangenome Reference Consortium, VDJ-Insights identified 652 and 275 novel IG and TCR alleles, respectively, highlighting its scalability for large immunogenetic studies. AVAILABILITY AND IMPLEMENTATION: Datasets and software package are available in the VDJ-insights repository, https://github.com/BPRC-Bioinfo and https://doi.org/10.5281/zenodo.17588835. Additional intermediate datasets used and analyzed during the current study are available from the corresponding authors upon reasonable request.

Software

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans

The MYC/TXNIP axis mediates NCL-Suppressed CD8+T cell immune response in lung adenocarcinoma.

BACKGROUND: Lung adenocarcinoma is a deadly malignancy with immune evasion playing a key role in tumor progression. Glucose metabolism is crucial for T cell function, and the nucleolar protein NCL may influence T cell glucose metabolism. This study aims to investigate NCL's role in T cell glucose metabolism and immune evasion by lung adenocarcinoma cells. METHODS: Utilizing single-cell RNA sequencing (scRNA-seq) data from the Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA), we analyzed cell clustering, annotation, and prognosis. In vitro experiments involved manipulating NCL expression in CD8+ T cells to study immune function and glucose metabolism. In vivo studies using an orthotopic transplant mouse model monitored NCL's impact on CD8+ T cell glucose metabolism and anti-tumor immune function. RESULTS: NCL was associated with T cell dysfunction and glucose metabolism. NCL silencing enhanced CD8+ T cell glucose metabolism, cytotoxicity, and infiltration, while NCL overexpression had the opposite effect. NCL overexpression relieved MYC-mediated transcriptional repression of TXNIP, reducing CD8+ T cell glucose metabolism. In vivo, NCL inhibited CD8+ T cell glucose metabolism through the MYC/TXNIP axis, hindering anti-tumor immune function. CONCLUSIONS: NCL overexpression suppresses CD8+ T cell glucose metabolism and anti-tumor immune function, promoting lung adenocarcinoma progression via the MYC/TXNIP axis.

CD8-Positive T-Lymphocytes

MaxComp: Predicting single-cell chromatin compartments from 3D chromosome structures.

The genome is organized into distinct chromatin compartments with at least two main classes, a transcriptionally active A and an inactive B compartment, broadly corresponding to euchromatin and heterochromatin. Chromatin regions within the same compartment preferentially interact with each other over regions in the opposite compartment. A/B compartments are traditionally identified from ensemble Hi-C contact frequency matrices using principal component analysis of their covariance matrices. However, defining compartments at the single-cell level from sparse single-cell Hi-C data is challenging, especially since homologous copies are often not resolved. To address this, we present MaxComp, an unsupervised method, for inferring single-cell A/B compartments based on 3D geometric considerations in single-cell chromosome structures-derived either from multiplexed FISH-omics imaging or 3D structure models derived from Hi-C data. By representing each 3D chromosome structure as an undirected graph with edge-weights encoding structural information, MaxComp reformulates compartment prediction as a variant of the Max-cut problem, solved using semidefinite graph programming (SPD) to optimally partition the graph into two structural compartments. Our results show that the population average of MaxComp single-cell compartment annotations closely matches those derived from ensemble Hi-C principal component analysis, demonstrating that compartmentalization can be recovered from geometric principles alone, using only the 3D coordinates and nuclear microenvironment of chromatin regions. Our approach reveals widespread cell-to-cell variability in compartment organization, with substantial heterogeneity across genomic loci. When applied to multiplexed FISH imaging data, MaxComp also uncovers relationships between compartment annotations and transcriptional activity at the single-cell level. In summary, MaxComp offers a new framework for understanding chromatin compartmentalization in single cells, connecting 3D genome architecture, and transcriptional activity with the cell-to-cell variations of chromatin compartments.

Chromatin

DNA methylation landscape of cerebrospinal fluid cells in multiple sclerosis: an epigenome-wide association study.

BACKGROUND: Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system in which DNA methylation may link genetic and environmental risk factors. METHODS: We profiled genome-wide DNA methylation in cerebrospinal fluid (CSF) cells from people with MS (pwMS) and matched controls. Differentially methylated positions (DMPs) and regions (DMRs) were integrated with transcriptomic data, T-cell chromatin annotations, and pathway analyses. Protocadherin gamma (PCDHγ) expression was assessed in primary CD4+ T-cell subsets and confirmed by flow cytometry. FINDINGS: We identified 2710 DMPs and 4330 DMRs associating with genes that were enriched in immune signalling, adhesion and migration processes, and were accompanied by corresponding RNA changes. MS-associated methylation changes enriched in the cohesin chromatin-regulation pathway localised to T-cell regulatory regions, and this pathway included multiple protocadherin (PCDH) genes, which displayed consistent methylation and expression changes in CSF cells of pwMS compared to controls. PCDHγ cluster gene expression was detected in CD4+ T-cell subsets, and flow cytometry confirmed PCDHγ protein expression in peripheral blood T cells. Moreover, co-expression analysis suggests a role of PCDH genes in aryl hydrocarbon receptor (AHR) signalling. Protein-level validation showed fewer PCDHγ-positive CD4+ T cells in pwMS and activation-induced PCDHγ upregulation after T-cell stimulation. INTERPRETATION: DNA methylation changes in CSF resident cells reflect dysregulated T cell activation and migration in pwMS and suggest involvement of protocadherin molecules in MS pathogenesis. FUNDING: European Research Council, Swedish Research Council, Swedish Brain Foundation, Swedish MS Foundation, Knut and Alice Wallenberg Foundation, European Union and others.

Humans

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software