Search PubMedSearch

SEARCH · Search PubMed

Results for “Cell typing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Spatial isoform sequencing at single-cell resolution reveals cell-type-specific spatial isoform variability in multiple brain cell types.

Spatial long-read technologies are increasingly common but usually lack single-cell resolution. This leaves unanswered whether spatially variable isoforms reflect variability within one cell type or differences in region-specific cell-type composition. Here, we developed Spl-ISO-Seq2 (500-nm resolution) and accompanying software, Spl-IsoQuant-2 and Spl-IsoFind, enabling long-read sequencing of >450 million barcodes versus 80,000 previously. Applying this to the adult mouse brain, we compared differential isoform abundance between known regions and spatial isoform patterns independent of predefined regions. Both identified overlapping hits, for example, Rps24 in oligodendrocytes. For known Snap25 spatial isoform variation, we show that it occurs in excitatory neurons. The region-agnostic approach also uncovered patterns missed by region-based comparisons, for example, for Ighm. Notably, many spatial isoform signals are not driven by cell-type composition alone. Finally, our software is applicable to many spatial and single-cell protocols, demonstrating reproducibility between platforms (for example, Visium HD/Stereo-seq). Overall, our experimental/analytical methods enable a submicron-resolution-isoform view and open avenues for spatial isoform disease research.

Animals

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However, the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

scATAnno: Automated Cell Type Annotation for Single-cell ATAC-seq Data.

Recent advances in single-cell epigenomic techniques have increased the demand for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) analysis. One key analytical task is to determine cell type identity based on epigenetic data. Here, we introduce scATAnno, a Python package designed to automatically annotate scATAC-seq data using large-scale scATAC-seq reference atlases. This workflow generates reference atlases from publicly available datasets, enabling accurate cell type annotation by integrating query data with reference atlases without the use of single-cell RNA sequencing (scRNA-seq) data. To enhance annotation accuracy, we incorporated k-nearest neighbors (KNN)-based and weighted distance-based uncertainty scores to effectively detect cell populations within the query data that are distinct from all cell types in the reference data. We compared and benchmarked scATAnno against five other published cell annotation approaches, demonstrating its superior performance across multiple datasets and metrics. We further showcased the utility of scATAnno across multiple datasets, including peripheral blood mononuclear cells (PBMCs), triple-negative breast cancer (TNBC), and basal cell carcinoma (BCC), and demonstrated that scATAnno accurately annotates cell types across diverse biological conditions. Overall, scATAnno is a useful tool for scATAC-seq reference atlas construction and cell type annotation and can facilitate the interpretation of new scATAC-seq datasets in complex biological systems. scATAnno is publicly available at https://scatanno-main.readthedocs.io/.

Single-Cell Analysis

Defining breast epithelial cell types in the single-cell era.

Single-cell studies on breast tissue have contributed to a change in our understanding of breast epithelial diversity that has, in turn, precipitated a lack of consensus on breast cell types. The confusion surrounding this issue highlights a possible challenge for advancing breast atlas efforts. In this perspective, we present our consensus on the identities, properties, and naming conventions for breast epithelial cell types and propose goals for future atlas endeavors. Our proposals and their underlying thought processes aim to catalyze the adoption of a shared model for this tissue and to serve as guidance for other investigators facing similar challenges.

Humans

Cell type resolved MR based on brain single cell eQTLs corroborated by single cell RNA sequencing uncovers neuroimmune and vascular programs in intracerebral hemorrhage.

BACKGROUND: Intracerebral hemorrhage (ICH) lacks effective neuroprotective therapies. We integrated cell type–resolved genetic inference with single-cell profiling to map putative causal programs and multicellular circuitry relevant to ICH. METHODS: Cis-eQTLs from eight human brain cell types were used as instruments for two-sample Mendelian randomization (MR), with an ICH meta-analysis from large biobanks and a stroke consortium as the outcome. Instruments were LD-pruned and restricted to strong variants (F > 10). Inverse-variance weighting (IVW) was the primary estimator, supported by robustness methods, heterogeneity/pleiotropy diagnostics, and false discovery rate control. Experimental validation used mouse collagenase ICH single-cell RNA-seq at 24 h (n = 3 sham; n = 3 ICH) with Seurat integration, composition testing, Slingshot pseudotime, and CellChat. An independent mouse cohort underwent qRT–PCR for selected genes. RESULTS: The ICH meta-analysis showed acceptable genomic control, supporting downstream MR. We identified 524 nominal gene–cell type associations, with a glia-weighted signal landscape. Enrichment implicated autophagy/mitophagy, antigen processing, cytoskeletal and vesicular trafficking, endothelial matrix–adhesion programs, ferroptosis, and myelin stress pathways. In mouse scRNA-seq, disease-associated microglia expanded with reciprocal loss of homeostatic microglia and increased neutrophils and T cells. Prioritized genes showed directional concordance; qRT–PCR confirmed ARPC3 and EIF2AK2 upregulation and TBCK and SPECC1 downregulation in ICH versus sham. Pseudotime supported a shift toward disease-associated microglial states, and CellChat indicated increased network interaction strength with microglia and endothelium as hubs. CONCLUSIONS: Cell type–specific MR combined with single-cell validation highlights neuroimmune and neurovascular programs in ICH and links genetic signals to state transitions and inferred intercellular communication.

Animals

Sparse deconvolution of cell type medleys in spatial transcriptomics.

Mapping cell distributions across spatial locations with whole-genome coverage is essential for understanding cellular responses and signaling However, current deconvolution models aim to estimate the proportions of distinct cell types in each spatial transcriptomics spot by integrating reference single-cell data. These models often assume strong overlap between the reference and spatial datasets, neglecting biology-grounded constraints such as sparsity and cell-type variations, as well as technical sparsity. As a result, these methods rely on over-permissive algorithms that ignore given constraints leading to inaccurate predictions, particularly in heterogeneous or unmatched datasets. We introduce Weight-Induced Sparse Regression (WISpR), a machine learning algorithm that integrates spot-specific hyperparameters and sparsity-driven modeling. Unlike conventional approaches that neglect biology-grounded constraints, WISpR accurately predicts cell-type distributions while preserving biological coherence, i.e., spatially and functionally consistent cell-type localization, even in unmatched datasets. Benchmarking against five alternative methods across ten datasets, WISpR consistently outperformed competitors and predicted cellular landscapes in both normal and cancerous tissues. By leveraging sparse cell-type arrangements, WISpR provides biologically informed, high-resolution cellular maps. Its ability to decode tissue organization in both healthy and diseased states highlights WISpR's practical utility for spatial transcriptomics, particularly in challenging settings involving noise, sparsity, or reference mismatches.

Humans

Systematic identification of oscillatory gene expression in single cell types.

Many biological cycles are driven by oscillatory gene expression coordinated across cell types. For example, larval development in Caenorhabditis elegans involves coordinated cyclic changes in cell division, behavior, and growth, the latter requiring production of a structured extracellular matrix called the cuticle. Here, we combine single-cell RNA sequencing and novel computational approaches to identify oscillatory gene expression in individual cell types. We find that many cell types exhibit looping structures in PCA and UMAP space that correspond to transcriptional oscillations at each larval stage. Oscillatory gene expression is found in all cuticle-producing cell types, including glia, but not detected in neurons or muscle. We develop rigorous statistical approaches for de novo identification of oscillatory genes and cell types, yielding >5,000 genes. While many oscillatory genes relate to cuticle production, each cell type expresses largely distinct genes, suggesting that cuticle production is a patchwork of cell-type-specific programs. Finally, we derive a potential set of regulatory transcription factors that can explain coordinated oscillatory gene expression and find that shared upstream factors likely control gene timing across cell types. Together, our results suggest that shared regulators control cell-type-specific oscillatory gene expression, including in previously overlooked cell types such as glia.

Journal Article

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans

Detection of cell-type-specific differentially methylated regions in epigenome-wide association studies.

MOTIVATION: DNA methylation at cytosine-phosphate-guanine (CpG) sites is one of the most important epigenetic markers. Therefore, epidemiologists are interested in investigating DNA methylation in large cohorts through epigenome-wide association studies (EWAS). However, the observed EWAS data are bulk data with signals aggregated from distinct cell types. Deconvolution of cell-type-specific signals from EWAS data is challenging because phenotypes can affect both cell-type proportions and cell-type-specific methylation levels. Recently, there has been active research on detecting cell-type-specific risk CpG sites for EWAS data. However, existing methods all assume that the methylation levels of different CpG sites are independent and perform association detection for each CpG site separately. Although these methods significantly improve the detection at the aggregated-level-identifying a CpG site as a risk CpG site as long as it is associated with the phenotype in any cell type, they have low power in detecting cell-type-specific associations for EWAS with typical sample sizes. RESULTS: Here, we develop a new method, Fine-scale inference for Differentially Methylated Regions (FineDMR), to borrow strengths of nearby CpG sites to improve the cell-type-specific association detection. Via a Bayesian hierarchical model built upon Gaussian process functional regression, FineDMR takes advantage of the spatial dependencies between CpG sites. FineDMR can provide cell-type-specific association detection as well as output subject-specific and cell-type-specific methylation profiles for each subject. Simulation studies and real data analysis show that FineDMR substantially improves the power in detecting cell-type-specific associations for EWAS data. AVAILABILITY AND IMPLEMENTATION: FineDMR is freely available at https://github.com/JiaRuofan/Detection-of-Cell-type-specific-DMRs-in-EWAS.

DNA Methylation

Cell-type-specific genetic associations in Lewy body dementia identified using single-cell eQTL-based Mendelian randomization.

BACKGROUND: Lewy body dementia (LBD) is a complex neurodegenerative disorder marked by α-synuclein aggregation and dual impairment of cognitive and motor function.While genome-wide association studies have identified risk loci, the cellular mechanisms linking genetic variation to disease susceptibility remain largely unexplored. METHODS: We performed single-cell transcriptome-wide Mendelian randomization using brain cell-type-specific eQTLs across eight major cell types. Genetic associations were evaluated using inverse-variance weighted models, followed by Bayesian colocalization analysis. Replication was performed in independent stratified LBD cohorts based on APOE ε4 carrier status. Phenome-wide association analysis was included as a supplementary, descriptive assessment of cross-trait associations. RESULTS: Expression of ANKRD65 in excitatory neurons was significantly associated with reduced LBD risk (odds ratio = 0.65, 95 % CI: 0.52-0.81, p = 0.00013). This association passed a false discovery rate of 0.1 and showed strong evidence of colocalization (posterior probability = 0.93). Effect direction was consistent across APOE ε4+ and ε4- LBD subgroups in independent cohorts. No genome-wide significant associations were observed with non-neurological traits in the phenome-wide analysis. CONCLUSIONS: Our findings identify a genetically supported, cell-type-resolved association between ANKRD65 expression in excitatory neurons and LBD risk. This study demonstrates the value of integrating cell-resolved transcriptomic regulation with genetic inference to pinpoint functionally relevant targets in neurodegenerative diseases.

Humans

Beyond Bulk: Cell-Type-Resolved Epigenomics as the Path Forward in Alzheimer's Disease Research.

Alzheimer's disease (AD) is a complex neurodegenerative disorder in which most risk variants are noncoding and are enriched at gene regulatory regions, implicating epigenetic mechanisms as central mediators of disease pathogenesis. For most of the history of AD epigenetics research, bulk tissue analysis has dominated, obscuring the fundamentally distinct epigenomic landscapes of individual brain cell types and masking cell-type-specific contributions to disease. Advances in single-cell and single-nucleus sequencing, fluorescence-activated nuclei sorting and multiplexed epigenomic platforms have transformed this landscape, enabling cell-type-resolved profiling of chromatin accessibility, DNA methylation, histone modifications and transcription across the major neuronal, glial and neurovascular populations of the human brain. Here, we review these advances, structured around the argument that cell-type resolution is not a methodological refinement but a conceptual necessity. We describe the distinct epigenomic programs disrupted in neurons, microglia, astrocytes, oligodendrocytes and neurovascular cells in AD, highlighting how each cell type responds to pathology. We discuss the discovery of epigenomic erosion, the progressive loss of cell-type-specific epigenomic identity across virtually all brain cell populations as AD advances, as a unifying disease mechanism linking chromatin dysregulation to cognitive decline. Finally, we identify critical gaps in current knowledge, including the near-complete absence of cell-type-resolved histone modification and DNA methylation data for most brain cell types, the underrepresentation of rare populations in standard preparations and the untapped potential of metabolic acylation marks as indicators of the epigenome-metabolism interface in neurodegeneration.

Humans

Identifying independent causal cell types for human diseases and risk variants.

The SNP-heritability of human diseases is extremely enriched in candidate regulatory elements (cREs) from disease-relevant cell types. Critical next steps are to understand whether these enrichments are driven by multiple causal cell types and whether individual variants impact disease risk via a single or multiple of cell types. Here, we propose CT-FM and CT-FM-SNP, 2 methods accounting for cREs shared across cell types to identify independent sets of causal cell types for a trait and its candidate causal variants, respectively. We applied CT-FM to 63 GWAS summary statistics (average N = 417K) using 924 cRE annotations, primarily from ENCODE4. CT-FM inferred 79 sets of causal cell types, with corresponding SNP-annotations explaining 39.0 ± 1.8% of trait SNP-heritability. It identified 14 traits with independent causal cell types, uncovering previously unexplored cellular mechanisms in height, schizophrenia and autoimmune diseases. We applied CT-FM-SNP to 39 UK Biobank traits and predicted high-confidence causal cell types for 3,091 candidate causal non-coding SNPs-trait pairs. Our results suggest that most SNPs affect a phenotype via a single set of cell types, whereas pleiotropic SNPs might target different cell types depending on the phenotype context. Altogether, CT-FM and CT-FM-SNP shed light on how genetic variants act collectively and individually at the cellular level to affect disease risk.

Journal Article

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis

Enhancing and accelerating cell type deconvolution of large-scale spatial transcriptomics slices with dual network model.

MOTIVATION: Cell type deconvolution deciphers spatial distribution of mRNA transcripts at single cell level by integrating single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data to infer mixture of cell types of spots in slices. Current algorithms are criticized for neglecting connection between scRNA-seq and spatial transcriptomics data, as well as time-consuming, hampering their application to large-scale datasets. RESULTS: In this study, we propose a joint learning nonnegative matrix factorization algorithm for fast cell type deconvolution (aka jMF2D), which integrates scRNA-seq and spatial transcriptomics data with network models. To bridge scRNA-seq and spatial transcriptomics data, jMF2D jointly learns cell type similarity network to enhance quality of signatures of cell types, thereby promoting accuracy and efficiency of deconvolution. Experiments demonstrate that jMF2D outperforms state-of-the-art baselines in terms of accuracy by saving about 90% running time on various datasets generated by different platforms. Furthermore, it can also facilitates the identification of spatial domains and bio-marker genes, providing an efficient and effective model for analyzing spatial transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The software is coded using python, and is free available for academic https://github.com/xkmaxidian/jMF2D.

Algorithms

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

ScGOclust: leveraging gene ontology to find functionally analogous cell types between distant species.

MOTIVATION: Basic biological processes are shared across animal species, yet their cellular mechanisms are profoundly diverse. Comparing cell-type gene expression between species reveals conserved and divergent cellular functions. However, as phylogenetic distance increases, gene-based comparisons become less informative. The gene ontology (GO) knowledgebase offers a solution by serving as the most comprehensive resource of gene functions across a vast diversity of species, providing a bridge for distant species comparisons. RESULTS: Here, we present scGOclust, a computational tool that constructs de novo cellular functional profiles using GO terms, facilitating systematic and robust comparisons within and across species. We applied scGOclust to analyse and compare the heart, gut, and kidney between mouse and fly, and whole-body data from Caenorhabditis elegans and Hydra vulgaris. We show that scGOclust effectively recapitulates the function spectrum of different cell types, characterizes functional similarities between homologous cell types, and reveals functional convergence between unrelated cell types. Additionally, we identified subpopulations within the fly crop that show circadian rhythm-regulated secretory properties and hypothesize an analogy between fly principal cells from different segments and distinct mouse kidney tubules. We envision scGOclust as an effective tool for uncovering functionally analogous cell types or organs across distant species, offering fresh perspectives on evolutionary and functional biology. AVAILABILITY AND IMPLEMENTATION: ScGOclust is publicly available on CRAN: https://cran.r-project.org/web/packages/scGOclust/index.html and development versions are available on GitHub: github.com/Papatheodorou-Group/scGOclust/.

Animals