Search PubMedSearch

Biomedical subjects

Dijun Chen

Publications and source records attributed to Dijun Chen.

2 recordsLinked to original sources

scPlantLLM: A Foundation Model for Exploring Single-cell Expression Atlases in Plants.

Single-cell RNA sequencing (scRNA-seq) provides unprecedented insights into plant cellular diversity by enabling high-resolution analyses of gene expression at the single-cell level. However, the complexity of scRNA-seq data, including challenges in batch integration, cell type annotation, and gene regulatory network (GRN) inference, demands advanced computational approaches. To address these challenges, we developed scPlantLLM, a Transformer model trained on millions of plant single-cell data points. Using a sequential pretraining strategy incorporating masked language modeling and cell type annotation tasks, scPlantLLM generates robust and interpretable single-cell data embeddings. When applied to Arabidopsis thaliana datasets, scPlantLLM excels in clustering, cell type annotation, and batch integration, achieving an accuracy of up to 0.91 in zero-shot learning scenarios. Furthermore, the model demonstrates an ability to identify biologically meaningful GRNs and subtle cellular subtypes, showcasing its potential to advance plant biology research. Compared to traditional methods, scPlantLLM outperforms in key metrics such as adjusted rand index (ARI), normalized mutual information (NMI), and silhouette score (SIL), highlighting its superior clustering accuracy and biological relevance. scPlantLLM represents a foundation model for exploring plant single-cell expression atlases, offering unprecedented capabilities to resolve cellular heterogeneity and regulatory dynamics across diverse plant systems. The code used in this study is available at https://github.com/compbioNJU/scPlantLLM.

Single-Cell Analysis

Genetic effects on chromatin accessibility reveal the molecular mechanisms of complex traits in maize.

Cis-regulatory elements (CREs) are critical for modulating gene expression and phenotypic diversity in maize. While genome-wide association study (GWAS) hits and expression quantitative trait loci (eQTLs) are often enriched in CREs, their molecular mechanisms remain poorly understood. Characterizing CREs within accessible chromatin regions (ACRs) offers a powerful approach to link noncoding variants to chromatin structure alterations and phenotypic variation. Here, we generated ATAC-seq profiles from seedling leaves of 214 maize inbred lines, identifying 82 174 consensus ACRs. Notably, 39.55% of these ACRs exhibited significant population-wide chromatin accessibility variation. By mapping chromatin accessibility quantitative trait loci (caQTLs), we discovered 27 004 loci, including 1398 predicted to disrupt transcription factor (TF)-binding sites. Integration with multi-omics data revealed 7405 caACR-target gene pairs and linked 56 caACRs to GWAS signals for 51 agronomic traits, with significant enrichment in flowering-related pathways. Functional candidates such as ZmZIM30 - putatively regulated by caACRs - emerged as key regulators of flowering time. At the fad7 locus associated with linolenic acid content, allelic variants overlapping a caQTL showed differential chromatin accessibility. Our study provides a high-resolution cis-elements of maize leaves, deciphers the genetic basis of chromatin accessibility variation, and bridges noncoding caQTLs to molecular mechanisms underlying GWAS hits.

Zea mays