Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Cell annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans↗

Regulation of gene expression in RAW 264.7 macrophage cell line by interferon-gamma.

Macrophages play an important role in immune responses and in inflammatory disease states such as atherosclerosis. Interferon-gamma (IFN-gamma) is a major cytokine involved in the activation of macrophages. To elucidate the primary response of various genes and biological pathways regulated by IFN-gamma in macrophage, we analyzed the gene expression profile in RAW 264.7 macrophage cells treated with IFN-gamma for 4h. Microarray analysis revealed that about 400 genes were differentially expressed, of which about 250 genes were up-regulated and 150 were down-regulated. Functional organization of the transcriptome revealed that induced genes are involved in antimicrobial and antiviral responses, antigen presentation, chemokine and cytokine signaling, and inhibition of cell growth. We also found that expression of genes involved in cell-cycle control, DNA repair, and lipid metabolism was suppressed by IFN-gamma. We also identified induction of multiple transcription factors by IFN-gamma in RAW 264.7 cells. Functional annotation of genes regulated by IFN-gamma in RAW 264.7 cells may provide novel insights into the role of macrophages in immunity and in inflammatory disease.

Animals↗

The MYC/TXNIP axis mediates NCL-Suppressed CD8+T cell immune response in lung adenocarcinoma.

BACKGROUND: Lung adenocarcinoma is a deadly malignancy with immune evasion playing a key role in tumor progression. Glucose metabolism is crucial for T cell function, and the nucleolar protein NCL may influence T cell glucose metabolism. This study aims to investigate NCL's role in T cell glucose metabolism and immune evasion by lung adenocarcinoma cells. METHODS: Utilizing single-cell RNA sequencing (scRNA-seq) data from the Gene Expression Omnibus (GEO) and The Cancer Genome Atlas (TCGA), we analyzed cell clustering, annotation, and prognosis. In vitro experiments involved manipulating NCL expression in CD8+ T cells to study immune function and glucose metabolism. In vivo studies using an orthotopic transplant mouse model monitored NCL's impact on CD8+ T cell glucose metabolism and anti-tumor immune function. RESULTS: NCL was associated with T cell dysfunction and glucose metabolism. NCL silencing enhanced CD8+ T cell glucose metabolism, cytotoxicity, and infiltration, while NCL overexpression had the opposite effect. NCL overexpression relieved MYC-mediated transcriptional repression of TXNIP, reducing CD8+ T cell glucose metabolism. In vivo, NCL inhibited CD8+ T cell glucose metabolism through the MYC/TXNIP axis, hindering anti-tumor immune function. CONCLUSIONS: NCL overexpression suppresses CD8+ T cell glucose metabolism and anti-tumor immune function, promoting lung adenocarcinoma progression via the MYC/TXNIP axis.

CD8-Positive T-Lymphocytes↗

MaxComp: Predicting single-cell chromatin compartments from 3D chromosome structures.

The genome is organized into distinct chromatin compartments with at least two main classes, a transcriptionally active A and an inactive B compartment, broadly corresponding to euchromatin and heterochromatin. Chromatin regions within the same compartment preferentially interact with each other over regions in the opposite compartment. A/B compartments are traditionally identified from ensemble Hi-C contact frequency matrices using principal component analysis of their covariance matrices. However, defining compartments at the single-cell level from sparse single-cell Hi-C data is challenging, especially since homologous copies are often not resolved. To address this, we present MaxComp, an unsupervised method, for inferring single-cell A/B compartments based on 3D geometric considerations in single-cell chromosome structures-derived either from multiplexed FISH-omics imaging or 3D structure models derived from Hi-C data. By representing each 3D chromosome structure as an undirected graph with edge-weights encoding structural information, MaxComp reformulates compartment prediction as a variant of the Max-cut problem, solved using semidefinite graph programming (SPD) to optimally partition the graph into two structural compartments. Our results show that the population average of MaxComp single-cell compartment annotations closely matches those derived from ensemble Hi-C principal component analysis, demonstrating that compartmentalization can be recovered from geometric principles alone, using only the 3D coordinates and nuclear microenvironment of chromatin regions. Our approach reveals widespread cell-to-cell variability in compartment organization, with substantial heterogeneity across genomic loci. When applied to multiplexed FISH imaging data, MaxComp also uncovers relationships between compartment annotations and transcriptional activity at the single-cell level. In summary, MaxComp offers a new framework for understanding chromatin compartmentalization in single cells, connecting 3D genome architecture, and transcriptional activity with the cell-to-cell variations of chromatin compartments.

Chromatin↗

DNA methylation landscape of cerebrospinal fluid cells in multiple sclerosis: an epigenome-wide association study.

BACKGROUND: Multiple sclerosis (MS) is a chronic inflammatory disease of the central nervous system in which DNA methylation may link genetic and environmental risk factors. METHODS: We profiled genome-wide DNA methylation in cerebrospinal fluid (CSF) cells from people with MS (pwMS) and matched controls. Differentially methylated positions (DMPs) and regions (DMRs) were integrated with transcriptomic data, T-cell chromatin annotations, and pathway analyses. Protocadherin gamma (PCDHγ) expression was assessed in primary CD4+ T-cell subsets and confirmed by flow cytometry. FINDINGS: We identified 2710 DMPs and 4330 DMRs associating with genes that were enriched in immune signalling, adhesion and migration processes, and were accompanied by corresponding RNA changes. MS-associated methylation changes enriched in the cohesin chromatin-regulation pathway localised to T-cell regulatory regions, and this pathway included multiple protocadherin (PCDH) genes, which displayed consistent methylation and expression changes in CSF cells of pwMS compared to controls. PCDHγ cluster gene expression was detected in CD4+ T-cell subsets, and flow cytometry confirmed PCDHγ protein expression in peripheral blood T cells. Moreover, co-expression analysis suggests a role of PCDH genes in aryl hydrocarbon receptor (AHR) signalling. Protein-level validation showed fewer PCDHγ-positive CD4+ T cells in pwMS and activation-induced PCDHγ upregulation after T-cell stimulation. INTERPRETATION: DNA methylation changes in CSF resident cells reflect dysregulated T cell activation and migration in pwMS and suggest involvement of protocadherin molecules in MS pathogenesis. FUNDING: European Research Council, Swedish Research Council, Swedish Brain Foundation, Swedish MS Foundation, Knut and Alice Wallenberg Foundation, European Union and others.

Humans↗

Automatic discovery of regulatory patterns in promoter regions based on whole cell expression data and functional annotation.

MOTIVATION: The whole genomes submitted to GenBank contain valuable information about the function of genes as well as the upstream sequences and whole cell expression provides valuable information on gene regulation. To utilize these large amounts of data for a biological understanding of the regulation of gene expression, new automatic methods for pattern finding are needed. RESULTS: Two word-analysis algorithms for automatic discovery of regulatory sequence elements have been developed. We show that sequence patterns correlated to whole cell expression data can be found using Kolmogorov-Smirnov tests on the raw data, thereby eliminating the need for clustering co-regulated genes. Regulatory elements have also been identified by systematic calculations of the significance of correlations between words found in the functional annotation of genes and DNA words occurring in their promoter regions. Application of these algorithms to the Saccharomyces cerevisiae genome and publicly available DNA array data sets revealed a highly conserved 9-mer occurring in the upstream regions of genes coding for proteasomal subunits. Several other putative and known regulatory elements were also found. AVAILABILITY: Upon request.

Algorithms↗

Expressed sequence tags from the midgut and an epithelial cell line of Chironomus tentans: annotation, bioinformatic classification of unknown transcripts and analysis of expression levels.

Expressed sequence tags (ESTs) were generated from two Chironomus tentans cDNA libraries, constructed from an embryo epithelial cell line and from larva midgut tissue. 8584 5'-end ESTs were generated and assembled into 3110 tentative unique transcripts, providing the largest contribution of C. tentans sequences to public databases to date. Annotation using Blast gave 1975 (63.5%) transcripts with a significant match in the major gene/protein databases, 1170 with a best match to Anopheles gambiae and 480 to Drosophila melanogaster. 1091 transcripts (35.1%) had no match to any database. Studies of open reading frames suggest that at least 323 of these contain a coding sequence, indicating that a large proportion of the genes in C. tentans belong to previously unknown gene families.

Animals↗

Pfarao: a web application for protein family analysis customized for cytoskeletal and motor proteins (CyMoBase).

BACKGROUND: Annotation of protein sequences of eukaryotic organisms is crucial for the understanding of their function in the cell. Manual annotation is still by far the most accurate way to correctly predict genes. The classification of protein sequences, their phylogenetic relation and the assignment of function involves information from various sources. This often leads to a collection of heterogeneous data, which is hard to track. Cytoskeletal and motor proteins consist of large and diverse superfamilies comprising up to several dozen members per organism. Up to date there is no integrated tool available to assist in the manual large-scale comparative genomic analysis of protein families. DESCRIPTION: Pfarao (Protein Family Application for Retrieval, Analysis and Organisation) is a database driven online working environment for the analysis of manually annotated protein sequences and their relationship. Currently, the system can store and interrelate a wide range of information about protein sequences, species, phylogenetic relations and sequencing projects as well as links to literature and domain predictions. Sequences can be imported from multiple sequence alignments that are generated during the annotation process. A web interface allows to conveniently browse the database and to compile tabular and graphical summaries of its content. CONCLUSION: We implemented a protein sequence-centric web application to store, organize, interrelate, and present heterogeneous data that is generated in manual genome annotation and comparative genomics. The application has been developed for the analysis of cytoskeletal and motor proteins (CyMoBase) but can easily be adapted for any protein.

Amino Acid Sequence↗

scGPA: an LLM-assisted workflow for directional virtual gene perturbation analysis from single-cell transcriptomes.

BACKGROUND: Existing virtual perturbation methods can often infer directional changes by comparing predicted post-perturbation expression profiles with control cells. However, workflows that directly return direction-specific downstream candidate genes together with confidence scores, evidence support and interpretable summaries remain limited. We developed scGPA, an LLM-assisted workflow system for directional single-cell virtual gene perturbation analysis. METHODS: scGPA starts from raw single-cell RNA sequencing data and performs quality control, normalization, dimensionality reduction, clustering and cell-group selection. It then constructs cell-group-specific wild-type regulatory networks using repeated subsampling, principal component regression (PCR)/Ridge-based network inference and CP tensor denoising. Based on these networks, scGPA simulates dose-aware virtual knockdown of the target gene and applies signed perturbation propagation to estimate the magnitude and direction of downstream transcriptional responses. LLM assistance is used for marker-based cell-type annotation, evidence-guided candidate prioritization and user-facing biological summarization. RESULTS: We benchmarked scGPA across five public Perturb-seq datasets and compared its performance with GEARS, scGPT and a random baseline. The overall correct prediction rate of scGPA was 23.0%, exceeding those of GEARS (20.7%), scGPT (15.1%) and the random baseline (13.6%). These results indicate that scGPA achieved a higher correct prediction rate than the two comparator models and the random baseline. We subsequently evaluated scGPA using a public osteosarcoma single-cell dataset and performed qRT-PCR validation in 143B osteosarcoma cells. Among genes with significant experimental changes, scGPA achieved a directional concordance of 76.9%. When all tested downstream genes were counted, 37.0% were directionally correct, 51.9% showed no significant change and 11.1% changed in the opposite direction. CONCLUSIONS: scGPA provides a practical workflow system for predicting and prioritizing direction-specific downstream transcriptional responses after target-gene perturbation. By integrating single-cell regulatory network inference, signed virtual perturbation and LLM-assisted interpretation, scGPA supports target-gene function inference and downstream mechanistic investigation from single-cell transcriptomic data.

Single-Cell Gene Expression Analysis↗

Improved method for predicting linear B-cell epitopes.

BACKGROUND: B-cell epitopes are the sites of molecules that are recognized by antibodies of the immune system. Knowledge of B-cell epitopes may be used in the design of vaccines and diagnostics tests. It is therefore of interest to develop improved methods for predicting B-cell epitopes. In this paper, we describe an improved method for predicting linear B-cell epitopes. RESULTS: In order to do this, three data sets of linear B-cell epitope annotated proteins were constructed. A data set was collected from the literature, another data set was extracted from the AntiJen database and a data sets of epitopes in the proteins of HIV was collected from the Los Alamos HIV database. An unbiased validation of the methods was made by testing on data sets on which they were neither trained nor optimized on. We have measured the performance in a non-parametric way by constructing ROC-curves. CONCLUSION: The best single method for predicting linear B-cell epitopes is the hidden Markov model. Combining the hidden Markov model with one of the best propensity scale methods, we obtained the BepiPred method. When tested on the validation data set this method performs significantly better than any of the other methods tested. The server and data sets are publicly available at http://www.cbs.dtu.dk/services/BepiPred.

Journal Article↗

scSNViz: visualization and analysis of cell-specific expressed SNVs.

MOTIVATION: Accurately characterizing expressed genetic variation at the single-cell level is essential for understanding transcriptional heterogeneity, allelic regulation, and mutational dynamics within complex tissues. However, few tools enable comprehensive visualization and quantitative analysis of expressed variants across individual cells. RESULTS: scSNViz is an R package for the exploration, quantification, and visualization of expressed single-nucleotide variants (SNVs) from cell-barcoded single-cell RNA sequencing (scRNA-seq) data. The software supports estimation of variant allele fractions, clustering of SNV expression profiles, and 2D and 3D visualization of individual SNVs or user-defined SNV groups. Beyond visualization, scSNViz facilitates investigation of cell-, cluster-, or lineage-specific variant expression patterns, as well as allelic dynamics including imprinting, random allele inactivation, and transcriptional bursting. It interoperates seamlessly with established single-cell frameworks-Seurat for clustering, Slingshot for trajectory inference, scType for cell-type annotation, and CopyKat for copy-number profiling-enabling integrative multi-omic analyses of expressed variation. AVAILABILITY AND IMPLEMENTATION: scSNViz is implemented in R and freely available at https://github.com/HorvathLab/scSNViz (DOI: 10.5281/zenodo.17307516). The package includes comprehensive documentation and example workflows designed for users with limited bioinformatics experience.

Software↗

NMPdb: Database of Nuclear Matrix Proteins.

The nuclear matrix (NM) is a structure resulting from the aggregation of proteins and RNA in the nucleus of eukaryotic cells; it is the 'sticky bit' that remains after aggressive DNAse digestion and salt extraction protocols. Owing to the important role of the NM in DNA replication, DNA transcription and RNA splicing, the expression pattern of NM proteins has become an important early indicator for numerous cancers/tumors. Recent descriptions of the NM structure distinguish between a network-like 'internal nuclear matrix' (INM) and a 'nuclear shell' that connects the INM to the inner and outer nuclear membranes. A cautious NM preparation protocol reveals a coat of proteins on top of the INM; these proteins are usually referred to as the 'nuclear matrix-associated proteins'. Here, we describe a new database (NMPdb at http://www.rostlab.org/db/NMPdb/) that currently contains details of 398 NM proteins. We collected these data through a semi-automated analysis of over 3000 scientific articles in PubMed. We could match these 398 proteins to 302 protein sequences in UniProt or GenBank. Our NMPdb repository annotates these links along with the following annotations: organism, cell type, PubMed identifier, sequence-based predictions of structural and functional features and for some entries the explicit sequence segment that is responsible for localization (nuclear matrix targeting signal).

Amino Acid Sequence↗

Molecular processes during fat cell development revealed by gene expression profiling and functional annotation.

BACKGROUND: Large-scale transcription profiling of cell models and model organisms can identify novel molecular components involved in fat cell development. Detailed characterization of the sequences of identified gene products has not been done and global mechanisms have not been investigated. We evaluated the extent to which molecular processes can be revealed by expression profiling and functional annotation of genes that are differentially expressed during fat cell development. RESULTS: Mouse microarrays with more than 27,000 elements were developed, and transcriptional profiles of 3T3-L1 cells (pre-adipocyte cells) were monitored during differentiation. In total, 780 differentially expressed expressed sequence tags (ESTs) were subjected to in-depth bioinformatics analyses. The analysis of 3'-untranslated region sequences from 395 ESTs showed that 71% of the differentially expressed genes could be regulated by microRNAs. A molecular atlas of fat cell development was then constructed by de novo functional annotation on a sequence segment/domain-wise basis of 659 protein sequences, and subsequent mapping onto known pathways, possible cellular roles, and subcellular localizations. Key enzymes in 27 out of 36 investigated metabolic pathways were regulated at the transcriptional level, typically at the rate-limiting steps in these pathways. Also, coexpressed genes rarely shared consensus transcription-factor binding sites, and were typically not clustered in adjacent chromosomal regions, but were instead widely dispersed throughout the genome. CONCLUSIONS: Large-scale transcription profiling in conjunction with sophisticated bioinformatics analyses can provide not only a list of novel players in a particular setting but also a global view on biological processes and molecular networks.

3T3-L1 Cells↗

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans↗

Genome-wide annotation and expression profiling of cell cycle regulatory genes in Chlamydomonas reinhardtii.

Eukaryotic cell cycles are driven by a set of regulators that have undergone lineage-specific gene loss, duplication, or divergence in different taxa. It is not known to what extent these genomic processes contribute to differences in cell cycle regulatory programs and cell division mechanisms among different taxonomic groups. We have undertaken a genome-wide characterization of the cell cycle genes encoded by Chlamydomonas reinhardtii, a unicellular eukaryote that is part of the green algal/land plant clade. Although Chlamydomonas cells divide by a noncanonical mechanism termed multiple fission, the cell cycle regulatory proteins from Chlamydomonas are remarkably similar to those found in higher plants and metazoans, including the proteins of the RB-E2F pathway that are absent in the fungal kingdom. Unlike in higher plants and vertebrates where cell cycle regulatory genes have undergone extensive duplication, most of the cell cycle regulators in Chlamydomonas have not. The relatively small number of cell cycle genes and growing molecular genetic toolkit position Chlamydomonas to become an important model for higher plant and metazoan cell cycles.

Algal Proteins↗

Genomic dimensions deconstruct the clinical heterogeneity of bipolar disorder.

Bipolar disorder's (BD) clinical heterogeneity has an unresolved genetic basis. We meta-analyzed genome-wide association studies (GWAS) of 16 BD subphenotypes in 226,032 individuals from 57 cohorts (38,022 cases); 10 advanced to multivariate and multi-trait analyses. Four factors (compulsive, psychotic, dysregulated, internalizing) explained 82.8% of shared genetic variance. BD1 and BD2 loaded on distinct factors despite a high genetic correlation; 87.0% of common-factor loci were significant in neither subtype. Unipolar mania aligned with psychosis over internalizing, and was distinguishable from BD1, and rapid cycling showed heritable cross-domain liability. We identified 356 risk loci, 158 novel, including the first univariate-GWAS associations for psychosis, unipolar mania, rapid cycling and schizoaffective disorder-and 249 credible genes (89 high-confidence), 12 with approved-drug or clinical-phase annotations. Cell-type association showed a midbrain dopaminergic-GABAergic gradient along the psychotic factor. BD's genetic architecture appears hierarchical-a general liability resolving into dimensions of course and comorbidity, beyond subtypes.

Journal Article↗

dbscATAC: a resource of single-cell super-enhancers/enhancers and gene markers derived from scATAC-seq data.

MOTIVATION: scATAC-seq enables high-resolution mapping of cis-regulatory elements. It has been widely applied to uncover cell-type-specific regulatory networks and complement scRNA-seq analysis in numerous studies. However, a large number of datasets generated by scATAC-seq remain underutilized due to limited exploration of super-enhancers/typical enhancers and gene markers. A comprehensive resource enabling cell-type-specific annotation of cis-regulatory elements and their dynamic enhancer-gene linkages remains an urgent unmet need for scATAC-seq. RESULTS: We present dbscATAC, a specialized single-cell database for annotating super-enhancers, gene markers, and enhancer-gene interactions derived from scATAC-seq data. Using improved machine learning algorithms, we identified 213 835 super-enhancers across 520 tissue/cell types from three species, as well as 347 484 gene markers, 13 470 526 enhancers, and 10 402 346 enhancer-gene interactions derived from 1 668 076 single cells spanning 1028 tissue/cell types in 13 species. An easy-to-use online platform with multiple analytic modules and hierarchical query options was developed for searching, browsing and visualizing single-cell super-enhancers, enhancers, and gene markers. dbscATAC provides a comprehensive resource to facilitate the exploration of enhancer landscapes, gene regulation, and cell-type-specific characteristics in single-cell epigenomics. AVAILABILITY AND IMPLEMENTATION: The database with all the super-enhancer/enhancer annotation data is available at http://singlecelldb.com/dbscATAC/index.php. And the source code of dbscATAC for prediction of SEs, enhancers, and gene markers are available at https://github.com/EvansGao/dbscATAC. The source code, tissue/cell type description, and data summary can be downloaded at DOI: 10.6084/m9.figshare.28706414.scATAC-seq, Database, Super-enhancers/enhancers, Gene markers.

Enhancer Elements, Genetic↗

Association of MPO Expression with the Immune Microenvironment in Breast Cancer: Insights from Bioinformatics and Single-Cell Analyses.

Breast cancer remains a major cause of cancer-related mortality, and exploratory computational workflows can help prioritize immune-associated markers for further investigation. Here, we used the cancer genome atlas breast invasive carcinoma (TCGA-BRCA) bulk transcriptomic data and the public single-cell dataset GSE161529 to examine associations between myeloperoxidase (MPO) expression, clinical outcomes, immune infiltration, methylation, upstream-regulator annotations, single-cell expression patterns, virtual knockdown sensitivity outputs, drug-gene interaction retrieval, and absorption, distribution, metabolism, excretion, and toxicity (ADMET) annotation. MPO expression was lower in breast cancer tissues than in adjacent non-tumor tissues. Higher MPO expression was associated with a longer progression-free interval, whereas its associations with overall survival and disease-specific survival were not statistically significant. Receiver operating characteristic (ROC) analysis suggested tumor-normal separation within the analyzed public dataset, but this should not be interpreted as clinical diagnostic validation. Immune deconvolution and enrichment analyses indicated that MPO expression mainly tracked with immune- and myeloid-related transcriptional features, rather than establishing tumor-intrinsic regulation of the immune microenvironment. At single-cell resolution, the MPO signal was sparse, with only 85 MPO-positive cells detected before k-nearest neighbor (KNN)-based neighborhood expansion. Detectable MPO signal and MPO-associated scores were interpreted cautiously because they may be influenced by sparse expression, cell-type annotation uncertainty, dropout, doublets, or ambient RNA. In silico virtual knockdown suggested candidate immune- and inflammatory-related transcriptional changes, but these results were considered exploratory and require validation. Drug-gene interaction database (DGIdb)-based drug-gene retrieval and ADMET annotation were used only as preliminary chemical annotations and were not interpreted as therapeutic evidence. Overall, this study provides a reproducible in silico workflow for generating hypotheses about MPO-associated immune/myeloid features in breast cancer, which require external cohort validation and experimental confirmation.

Humans↗