Search PubMedSearch

SEARCH · Search PubMed

Results for “unsupervised clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Protein Profiling Identifies Biomarkers for Predicting Disease Severity in Anti-NMDAR Encephalitis.

Anti-N-methyl-D-aspartate receptor (NMDAR) encephalitis is a severe autoimmune neurological disorder characterized by pathogenic antibodies against the NMDAR. A systematic protein profiling approach is warranted to identify biomarkers capable of predicting disease status. An Olink proximity extension assay (PEA) profiled 91 inflammation-related proteins from anti-NMDAR encephalitis patients. Disease severity or prognosis were assessed by CASE score or mRS score at 6-month follow-up. Patients were stratified into distinct molecular clusters using unsupervised clustering. Logistic regression models incorporating selected biomarkers were developed to predict disease severity and prognosis, followed by absolute quantification using ELISA. Patients were classified into four consensus clusters. Clusters 1 and 2 corresponded to the mild group, while Cluster 3 represented the severe group, consistent with CASE score above 6. Cluster 4 showed heterogeneous clinical features. Elevated serum levels of IL-10, IL-6, and SIRT2, as well as increased CSF levels of CXCL10, CXCL11, and MMP10, were positively associated with severe disease. Conversely, several proteins including LTA and CCL11, CCL8, TGFB1, CXCL6 were associated with severe disease or unfavorable 6-month outcomes. A logistic regression model combining serum CXCL6 and CCL11 with CSF MMP10 achieved an area under the curve (AUC) of 0.95 for predicting disease severity. Serum CCL11 alone showed predictive value for 6-month prognosis, with an AUC of 0.79. These findings delineate distinct protein signatures associated with clinical heterogeneity of anti-NMDAR encephalitis. Prediction models incorporating multiple biomarkers may provide an approach for disease severity stratification and prognosis forecast.

Humans

Unsupervised multiscale clustering of single-cell transcriptomes to identify hierarchical structures of cell subtypes.

BACKGROUND: Cell clustering is an essential step in uncovering cellular architectures in single-cell RNA sequencing (scRNA-seq) data. However, the existing cell clustering approaches are not well designed to dissect complex structures of cellular landscapes at a finer resolution. RESULTS: Here, we develop a multiscale clustering (MSC) approach to construct a sparse cell-cell correlation network for unsupervised identification of de novo cell types and subtypes across multiple resolutions. Based upon simulated silver- and gold-standard data as well as real scRNA-seq data in diseases, MSC demonstrates significantly improved performance compared to established benchmark methods and reveals a biologically meaningful cell hierarchy to facilitate the discovery of novel disease-associated cell subtypes and mechanisms. CONCLUSIONS: We present MSC as a new single-cell multiscale clustering framework as a powerful tool for advancing discoveries in disease-associated cell populations using single-cell sequencing data.

Single-Cell Analysis

Unsupervised characterization of 100,272 EHR patients identifies high-risk groups and comorbidities linked to premature aging.

Electronic health records (EHRs) contain extensive multidimensional patient data, presenting challenges for the discovery of novel and meaningful clinical patterns. Unsupervised clustering of high-dimensional clinical data holds great potential for identifying novel clinical patterns. Here, we performed unsupervised clustering and characterized 100,272 patients in the Electronic Medical Records and GEnomics (eMERGE) Network. We identified 70 clusters defined by distinct comorbidity patterns. Meanwhile, age and sex are also strongly associated with patient stratification, influencing phenotype prevalence and onset time. Notably, phenotype onset time accurately predicted chronological age and was significantly associated with overall mortality risk. Besides age and sex, we assessed the contribution of genetic variation to phenotype development and observed evidence of cross-phenotype associations influencing cluster membership and comorbidity patterns. However, the role of genetics recedes during aging. We also identified several high-risk clusters with elevated Charlson Comorbidity Index (CCI) scores and validated these findings in an independent cohort. Further analysis of these clusters revealed phenotypes linked to premature aging and highlighted a survival selection among older participants in observational studies. Overall, this study enables phenome-wide unsupervised patient stratification for multimorbidity discovery in largely unannotated clinical data, offering valuable insights into patient stratification, comorbidity analysis, aging, and health outcomes.

Journal Article

Notch pathway defines an aggressive and immune-suppressive phenotype associated with checkpoint inhibitor resistance in pan-gastrointestinal adenocarcinomas.

The Notch pathway regulates the homeostasis and tumorigenesis of gastrointestinal epithelium. Given its roles in cancer stem cell capacity and cancer immunity, we hypothesized that Notch activation can predict poor prognosis and resistance to immune checkpoint inhibitors (ICIs) in gastrointestinal adenocarcinoma (GIAC). The mRNA expression and genomic alterations of Notch pathway were characterized in esophagus (ESAD), stomach (STAD), colon (COAD), or rectum (READ) adenocarcinomas from The Cancer Genome Atlas (TCGA) dataset. The prognostic model (mRNA-score) was constructed using the TCGA dataset (the training set) and was validated in 3 independent sets (GSE19417 [ESAD], GSE84437 [STAD], and GSE40967 [COAD]). The associations of the mRNA-score with drug sensitivity, immune cell infiltration, and immunotherapy efficacy were, respectively, analyzed using the Genomics of Drug Sensitivity in Cancer (GDSC) database, the TCGA dataset, and multiple clinical cohorts including GSE165252, PRJEB25780, IMvigor210, and CheckMate-009/010/025. Notch pathway genes exhibited conserved genomic/transcriptomic features across four GIAC subtypes. Three pan-GIAC clusters were determined by unsupervised clustering, and the cluster with higher expression of the Notch pathway genes had shorter overall survival (OS), immunosuppressive microenvironment, and higher scores of the signatures concerning angiogenesis, cell cycle, PI3K-AKT-mTOR, TGF-&#x3b2;, glycolysis, etc. A prognostic algorithm (mRNA-score) was constructed, which was correlated with poor OS in the training set (TCGA, P&#x2009;<&#x2009;0.001) and three validation sets (GSE19417, P&#x2009;=&#x2009;0.025; GSE84437, P&#x2009;=&#x2009;0.001, GSE40967, P&#x2009;=&#x2009;0.007). A high mRNA-score was linked with more "resting"/ "anti-inflammatory" rather than "activated"/ "pro-inflammatory" tumor-infiltrating immune cells and ICI resistance in GIACs (GSE165252, P&#x2009;=&#x2009;0.047; PRJEB25780, P&#x2009;=&#x2009;0.047) and other solid tumors such as urothelial carcinoma and clear cell renal cell carcinoma. Our findings demonstrate the utility of the Notch pathway in predicting prognosis and ICI resistance. Further studies are warranted to explore the efficacy of Notch inhibitors as immunotherapeutic adjuvants to overcome ICI resistance.

Humans

Beyond enrichment: pharmacogenetic heterogeneity in treatment-resistant depression.

OBJECTIVES: Genetic variation has been proposed as a potential contributor to antidepressant nonresponse, but its role in treatment-resistant depression (TRD) remains unclear. This study used pharmacogenetics (PGx) to characterize genetic variation in TRD and determine whether actionable PGx variation and drug-gene interaction (DGI) mismatch were associated with antidepressant nonresponse and TRD burden. METHODS: This observational study included 158 individuals with TRD recruited from outpatient clinics in Western Australia. Genotype and genotype-predicted phenotypes for CYP2B6, CYP2C19, and CYP2D6 were derived from commercial PGx testing and compared with ethnicity-matched reference populations from ClinPGx. Antidepressant-specific DGIs were classified as actionable or nonactionable according to Clinical Pharmacogenetics Implementation Consortium guidelines, and unsupervised clustering was used to identify clusters based on these actionability profiles. Analyses were performed to determine if actionable PGx variation, cluster membership, or PGx mismatch was associated with TRD burden (number of failed antidepressant trials). RESULTS: PGx variation in the TRD cohort was consistent with population expectations, with no evidence of enrichment for actionable PGx variants. Clustering identified six clusters with distinct and gene-specific patterns of PGx variation independent of demographic and clinical characteristics. However, neither PGx mismatch nor cluster membership were associated with TRD burden. CONCLUSION: These findings suggest that actionable PGx phenotypes are neither enriched in TRD nor associated with greater TRD severity. Rather, the results indicate that TRD does not represent a single, unified PGx-predicted 'poor pharmacological responder' phenotype but instead reflects a biologically heterogeneous collection of distinct PGx profiles.

antidepressants

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Integrated Genomic and Proteomic Analysis Reveals T-B Lymphocyte Signatures in the MYCN Driven "Immune Desert" of Specific Neuroblastoma Subtypes.

AIMS: This study aims to systematically dissect how MYCN amplification shapes the immunosuppressive tumor microenvironment (TME) in high-risk neuroblastoma, elucidating key mechanisms underlying immune evasion. METHODS: We performed an integrated multi-omics analysis of bulk RNA-seq (n&#x2009;=&#x2009;721), single-cell RNA-seq (n&#x2009;=&#x2009;9), proteomic data (n&#x2009;=&#x2009;49) and spatial transcriptomics (Visium, with external validation in melanoma). Analyses included unsupervised clustering, cell-cell communication inference, transcriptional regulatory network reconstruction, and spatial proximity assessment to map the immune landscape. RESULTS: A distinct molecular subtype (Class C), defined by MYCN amplification and poor prognosis, exhibited a comprehensive "immune desert" phenotype characterized by low immune scores and minimal leukocyte infiltration. Single-cell analysis confirmed significant depletion of T and B lymphocytes within the Class C TME. Dysregulated transcriptional networks were identified, including upregulation of REL and EOMES in T cells-with EOMES potentially driving exhaustion via regulation of Transient Receptor Potential (TRP) genes, and REL inhibition enhancing cytotoxic function in&#xa0;vitro. A unique immunosuppressive B-cell subset (B7) engaged in enhanced crosstalk with exhausted T cells and harbored a MYC-centered network linked to cell cycle dysregulation and poor survival. Spatial transcriptomics revealed significant proximity between B7-active regions and Treg/exhaustion-enriched areas, externally validated in melanoma. Proteomic data validated elevated REL expression in MYCN-amplified tumors. CONCLUSION: This work delineates the immunosuppressive architecture of MYCN-driven neuroblastoma, revealing novel regulatory nodes within specific lymphocyte compartments. Integrating single-cell, spatial, and proteomic evidence, we propose REL inhibition as a therapeutic candidate, the EOMES/TRP axis as a bioinformatically supported hypothesis, and the B7/MYC hub as a hypothesis supported by transcriptomic and spatial evidence.

Humans

Integrated phytochemical and bioactivity profiling of Xanthium strumarium fruits from Korea and China: Implications for origin-specific quality specification.

BACKGROUND: Geographic origin influences the phytochemical composition and biological activities of medicinal plant resources. Xanthium strumarium L. (XS) fruit is widely used in East Asian traditional medicine. However, current pharmacopeial standards primarily recognize Chinese-derived material, despite the availability and traditional use of XS in Korea. To address this gap and support origin-informed quality specification, we compared fruits from Korea (XS-K) and China (XS-C) using chloroplast genome sequencing, targeted phytochemical profiling (high-performance liquid chromatography (HPLC) for selected phenolics and gas chromatography-flame ionization detection (GC-FID) for fatty acids and phytosterols, and multivariate chemometric analysis. RESULTS: Chloroplast genome analysis revealed high overall similarity but localized divergence around the rpoC2 locus and a greater mutation burden in XS-C, supporting origin-associated genomic differentiation. Phytochemical profiling revealed distinct origin-dependent metabolic signatures. XS-K showed higher levels of phytosterols, chlorogenic acid, 4,5-dicaffeoylquinic acid (4,5-DCQ), and xanthatin was detected only in XS-K, whereas XS-C exhibited greater abundance of total fatty acids, particularly oleic acid. Unsupervised clustering and log2 fold-change ranking confirmed clear compositional separation, and variable importance in projection (VIP) analysis identified chlorogenic acid, &#x3b2;-sitosterol, oleic acid, 4,5-DCQ, and xanthatin as major discriminators between origins. Bioactivity assays demonstrated that XS-K exerted stronger antioxidant effects in ABTS, DPPH and FRAP assays, stronger skin-related enzyme inhibition, and greater antibacterial activity against Staphylococcus aureus, consistent with its enriched phenolic and sterol profile. CONCLUSION: Together, chloroplast sequence variation, targeted metabolite quantification, and screening bioassays consistently distinguished XS-K from XS-C. These findings support the use of candidate markers for the origin-based authentication and quality control of XS fruit-derived ingredients. &#xa9; 2026 The Author(s). Journal of the Science of Food and Agriculture published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.

Fruit

Single-cell transcriptomic landscape of the southern green stink bug (Nezara viridula) midgut.

BACKGROUND: The southern green stink bug (SGSB), Nezara viridula, is a globally distributed hemipteran pest that damages many economically important crops. Its midgut supports digestion, defense, symbiosis, and interactions with orally delivered control agents, yet the cellular composition of this tissue remains poorly characterized. We therefore developed a single-cell transcriptomic atlas of the N. viridula midgut. RESULTS: Single-cell RNA sequencing of two biological replicates yielded a quality-filtered data set of 13,763 cells. Unsupervised clustering identified 12 transcriptionally distinct populations with putative annotations, including a stem cell/enteroblast (SC/EB)-like population, seven enterocyte-related populations, goblet-like cells, enteroendocrine cells, visceral muscle cells, and an extracellular-matrix-associated epithelial population. Enterocyte-related populations accounted for more than 77% of recovered cells. Putative annotations were assigned primarily from marker gene enrichment and homology to markers reported in other insects. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes analyses identified population-associated functional enrichment patterns, and pseudotime analysis suggested transcriptional relationships between the SC/EB-like population and several enterocyte- and secretory-associated populations without establishing developmental lineages. Immune- and defense-associated transcripts were preferentially enriched in the pEC2 population, and genes associated with symbiont recognition, insecticide action, xenobiotic transport, and orally delivered double-stranded RNA showed population-biased expression. Descriptive comparisons with published insect midgut data sets identified shared and data-set-specific patterns among annotated populations. CONCLUSION: This atlas provides the first single-cell transcriptomic resource for a stink bug midgut and establishes a descriptive cellular framework for SGSB midgut biology. The dataset prioritizes candidate genes and cell populations for future spatial validation, functional testing, and studies of hemipteran midgut physiology, symbiosis, immunity, and pest-management-relevant traits. &#xa9; 2026 Society of Chemical Industry.

Nezara viridula

An immune exhaustion signature predicts prognosis and identifies patients with diffuse large B-cell lymphoma (DLBCL) who derive preferential benefit from chimeric antigen receptor (CAR)-T cell therapy.

BACKGROUND: The tumor microenvironment (TME) is a key determinant of prognosis in diffuse large B-cell lymphoma (DLBCL). While T-cell exhaustion is implicated in therapeutic failure, its precise molecular hallmarks and utility for predicting response to modern immunotherapies, such as chimeric antigen receptor (CAR)-T cell therapy, remain unclear. METHODS: We performed an integrative analysis of transcriptomic and clinical data from multiple DLBCL cohorts (The Cancer Genome Atlas [TCGA], GSE181063, GSE10846, GSE248835, GSE182434). We used unsupervised clustering, exploratory analysis of single-cell RNA sequencing data, and the least absolute shrinkage and selection operator for variable selection (LASSO-Cox) regression to characterize the exhausted TME, construct a prognostic model, and evaluate its predictive value for CAR-T cell therapy. The model's dynamic behavior was assessed in a proof-of-concept longitudinal cohort of patients treated with the T-cell-engaging bispecific antibody glofitamab. RESULTS: We identified a "high-exhaustion" subtype associated with significantly poorer overall survival (OS; log-rank P = 0.016). Based on this, we developed a five-gene immune exhaustion-Related Prognostic Score (IERPS) that served as a robust independent predictor of poor OS across multiple cohorts. Critically, in a cohort of 256 relapsed/refractory patients, the IERPS was strongly prognostic for event-free survival (EFS) in the standard-of-care (SOC) arm (HR = 2.02, 95% confidence interval [95% CI]: 1.07-3.81, P = 0.029) but lost prognostic significance in the CAR-T arm (HR = 0.70, 95 % CI: 0.35-1.40, P = 0.314). This significant interaction suggests that CAR-T cell therapy may abrogate the poor prognosis associated with a high IERPS. Biologically, exploratory single-cell analysis (n = 4 samples) defined the high-IERPS state by hallmarks of classical T-cell exhaustion, and a descriptive case study showed the score dynamically tracked clinical response to glofitamab. CONCLUSIONS: A state of active T-cell exhaustion and a suppressive TME drive the adverse immune phenotype in DLBCL. Our IERPS model captures this dysfunctional state, acting as a powerful prognostic tool and, more importantly, as a potential predictive biomarker to identify high-risk patients who appear to overcome their inherently poor prognosis through CAR-T cell therapy.

Biomarkers

Integrated single-cell and bulk transcriptomic analysis identifies a novel senescent fibroblast subtype associated with poor prognosis in acral melanoma.

BACKGROUND: Acral melanoma (AM) exhibits significant intratumoral heterogeneity, but its tumor microenvironment (TME) and immune regulation remain unclear. This study aims to dissect TME heterogeneity and establish a prognostic model based on key cell subpopulations. METHODS: We collected AM single-cell RNA sequencing (scRNA-seq) and bulk RNA-seq data from the Gene Expression Omnibus (GEO) and the Cancer Genome Atlas (TCGA). Unsupervised clustering, CellChat, and Scissor analysis were performed to characterize cellular heterogeneity, cell-cell communication, and prognosis-related cell subpopulations. Kaplan-Meier analysis was used to assess the prognostic value of key genes, which were further validated by multiplex immunohistochemistry (mIHC). RESULTS: In AM, Mel_C2, C7, and C9 with high SEMA6A and KIT expression were strongly linked to poor prognosis. We further identified a senescent fibroblast subpopulation (sCAF_CDKN2A) characterized by high fibroblast senescence signature (FSS) scores. Integrating Scissor analysis of fibroblast subtypes with bulk prognostic data, we identified COL3A1, VCAN, and KIT as prognosis-associated genes upregulated in poor-outcome-related fibroblast subsets. Cell-cell communication analysis revealed that sCAF_CDKN2A engages in an immunosuppressive network, interacting with regulatory T cells (Tregs) via MIF signaling and receiving signals from exhausted CD8+ T cells through PPIA-BSG interactions. Using transcription factor expression patterns from these fibroblast subtypes, we constructed a prognostic model that effectively stratified patients into distinct risk groups with significant differences in overall survival (OS). mIHC confirmed significantly higher protein levels of SEMA6A and COL3A1 in tumor tissues compared to matched normal tissues. CONCLUSIONS: We established a novel prognostic model for AM and identified sCAF_CDKN2A as an immunosuppressive senescent fibroblast subpopulation driving poor prognosis.

Acral melanoma

Methylation profiling of normal tissue adjacent to breast tumors reveals two distinct groups with divergent tumor microenvironment features.

We previously identified diverse genetic evolutionary patterns in whole-genome sequencing of paired normal tissue adjacent to tumor (NAT) and tumor tissues from Hong Kong breast cancer (HKBC) patients. Here, we investigated whether DNA methylation (DNAm) contributes to NAT heterogeneity and shapes the tumor microenvironment (TME). Genome-wide DNAm profiling was performed on paired NAT and tumor tissues from 188 HKBC patients using the Infinium 850&#x2009;K array. RNA-seq data were available for 76 NATs and 177 tumors. Cellular composition was inferred using MethylCIBERSORT, CIBERSORTx, and EpiDISH, and histopathologic features were assessed on 115 H&E-stained sections. Unsupervised clustering identified two distinct NAT subtypes with divergent TME characteristics. Cluster 1 (N&#x2009;=&#x2009;139) showed higher epithelial and fibroblast content and enrichment of estrogen response pathways. Cluster 2 (N&#x2009;=&#x2009;49) exhibited an immune-metabolic phenotype characterized by increased fat and immune cells, stromal disruption, inflammatory pathway activation, and greater macrophage infiltration. Cluster 2 patients also demonstrated significantly younger epigenetic age estimated using multiple epigenetic clocks. These DNAm-defined NAT subtypes and associated TME features were validated in 97 NAT samples from TCGA breast cancer patients. Overall, our findings identify DNAm-driven NAT heterogeneity with distinct TME landscapes, providing new insights into field cancerization and tumor evolution in breast cancer.

Journal Article

Quantitative essentiality in a reduced genome: a functional, regulatory and structural fitness map.

Essentiality studies have traditionally focused on coding regions, often overlooking other small genetic regulatory elements. To address this, we combined transposon libraries containing promoter or terminator sequences to obtain a high-resolution essentiality map of a genome-reduced bacterium, at near-single-nucleotide precision when considering non-essential genes. By integrating temporal transposon-sequencing data by k-means unsupervised clustering, we present a novel essentiality assessment approach, providing dynamic and quantitative information on the fitness contribution of different genomic regions. We compared the insertion tolerance and persistence of the two engineered libraries, assessing the local impact of transcription and termination on cell fitness. Essentiality assessment at the local base-level revealed essential protein domains and small genomic regions that are either essential or inaccessible to transposon insertion. We also identified structural regions within essential genes that tolerate transposon disruptions, resulting in functionally split proteins. Overall, this study presents a nuanced view of gene essentiality, shifting from static and binary models to a more accurate perspective. Additionally, it provides valuable insights for genome engineering and enhances our understanding of the biology of genome-reduced cells.

DNA Transposable Elements

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease

PreDigs: A Database of Context-specific Cell Type Markers and Precise Cell Subtypes for Digestive Cell Annotation.

Research on cell type markers helps investigators explore the diverse cellular composition of gastrointestinal tumors, thereby enhancing our understanding of tumor heterogeneity and its impact on disease progression and treatment response. However,&#xa0;the integration of large-scale datasets and the standardization of cell type identification remain challenging. Here, we developed PreDigs, a user-friendly database of predicted signatures for the digestive system, which offers 124 curated single-cell RNA sequencing datasets, covering over 3.4 million cells, all available for download. After unsupervised clustering, we unified the identification and nomenclature of cell subtype labels, constructing a cell ontology tree with 142 cell types across 8 hierarchical levels. Meanwhile, we calculated three different context-specific cell type markers, including "Cell Markers", "Subtype Markers", and "TPN Markers", based on various application requirements within or across tissues. Through the integrated analysis of PreDigs data, we identified distinct cell subpopulations exclusive to tumors, one of which corresponds to tumor-specific endothelial cells. Additionally, PreDigs offers online cell annotation tools, allowing users to classify single cells with greater flexibility. PreDigs is accessible at https://www.biosino.org/predigs/.

Humans

DNA copy number patterns reveal prognostic markers and elucidate mechanisms of evolution in IDH-mutant astrocytoma.

BACKGROUND: Current literature suggestsisocitrate dehydrogenase (IDH)-mutant astrocytoma contains several molecular subgroups. In this study, we are interested in determining the connection between different molecular subgroups with grade and/or survival. METHODS: A cohort of 470 Mayo Clinic adult patients (&#x2265;18 years, 56.2% male) with primary IDH-mutant astrocytoma diagnosed by World Health Organization (WHO) 2021 criteria were examined. Results were validated in an independent cohort of 614 Mayo Clinic Neuropathology consult patients and 235 The Cancer Genome Atlas (TCGA) patients. RESULTS: The Mayo Clinic Practice cohort confirmed the association of CDKN2A/B deletion with overall survival (OS, homozygous vs hemizygous vs intact, 2.7 vs 9.6 vs 17.2 years, P&#x2009;<&#x2009;.001). Phosphatase and tensin homolog (PTEN) deletion was also associated with poor OS (7.3 vs 17.4 years, P&#x2009;<&#x2009;.001). Increased number of copy number alterations was associated with OS (continuous variable, HR&#x2009;=&#x2009;1.027, P&#x2009;<&#x2009;.001). Carrying one or more copies of the germline risk allele at rs55705857 was associated with earlier age of onset (median age 33 vs 35 years, P&#x2009;=&#x2009;.01), and a shorter OS after adjusting for age, grade, sex and treatment (HR&#x2009;=&#x2009;1.81, P&#x2009;=&#x2009;.007). The Mayo Clinic Neuropathology Consult cohort and TCGA were utilized to validate age of onset and survival, respectively. Unsupervised clustering of the copy number alterations identified several clinically significant groups that may define pathways to disease progression. Losses of chromosomes 11p, 13q, 1p, and 10q were all associated with reduced overall survival in the Mayo Clinic cohort. CONCLUSIONS: Patients with hemizygous loss of CDKN2A/B, loss of PTEN, increased number of copy number alterations, specific chromosomal arm losses or rs55705857 germline risk allele have reduced overall survival.

Humans

Metabolism pathway-based subtyping in pancreatic adenocarcinoma: an integrated study by bulk RNA-sequence and machine learning algorithms.

BACKGROUND: Pancreatic adenocarcinoma (PAAD) is highly aggressive, and its tumor microenvironment has significant metabolic and immune microenvironment complexity and genomic instability. In this study, by integrating the metabolic pathway activity score and clinical data, we constructed a novel risk assessment model to reveal the unique biological behavior and clinical significance behind different PAAD subtypes. METHODS: In this study, the transcriptome and clinical data of TCGA and GSE57495 databases were integrated to explore the interaction between metabolic pathways. Based on unsupervised clustering analysis of pathway activity and survival prognosis, patients with PAAD were classified into metabolic subtypes with significant prognostic differences. Subsequently, we assessed the heterogeneity of these subtypes in terms of clinical outcomes, genomic characteristics, and immune microenvironment composition. Based on the differentially expressed genes (DEGs) among metabolic subtypes, a clinical prognostic risk model and nomogram were constructed, which were double-validated by GSE57495-independent cohort and GSE57495&#xa0;+&#xa0;TCGA-PAAD combined cohort. Finally, the correlations between risk scores (RSs) and signaling pathway activity and tumor immune microenvironment characteristics were evaluated. RESULTS: Based on metabolic pathway correlation and prognostic information, 240 patients in the TCGA-PAAD and GSE57495 datasets were divided into three subgroups. There were significant differences between subgroups in gene expression, pathway activity, clinical prognosis, and immune infiltration characteristics among the subtypes. Using machine learning algorithms, an RS model was constructed from DEGs among the subgroups, with the random forest method showing the best performance. A nomogram integrating the RS and clinical indicators demonstrated excellent predictive accuracy for 1-, 3-, and 5-year survival rates, confirming the RS as an independent prognostic factor. High- and low-risk groups exhibited significant differences in immune infiltration, pathway activity, and gene mutations. Drug sensitivity analysis showed that the high-risk group was more sensitive to AZD6244, ABT737, and other drugs. CONCLUSION: This study stratified patients with PAAD into three subgroups based on metabolic pathways and prognostic information, revealing significant differences in clinical outcomes, immune characteristics, and genetic mutations. The robust RS model developed from these findings demonstrated strong predictive power for patient survival and identified promising therapeutic strategies, providing valuable insights for advancing precision medicine in PAAD.

immune microenvironment

A multifaceted investigation into the impact of m6A methylation-related genes on pancreatic cancer, integrating insights from various databases and foundational experimental research.

BACKGROUND: Despite advances in surgical techniques, immunotherapy, the mortality rate associated with pancreatic cancer (PC) has been on the rise in recent years. Understanding the importance of RNA N6-methyladenosine (m6A) in PC is critical for prognosis, tumor microenvironment, and immunotherapy efficacy. The study aims to identify m6A methylation regulators that play an important role in the development and progression of PC by mining databases. The effect of insulin-like growth factor-binding protein 3 (IGFBP3) on pancreatic tumors was explored, and the related mechanisms were explored. METHODS: We analyzed the expression of m6A regulators in PC by digging deeper into the datasets of The Cancer Genome Atlas and Gene Expression Omnibus (GEO) databases, and analyzed its relationship with the prognosis of patients with PC, looking for m6A methylation regulators that play an important role in the development and progression of PC. Reuse the ConsensusClusterPlus package, Cox analysis, and unsupervised clustering to delineate three distinct m6A clusters - designated as m6A cluster A, m6A cluster B, and m6A cluster C single-sample gene set enrichment analysis, gene set variation analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes (KEGG) analyses evaluated the different pathway roles of these clusters in the development and progression of PC. Finally, the cell lines with IGFBP3 overexpression and knockdown were constructed by lentivirus transfection, the transfection effect was identified by WB, and the effects of IGFBP3 overexpression/knockdown on the survival and growth of PC cell lines were verified by cell cloning experiments and cell counting kit-8 experiments, and the possible related pathways were explored by KEGG. RESULTS: Most m6A regulatory factors are highly expressed in PC, and their high expression is negatively correlated with the prognosis of patients with PC. Furthermore, m6A regulatory factors may influence the occurrence and development of PC through metabolic pathways, stroma activation pathways, immune regulatory processes, and the immune microenvironment. Finally, the overexpression of IGFBP3 promoted the growth of PC cells, and vice versa. CONCLUSIONS: Most m6A regulatory factors are differentially expressed in PC and are associated with the prognosis of patients with PC, potentially influencing the occurrence and development of PC through pathways such as the immune microenvironment. The overexpression of IGFBP3 can promote the growth of PC cells and vice versa.

IGFBP3