Search PubMedSearch

SEARCH · Search PubMed

Results for “Unsupervised classification”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

13 recordsLinked to original sources

Integrative multi-omics analysis proposes a metabolic classification of gliomas: distinct metabolic states, immune infiltration, and prognosis.

BACKGROUND: The tumor microenvironment (TME) of glioma harbors diverse cell types; however, cell metabolic heterogeneity remains to be explored. This study aims to characterize the metabolic features of different cell types in the TME by integrating multiple datasets, including genomics, bulk and single-cell transcriptomics, and metabolomics. METHODS: Unsupervised machine learning was used to construct an energy metabolic classifier based on the metabolic pathways identified from bulk RNA-seq of gliomas in the TCGA dataset. The classifier was externally validated using multiple datasets, including genomics, bulk RNA-seq, snRNA-seq, and the metabolomics data. Furthermore, metabolic heterogeneity associated with the classifier was further characterized at single-cell resolution. RESULTS: The energy metabolism-based classifier stratified patients into two prognostic clusters: patients in cluster 1 were characterized by high pathway activity of glycolysis, the pentose phosphate pathway (PPP), and fatty acid oxidation (FAO), whereas patients in cluster 2 exhibited higher activity in glutaminolysis. This metabolic classifier revealed both intratumoral and intertumoral metabolic heterogeneity, and the complexity was further validated by the metabolomics profiling and snRNA-seq data from the CPTAC dataset. Notably, OSMR, highly expressed in cluster 1, showed significant co-expression with key glycolytic enzyme genes. The OSM/OSMR/JAK1/STAT3 axis potently drives malignant progression of glioma cells, specially enhancing their invasive and migratory capabilities. Single-cell resolution analyses demonstrated that tumor metabolic heterogeneity is primarily driven by malignant cells rather than non-malignant components, while tumor microenvironment (TME) factors were also found to modulate malignant cell metabolism. Significantly, glycolytic activity in glioma cells increased during the phenotypic transition from PN (proneural) to MES (mesenchymal), with cluster 1 metabolic phenotypes predominating in the tumor core. Compared to cluster 2, cluster 1 patients exhibited higher mRNA expression of immunosuppressive checkpoint genes, which correlated with pronounced immunosuppression in the TME. Furthermore, various immune cells demonstrated distinct metabolic preferences at single-cell resolution. CONCLUSIONS: This study developed an energy metabolic-based classifier for gliomas with prognostic and therapeutic potential. Metabolic reprogramming was linked with the PN-to-MES transition of glioma cells and immunosuppression in the tumor microenvironment. Multi-omics data, especially snRNA-seq, offered insights into metabolism heterogeneity at single-cell resolution, enabling personalized treatment strategies.

Humans

Integrative multi-omics profiling of insomnia-related molecular features reveals microbiome, immune, and therapy-relevant heterogeneity in colorectal cancer.

Emerging evidence implicates insomnia as a potential risk factor in carcinogenesis, potentially involving systemic inflammation, circadian disruption, and microbiome alterations. However, the molecular associations linking insomnia-related features to colorectal cancer (CRC), particularly with respect to tumor biology, immune microenvironmental states, and therapy-relevant phenotypes, remain largely unexplored. Multi-omics integration of genomic, transcriptomic, and microbiome data from 3,026 CRC patients across seven independent cohorts, including a large, well-annotated Clinical Omics study of Colorectal Cancer in China (COCC) cohort, enabled insomnia-based molecular classification through unsupervised non-negative matrix factorization (NMF) clustering. The insomnia subtype (IS) was biologically characterized via pathway enrichment, immune deconvolution, microbial profiling, and single-cell transcriptomics. Furthermore, an insomnia score (ISscore) was developed and validated in multiple cohorts for risk stratification and assessment of treatment-response-related indicators in CRC. Unsupervised clustering revealed two distinct molecular subtypes (IS1/IS2), with IS2 demonstrating significantly poorer survival. IS2 exhibited marked activation of EMT/angiogenesis pathways versus cell cycle activation in IS1. The IS2 microenvironment showed increased immunosuppression-related infiltration and exhausted T cell signatures, together with intratumoral microbiome variation characterized by depletion of Ruminococcaceae UCG-002 and enrichment of Hungatella/Selenomonas. The ISscore system stratified survival risk and was associated with computational indicators of immunotherapy response. Single-cell analysis nominated PPIA-BSG as a potential cell-cell communication signal involving high-ISscore tumor cells, CXCL12+ endothelial cells, and CLEC9A+ dendritic cell subsets. This multi-omics characterization of insomnia-CRC interplay suggests that insomnia-related molecular features are associated with an immunologically distinct and microbiome-altered tumor ecosystem. The ISscore provides a reproducible framework for capturing insomnia-related molecular heterogeneity, supporting risk stratification and future evaluation of therapy-relevant phenotypes.IMPORTANCEChronic insomnia affects millions, but it is not typically considered a cancer risk factor. Our study, analyzing vast biological data from over 3,000 colorectal cancer patients, uncovers a potential link between a person's predisposition to insomnia and their risk of developing this disease. This suggests that the biological pathways related to sleep may play a role in cancer development. Understanding this connection opens up new avenues for identifying individuals at higher risk and developing novel prevention strategies for colorectal cancer.

colorectal cancer

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Neoadjuvant Immunotherapy Promotes the Formation of Mature Tertiary Lymphoid Structures in a Remodeled Pancreatic Tumor Microenvironment.

Pancreatic ductal adenocarcinoma (PDAC) is a rapidly progressing cancer that responds poorly to immunotherapies. Intratumoral tertiary lymphoid structures (TLS) have been associated with rare long-term PDAC survivors, but the role of TLS in PDAC and their spatial relationships within the context of the broader tumor microenvironment remain unknown. In this study, we report the generation of a spatial multiomic atlas of PDAC tumors and tumor-adjacent lymph nodes from patients treated with combination neoadjuvant immunotherapies. Using machine learning-enabled hematoxylin and eosin image classification models, imaging mass cytometry, and unsupervised gene expression matrix factorization methods for spatial transcriptomics, we characterized cellular states within and adjacent to TLS spanning distinct spatial niches and pathologic responses. Unsupervised learning identified TLS-specific spatial gene expression signatures that are significantly associated with improved survival in patients with PDAC. We identified spatial features of pathologic immune responses, including intratumoral TLS-associated B-cell maturation colocalizing with IgG dissemination and extracellular matrix remodeling. Our findings offer insights into the cellular and molecular landscape of TLS in PDACs during immunotherapy treatment.

Humans

Alzheimer's subtypes A supervised, unsupervised, multimodal, multilayered embedded recursive (SUMMER) AI study.

Since Alzheimer's disease (AD) is a heterogeneous disease, different subtypes may have distinct biological, genetic, and clinical characteristics, requiring tailored interventions. While several proposed subtypes of AD exist, there is still no clear consensus on a definitive classification. By leveraging complementary AI approaches, including supervised and unsupervised learning, within a recursive pipeline (SUMMER) that integrates multimodal datasets encompassing MRI measurements, phenotypes, and genetic data, our goal was to generate robust scientific evidence for identifying AD subtypes. Data was downloaded from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database and included neuroimaging data (MRI), genetics (SNPs), clinical diagnosis, and demographics. 1133 European American participants' images, aged 55-95, were included in this study. The analysis was multi-fold, where the first step involved applying an unsupervised application to a subset of the MRI sample (AD + cognitively normal (CN) aged matched groups, 100 men aged 68-85 years, and 76 women aged 68-85 years). The MRI brain gray matter was segmented into 44 regions of interest (ROIs) according to a standard atlas, and 618 features were extracted, including ROI voxel intensity measurements such as minimum, maximum, and histogram variables. Results identified a cluster of subtype AD men and a cluster of subtype AD women that were distinct from the rest of their respective samples. In the next step, the integrity of the identified subtype AD clusters was investigated using the XGBoost supervised machine learning application with genetic features (SNPs, N=36,724) and labels: the identified subtype AD cluster vs. the rest of the sample, stratified by sex. A significant AD subtype men model (accuracy=0.85, F1=0.72, AUC=0.83) and a significant women AD subtype model (accuracy=0.81, F1=0.81, AUC=0.81) were built, confirming the homogeneity of the isolated AD subtype clusters. Discriminative biomarkers were extracted from the significant models, including selected ROIs and SNPs. Finally, the subtype models were tested on an unseen subset of ADNI data. The genetic-based models identified clusters of AD subtype participants consisting of 34% of the men AD group and 47% of the women AD group. Phenotypic analysis indicates that lower body weight was associated with the women's AD subtype. Complex diseases like AD demand a sophisticated, multimodal approach for precise diagnosis. Effectively identifying disease subtypes enhances the potential for personalized treatment, ultimately improving patient outcomes.

Journal Article

N6-methyladenine identification using deep learning and discriminative feature integration.

N6-methyladenine (6 mA) is a pivotal DNA modification that plays a crucial role in epigenetic regulation, gene expression, and various biological processes. With advancements in sequencing technologies and computational biology, there is an increasing focus on developing accurate methods for 6 mA site identification to enhance early detection and understand its biological significance. Despite the rapid progress of machine learning in bioinformatics, accurately detecting 6 mA sites remains a challenge due to the limited generalizability and efficiency of existing approaches. In this study, we present Deep-N6mA, a novel Deep Neural Network (DNN) model incorporating optimal hybrid features for precise 6 mA site identification. The proposed framework captures complex patterns from DNA sequences through a comprehensive feature extraction process, leveraging k-mer, Dinucleotide-based Cross Covariance (DCC), Trinucleotide-based Auto Covariance (TAC), Pseudo Single Nucleotide Composition (PseSNC), Pseudo Dinucleotide Composition (PseDNC), and Pseudo Trinucleotide Composition (PseTNC). To optimize computational efficiency and eliminate irrelevant or noisy features, an unsupervised Principal Component Analysis (PCA) algorithm is employed, ensuring the selection of the most informative features. A multilayer DNN serves as the classification algorithm to identify N6-methyladenine sites accurately. The robustness and generalizability of Deep-N6mA were rigorously validated using fivefold cross-validation on two benchmark datasets. Experimental results reveal that Deep-N6mA achieves an average accuracy of 97.70% on the F. vesca dataset and 95.75% on the R. chinensis dataset, outperforming existing methods by 4.12% and 4.55%, respectively. These findings underscore the effectiveness of Deep-N6mA as a reliable tool for early 6 mA site detection, contributing to epigenetic research and advancing the field of computational biology.

Deep Learning

Integrating Imaging-Derived Clinical Endotypes with Plasma Proteomics and External Polygenic Risk Scores Enhances Coronary Microvascular Disease Risk Prediction.

Coronary microvascular disease (CMVD) is an underdiagnosed but significant contributor to the burden of ischemic heart disease, characterized by angina and myocardial infarction. The development of risk prediction models such as polygenic risk scores (PRS) for CMVD has been limited by a lack of large-scale genome-wide association studies (GWAS). However, there is significant overlap between CMVD and enrollment criteria for coronary artery disease (CAD) GWAS. In this study, we developed CMVD PRS models by selecting variants identified in a CMVD GWAS and applying weights from an external CAD GWAS, using CMVD-associated loci as proxies for the genetic risk. We integrated plasma proteomics, clinical measures from perfusion PET imaging, and PRS to evaluate their contributions to CMVD risk prediction in comprehensive machine and deep learning models. We then developed a novel unsupervised endotyping framework for CMVD from perfusion PET-derived myocardial blood flow data, revealing distinct patient subgroups beyond traditional case-control definitions. This imaging-based stratification substantially improved classification performance alongside plasma proteomics and PRS, achieving AUROCs between 0.65 and 0.73 per class, significantly outperforming binary classifiers and existing clinical models, highlighting the potential of this stratification approach to enable more precise and personalized diagnosis by capturing the underlying heterogeneity of CMVD. This work represents the first application of imaging-based endotyping and the integration of genetic and proteomic data for CMVD risk prediction, establishing a framework for multimodal modeling in complex diseases.

Cardiovascular Disease

Methylome Profiling of Cartilage Tumors: A Promising New Diagnostic Tool?

DNA methylation and copy number variation (CNV) profiling has emerged as a promising tool for the classification of bone and soft tissue tumors. We evaluated its utility in cartilage tumors, where distinguishing low-grade from high-grade conventional central chondrosarcomas (CSs) and atypical cartilaginous tumors (ACTs) from enchondromas (ECs) is a frequent diagnostic challenge, particularly on biopsy material. We analyzed 214 chondrogenic tumors, including ECs, ACTs, conventional CSs, dedifferentiated chondrosarcomas (DDCSs), and clear cell CSs, and determined their IDH1/2 mutation status. Unsupervised dimensionality reduction of genome-wide DNA methylation patterns revealed 4 clusters among IDH-mutant (MUT) tumors (IDH-MUT-1: mostly ECs and ACTs and some high-grade CSs; IDH-MUT-2: predominantly high-grade CSs; IDH-MUT-3: largely DDCSs; and IDH-MUT-SB: distinct skull base group with a markedly different methylation pattern) and 2 clusters among IDH-wild-type (WT) tumors (IDH-WT-1 and IDH-WT-2: both primarily high-grade CSs, with IDH-WT-2 showing higher tumor grade and more extensive CNVs). Clear cell CSs formed a separate cluster. The amount of CNVs, including loss of CDKN2A, increased with tumor grade, reflecting increased genomic instability during chondrosarcoma progression. Supervised classifiers trained separately, both on methylation and CNV data, and distinguished low-grade and high-grade cartilaginous tumors with area under the curve values of 0.87 to 0.97 and 85% to 90% accuracy. Furthermore, we tested whether DDCSs can be distinguished from metastatic carcinomas and other high-grade sarcomas of the bone. Across 246 reference samples, a supervised classifier achieved 97.2% accuracy (area under the curve, 99.8%) and correctly identified 30 of 32 DDCSs (93.8%). These results indicate that DNA methylation and CNV data analysis provide a valuable tool for distinguishing most low- and high-grade CSs, with additional utility also in differentiating DDCS from morphologic mimics.

cartilaginous tumors

Molecular subtyping of adrenocortical carcinoma reveals distinct subtypes with prognostic and therapeutic implications.

Adrenocortical carcinoma (ACC) is a rare but aggressive malignancy with poor survival and limited treatment options. To comprehensively characterize its molecular landscape and identify clinically relevant subtypes, we performed an integrated genomic analysis - including whole-exome sequencing, RNA sequencing, and copy number variation profiling - on 61 Chinese patients with ACC. We identified recurrent mutations in TP53 (25%), CTNNB1 (15%), ZNRF3 (10%), and MEN1 (8%). Unsupervised clustering of transcriptomic data revealed four distinct molecular subtypes: cortisol-driven (CD, 14%), immune-suppressed (IS, 40%), cell cycle-altered (CCA, 22%), and immunomodulatory (IM, 24%). The CD subtype exhibited steroidogenic pathway activation; the IS subtype showed T cell receptor downregulation and the worst disease-free survival; the CCA subtype was marked by chromosomal instability and cell cycle gene overexpression; and the IM subtype displayed enriched immune signaling and favorable outcomes. Copy number analysis further uncovered focal amplifications (e.g. TERT, CDK4) and HLA-II deletions. This study establishes a novel molecular classification of ACC, providing a framework for subtype-specific therapeutic strategies, such as CDK4/6 inhibition for CCA and immunotherapy for IM tumors, while highlighting the clinical challenges of immune-cold IS tumors.

Humans

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans

The machine-learning classifier ALLCatchR2 identifies 20 T-ALL subtypes across cohorts and age groups.

T-cell acute lymphoblastic leukemia (T-ALL) comprises molecularly diverse subtypes, but robust cross-cohort validations and operational gene-expression definitions are lacking. To establish a gene-expression-anchored framework for T-ALL subtyping, we aggregated 2314 transcriptomes (15 cohorts, age: 0.8-90.8 years). An extended unsupervised approach defined 17 main clusters and 3 subclusters in samples with high blast fractions. Supervised analyses added an overarching immature T-ALL (early T cell precursor [ETP]-like) definition and resolved the LMO2 &#x3b3;&#x3b4;-like subtype. All clusters contained samples from at least two cohorts. Characteristic genomic driver enrichments were consistent across cohorts, while gene-expression clusters did not correspond exclusively to single driver events but also reflected developmental origins. A machine-learning classifier based on ALLCatchR, our B-cell acute lymphoblastic leukemia (B-ALL) classifier, identified these 20 transcriptomic subtypes and the immature T-ALL (ETP-like) signature with 0.995-1.0 accuracy in a validation set (n&#x2009;=&#x2009;203). Testing the classifier on a second hold-out data set (n&#x2009;=&#x2009;265 samples) showed that 92.7% of predictions matched with corresponding driver alterations. Across all samples, 83.2% of cases received high-confidence predictions, 7.3% candidate predictions, and 9.5% remained unclassified, largely because of low blast fractions. We identified a novel gene-expression cluster markedly enriched (P&#x2009;<&#x2009;0.001) for clonal hematopoiesis mutations (IDH2 R140Q, DNMT3A) and a stem-/progenitor cell-like gene expression. This novel clonal hematopoiesis-related T-ALL subtype was observed in six cohorts and accounted for 8.9% of adults and 39.5% of patients aged >50 years. We extended&#xa0;ALLCatchR into ALLCatchR2, a free R package that now enables B-/T-lineage separation, gene-expression subtyping, blast estimation, and developmental annotation to harmonize T-ALL classification across studies and clinical contexts.

Journal Article