Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning.”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Transcriptome Analysis, Machine Learning, and Experimental Identification of CDK7 Affecting the Progression of Pregnancy-induced Hypertension by Influencing Macrophage Polarization.

INTRODUCTION: Pregnancy-induced hypertension (PIH) is a severe pregnancy complication characterized by placental insufficiency, abnormal vascular remodeling, and immune dysregulation, but personalized therapeutic markers remain unclear. This study aimed to identify key genes and explore immune mechanisms in PIH using transcriptome analysis, machine learning, and experimental validation. METHODS: We analyzed the GSE204835 transcriptomic dataset to screen differentially expressed genes (DEGs) and performed Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and Gene Set Enrichment Analysis (GSEA) for functional annotation. Immune infiltration analysis was also performed to examine the immune landscape in PIH. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified key genes, which were validated in a PIH cell model. Flow cytometry and immunofluorescence assays assessed the effect of CDK7 knockdown on macrophage polarization. RESULTS: A total of 1,598 DEGs (1,123 upregulated, 475 downregulated) were identified. Enrichment analyses highlighted associations with embryonic organ development, oxidative phosphorylation, angiogenesis, and oxidative stress. Immune infiltration analysis revealed altered eosinophil and macrophage polarization in PIH. LASSO regression selected 12 key genes, with CDK7 showing the most significant upregulation in the PIH model. CDK7 knockdown promoted macrophage polarization toward the anti-inflammatory M2 phenotype. DISCUSSION: These findings link CDK7 to immune dysregulation in PIH by modulating macrophage polarization, expanding our understanding of PIH's molecular mechanisms. The study's limitations include reliance on public datasets and in vitro models, warranting in vivo validation. CONCLUSION: CDK7 emerges as a potential therapeutic target for PIH, offering new insights into immunoregulatory interventions for this complication.

Female

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans

Integrative multi-omics and machine learning identify the SPI1-METTL16-PLIN4 axis as a candidate driver of steatosis in HepG2 cells.

BACKGROUND: Non-alcoholic fatty liver disease (NAFLD) is a prevalent metabolic disorder with limited therapeutic options. This study aimed to identify potential regulators and explore their functional roles in a cellular model of NAFLD. METHODS: WGCNA was performed on the hepatic transcriptomic dataset GSE126848 (31 NAFLD vs. 26 controls), followed by integration with serum proteomic data from 12 NAFLD patients and 12 healthy controls. Hub genes were prioritized using three machine learning algorithms. Functional validation was conducted in a HepG2 cellular steatosis model induced by high fructose (3.2&#x202f;g/L) and oleic acid (400&#x202f;&#x3bc;M) for 48&#x202f;h. Lipid accumulation was assessed by Oil Red O staining and triglyceride/total cholesterol measurement. Inflammation was evaluated by TNF-&#x3b1; and IL-6 secretion (ELISA), and oxidative stress by ROS levels (flow cytometry). The binding interaction between METTL16 and PLIN4 mRNA was validated by RNA immunoprecipitation (RIP)-quantitative PCR. METTL16-mediated m6A modification of PLIN4 was assessed by Methylated RIP (MeRIP)-quantitative PCR. Transcriptional regulation of METTL16 by SPI1 was examined by chromatin immunoprecipitation (ChIP) and dual-luciferase reporter assays. RESULTS: Integrative analysis identified PLIN4 as a core hub gene. PLIN4 was upregulated in the HepG2 steatosis model (P&#x202f;<&#x202f;0.001). PLIN4 knockdown alleviated lipid droplet accumulation (P&#x202f;<&#x202f;0.001), reduced TNF-&#x3b1; and IL-6 secretion (P&#x202f;<&#x202f;0.01), and decreased ROS levels (P&#x202f;<&#x202f;0.001) in fructose/oleic acid-treated HepG2 cells. Mechanistically, METTL16 mediated its m6A modification to enhance PLIN4 mRNA stability. Furthermore, SPI1 was found to transcriptionally activate METTL16 by binding to its promoter (P&#x202f;<&#x202f;0.001). PLIN4 re-expression partially reversed the protective effects of SPI1 knockdown on lipid accumulation (P&#x202f;=&#x202f;0.01), inflammation (P&#x202f;<&#x202f;0.05), and oxidative stress (P&#x202f;<&#x202f;0.001). CONCLUSION: This study identifies the SPI1/METTL16/PLIN4 axis as a potential regulatory mechanism contributing to in vitro steatosis, inflammation, and oxidative stress in steatotic HepG2 cells.

Humans

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I&#xb2; = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8&#x207a; T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

Multi-level Transcriptomic and Machine-learning Analyses Identify MZT1 as a Proliferation-associated Prognostic Marker in Lung Adenocarcinoma.

BACKGROUND/AIM: Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity and variable clinical outcomes, highlighting the need for biomarkers that reflect core tumor biological processes. Centrosome-associated proteins regulate mitotic fidelity and genome stability, yet their roles in LUAD remain incompletely defined. In this study, we systematically characterized mitotic spindle organizing protein 1 (MOZART1; MZT1) and related family members in LUAD. MATERIALS AND METHODS: We performed integrated analyses combining bulk transcriptomic datasets, survival modeling, gene set enrichment, immune deconvolution, machine-learning based prognostic modeling, and single-cell RNA sequencing. Expression patterns and clinical associations of MZT family genes were evaluated across pan-cancer and LUAD cohorts. RESULTS: MZT family genes were consistently upregulated in tumor tissues, with MZT1 showing the most robust expression pattern. Elevated MZT1 expression was significantly associated with reduced overall survival. Functional analyses revealed coordinated activation of proliferative and genome maintenance pathways, including G2/M checkpoint regulation, E2F and MYC signaling, and DNA repair. A multivariable analysis indicated that the prognostic association of MZT1 was reduced after adjusting for canonical proliferation markers, suggesting partial overlap with established proliferation signals. The LASSO-based Cox model demonstrated stable time-dependent predictive performance at 1-, 3-, and 5-year survival. Immune analyses indicated associations between MZT1 expression and tumor microenvironmental features. Single-cell analysis showed that MZT1 expression was predominantly enriched in malignant epithelial cells and associated with proliferative cellular states. Protein-level validation supported concordance with transcriptomic findings. CONCLUSION: MZT1 is a proliferation-associated marker that integrates clinical risk, transcriptional programs, cellular heterogeneity, and predictive modeling in LUAD, providing a potential framework for biomarker development and risk stratification.

Humans

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans

Exploring the Genetic Link between Irritable Bowel Syndrome and Polycystic Ovary Syndrome: Bidirectional Mendelian Randomization and Machine Learning Approaches.

BACKGROUND: Research has shown a certain correlation between polycystic ovary syndrome (PCOS) and irritable bowel syndrome (IBS). The study aims to determine the directionality and underlying biological processes influencing the relationship between these two disorders. METHODS: We explored the causal relationship between IBS and PCOS by conducting a comprehensive bidirectional Mendelian randomization (MR) analysis using five different methods and conducted robustness assessments. We extracted differentially expressed genes from the IBS and PCOS datasets for Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. Additionally, we developed a protein-protein interaction (PPI) network and applied the Least Absolute Shrinkage and Selection Operator (LASSO) and Support Vector Machine (SVM) methodologies to pinpoint key diagnostic markers. Diagnostic efficacy was further assessed through Receiver Operating Characteristic (ROC) curve analysis for selected genes. Finally, single-sample gene set enrichment analysis (ssGSEA) was carried out to examine immune cell infiltration in IBS and PCOS. RESULTS: MR analysis identified a causal effect of PCOS on IBS (IVW, OR = 1.034, 95% CI: 1.003-1.065, P = 0.029). Conversely, no relationship between IBS and PCOS was observed in the reverse analysis. Furthermore, integrative bioinformatics and machine learning analyses identified CD14 and CASP1 as key diagnostic biomarkers for both IBS and PCOS, which were significantly associated with immune cell infiltration. CONCLUSION: MR analysis has demonstrated a significant positive causal relationship between PCOS and IBS, though the reverse causality from IBS to PCOS appeared non-significant. The genes CD14 and CASP1 emerged as potential shared diagnostic markers between these two conditions.

Polycystic Ovary Syndrome

A machine learning approach to identify key epigenetic transcripts for ageing research in human blood (Epitage).

DNA methylation is an established biomarker of human ageing and is used by a variety of tools to identify meaningful epigenetic signals. We investigated whether analysing CpGs grouped by transcript as functional units could generate a ranked list of transcripts most correlated with age that might otherwise be overlooked in genome-wide CpG-based studies. Here we present Epitage ( https://github.com/a00s/epitage ), a continuously updated ranked list of transcripts built from the GSE87571 dataset (714 whole-blood samples, ages 14-94 years) through intensive testing with machine-learning models. To support reproducible analyses, we developed ugPlot ( https://github.com/a00s/ugplot ), an open-source R/Shiny tool with a graphical user interface that automates model training, testing, and comparison. Initially, we identified 48 transcripts across 13 genes, with some transcripts from the genes OBSCN, PRRT1, and SPTBN4 showing better predictive performance when multiple associated CpGs were analysed together rather than individually. In contrast, for the majority of transcripts, a dominant individual CpG still showed a higher Spearman correlation with age, as seen in established ageing genes such as ELOVL2, FHL2, and TRIM59. Epitage is a transcript-ranking list based on the methylation patterns observed in the analysed dataset. It provides a reproducible framework for prioritising transcripts associated with human ageing and for guiding future epigenetic studies.

Humans

Machine learning algorithm-based biomarker exploration and validation of mitochondria-related diagnostic genes in osteoarthritis.

The role of mitochondria in the pathogenesis of osteoarthritis (OA) is significant. In this study, we aimed to identify diagnostic signature genes associated with OA from a set of mitochondria-related genes (MRGs). First, the gene expression profiles of OA cartilage GSE114007 and GSE57218 were obtained from the Gene Expression Omnibus. And the limma method was used to detect differentially expressed genes (DEGs). Second, the biological functions of the DEGs in OA were investigated using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. Wayne plots were employed to visualize the differentially expressed mitochondrial genes (MDEGs) in OA. Subsequently, the LASSO and SVM-RFE algorithms were employed to elucidate potential OA signature genes within the set of MDEGs. As a result, GRPEL and MTFP1 were identified as signature genes. Notably, GRPEL1 exhibited low expression levels in OA samples from both experimental and test group datasets, demonstrating high diagnostic efficacy. Furthermore, RT-qPCR analysis confirmed the reduced expression of Grpel1 in an in vitro OA model. Lastly, ssGSEA analysis revealed alterations in the infiltration abundance of several immune cells in OA cartilage tissue, which exhibited correlation with GRPEL1 expression. Altogether, this study has revealed that GRPEL1 functions as a novel and significant diagnostic indicator for OA by employing two machine learning methodologies. Furthermore, these findings provide fresh perspectives on potential targeted therapeutic interventions in the future.

Humans

A Risk Score for Polycystic Ovary Syndrome Based on Meta-Analysis and Machine Learning of Gut Microbiota Signatures.

Polycystic Ovary Syndrome (PCOS) is a prevalent endocrine and metabolic disorder among reproductive-age women, in which emerging evidence suggests a substantial role played by the gut microbiota. To comprehensively evaluate gut microbiota alterations in PCOS and identify microbial biomarkers through integrated analysis, a systematic search of PubMed, Web of Science, and Embase was conducted for studies employing 16S rRNA gene sequencing of fecal samples from PCOS cohorts. Ten eligible PCOS cohorts, comprising 858 individuals, were included in the study, from which a risk score was derived using a 20-gene gut microbial signature associated with PCOS. Meta-analysis at the genus level identified that Subdoligranulum, NK4A214_group, and Collinsella significantly decreased, and Bacteroides increased in PCOS across multiple cohorts. Machine learning analysis identified a 20-genus microbial signature using the least absolute shrinkage and selection operator (LASSO) method, which was used to construct a risk score with an AUC of 0.835 in diagnosis prediction. Network analysis further identified Negativibacillus and Lachnospiraceae_UCG_010 as potential driver microbes in PCOS. The analysis in this study highlights key alterations in the gut microbiota across PCOS cohorts. The identified gut microbial signature and derived LASSO-based risk model offer novel insights and a potential tool for PCOS diagnosis.

Polycystic Ovary Syndrome

Identification of key immune-related genes and potential therapeutic drugs in diabetic nephropathy based on machine learning algorithms.

BACKGROUND: Diabetic nephropathy (DN) is a major contributor to chronic kidney disease. This study aims to identify immune biomarkers and potential therapeutic drugs in DN. METHODS: We analyzed two DN microarray datasets (GSE96804 and GSE30528) for differentially expressed genes (DEGs) using the Limma package, overlapping them with immune-related genes from ImmPort and InnateDB. LASSO regression, SVM-RFE, and random forest analysis identified four hub genes (EGF, PLTP, RGS2, PTGDS) as proficient predictors of DN. The model achieved an AUC of 0.995 and was validated on GSE142025. Single-cell RNA data (GSE183276) revealed increased hub gene expression in epithelial cells. CIBERSORT analysis showed differences in immune cell proportions between DN patients and controls, with the hub genes correlating positively with neutrophil infiltration. Molecular docking identified potential drugs: cysteamine, eltrombopag, and DMSO. And qPCR and western blot assays were used to confirm the expressions of the four hub genes. RESULTS: Analysis found 95 and 88 distinctively expressed immune genes in the two DN datasets, with 14 consistently differentially expressed immune-related genes. After machine learning algorithms, EGF, PLTP, RGS2, PTGDS were identified as the immune-related hub genes associated with DN. In addition, the mRNA and protein levels of them were obviously elevated in HK-2 cells treated with glucose for 24&#xa0;h, as well as their mRNA expressions in kidney tissues of mice with DN. CONCLUSION: This study identified 4 hub immune-related genes (EGF, PLTP, RGS2, PTGDS), as well as their expression profiles and the correlation with immune cell infiltration in DN.

Diabetic Nephropathies

Machine learning analysis of the human initiator region reveals key features of different types of core promoters.

The initiator (Inr) is the starting point for the transcription of many genes. Here, we generated highly predictive machine learning models of the human Inr region, and determined that the Inr is present in &#x223c;60% of focused human promoters, identified a novel TATA-specific Inr, and detected the overlapping but functionally distinct TCT motif. Quantitative genome-wide analyses revealed a strict and synergistic interaction between the Inr and DPR, an inverse relationship between the TATA and DPR, a flexible and sometimes independent function of the TATA box in relation to the Inr, and different properties of the TCT motif in humans versus Drosophila.

Humans

Stochastic epigenetic mutation profiles as biomarkers of clinical activity in juvenile idiopathic arthritis: a multi-omic machine learning approach for gene prioritization.

BACKGROUND: Juvenile idiopathic arthritis (JIA) is a rare autoimmune disease arising from a complex interplay between genetic and environmental factors. Epigenetic modifications such as DNA methylation (DNAm) have been described as potential mediators in gene-environment interactions, contributing to immune system dysregulation. Emerging evidence suggests that DNAm profiles also predict therapeutic responses in autoimmune diseases. This study aims to identify epigenetic biomarkers and epigenetic-driven gene expression changes associated with JIA clinical activity. METHODS: We reanalyzed a publicly available dataset of 44 JIA patients, with whole-genome DNAm and gene expression from CD4&#x2009;+&#x2009;T cells measured at two points: at anti-TNF therapy withdrawal (T0) and eight months later (Tend). At Tend, 30 patients maintained inactive disease (ID) while 14 did not (NO ID). We investigated differences between ID and NO ID patients in the epigenetic mutation load and various epigenetic clocks through linear regression models, and prioritized genomic regions with significantly higher number of epimutations in NO ID patients through machine learning. RESULTS: We found a higher mutation load in NO ID than ID patients, both at T0 and at Tend, with the differences at Tend reaching statistical significance (p&#x2009;=&#x2009;0.02). In contrast, we found no evidence of association between epigenetic clocks and JIA clinical activity. Using a multi-omic approach, we identified a List of candidate epigenetically-driven differentially expressed genes, 80 up-regulated and 77 down-regulated, in NO ID patients. Finally, comparing our candidate gene list with the Connectivity Map database, we identified new candidate potential therapeutic targets. Key findings were validated in independent datasets: DNAm profiles from CD4&#x2009;+&#x2009;T cells (56 JIA patients, 57 controls) and transcriptomic data from PBMCs of JIA patients with active or inactive disease, confirming dysregulation of pathways such as TNF-&#x3b1; signaling via NF-kB and TGF-&#x3b2; signaling among others. CONCLUSIONS: We described a significant association of epigenetic mutations with JIA clinical activity, indicating that epigenetic changes might precede clinical symptoms and may serve as biomarkers for early disease monitoring. Further, our results shed light on biomolecular mechanisms of JIA, supporting the development of more effective treatments.

Humans

Identification of potential biomarkers and mechanisms for keloid disorder based on comprehensive bioinformatics analysis and machine learning algorithms.

BACKGROUND: Keloid disorder (KD) encompasses a spectrum of fibroproliferative dermal conditions, the pathogenesis remains complex and incompletely understood. This study sought to identify biomarkers and potential therapeutic targets for KD through an integrative bioinformatics approach and machine learning analysis of RNA sequencing data. METHODS: RNA sequencing was performed on skin tissue samples from 13 patients with KD and 14 healthy controls. Using weighted gene co-expression network analysis and differential expression analysis revealed differentially expressed key module genes, and the CytoHubba plugin identified candidate genes. Subsequently analyzed using least absolute shrinkage and selection operator (LASSO) and support vector machine recursive feature elimination (SVM-RFE) methods to pinpoint feature genes associated with KD. Following this, biomarkers were determined through expression level validation, enrichment analysis, and immune infiltration analysis. RESULTS: A total of 420 differentially expressed key module genes were identified, and the top 10 genes with DMNC values were selected as candidate genes. Five feature genes were selected through LASSO and SVM-RFE, with NID2, MFAP2, COL8A1, and P4HA3 showing significant expression differences between KD and control samples, along with consistent expression patterns across datasets, identified as potential biomarkers. These four biomarkers were proved to possess high diagnostic potential, and they were found to exhibit significant positive correlations with one another. Functional enrichment analysis indicated that the primary KEGG pathways associated with these biomarkers included "steroid hormone biosynthesis" and "cytokine-cytokine receptor interaction." Moreover, immune infiltration analysis revealed that the four biomarkers were negatively correlated with type 17 T helper cells and positively correlated with 15 immune cell types, including activated B cells and central memory CD4 T cells. CONCLUSION: In conclusion, NID2, MFAP2, COL8A1, and P4HA3 were identified as key biomarkers for KD, offering new avenues for more targeted and effective diagnostic and therapeutic strategies for managing this condition.

Humans

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans