Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Identification of key immune-related genes and potential therapeutic drugs in diabetic nephropathy based on machine learning algorithms.

BACKGROUND: Diabetic nephropathy (DN) is a major contributor to chronic kidney disease. This study aims to identify immune biomarkers and potential therapeutic drugs in DN. METHODS: We analyzed two DN microarray datasets (GSE96804 and GSE30528) for differentially expressed genes (DEGs) using the Limma package, overlapping them with immune-related genes from ImmPort and InnateDB. LASSO regression, SVM-RFE, and random forest analysis identified four hub genes (EGF, PLTP, RGS2, PTGDS) as proficient predictors of DN. The model achieved an AUC of 0.995 and was validated on GSE142025. Single-cell RNA data (GSE183276) revealed increased hub gene expression in epithelial cells. CIBERSORT analysis showed differences in immune cell proportions between DN patients and controls, with the hub genes correlating positively with neutrophil infiltration. Molecular docking identified potential drugs: cysteamine, eltrombopag, and DMSO. And qPCR and western blot assays were used to confirm the expressions of the four hub genes. RESULTS: Analysis found 95 and 88 distinctively expressed immune genes in the two DN datasets, with 14 consistently differentially expressed immune-related genes. After machine learning algorithms, EGF, PLTP, RGS2, PTGDS were identified as the immune-related hub genes associated with DN. In addition, the mRNA and protein levels of them were obviously elevated in HK-2 cells treated with glucose for 24 h, as well as their mRNA expressions in kidney tissues of mice with DN. CONCLUSION: This study identified 4 hub immune-related genes (EGF, PLTP, RGS2, PTGDS), as well as their expression profiles and the correlation with immune cell infiltration in DN.

Diabetic Nephropathies↗

Artificial intelligence (AI) uses in stereotactic radiosurgery (SRS): diagnosis with brain metastasis (BM) - A systematic review.

BACKGROUND: Brain metastases (BM) are the most common intracranial tumors in adults, and stereotactic radiosurgery (SRS) has become a mainstay of management. However, several diagnostic challenges persist in the SRS pathway, particularly the differentiation of radiation necrosis (RN) from true tumor progression, which conventional MRI and even advanced imaging techniques often cannot reliably resolve. Recent advances in artificial intelligence (AI) offer the potential to address these diagnostic limitations. This systematic review synthesizes current literature on AI applications for MRI-based diagnostic decision support in BM patients undergoing SRS, with a focus on radiomics and deep learning tools for distinguishing RN from progression, classifying molecular and histologic subtypes, and predicting treatment response. METHODS: A systematic review was performed in accordance with PRISMA guidelines. PubMed, Web of Science, and Scopus were searched using a targeted query combining terms related to AI, brain metastasis, diagnosis or imaging, and SRS. After screening 483 records and applying strict inclusion and exclusion criteria, 18 studies published between 2015 and 2025 were included. Data were extracted on study design, cohort characteristics, imaging modality, AI methodology, validation strategy, and reported diagnostic performance. RESULTS: Among the 18 included studies, AI models demonstrated strong performance across diagnostic tasks in the BM-SRS pathway. The differentiation of RN from true tumor progression was the most extensively studied application, addressed by 14 of 18 studies, with reported AUCs ranging from 0.71 to 0.94. Support vector machines, random-forest ensembles, convolutional neural networks, and transformer-based multimodal architectures were widely used. The literature evolved from single-sequence radiomic classifiers in 2018 to multimodal deep learning frameworks fusing imaging with clinical and genomic data in 2025. Contrast-enhanced T1-weighted MRI was the dominant imaging input, and texture-based radiomic features (GLCM, GLSZM, GLDM, and wavelet-derived features) were the most consistently predictive. The highest-performing models reached AUCs of 0.85-0.91 through multimodal integration of imaging with clinical and genomic features, and consistently outperformed expert neuroradiologist read on matched cases. Remaining studies addressed longitudinal segmentation-based detection of local failure and adverse radiation effects, BRAF mutation status in melanoma BM, early Gamma Knife treatment response, and primary tumor histology classification, with more variable performance. CONCLUSION: AI models, particularly those integrating MRI-derived radiomic features with clinical and genomic data, show high accuracy in supporting diagnostic decisions for BM patients treated with SRS. The post-SRS differentiation of radiation necrosis from true tumor progression has reached the greatest level of maturity and is closest to clinical translation, with potential to reduce unnecessary biopsies, personalize surveillance intervals, and rationalize treatment-pathway decisions. Other diagnostic applications, including molecular subtyping and primary tumor histology classification, remain exploratory and require further multicenter validation. Integration of AI tools into multidisciplinary tumor-board workflows, combined with prospective validation and standardized reporting, will be essential to realize the full clinical benefits of AI in SRS for brain metastases.

Humans↗

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans↗

Triage and workflow optimization with artificial intelligence in pediatric imaging.

Artificial intelligence (AI) is being increasingly utilized in various aspects by the radiology department. With an ever-increasing burden on the healthcare system, particularly in emergency units, the need to incorporate AI in patient triage and workflow optimization cannot be overstated. Machine learning (ML)-based algorithms form the core of AI-based software, aiding healthcare professionals at nearly every step in delivering appropriate patient care. Regarding the radiology section of the hospital, AI-based algorithms have proven exceptionally useful in assisting radiologists and technicians with image acquisition. From accurate clinical referrals to scheduling computed tomography/magnetic resonance imaging scan appointments, from ensuring the lowest radiation exposure to offering timely follow-up reminders, ML-based software has indeed revolutionized the concept of modern image acquisition, especially in the pediatric radiology section. Although the implementation of these algorithms is swift, several technical challenges and the limited availability of pediatric datasets preclude their widespread use. The utility of multimodal pediatric datasets, which combine imaging, genomics, and clinical data, for comprehensive AI triage models can help AI systems evolve toward greater adaptability and integration, resulting in enhanced efficiency, reduced turnaround times, and improved patient outcomes in pediatric radiology departments in the future. In this article, we highlight and review the utility of AI and machine learning-based algorithms in efficiently aiding triage and streamlining the workflow in the pediatric radiology section, thereby ensuring an overall improvement in the departmental workflow.

Triage↗

Machine learning algorithm-based biomarker exploration and validation of mitochondria-related diagnostic genes in osteoarthritis.

The role of mitochondria in the pathogenesis of osteoarthritis (OA) is significant. In this study, we aimed to identify diagnostic signature genes associated with OA from a set of mitochondria-related genes (MRGs). First, the gene expression profiles of OA cartilage GSE114007 and GSE57218 were obtained from the Gene Expression Omnibus. And the limma method was used to detect differentially expressed genes (DEGs). Second, the biological functions of the DEGs in OA were investigated using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. Wayne plots were employed to visualize the differentially expressed mitochondrial genes (MDEGs) in OA. Subsequently, the LASSO and SVM-RFE algorithms were employed to elucidate potential OA signature genes within the set of MDEGs. As a result, GRPEL and MTFP1 were identified as signature genes. Notably, GRPEL1 exhibited low expression levels in OA samples from both experimental and test group datasets, demonstrating high diagnostic efficacy. Furthermore, RT-qPCR analysis confirmed the reduced expression of Grpel1 in an in vitro OA model. Lastly, ssGSEA analysis revealed alterations in the infiltration abundance of several immune cells in OA cartilage tissue, which exhibited correlation with GRPEL1 expression. Altogether, this study has revealed that GRPEL1 functions as a novel and significant diagnostic indicator for OA by employing two machine learning methodologies. Furthermore, these findings provide fresh perspectives on potential targeted therapeutic interventions in the future.

Humans↗

Paternally Expressed Gene 10 Promoter Methylation Level as a Predictor of HBeAg Seroconversion in Chronic Hepatitis B Patients.

The management of chronic hepatitis B (CHB) encounters challenges like suboptimal antiviral response and the lack of predictive biomarkers. In this study, the role of paternally expressed gene 10 (PEG10) in hepatitis B e antigen (HBeAg) seroconversion (HBeAg SC) was explored to identify a therapeutic target and predictive model. In total, 349 participants were recruited, and 141 HBeAg-positive patients were followed up after 48 weeks of antiviral therapy. Key genes were screened by machine learning algorithms (BORUTA, RF and LASSO). PEG10 mRNA, promoter methylation and plasma levels were examined. The effect of PEG10 was assessed by logistic regression, and HBeAg SC was predicted by nomograms. HBeAg-positive patients showed markedly elevated PEG10 mRNA expression (p&#x2009;<&#x2009;0.001), which correlated strongly with major virological markers such as HBV DNA (r&#x2009;=&#x2009;0.520, p&#x2009;<&#x2009;0.001), HBeAg (r&#x2009;=&#x2009;0.490, p&#x2009;<&#x2009;0.001) and HBsAg (r&#x2009;=&#x2009;0.400, p&#x2009;<&#x2009;0.001). In addition, HBeAg-positive patients exhibited a significant reduction in PEG10 promoter methylation levels compared with controls (p&#x2009;<&#x2009;0.001). According to logistic regression analysis, PEG10 promoter methylation status was an independent predictor of HBeAg SC. The predictive nomogram incorporating PEG10 promoter methylation ratio (PMR), albumin (ALB), aspartate aminotransferase (AST) and HBeAg demonstrated excellent clinical predictive value (area under curve (AUC)&#x2009;=&#x2009;0.895,95% confidence interval (CI): 0.808&#x2009;~&#x2009;0.963). The methylation status of the PEG10 promoter represents a promising biomarker for the prediction of HBeAg SC in patients with CHB. CLINICAL TRIAL REGISTRATION: Not applicable.

Humans↗

A machine learning approach to computer-aided molecular design.

Preliminary results of a machine learning application concerning computer-aided molecular design applied to drug discovery are presented. The artificial intelligence techniques of machine learning use a sample of active and inactive compounds, which is viewed as a set of positive and negative examples, to allow the induction of a molecular model characterizing the interaction between the compounds and a target molecule. The algorithm is based on a twofold phase. In the first one--the specialization step--the program identifies a number of active/inactive pairs of compounds which appear to be the most useful in order to make the learning process as effective as possible and generates a dictionary of molecular fragments, deemed to be responsible for the activity of the compounds. In the second phase--the generalization step--the fragments thus generated are combined and generalized in order to select the most plausible hypothesis with respect to the sample of compounds. A knowledge base concerning physical and chemical properties is utilized during the inductive process.

Amino Acid Sequence↗

Information retrieval: an overview of system characteristics.

The paper gives an overview of characteristics of information retrieval (IR) systems. The characteristics are identified from the descriptions of 23 IR systems. Four IR models are discussed: the Boolean model, the vector model, the probabilistic model and the connectionistic model. Twelve other characteristics of IR models are identified: search intermediary, domain knowledge, relevance feedback, natural language interface, graphical query language, conceptual queries, full-text IR, field searching, fuzzy queries, hypertext integration, machine learning, and ranked output. Finally, the relevance of IR systems for the World Wide Web is established.

Algorithms↗

mamp-ml: A deep learning approach to epitope immunogenicity in plants.

Eukaryotes detect biomolecules through surface-localized receptors, key signaling components. A subset of receptors survey for pathogens, induce immunity, and restrict pathogen growth. Comparative genomics of both hosts and pathogens has unveiled vast sequence variation in receptors and potential ligands, creating an experimental bottleneck. We have developed mamp-ml, a machine learning framework for predicting plant receptor-ligand interactions. We leveraged existing functional data from over two decades of foundational research, together with the large protein language model ESM-2, to build a pipeline and model that predicts immunogenic outcomes using a combination of receptor-ligand features. Our model achieves 73% prediction accuracy on a held-out test set, even when an experimental structure is lacking. Our approach enables high-throughput screening of LRR receptor-ligand combinations and provides a computational framework for engineering plant immune systems.

Journal Article↗

Gradient-based optimization of hyperparameters.

Many machine learning algorithms can be formulated as the minimization of a training criterion that involves a hyperparameter. This hyperparameter is usually chosen by trial and error with a model selection criterion. In this article we present a methodology to optimize several hyperparameters, based on the computation of the gradient of a model selection criterion with respect to the hyperparameters. In the case of a quadratic training criterion, the gradient of the selection criterion with respect to the hyperparameters is efficiently computed by backpropagating through a Cholesky decomposition. In the more general case, we show that the implicit function theorem can be used to derive a formula for the hyperparameter gradient involving second derivatives of the training criterion.

Algorithms↗

Combination of computational techniques and RNAi reveal targets in Anopheles gambiae for malaria vector control.

Increasing reports of insecticide resistance continue to hamper the gains of vector control strategies in curbing malaria transmission. This makes identifying new insecticide targets or alternative vector control strategies necessary. CLassifier of Essentiality AcRoss EukaRyote (CLEARER), a leave-one-organism-out cross-validation machine learning classifier for essential genes, was used to predict essential genes in Anopheles gambiae and selected predicted genes experimentally validated. The CLEARER algorithm was trained on six model organisms: Caenorhabditis elegans, Drosophila melanogaster, Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Schizosaccharomyces pombe, and employed to identify essential genes in An. gambiae. Of the 10,426 genes in An. gambiae, 1,946 genes (18.7%) were predicted to be Cellular Essential Genes (CEGs), 1716 (16.5%) to be Organism Essential Genes (OEGs), and 852 genes (8.2%) to be essential as both OEGs and CEGs. RNA interference (RNAi) was used to validate the top three highly expressed non-ribosomal predictions as probable vector control targets, by determining the effect of these genes on the survival of An. gambiae G3 mosquitoes. In addition, the effect of knockdown of arginase (AGAP008783) on Plasmodium berghei infection in mosquitoes was evaluated, an enzyme we computationally inferred earlier to be essential based on chokepoint analysis. Arginase and the top three genes, AGAP007406 (Elongation factor 1-alpha, Elf1), AGAP002076 (Heat shock 70kDa protein 1/8, HSP), AGAP009441 (Elongation factor 2, Elf2), had knockdown efficiencies of 91%, 75%, 63%, and 61%, respectively. While knockdown of HSP or Elf2 significantly reduced longevity of the mosquitoes (p<0.0001) compared to control groups, Elf1 or arginase knockdown had no effect on survival. However, arginase knockdown significantly reduced P. berghei oocytes counts in the midgut of mosquitoes when compared to LacZ-injected controls. The study reveals HSP and Elf2 as important contributors to mosquito survival and arginase as important for parasite development, hence placing them as possible targets for vector control.

Animals↗

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans↗

A Risk Score for Polycystic Ovary Syndrome Based on Meta-Analysis and Machine Learning of Gut Microbiota Signatures.

Polycystic Ovary Syndrome (PCOS) is a prevalent endocrine and metabolic disorder among reproductive-age women, in which emerging evidence suggests a substantial role played by the gut microbiota. To comprehensively evaluate gut microbiota alterations in PCOS and identify microbial biomarkers through integrated analysis, a systematic search of PubMed, Web of Science, and Embase was conducted for studies employing 16S rRNA gene sequencing of fecal samples from PCOS cohorts. Ten eligible PCOS cohorts, comprising 858 individuals, were included in the study, from which a risk score was derived using a 20-gene gut microbial signature associated with PCOS. Meta-analysis at the genus level identified that Subdoligranulum, NK4A214_group, and Collinsella significantly decreased, and Bacteroides increased in PCOS across multiple cohorts. Machine learning analysis identified a 20-genus microbial signature using the least absolute shrinkage and selection operator (LASSO) method, which was used to construct a risk score with an AUC of 0.835 in diagnosis prediction. Network analysis further identified Negativibacillus and Lachnospiraceae_UCG_010 as potential driver microbes in PCOS. The analysis in this study highlights key alterations in the gut microbiota across PCOS cohorts. The identified gut microbial signature and derived LASSO-based risk model offer novel insights and a potential tool for PCOS diagnosis.

Polycystic Ovary Syndrome↗

Machine Learning-Based Preoperative Predicting TERT Promoter Mutation and EGFR Gene Amplification Phenotype in IDH Wild-Type Glioblastoma Using Advanced MR Habitat Imaging.

BACKGROUND AND PURPOSE: The telomerase reverse transcriptase (TERT) gene promoter mutation is a crucial factor for identifying an isocitrate dehydrogenase (IDH) wild-type glioblastoma with poor prognosis, and the epidermal growth factor receptor (EGFR) amplification may be a potential prognostic factor. The purpose of this study was to investigate the value of the tumor habitats imaging model on advanced MRI in predicting TERT promoter mutation and EGFR gene amplification phenotype of IDH wild-type glioblastoma. MATERIALS AND METHODS: One hundred seventy-nine patients with pretreatment conventional MRI, DWI, and DSC-PWI were included. The data were divided into the training set (n=112), test set (n=29), and time-independent validation set (n=38). Based on the ADC and CBV map, the solid tumor area was split into several habitat subregions using the k-means clustering algorithm (hypovascular hypercellular area, hypervascular area, and hypovascular hypocellular area). In the training set, TERT promoter mutation and EGFR gene amplification phenotype prediction models were constructed using the random forest method. The reliability of prediction models was validated in the test and the time-independent validation sets. Receiver operating characteristic (ROC) curve analysis, calibration curve, and decision curve analysis (DCA) were used. RESULTS: The area under the curve (AUC) of the training, test, and validation sets of the TERT promoter prediction model was 0.877, 0.783, and 0.796, respectively. The accuracy of the TERT promoter prediction model was 82.1%, 75.9%, and 76.3%, respectively. The AUCs of the 3 sets for the EGFR gene amplification status prediction model were 0.877, 0.784, and 0.878, respectively. The accuracy of the EGFR gene amplification status prediction model was 79.5%, 75.9%, and 89.5%, respectively. Moreover, the prediction probability of these models was in good agreement with the actual result. CONCLUSIONS: The tumor habitat imaging model based on advanced MRI was useful for accurately predicting TERT promoter mutation and EGFR amplification status in IDH wild-type glioblastoma.

Humans↗

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article↗

Machine learning of motor vehicle accident categories from narrative data.

Bayesian inferencing as a machine learning technique was evaluated for identifying pre-crash activity and crash type from accident narratives describing 3,686 motor vehicle crashes. It was hypothesized that a Bayesian model could learn from a computer search for 63 keywords related to accident categories. Learning was described in terms of the ability to accurately classify previously unclassifiable narratives not containing the original keywords. When narratives contained keywords, the results obtained using both the Bayesian model and keyword search corresponded closely to expert ratings (P(detection) > or = 0.9, and P (false positive) < or = 0.05). For narratives not containing keywords, when the threshold used by the Bayesian model was varied between p > 0.5 and p > 0.9, the overall probability of detecting a category assigned by the expert varied between 67% and 12%. False positives correspondingly varied between 32% and 3%. These latter results demonstrated that the Bayesian system learned from the results of the keyword searches.

Accidents, Traffic↗

Phylogenetic Methods Meet Deep Learning.

Deep learning (DL) has been widely used in various scientific fields, but its integration into phylogenetics has been slower, primarily due to the complex nature of phylogenetic data. The studies that apply DL to sequencing data often limit analyses to four-taxon trees. Many of these studies serve as "proof of principle" and perform similarly to traditional phylogeny reconstruction methods. New ways of using training data, such as encoding with compact bijective ladderized vectors or transformers, enable the handling of much larger trees and genomic data sets. This short perspective focuses on the application of DL in phylogenetics, introducing prevalent DL architectures. We highlight potential problems in the field by discussing the risks of using simulation-based training data and emphasize the importance of reproducibility and robustness in computational estimates. Finally, we explore promising research areas, including the combination of phylogenetics and population genetics in DL, the analysis of neighbor dependencies, and the potential to significantly reduce computational cost compared to traditional methods. This perspective illustrates the potential of DL in complementing traditional phylogeny reconstruction methods and aiding the advancement of phylogenetic analysis, especially in performing computationally demanding tasks such as model selection or estimating branch support values.

Humans↗