Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Prognostic significance of DNA damage response-related markers in esophageal squamous cell carcinoma using machine learning approaches.

BACKGROUND: Esophageal squamous cell carcinoma (ESCC) lacks reliable prognostic biomarkers. Homologous recombination deficiency (HRD) has been implicated in genomic instability across multiple cancers, but its prognostic significance in ESCC remains unexplored. This study aimed to evaluate HRD score as a prognostic biomarker and develop a machine learning-based predictive model for ESCC. METHODS: Transcriptomic and clinical data from 78 ESCC patients were obtained from The Cancer Genome Atlas (TCGA) and randomly split into training (70%) and test (30%) cohorts. Prognostic models were constructed using 112 machine learning algorithm combinations based on DNA damage response (DDR)-related genes. Gene set enrichment analysis (GSEA), somatic mutation profiling, and immune cell infiltration estimation via CIBERSORT were performed to characterize HRD-associated molecular features. RESULTS: High HRD scores were significantly associated with poorer overall survival (P<0.05). Among 112 algorithm combinations, the survival support vector machine (Survival-SVM) model demonstrated optimal performance [training concordance index (C-index): 0.741; test C-index: 0.708], identifying six hub genes: PARP1, MBD4, TELO2, NSMCE3, SMUG1, and BABAM1. A nomogram incorporating risk score (RS) and clinical variables achieved strong predictive accuracy for 1- to 3-year survival [area under the curve (AUC) >0.7]. High-HRD tumors exhibited distinct mutational patterns (TP53 and TTN) and enriched glutathione metabolism and cytochrome P450 pathways. Immune infiltration analysis revealed significant differences in plasma cell and neutrophil infiltration between risk groups (P<0.05), suggesting HRD-associated immune microenvironment remodeling. CONCLUSIONS: We developed a novel HRD-based prognostic model incorporating six DDR-related genes that demonstrates robust predictive performance in ESCC. HRD score is identified as an independent prognostic factor associated with genomic instability, immune microenvironment alterations, and clinical outcomes. These findings provide a theoretical basis for personalized treatment strategies, including potential applications of PARP inhibitors and immunotherapy in ESCC.

Esophageal squamous cell carcinoma (ESCC)↗

Predicting natural variation in the yeast phenotypic landscape with machine learning.

Most organismal traits result from the complex interplay of many genetic and environmental factors, making their prediction difficult. Here, we used machine learning (ML) models to explore phenotype predictions for 223 traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains isolated worldwide. We benchmarked a ML pipeline with multiple linear and non-linear models to predict phenotypes from genotypes and gene expression, and determined gradient boosting machines as the best-performing model. Gene function disruption scores and gene presence/absence emerged as best predictors, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The prediction accuracy broadly varied among phenotypes, with stress resistance being easier to predict compared to growth across nutrients. ML identified relevant genomic features linked to phenotypes, including high-impact variants with established relationships to phenotypes, despite these being rare in the population. Near-perfect accuracies were achieved when other phenomics data mostly in similar conditions were used, suggesting that useful information can be conveyed across phenotypes. Overall, our study underscores the power of ML to interpret the functional outcome of genetic variants.

Genetic Variation↗

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in&#xa0;gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n&#x2009;=&#x2009;42, Healthy: n&#x2009;=&#x2009;54), Franzosa et al. (IBD: n&#x2009;=&#x2009;164, Healthy: n&#x2009;=&#x2009;56), and Yachida et al. (CRC: n&#x2009;=&#x2009;150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans↗

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning↗

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal&#x2011;Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans↗

NRG-P0074 Viral Sample RU1 from Unclassified Mosigvirus Genomic Characterization and Host Range Analysis.

BACKGROUND: Machine learning models for phage-host range prediction and design require comprehensive training data on phage genomes and host ranges to predict phage-host interactions effectively. MATERIALS AND METHODS: This study characterizes phage sample NRG-P0074 viral sample RU1 from unclassified Mosigvirus, originally isolated by the Betty Kutter. The complete genome of NRG-P0074 was sequenced, annotated, and analyzed using various bioinformatic tools. Host range analysis was conducted using the Escherichia coli Reference (ECOR) Library and nine Escherichia coli (E. coli) K12 strains (Keio Knockout Collection) with single nonessential gene deletions. RESULTS: The genome of NRG-P0074 spans 168,357 base pairs with a guanine-cytosine (GC) content of 37.5%. NRG-P0074 exhibited permissiveness in 15.28% of the ECOR isolates and all 9 Keio knockout strains. Comparative genomic analysis revealed that NRG-P0074 is closely related to E. coli phage a20. Its genome is comprised of 270 coding sequences, 153 known genes, 16 terminators, 3 ribosomal-binding sites, 0 tRNAs, and 117 hypothetical proteins. CONCLUSIONS: This research provides valuable data for developing machine learning models to predict phage-host interactions, aiding the development of targeted phage therapies against antibiotic-resistant bacteria.

ECOR Library↗

Engineering Bacillus Subtilis for Efficient Biosynthesis of Riboflavin: Current Knowledge and Future Perspectives.

Riboflavin is an essential water-soluble vitamin that serves as a precursor for the biosynthesis of the flavin cofactors FMN and FAD, which play pivotal roles in numerous redox and energy metabolism reactions. With the growing global demand for sustainable vitamin production, microbial fermentation has become an attractive alternative to chemical synthesis due to its environmental and economic advantages. Among microbial hosts, Bacillus subtilis has emerged as a leading cell factory for riboflavin production owing to its GRAS status, well-characterized genetics, and efficient protein secretion system. This review provides a comprehensive overview of recent advances in metabolic engineering strategies to enhance riboflavin biosynthesis in B. subtilis. Key topics include strengthening biosynthetic and precursor pathways, relieving feedback inhibition, balancing metabolic flux and cell growth, employing adaptive laboratory evolution, and utilizing omics-guided optimization and 13C metabolic flux analysis. Moreover, the integration of synthetic biology tools such as riboswitch engineering, regulatory element design, and high-throughput screening has significantly accelerated strain improvement. Despite remarkable progress, challenges remain in achieving precise regulatory control, optimizing multi-gene expression, and enhancing genome integration efficiency. Future research combining multi-omics data, synthetic regulatory design, and machine learning-driven predictive modeling is expected to further advance the development of intelligent B. subtilis cell factories. However, the practical implementation of these systems remains constrained by the metabolic burden of overproduction and the lack of universal regulatory models that can predict strain performance across varying industrial scales.

Bacillus subtilis↗

Integrative multi-omics profiling deciphers tumor microenvironment heterogeneity and immunotherapy vulnerabilities in lung neuroendocrine carcinomas.

INTRODUCTION: Lung neuroendocrine carcinomas (Lu-NECs) are rare, highly aggressive lung tumors with poor prognosis and limited therapeutic options. Understanding the tumor immune microenvironment (TIME) is crucial towards personalized therapeutic strategies. OBJECTIVES: This study aims to systematically characterize the heterogeneity and complexity of the TIME in Lu-NECs by integrating proteomic, transcriptomic, and genomic data. METHODS: We performed comprehensive immune-proteomic profiling of 76 Lu-NECs across diverse histopathological subtypes to elucidate intra-tumoral TIME heterogeneity at the proteomic level. Validation was conducted in multiple independent cohorts, including 112 Lu-NECs using immunohistochemistry, 147 Lu-NECs, and 17 small cell lung carcinoma samples using transcriptomics. We integrated proteomic, transcriptomic, genomic, and clinical data to assess molecular, immunological, and clinical features, as well as therapeutic vulnerabilities across different immune subtypes. RESULTS: We delineated the immuno-proteomic landscape of Lu-NECs and identified two major immuno-proteomic clusters with distinct immunological, molecular, and clinical characteristics. IPC1 was characterized by high immune cell infiltration, while IPC2 exhibited sparse immune cell presence. Genomic analysis revealed distinct mutational patterns, with IPC1 showing a higher incidence of APOBEC-associated mutation signatures and IPC2 being enriched for mutations associated with defective DNA mismatch repair and tobacco-related mutagens. Functional analyses indicated that IPC1 was related to immune and oncogenic signaling activity, whereas IPC2 was associated with cancer stemness and proliferation-related features. Furthermore, IPC1 and IPC2 demonstrated histological subtype-specific clinical benefits from postoperative chemotherapy. Finally, we developed a machine learning model (iPROM) to predict Lu-NECs immune classification and improve risk stratification, which was validated across multiple independent cohorts. CONCLUSIONS: This study advances the understanding of the tumor immune microenvironment in Lu-NECs through multi-omics characterization and highlights potential personalized therapeutic vulnerabilities tailored to the specific immune landscapes of Lu-NECs.

Humans↗

Machine learning in sedimentation modelling.

The paper presents machine learning (ML) models that predict sedimentation in the harbour basin of the Port of Rotterdam. The important factors affecting the sedimentation process such as waves, wind, tides, surge, river discharge, etc. are studied, the corresponding time series data is analysed, missing values are estimated and the most important variables behind the process are chosen as the inputs. Two ML methods are used: MLP ANN and M5 model tree. The latter is a collection of piece-wise linear regression models, each being an expert for a particular region of the input space. The models are trained on the data collected during 1992-1998 and tested by the data of 1999-2000. The predictive accuracy of the models is found to be adequate for the potential use in the operational decision making.

Algorithms↗

Clinical translation of senescence-related pan-cancer multi-omics: tools for assessment and immunotherapy prediction.

Cellular senescence (CS) exerts dual roles in tumorigenesis, yet its pan-cancer molecular characteristics and clinical value remain unclear, hindering its translation to oncology and personalized therapy. To address the lack of specific and universal tools for senescence assessment and immunotherapy response prediction, this study systematically analyzed 1259 CS-related genes from the CellAge database across 31 cancer types by integrating multi-omics data, including bulk RNA-seq, single-cell/spatial transcriptomics, and CRISPR screening. We developed a rank-based algorithm SenScoreR (publicly available at https://gxhub.shinyapps.io/SenScoreR/ ) for senescence quantification, validated with 10 independent datasets, and constructed a machine learning-based predictive model CS.Sig for immunotherapy response. Results showed that tumors had significantly lower Rank-based Senescence Score (RSS) than normal tissues across 31 cancers (average diagnostic AUC&#x2009;=&#x2009;0.895), with low RSS linked to poor survival; high RSS correlated with reduced genomic instability, enriched CD8&#x207a; T/NK cell/macrophage infiltration, upregulated PD-L1 expression, and elevated immune cytolytic activity. CS.Sig demonstrated robust performance in predicting ICI response (AUC&#x2009;=&#x2009;0.716 across 10 cohorts), outperforming 13 existing signatures, while CRISPR screening identified 17 senescence-related targets (e.g., CEP55, PPP1CC) whose knockout enhanced anti-tumor immunity. Our findings clarify CS's role in maintaining tumor genomic stability and shaping immune microenvironments, and the developed SenScoreR, CS.Sig, and identified targets bridge basic CS research with clinical oncology, providing a translational resource and hypothesis basis for future experimental and clinical validation.

Journal Article↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

Molecular hashkeys: a novel method for molecular characterization and its application for predicting important pharmaceutical properties of molecules.

We define a novel numerical molecular representation, called the molecular hashkey, that captures sufficient information about a molecule to predict pharmaceutically interesting properties directly from three-dimensional molecular structure. The molecular hashkey represents molecular surface properties as a linear array of pairwise surface-based comparisons of the target molecule against a common 'basis-set' of molecules. Hashkey-measured molecular similarity correlates well with direct methods of measuring molecular surface similarity. Using a simple machine-learning technique with the molecular hashkeys, we show that it is possible to accurately predict the octanol-water partition coefficient, log P. Using more sophisticated learning techniques, we show that an accurate model of intestinal absorption for a set of drugs can be constructed using the same hashkeys used in the aforementioned experiments. Once a set of molecular hashkeys is calculated, its use in the training and testing of property-based models is very fast. Further, the required amount of data for model construction is very small. Neural network-based hashkey models trained on data sets as small as 30 molecules yield statistically significant prediction of molecular properties. The lack of a requirement for large data sets lends itself well to the prediction of pharmaceutically relevant molecular parameters for which data generation is expensive and slow. Molecular hashkeys coupled with machine-learning techniques can yield models that predict key pharmacological aspects of biologically important molecules and should therefore be important in the design of effective therapeutics.

Drug Design↗

Data-driven consideration of genetic disorders for global genomic newborn screening programs.

PURPOSE: Over 30 international studies are exploring newborn sequencing (NBSeq) to expand the range of genetic disorders included in newborn screening. Substantial variability in gene selection across programs exists, highlighting the need for a systematic approach to prioritize genes. METHODS: We assembled a data set comprising 25 characteristics about each of the 4390 genes included in 27 NBSeq programs. We used regression analysis to identify several predictors of inclusion and developed a machine learning model to rank genes for public health consideration. RESULTS: Among 27 NBSeq programs, the number of genes analyzed ranged from 134 to 4299, with only 74 (1.7%) genes included by over 80% of programs. The most significant associations with gene inclusion across programs were presence on the US Recommended Uniform Screening Panel (inclusion increase of 74.7%, CI: 71.0%-78.4%), robust evidence on the natural history (29.5%, CI: 24.6%-34.4%), and treatment efficacy (17.0%, CI: 12.3%-21.7%) of the associated genetic disease. A boosted trees machine learning model using 13 predictors achieved high accuracy in predicting gene inclusion across programs (area under the curve = 0.915, R2 = 84%). CONCLUSION: The machine learning model developed here provides a ranked list of genes that can adapt to emerging evidence and regional needs, enabling more consistent and informed gene selection in NBSeq initiatives.

Humans↗

Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis.

PURPOSE: Given the existing uncertainties regarding the link between Di(2-ethylhexyl) phthalate (DEHP) exposure and gastric cancer (GC) progression, this study aimed to clarify their association, identify the toxic targets of DEHP, and elucidate the underlying molecular mechanisms. METHODS: Multiple integrated approaches were employed, including Gene Expression Omnibus (GEO) data analysis, network toxicology, molecular docking, and machine learning. STRING and Cytoscape tools were utilized to identify key targets, while Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the functional enrichment of intersecting targets. Machine learning and SHAP analysis were applied to screen core targets in GC. Molecular docking was performed to evaluate the binding affinity of DEHP toward core targets, and 200 ns molecular dynamics simulations were further conducted for representative complexes to validate their dynamic stability. RESULTS: A total of 18 key targets were identified using STRING and Cytoscape. GO and KEGG enrichment analyses demonstrated that these intersecting targets were primarily enriched in the extracellular region, as well as the Calcium signaling pathway and cAMP signaling pathway. Through machine learning analyses, 7 key genes (ADRB2, ESRRG, GRIA4, IL13RA2, NR3C2, PLA2G1B, and SULT2A1) were identified as core targets in GC through machine learning analyses. Molecular docking simulations revealed strong binding specificity between DEHP and the target proteins. Among them, NR3C2 and ADRB2 exhibited relatively high predictive importance in the machine learning models. DEHP showed favorable binding affinity toward these core targets, and molecular dynamics simulations further confirmed that ADRB2-DEHP and NR3C2-DEHP complexes maintained stable conformations throughout the simulation. CONCLUSIONS: Our findings identified GC associated genes that were computationally predicted as potential targets of DEHP. These results indicated structural compatibility between DEHP and its target proteins but did not prove that DEHP exposure accounts for the gene expression changes in GC.

Molecular Docking Simulation↗

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (&#x2264;&#x2009;12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n&#x2009;=&#x2009;121, 19 events) for training and centers 2-7 (n&#x2009;=&#x2009;207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans↗

Machine learning approaches to supporting the identification of photoreceptor-enriched genes based on expression data.

BACKGROUND: Retinal photoreceptors are highly specialised cells, which detect light and are central to mammalian vision. Many retinal diseases occur as a result of inherited dysfunction of the rod and cone photoreceptor cells. Development and maintenance of photoreceptors requires appropriate regulation of the many genes specifically or highly expressed in these cells. Over the last decades, different experimental approaches have been developed to identify photoreceptor enriched genes. Recent progress in RNA analysis technology has generated large amounts of gene expression data relevant to retinal development. This paper assesses a machine learning methodology for supporting the identification of photoreceptor enriched genes based on expression data. RESULTS: Based on the analysis of publicly-available gene expression data from the developing mouse retina generated by serial analysis of gene expression (SAGE), this paper presents a predictive methodology comprising several in silico models for detecting key complex features and relationships encoded in the data, which may be useful to distinguish genes in terms of their functional roles. In order to understand temporal patterns of photoreceptor gene expression during retinal development, a two-way cluster analysis was firstly performed. By clustering SAGE libraries, a hierarchical tree reflecting relationships between developmental stages was obtained. By clustering SAGE tags, a more comprehensive expression profile for photoreceptor cells was revealed. To demonstrate the usefulness of machine learning-based models in predicting functional associations from the SAGE data, three supervised classification models were compared. The results indicated that a relatively simple instance-based model (KStar model) performed significantly better than relatively more complex algorithms, e.g. neural networks. To deal with the problem of functional class imbalance occurring in the dataset, two data re-sampling techniques were studied. A random over-sampling method supported the implementation of the most powerful prediction models. The KStar model was also able to achieve higher predictive sensitivities and specificities using random over-sampling techniques. CONCLUSION: The approaches assessed in this paper represent an efficient and relatively inexpensive in silico methodology for supporting large-scale analysis of photoreceptor gene expression by SAGE. They may be applied as complementary methodologies to support functional predictions before implementing more comprehensive, experimental prediction and validation methods. They may also be combined with other large-scale, data-driven methods to facilitate the inference of transcriptional regulatory networks in the developing retina. Furthermore, the methodology assessed may be applied to other data domains.

Animals↗

Case-based explanation of non-case-based learning methods.

We show how to generate case-based explanations for non-case-based learning methods such as artificial neural nets or decision trees. The method uses the trained model (e.g., the neural net or the decision tree) as a distance metric to determine which cases in the training set are most similar to the case that needs to be explained. This approach is well suited to medical domains, where it is important to understand predictions made by complex machine learning models, and where training and clinical practice makes users adept at case interpretation.

Artificial Intelligence↗

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4&#x207a; memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC&#x2009;>&#x2009;0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics↗