Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Proteomic signature of dementia risk in type 2 diabetes.

INTRODUCTION: Type 2 diabetes (T2D) significantly increases dementia risk, yet the molecular mechanisms underlying this association remain unclear. OBJECTIVES: This study aimed to identify protein signatures that distinguish dementia risk in T2D patients, develop a proteomic prediction model, and elucidate biological pathways connecting T2D and dementia. METHODS: We analyzed 2,920 plasma proteins from 52,958 participants (including 3,292 with T2D) in the UK Biobank Pharma Proteomics Project with a median follow-up of 14.6 years. Cox regression models with interaction terms identified T2D-specific protein associations with dementia risk. Machine learning models were developed to predict dementia in T2D patients. Pathway analysis and weighted gene co-expression network analysis identified biological mechanisms linking T2D and dementia. RESULTS: We identified 471 proteins with significant interaction effects between T2D and dementia risk. In non-T2D individuals, elevated levels of neuronal pentraxin receptor (NPTXR, HR = 0.74, 95 %CI:0.66-0.83) and carbonic anhydrase 14 (CA14, HR = 0.67, 95 %CI:0.60-0.75) were exclusively associated with decreased dementia risk. Conversely, in T2D patients, elevated rho guanine nucleotide exchange factor 12 (ARHGEF12, HR = 1.45, 95 %CI:1.10-1.91) was specifically associated with increased dementia risk. A 51-protein model accurately predicted 15-year dementia risk in T2D patients (AUC = 0.835, C-index = 0.829), outperforming conventional clinical risk scores and maintaining high accuracy for Alzheimer's disease and vascular dementia. Pathway analysis revealed enrichment of IL6-JAK-STAT3 signaling in T2D-related dementia, while dysregulation of fatty acid metabolism was specific to T2D-associated Alzheimer's disease. CONCLUSIONS: This large-scale proteomic analysis identifies specific molecular signatures that differentiate dementia risk in diabetic and non-diabetic populations, with potential applications for early risk stratification and targeted interventions. The identified pathways provide novel insights into the pathophysiological processes connecting T2D and dementia and suggest potential therapeutic targets.

Humans

Predicting food taste with bound-driven optimization.

The prediction of sensory attributes from ingredient-level formulations is an emerging challenge at the intersection of food science and artificial intelligence. We address the fundamental question of whether the taste of a food can be predicted from its ingredients by treating recipes as composite materials. We apply Hashin-Shtrikman (HS) and Reuss-Voigt (RV) bounds, techniques originally developed for elastic moduli, as a null-hypothesis additive baseline for five taste dimensions (sweetness, sourness, bitterness, umami, saltiness) on a curated dataset of 70 recipes decomposed into 115 distinct ingredients scored against a library of 209 ingredient-level taste references with trained-panel ground truth. This baseline systematically under-predicts perceived taste: 77% of actual taste values exceeded the HS upper bound, with the exceedance rate ranging from 26% (bitterness) to 97% (saltiness). We traced this gap to specific processing chemistry (Maillard reactions, caramelization, evaporative concentration, protein hydrolysis, and nucleotide synergy) and introduced a hybrid model that augments the HS baseline with eight chemistry-proxy features encoding these mechanisms. Our results show that our interpretable hybrid model eliminates the systematic bias and reduces mean absolute error by 27%-62% for sweetness, sourness, umami, and saltiness while using only 10 interpretable features, achieving performance comparable to a black-box Lasso regression on 115 per-ingredient features. We further demonstrate constrained inverse design via Differential Evolution, recovering ingredient formulations that match target taste profiles subject to compositional bounds. Our work demonstrates how key chemical processes during food preparation can inform and augment physics-based and machine learning models, providing a quantitative fingerprint of processing chemistry's contribution to taste perception and paving the way for model-driven food formulation with targeted sensory characteristics.

Composite material bounds

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4⁺ memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC > 0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides

Machine learning-based clinical tool for identifying factors associated with symptomatic knee osteoarthritis: the Nagahama study.

BACKGROUND: A clinical tool that evaluates factors associated with symptomatic knee osteoarthritis (OA) based on modifiable factors is lacking. This study aimed to develop a machine learning-based clinical assessment tool using modifiable factors to identify factors associated with symptomatic knee OA and to determine its accuracy. METHODS: This study included 429 participants (81.8% women; age, 69.0&#xa0;&#xb1;&#xa0;5.3 years) from the Nagahama Study who were &#x2265;60&#xa0;years old and had radiographically confirmed knee OA. A Knee Society Knee Scoring System 2011 symptom score of <23 points defined symptomatic knee OA. Participants were randomly assigned to training (70%) and test (30%) datasets. A machine learning model was developed using Extreme Gradient Boosting with 27 variables, and the SHapley Additive exPlanation (SHAP) values were used to assess feature importance. The top 8 features were translated into a 100-point clinical scoring tool weighted by their SHAP contributions. The cutoff value indicating symptomatic knee OA in the clinical assessment tool was determined using receiver operating characteristic analysis, and model performance was evaluated in both datasets. RESULTS: The clinical assessment tool consisted of low back pain, OA severity, depressive tendencies, knee flexion/extension range of motion, knee extension and hip abduction strength, and lower limb muscle quality. The model showed moderate discriminative performance (AUC 0.771 and 0.773 in the training and test datasets, respectively), with a cutoff point of 47. CONCLUSION: The proposed clinical assessment tool may provide a structured framework for assessing modifiable factors associated with symptomatic knee OA, reflecting their contribution to current symptom status.

Humans

PGS-GS: a framework integrating polygenic scores and genomic selection in animal breeding.

Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.

Animals

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, &#x3008;MAE&#x3009; = 0.11 eV, and &#x3008;RMSE&#x3009; = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by &#x223c;10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature

Machine Learning and Metabolomics to Characterize Warburg-Like Metabolic Subtypes in Human Retinal Endothelial Cells Exposed to Risk Factors Associated With Proliferative Diabetic Retinopathy.

PURPOSE: High glucose (HG), hypoxia (Hyp), and their combination are major risk factors for proliferative diabetic retinopathy (PDR). Although these conditions induce features of the Warburg-like metabolic reprogramming in human retinal endothelial cells (HRECs), it remains unclear whether they produce distinct metabolic and angiogenic subtypes. This study aimed to characterize the Warburg-like-associated metabolic heterogeneity induced by these PDR-related risk factors and evaluate the ability of supervised machine-learning models to distinguish these subtypes. METHODS: HRECs were cultured under normoglycemic, HG, Hyp (2% O2), and combined HG-Hyp conditions. Untargeted LC-MS/MS metabolomics quantified metabolites spanning carbohydrates, amino acids, nucleotides, and lipids. Principal component analysis (PCA) assessed overall metabolic variation, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis identified metabolic pathways associated with angiogenesis. In vitro angiogenesis assays measured endothelial tube formation and branching. Nine supervised classifiers (decision tree, logistic regression, na&#xef;ve Bayes, random forest, K-Nearest Neighbors, neural network, gradient boosting, AdaBoost, and Support Vector Machine) were trained on the highest-ranked metabolites selected by the Information Gain Ratio feature-ranking approach. Model performance was evaluated using 10-fold cross-validation, leave-one-out cross-validation (LOOCV), permutation testing, and a classifier stability analysis under biologically meaningful distributional shift using an independent chemically induced hypoxia model (CoCl2). RESULTS: PCA revealed partial separation of metabolic profiles across conditions, indicating different Warburg-like metabolic subtypes. The combined HG-Hyp condition exhibited enhanced angiogenic potential relative to either HG or Hyp alone. KEGG pathway enrichment analysis identified fatty acid biosynthesis and elongation among the most significantly enriched pathways in HRECs under combined HG-Hyp conditions, alongside amino sugar and nucleotide sugar metabolism, glycerophospholipid metabolism, the pentose phosphate pathway, and glycolysis/gluconeogenesis. Supervised machine-learning classifiers distinguished these metabolic subtypes, with AdaBoost and gradient Boosting showing the most balanced, reproducible performance across 10-fold cross-validation, LOOCV, and permutation testing, and remaining the most reliable classifiers under domain-shift testing (area under the curve = 0.88, P = 0.0061). CONCLUSIONS: In this exploratory analysis, HG, Hyp, and their combination drive metabolically and functionally distinct subtypes of Warburg-like metabolic reprogramming in HRECs, with HG-Hyp in combination producing a highly angiogenic phenotype. Boosting-based ensemble classifiers provide a promising framework for detecting these subtypes even under domain-shift conditions, warranting validation in larger independent datasets. TRANSLATIONAL RELEVANCE: Integrating metabolomics with machine-learning classification offers a strategy to identify Warburg-like metabolic subtypes in retinal endothelial cells, providing insights into angiogenic mechanisms and guiding the development of targeted diagnostics or therapeutics for PDR.

Humans

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

Quality over quantity: biopsy-anchored CT radiogenomics models outperform all-lesion training in a multi-tumour cohort despite a smaller sample size.

OBJECTIVE: Radiogenomics aims to non-invasively predict tumour genotypes from imaging, but most studies assume molecular homogeneity by assigning a single biopsy-derived label to all lesions within a patient. This approach risks substantial label noise given well-documented interlesional heterogeneity. We investigated whether anchoring training to biopsy-confirmed lesions improves radiogenomic model performance and generalisability. MATERIALS AND METHODS: We retrospectively analysed 1646 patients (11473 segmented lesions) with contrast-enhanced CT and EGFR mutation status from next-generation sequencing at the Netherlands Cancer Institute, alongside an external NSCLC radiogenomics cohort (n&#x2009;=&#x2009;158). All visible lesions were segmented, and the exact biopsy site was matched to its segmentation. Radiomic features were extracted, and machine learning models were trained with three lesion selection strategies: all lesions, non-biopsied lesions only, and biopsy-confirmed lesions only. To disentangle label quality from sample size, we created size-matched variants (one lesion per patient) for all-lesion and non-biopsied strategies. RESULTS: All models achieved significant discrimination of EGFR status on internal validation (AUC&#x2009;=&#x2009;0.62-0.68). However, performance of the all-lesion and non-biopsied models declined on external validation (AUC&#x2009;=&#x2009;0.55-0.63), while the biopsy-anchored model maintained stable performance (AUC&#x2009;=&#x2009;0.62), despite having only 1/10th of the training sample size. When training sets were size-matched, the biopsy-anchored approach significantly outperformed a model trained on all available lesions on external validation (p&#x2009;=&#x2009;0.037). CONCLUSIONS: Radiogenomic models trained on biopsy-confirmed lesions outperform conventional all-lesion strategies in external validation, despite using an order of magnitude fewer samples. Prioritising lesion-level label fidelity can mitigate heterogeneity-driven noise, enhancing robustness and clinical translation of imaging-based genomic prediction. KEY POINTS: Question Does assigning biopsy-derived molecular labels to all lesions introduce heterogeneity-driven label noise that reduces the generalisability of radiogenomic models? Findings Models trained exclusively on biopsy-confirmed lesions demonstrated superior external generalisability compared with all-lesion approaches, despite being trained on substantially fewer samples. Clinical relevance Biopsy-anchored radiogenomics improves the reliability of non-invasive mutation prediction by accounting for tumour heterogeneity, potentially supporting clinical decision-making when tissue sampling is limited or molecular results are discordant across lesions.

Humans

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data.

Long noncoding RNAs (lncRNAs) play critical regulatory roles in human development and disease. Although there are over 100,000 samples with available RNA sequencing (RNA-seq) data, many lncRNAs have yet to be annotated. The conventional approach to identifying novel lncRNAs from RNA-seq data is to find transcripts without coding potential but this approach has a false discovery rate of 30-75%. Other existing methods either identify only multi-exon lncRNAs, missing single-exon lncRNAs, or require transcriptional initiation profiling data (such as H3K4me3 ChIP-seq data), which is unavailable for many samples with RNA-seq data. Because of these limitations, current methods cannot accurately identify novel lncRNAs from existing RNA-seq data. To address this problem, we have developed software, Flnc, to accurately identify both novel and annotated full-length lncRNAs, including single-exon lncRNAs, directly from RNA-seq data without requiring transcriptional initiation profiles. Flnc integrates machine learning models built by incorporating four types of features: transcript length, promoter signature, multiple exons, and genomic location. Flnc achieves state-of-the-art prediction power with an AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the conventional approach to over 85%. Flnc is available via GitHub platform.

RNA-seq

The molecular landscape of chordoma: Current frontiers from multi-omics to artificial intelligence.

Chordoma is a rare and aggressive malignant bone tumor of the axial skeleton that has historically challenged clinicians due to its complex anatomical locations and a high recurrence rate of up to 85%. This review synthesizes the most recent advances in chordoma research and offers an overview of how multi-omics, advanced immunology, and artificial intelligence are reshaping the treatment paradigm. Central to its pathogenesis is the T-box transcription factor Brachyury, which this review highlights as both the pathognomonic diagnostic marker and the primary therapeutic vulnerability. Cutting-edge innovations targeting this driver include covalent small-molecule binders, targeted protein degradation, and peptide-centric CAR-T cells designed to attack the intracellular oncoprotein. The tumor immune microenvironment is functionally dynamic, and new dimensions in cellular therapy, such as dual-specific CAR constructs and NK-cell platforms, are being engineered to neutralize immunosuppressive factors. Beyond biological insights, the review emphasizes the role of computational biology, specifically how deep-learning and machine-learning models achieve expert-level precision in tumor segmentation and personalized survival forecasting. By integrating genomic, transcriptomic, epigenomic, and proteomic data, multiomics approaches can fully elucidate chordoma subtypes and underlying resistance mechanisms, ultimately paving the way for more precise and personalized therapeutic strategies.

Humans

Ribo-ITP enables identification of translons from limited input samples.

In the last decade, an unexpectedly large number of translated regions (translons) have been discovered using ribosome profiling and proteomics. Translons can act as regulatory elements or encode functional micropeptides. However, identification of translons has been limited to cell lines or large organs due to high input requirements for conventional ribosome profiling and mass spectrometry. Here, we address this input limitation using Ribo-ITP on difficult-to-collect samples such as microdissected hippocampal tissues and single preimplantation embryos to identify thousands of translons. To test the translational capacity of the identified translons, we engineer a translon-dependent GFP reporter system and detect expression of translons initiating at ATG and near-cognate start codons in mouse embryonic stem cells (mESCs). We identify distinct expression patterns of translons using a comparative analysis of more than a thousand ribosome profiling datasets across a wide range of cell types. Further, using a machine learning model, we predict that specific upstream translons in synaptically enriched mRNAs regulate translation efficiency of the annotated coding region. Taken together, we present a proof-of-concept study to identify non-canonical translation events from low input samples which can be applied to cell and tissue types inaccessible to conventional methods.

Animals

Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis.

PURPOSE: Given the existing uncertainties regarding the link between Di(2-ethylhexyl) phthalate (DEHP) exposure and gastric cancer (GC) progression, this study aimed to clarify their association, identify the toxic targets of DEHP, and elucidate the underlying molecular mechanisms. METHODS: Multiple integrated approaches were employed, including Gene Expression Omnibus (GEO) data analysis, network toxicology, molecular docking, and machine learning. STRING and Cytoscape tools were utilized to identify key targets, while Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the functional enrichment of intersecting targets. Machine learning and SHAP analysis were applied to screen core targets in GC. Molecular docking was performed to evaluate the binding affinity of DEHP toward core targets, and 200 ns molecular dynamics simulations were further conducted for representative complexes to validate their dynamic stability. RESULTS: A total of 18 key targets were identified using STRING and Cytoscape. GO and KEGG enrichment analyses demonstrated that these intersecting targets were primarily enriched in the extracellular region, as well as the Calcium signaling pathway and cAMP signaling pathway. Through machine learning analyses, 7 key genes (ADRB2, ESRRG, GRIA4, IL13RA2, NR3C2, PLA2G1B, and SULT2A1) were identified as core targets in GC through machine learning analyses. Molecular docking simulations revealed strong binding specificity between DEHP and the target proteins. Among them, NR3C2 and ADRB2 exhibited relatively high predictive importance in the machine learning models. DEHP showed favorable binding affinity toward these core targets, and molecular dynamics simulations further confirmed that ADRB2-DEHP and NR3C2-DEHP complexes maintained stable conformations throughout the simulation. CONCLUSIONS: Our findings identified GC associated genes that were computationally predicted as potential targets of DEHP. These results indicated structural compatibility between DEHP and its target proteins but did not prove that DEHP exposure accounts for the gene expression changes in GC.

Molecular Docking Simulation

usiGrabber: automating the curation of proteomics spectra data at scale, making large datasets ready for use in machine learning systems.

MOTIVATION: An unprecedented amount of mass spectrometry-based proteomics data is publicly available through repositories such as the PRoteomics IDEntifications Database (PRIDE), and the field is increasingly leveraging machine-learning approaches. However, the available data is not ready to be reused in a scalable way beyond the original acquisition purpose. Existing machine learning models commonly rely on a few manually curated datasets that require deep domain expertise and tedious technical work to construct. Importantly, these datasets have not been updated in recent years, so that newly published data remains inaccessible. We present usiGrabber, a scalable framework for assembling large proteomic datasets. usiGrabber is designed around portability and extensibility. It extracts spectra identification data from mzIdentML files, stores additional project-level metadata retrieved through the PRIDE API, indexes raw spectra using Universal Spectrum Identifiers (USIs), and offers download utilities to retrieve spectra data at scale. RESULTS: Within 49&#x2009;h, we parsed over 800 million peptide spectrum matches and corresponding USIs from over 1200 projects. As a proof of concept, we used usiGrabber to construct a phosphorylation-specific training dataset of nearly 11 million spectra in under 2 days and used it to retrain a binary phosphorylation classifier based on the AHLF model architecture. With a balanced accuracy of 0.78, our model achieves comparable performance to the original model on an independent test set, showing that automated data extraction is an alternative to manual curation of static datasets. AVAILABILITY AND IMPLEMENTATION: All code is available at https://github.com/usiGrabber/usiGrabber; the data are available at https://zenodo.org/records/18853258.

Machine Learning

Machine learning analysis of the human initiator region reveals key features of different types of core promoters.

The initiator (Inr) is the starting point for the transcription of many genes. Here, we generated highly predictive machine learning models of the human Inr region, and determined that the Inr is present in &#x223c;60% of focused human promoters, identified a novel TATA-specific Inr, and detected the overlapping but functionally distinct TCT motif. Quantitative genome-wide analyses revealed a strict and synergistic interaction between the Inr and DPR, an inverse relationship between the TATA and DPR, a flexible and sometimes independent function of the TATA box in relation to the Inr, and different properties of the TCT motif in humans versus Drosophila.

Humans