Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Machine learning for survival analysis: a case study on recurrence of prostate cancer.

Machine learning techniques have recently received considerable attention, especially when used for the construction of prediction models from data. Despite their potential advantages over standard statistical methods, like their ability to model non-linear relationships and construct symbolic and interpretable models, their applications to survival analysis are at best rare, primarily because of the difficulty to appropriately handle censored data. In this paper we propose a schema that enables the use of classification methods--including machine learning classifiers--for survival analysis. To appropriately consider the follow-up time and censoring, we propose a technique that, for the patients for which the event did not occur and have short follow-up times, estimates their probability of event and assigns them a distribution of outcome accordingly. Since most machine learning techniques do not deal with outcome distributions, the schema is implemented using weighted examples. To show the utility of the proposed technique, we investigate a particular problem of building prognostic models for prostate cancer recurrence, where the sole prediction of the probability of event (and not its probability dependency on time) is of interest. A case study on preoperative and postoperative prostate cancer recurrence prediction shows that by incorporating this weighting technique the machine learning tools stand beside modern statistical methods and may, by inducing symbolic recurrence models, provide further insight to relationships within the modeled data.

Artificial Intelligence↗

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning↗

Predictive models for protein crystallization.

Crystallization of proteins is a nontrivial task, and despite the substantial efforts in robotic automation, crystallization screening is still largely based on trial-and-error sampling of a limited subset of suitable reagents and experimental parameters. Funding of high throughput crystallography pilot projects through the NIH Protein Structure Initiative provides the opportunity to collect crystallization data in a comprehensive and statistically valid form. Data mining and machine learning algorithms thus have the potential to deliver predictive models for protein crystallization. However, the underlying complex physical reality of crystallization, combined with a generally ill-defined and sparsely populated sampling space, and inconsistent scoring and annotation make the development of predictive models non-trivial. We discuss the conceptual problems, and review strengths and limitations of current approaches towards crystallization prediction, emphasizing the importance of comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high throughput crystallography initiatives, exchange of information and standardization should be encouraged, aiming to effectively integrate data mining and machine learning efforts into a comprehensive predictive framework for protein crystallization. Similar experimental design and knowledge discovery strategies should be applied to valid analysis and prediction of protein expression, solubilization, and purification, as well as crystal handling and cryo-protection.

Bayes Theorem↗

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29 709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85) and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans↗

Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics.

BACKGROUND: Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction. METHODS: Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32 330) and internal validation (n=13 857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel. RESULTS: Across cohorts, the median age was 58 years and ∼45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information. CONCLUSIONS: Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.

Humans↗

PathMED: an R toolkit for single-sample molecular scoring and machine learning with omics data.

MOTIVATION: Molecular scoring is a popular approach for studying pathway-level functional alterations with omics data. Using molecular scores for tasks such as single-sample molecular characterisation, phenotype prediction or disease stratification has several advantages compared to using omics data directly. Molecular scores provide biological interpretability and are more generalisable across datasets, facilitating data integration and machine learning applications. However, numerous scoring methods are available through different software packages, and currently there is a lack of tools to easily use these scores for model training and prediction. RESULTS: We developed pathMED, an R/Bioconductor package that unifies various scoring methods in a simple framework. Furthermore, pathMED also contains a machine learning module to train and test models that use the calculated molecular scores to predict clinical outcomes. We demonstrate some of its potential applications in three use cases using public omics data. We showed the generalisability of machine learning models trained on transcriptomic scores in predicting clinical outcomes when deploying on proteomic scores. We also demonstrated the application of transcriptomics scores in predicting breast cancer treatment response and identifying pathways strongly associated to tumour biology and treatment response. Finally, we demonstrated the benefit of integrating a novel gene set dissection step into the analysis pipeline to resolve disease heterogeneity at the pathway level. AVAILABILITY: PathMED is freely available in the Bioconductor repository (https://bioconductor.org/packages/release/bioc/html/pathMED.html). Code to reproduce the analyses is publicly available at https://github.com/GENyO-BioInformatics/pathMED_article.

Software↗

Predicting food taste with bound-driven optimization.

The prediction of sensory attributes from ingredient-level formulations is an emerging challenge at the intersection of food science and artificial intelligence. We address the fundamental question of whether the taste of a food can be predicted from its ingredients by treating recipes as composite materials. We apply Hashin-Shtrikman (HS) and Reuss-Voigt (RV) bounds, techniques originally developed for elastic moduli, as a null-hypothesis additive baseline for five taste dimensions (sweetness, sourness, bitterness, umami, saltiness) on a curated dataset of 70 recipes decomposed into 115 distinct ingredients scored against a library of 209 ingredient-level taste references with trained-panel ground truth. This baseline systematically under-predicts perceived taste: 77% of actual taste values exceeded the HS upper bound, with the exceedance rate ranging from 26% (bitterness) to 97% (saltiness). We traced this gap to specific processing chemistry (Maillard reactions, caramelization, evaporative concentration, protein hydrolysis, and nucleotide synergy) and introduced a hybrid model that augments the HS baseline with eight chemistry-proxy features encoding these mechanisms. Our results show that our interpretable hybrid model eliminates the systematic bias and reduces mean absolute error by 27%-62% for sweetness, sourness, umami, and saltiness while using only 10 interpretable features, achieving performance comparable to a black-box Lasso regression on 115 per-ingredient features. We further demonstrate constrained inverse design via Differential Evolution, recovering ingredient formulations that match target taste profiles subject to compositional bounds. Our work demonstrates how key chemical processes during food preparation can inform and augment physics-based and machine learning models, providing a quantitative fingerprint of processing chemistry's contribution to taste perception and paving the way for model-driven food formulation with targeted sensory characteristics.

Composite material bounds↗

ProMeta: a meta-learning framework for robust disease diagnosis and prediction from plasma proteomics.

MOTIVATION: The plasma proteome offers a dynamic window of human health, capturing the real-time intersections between genetics and physiology. However, the application of deep learning to proteomics is currently hindered by a reliance on large-scale labeled datasets, rendering standard models ineffective for rare or novel diseases where patient samples are inherently scarce. RESULTS: Here, we present ProMeta, a meta-learning framework designed to enable robust disease modeling under extreme data restrictions. By integrating knowledge-guided pathway encoding with bi-level meta-optimization, ProMeta projects unstructured proteomic profiles into biologically interpretable functional tokens. This architecture allows the model to learn a global initialization containing transferable biological priors from biobank-scale data, facilitating rapid adaptation to novel tasks. Through comprehensive benchmark experiments, ProMeta consistently outperformed transfer learning and traditional machine learning baselines in both disease diagnosis and prediction tasks. In the most challenging 4-shot scenarios (utilizing only 2 cases and 2 controls), the model achieved robust generalization with an average AUROC of ∼0.69, representing a 24.6% relative improvement over the best-performing baseline methods. Mechanistic investigation revealed that ProMeta disentangles cases from controls in the latent space prior to task-specific adaptation, confirming the acquisition of universal biological rules rather than rote memorization. Furthermore, gradient-based interpretation identified disease-specific protein biomarkers and functional pathways consistent with known pathophysiology. Collectively, ProMeta overcomes the data-scarcity bottleneck in precision medicine, providing a scalable, interpretable framework for characterizing the full spectrum of human diseases, particularly for rare conditions lacking extensive clinical cohorts. AVAILABILITY AND IMPLEMENTATION: The source code of ProMeta is available at GitHub (https://github.com/lihan97/ProMeta).

Proteomics↗

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements↗

Computational models for predicting interactions with cytochrome p450 enzyme.

Cytochrome p450 (CYP) enzymes are predominantly involved in Phase 1 metabolism of xenobiotics. As only 6 isoenzymes are responsible for approximately 90 % of known oxidative drug metabolism, a number of frequently prescribed drugs share the CYP-mediated metabolic pathways. Competing for a single enzyme by the co-administered therapeutic agents can substantially alter the plasma concentration and clearance of the agents. Furthermore, many drugs are known to inhibit certain p450 enzymes which they are not substrates for. Because some drug-drug interactions could cause serious adverse events leading to a costly failure of drug development, early detection of potential drug-drug interactions is highly desirable. The ultimate goal is to be able to predict the CYP specificity and the interactions for a novel compound from its chemical structure. Current computational modeling approaches, such as two-dimensional and three-dimensional quantitative structure-activity relationship (QSAR), pharmacophore mapping and machine learning methods have resulted in statistically valid predictions. Homology models have been often combined with 3D-QSAR models to impose additional steric restrictions and/or to identify the interaction site on the proteins. This article summarizes the available models, methods, and key findings for CYP1A2, 2A6, 2C9, 2D6 and 3A4 isoenzymes.

Computational Biology↗

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans↗

A multi-scale fusion model based on multi-phase contrast-enhanced CT for predicting pancreatic cancer resectability.

Purpose.Develop a multi-scale fusion model (MSFM) based on multi-phase contrast-enhanced computed tomography (CECT) to predict pancreatic cancer (PC) resectability, thereby assisting expert decision-making.Methods.This retrospective study enrolled 280 patients with PC from four institutions, which were randomly divided into a training cohort (202 patients) and an independent test cohort (78 patients). Three-phase CECT images (arterial, venous, and delayed phases) were used for modeling. The MSFM comprises two sub-networks: (1) a multi-phase fusion network for extracting cross-phase shared fusion features, (2) a phase-specific branch network for capturing phase-specific features; and a post-fusion strategy to generate the final predictive score by integrating the shared fusion features and three groups of phase-specific features. Additionally, a human-machine fusion deep learning model (HMfDL) was constructed by fusing the predictive score of the MSFM with expert assessments.Results.In the independent test, the MSFM achieved an AUC (area under the receiver operating characteristic curve) of 0.8385 (95% CI: 0.7521-0.9249), accuracy of 84.62%, sensitivity of 72.00%, and specificity of 90.57%. This performance outperformed single-phase models (AUC range: 0.7638-0.7781), two-phase models (AUC range: 0.7826-0.7864), and ten states-of-the-art classifiers (AUC range: 0.7404-0.7796). The HMfDL further improved the performance, reaching an AUC of 0.8626 (95% CI: 0.7853-0.9400), accuracy of 91.03%, sensitivity of 80.00%, and specificity of 96.23%. Notably, the HMfDL corrected 58.82% of misdiagnosis made by experts.Conclusions. The MSFM effectively fuses multi-phase CECT to enable highly accurate predictions of PC resectability, and provides valuable support for expert decision-making through HMfDL.

Humans↗

Artificial intelligence in treatment prediction for skeletal Class III malocclusion: A systematic review.

In skeletal Class III patients, treatment options range from orthodontics to orthognathic surgery. Choosing the optimal approach requires a comprehensive clinical evaluation, which may be supported by AI tools. The aim of this study was to assess the performance of AI models in predicting the need for orthognathic surgery and in identifying predictors influencing treatment decisions. A PRISMA-guided electronic database search (PubMed, Web of Science; 2009-2024; English/French) was performed to identify studies using machine learning (ML) or deep learning (DL) on cephalometric and clinical data. After screening and assessment for eligibility, 15 studies were critically appraised. Model performance was summarized using accuracy, sensitivity, specificity, and the area under the curve (AUC). ML algorithms (particularly Random Forest and XGBoost) and DL models (ResNet-based convolutional neural networks (CNNs)) achieved high accuracy for predicting surgical need. Frequently selected predictors included Wits appraisal, ANB angle, the maxillomandibular ratio (Mx/Md), overjet, and the divergence of the lower gonial angle. AI methods show promise for assisting treatment decisions in Class III malocclusion, with Random Forest and XGBoost performing well on tabular cephalometric data and CNNs on imaging. Larger, multicentre datasets and external validation are needed to improve reliability, address bias, and support clinical implementation.

Humans↗

Neural pattern formation via a competitive Hebbian mechanism.

In this contribution we investigate a simple pattern formation process [9,10] based on Hebbian learning and competitive interactions within cortex. This process generates spatial representations of afferent (sensory) information which strongly resemble patterns of response properties of neurons commonly called brain maps. For one of the most thoroughly studied phenomena in cortical development, the formation of topographic maps, orientation and ocular dominance columns in macaque striate cortex, the process, for example, generates the observed patterns of receptive field properties including the recently described correlations between orientation preference and ocular dominance. Competitive Hebbian learning has not only proven to be a useful concept in the understanding of development and plasticity in several brain areas, but the underlying principles have have been successfully applied to problems in machine learning [22]. The model's universality, simplicity, predictive power, and usefulness warrants a closer investigation.

Animals↗

Splicing-site recognition of rice (Oryza sativa L.) DNA sequences by support vector machines.

MOTIVATION: It was found that high accuracy splicing-site recognition of rice (Oryza sativa L.) DNA sequence is especially difficult. We described a new method for the splicing-site recognition of rice DNA sequences. METHOD: Based on the intron in eukaryotic organisms conforming to the principle of GT-AG, we used support vector machines (SVM) to predict the splicing sites. By machine learning, we built a model and used it to test the effect of the test data set of true and pseudo splicing sites. RESULTS: The prediction accuracy we obtained was 87.53% at the true 5' end splicing site and 87.37% at the true 3' end splicing sites. The results suggested that the SVM approach could achieve higher accuracy than the previous approaches.

Algorithms↗

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer’s disease↗

A multimarker model to predict outcome in tamoxifen-treated breast cancer patients.

PURPOSE: This study was designed to produce a model to predict outcome in tamoxifen-treated breast cancer patients based on clinicopathologic features and multiple molecular markers. EXPERIMENTAL DESIGN: This was a retrospective study of 324 stage I to III female breast cancer patients treated with tamoxifen for whom standard clinicopathologic data and tumor tissue microarrays were available. Nine molecular markers were studied by semiquantitative immunohistochemistry and/or fluorescence in situ hybridization. Cox proportional hazards analysis was used to determine the contributions of each variable to disease-specific and overall survival, and machine learning was used to produce a model to predict patient outcome. RESULTS: On a univariate basis, the following features were significantly associated with worse survival: high pathologic tumor or nodal class, histologic grade, epidermal growth factor receptor, ERBB2, MYC, or TP53; absent estrogen receptor (ER) or progesterone receptor; and low BCL2. CCND1 and CDKN1B did not reach statistical significance. On a multivariate basis, nodal class, ER, and MYC were statistically significant as independent factors for survival. However, the benefit of ER-positive status was moderated by BCL2, ERBB2, and progesterone receptor. BCL2 and TP53 also interacted as an independent risk factor. A kernel partial least squares polynomial model was developed with an area under the receiver operating characteristic curve of 0.90. CONCLUSIONS: Our data show the predictive value of BCL2, ERBB2, MYC, and TP53 in addition to the standard hormone receptors and clinicopathologic features, and they show the importance of conditional interpretation of certain molecular markers. Our multimarker predictive model performed significantly better than standard guidelines.

Aged↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗