Search PubMedSearch

SEARCH · Search PubMed

Results for “Machine learning model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Multi-level Transcriptomic and Machine-learning Analyses Identify MZT1 as a Proliferation-associated Prognostic Marker in Lung Adenocarcinoma.

BACKGROUND/AIM: Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity and variable clinical outcomes, highlighting the need for biomarkers that reflect core tumor biological processes. Centrosome-associated proteins regulate mitotic fidelity and genome stability, yet their roles in LUAD remain incompletely defined. In this study, we systematically characterized mitotic spindle organizing protein 1 (MOZART1; MZT1) and related family members in LUAD. MATERIALS AND METHODS: We performed integrated analyses combining bulk transcriptomic datasets, survival modeling, gene set enrichment, immune deconvolution, machine-learning based prognostic modeling, and single-cell RNA sequencing. Expression patterns and clinical associations of MZT family genes were evaluated across pan-cancer and LUAD cohorts. RESULTS: MZT family genes were consistently upregulated in tumor tissues, with MZT1 showing the most robust expression pattern. Elevated MZT1 expression was significantly associated with reduced overall survival. Functional analyses revealed coordinated activation of proliferative and genome maintenance pathways, including G2/M checkpoint regulation, E2F and MYC signaling, and DNA repair. A multivariable analysis indicated that the prognostic association of MZT1 was reduced after adjusting for canonical proliferation markers, suggesting partial overlap with established proliferation signals. The LASSO-based Cox model demonstrated stable time-dependent predictive performance at 1-, 3-, and 5-year survival. Immune analyses indicated associations between MZT1 expression and tumor microenvironmental features. Single-cell analysis showed that MZT1 expression was predominantly enriched in malignant epithelial cells and associated with proliferative cellular states. Protein-level validation supported concordance with transcriptomic findings. CONCLUSION: MZT1 is a proliferation-associated marker that integrates clinical risk, transcriptional programs, cellular heterogeneity, and predictive modeling in LUAD, providing a potential framework for biomarker development and risk stratification.

Humans

Spatiotemporal genomic analysis and risk assessment of the plasmids carrying blaOXA-48-like genes based on a large-scale international dataset.

BACKGROUND: The spread of OXA-48-like carbapenemases represents a major public health challenge. Although previous studies have investigated OXA-48-like carbapenemases risk factors, nosocomial dissemination, and plasmid dynamics, an integrated plasmid-centered framework combining complete plasmid mining, transmission-unit analysis, phylogenetic reconstruction, and machine learning-based risk assessment remains limited. METHODS: We systematically collected 747 complete plasmid sequences carrying blaOXA-48-like genes from the NCBI database, establishing the largest collections of complete plasmid sequences to date. Using an integrative framework of population genomics, phylogenetic dating, and machine learning, this study aimed to characterize the dissemination patterns, plasmid replicon diversity, transmission units, mobile genetic elements, co-resistance profiles, and risk classification of these plasmid. RESULTS: Plasmids carrying blaOXA-48-like genes were detected across 50 countries on six continents, with blaOXA-48 predominating in Europe, blaOXA-181 in South Asia, and blaOXA-232 largely in Asia. IncL and ColKP3/IncX3 replicons, together with Tn1999.2 and other MGEs, were central drivers of plasmid maintenance and spread. Sixteen transmission units were defined, with AA068_Cluster3 estimated to have originated in the Netherlands around 2005 before expanding to Europe, the Middle East, Asia, and North America. Co-resistance analyses revealed frequent modules involving aminoglycoside and quinolone resistance, with qnrS1 and aph(3'')-Ib most prevalent. Notably, high-risk transposon structures were often identified in non-clinical environments, underscoring their cross-ecological transmission potential. Machine learning-based classification models showed good internal performance for predefined composite-risk categories, with plasmid mobility, clinical/non-clinical source composition, and host background contributing to the classification results. CONCLUSIONS: This study provides a large-scale plasmid-centered genomic analysis of publicly available complete plasmid sequences carrying blaOXA-48-like genes, integrating transmission-unit inference, phylogeographic reconstruction, mobile genetic element and co-resistance profiling, and composite genomic risk stratification. This gene-centered framework may support future One Health-oriented antimicrobial resistance surveillance and prioritization of plasmids with higher dissemination and resistance potential.

Plasmids

Predicting training outcomes for developmental dyslexia from EEG data.

Developmental dyslexia (DD) is characterised by lower-than-average reading abilities and is diagnosed in approximately 10% of individuals. The societal barriers may limit professional fulfilment and psychological wellbeing of individuals with DD, calling for the development of effective interventions to counteract them. As DD is associated with challenges in both phonological and visuo-attentional domains, different longitudinal training approaches were developed to strengthen them. However, they require a considerable amount of personal, social and economic resources and the outcomes may vary depending on individual differences in behavioural and neurophysiological functionality. Hence, predicting training outcomes might help in developing personalised treatment protocols and optimising the use of resources. In the present work we applied machine learning to resting-state EEG to predict longitudinal training outcomes in adults with DD enrolled in a randomized clinical trial. In particular, one group received a visuo-attentional training combined with transcranial alternating current stimulation (tACS), another group received visuo-attentional training with sham/placebo stimulation, and the third group received a phonological training with sham/placebo stimulation. The improvement in text reading speed was associated with spectral power in low-beta and individual frequencies in the alpha (IAF) and beta (IBF) bands, while the improvement in pseudoword reading was associated with IBF. The findings highlight the potential of capturing neural markers of treatment responsiveness in DD. Future studies should focus on the generalisability of predictive models to real-world settings, while investigating whether specific EEG markers predict responsiveness to distinct remediation protocols, thus supporting the development of personalised interventions.

Humans

Proteomic profiling of bone for the estimation of post-mortem interval and post-mortem submersion interval: a systematic review.

Accurate estimation of the Post-Mortem Interval (PMI) and Post-Mortem Submersion Interval (PMSI) remains a persistent challenge in forensic science, especially when traditional morphological and entomological methods fail due to advanced decomposition or in aquatic environments. Proteomic profiling of bone tissues has recently emerged as a promising approach, leveraging the predictable degradation patterns of bone proteins to estimate time since death more reliably. This systematic review, conducted in accordance with PRISMA guidelines, analyzed 24 peer-reviewed studies focusing on the application of proteomic techniques to bone tissue for PMI and PMSI estimation. The included studies were evaluated based on sample type, analytical techniques used, identified biomarkers, environmental conditions assessed, and the overall reliability and reproducibility of the findings. The review found that specific bone proteins, particularly collagen, osteocalcin, fetuin-A, etc. exhibited consistent degradation patterns that correlated strongly with elapsed post-mortem time. Cortical bone was identified as a more stable and informative matrix compared to trabecular bone. Mass spectrometry, especially LC-MS/MS, emerged as the predominant analytical technique due to its high sensitivity and accuracy in detecting low-abundance proteins over extended PMIs and PMSIs. However, protein degradation rates were significantly influenced by environmental variables such as temperature, humidity, soil pH, and microbial activity. This review also emphasizes the transformative role of bone proteomics in advancing forensic science while identifying key gaps that must be addressed to achieve global standardization and practical implementation in diverse forensic contexts. The integration of proteomics with other emerging technologies, such as machine learning algorithms and computational modeling, may further enhance the precision of PMI and PMSI estimation in future applications.

Postmortem Changes

Engineering Bacillus Subtilis for Efficient Biosynthesis of Riboflavin: Current Knowledge and Future Perspectives.

Riboflavin is an essential water-soluble vitamin that serves as a precursor for the biosynthesis of the flavin cofactors FMN and FAD, which play pivotal roles in numerous redox and energy metabolism reactions. With the growing global demand for sustainable vitamin production, microbial fermentation has become an attractive alternative to chemical synthesis due to its environmental and economic advantages. Among microbial hosts, Bacillus subtilis has emerged as a leading cell factory for riboflavin production owing to its GRAS status, well-characterized genetics, and efficient protein secretion system. This review provides a comprehensive overview of recent advances in metabolic engineering strategies to enhance riboflavin biosynthesis in B. subtilis. Key topics include strengthening biosynthetic and precursor pathways, relieving feedback inhibition, balancing metabolic flux and cell growth, employing adaptive laboratory evolution, and utilizing omics-guided optimization and 13C metabolic flux analysis. Moreover, the integration of synthetic biology tools such as riboswitch engineering, regulatory element design, and high-throughput screening has significantly accelerated strain improvement. Despite remarkable progress, challenges remain in achieving precise regulatory control, optimizing multi-gene expression, and enhancing genome integration efficiency. Future research combining multi-omics data, synthetic regulatory design, and machine learning-driven predictive modeling is expected to further advance the development of intelligent B. subtilis cell factories. However, the practical implementation of these systems remains constrained by the metabolic burden of overproduction and the lack of universal regulatory models that can predict strain performance across varying industrial scales.

Bacillus subtilis

Clinical translation of senescence-related pan-cancer multi-omics: tools for assessment and immunotherapy prediction.

Cellular senescence (CS) exerts dual roles in tumorigenesis, yet its pan-cancer molecular characteristics and clinical value remain unclear, hindering its translation to oncology and personalized therapy. To address the lack of specific and universal tools for senescence assessment and immunotherapy response prediction, this study systematically analyzed 1259 CS-related genes from the CellAge database across 31 cancer types by integrating multi-omics data, including bulk RNA-seq, single-cell/spatial transcriptomics, and CRISPR screening. We developed a rank-based algorithm SenScoreR (publicly available at https://gxhub.shinyapps.io/SenScoreR/ ) for senescence quantification, validated with 10 independent datasets, and constructed a machine learning-based predictive model CS.Sig for immunotherapy response. Results showed that tumors had significantly lower Rank-based Senescence Score (RSS) than normal tissues across 31 cancers (average diagnostic AUC = 0.895), with low RSS linked to poor survival; high RSS correlated with reduced genomic instability, enriched CD8⁺ T/NK cell/macrophage infiltration, upregulated PD-L1 expression, and elevated immune cytolytic activity. CS.Sig demonstrated robust performance in predicting ICI response (AUC = 0.716 across 10 cohorts), outperforming 13 existing signatures, while CRISPR screening identified 17 senescence-related targets (e.g., CEP55, PPP1CC) whose knockout enhanced anti-tumor immunity. Our findings clarify CS's role in maintaining tumor genomic stability and shaping immune microenvironments, and the developed SenScoreR, CS.Sig, and identified targets bridge basic CS research with clinical oncology, providing a translational resource and hypothesis basis for future experimental and clinical validation.

Journal Article

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans

Discovery and validation of a multi-protein panel for predicting non-fatal major adverse cardiovascular events in diabetic kidney disease.

OBJECTIVE: To identify plasma protein biomarkers associated with incident non-fatal major adverse cardiovascular events (MACE) in diabetic kidney disease (DKD) patients. RESEARCH DESIGN AND METHODS: We analyzed 317 DKD patients from the UK Biobank. Plasma proteomics and clinical data (demographics, metabolism, renal function) were integrated. In an exploratory discovery phase, three sequential Cox regression models (crude, socio-demographic-adjusted, socio-demographic-metabolic adjusted) screened non-fatal MACE-associated proteins. To prevent information leakage, the cohort was then randomly split into training (70%) and testing (30%) sets; machine-learning feature selection, hyperparameter optimization, and final model development were performed exclusively within the training set. The associated proteins were input into the four-step machine-learning pipeline (LASSO-Cox, random survival forest, Boruta, XGBoost-Cox). Predictive performance was validated using Kaplan-Meier survival analyses, longitudinal trajectory modeling, and ROC benchmarking. An interactive web application was deployed for clinical implementation. RESULTS: Of 1,463 plasma proteins, 561 were associated with non-fatal MACE across Cox models, with 14 overlapping proteins. Nine core proteins (ANG, IL1R1, CXCL14, ESAM, PTGDS, HAVCR1, FGFR2, IGSF8, CCL3) were validated: ANG showed the strongest non-fatal MACE association (HR&#xa0;=&#xa0;3.88, 95%CI 2.33-6.48, p<0.001), and all high-expression groups had elevated non-fatal MACE risk. GO/KEGG enrichment highlighted inflammatory-immune pathways like positive regulation of MAPK cascade, Cytokine-cytokine receptor interaction and PI3K-Akt signaling pathway as key mechanisms. The model integrating proteins, demographic factors, and clinical variables achieved the highest predictive performance across non-fatal MACE (AUC&#xa0;=&#xa0;0.768), myocardial infarction (MI) (0.808), and stroke (0.816) outcomes, with superior stability in cross-validation. CoxBoost + Elastic Net framework was selected as the optimal framework via benchmarking of 101 algorithms. The model demonstrated favorable calibration in high-risk patients and yielded positive net clinical benefit across decision thresholds of 5% to 45%. The web tool (https://jiangli2941.github.io/MACE-prediction-v2/) enables input of 28 variables, outputs non-fatal MACE risk status, risk probability, and highlights abnormal indicators. CONCLUSION: Plasma proteomics combined with machine learning identifies robust non-fatal MACE predictors in DKD.

Humans

A shape-based machine learning tool for drug design.

Building predictive models for iterative drug design in the absence of a known target protein structure is an important challenge. We present a novel technique, Compass, that removes a major obstacle to accurate prediction by automatically selecting conformations and alignments of molecules without the benefit of a characterized active site. The technique combines explicit representation of molecular shape with neural network learning methods to produce highly predictive models, even across chemically distinct classes of molecules. We apply the method to predicting human perception of musk odor and show how the resulting models can provide graphical guidance for chemical modifications.

Algorithms

Non-destructive prediction of lead content in oilseed rape leaves by fluorescence hyperspectral technology based on neural network.

Based on fluorescence hyperspectral imaging (FHSI), this study targeted rapid, non-destructive quantification of lead (Pb) content in oilseed rape leaves treated with varying silicon (Si) concentrations, acquiring fluorescence spectra over the 484.43-1001.61&#xa0;nm wavelength range. To optimize spectral data quality, preprocessing methods (Savitzky-Golay smoothing, first derivative, detrending) were comprehensively compared. Characteristic wavelengths were then selected via interval variable iterative shrinkage, which effectively compressed data dimensionality and reduced computational load. A hybrid SE-CL1DA model, fusing a 1D convolutional neural network, a long short-term memory network and SE attention mechanism was constructed, with Bayesian optimization tuning hyperparameters to boost stability. The BO-SE-CL1DA outperformed both traditional machine learning and insufficiently optimized deep learning model (Rp2=0.9609, RMSE&#xa0;=&#xa0;0.0377&#xa0;mg/kg, RPD&#xa0;=&#xa0;5.1736), thus enabling accurate Pb estimation, supporting Si-regulated heavy metal stress management and facilitating agricultural contamination monitoring.

Plant Leaves

Instrumented Walkway Gait Analysis Predicts Fallers in Neurological Disorders: Identifying Digital Biomarkers for Balance Monitoring.

Assessing balance is crucial in neurological rehabilitation, yet while wearable sensors enable real-world monitoring, identifying reliable digital biomarkers remains challenging. This study utilized a high-fidelity instrumented walkway to determine which gait parameters best predict balance impairment, providing robust targets for future wearable applications. We analyzed 49 steady-state gait metrics from 140 individuals with diverse neurological conditions. Using statistical analysis and machine learning, we evaluated these parameters against objective force plate sway scores and clinical fall-history labels. Group analysis identified 16 parameters significantly distinguishing fallers from non-fallers, and a neural network classified fallers with an area under the curve of 0.75. Across all analytical approaches, overall gait variability, e.g., Stride Width S.D. and the Gait Variability Index, emerged as a universal predictor of balance impairment and fall risk. Furthermore, while traditional linear models emphasized spatial postural control, machine learning classification uniquely identified inter-limb asymmetry as a premier driver of fall prediction. These findings indicate that instrumented gait analysis effectively identifies digital biomarkers for balance deficits. Isolating these specific metrics provides a clear blueprint for meaningful metrics required for continuous objective monitoring and future development of personalized, adaptive rehabilitation strategies.

Humans

Artificial Intelligence and Machine Learning Applications in Fibromuscular Dysplasia: Transforming Diagnosis, Risk Stratification, and Clinical Decision-Making.

Fibromuscular dysplasia (FMD) is a non-atherosclerotic vascular disorder with heterogeneous presentations, making diagnosis and management highly dependent on imaging and clinical expertise. This narrative review examines how artificial intelligence (AI) and machine learning (ML) are transforming FMD care. AI-enhanced imaging, particularly convolutional neural network-based analysis, improves detection of the characteristic "string-of-beads" pattern on CT angiography, magnetic resonance angiography, and ultrasound, although FMD-specific validation remains limited. ML models facilitate risk stratification, prediction of disease progression, and early identification of complications such as aneurysms and stroke by integrating clinical, imaging, and genomic data. AI-driven clinical decision support systems further enable personalized treatment selection through pharmacogenomic insights and robot-assisted interventions. Despite promising real-world applications, challenges persist, including limited large-scale datasets, workflow integration, regulatory barriers, and algorithmic bias affecting underrepresented populations. Future advances in explainable AI, federated learning, and digital health integration may enable a shift toward predictive, patient-centered FMD management.

Humans

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform

An encyclopedia of human enhancer-gene regulatory interactions.

Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92&#x2009;million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study&#xa0;(GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that&#xa0;guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.

Humans

KAVAS-2: Knowledge Acquisition, Visualization and Assessment System.

The objective of KAVAS-2 is the development of a tool, named KAVIAR, with which domain experts can make their knowledge explicit. It contains components for (computer assisted) knowledge elicitation and for machine learning. A key issue in KAVAS is the assessment of the quality of the classification and domain models built. Various quality measures are available and implemented in KAVIAR to assess the quality of models, specifically those developed from data bases by machine learning techniques.

Computer Simulation

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (&#x2264;&#x2009;12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n&#x2009;=&#x2009;121, 19 events) for training and centers 2-7 (n&#x2009;=&#x2009;207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans