Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “machine learning prediction model”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Exploring predictive and reproducible modeling with the single-subject FIAC dataset.

Predictive modeling of functional magnetic resonance imaging (fMRI) has the potential to expand the amount of information extracted and to enhance our understanding of brain systems by predicting brain states, rather than emphasizing the standard spatial mapping. Based on the block datasets of Functional Imaging Analysis Contest (FIAC) Subject 3, we demonstrate the potential and pitfalls of predictive modeling in fMRI analysis by investigating the performance of five models (linear discriminant analysis, logistic regression, linear support vector machine, Gaussian naive Bayes, and a variant) as a function of preprocessing steps and feature selection methods. We found that: (1) independent of the model, temporal detrending and feature selection assisted in building a more accurate predictive model; (2) the linear support vector machine and logistic regression often performed better than either of the Gaussian naive Bayes models in terms of the optimal prediction accuracy; and (3) the optimal prediction accuracy obtained in a feature space using principal components was typically lower than that obtained in a voxel space, given the same model and same preprocessing. We show that due to the existence of artifacts from different sources, high prediction accuracy alone does not guarantee that a classifier is learning a pattern of brain activity that might be usefully visualized, although cross-validation methods do provide fairly unbiased estimates of true prediction accuracy. The trade-off between the prediction accuracy and the reproducibility of the spatial pattern should be carefully considered in predictive modeling of fMRI. We suggest that unless the experimental goal is brain-state classification of new scans on well-defined spatial features, prediction alone should not be used as an optimization procedure in fMRI data analysis.

Artifacts↗

The prediction of human oral absorption for diffusion rate-limited drugs based on heuristic method and support vector machine.

Support vector machine (SVM), as a novel machine learning technique, was used for the prediction of the human oral absorption for a large and diverse data set using the five descriptors calculated from the molecular structure alone. The molecular descriptors were selected by heuristic method (HM) implemented in CODESSA. At the same time, in order to show the influence of different molecular descriptors on absorption and to well understand the absorption mechanism, HM was used to build several multivariable linear models using different numbers of molecular descriptors. Both the linear and non-linear model can give satisfactory prediction results: the square of correlation coefficient R(2) was 0.78 and 0.86 for the training set, and 0.70 and 0.73 for the test set respectively. In addition, this paper provides a new and effective method for predicting the absorption of the drugs from their structures and gives some insight into structural features related to the absorption of the drugs.

Administration, Oral↗

Improving the reliability of medical software by predicting the dangerous software modules.

Software reliability analysis is inevitable for modern medical systems, since a large amount of medical system functionality is now dependent on software, and software does contribute to system failures. Most software reliability models are based on software failure data collected from the project. This creates a problem for the designers since, during the early stage, software failure data are not available. However, a valuable knowledge can be learned from the analysis of previous projects and applied to the new ones. This paper presents the approach that predicts the potentially dangerous software modules under development based on the analysis of the already finished modules using the machine-learning techniques. On the basis of the prediction given by our method software designers are able to devote more testing effort to the dangerous parts of the system, which results in a more reliable medical software system.

Algorithms↗

Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives.

Healthcare systems worldwide face growing challenges, including rising costs, workforce shortages, and disparities in access and quality, particularly in low- and middle-income countries. Artificial intelligence (AI) has emerged as a transformative tool capable of addressing these issues by enhancing diagnostics, treatment planning, patient monitoring, and healthcare efficiency. AI's role in modern medicine spans disease detection, personalized care, drug discovery, predictive analytics, telemedicine, and wearable health technologies. Leveraging machine learning and deep learning, AI can analyze complex data sets, including electronic health records, medical imaging, and genomic profiles, to identify patterns, predict disease progression, and recommend optimized treatment strategies. AI also has the potential to promote equity by enabling cost-effective, resource-efficient solutions in low-resource and remote settings, such as mobile diagnostics, wearable biosensors, and lightweight algorithms. Successful deployment requires addressing critical challenges, including data privacy, algorithmic bias, model interpretability, regulatory oversight, and maintaining human clinical oversight. Emphasizing scalable, ethical, and evidence-driven implementation, key strategies include clinician training in AI literacy, adoption of resource efficient tools, global collaboration, and robust regulatory frameworks to ensure transparency, safety, and accountability. By complementing rather than replacing healthcare professionals, AI can reduce errors, optimize resources, improve patient outcomes, and expand access to quality care. This review emphasizes the responsible integration of AI as a powerful catalyst for innovation, sustainability, and equity in healthcare delivery worldwide.

Humans↗

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models↗

CSGL: chemical synthesis graph learning for molecule representation.

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Machine Learning↗

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform↗

Radiogenomic MRI biomarkers for noninvasive prediction of GPC3 expression and tumor microenvironment in hepatocellular carcinoma.

BACKGROUND: Glypican-3 (GPC3) is frequently overexpressed in hepatocellular carcinoma (HCC) and plays a key role in immune and metabolic remodeling of the tumor microenvironment. Reliable noninvasive biomarkers for predicting GPC3 status could improve patient stratification and support precision immunotherapy. METHODS: This multicenter retrospective study included 274 patients with pathologically confirmed hepatocellular carcinoma from three institutions, 34 external cases with MRI from The Cancer Imaging Archive, and 363 transcriptomic profiles from The Cancer Genome Atlas. Contrast-enhanced T1-weighted imaging and diffusion-weighted imaging were analyzed. Tumor and peritumoral regions were segmented manually and radiomic features extracted using PyRadiomics. Feature selection was performed with correlation filtering and least absolute shrinkage and selection operator regression. Machine learning classifiers including logistic regression, random forest, support vector machine, k-nearest neighbor, and decision tree were trained with 10-fold cross-validation and tested on independent external cohorts. A radiomics score was calculated for each patient. Radiogenomic analysis correlated radiomics scores with transcriptomic data using weighted gene co-expression network analysis. Hub genes and enriched pathways were identified, and immune infiltration and predicted immunotherapy response were assessed using computational methods. RESULTS: The random forest model using contrast-enhanced T1-weighted imaging achieved an area under the curve of 0.966 in training and 0.935 in internal validation. The integrated contrast-enhanced T1-weighted imaging plus diffusion-weighted imaging model reached an internal validation area under the curve of 0.979. In external testing, the best performance was obtained with a support vector machine model (area under the curve 0.756). Radiomics scores were significantly correlated with GPC3 expression (R&#x2009;=&#x2009;0.78, p&#x2009;<&#x2009;0.05). Transcriptomic analysis identified a 10-gene signature enriched in hypoxia and lipid metabolism pathways that stratified patients into prognostic subgroups (concordance index 0.720, hazard ratio 4.07, p&#x2009;<&#x2009;0.0001). High-risk patients had greater immune infiltration and a lower predicted immune evasion score, suggesting a potential benefit from immunotherapy. CONCLUSIONS: MRI-based radiomics models can noninvasively predict GPC3 expression in hepatocellular carcinoma. Radiomics scores reflect underlying hypoxia and lipid metabolism pathways and stratify patients by prognosis and predicted immunotherapy response. These findings support radiogenomics as a translational approach to imaging-guided precision treatment in hepatocellular carcinoma.

Humans↗

Causal protein-signaling networks derived from multiparameter single-cell data.

Machine learning was applied for the automated derivation of causal influences in cellular signaling networks. This derivation relied on the simultaneous measurement of multiple phosphorylated protein and phospholipid components in thousands of individual primary human immune system cells. Perturbing these cells with molecular interventions drove the ordering of connections between pathway components, wherein Bayesian network computational methods automatically elucidated most of the traditionally reported signaling relationships and predicted novel interpathway network causalities, which we verified experimentally. Reconstruction of network models from physiologically relevant primary single cells might be applied to understanding native-state tissue signaling biology, complex drug actions, and dysfunctional signaling in diseased cells.

Algorithms↗

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI↗

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2&#xd7;2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I&#xb2;=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans↗

Transcriptome-wide analysis reveals potential roles of CFD and ANGPTL4 in fibroblasts regulating B cell lineage for extracellular matrix-driven clustering and novel avenues for immunotherapy in breast cancer.

BACKGROUND: The remodeling of the extracellular matrix (ECM) plays a pivotal role in tumor progression and drug resistance. However, the compositional patterns of ECM in breast cancer and their underlying biological functions remain elusive. METHODS: Transcriptome and genome data of breast cancer patients from TCGA database was downloaded. Patients were classified into different clusters by using non-negative matrix factorization (NMF) based on signatures of ECM components and regulators. Weighted Gene Co-expression Network Analysis (WGCNA) was used to identify core genes related to ECM clusters. Additional 10 independent public cohorts including Metabric, SCAN_B, GSE12276, GSE16446, GSE19615, GSE20685, GSE21653, GSE58644, GSE58812, and GSE88770 were collected to construct Training or Testing cohort, following machine learning calculating ECM correlated index (ECI) for survival analysis. Pathway enrichment and correlation analysis were used to explore the relationship among ECM clusters, ECI and TME. Single-cell transcriptome data from GSE161529 was processed for uncovering the differences among ECM clusters. RESULTS: Using NMF, we identified three ECM clusters in the TCGA database: C1 (Neuron), C2 (ECM), and C3 (Immune). Subsequently, WGCNA was employed to pinpoint cluster-specific genes and develop a prognostic model. This model demonstrated robust predictive power for breast cancer patient survival in both the Training cohort (n&#x2009;=&#x2009;5,392, AUC&#x2009;=&#x2009;0.861) and the Testing cohort (n&#x2009;=&#x2009;1,344, AUC&#x2009;=&#x2009;0.711). Upon analyzing the tumor microenvironment (TME), we discovered that fibroblasts and B cell lineage were the core cell types associated with the ECM cluster phenotypes. Single-cell RNA sequencing data further revealed that angiopoietin like 4 (ANGPTL4)+ fibroblasts were specifically linked to the C2 phenotype, while complement factor D (CFD)+ fibroblasts characterized the other ECM clusters. CellChat analysis indicated that ANGPTL4+ and CFD+ fibroblasts regulate B cell lineage via distinct signaling pathways. Additionally, analysis using the Kaplan-Meier Plotter website showed that CFD was favorable for immunotherapy response, whereas ANGPTL4 negatively impacted the outcomes of cancer patients receiving immunotherapy. CONCLUSION: We identified distinct ECM clusters in breast cancer patients, irrespective of molecular subtypes. Additionally, we constructed an effective prognostic model based on these ECM clusters and recognized ANGPTL4+ and CFD+ fibroblasts as potential biomarkers for immunotherapy in breast cancer.

Humans↗

[Fish-based assessment methods for the ecological status of aquatic systems].

A short overview of fish-based assessment methods for aquatic systems is presented. Multimetric indices, as, e.g., the index of biotic integrity (IBI), firstly developed in USA and later adapted for European river basins and other countries, are shortly described. Non-multimetric indices are also discussed, e.g. the ichthyological index (II) and the index of ecological status of fish communities (ISECI), both proposed for monitoring Italian rivers. Moreover, statistical and predicting methods based on machine learning techniques are described. Finally, a new approach for developing standardised fish-based methods useful to assess ecological status of Italian rivers is proposed. Although rather complex, the use of bony fish in biomonitoring is promising and requires a multidisciplinary approach to be adopted.

Animals↗

DNA splice site detection: a comparison of specific and general methods.

In an era when whole organism genomes are being routinely sequenced, the problem of gene finding has become a key issue on the road to understanding. For eukaryotic organisms a large part of locating the genes is accomplished by predicting the likely location of splice sites on a DNA strand. This problem of splice site location has been ap- proached using a number of machine learning or statistical methods tailored more or less specifically to the nature of the problem. Recently large margin classifiers and boosting methods have been found to give improvements over more traditional methods in a number of areas. Here we compare large margin classifiers (SVM and CMLS) and boosted decision trees with the three most common models used for splice site detection (WMM, WAM, and MDT). We find that the newer methods compare favorably in all cases and can yield significant improvement in some cases.

Algorithms↗

Tumour class prediction and discovery by microarray-based DNA methylation analysis.

Aberrant DNA methylation of CpG sites is among the earliest and most frequent alterations in cancer. Several studies suggest that aberrant methylation occurs in a tumour type-specific manner. However, large-scale analysis of candidate genes has so far been hampered by the lack of high throughput assays for methylation detection. We have developed the first microarray-based technique which allows genome-wide assessment of selected CpG dinucleotides as well as quantification of methylation at each site. Several hundred CpG sites were screened in 76 samples from four different human tumour types and corresponding healthy controls. Discriminative CpG dinucleotides were identified for different tissue type distinctions and used to predict the tumour class of as yet unknown samples with high accuracy using machine learning techniques. Some CpG dinucleotides correlate with progression to malignancy, whereas others are methylated in a tissue-specific manner independent of malignancy. Our results demonstrate that genome-wide analysis of methylation patterns combined with supervised and unsupervised machine learning techniques constitute a powerful novel tool to classify human cancers.

Algorithms↗

PicSOr: an objective test of perceptual skill that predicts laparoscopic technical skill in three initial studies of laparoscopic performance.

BACKGROUND: Laparoscopic surgery requires surgeons to infer the shape of 3-D structures, such as the internal organs of patients, from 2-D displays on a video monitor. Recent evidence indicates that the issue is not resolved by the use of contemporary 3-D camera systems. It is therefore crucial to find ways of measuring differences in aptitude for recovering 3-D structure from 2-D images, and assessing its impact on performance. Our aim was to test empirically for a relationship between laparoscopic ability and the perceptual skill of recovering information about 3-D structures from 2-D monitor displays. METHODS: Participants in three studies completed a simulated laparoscopic cutting task as well as the Pictorial Surface Orientation (PicSOr)3 Test. In studies 1 (n = 48) and 2 (n = 32) both groups were laparoscopic novices, and in study 3 (n = 34) 18 of the participants were experienced laparoscopic surgeons. FINDINGS: All three studies showed that PicSOr consistently predicted the laparoscopic performance of participants on the laparoscopic cutting task (study 1, r = 0.5, p < 0.0003; study 2, r = 0.5, p < 0.004; and study 3, r = 0.42, p = 0.017). Furthermore, it was also a significant predictor of laparoscopic surgeons' performance (r = 0.54, p = 0.047). INTERPRETATIONS: This is the first objective perceptual psychometric test to reliably predict laparoscopic technical skills. PicSOr provides a tool for assessing which trainees have the potential to learn minimal access surgery.

Adult↗

Investigating the mechanisms of PhIP-induced colorectal cancer through network toxicology, machine learning, and molecular dynamics simulation.

BACKGROUND: Over the past few years, 2-amino-1-methyl-6-phenylimidazo[4,5-b]pyridine (PhIP)- a compound from grilled or processed meats-has emerged as a major player in cancer development, especially colorectal cancer (CRC). This work dives into its potential links to CRC and uncovers the key genes that bridge this connection. METHODS: We tapped into various databases to pinpoint target genes tied to PhIP and CRC, then ran protein-protein interaction (PPI) analyses for visualization. Next, we explored underlying mechanisms through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment. To nail down predictions, we tested 107 machine learning pipelines and picked the best one, validating its accuracy and the core genes' prognostic value across datasets. Next, molecular docking and dynamics simulations probed the interactions between these genes and PhIP. Finally, cell proliferation was assessed using Cell Counting Kit-8 (CCK-8) and 5-ethynyl-2'-deoxyuridine (EdU) assays, and polymerase chain reaction (PCR) was performed to validate the expression levels of the hub genes. RESULTS: Our analysis identified 39 overlapping genes, from which a machine learning model (glmBoost + Enet) identified six candidate targets: CDK4, CEBPB, COMT, SOX9, TIMP1, and TOP2A. To prioritize these, a hierarchical screening framework was applied. Molecular docking and dynamics simulations identified CDK4, COMT, and TIMP1 as the most stable interactors with PhIP. Functional assays confirmed that PhIP treatment significantly enhanced the proliferation of CRC cells. Crucially, quantitative PCR (qPCR) validation in multiple CRC cell lines identified TIMP1 as the primary target, showing the most consistent and significant upregulation upon PhIP exposure. CONCLUSIONS: In essence, these genes drive PhIP is role in CRC, offering novel insights into its molecular pathways. This could reshape how we tackle food-related pollutants, paving the way for better prevention and targeted therapies.

Colorectal cancer (CRC)↗

A validated, modifiable proteomic score from the EXSCEL trial predicts cardiovascular events in diabetes.

BACKGROUNDAdults with type 2 diabetes mellitus (T2DM) are at increased risk for stroke, myocardial infarction, and cardiovascular death, yet individual risk is heterogeneous and incompletely captured by clinical models.METHODSIn the Exenatide Study of Cardiovascular Event Lowering (EXSCEL), adults with T2DM were randomized to a GLP-1 RA (exenatide) or a placebo and followed longitudinally for major adverse cardiovascular events (MACE). High-throughput discovery proteomics was done in plasma collected at baseline and 12 months. Proteins associated with time to MACE were identified using multivariable regression and incorporated into supervised machine learning models. A multi-protein score was developed and externally validated in 2 independent population-based and trial cohorts.RESULTSThe proteomic score showed incremental improvement in cardiovascular risk discrimination beyond clinical factors alone, and several proteins were consistently prioritized across modeling approaches. The protein score and a top-ranked protein, tetranectin, were modified by GLP-1 RA treatment, and a decrease in protein score was associated with improved outcomes, supporting modifiability of MACE risk.CONCLUSIONExternal validation confirmed generalizability across cohorts with and without diabetes. Together, these findings demonstrate that plasma proteomic signatures can enhance cardiovascular risk stratification and identify treatment-responsive biomarkers in T2DM, supporting their potential role in precision prevention strategiesFUNDINGThe EXSCEL study was funded by Amylin Pharmaceuticals. This research was supported by contracts HHSN268201200036C, HHSN268200800007C, HHSN268201800001C, N01HC55222, N01HC85079, N01HC85080, N01HC85081, N01HC85082, N01HC85083, N01HC85086, 75N92021D00006, and grants R01HL146145, U01HL080295, U01HL130114, R01HL172803, and R01HL144483 from the National Heart, Lung, and Blood Institute, with additional contribution from the National Institute of Neurological Disorders and Stroke. Additional support was provided by R01AG023629 from the National Institute on Aging.

Aged↗