Search PubMedSearch

SEARCH · Search PubMed

Results for “Predictive”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n = 907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n = 35), colorectal cancer (n = 21), and pancreatic cancer (n = 9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans

Polygenic Risk Scores for Preeclampsia Prediction Beyond Gold-Standard Clinical Models in Multiethnic Populations.

BACKGROUND: Preeclampsia is a major cause of maternal and fetal mortality and morbidity. Early risk stratification enables timely preventative therapy in high-risk women. Polygenic risk scores (PGS) improve prediction in complex diseases, but their added value for preeclampsia remains unclear, particularly in comparison to gold-standard first-trimester prediction models and across non-European ancestries. METHODS: We evaluated the performance of both a preeclampsia and systolic blood pressure PGS in 2 prospective pregnancy cohorts with detailed phenotyping: the Fetal Medicine Foundation study (n=5207; 2127 cases) and the Pregnancy Outcome Prediction study (n=3659; 228 cases). Risk models included (1) clinical factors; (2) clinical factors plus PGS; (3) advanced model including first-trimester mean arterial pressure, PAPP-A (pregnancy-associated plasma protein-A), and uterine artery pulsatility index; and (4) advanced model plus PGS. Discriminative performance, measured by the area under the receiver operating characteristic curve, was assessed overall and by ancestry. RESULTS: The preeclampsia PGS was independently associated with preeclampsia (odds ratio per SD, 1.24 [95% CI, 1.17-1.31]; P<0.001). It modestly improved prediction over clinical models (area under the receiver operating characteristic curve 0.746 versus 0.750; P=0.017) but not over the advanced model (area under the receiver operating characteristic curve 0.817 versus 0.818; P=0.326). The systolic blood pressure PGS showed stronger performance, improving prediction over both models in women of European ancestry. No improvement was observed with either score in women of African ancestry. CONCLUSIONS: PGSs for preeclampsia and SBP provide modest added predictive value beyond clinical risk factors in European ancestry women. Limited utility in African ancestry women reflects underrepresentation in the genome-wide association studies used to develop current scores. As cohort sizes grow and models are refined, PGSs may become important tools for equitable risk stratification in maternal health.

Adult

Enhanced Prediction of Peripheral Artery Disease Using Plasma Proteomics Among Individuals Without Diabetes.

BACKGROUND: Although peripheral artery disease (PAD) is an important diabetes complication, a substantial proportion of cases occur among individuals without diabetes. This study aimed to assess the predictive value of plasma proteomics in the long-term risk of PAD among individuals initially free of diabetes. METHODS: Included were 46&#x2009;508 participants (6046 with prediabetes) without diabetes or major cardiovascular disease at recruitment of the UK Biobank. Using multivariable Cox regression models, a total of 2923 unique plasma proteins were assessed for the associations with incident PAD. Significant proteins were subsequently processed by a trained light gradient boosting machine classifier to determine important proteins. Using receiver operating characteristic analyses, the performance of these important proteins in predicting incident PAD were evaluated, in the whole sample and by glycemic status (normoglycemia and prediabetes). RESULTS: During a median follow-up of 12.7&#x2009;years, 461 participants developed PAD. There were 107 proteins associated with incident PAD, with 103 positive associations. The LGBM approach identified 9 proteins (eg, WFDC2 [WAP 4-disulfide core domain protein 2], MMP12 [macrophage metalloelastase], and GDF15 [growth differentiation factor 15]) as the top-ranked proteins based on their importance ordering. Whereas glycated hemoglobin showed very modest predictive accuracy, a panel incorporating these top proteins showed good performance in the prediction of PAD risk (area under the curve 0.820), and it significantly enhanced the prediction beyond traditional risk factors (raising area under the curve from 0.803 to 0.837, DeLong test P=5.21&#xd7;10-3). These observations were consistent for participants with normoglycemia or prediabetes. CONCLUSIONS: Plasma protein biomarkers enhance the prediction of long-term risk for PAD among individuals without diabetes, regardless of glycemic status.

Humans

Beyond Morphology: Reframing Lymph-Node Metastasis Prediction Through Clonal Ecology-Decades-Long Genomic Instability and Polyclonal-to-Monoclonal Transitions as the Missing Dimension in Cancer.

Recent whole-genome, lineage-tracing, single-cell, and spatial studies have reshaped our understanding of tumor evolution, revealing that cancers can arise from polyclonal populations, undergo decades-long genomic instability before clinical detection, and progress through dynamic changes in subclonal composition, cellular state, and ecological organization. These findings challenge the assumption underlying morphology-based prediction models that metastatic risk can be inferred from static histological features alone. Here, we revisit lymph-node metastasis prediction in colorectal cancer through clonal ecology, integrating computational pathology with evolutionary oncology. Drawing on the subclonal switchboard model proposed in 2012 and subsequent artificial intelligence (AI)-enabled approaches for tracking dominant and dormant subclones, we synthesize evidence that metastatic potential reflects clonal ancestry, evolutionary timing, spatial niche architecture, cellular plasticity, intercellular interactions, dormancy, and treatment-driven shifts in subclonal fitness. We define five complementary methodological pillars for operationalizing clonal ecology: single-cell transcriptomics for resolving rare subclones, evolutionary trajectories, and adaptive cell states; lineage tracing and phylogenetics for reconstructing clonal ancestry and divergence; spatial transcriptomics and genomics for mapping subclonal geography and tumor-stromal-immune interactions; longitudinal liquid biopsy surveillance for monitoring residual disease, clonal turnover, and emerging resistance; and AI-enabled multimodal integration for connecting histopathology, genomics, spatial biology, and longitudinal data into predictive ecological-state models. Multiple-instance learning and pathology foundation models provide scalable computational foundations for evolution-aware prediction. Translationally, dormant subclones represent actionable reservoirs of recurrence. A longitudinal clinical and experimental study of KMT2A-rearranged acute myeloid leukemia further supports central predictions of the subclonal switchboard framework by demonstrating treatment-associated shifts in subclonal dominance, persistence of cryptic adaptive programs, and ecological rewiring during resistance and relapse. We propose clonal ecology as a measurable dimension for extending morphology-driven prediction toward integrative models that anticipate evolutionary transitions, identify therapeutic windows, and proactively constrain adaptive tumor ecosystems before resistant or metastatic subclones achieve clinical dominance.

Humans

DPAS-Graph: adaptive spatial-feature relation learning for spatial RNA-to-protein prediction and virtual protein profiling.

Paired spatial multi-omics provides a supervised basis for learning RNA-protein correspondence in situ, but predicting protein abundance from spatial transcriptomic data alone remains challenging across tissue contexts and protein panels. Here, we present DPAS-Graph, an adaptive relation-learning framework for spatial RNA-to-protein prediction. Rather than directly merging spatial proximity and transcriptomic similarity as fixed graph priors, DPAS-Graph represents them as two relation channels on a shared edge support and updates their contributions during representation learning for protein prediction. Its Niche-Coupled Field Encoder combines layer-wise edge-relation modeling, intra-branch relation refinement, and cross-branch residual correction to learn spot representations for protein abundance prediction. In a leave-one-dataset-out benchmark across seven paired spatial multi-omics datasets, DPAS-Graph achieved lower aggregate prediction errors and improved spot-level agreement of protein expression profiles, with gains mainly reflected in error-based metrics and PCC-Spot. Spatial autocorrelation and protein-derived domain agreement analyses were further used to characterize the spatial behavior of the predicted protein maps. When applied to external RNA-only spatial sections, DPAS-Graph generated qualitatively interpretable marker-level virtual protein maps, illustrating its use as a complementary tool for protein-level interpretation of transcriptomics-only spatial data.

RNA

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core C&#x3b1; RMSDs of 0.62&#xa0;&#xc5; for MHC-I and 0.33&#xa0;&#xc5; for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63&#xa0;817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides

Proteomics-enabled learning machine algorithms enhance the prediction of cardiovascular diseases in patients with type 2 diabetes mellitus.

BACKGROUND AND AIMS: Estimating the risk of cardiovascular disease (CVD) complications in type 2 diabetes mellitus (T2DM) patients is critical in the medical decision-making process. This study aimed to use a machine learning technique combined with proteomics to develop personalized models for predicting CVD in patients with T2DM. METHODS AND RESULTS: In total, 874 patients with T2DM and 2,920 Olink proteins obtained from the UK Biobank were used in this study. Proteins were screened using Cox regression and LASSO regression. A basic model containing clinical features and a full model combining proteome and clinical features were constructed using the random survival forest algorithm. The area under the receiver operating characteristic (ROC) curve (AUC) was used to evaluate the predictive performance of the models and compare them with other CVD predictive models. Compared with the basic model, the full model performed better in predicting CVD, with time-dependent AUCs of 0.81 (3&#x2009;years), 0.74 (5&#x2009;years) and 0.74 (10&#x2009;years) (0.77, 0.69 and 0.67). We calculated the risk scores of the Framingham, ASCVD and Score2-Diabetes models. The results revealed that the prediction performance of the full model was also better than that of the abovementioned models. In terms of differentiation accuracy, the results of the net reclassification improvement index and integrated discrimination improvement index showed that the full model can identify high-risk individuals more accurately (accuracy rate: 79% vs. 69%). CONCLUSIONS: Proteomics can be used to predict cardiovascular complications in diabetic patients. It is also necessary to consider the applicability of the model due to the limitations of the sample size and the constraints of proteomics in clinical applications.

Humans

Predictive evolutionary genomics: principles, validation, and practice.

Climate change and habitat loss are driving rapid evolutionary responses in populations world-wide, which creates an urgent need for evolutionary forecasting in conservation and agriculture. Such forecasting can be categorized into three time scales: trait-based models that use multivariate quantitative genetic equations to project correlated phenotypic responses up to c.&#xa0;20 generations, allele-based analyses that model allele frequency dynamics up to 100 generations, and composite adaptation scores that aggregate many small effects to yield predictions across longer horizons. However, these approaches have remained largely disconnected. Here, we present a Bayesian framework that integrates these three complementary approaches for evolutionary prediction. Our framework combines genomic, phenotypic, and environmental data to yield probabilistic predictions with explicit uncertainty. We show how predictive evolutionary forecasts can be validated with experimental evolution, field experimentation, historical specimens, and reciprocal transplants. These validated forecasts can help advance conservation and agricultural programmes by helping predict which populations are at risk of future extinction, optimizing breeding programmes for future climates, and planning ecosystem management under environmental change. By supporting a shift towards more predictive approaches in evolutionary biology, this framework may help improve our ability to manage biodiversity and food security in a changing world.

Genomics

Utilizing evolutionary conservation to detect deleterious mutations and improve genomic prediction in cassava.

INTRODUCTION: Cassava (Manihot esculenta) is an annual root crop which provides the major source of calories for over half a billion people around the world. Since its domestication ~10,000 years ago, cassava has been largely clonally propagated through stem cuttings. Minimal sexual recombination has led to an accumulation of deleterious mutations made evident by heavy inbreeding depression. METHODS: To locate and characterize these deleterious mutations, and to measure selection pressure across the cassava genome, we aligned 52 related Euphorbiaceae and other related species representing millions of years of evolution. With single base-pair resolution of genetic conservation, we used protein structure models, amino acid impact, and evolutionary conservation across the Euphorbiaceae to estimate evolutionary constraint. With known deleterious mutations, we aimed to improve genomic evaluations of plant performance through genomic prediction. We first tested this hypothesis through simulation utilizing multi-kernel GBLUP to predict simulated phenotypes across separate populations of cassava. RESULTS: Simulations showed a sizable increase of prediction accuracy when incorporating functional variants in the model when the trait was determined by<100 quantitative trait loci (QTL). Utilizing deleterious mutations and functional weights informed through evolutionary conservation, we saw improvements in genomic prediction accuracy that were dependent on trait and prediction. CONCLUSION: We showed the potential for using evolutionary information to track functional variation across the genome, in order to improve whole genome trait prediction. We anticipate that continued work to improve genotype accuracy and deleterious mutation assessment will lead to improved genomic assessments of cassava clones.

cassava (Manihot esculenta)

Construction of precision clinical-proteomics risk model based on machine learning for predicting heart failure in type II diabetes mellitus.

BACKGROUND AND AIMS: Heart failure (HF) is a severe complication in type 2 diabetes mellitus (T2DM), but current risk stratification scores have limited predictive accuracy. We aimed to develop novel prediction tools integrating clinical variables with proteomics to improve risk stratification of hospitalization for HF in T2DM. METHODS AND RESULTS: In this study, we included 2111 UK Biobank participants with T2DM but no prior HF, and profiled 2920 proteins to predict 10-year incident HF hospitalization. Participants were randomly divided into training (70%), tuning (10%), and validation (20%) sets.Three prediction models were developed: a Clinical model based on demographic characteristics, comorbidities, medication use, and laboratory indices; a Protein model based on 40 proteins selected by the Light Gradient Boosting Machine (LGBM); and the Clinical OMics and Protein ASSessment for Heart Failure (COMPASS-HF) model, which integrated both clinical variables and the LGBM-selected proteins. Models were evaluated for area under the curve (AUC), sensitivity, and specificity. During follow-up, 168 participants (7.96%) developed incident HF. The COMPASS-HF model showed better discrimination than the Clinical model, with an AUC of 0.897 (95% CI: 0.850-0.945) versus 0.790 (95% CI: 0.723-0.856). It also demonstrated higher sensitivity (0.882; 95% CI: 0.725-0.967) and consistent performance in subgroups. COMPASS-HF effectively stratified risk of hospitalization for HF, with cumulative incidence rates of 31.9% in the high-risk group and 1.2% in the low-risk group. CONCLUSIONS: By combining clinical and proteomic variables, we developed a high-performance HF prediction model for T2DM, enabling precise risk stratification and informing early intervention strategies.

Humans

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans

The magnitude of exercise-induced ST segment depression and the predictive value of exercise testing.

The assess whether the magnitude of exercise induced ST segment depression improves the predictive values of symptom limited exercise tests, and helps in the recognition of patients with more severe coronary heart disease, 90 consecutive patients with positive treadmill tests who also underwent selective coronary arteriography were reviewed. The predictive value improved progressively with the increasing ST depression and was most reliable in a select group of patients with normal electrocardiographic baseline who were not receiving digitalis (73% with ST depression greater than or equal to 1 mm to 100% with ST depression greater than or equal to 4 mm). The incidence of 2 and 3 vessel disease increased from 61% with ST depression greater than or equal to 1 mm in the overall population to 100% with ST depression greater than or equal to 4 mm in the select group, and the incidence of left main trunk lesions increased, respectively from 6 to 30%. The prediction of 2 and 3 vessels disease was found to be significantly greater when patients were dichotomized into those with ST depression greater than or equal to 4 mm compared to less than 4 mm. It is concluded that the magnitude of ST segment depression definitely improves the predictive values of exercise tests as well as the ability to recognize the patients with more severe disease. However, the markedly positive exercise tests cannot be utilized to accurately predict the presence of 2 or 3 vessel disease in individual cases unless ST depression attains 4 mm or more in patients with normal electrocardiographic baseline who are not taking digitalis. In this group, the ability to predict left main trunk lesion is approximately 30%.

Coronary Angiography

Prediction of Australian wheat genotype by environment interactions and mega-environments.

Latent environmental effects of genotype by environment interactions could be predicted from observed environmental covariates. Predictions into the wider target population of environments revealed greater insights. Wheat is grown across a diverse range of environments in Australia with contrasting environmental constraints. Targeted breeding to optimise genotypes in target environments is hindered by large and ubiquitous genotype by environment interactions (GEI). Common GEI in multi-environment trial experiments, which sample the target population of environments, can be efficiently modelled using latent environmental effects from factor analytic mixed models. However, generalised prediction into the full target population of environments is difficult without a clear link to observed environmental covariates (ECs) that are defined from high-resolution weather and soil data. Here, we used a large wheat multi-environment trial dataset and demonstrated that latent environmental effects can be associated with and predicted from observed ECs. We found GEI-based environment classes could be defined by combinations of key ECs. Prediction of main and latent effects in a wider set of environments covering the full TPE across the Australian grain belt over 13&#xa0;years revealed the complex trends of environmental effects and GEI over regional scales demonstrating high year-to-year variability. Regional environment types often shifted year-to-year. Cross-validation of forward genomic prediction into untested year environments demonstrated that increased accuracy is possible if estimated genetic effects are also accurate and ECs of new environments are known. These findings may guide Australian wheat breeders to better target specifically adapted material to mega-environments defined by static GEI while also considering broad adaptability and non-static GEI resulting from year-to-year variability.

Triticum

Genomic prediction of agronomic traits in perennial ryegrass (Lolium perenne L.) and genotype x environment interactions at the limit of the species distribution.

KEY MESSAGE: Perennial ryegrass shows extensive genotype x environment interactions at the limit of its ecological niche. Accounting for GxE may improve prediction even when environmental and genetic samples are highly diverse. BACKGROUND: In breeding the aim is to identify and accumulate beneficial variants. However, detection of these variants may be challenging in the presence of extensive genotype x environment interactions (GxE). METHODS: The study assesses the performance of 264 diploid perennial ryegrass accessions in a multi-environment field trial. We investigate the extent of GxE, for yield (total dry matter) and persistence traits under environmental conditions experienced in Nordic and Baltic regions at the limit of the species distribution. Two different approaches to modelling GxE were tested and validated under three different breeding scenarios. RESULTS: Our analysis documented the presence of significant GxE for all traits. Validation showed improvements in prediction accuracy when accounting for GxE: up to 4% for yield when predicting in unobserved environments, and up to 22% and 9% for spring cover and winter kill, respectively, when predicting unobserved germplasm. Genome-wide-association-studies (GWAS) were utilized to detect genetic variants with marginal effects (environment-independent effect) and conditional effects (environment-dependent effects). Results showed the presence of large-effect genetic variants with marginal effects, in addition to few Quantitative Trait Loci (QTL) whose effects were adaptive under specific environmental conditions while neutral or deleterious under different environmental conditions. CONCLUSION: This study demonstrates the usefulness and limitations of genomic prediction models for predicting GxE in highly diverse samples and describes the extent of GxE at the limit of species distribution for perennial ryegrass. Our study points towards adaptive variation which may enhance persistence of perennial ryegrass populations in Nordic and Baltic growing conditions.

Lolium

Prediction of infarct size from serial CK determinations: evaluation by clinical studies and computer simulation.

To assess reduction of infarct size by therapeutic intervention, a high predictive accuracy is mandatory. The CK release in the circulation (CKr) was studied in 12 consecutive patients after uncomplicated myocardial infarction, admitted within 5 h after onset of symptoms. Despite improvement of existing methods, such as a more frequent sampling, CK-MB determination instead of total CK determination and use of a gamma-exponential instead of a log-normal curve-fitting technique, the correlation between CKr predicted from measurements within 7 h after the start of CK rise and CKr calculated after completion of the CK curve remained poor. Computer simulations were done to investigate measurement errors as a cause of this failure. Normally distributed noise, with standard deviations ranging from 0.2% to 8.0% of peak CK-MB, was added to the first points of an ideal gamma-exponential CK-MB curve and predictions were made from these "noisy" points. A small noise already produced a great variation in prediction: 0.8% noise resulted in a deviation of predicted CKr from calculated CKr ranging from --20 to +6%. It is concluded that adequate prediction of infarct size from serial CK determinations in the first 7 h after onset of the CK rise must fail if the precision of the biochemical determination is not less than 0.4%.

Clinical Enzyme Tests

Deep Learning on Histologic Slides Accurately Predicts Consensus Molecular Subtypes and Spatial Heterogeneity in Colon Cancer.

Colon cancer (CC) is the third most prevalent cancer type. It is highly heterogeneous, particularly in terms of molecular profiles, which have both prognostic and predictive impacts on the treatment efficacy. However, CC treatment in adjuvant situations is currently guided solely by T and N staging. In this context, consensus molecular subtypes (CMSs) were introduced to stratify patients with CC based on molecular profiles. Recent studies have shown that CMS can be heterogeneous in CC, leading to a worse prognosis. This study focused on predicting CMS and its heterogeneity in CC using deep learning on digitized hematoxylin and eosin &#xb1; saffron-stained whole-slide images. Data and whole-slide images of 1996 patients from the PETACC-8, The Cancer Genome Atlas-COAD, and PRODIGE-13 cohorts were used. The model is trained to predict a 4-dimensional CMS vector, reflecting intratumor heterogeneity (ITH). It comprises a self-supervised model for embedding image patches into vectors and a weakly supervised model predicting CMS calls. Ground-truth CMS scores are obtained with the CMSclassifier package. Interpretability analyses are performed at the slide and patch levels. For homogeneous tumors, the model trained on PETACC-8 achieves 93.0% (&#xb1;1.4%) macroaverage area under the curve in internal cross-validation and 94.4% macroaverage area under the curve in external validation over PRODIGE-13, whereas the The Cancer Genome Atlas-COAD model reaches 85.4% (&#xb1;3.0%) in cross-validation and 92.4% over PRODIGE-13. The trained models also provide spatial distributions of CMS across tumor slides and associate specific histologic features with each CMS. Finally, the models are able to predict ITH. The results show that a deep learning model trained on routine histology slides is capable of providing an efficient and robust method for predicting CMS and characterizing a patient's ITH, paving the way for the routine consideration of CMS/ITH in clinical decision making in the adjuvant setting.

Humans