Search PubMedSearch

PubMed · 42552668

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Abstract

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Explore related subjects

Keep this discovery

BibTeXRIS

Jiahui Wang, Yong Zhang, Bo Li, Xinglin Piao, Xiangyu Zhao, Dongfeng Zhang, Aiwen Wang, Bob Zhang, Kaiyi Wang. 2026-08-04. Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.. https://doi.org/10.1111/nph.71461

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

Predictive evolutionary genomics: principles, validation, and practice.

Climate change and habitat loss are driving rapid evolutionary responses in populations world-wide, which creates an urgent need for evolutionary forecasting in conservation and agriculture. Such forecasting can be categorized into three time scales: trait-based models that use multivariate quantitative genetic equations to project correlated phenotypic responses up to c. 20 generations, allele-based analyses that model allele frequency dynamics up to 100 generations, and composite adaptation scores that aggregate many small effects to yield predictions across longer horizons. However, these approaches have remained largely disconnected. Here, we present a Bayesian framework that integrates these three complementary approaches for evolutionary prediction. Our framework combines genomic, phenotypic, and environmental data to yield probabilistic predictions with explicit uncertainty. We show how predictive evolutionary forecasts can be validated with experimental evolution, field experimentation, historical specimens, and reciprocal transplants. These validated forecasts can help advance conservation and agricultural programmes by helping predict which populations are at risk of future extinction, optimizing breeding programmes for future climates, and planning ecosystem management under environmental change. By supporting a shift towards more predictive approaches in evolutionary biology, this framework may help improve our ability to manage biodiversity and food security in a changing world.

Genomics

Genomic insights into end-use grain quality and nutritional traits of an ancient Indian dwarf wheat ( Triticum sphaerococcum Percival) population using a multi-locus genome-wide association study.

BACKGROUND: Triticum sphaerococcum, an ancient hexaploid wheat species, is renowned for its stress resilience and superior nutritional quality. A panel of 116 T. sphaerococcum accessions (the largest known collection at a single site globally), with six bread wheat released varieties, was evaluated for its potential for genetic quality improvement. Field experiments were conducted under standard, heat and moisture-deficit conditions across two cropping seasons for ten grain end-use quality and nutritional traits. RESULTS: Genotypes showed highly significant differences (P ≤ 0.001) for measured traits, with high broad-sense heritability resulting from substantial genotypic variance contributions. Triticum sphaerococcum consistently outperformed T. aestivum across environments, with moisture-deficit stress proving more detrimental to quality parameters than heat stress, while micronutrient content increased under stressed conditions. Trait correlations revealed that the gluten index (GI) correlated negatively with the grain hardness index (GHI), wet gluten (WG), and water-binding capacity (WB), while positively correlating with dry gluten (DG) and protein content (PRO), whereas grain iron (GFE), zinc (GZN), and protein showed consistent positive interrelationships. Two superior accessions, PAUTS10 (WG 35.13%, DG 13.71%, PRO 16.42%, GZN 50.89 ppm) and Sonamoti (WG 33.33%, DG 12.92%, PRO 16.27%, GZN 56.03 ppm), were identified, surpassing the best check variety HD3226 for quality and nutritional parameters. Multi-locus genome-wide association studies identified 30 stable quantitative trait nucleotides across environments, with candidate gene analysis revealing genes involved in transcription regulation, biosynthetic processes, metal ion homeostasis, and transport. CONCLUSIONS: Triticum sphaerococcum demonstrated superior grain quality and micronutrient potential compared with modern wheat, highlighting its value as a genetic resource for biofortification. The identification of elite accessions and stable quantitative trait nucleotides (QTNs) provides useful targets for breeding programs aimed at improving protein and micronutrient content. Integrating ancient germplasm with modern genomic tools can accelerate the development of nutritionally enhanced wheat varieties. © 2026 Society of Chemical Industry.

Triticum