Search PubMedSearch

PubMed · 41702378

ReverseGWAS identifies combined phenotypes associated with a genotype in GWA studies.

Abstract

MOTIVATION: Traditional genome-wide association studies (GWAS) aim to uncover the genetic variants associated with a single phenotype of interest (typically a disease), and to elucidate its genotypic architecture. However, many of today's GWAS simultaneously measure multiple related phenotypes, leading to the possibility of pursuing the reverse aim of elucidating the "phenotypic architecture" of a single genetic variant. In other words, we may ask what combination of measured phenotypes is associated with a given genotypic variant. ReverseGWAS is an algorithmic platform for answering such questions in the context of large-scale multi-phenotype GWAS. RESULTS: We demonstrate the effectiveness of ReverseGWAS on simulated data, showing its ability to identify logical combinations of phenotypes with a reasonable amount of noise. We then apply it to a selection of combined phenotypes from the UK Biobank, obtaining 719 candidate associations using autoimmune diseases and 205 using common ICD10 codes. We find that the majority of these associations (546/719 and 111/205, respectively) successfully replicate in an independent cohort, FinnGen. AVAILABILITY AND IMPLEMENTATION: The source code of ReverseGWAS is freely available to non-commercial users as an installable R package at https://github.com/Leonardini/rgwas.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Leonid Chindelevitch, Åsa K Hedman, Dmitri Bichko, Daniel Ziemek. 2026-02-28. ReverseGWAS identifies combined phenotypes associated with a genotype in GWA studies.. https://doi.org/10.1093/bioinformatics%2Fbtag079

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A sequence-based classifier distinguishes phenotype-associated genes from other gene models in plants.

Only a small fraction of annotated plant genes possess experimentally validated associations with specific phenotypes. Phenotype-associated genes have distinct structural, molecular, and evolutionary characteristics compared with nonvalidated gene models. Here, we develop a simple classifier that uses sequence and evolutionary features, which can be generated for any species with an annotated reference genome assembly, to accurately distinguish phenotype-associated genes from both the overall population of annotated gene models and a specific set of genes identified as being tolerant of premature stop mutations. A model trained solely on genes from maize (Zea mays) identifies and prioritizes rice (Oryza sativa) and Arabidopsis (Arabidopsis thaliana) genes that are highly enriched in genes with experimentally validated links to phenotypes in both of these evolutionarily distant species. Gene models predicted to have a higher probability of being linked to phenotypes display patterns consistent with known biological properties of phenotype-associated genes. Notably, the sets of genes predicted to have a high probability of being linked to phenotype variation do not consist exclusively of well-characterized gene families but included many uncharacterized gene families carrying domains of unknown function. The quantitative scores generated by this model offer a valuable resource for prioritizing and exploring the vast number of uncharacterized gene models in plants, reducing the risk of failure in future reverse genetic efforts and potentially accelerating gene discovery and functional annotation in crops.

Phenotype

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Integrating plant phenotypic and genotypic data in the AGENT project: a BrAPI service implementation.

MOTIVATION: The AGENT project established a network of actively cooperating European genebanks, integrating genomic and phenotypic data from accessions of wheat and barley. Due to specific storage demands for phenotypic and genotypic data, the project used separate database instances and backend technologies to manage integrated phenotypic and genotypic data. RESULTS: We discuss the challenges encountered when integrating dispersed data to serve through a single interface such as the Plant Breeding Application Programming Interface, BrAPI. We examine how the consistent mappability of genebank data to the BrAPI model can enable the implementation of effective services. The advantages of BrAPI in transparently linking distributed data entities through embedded, unique identifiers are highlighted. We present a technical solution involving a BrAPI proxy, which combines and merges separate BrAPI endpoints. Finally, we demonstrate the AGENT BrAPI implementation with an illustrative example that validates a suggested SNP for a trait from the literature by linking phenotypic, genotypic and passport data. AVAILABILITY AND IMPLEMENTATION: The BrAPI proxy implementation and documentation is available at the Python Package Index (https://pypi.org/project/brapi-proxy) and archived in Zenodo (doi: 10.5281/zenodo.19436445). SUPPLEMENTARY INFORMATION: A Jupyter Notebook file for the validation example using a marker-trait relationship found in the literature.

Phenotype