Search PubMedSearch

PubMed · 41002190

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

Abstract

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pavel P Kuksa, Matei Ionita, Luke Carter, Jeffrey Cifello, Prabhakaran Gangadharan, Kaylyn Clark, Otto Valladares, Yuk Yee Leung, Li-San Wang. 2025-10-02. BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.. https://doi.org/10.1093/bioinformatics%2Fbtaf509

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Pinpointing genomic regions conferring herbicide tolerance in cassava via genome-wide association mapping.

Cassava (Manihot esculenta Crantz) is a tropical crop of major socioeconomic importance, whose productivity can be limited by sensitivity to herbicides used for weed management. This study aimed to perform a genome-wide association study (GWAS) in 194 cassava genotypes to identify genomic regions associated with tolerance to the herbicides mesotrione, S-metolachlor, and chloransulam-methyl. The evaluations performed at 3, 6, 9, 15, and 30 days after application (DAA) were used to characterize the temporal progression of phytotoxicity. Based on this analysis, the phenotype obtained at 9 days after application (PhytoX9DAA) was selected for genome-wide association analyses because it represented the period of greatest symptom expression and the highest discrimination among genotypes. GWAS analyses were performed using de-regressed BLUPs and the MLM, MLMM, and BLINK models, incorporating kinship (K) and population structure (Q) matrices. Significant markers were detected across multiple chromosomes, and the corresponding genomic windows contained candidate genes with functional annotations related to herbicide response. The predominant functional categories included membrane transport, channel activity, signal peptide processing, protein phosphorylation, cellular signaling, and metabolic regulation. Key candidate genes included Manes.02G151900 and Manes.02G152700 (chromosome 2), associated with transmembrane transport and signal peptide processing; Manes.09G060900 (chromosome 9), associated with protein kinase activity, ATP binding, and protein phosphorylation; and Manes.15G083800 and Manes.15G084000 (chromosome 15), associated with S-adenosylmethionine-dependent methyltransferase activity, membrane-related functions, and protein phosphorylation. These genes participate in biochemical pathways involved in cellular signaling, membrane transport, and metabolic regulation that may contribute to herbicide tolerance. Overall, the results demonstrate that herbicide tolerance in cassava is a quantitative and polygenic trait governed by numerous small-effect loci. The integration of cellular signaling, metabolic regulation, and membrane transport supports the physiological resilience of the species under chemical exposure, providing valuable insights for breeding strategies and marker-assisted selection.

Genome-Wide Association Study

Dissecting genetic architecture of growth and yield traits in horsegram using GWAS.

Horsegram (Macrotyloma uniflorum), a member of the Fabaceae family, is a nutritious and low-cost legume used for both grain and fodder. This study employed a genome-wide association approach to identify loci linked to key agronomic traits in horsegram. Plant height, seed size, and shoot fresh weight were evaluated in a panel of 96 diverse genotypes. GBS was performed using the Illumina HiSeq platform, yielding 20,241 high-quality SNPs after filtering at a 5% minor allele frequency. Population structure analysis classified genotypes into three admixed subgroups. Phenotyping was conducted over three consecutive years at two locations in Himachal Pradesh (Palampur and Bajaura) using a randomized block design with two replications. GWAS analyses using GLM, MLM, FarmCPU, and BLINK models identified eight markers for plant height, three for seed size, and five for shoot fresh weight across different chromosomes. These markers provide valuable tools for accelerating trait improvement in future horsegram breeding programs.

Genome-Wide Association Study

A module-based approach for post-omics, post-GWAS network-based gene classification.

MOTIVATION: Complex traits and diseases are highly polygenic and understanding the full set of genes involved is a central challenge in biomedicine. However, due to sample size limitations and noise (technical and biological), experimental approaches for disease-gene discovery such as transcriptomics and GWAS result in long, noisy, heterogeneous gene lists, which may be trimmed to a subset of likely relevant genes while leaving several false negatives. Computational gene classification approaches, especially those using genome-scale molecular interaction networks, are promising avenues for complementing such experimental findings by analytically expanding observed gene lists based on the functional relatedness between genes. We previously introduced the network-based gene classification approach, GenePlexus, which was rigorously benchmarked to show state-of-the-art performance, especially for predicting novel genes associated with biological processes and fine-grained phenotypes. Network-based gene classification performance,however, declines for diseases, especially when the inputs are omics and GWAS-based long gene lists. RESULTS: Here, we show that these disease gene lists span multiple biological processes spread across the molecular network, and we propose ModGenePlexus, a new network-based gene classification method that takes a two-stage approach. First, clustering and semi-supervised learning decomposes the input gene list into coherent, denoised network gene modules. Then, ModGenePlexus trains supervised (GenePlexus) classifiers for each module and aggregates predictions to return genome-wide rankings. We benchmarked ModGenePlexus across simulated data, transcriptomic signatures, and GWAS datasets (together spanning hundreds of diseases), showing improved recovery of known disease genes compared to GenePlexus. Beyond improved classification, the results of enrichment analysis of ModGenePlexus outputs are much more interpretable by virtue of revealing nuanced biological processes. Together, these results establish ModGenePlexus as a scalable, interpretable tool for gene classification of GWAS- and omics-derived gene lists across diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: ModGenePlexus is freely available on GitHub at https://github.com/krishnanlab/ModGenePlexus, and the full source code and results supporting this study are available on Zenodo at https://zenodo.org/records/19857910.

Genome-Wide Association Study