Search PubMedSearch

Biomedical subjects

Pavel P Kuksa

Publications and source records attributed to Pavel P Kuksa.

3 recordsLinked to original sources

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

Multi-ancestry genome-wide meta-analysis of 56,241 individuals identifies known and novel cross-population and ancestry-specific associations as novel risk loci for Alzheimer's disease.

BACKGROUND: Limited ancestral diversity has impaired our ability to detect risk variants more prevalent in ancestry groups of predominantly non-European ancestral background in genome-wide association studies (GWAS). We construct and analyze a multi-ancestry GWAS dataset in the Alzheimer's Disease Genetics Consortium (ADGC) to test for novel shared and population-specific late-onset Alzheimer's disease (LOAD) susceptibility loci and evaluate underlying genetic architecture in 37,382 non-Hispanic White (NHW), 6728 African American, 8899 Hispanic (HIS), and 3232 East Asian individuals, performing within ancestry fixed-effects meta-analysis followed by a cross-ancestry random-effects meta-analysis. RESULTS: We identify 13 loci with cross-population associations including known loci at/near CR1, BIN1, TREM2, CD2AP, PTK2B, CLU, SHARPIN, MS4A6A, PICALM, ABCA7, APOE, and two novel loci not previously reported at 11p12 (LRRC4C) and 12q24.13 (LHX5-AS1). We additionally identify three population-specific loci with genome-wide significance at/near PTPRK and GRB14 in HIS and KIAA0825 in NHW. Pathway analysis implicates multiple amyloid regulation pathways and the classical complement pathway. Genes at/near our novel loci have known roles in neuronal development (LRRC4C, LHX5-AS1, and PTPRK) and insulin receptor activity regulation (GRB14). CONCLUSIONS: Using cross-population GWAS meta-analyses, we identify novel LOAD susceptibility loci in/near LRRC4C and LHX5-AS1, both with known roles in neuronal development, as well as several novel population-unique loci. Reflecting the power of diverse ancestry in GWAS, we detect the SHARPIN locus with only 13.7% of the sample size of the NHW GWAS study (n = 409,589) in which this locus was first observed. Continued expansion into larger multi-ancestry studies will provide even more power for further elucidating the genomics of late-onset Alzheimer's disease.

Humans

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study