Search PubMedSearch

Biomedical subjects

Yuk Yee Leung

Publications and source records attributed to Yuk Yee Leung.

7 recordsLinked to original sources

Dissecting the relationship between haplotypes around ATXN2 CAG repeats and the number of CAA interruptions by long-read sequencing.

BACKGROUND: CAG repeat expansions in ATXN2 are implicated as risk factors for several neurological diseases, including spinocerebellar ataxia type 2 (SCA2) when >=33 CAG repeats are present, and amyotrophic lateral sclerosis (ALS) when 27-33 CAG repeats are present. However, how haplotypes around the repeats and CAA interruptions within the repeats are associated with disease phenotypes remains poorly understood. Previous studies on haplotypes around ATXN2 were limited to SNPs very close to the repeats (<5kb) or were based on statistical inference only. METHODS: Here, we used long-read sequencing on the Oxford Nanopore Technologies (ONT) platform to simultaneously infer haplotypes around ATXN2, the number of CAG repeats, and the number of CAA interruptions, along with NYGC ALS Consortium NGS dataset. We further sequenced 41 individuals (EUR = 39) with neurological diseases with intermediate repeats by ONT. RESULTS: We found that haplotypes around ATXN2 and the number of interruptions show ethnicity-specific and ALS-specific distribution. Three CAA interruptions are present at low prevalence (~1%) in control populations in multiple ancestry groups, but high prevalence (~55%) in ALS individuals with intermediate repeats. Furthermore, we examined 159 individuals with ALS (~90% European ancestry) with intermediate ATXN2 repeats and found a unique haplotype in ALS individuals with three CAA interruptions, which can be tagged by an SNV, rs148019457. We also validated that the rs148019457-G allele is only present in haplotypes with three CAA interruptions. CONCLUSIONS: In summary, our study shows that 3 CAA interruptions are rarely seen in healthy controls but are common in those with expanded ATXN2 CAG repeats who have neurological disorders, and that rs148019457 tags a specific haplotype with 3 CAA interruptions within expanded ATXN2 CAG repeats in individuals of European ancestry. These results have implications for the development of precision genomic medicine for neurological disorders, and the tag SNP may help identify those with interruptions from existing population genotyping data.

ATXN2

Comprehensive adjudication identifies 111 high-confidence loci for Alzheimer's disease and related dementias.

BACKGROUND: The Alzheimer's Disease Sequencing Project Gene Verification Committee developed a systematic framework to adjudicate genetic evidence for AD and related dementias, addressing wide variation in association quality. METHODS: Phase 1 established tiered criteria by evaluating 23 nominated loci across study designs. Phase 2 applied this framework to 29 large-scale genome-wide studies published since 2015, tiering 163 unique loci. RESULTS: Phase 1 yielded 17 high-confidence loci (12 linked to specific genes), and Phase 2 identified 111 high-confidence loci/genes with replicated associations across ancestries and convergent single-variant/variant-set evidence. Prioritized loci highlight APP processing, microglial immunity, and lipid metabolism pathways, including genes not captured by existing resources like Agora or Open Targets. Summarized results can be viewed at https://topgenes.niagads.org/. CONCLUSION: This rigorously adjudicated catalog represents the most comprehensive AD/ADRD genetics resource to date, providing a foundation for functional validation and therapeutic discovery with broad applicability to complex diseases.

Journal Article

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100&#xd7; more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study

Multi-ancestry genome-wide meta-analysis of 56,241 individuals identifies known and novel cross-population and ancestry-specific associations as novel risk loci for Alzheimer's disease.

BACKGROUND: Limited ancestral diversity has impaired our ability to detect risk variants more prevalent in ancestry groups of predominantly non-European ancestral background in genome-wide association studies (GWAS). We construct and analyze a multi-ancestry GWAS dataset in the Alzheimer's Disease Genetics Consortium (ADGC) to test for novel shared and population-specific late-onset Alzheimer's disease (LOAD) susceptibility loci and evaluate underlying genetic architecture in 37,382 non-Hispanic White (NHW), 6728 African American, 8899 Hispanic (HIS), and 3232 East Asian individuals, performing within ancestry fixed-effects meta-analysis followed by a cross-ancestry random-effects meta-analysis. RESULTS: We identify 13 loci with cross-population associations including known loci at/near CR1, BIN1, TREM2, CD2AP, PTK2B, CLU, SHARPIN, MS4A6A, PICALM, ABCA7, APOE, and two novel loci not previously reported at 11p12 (LRRC4C) and 12q24.13 (LHX5-AS1). We additionally identify three population-specific loci with genome-wide significance at/near PTPRK and GRB14 in HIS and KIAA0825 in NHW. Pathway analysis implicates multiple amyloid regulation pathways and the classical complement pathway. Genes at/near our novel loci have known roles in neuronal development (LRRC4C, LHX5-AS1, and PTPRK) and insulin receptor activity regulation (GRB14). CONCLUSIONS: Using cross-population GWAS meta-analyses, we identify novel LOAD susceptibility loci in/near LRRC4C and LHX5-AS1, both with known roles in neuronal development, as well as several novel population-unique loci. Reflecting the power of diverse ancestry in GWAS, we detect the SHARPIN locus with only 13.7% of the sample size of the NHW GWAS study (n&#x2009;= 409,589) in which this locus was first observed. Continued expansion into larger multi-ancestry studies will provide even more power for further elucidating the genomics of late-onset Alzheimer's disease.

Humans

Structural variation detection and association analysis of whole-genome-sequence data from 16,543 Alzheimer's disease sequencing project subjects.

INTRODUCTION: The role of structural variations (SVs) in Alzheimer's disease (AD) remains understudied. METHODS: We analyzed whole-genome sequencing data from the Alzheimer's Disease Sequencing Project (N&#xa0;=&#xa0;16,543) and identified 400,234 (168,223 high-quality) SVs. Laboratory validation yielded a sensitivity of 82% (85% for high-quality). RESULTS: We found a burden of singletons (odds ratio [OR]&#xa0;=&#xa0;1.07, p&#xa0;=&#xa0;0.0017) and homozygous deletions (OR&#xa0;=&#xa0;1.14, p&#xa0;<&#xa0;0.0001) in cases. On AD genes, we observed the ultra-rare SVs associated with the disease, including protein-altering SVs in ABCA7, APP, PLCG2, and SORL1. Twenty-one SVs are in linkage disequilibrium (LD) with known AD-risk variants, exemplified by a 5k deletion in LD (R2&#xa0;=&#xa0;0.99) with rs143080277 in NCK2. We identified a rare deletion near RNA5SP293 associated with AD (OR&#xa0;=&#xa0;1.99, p&#xa0;=&#xa0;1.3&#xa0;&#xd7;&#xa0;10-5), which was replicated using an independent dataset. DISCUSSION: This study highlights the pivotal role of SVs in AD genetics. HIGHLIGHTS: Observed a significant burden of singletons and homozygous deletions in Alzheimer's disease (AD) patients. Identified rare protein-altering structural variations (SVs) in ABCA7, APP, PLCG2, and SORL1. Established linkages between SVs and AD risk-associated single nucleotide variants (SNVs). Discovered a novel deletion near RNA5SP293 linked to AD, replicated independently. Uncovered over-representation of SVs in neuronal function pathways.

Humans

Association of common and rare variants with Alzheimer's disease in more than 13,000 diverse individuals with whole-genome sequencing from the Alzheimer's Disease Sequencing Project.

INTRODUCTION: Alzheimer's disease (AD) is a common disorder of the elderly that is both highly heritable and genetically heterogeneous. METHODS: We investigated the association of AD with both common variants and aggregates of rare coding and non-coding variants in 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. RESULTS: Pooled-population analyses of all individuals identified genetic variants at apolipoprotein E (APOE) and BIN1 associated with AD (p&#xa0;<&#xa0;5&#xa0;&#xd7;&#xa0;10-8). Subgroup-specific analyses identified a haplotype on chromosome 14 including PSEN1 associated with AD in Hispanics, further supported by aggregate testing of rare coding and non-coding variants in the region. Common variants in LINC00320 were observed associated with AD in Black individuals (p&#xa0;=&#xa0;1.9&#xa0;&#xd7;&#xa0;10-9). Finally, we observed rare non-coding variants in the promoter of TOMM40 distinct of APOE in pooled-population analyses (p&#xa0;=&#xa0;7.2&#xa0;&#xd7;&#xa0;10-8). DISCUSSION: We observed that complementary pooled-population and subgroup-specific analyses offered unique insights into the genetic architecture of AD. HIGHLIGHTS: We determine the association of genetic variants with Alzheimer's disease (AD) using 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. We identified genetic variants at apolipoprotein E (APOE), BIN1, PSEN1, and LINC00320 associated with AD. We observed rare non-coding variants in the promoter of TOMM40 distinct of APOE.

Humans

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study