Search PubMedSearch

Biomedical subjects

Yinan Zheng

Publications and source records attributed to Yinan Zheng.

4 recordsLinked to original sources

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

PathwayVote: an R package for robust pathway enrichment analysis for DNA methylation data using a consensus-based voting framework.

MOTIVATION: Pathway enrichment analysis is commonly used to interpret epigenomewide association studies, yet conventional methods often rely on arbitrary thresholds and simplified CpG-gene mappings, making them sensitive to analytical choices and unable to fully leverage CpG-gene relationships Recent advances in expression quantitative trait methylation (eQTM) studies offer a rich resource to refine these mappings, but are rarely utilized in DNA methylation enrichment pipelines. RESULTS: We developed PathwayVote, an R package that implements a voting-based consensus approach and leverages eQTM data to identify robustly enriched pathways. PathwayVote reduces dependence on arbitrary cutoffs and improves sensitivity and reproducibility of enrichment results. AVAILABILITY AND IMPLEMENTATION: PathwayVote is freely available on GitHub (https://github.com/YinanZheng/PathwayVote) under the GPL-3 license and CRAN: https://CRAN.R-project.org/package=PathwayVote. The version of the code corresponding to this manuscript has been archived on Zenodo (https://doi.org/10.5281/zenodo.17209507).

Humans

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

Alterations in DNA Methylation, Proteomic, and Metabolomic Profiles in African Ancestry Populations with APOL1 Risk Alleles.

KEY POINTS: We aimed to elucidate potential methylation, proteomic, and metabolomic mechanisms by which APOL1 variants may be linked to kidney disease. We report distinct methylation profiling between APOL1 risk allele carriers and noncarriers, many near APOL gene family. We report higher APOL1 protein and lower C18:1 cholesteryl ester in two risk allele carriers. BACKGROUND: The APOL1 high-risk haplotype has been associated with CKD and the deterioration of kidney function, particularly in populations with West African ancestry. However, the mechanisms by which APOL1 risk variants increase the risk for kidney disease and its progression have not been fully elucidated. METHODS: We compared methylation (N=3191; 715 [22%] carriers), proteomic (N=1240; 169 [14%] carriers), and metabolomic (N=6309; 674 [11%] carriers) profiles in African and Hispanic/Latino carriers of two APOL1 high-risk alleles (G1/G1, G2/G2, G1/G2) and noncarriers (G0/G0), excluding heterozygotes (G0/G1, G0/G2), from the Population Architecture using Genomics and Epidemiology Consortium and UK Biobank. In each study, the associations between the APOL1 high-risk haplotype and up to 722,719 cytosine-phosphate-guanine (CpG) sites, 2923 proteins, or 836 metabolites were estimated using covariate-adjusted linear regression models, followed by fixed-effects sample size&#x2013;weighted meta-analyses. RESULTS: Significant associations were observed between APOL1 high-risk haplotype and methylation at 52 CpG sites, with 48 located on chromosome 22 and 18 in the vicinity of APOL1&#x2013;4 and MYH9. All significant CpG sites near APOL2 were hypomethylated, whereas those near APOL3 and APOL4 were hypermethylated. APOL1-associated CpG sites were also identified in genes involved in ion transport and mitochondrial stress pathways. Sensitivity analyses indicated consistent yet attenuated effects among heterozygotes, supporting an additive effect of APOL1 risk alleles. Further analyses of the 52 CpG sites identified two near APOL4 exhibiting G1-specific effects, eight associated with CKD but none with eGFR, and three showing heterogeneity by CKD status. In addition, carrying two APOL1 risk alleles was associated with higher plasma APOL1 protein (&#x3b2;=1.12, PFDR = 2.26e-70) and lower C18:1 cholesteryl ester metabolite (Z=&#x2212;4.50, PFDR = 4.83e-3). CONCLUSIONS: Our results demonstrate differential methylation, proteomic, and metabolomic profiles associated with APOL1 high-risk haplotypes.

APOL1