Search PubMedSearch

Biomedical subjects

Laura M Raffield

Publications and source records attributed to Laura M Raffield.

18 recordsLinked to original sources

Inherited Predisposition to Increased Systemic Inflammation Predicts a Broad Class of Disease Phenotypes.

Chronic, low-grade systemic inflammation is a polygenic trait captured with the INFLA-score, a composite of C-reactive protein, platelet count, leukocyte count, and granulocyte-to-lymphocyte ratio. We derived a polygenic risk score from the INFLA-score (iPRS) in a multi-ancestry population from the UK Biobank (n=421,368), then evaluated and used it in a phenome-wide association study among participants in the All of Us Research Program (AoU). The multi-ancestry iPRS was tested for association with the INFLA-score in AoU (N=4,833 with biomarker data) via linear regression, adjusting for age, sex, and genetically-determined principal components (PCs) and with 2,821 phecodeX-defined phenotypes in AoU (N=265,068) via logistic regression, adjusting for sex, age, EHR length, race, ethnicity and PCs. The iPRS predicted the INFLA-score (R-squared=0.026, beta=0.980, p<2x10-16) and was associated with 47 phenotypes (Bonferroni-corrected p<0.05). The strongest associations were with blood-related phenotypes: elevated white blood cell count (OR=1.19, p=3.85x10-66), thrombocytopenia (OR=0.86, p=5.70x10-44), platelet defects (OR=0.86, p=2.47x10-43), neutropenia (OR= 0.86, p=5.52x10-18), myeloproliferative disorder (OR= 1.2, p=2.77x10-15). Others included celiac disease (OR=0.713, p=2.98x10-46), ankylosing spondylitis (OR=1.4, p=1.33 x 10-17), hypertension (OR=1.04, p=4.56x10-15), rheumatoid arthritis (OR=1.09, p=1.02x10-13), hematuria (OR=1.05, p=1.96x10-10). Removing major-histocompatibility-complex SNPs abolished associations with known autoimmune diseases, while all other associations remained. We replicated 17 (42.5%) of 40 significant phenotypes available in the Vanderbilt University Medical Center's BioVU. Our findings demonstrate that systemic inflammation can be predicted using the iPRS across multiple ancestries, and the iPRS is associated with numerous clinical endpoints. This multi-ancestry iPRS may have future utility in stratifying risk for inflammation-driven conditions across diverse populations.

Journal Article

Proteomic pathways mediating low socioeconomic status and cardiovascular events in older adults in CHS and ARIC.

BACKGROUND AND AIMS: Many studies have linked socioeconomic status (SES) and cardiovascular outcomes, yet the biologic mechanisms mediating these associations are only partially understood. The objective of this study was to identify molecular mediators of the association of low SES with coronary heart disease (CHD) and stroke. METHODS: This research was conducted in 2942 Black and White adults in the Cardiovascular Health Study (mean age 76.2 years) and 10,689 Black and White adults in the Atherosclerosis Risk in Communities Study (mean age 60.0 years). We used factor analysis to create a composite measure of low educational attainment, low-income, and blue-collar occupation. Approximately 5000 proteins were measured with an aptamer-based method, and CHD and stroke events were adjudicated. Results were stratified by race, which was conceptualized as a social factor. RESULTS: Low SES was associated with 44 and 262 proteins, in Black and White adults, respectively. No protein met the Bonferroni adjusted threshold for statistically significantly mediation among Black participants. Among White participants, 23 proteins mediated the association between SES adversity and CHD and 5 mediated the association between SES adversity and stroke. The strongest mediating associations for CHD included PTPRS, SCG3, and MMP12. The strongest mediating associations for stroke included NCAN, FAM20B, and APLP1. SPARCL1 and CDCP1 remained the strongest mediators of the association between SES adversity and CHD, after adjusting for potential confounders and traditional cardiovascular risk factors. CONCLUSION: We identified several biomarkers that characterize the biologic risk of SES adversity on CHD and stroke.

Aged

Rethinking immune studies: population-level immune variations and the path forward.

Shifts in immune cell proportions underlie disease progression and immunotherapy response, positioning them as promising diagnostic biomarkers and therapeutic targets. However, these shifts also occur naturally across the lifespan and vary with demographic factors such as age and sex, which may complicate their interpretation and clinical utility. While demographic associations have been explored previously, there have been mixed results likely due to small sample sizes and cross-cohort population-specific differences. To address these limitations, we conducted a meta-analysis across 3 large, diverse cohort studies to evaluate associations between 20 immune cell subtypes, 3 informative cell ratios, and a range of sociodemographic variables such as age, sex, self-identified race and ethnicity (SIRE), and socioeconomic status. We find consistent and significant associations across all sociodemographic dimensions. Cytomegalovirus (CMV)-a key driver of immune senescence-emerged as a major contributor to variation in immune composition and CMV antibody levels were higher among women, individuals of lower socioeconomic status, and marginalized racial and ethnic groups. In addition, male sex showed similar patterns of association with immune profiles as aging, whereas race did not. These findings underscore the need to account for diverse sociodemographic factors in immunology study design and participant recruitment to avoid population-specific biases and ensure broadly generalizable results.

Humans

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

Genome-wide gene-sleep interaction study identifies novel lipid loci in 732,564 participants.

BACKGROUND AND AIMS: Deviations from the population mean in sleep duration have been associated with increased risk for developing dyslipidemia and atherosclerotic cardiovascular disease, but the mechanism of effect is poorly characterized. We performed large-scale genome-wide gene-sleep interaction analyses of lipid levels to identify genetic variants underpinning the biomolecular pathways of sleep-associated lipid disturbances and to suggest possible druggable targets. METHODS: We collected data from 55 cohorts with a combined sample size of 732,564 participants (87&#xa0;% European ancestry) with data on lipid traits (high-density lipoprotein [HDL-c] and low-density lipoprotein [LDL-c] cholesterol and triglycerides [TG]). Short (STST) and long (LTST) total sleep time were defined by the extreme 20&#xa0;% of the age- and sex-standardized values within each cohort. Based on cohort-level summary statistics data, we performed meta-analyses for one-degree of freedom tests of interaction and two-degree of freedom joint tests of the SNP-main and -interaction effect on lipid levels. RESULTS: The one-degree of freedom variant-sleep interaction test identified 10 novel loci (Pint<5.0e-9), and we additionally identify 7 loci within the two-degree of freedom analyses (Pjoint<5.0e-9 in combination with Pint<6.6e-6). Multiple loci, including those mapped to APSH (target for aspartic and succinic acid) and SLC8A1 showed biological plausibility and druggability potential based on literature. CONCLUSIONS: Collectively, the 17 (9 with short and 8 with long sleep) loci provided evidence into the biomolecular mechanisms underlying sleep-associated lipid changes, including potential involvement of the vitamin D receptor pathway. Collectively, these findings may contribute developing novel interventions for treating dyslipidemia in people with sleep disturbances.

Humans

Admixture-mapping analysis reveals genetic determinants of the human plasma proteome.

Protein profiling and genetic findings can be integrated to define the genetic architecture of the circulating proteome in chronic diseases. Most self-identified African American (AA) individuals have both African and European genetic ancestry. Admixture mapping can detect genomic association regions in which causal variants exist with substantial differences in allele frequency or effect sizes between genetic ancestries. We performed admixture mapping of the circulating proteome in 1,989 participants from the Jackson Heart Study (JHS), investigating the relation of local African ancestry within genomic regions with levels of circulating proteins. We conditioned protein-local ancestry association models on variants previously found to be associated with those proteins in genome-wide association studies (GWASs). We replicated findings in 196 AA participants from the Multi-Ethnic Study of Atherosclerosis (MESA). 62 proteins were associated with local African ancestry. 21 of 62 remained statistically significant after conditioning on protein-associated variants observed in previous GWASs. 48 of 54 available protein-local ancestry associations were replicated in the MESA. Proteins associated with local African ancestry included chemokines, factors associated with vascular biology and inflammation, and other biologically interesting proteins. Admixture associations unexplained by previously reported protein-associated variants in conditional analysis suggest the existence of causal variants missed by standard GWAS techniques.

Aged

Genetic architecture and analysis practices of circulating metabolites in the NHLBI Trans-Omics for Precision Medicine Program.

Circulating metabolite levels partly reflect the state of human health and diseases and can be impacted by genetic determinants. Hundreds of loci associated with circulating metabolites have been identified; however, most findings focus on predominantly European ancestry or single-study analyses. Leveraging the rich metabolomics resources generated by the National Heart, Lung, and Blood Institute (NHLBI) Trans-Omics for Precision Medicine (TOPMed) Program, we harmonized and accessibly cataloged 1,729 circulating metabolites among 25,058 ancestrally diverse samples. From our comparison of multiple methods, we provided a set of reasonable strategies for outlier and imputation handling to process metabolite data and show that inverse normalization by study and half-minimum imputation provide mostly similar results for pooled or meta-analysis. Following the practical analysis framework, we further performed a genome-wide association analysis on 1,135 selected metabolites using whole-genome sequencing data from 16,359 individuals passing the quality-control filters and discovered 1,775 independent loci associated with 667 metabolites. Among 160 unreported locus-metabolite pairs, we identified associations with loci locating within previously implicated metabolite-associated genes, as well as associations with loci locating in genes such as GAB3 and VSIG4 (located on the X chromosome) that may play a role in metabolic regulation. In the sex-stratified analysis, we revealed 85 independent locus-metabolite pairs with evidence of sexual dimorphism, which were located in well-known metabolic genes such as FADS2, D2HGDH, SUGP1, and UGT2B17, strongly supporting the importance of exploring sex difference in the human metabolome. Taken together, our study depicted the genetic contribution to circulating metabolite levels, providing additional insight into the understanding of human health.

Humans

Estimating population structure using epigenome-wide methylation data.

INTRODUCTION: In epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. METHODS: We used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individuallevel data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. RESULTS: In the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. CONCLUSIONS: Methylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.

Journal Article

Polygenic scores for obstructive sleep apnoea reveal pathways contributing to cardiovascular disease.

BACKGROUND: Obstructive sleep apnoea (OSA) is a common chronic condition, with obesity its strongest risk factor. Polygenic scores (PGSs) summarise the genetic liability to phenotype and can provide insights into relationships between phenotypes. Recently, large datasets that include genetic data and OSA status became available, providing an opportunity to utilise PGS approaches to study the genetic relationship between OSA and other phenotypes, while differentiating OSA-specific from obesity-specific genetic factors. METHODS: Using race/ethnic diverse samples from over 1.2 million individuals from the Million Veteran Program, FinnGen, TOPMed, All of Us (AoU), Geisinger's MyCode, MGB Biobank, and the Human Phenotype Project, we developed and assessed PGSs for OSA, both without (BMIunadjOSA-PGS) and with adjustment for the genetic contributions of BMI (BMIadjOSA-PGS). FINDINGS: Adjusted odds ratios (ORs) for OSA per 1 standard deviation of the PGSs ranged from 1.38 to 2.75. The associations of BMIadjOSA- and BMIunadjOSA-PGSs with CVD outcomes in AoU shared both common and distinct patterns. Only BMIunadjOSA-PGS was associated with type 2 diabetes, heart failure, and coronary artery disease, while both BMIadjOSA- and BMIunadjOSA-PGSs were associated with hypertension and stroke. Sex stratified analyses revealed that BMIadjOSA-PGS association with hypertension was driven by females (OR = 1.1, p-value = 0.002, OR = 1.01 p-value = 0.2 in males). OSA PGSs were also associated with body fat measures with some sex-specific associations. INTERPRETATION: Distinct components of OSA genetic risk are related and independent of obesity. Sex-specific associations with body fat distribution measures may explain differing OSA risks and associations with cardiometabolic morbidities between sexes. FUNDING: R01AG080598.

Humans

Epigenetic mechanisms underlying variation of IL-6, a well-established inflammation biomarker and risk factor for cardiovascular disease.

BACKGROUND AND AIMS: Cardiovascular disease (CVD) is one of the leading causes of morbidity and mortality worldwide, yet the underlying molecular mechanisms remain less understood. Chronic low-grade inflammation is a complex immune response contributing to the pathophysiology of cardiovascular disease. This response is signaled in part by interleukin-6 (IL-6), a pleiotropic, pro-inflammatory cytokine. Phenotypic variance in circulating IL-6 level may be explained in part by DNA methylation which is increasingly being associated with cardiovascular effects. METHODS: In this study we evaluated methylated DNA (CpG sites) associated with blood IL-6 levels across &#x223c;4,400 ancestrally diverse individuals (81&#xa0;% self-reported White; 9&#xa0;% Black or African American, 8&#xa0;% Hispanic or Latino/a, and 2&#xa0;% Chinese American). RESULTS: We identified 178 CpG sites associated with IL-6 (p<0.05/&#x223c;395,000). Among the sites, cg04437762 is located within the transcription unit of IL6R, a current therapeutic target for inflammatory disease, and cg26692003 and cg00464927 were significant for IL6 and IL6ST trans-CpG-gene transcripts. Functional gene expression downstream of methylation identified cellular response to IL-6 and B-cell regulation and activation pathways. Four genes were linked with both a genetic component of cardiovascular disease and an IL-6 associated CpG site. Three CpG sites identified through Mendelian randomization analyses supported inference of a causal effect on IL-6 levels, including the LYN gene that regulates immune cell signaling and has been previously associated with atherosclerosis. CONCLUSIONS: Overall, we identified several novel IL-6-CpG sites and downstream pathways affected by methylation. Follow-up functional studies including the regulation of IL-6 would complement current knowledge of CVD pathophysiology and potential therapeutic targets.

Humans

Alterations in DNA Methylation, Proteomic, and Metabolomic Profiles in African Ancestry Populations with APOL1 Risk Alleles.

KEY POINTS: We aimed to elucidate potential methylation, proteomic, and metabolomic mechanisms by which APOL1 variants may be linked to kidney disease. We report distinct methylation profiling between APOL1 risk allele carriers and noncarriers, many near APOL gene family. We report higher APOL1 protein and lower C18:1 cholesteryl ester in two risk allele carriers. BACKGROUND: The APOL1 high-risk haplotype has been associated with CKD and the deterioration of kidney function, particularly in populations with West African ancestry. However, the mechanisms by which APOL1 risk variants increase the risk for kidney disease and its progression have not been fully elucidated. METHODS: We compared methylation (N=3191; 715 [22%] carriers), proteomic (N=1240; 169 [14%] carriers), and metabolomic (N=6309; 674 [11%] carriers) profiles in African and Hispanic/Latino carriers of two APOL1 high-risk alleles (G1/G1, G2/G2, G1/G2) and noncarriers (G0/G0), excluding heterozygotes (G0/G1, G0/G2), from the Population Architecture using Genomics and Epidemiology Consortium and UK Biobank. In each study, the associations between the APOL1 high-risk haplotype and up to 722,719 cytosine-phosphate-guanine (CpG) sites, 2923 proteins, or 836 metabolites were estimated using covariate-adjusted linear regression models, followed by fixed-effects sample size&#x2013;weighted meta-analyses. RESULTS: Significant associations were observed between APOL1 high-risk haplotype and methylation at 52 CpG sites, with 48 located on chromosome 22 and 18 in the vicinity of APOL1&#x2013;4 and MYH9. All significant CpG sites near APOL2 were hypomethylated, whereas those near APOL3 and APOL4 were hypermethylated. APOL1-associated CpG sites were also identified in genes involved in ion transport and mitochondrial stress pathways. Sensitivity analyses indicated consistent yet attenuated effects among heterozygotes, supporting an additive effect of APOL1 risk alleles. Further analyses of the 52 CpG sites identified two near APOL4 exhibiting G1-specific effects, eight associated with CKD but none with eGFR, and three showing heterogeneity by CKD status. In addition, carrying two APOL1 risk alleles was associated with higher plasma APOL1 protein (&#x3b2;=1.12, PFDR = 2.26e-70) and lower C18:1 cholesteryl ester metabolite (Z=&#x2212;4.50, PFDR = 4.83e-3). CONCLUSIONS: Our results demonstrate differential methylation, proteomic, and metabolomic profiles associated with APOL1 high-risk haplotypes.

APOL1

Rare damaging CCR2 variants are associated with lower lifetime cardiovascular risk.

BACKGROUND: Previous work has shown a role of CCL2, a key chemokine governing monocyte trafficking, in atherosclerosis. However, it remains unknown whether targeting CCR2, the cognate receptor of CCL2, provides protection against human atherosclerotic cardiovascular disease. METHODS: Computationally predicted damaging or loss-of-function (REVEL&#x2009;>&#x2009;0.5) variants within CCR2 were detected in whole-exome-sequencing data from 454,775 UK Biobank participants and tested for association with cardiovascular endpoints in gene-burden tests. Given the key role of CCR2 in monocyte mobilization, variants associated with lower monocyte count were prioritized for experimental validation. The response to CCL2 of human cells transfected with these variants was tested in migration and cAMP assays. Validated damaging variants were tested for association with cardiovascular endpoints, atherosclerosis burden, and vascular risk factors. Significant associations were replicated in six independent datasets (n&#x2009;=&#x2009;1,062,595). RESULTS: Carriers of 45 predicted damaging or loss-of-function CCR2 variants (n&#x2009;=&#x2009;787 individuals) were at lower risk of myocardial infarction and coronary artery disease. One of these variants (M249K, n&#x2009;=&#x2009;585, 0.15% of European ancestry individuals) was associated with lower monocyte count and with both decreased downstream signaling and chemoattraction in response to CCL2. While M249K showed no association with conventional vascular risk factors, it was consistently associated with a lower risk of myocardial infarction (odds ratio [OR]: 0.66, 95% confidence interval [CI]: 0.54-0.81, p&#x2009;=&#x2009;6.1&#x2009;&#xd7;&#x2009;10-5) and coronary artery disease (OR: 0.74, 95%CI: 0.63-0.87, p&#x2009;=&#x2009;2.9&#x2009;&#xd7;&#x2009;10-4) in the UK Biobank and in six replication cohorts. In a phenome-wide association study, there was no evidence of a higher risk of infections among M249K carriers. CONCLUSIONS: Carriers of an experimentally confirmed damaging CCR2 variant are at a lower lifetime risk of myocardial infarction and coronary artery disease without carrying a higher risk of infections. Our findings provide genetic support for the translational potential of CCR2-targeting as an atheroprotective approach.

Humans

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58&#x2009;706 SVs in a study sample of 11&#x2009;556 CAD cases and 42&#x2009;907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Soluble Immune Checkpoint Protein and Lipid Network Associations with All-Cause Mortality Risk: Trans-Omics for Precision Medicine (TOPMed) Program.

Adverse cardiovascular events are emerging with the use of immune checkpoint therapies in oncology. Using datasets in the Trans-Omics for Precision Medicine program (Multi-Ethnic Study of Atherosclerosis, Jackson Heart Study [JHS], and Framingham Heart Study), we examined the association of immune checkpoint plasma proteins with each other, their associated protein network with high-density lipoprotein cholesterol (HDL-C) and low-density lipoprotein cholesterol (LDL-C), and the association of HDL-C- and LDL-C-associated protein networks with all-cause mortality risk. Plasma levels of LAG3 and HAVCR2 showed statistically significant associations with mortality risk. Colocalization analysis using genome wide-association studies of HDL-C or LDL-C and protein quantitative trait loci from JHS and the Atherosclerosis Risk in Communities identified TFF3 rs60467699 and CD36 rs3211938 variants as significantly colocalized with HDL-C; in contrast, none colocalized with LDL-C. The measurement of plasma LAG3, HAVCR2, and associated proteins plus targeted genotyping may identify patients at increased mortality risk.

Journal Article

The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup.

Polygenic risk scores (PRSs) depend on genetic ancestry due to differences in allele frequencies between ancestral populations. This leads to implementation challenges in diverse populations. We propose a framework to calibrate PRS based on ancestral makeup. We define a metric called "expected PRS" (ePRS), the expected value of a PRS based on one's global or local admixture patterns. We further define the "residual PRS" (rPRS), measuring the deviation of the PRS from the ePRS. Simulation studies confirm that it suffices to adjust for ePRS to obtain nearly unbiased estimates of the PRS-outcome association without further adjusting for PCs. Using the TOPMed dataset, the estimated effect size of the rPRS adjusting for the ePRS is similar to the estimated effect of the PRS adjusting for genetic PCs. Similarly, we applied the ePRS framework to six cardiovascular-related traits in the All of Us dataset, and the results are consistent with those from the TOPMed analysis. The ePRS framework can protect from population stratification in association analysis and provide an equitable strategy to quantify genetic risk across diverse populations.

Journal Article

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency &#x2264;&#x2009;0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans

Association analysis of mitochondrial DNA heteroplasmic variants: Methods and application.

We rigorously assessed a comprehensive association testing framework for heteroplasmy, employing both simulated and real-world data. This framework employed a variant allele fraction (VAF) threshold and harnessed multiple gene-based tests for robust identification and association testing of heteroplasmy. Our simulation studies demonstrated that gene-based tests maintained an appropriate type I error rate at &#x3b1;&#x202f;=&#x202f;0.001. Notably, when 5&#x202f;% or more heteroplasmic variants within a target region were linked to an outcome, burden-extension tests (including the adaptive burden test, variable threshold burden test, and z-score weighting burden test) outperformed the sequence kernel association test (SKAT) and the original burden test. Applying this framework, we conducted association analyses on whole-blood derived heteroplasmy in 17,507 individuals of African and European ancestries (31&#x202f;% of African Ancestry, mean age of 62, with 58&#x202f;% women) with whole genome sequencing data. We performed both cohort- and ancestry-specific association analyses, followed by meta-analysis on both pooled samples and within each ancestry group. Our results suggest that mtDNA-encoded genes/regions are likely to exhibit varying rates in somatic aging, with the notably strong associations observed between heteroplasmy in the RNR1 and RNR2 genes (p&#x202f;<&#x202f;0.001) and advance aging by the Original Burden test. In contrast, SKAT identified significant associations (p&#x202f;<&#x202f;0.001) between diabetes and the aggregated effects of heteroplasmy in several protein-coding genes. Further research is warranted to validate these findings. In summary, our proposed statistical framework represents a valuable tool for facilitating association testing of heteroplasmy with disease traits in large human populations.

Humans

The Genetic Determinants and Genomic Consequences of Non-Leukemogenic Somatic Point Mutations.

Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as point mutations or mosaic chromosomal alterations (mCAs) in genes associated with hematologic malignancy, are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we used 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantified CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we developed a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We performed a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. After fine-mapping and variant-to-gene analyses, we identified seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1), and one locus associated with a sex-associated mutation pathway (SRGAP2C). We performed a secondary analysis excluding individuals with mCAs, finding that the genetic architecture was largely unaffected by their inclusion. Functional analyses of SMC4 and NRIP1 implicated altered HSC self-renewal and proliferation as the primary mediator of mutation burden in blood. We then performed comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we performed phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count and increased risk for incident peripheral artery disease, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS without a paired-tissue sample and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.

Journal Article