Search PubMedSearch

Biomedical subjects

Eric Boerwinkle

Publications and source records attributed to Eric Boerwinkle.

16 recordsLinked to original sources

Proteomic Profiling of Pulmonary Function and Cardiovascular Disease Risk in the Atherosclerosis Risk in Communities Study.

BACKGROUND: Pulmonary function is linked to cardiovascular disease risk; however, the underlying mechanisms remain unclear. We aimed to identify protein biomarkers associated with pulmonary function and examine their impact on incident chronic obstructive pulmonary disease, coronary heart disease, heart failure, and all-cause mortality. METHODS: Data from White and Black Americans in the Atherosclerosis Risk in Communities study (visit 2: N=11&#x2009;354, mean age=57 years; visit 5: N=3517, mean age=75 years), a prospective cohort, were analyzed. Linear regression assessed associations between protein levels and pulmonary function measures, including forced expiratory volume in 1 second and forced vital capacity. The impact of the identified proteins on incident chronic obstructive pulmonary disease, coronary heart disease, heart failure, and mortality was estimated using logistic regression and Cox proportional hazards models. Pathway enrichment and Mendelian randomization explored underlying biological functions and causal effects. RESULTS: Of 4766 proteins analyzed, 364 were cross-sectionally associated with forced expiratory volume in 1 second (and forced vital capacity (false discovery rate<0.05). Ninety-four and 270 proteins had concordant positive and negative effects, respectively. Five pathways related to pulmonary and cardiac function were enriched. Of the 364 proteins, 112 were linked to all 4 outcomes, where 86 were associated with increased risk (odds ratio/hazard ratio [OR/HR], 1.05-1.42) and 26 with reduced risk (OR/HR, 0.69-0.96). Six proteins (STAT3 [signal transducer and activator of transcription 3], MIC-1 [growth differentiation factor 15], apoA-II [apolipoprotein A-II], TPST1 [protein-tyrosine sulfotransferase 1], integrin a1b1 [integrin alpha-I: beta-1 complex], and BLC [C-X-C motif chemokine 13]) showed potential inverse causal effects on with forced expiratory volume in 1 second and forced vital capacity, and integrin a1b1 demonstrated consistent inverse associations with chronic obstructive pulmonary disease, coronary heart disease, and heart failure risks. CONCLUSIONS: Proteins associated with pulmonary function may influence CVD risk. Six proteins, including integrin a1b1, represent promising targets for future interventions.

Aged

Genome-wide gene-sleep interaction study identifies novel lipid loci in 732,564 participants.

BACKGROUND AND AIMS: Deviations from the population mean in sleep duration have been associated with increased risk for developing dyslipidemia and atherosclerotic cardiovascular disease, but the mechanism of effect is poorly characterized. We performed large-scale genome-wide gene-sleep interaction analyses of lipid levels to identify genetic variants underpinning the biomolecular pathways of sleep-associated lipid disturbances and to suggest possible druggable targets. METHODS: We collected data from 55 cohorts with a combined sample size of 732,564 participants (87&#xa0;% European ancestry) with data on lipid traits (high-density lipoprotein [HDL-c] and low-density lipoprotein [LDL-c] cholesterol and triglycerides [TG]). Short (STST) and long (LTST) total sleep time were defined by the extreme 20&#xa0;% of the age- and sex-standardized values within each cohort. Based on cohort-level summary statistics data, we performed meta-analyses for one-degree of freedom tests of interaction and two-degree of freedom joint tests of the SNP-main and -interaction effect on lipid levels. RESULTS: The one-degree of freedom variant-sleep interaction test identified 10 novel loci (Pint<5.0e-9), and we additionally identify 7 loci within the two-degree of freedom analyses (Pjoint<5.0e-9 in combination with Pint<6.6e-6). Multiple loci, including those mapped to APSH (target for aspartic and succinic acid) and SLC8A1 showed biological plausibility and druggability potential based on literature. CONCLUSIONS: Collectively, the 17 (9 with short and 8 with long sleep) loci provided evidence into the biomolecular mechanisms underlying sleep-associated lipid changes, including potential involvement of the vitamin D receptor pathway. Collectively, these findings may contribute developing novel interventions for treating dyslipidemia in people with sleep disturbances.

Humans

Wastewater Sequencing Reveals Persistent Circulation and Rising Prevalence of Several Oncogenic Viruses Across Texas.

BACKGROUND: Oncogenic viruses cause high-risk cancers in humans and are responsible for nearly 20% of all cancer cases worldwide. Currently, very limited data exists in the realm of wastewater-based viral epidemiology (WBE) of cancer-causing viruses, with existing studies using targeted approaches (i.e PCR-based approaches) which lack scalability. Our study aims to carry out WBE with hybrid-capture probes to detect and track multiple oncogenic viruses simultaneously in wastewater across Texas, USA, overcoming the drawbacks associated with targeted approaches. METHODS: Here, we used a hybrid-capture approach to detect, filter and sequence oncogenic virus signals from wastewater samples collected over a duration of three years, from May 2022 to May 2025. Once viral reads were sequenced, we utilized established computational tools to characterize reads into their respective virus of origin. Next, viral abundances of each characterized oncogenic virus were tracked over time and read coverage across their genomes was measured using read mapping techniques. FINDINGS: We detected six known oncogenic viruses, along with three suspected oncogenic viruses across all sampling locations within Texas. Over three years, viral abundance gradually increased, with distinct peaks and dips over the summer and winter months. The prevalence of high-risk viruses such as HPV and EBV rose sharply, with increases in abundance observed post-2024. We also obtained nearly 100% genome coverage with viral reads captured using a hybrid-capture technique for almost all oncogenic viruses and their types. INTERPRETATIONS: Our study shows that a hybrid-capture method can efficiently overcome the challenges faced with using targeted approaches for WBE. Using this method, we get broader read coverage, coupled with concurrent and consistent real-time tracking dynamics of multiple oncogenic viruses. Our findings also emphasize the persistent circulation and rising prevalence of high-risk cancer-causing viruses, underscoring the need for sustained public health interventions to protect communities and assess viral prevalence in high-risk populations. FUNDING: This work was supported by S.B. 1780, 87th Legislature, 2021 Reg. Sess. (Texas 2021), the Baylor College of Medicine and the Alkek Foundation Seed Funds.

Journal Article

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246&#xa0;K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246&#xa0;K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86&#xa0;K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

Genetic study of von Willebrand factor antigen levels &#x2264; 50 IU/dL identifies variants associated with increased risk of von Willebrand disease and bleeding.

BACKGROUND: von Willebrand disease (VWD) is a common inherited bleeding disorder caused by low levels or activity of circulating von Willebrand factor (VWF). Genetic susceptibility to VWF antigen (VWF:Ag) below normal (&#x2264; 50 IU/dL) in the general population is underexplored. OBJECTIVES: To identify genetic variants influencing VWF:Ag levels &#x2264; 50 IU/dL. METHODS: We performed a genome-wide association study in 926 cases with VWF:Ag levels &#x2264; 50 IU/dL and 12 846 controls from 7 studies from the Trans-Omics for Precision Medicine program. We then examined whether significant genome-wide findings were also associated with clinical diagnosis of VWD in 5 biobanks with 708 VWD cases and 1 286 069 controls, and with 6 bleeding and thrombotic disorders in FinnGen. RESULTS: Variants at 2 loci were associated (P < 5 &#xd7; 10-9) with VWF:Ag levels &#x2264; 50 IU/dL: ABO and VWF. The VWF index variant, p.Tyr1584Cys, is a rare (0.22%) missense variant with odds ratio (OR) of 78.58, while the ABO index variant is a common intronic variant with a smaller effect (OR = 2.52). Notably, both VWF (OR = 7.16) and ABO (OR = 1.57) variants were also associated (P < .025) with diagnosed VWD. Among p.Tyr1584Cys heterozygotes, the penetrance of VWF:Ag levels &#x2264; 50 IU/dL was 24.2% and the penetrance of diagnosed VWD was 0.3%. p.Tyr1584Cys was associated (P < .0042) with increased odds of heavy menstrual bleeding (OR = 1.27), iron deficiency anemia (OR = 1.55), and intrapartum hemorrhage (OR = 2.20), but decreased odds of deep vein thrombosis (OR = 0.54). CONCLUSIONS: Although there are currently conflicting interpretations of pathogenicity p.Tyr1584Cys, our results suggest that it is a low penetrance pathogenic variant that contributes to VWF:Ag levels &#x2264; 50 IU/dL, bleeding, and VWD.

Humans

Alterations in DNA Methylation, Proteomic, and Metabolomic Profiles in African Ancestry Populations with APOL1 Risk Alleles.

KEY POINTS: We aimed to elucidate potential methylation, proteomic, and metabolomic mechanisms by which APOL1 variants may be linked to kidney disease. We report distinct methylation profiling between APOL1 risk allele carriers and noncarriers, many near APOL gene family. We report higher APOL1 protein and lower C18:1 cholesteryl ester in two risk allele carriers. BACKGROUND: The APOL1 high-risk haplotype has been associated with CKD and the deterioration of kidney function, particularly in populations with West African ancestry. However, the mechanisms by which APOL1 risk variants increase the risk for kidney disease and its progression have not been fully elucidated. METHODS: We compared methylation (N=3191; 715 [22%] carriers), proteomic (N=1240; 169 [14%] carriers), and metabolomic (N=6309; 674 [11%] carriers) profiles in African and Hispanic/Latino carriers of two APOL1 high-risk alleles (G1/G1, G2/G2, G1/G2) and noncarriers (G0/G0), excluding heterozygotes (G0/G1, G0/G2), from the Population Architecture using Genomics and Epidemiology Consortium and UK Biobank. In each study, the associations between the APOL1 high-risk haplotype and up to 722,719 cytosine-phosphate-guanine (CpG) sites, 2923 proteins, or 836 metabolites were estimated using covariate-adjusted linear regression models, followed by fixed-effects sample size&#x2013;weighted meta-analyses. RESULTS: Significant associations were observed between APOL1 high-risk haplotype and methylation at 52 CpG sites, with 48 located on chromosome 22 and 18 in the vicinity of APOL1&#x2013;4 and MYH9. All significant CpG sites near APOL2 were hypomethylated, whereas those near APOL3 and APOL4 were hypermethylated. APOL1-associated CpG sites were also identified in genes involved in ion transport and mitochondrial stress pathways. Sensitivity analyses indicated consistent yet attenuated effects among heterozygotes, supporting an additive effect of APOL1 risk alleles. Further analyses of the 52 CpG sites identified two near APOL4 exhibiting G1-specific effects, eight associated with CKD but none with eGFR, and three showing heterogeneity by CKD status. In addition, carrying two APOL1 risk alleles was associated with higher plasma APOL1 protein (&#x3b2;=1.12, PFDR = 2.26e-70) and lower C18:1 cholesteryl ester metabolite (Z=&#x2212;4.50, PFDR = 4.83e-3). CONCLUSIONS: Our results demonstrate differential methylation, proteomic, and metabolomic profiles associated with APOL1 high-risk haplotypes.

APOL1

Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project.

BackgroundPrior studies examined variants within presenilin-2 (PSEN2), presenilin-1 (PSEN1), and amyloid precursor protein (APP) genes. However, previously-reported clinically-relevant variants and other predicted damaging missense (DM) variants have not been characterized in a newer release of the Alzheimer's Disease Sequencing Project (ADSP).ObjectiveTo characterize previously-reported clinically-relevant variants and DM variants in PSEN2, PSEN1, APP within the participants from the ADSP.MethodsWe identified rare variants (MAF&#x2009;<&#x2009;1%) in PSEN2, PSEN1, and APP in 14,641 individuals with whole genome sequencing and 16,849 individuals with whole exome sequencing available (Ntotal&#x2009;=&#x2009;31,490). We additionally curated variants from ClinVar, OMIM, and Alzforum and report carriers of variants in clinical databases as well as predicted DM variants in these genes.ResultsWe detected 31 previously-reported clinically-relevant variants with alternate alleles observed within the ADSP: 4 variants in PSEN2, 25 in PSEN1, and 2 in APP. The overall variant carrier rate for the 31 clinically-relevant variants in the ADSP was 0.3%. We observed that 79.5% of the variant carriers were cases compared to 3.9% were controls. In those with AD, the mean age of onset of AD among carriers of these clinically-relevant variants was 19.6&#x2009;&#xb1;&#x2009;1.4 years earlier compared with noncarriers (p&#x2009;=&#x2009;7.8&#x2009;&#xd7;&#x2009;10-57). Additionally, we identified 197 rare variants (MAF&#x2009;<&#x2009;1%) within ADSP participants not reported in known clinical databases.ConclusionsA small proportion of individuals in the ADSP are carriers of a previously-reported clinically-relevant variant allele for AD and these participants have significantly earlier age of AD onset compared to noncarriers.

Humans

Whole genome sequence-based association analysis of African American individuals with bipolar disorder and schizophrenia.

In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole genome sequencing of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls. To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of BD association with single-variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD GWAS loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.

Journal Article

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58&#x2009;706 SVs in a study sample of 11&#x2009;556 CAD cases and 42&#x2009;907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

The expected polygenic risk score (ePRS) framework: an equitable metric for quantifying polygenetic risk via modeling of ancestral makeup.

Polygenic risk scores (PRSs) depend on genetic ancestry due to differences in allele frequencies between ancestral populations. This leads to implementation challenges in diverse populations. We propose a framework to calibrate PRS based on ancestral makeup. We define a metric called "expected PRS" (ePRS), the expected value of a PRS based on one's global or local admixture patterns. We further define the "residual PRS" (rPRS), measuring the deviation of the PRS from the ePRS. Simulation studies confirm that it suffices to adjust for ePRS to obtain nearly unbiased estimates of the PRS-outcome association without further adjusting for PCs. Using the TOPMed dataset, the estimated effect size of the rPRS adjusting for the ePRS is similar to the estimated effect of the PRS adjusting for genetic PCs. Similarly, we applied the ePRS framework to six cardiovascular-related traits in the All of Us dataset, and the results are consistent with those from the TOPMed analysis. The ePRS framework can protect from population stratification in association analysis and provide an equitable strategy to quantify genetic risk across diverse populations.

Journal Article

Cardiovascular Risk Factors and Genetic Risk in Transthyretin V142I Carriers.

BACKGROUND: Nearly 3% to 4% of Black individuals in the United States carry the transthyretin V142I variant, which increases their risk of heart failure. However, the role of cardiovascular (CV) risk factors (RFs) in influencing the risk of clinical outcomes among V142I variant carriers is unknown. OBJECTIVES: This study aimed to assess the impact of CV RFs on the risk of heart failure in V142I carriers. METHODS: This study included self-identified Black individuals without prevalent heart failure from 6 TOPMed (Trans-Omics for Precision Medicine) cohorts, the REGARDS (Reasons for Geographic And Racial Differences in Stroke) study, and the All of Us Research Program. The cohort was stratified based on the V142I genotype and the number of CV RFs (hypertension, diabetes, obesity, and hypercholesterolemia). Adjusted Cox models were used to assess the association of heart failure with the V142I genotype and CV RF profile, taking noncarriers with a favorable CV RF profile as reference. RESULTS: The cross-sectional analysis, including 1,625 V142I carriers among 48,365 Black individuals, found that the prevalence of CV RFs did not vary by V142I carrier status. In the longitudinal analysis, there were 587 (3.2%) V142I carriers among 18,407 Black individuals (median age: 60 years [Q1-Q3: 52-68 years], 63.0% female). Among carriers, the heart failure risk was attenuated with a favorable (0 or 1 RF) CV RF profile (adjusted HR: 2.26; 95%&#xa0;CI: 1.58-3.23) compared with an unfavorable (3 or 4 RFs) CV RF profile (adjusted HR: 4.14; 95%&#xa0;CI: 2.79-6.14). CONCLUSIONS: A favorable CV RF profile lowers but does not abrogate V142I variant-associated heart failure risk. This study highlights the importance of having a favorable CV RF profile among V142I carriers for risk reduction of heart failure.

Aged

Association of common and rare variants with Alzheimer's disease in more than 13,000 diverse individuals with whole-genome sequencing from the Alzheimer's Disease Sequencing Project.

INTRODUCTION: Alzheimer's disease (AD) is a common disorder of the elderly that is both highly heritable and genetically heterogeneous. METHODS: We investigated the association of AD with both common variants and aggregates of rare coding and non-coding variants in 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. RESULTS: Pooled-population analyses of all individuals identified genetic variants at apolipoprotein E (APOE) and BIN1 associated with AD (p&#xa0;<&#xa0;5&#xa0;&#xd7;&#xa0;10-8). Subgroup-specific analyses identified a haplotype on chromosome 14 including PSEN1 associated with AD in Hispanics, further supported by aggregate testing of rare coding and non-coding variants in the region. Common variants in LINC00320 were observed associated with AD in Black individuals (p&#xa0;=&#xa0;1.9&#xa0;&#xd7;&#xa0;10-9). Finally, we observed rare non-coding variants in the promoter of TOMM40 distinct of APOE in pooled-population analyses (p&#xa0;=&#xa0;7.2&#xa0;&#xd7;&#xa0;10-8). DISCUSSION: We observed that complementary pooled-population and subgroup-specific analyses offered unique insights into the genetic architecture of AD. HIGHLIGHTS: We determine the association of genetic variants with Alzheimer's disease (AD) using 13,371 individuals of diverse ancestry with whole genome sequencing (WGS) data. We identified genetic variants at apolipoprotein E (APOE), BIN1, PSEN1, and LINC00320 associated with AD. We observed rare non-coding variants in the promoter of TOMM40 distinct of APOE.

Humans

Rare variant contribution to the heritability of coronary artery disease.

Whole genome sequences (WGS) enable discovery of rare variants which may contribute to missing heritability of coronary artery disease (CAD). To measure their contribution, we apply the GREML-LDMS-I approach to WGS of 4949 cases and 17,494 controls of European ancestry from the NHLBI TOPMed program. We estimate CAD heritability at 34.3% assuming a prevalence of 8.2%. Ultra-rare (minor allele frequency &#x2264;&#x2009;0.1%) variants with low linkage disequilibrium (LD) score contribute ~50% of the heritability. We also investigate CAD heritability enrichment using a diverse set of functional annotations: i) constraint; ii) predicted protein-altering impact; iii) cis-regulatory elements from a cell-specific chromatin atlas of the human coronary; and iv) annotation principal components representing a wide range of functional processes. We observe marked enrichment of CAD heritability for most functional annotations. These results reveal the predominant role of ultra-rare variants in low LD on the heritability of CAD. Moreover, they highlight several functional processes including cell type-specific regulatory mechanisms as key drivers of CAD genetic risk.

Humans

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans

Association analysis of mitochondrial DNA heteroplasmic variants: Methods and application.

We rigorously assessed a comprehensive association testing framework for heteroplasmy, employing both simulated and real-world data. This framework employed a variant allele fraction (VAF) threshold and harnessed multiple gene-based tests for robust identification and association testing of heteroplasmy. Our simulation studies demonstrated that gene-based tests maintained an appropriate type I error rate at &#x3b1;&#x202f;=&#x202f;0.001. Notably, when 5&#x202f;% or more heteroplasmic variants within a target region were linked to an outcome, burden-extension tests (including the adaptive burden test, variable threshold burden test, and z-score weighting burden test) outperformed the sequence kernel association test (SKAT) and the original burden test. Applying this framework, we conducted association analyses on whole-blood derived heteroplasmy in 17,507 individuals of African and European ancestries (31&#x202f;% of African Ancestry, mean age of 62, with 58&#x202f;% women) with whole genome sequencing data. We performed both cohort- and ancestry-specific association analyses, followed by meta-analysis on both pooled samples and within each ancestry group. Our results suggest that mtDNA-encoded genes/regions are likely to exhibit varying rates in somatic aging, with the notably strong associations observed between heteroplasmy in the RNR1 and RNR2 genes (p&#x202f;<&#x202f;0.001) and advance aging by the Original Burden test. In contrast, SKAT identified significant associations (p&#x202f;<&#x202f;0.001) between diabetes and the aggregated effects of heteroplasmy in several protein-coding genes. Further research is warranted to validate these findings. In summary, our proposed statistical framework represents a valuable tool for facilitating association testing of heteroplasmy with disease traits in large human populations.

Humans

The Genetic Determinants and Genomic Consequences of Non-Leukemogenic Somatic Point Mutations.

Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as point mutations or mosaic chromosomal alterations (mCAs) in genes associated with hematologic malignancy, are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we used 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantified CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we developed a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We performed a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. After fine-mapping and variant-to-gene analyses, we identified seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1), and one locus associated with a sex-associated mutation pathway (SRGAP2C). We performed a secondary analysis excluding individuals with mCAs, finding that the genetic architecture was largely unaffected by their inclusion. Functional analyses of SMC4 and NRIP1 implicated altered HSC self-renewal and proliferation as the primary mediator of mutation burden in blood. We then performed comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we performed phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count and increased risk for incident peripheral artery disease, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS without a paired-tissue sample and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.

Journal Article