Search PubMedSearch

Biomedical subjects

Joshua C Denny

Publications and source records attributed to Joshua C Denny.

3 recordsLinked to original sources

Robust replication of associations across patient-mediated and provider-sourced EHR data in the All of Us research program.

The All of Us Research Program is assembling a nationwide cohort with electronic health record (EHR) resources through two complementary pathways: healthcare provider organization (HPO)-sourced EHRs and patient-mediated EHR (PME) contributed through patient portal linkages. The comparative research utility of these two data sources has not been systematically evaluated. Here, we compared PME and HPO EHRs with respect to disease prevalence, phenotype-phenotype associations, and replication of established genotype-phenotype associations using data from 19,703 PME and 373,887 HPO participants. We benchmarked disease prevalence against national estimates, conducted phenome-wide association studies for 10 commonly studied diseases, and tested replication of more than 5000 established genotype-phenotype associations across multiple ancestral groups. Disease prevalence was consistently lower in PME than in HPO, although prevalence of most diseases in both cohorts exceeded national estimates. Both data sources reproduced known phenotype-phenotype associations and showed moderate-to-strong concordance in effect sizes across the phenome. The overall genotype-phenotype replication rate was 49.1% (5399/10,999) in HPO and 5.9% (381/6482) in PME across ancestral groups, with effect sizes strongly correlated among well-powered associations (R&#x2009;=&#x2009;0.84, P&#x2009;<&#x2009;0.001). To disentangle the impact of sample size from data quality, we performed 1:1 propensity score matching. After matching, the replication gap in genotype-phenotype associations narrowed from 8.3-fold to 1.3-fold, with equivalent replication rates among adequately powered associations and strongly concordant effect sizes; comorbidity patterns were also consistent across all 10 diseases tested. These findings demonstrate that both data sources are valuable for clinical and genomic research and can inform other cohorts integrating provider-derived and patient-mediated EHRs.

Computational biology and bioinformatics

Systematic common and rare variant association testing in 392,030 whole genomes in All of Us.

Large-scale genome-wide association studies (GWAS) and rare variant association studies (RVAS) from population biobanks provide valuable resources for gene discovery in complex human traits. We present an analysis of the All of Us Research Program v8 release, which includes whole genome sequencing data and harmonized phenotypic information of 392,030 participants after quality control, enabling a unified investigation of rare and common variants across a spectrum of human traits and diseases. We build an extensive phenome- and genome-wide ("All by All") computational framework to perform GWAS and RVAS on 3,602 phenotypes and identify 49,863 approximately independent, high-quality single-variant and gene-level associations. Meta-analyses of All of Us and UK Biobank, with sample sizes as large as 786,871 participants, further enhance statistical power and find 193 pLoF gene-phenotype associations that are not significant in either cohort alone, including 22 associations not highlighted by previous studies. We also present a public interactive browser that integrates association results for common and rare variants to facilitate interpretation and rapid querying of summary statistics, along with supporting documentation, and a Featured Workspace in the All of Us Researcher Workbench. Our framework will apply to iterative data releases as All of Us grows, empowering researchers worldwide to uncover insights into the functional effects of genetic components on complex traits and diseases.

Journal Article

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans