Search PubMedSearch

Biomedical subjects

Xihong Lin

Publications and source records attributed to Xihong Lin.

4 recordsLinked to original sources

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246 K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246 K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86 K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans

Causal Mediation Analysis for Integrating Exposure, Genomic, and Phenotype Data.

Causal mediation analysis provides an attractive framework for integrating diverse types of exposure, genomic, and phenotype data. Recently, this field has seen a surge of interest, largely driven by the increasing need for causal mediation analyses in health and social sciences. This article aims to provide a review of recent developments in mediation analysis, encompassing mediation analysis of a single mediator and a large number of mediators, as well as mediation analysis with multiple exposures and mediators. Our review focuses on the recent advancements in statistical inference for causal mediation analysis, especially in the context of high-dimensional mediation analysis. We delve into the complexities of testing mediation effects, especially addressing the challenge of testing a large number of composite null hypotheses. Through extensive simulation studies, we compare the existing methods across a range of scenarios. We also include an analysis of data from the Normative Aging Study, which examines DNA methylation CpG sites as potential mediators of the effect of smoking status on lung function. We discuss the pros and cons of these methods and future research directions.

causal inference

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans