Search PubMedSearch

Biomedical subjects

Haoyu Zhang

Publications and source records attributed to Haoyu Zhang.

4 recordsLinked to original sources

Identification of immune cell type-specific susceptibility genes in multiple cancers using transcriptome-wide association studies.

BACKGROUND: Transcriptome-wide association studies (TWAS) integrate gene expression and genome-wide association studies (GWAS) to identify disease susceptibility genes. Because gene expression varies substantially across cell types within tissues, cell type-specific prediction models may enhance the power of TWAS. METHODS: We conducted cell type-specific TWAS leveraging single-cell RNA sequencing data from the OneK1K cohort (14 immune cell types, 1.27 million cells) and GWAS summary statistics for 7 cancers (>290 000 cases in total). To improve prediction accuracy, we developed a modeling framework that incorporates shared gene expression effects across cell types. RESULTS: At a false discovery rate of 5%, we identified 106 (Bonferroni 5%: 13) previously unreported loci for breast cancer, 51 (4) loci for prostate cancer, 11 (4) loci for lung cancer, 39 (5) loci for melanoma, 9 (1) loci for ovarian cancer, and 2 (1) loci for diffuse large B-cell lymphoma, with most genes exhibiting cell type specificity. Gene set analyses confirmed joint associations of unreported genes with breast and prostate cancer risk in UK Biobank data. Additional lung tissue single-cell RNA sequencing data with 113 individuals validated 18 of 32 (56.3%) statistically significant genes for lung cancer. Across cancers, 139 statistically significant genes were shared by at least 2 cancer types and were primarily enriched in specific immune cell types. CONCLUSION: Cell type-specific TWAS improve the identification of novel cancer susceptibility loci and provide insights into the immune landscape of cancer etiology.

Humans

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

Cross-ancestry proteome-wide Mendelian randomization prioritizes 12 plasma protein candidates for breast cancer risk.

The plasma proteome provides a molecular bridge between genetic variation and disease risk, yet its contribution to breast cancer susceptibility across ancestries remains unclear. We conducted a proteome-wide Mendelian randomization (MR) study of 2,923 plasma proteins using cis-protein quantitative trait loci from 34,557 European participants in the UK Biobank Pharma Proteomics Project, integrated with genome-wide association studies of 156,901 breast cancer cases and 204,634 controls of European, East Asian, and African ancestries. Cross-ancestry meta-analysis identified 12 candidate proteins associated with breast cancer risk (P < 2.5&#xd7;10-5), including six previously reported and six newly implicated in MR studies. DNPH1 showed cross-ancestry heterogeneity, with a risk-increasing association in European populations and a nominally inverse association in East Asian populations. CASP8, RALB, and USP28 displayed subtype-differentiated associations. Orthogonal validation provided variable support: six demonstrated strong evidence of statistical colocalization; four replicated in an independent European proteomic dataset (deCODE, n = 35,559); two replicated in an independent East Asian proteomic dataset (JCTF, n = 1,384); and four were supported by polygenic-score analyses in the ancestrally diverse All of Us cohort (9,250 cases, 214,857 controls). These findings prioritize a high-confidence subset of plasma proteins, including LRRC25, PARK7, and LRRC37A2, for future mechanistic and translational investigation.

Mendelian randomization

DiscoDivas: Leveraging genetic ancestry continuum information to interpolate PRS for admixed populations.

The relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits polygenic risk score (PRS) development and equitable clinical translation for admixed populations. Under the assumption that the most informative PRS model for a genetically homogeneous sample varies linearly in an ancestry continuum space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, genetic ancestries, leveraging multiple PRS models fine-tuned within existing samples, which are mostly of single ancestry, and genetic distance. DiscoDivas treats genetic ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar genetic ancestry samples. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed individuals.

PRS harmonization