Search PubMedSearch

Biomedical subjects

Yan Cui

Publications and source records attributed to Yan Cui.

3 recordsLinked to original sources

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

RAS signaling in lung adenocarcinoma is defined by lineage context and DUSP4 loss.

BACKGROUNDThe molecular landscape of lung adenocarcinoma (LUAD) is often illustrated as a driver-oncogene pie chart, but identical mutations exhibit heterogeneous signaling shaped by comutations, transcriptional programs, and lineage context. We propose a lineage-integrated signaling framework using an EGFR mutation signature (mSig).METHODSWe defined EGFR mSig using differentially expressed genes in EGFR-mutant (EGFR-mt) LUADs. Semisupervised clustering and machine learning models were used to test reproducibility in different combinations of datasets. We analyzed molecular subtypes, lineage markers, co-occurring mutations, and EGFR copy number alterations in EGFR mSig-defined subtypes of LUAD.RESULTSEGFR mSig showed robust classification performance (area under receiver operating characteristic curve = 0.83-0.95; mean negative predictive value = 96.3%). Validated gene expression subtypes and lung lineage markers were closely aligned with EGFR mSig status. Most EGFR mSig+ tumors, including many without EGFR mutations, belonged to the bronchioid subtype. A subset of canonical RAS mutations were mSig+ and mirrored the EGFR mutation pattern. EGFR WT/mSig- tumors were enriched for nonbronchioid subtypes and had comutations in TP53 or RAS/RAF/RTKs. We highlight a parsimonious collection of coordinated mutations, including RAS, KEAP1, STK11, TP53, and CDKN2A, that taken together suggest coordination of tumor signaling previously suggested but now reproduced and expanded.CONCLUSIONA potentially novel EGFR mSig that captures the transcriptional footprint of EGFR activation revealed a subset of EGFR WT LUADs with mt-like features. mSig refines LUAD taxonomy beyond mutation-only pie-chart models by incorporating lineage and comutation context. Lineage-directed stratification with coalteration identifies clinically relevant groups across EGFR and RAS states and highlights treatment opportunities for patients currently considered oncogene-negative.FUNDINGNational Cancer Institute (NCI) U01CA272541, R01CA262296, U24CA264021, UG1CA233333, R01CA211939.

Humans