Search PubMedSearch

SEARCH · Search PubMed

Results for “Genetic ancestry”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Leveraging local ancestry and cross-ancestry genetic architecture to improve genetic prediction of complex traits in admixed populations.

The broader application of polygenic risk score (PRS) is hindered by the limited transferability of PRS developed in Europeans to non-European populations. While many statistical methods have been developed to improve the performance of PRS in non-European populations, most of them focused on discrete genetic ancestry clusters and did not consider admixed individuals. Admixed individuals pose a unique challenge for PRS calculation due to the complexity of local ancestry and cross-ancestry effect sizes. Here, we present a statistical method called SDPR_admix for calculating PRS in admixed individuals. SDPR_admix characterizes the joint distribution of the effect sizes of a genetic variant with two ancestries to be both zero, ancestry enriched, or shared with correlation. SDPR_admix outperformed other methods in simulations and improved the prediction of real traits in European-African admixed individuals in UK Biobank when trained on the Population Architecture using Genomics and Epidemiology (PAGE) dataset (N = 13,000). Deployment of SDPR_admix on All of Us (N = 52,000) further increased the prediction accuracy by approximately 5-fold on average compared with training on PAGE. This enhancement was achieved with manageable computational time and cost, demonstrating the feasibility of training PRS models on large-scale All of Us data. We provided several examples demonstrating that both ancestral-enriched and shared effects, as included in the SDPR_admix prediction model, are helpful for improving polygenic prediction in admixed populations. We also applied SDPR_admix to construct PRS for admixed Americans with mixture of European and Amerindigenous ancestries and showed that SDPR_admix overall outperformed other methods.

Humans

The Landscape of Genomic and Socioeconomic Variables in Patients with Colorectal Cancer Based on Genetic Ancestry.

BACKGROUND: Despite differences in tumor alterations across genetic ancestries, investigations of the colorectal cancer molecular landscape have used self-reported ethnicity instead of genetic ancestry. METHODS: We used tumor and matched normal whole-exome sequencing data from 16,388 patients with stage I to IV colorectal cancer to investigate colorectal cancer's germline and somatic molecular landscape and the potential influence of socioeconomic factors (Distressed Communities Index, DCI) across diverse genetic ancestries. Genetic ancestry determined via supervised local ancestry inference included African (AFR, N = 1,697), Native American (AMR, N = 1,291), East Asian (EAS, N = 2,247), European (EUR, N = 9,726), Levantine Middle Eastern (LME, N = 1,192), and South Asian (SAS, N = 184). RESULTS: Microsatellite instability (MSI) was the most common form of hypermutation (80.8%), higher in the EUR genetic ancestry than in the AFR, AMR, and EAS genetic ancestry. Among germline findings, positive results were most common in high-penetrance genes associated with Lynch syndrome. Enrichment patterns included MLH1 (SAS) and PMS2 (AFR). There were significant differences in the frequency of driver mutations in APC, BRAF, KRAS, TP53, and PIK3CA between the EUR and other ancestry groups in both MSI and microsatellite stable tumors. Mutational signatures suggested enrichment of reactive oxygen species and POLE in AFR, colibactin in EAS, and aflatoxin and NTHL1 in SAS. DCI scores differed by ancestry (higher distress in AFR/AMR than in EUR), but driver mutation frequencies did not vary across DCI quintiles. CONCLUSIONS: Genetic ancestry shapes hereditary risk, tumor biology, and environmental exposures. IMPACT: These findings suggest that incorporating ancestry into screening, trials, and precision oncology may improve equity, though outcome-linked prospective studies and implementation research are warranted.

Aged

Genetic Ancestry and Colorectal Cancer in the All of Us Dataset.

IMPORTANCE: Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited. OBJECTIVE: To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model. DESIGN, SETTING, AND PARTICIPANTS: This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome. EXPOSURES: Genetically inferred ancestry categories and principal components. MAIN OUTCOMES AND MEASURES: Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost). RESULTS: Among 316 624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172 327 [54.4%] of European ancestry; 191 705 female [61.2%]; 121 585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated. CONCLUSIONS AND RELEVANCE: In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.

Aged

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+ approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+ led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+ also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+ approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+ can be integrated into existing R pipelines for commercial methylation arrays, such as 450 K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation

Assessment of Genetic Correlations Between Tobacco or Alcohol Use and Neurodegenerative Diseases Using East Asian Genetic Ancestry Genome-Wide Association Study Results.

Alzheimer's disease (AD) and Parkinson's disease (PD) are the most prevalent late-onset neurodegenerative diseases worldwide. Both are influenced in part by genetic factors and are currently incurable. Tobacco and alcohol, the two most common substances used among the general adult population, are potential AD/PD risk factors and are also heritable. Although important progress has been made, most existing research on the genetics of AD and PD has been carried out in individuals of European genetic ancestry. Investigations in a broad range of groups are crucial to understand disease mechanisms. Given the current availability of ancestry-specific tobacco and alcohol use as well as AD and PD genome-wide association study summary statistics, we performed global and local genetic correlation analyses using East Asian datasets. Genes within the correlated genetic regions were subsequently used to identify potentially enriched biological pathways between substance use and neurodegenerative diseases. We identified a global genetic correlation between smoking cessation and PD, which we confirmed in complementary European genetic ancestry data. Gene set enrichment analyses highlighted potentially shared genetic mechanisms between breast cancer and AD, which warrants further exploration. This work aims to promote further analyses across genetic ancestry groups.

Female

AncestryGeni: a novel genetic ancestry classification pipeline for small and noisy sequence data.

MOTIVATION: Efforts to address health disparities are often limited by the lack of robust computational tools for inferring genetic ancestry by calculating an individual's genetic similarity to continental groups. We have already shown that a preferred alternative to self-described race is using ancestry-informative markers (AIMs) that can be classified into ancestral components and used to estimate their similarity to those of known populations to identify continental groups. However, real-world genomic data can present challenges, including limited availability of germline DNA, a small number of AIMs for each sample, and the use of different variant calling software, limiting the application of existing solutions. RESULTS: Here, we describe a novel supervised machine-learning tool AncestryGeni, which infers genetic ancestry for samples with even a hundred markers and is applicable to any genomic data, including whole exome sequencing (WES) and RNA sequencing (RNA-Seq) data. Applying AncestryGeni to a real-world genomic dataset obtained from the Multiple Myeloma Research Foundation (MMRF) CoMMpass study, we show that it is more accurate than the commonly used FastNGSadmix when using nonstandard genomic material. We also demonstrate that when using AncestryGeni, the tumor-derived sequence obtained from WES and RNA-Seq can be a robust data source to accurately estimate an individual's genetic similarity to a continental group. AVAILABILITY AND IMPLEMENTATION: AncestryGeni pipeline is available at https://github.com/eelhaik/AncestryGeni/tree/main.

Humans

DiscoDivas: Leveraging genetic ancestry continuum information to interpolate PRS for admixed populations.

The relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits polygenic risk score (PRS) development and equitable clinical translation for admixed populations. Under the assumption that the most informative PRS model for a genetically homogeneous sample varies linearly in an ancestry continuum space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, genetic ancestries, leveraging multiple PRS models fine-tuned within existing samples, which are mostly of single ancestry, and genetic distance. DiscoDivas treats genetic ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar genetic ancestry samples. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed individuals.

PRS harmonization

The impact of sex, age, and genetic ancestry on DNA methylation across tissues.

Understanding the consequences of individual DNA methylation variation is crucial for advancing our knowledge of human biology and disease, yet the collective impact of individual traits on DNA methylation and their downstream effects on gene expression across human tissues remains poorly understood. Here, we quantify the contributions of sex, age, genetic ancestry, and BMI on autosomal DNA methylation variation across nine human tissues and 424 individuals from the Genotype-Tissue Expression project. We show that genetic ancestry and age have a greater impact on DNA methylation compared with sex, with aging effects being more widespread but less pronounced. On average, <10% of the gene expression variation in sex, age, and ancestry is mediated by DNA methylation differences, with ancestry showing the largest proportion of mediation. We further show that ancestry-associated DNA methylation differences accumulate at CpG sites with extreme methylation states and are largely under genetic control. The female autosomal genome exhibits consistent hypermethylation across tissues at Polycomb-repressed regions. Ultimately, we show that age-related Polycomb target hypermethylation is observed across multiple tissues but not in the gonads. Our multi-individual, multitissue approach defines the key drivers of human DNA methylation variation in healthy conditions, establishing a baseline for the interpretation of DNA methylation changes in disease contexts.

Humans

Reduced FOXP3 expression and its association with genetic ancestry in neuromyelitis Optica spectrum disorder: a Colombian cohort.

BACKGROUND: Autoimmune disorders are characterized by impaired immune tolerance, largely mediated by CD4&#x207a; regulatory T cells (Tregs), whose function depends on the transcription factor FOXP3. OBJECTIVES: To compare FOXP3 expression levels between patients with neuromyelitis optica spectrum disorder (NMOSD) and healthy controls from Bogot&#xe1;, Colombia, and to explore their association with genetic ancestry. METHODS: FOXP3 expression was quantified from peripheral blood mononuclear cells using RNA-based analysis. Genomic ancestry proportions were estimated using ancestry-informative markers. RESULTS: FOXP3 mRNA expression was significantly reduced in NMOSD patients compared with controls (&#x223c;2.4-fold decrease in median expression; p&#x202f;=&#x202f;0.027). Lower FOXP3 mRNA expression was associated with optic neuritis (p&#x202f;=&#x202f;0.0059), but not with historical or current AQP4-IgG seropositivity. In a sensitivity analysis restricted to patients with documented historical AQP4-IgG seropositivity, the direction of reduced FOXP3 mRNA expression was preserved but did not reach statistical significance. In pre-specified exploratory ancestry-related analyses, African (p&#x202f;=&#x202f;0.035) and Amerindian ancestry (p&#x202f;=&#x202f;0.012) were associated with FOXP3 mRNA expression levels. CONCLUSIONS: Reduced FOXP3 mRNA expression in Peripheral Blood Mononuclear Cells (PBMCs) suggests an altered immune regulatory profile in NMOSD, potentially involving mechanisms beyond antibody-mediated immunity. The association between genetic ancestry and FOXP3 mRNA expression suggests that population-specific genetic background may influence immune regulatory pathways. These findings should be interpreted as exploratory and require validation in larger, clinically homogeneous cohorts.

Adolescent

Is there are relationship between polymorphisms TSHR gene frequencies and genetic ancestry markers in patients with Primary Congenital Hypothyroidism?

The literature shows a correlation between ethnicity and pathogenic variants of the thyroid stimulating hormone receptor (TSHR) gene. Some of these polymorphisms may be risk factors for the development of primary congenital hypothyroidism (PCH). In this study, we investigated the relationship between the frequency of TSHR gene polymorphisms and the genetic influence of African, Amerindian, and European ancestry-informative markers in patients from an Amazonian population in Brazil who were diagnosed with PCH. The study was conducted on samples from 106 patients who were diagnosed with PCH. Genomic DNA was isolated from peripheral blood samples, and 10 exons from the TSHR gene were automatically sequenced. Ancestry-informative marker identification was performed using a panel of 48 markers, and the results were compared with parental Amerindian, Western European, and Sub-Saharan African populations using Structure v2.3.4 software. Four nucleotide alterations were identified among 49 patients. The distribution of tested ancestry markers among the 106 patients indicated a significant difference in the percentages of Amerindian (25.90 %), European (41.80 %), and African (32.20 %) ancestry. Logistic regression analysis revealed no significant association between the rs2075179 and rs1991517 polymorphisms and genetic ancestry. This study revealed no evidence of a relationship between polymorphic TSHR gene variants and genetic ancestry markers in patients with PCH.

Journal Article

The paradoxical extinction: Exploring signatures of assortative mating as a possible mechanism that maintains canonical Red Wolf genetic ancestry in the American Gulf Coast canids.

Admixed genomes, particularly those with an evolutionary history of genetic exchange with an endangered or extinct species, are valued for innovative and unconventional conservation actions. Here, we show the substantial conservation value that the admixed canids of the Gulf Coast have as they retain high amounts of contemporary Red Wolf ancestry and unique genetic variation of past Red Wolf lineages (e.g. ghost ancestry). We analyzed 54,439 loci genotyped across the genome of 413 North American canids and investigated the role that assortative mating with respect to ancestry proportions played in the retention of endangered genetic variation. We report high correlations of inter-chromosomal ancestry proportions that varied with geographic location along Texas and Louisiana Gulf Coast populations, with the stronger signatures reported in the latter. We found that models of assortative mating promoted greater ancestry variance compared with random mating leading to increased efficiency of selection for Red Wolf and ghost alleles. Despite the Red Wolf being extinct in the wild, original, and ghost genomic variation persists in Gulf Coast admixed canids. We suggest two conservation strategies that value and preserve this unique and endangered genomic variation through designed breeding programs. Ultimately the incorporation of this ghost genetic variation would be valuable to boost the genetic viability of the ex situ Red Wolf breeding program, create in situ redundancy, and avoid extinction for this endemic American wolf species.

Animals

Genetic Ancestry and Carrier Variant Frequency Enrichment in a Colombian Andean Population: Insights From the Eje Cafetero.

Colombia is one of the most genetically diverse populations in Latin America, and its demographic process has promoted the persistence and local enrichment of deleterious alleles, increasing the frequency of autosomal recessive disorders, particularly in semi-isolated Andean populations such as the Eje Cafetero. However, exome-based reference data from this region remain scarce, limiting ancestry-aware variant interpretation and carrier screening strategies. We aimed to characterize the ancestry proportions of this population using exome data, and to estimate the carrier frequency and distribution of pathogenic and likely pathogenic (P/LP) variants in clinically relevant recessive genes. We conducted a cross-sectional study with whole-exome sequencing (WES) in 316 unrelated individuals from the Colombian Eje Cafetero. P/LP variants were evaluated in 454 genes associated with autosomal recessive disorders. The global ancestry proportions were estimated using a validated panel of 250 exome-compatible ancestry-informative markers. Carrier frequencies were compared against Non-Finnish Europeans (NFE) and Admixed Americans (AMX) from gnomAD v4. The cohort showed predominant European ancestry (mean 51%), followed by Native American (36%) and African (13%) components. We identified 151 carriers of 89 distinct pathogenic variants across autosomal recessive genes. The most frequent variants were SERPINA1 c.863A>T (5.5%), CFTR c.1210-11T>G (3.5%), and PYGM c.1094C>T (1.5%). Also, recurrent variants were significantly enriched compared with both NFE and AMX populations, supporting regional founder effects. This study represents one of the most comprehensive exome-based genetic characterizations of the Colombian Eje Cafetero, revealing ancestry-specific enrichment of clinically relevant autosomal recessive variants driven by founder effects.

Female

Multi-ancestry genetic architecture of heart failure subtypes.

Heart failure (HF) affects 6.7 million people in the US and includes two major subtypes, HF with reduced ejection fraction (HFrEF) and HF with preserved ejection fraction (HFpEF), with distinct genetic architectures. We meta-analyze genome-wide association studies (GWAS) of 38,781 HFrEF cases, 38,163 HFpEF cases, and 526,135 controls across European, African, Hispanic, and Asian ancestries using the Million Veteran Program and Vanderbilt University DNA Databank (BioVU). We identify 46 genome-wide significant loci for HFrEF (9 novel) and 3 loci for HFpEF (1 novel). Four HFrEF loci are detected in African ancestry participants near CD36, SPI1, TRIM48, and SPNS3, with lead SNPs showing low risk-allele frequencies in European populations. In the all-cause HF meta-analysis (200,070 cases, 2,076,466 controls), we identify 136 loci (12 novel). Gene-based tests, tissue enrichment, transcriptome-wide association, and fine-mapping implicate vascular, metabolic, and TGF-&#x3b2;/Smad signaling pathways and nominate candidate causal genes, clarifying shared and subtype-specific risk across ancestries.

Humans

Large-scale admixture mapping in the All of Us Research Program improves the characterization of cross-population phenotypic differences.

Admixed individuals have been understudied in medical research largely due to their complex genetic ancestries. However, the consideration of admixture can identify ancestry-enriched genetic associations, delineating genetic underpinnings of cross-population phenotypic variation. Here, we performed admixture mapping in individuals with inferred admixture from African and European populations (N&#x2009;=&#x2009;48,921). Across 22 traits, we identified 71 ancestry-trait associations, including loci where ancestral haplotypes explained phenotypic variation yet were missed by single-variant association testing due to their stricter multiple testing burden. One such locus where inferred local AFR ancestries are associated with increased hemoglobin A1c (HbA1c) was 12q14.3, highlighting its potential role in explaining differences between populations. Together, our results expand upon the phenotypic differences between populations and characterize loci where genetic ancestries play a critical role in the architecture of disease.

Humans

Large-scale admixture mapping in the All of Us Research Program improves the characterization of cross-population phenotypic differences.

Admixed individuals have largely been understudied in medical research due to their complex genetic ancestries. However, the consideration of admixture can help identify ancestry-enriched genetic associations, delineating some of the genetic underpinnings of cross-population phenotypic variation. To this end, we performed local ancestry inference within the All of Us Research Program to identify individuals with recent admixture between African (AFR) and European (EUR) populations (N=48,921). We identified evidence of local AFR ancestry enrichment at the HLA locus, suggestive of putative selection since admixture. Furthermore, we performed the largest admixture mapping (ADM) efforts in AFR-EUR Admixed individuals for 22 traits, identifying 71 associations between inferred local AFR ancestries and a trait. Variants from published GWAS could only account for 18 (25%) of the ADM associations, highlighting novel loci where ancestral haplotypes explained some phenotypic variation. Previous studies likely have not identified these loci due to the low availability of high-powered GWAS in populations genetically similar to AFR. One such loci was 9q21.33, associated with 1.4-fold risk of end-stage kidney disease (ESKD) for carriers of inferred local AFR ancestries at the region. This locus contains the gene SLC28A3, which has previously been linked to kidney function but has never been associated with cross-population ESKD prevalence differences. Together, our results expand upon the existing literature on phenotypic differences between populations, highlighting loci where genetic ancestries play a critical role in the genetic architecture of disease.

Journal Article

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

Genetic analysis in African ancestry populations reveals genetic contributors to lung cancer susceptibility.

Striking disparities in lung cancer exist, with Black/African American individuals disproportionately affected by lung cancer, yet the genetic architecture in African ancestry individuals is poorly understood. We aimed to address this by performing a comprehensive genetic association study of lung cancer, incorporating local ancestry, across 6,490 African ancestry individuals (2,390 individuals with lung cancer and 4,100 control subjects). We identified a single genome-wide significant (p < 5 &#xd7; 10-8) locus, 15q25.1 (lead SNP rs17486278, OR [95% CI] = 1.34 [1.23-1.45], p = 4.52 &#xd7; 10-12), that has consistently shown a strong association with lung cancer across populations. Additionally, we identified nine suggestive (p < 1 &#xd7; 10-6) loci. Four of these loci (3p12.1, 8q22.2, 14q11.2, and 18q22.3) have no prior reported associations with lung cancer. We performed a multi-ancestry lung cancer meta-analysis using prior large-scale summary statistics from European and Asian ancestry populations, incorporating our African ancestry results. The meta-analysis identified 17 genome-wide significant loci, including an association with locus 4q35.2 (p = 1.22 &#xd7; 10-8), a genomic region that has been previously linked to forced expiratory volume. Genome-wide SNP-based heritability for lung cancer was 16% among African ancestry individuals. Follow-up in silico functional analyses identified genetically regulated gene expression (GReX) of nine genes (AC012184.3, ADK, CCDC12, CHRNA3, EML4, PSMA4, SNRNP200, TMEM50A, and ZYG11A) associated with lung cancer risk and biological pathways relevant to cancer and lung function. Cumulatively, these findings further elucidate the genetic architecture of lung cancer in African ancestry individuals, confirming prior loci and revealing new loci.

Female