Search PubMedSearch

SEARCH · Search PubMed

Results for “ancestry”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10 250 participants from the HostSeq cohort.

Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10 250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score ± SD: 0.012 ± 0.052), Black (0.013 ± 0.042), White (0.055 ± 0.022), and South Asian (0.057 ± 0.098) participants, whereas higher scores among Hispanic (0.083 ± 0.060) and Middle Eastern or Central Asian (0.122 ± 0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.

Humans

Leveraging local ancestry and cross-ancestry genetic architecture to improve genetic prediction of complex traits in admixed populations.

The broader application of polygenic risk score (PRS) is hindered by the limited transferability of PRS developed in Europeans to non-European populations. While many statistical methods have been developed to improve the performance of PRS in non-European populations, most of them focused on discrete genetic ancestry clusters and did not consider admixed individuals. Admixed individuals pose a unique challenge for PRS calculation due to the complexity of local ancestry and cross-ancestry effect sizes. Here, we present a statistical method called SDPR_admix for calculating PRS in admixed individuals. SDPR_admix characterizes the joint distribution of the effect sizes of a genetic variant with two ancestries to be both zero, ancestry enriched, or shared with correlation. SDPR_admix outperformed other methods in simulations and improved the prediction of real traits in European-African admixed individuals in UK Biobank when trained on the Population Architecture using Genomics and Epidemiology (PAGE) dataset (N = 13,000). Deployment of SDPR_admix on All of Us (N = 52,000) further increased the prediction accuracy by approximately 5-fold on average compared with training on PAGE. This enhancement was achieved with manageable computational time and cost, demonstrating the feasibility of training PRS models on large-scale All of Us data. We provided several examples demonstrating that both ancestral-enriched and shared effects, as included in the SDPR_admix prediction model, are helpful for improving polygenic prediction in admixed populations. We also applied SDPR_admix to construct PRS for admixed Americans with mixture of European and Amerindigenous ancestries and showed that SDPR_admix overall outperformed other methods.

Humans

Multi-ancestry genome-wide meta-analysis of 56,241 individuals identifies known and novel cross-population and ancestry-specific associations as novel risk loci for Alzheimer's disease.

BACKGROUND: Limited ancestral diversity has impaired our ability to detect risk variants more prevalent in ancestry groups of predominantly non-European ancestral background in genome-wide association studies (GWAS). We construct and analyze a multi-ancestry GWAS dataset in the Alzheimer's Disease Genetics Consortium (ADGC) to test for novel shared and population-specific late-onset Alzheimer's disease (LOAD) susceptibility loci and evaluate underlying genetic architecture in 37,382 non-Hispanic White (NHW), 6728 African American, 8899 Hispanic (HIS), and 3232 East Asian individuals, performing within ancestry fixed-effects meta-analysis followed by a cross-ancestry random-effects meta-analysis. RESULTS: We identify 13 loci with cross-population associations including known loci at/near CR1, BIN1, TREM2, CD2AP, PTK2B, CLU, SHARPIN, MS4A6A, PICALM, ABCA7, APOE, and two novel loci not previously reported at 11p12 (LRRC4C) and 12q24.13 (LHX5-AS1). We additionally identify three population-specific loci with genome-wide significance at/near PTPRK and GRB14 in HIS and KIAA0825 in NHW. Pathway analysis implicates multiple amyloid regulation pathways and the classical complement pathway. Genes at/near our novel loci have known roles in neuronal development (LRRC4C, LHX5-AS1, and PTPRK) and insulin receptor activity regulation (GRB14). CONCLUSIONS: Using cross-population GWAS meta-analyses, we identify novel LOAD susceptibility loci in/near LRRC4C and LHX5-AS1, both with known roles in neuronal development, as well as several novel population-unique loci. Reflecting the power of diverse ancestry in GWAS, we detect the SHARPIN locus with only 13.7% of the sample size of the NHW GWAS study (n = 409,589) in which this locus was first observed. Continued expansion into larger multi-ancestry studies will provide even more power for further elucidating the genomics of late-onset Alzheimer's disease.

Humans

Genetic Ancestry and Colorectal Cancer in the All of Us Dataset.

IMPORTANCE: Genetic ancestry may complement biological, behavioral, and clinical factors in understanding colorectal cancer (CRC) disparities; yet, ancestry-informed analyses in CRC remain limited. OBJECTIVE: To characterize associations of genetic ancestry with CRC burden, age at diagnosis, and age-specific risk, and to develop a multiethnic CRC risk-prediction model. DESIGN, SETTING, AND PARTICIPANTS: This retrospective cohort study used All of Us data from July 1986 to October 2023, with follow-up through last visit or death (median [IQR], 133.1 [57.1-186.5] months); analyses were conducted from February to June 2026. All of Us is a US research cohort with linked electronic health record (EHR) and short-read whole-genome sequencing (srWGS) data. All of Us Research Program participants with srWGS and linked EHR data were included, except those with hereditary polyposis or Lynch syndrome. EXPOSURES: Genetically inferred ancestry categories and principal components. MAIN OUTCOMES AND MEASURES: Any CRC was the primary outcome. Associations were evaluated using Fisher exact tests, cumulative incidence functions with Gray tests, cause-specific and Fine-Gray subdistribution hazard models, and pooled multivariable logistic regression. Prediction models used penalized least absolute shrinkage and selection operator and extreme gradient boosting (XGBoost). RESULTS: Among 316 624 participants (median [IQR] age, 56.3 [40.2-68.2] years; 172 327 [54.4%] of European ancestry; 191 705 female [61.2%]; 121 585 male [38.8%]), 2914 (0.9%) developed CRC. European ancestry was associated with higher odds of CRC vs all other ancestries combined (odds ratio, 1.50; 95% CI, 1.39-1.62). The median age at CRC diagnosis was older in European (63.4 [53.9-71.2] years) than in American admixed-Latino, African, East Asian, and Other ancestry groups. In cause-specific hazard models on the attained-age scale, American admixed-Latino (hazard ratio, 1.30; 95% CI, 1.14-1.47) and East Asian (hazard ratio, 1.43; 95% CI, 1.06-1.94) ancestry had higher age-specific CRC hazard than European ancestry, with consistent findings on the subdistribution scale accounting for competing death. The multiethnic XGBoost model performed best (receiver operating characteristic area under the curve, 0.898; 95% CI, 0.882-0.912; precision-recall area under the curve, 0.338; 95% CI, 0.296-0.379) and was well calibrated. CONCLUSIONS AND RELEVANCE: In this cohort study, genetic ancestry was associated with meaningful differences in CRC burden and age-specific risk. These findings suggest that a multiethnic XGBoost model may complement CRC screening as a risk-enrichment tool.

Aged

The Landscape of Genomic and Socioeconomic Variables in Patients with Colorectal Cancer Based on Genetic Ancestry.

BACKGROUND: Despite differences in tumor alterations across genetic ancestries, investigations of the colorectal cancer molecular landscape have used self-reported ethnicity instead of genetic ancestry. METHODS: We used tumor and matched normal whole-exome sequencing data from 16,388 patients with stage I to IV colorectal cancer to investigate colorectal cancer's germline and somatic molecular landscape and the potential influence of socioeconomic factors (Distressed Communities Index, DCI) across diverse genetic ancestries. Genetic ancestry determined via supervised local ancestry inference included African (AFR, N = 1,697), Native American (AMR, N = 1,291), East Asian (EAS, N = 2,247), European (EUR, N = 9,726), Levantine Middle Eastern (LME, N = 1,192), and South Asian (SAS, N = 184). RESULTS: Microsatellite instability (MSI) was the most common form of hypermutation (80.8%), higher in the EUR genetic ancestry than in the AFR, AMR, and EAS genetic ancestry. Among germline findings, positive results were most common in high-penetrance genes associated with Lynch syndrome. Enrichment patterns included MLH1 (SAS) and PMS2 (AFR). There were significant differences in the frequency of driver mutations in APC, BRAF, KRAS, TP53, and PIK3CA between the EUR and other ancestry groups in both MSI and microsatellite stable tumors. Mutational signatures suggested enrichment of reactive oxygen species and POLE in AFR, colibactin in EAS, and aflatoxin and NTHL1 in SAS. DCI scores differed by ancestry (higher distress in AFR/AMR than in EUR), but driver mutation frequencies did not vary across DCI quintiles. CONCLUSIONS: Genetic ancestry shapes hereditary risk, tumor biology, and environmental exposures. IMPACT: These findings suggest that incorporating ancestry into screening, trials, and precision oncology may improve equity, though outcome-linked prospective studies and implementation research are warranted.

Aged

Tractor workflow: a scalable Nextflow framework for local ancestry-aware genome-wide association studies.

MOTIVATION: The routine exclusion of admixed individuals from traditional genome-wide association studies (GWAS) due to concerns about spurious associations has limited multi-ancestry genetic discovery. Tractor addresses this issue by incorporating local ancestry into association testing, enabling the identification of ancestry-enriched signals and generating ancestry-specific summary statistics. However, adoption has been constrained by the complexity of prerequisite steps, including phasing and local ancestry inference, which require substantial bioinformatics expertise and introduce key analytical decision points. RESULTS: We developed a scalable, automated Nextflow workflow that integrates phasing, local ancestry inference, and Tractor association testing into a reproducible end-to-end pipeline. To demonstrate its utility, we applied the workflow to 32 blood biomarkers in 6245 two-way African-European admixed individuals from the UK Biobank. This pipeline performed efficiently at scale, replicating known associations and uncovering key ancestry-specific loci. These associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously masked genetic signals. AVAILABILITY AND IMPLEMENTATION: The workflow is modular, customizable, and compatible with commonly used phasing and local ancestry tools, minimizing manual intervention while preserving analytical flexibility. By lowering technical barriers to implementation, this framework facilitates broader adoption of local ancestry-aware GWAS, paving the way for expanded genetic discovery.

Humans

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

A genealogy-based approach for revealing ancestry-specific structures in admixed populations.

Elucidating ancestry-specific structures in admixed populations is crucial for comprehending population history and mitigating confounding effects in genome-wide association studies. Existing methods to reveal the ancestry-specific structures generally rely on frequency-based estimates of genetic relationship matrix (GRM) among admixed individuals after masking segments from ancestry components not being targeted for investigation. However, these approaches disregard linkage information between markers, potentially limiting their resolution in revealing structure within an ancestry component. We introduce ancestry-specific expected GRM (as-eGRM), a novel framework for estimating the relatedness within ancestry components between admixed individuals. The key design of as-eGRM consists of defining ancestry-specific pairwise relatedness between individuals based on genealogical trees encoded in the ancestral recombination graph (ARG) and local ancestry calls and then computing the expectation of the ancestry-specific relatedness across the genome. Comprehensive evaluations using both simulated stepping-stone models of population structure and empirical datasets based on three-way admixed Latino cohorts showed that analysis based on as-eGRM robustly outperforms existing methods in revealing the structure in admixed populations with diverse demographic histories, which in turn improves the robustness against confounding due to population structure in association testing.

Humans

Genome-Wide and Rare Variant Association Studies of Amblyopia in Admixed American and African Ancestry Groups.

OBJECTIVE: To identify genetic variants associated with amblyopia in African (AFR) and Admixed American (AMR) ancestry groups, expanding on previous studies conducted in European ancestry. DESIGN: Retrospective ancestry-stratified genome-wide association study (GWAS) and gene-level rare variant association study (RVAS). PARTICIPANTS: Participants in the All of Us Research Program from AFR and AMR ancestry groups who had whole-genome sequencing available. Cases and controls were distinguished based on the presence of International Classification of Diseases 9/10/SNOMED diagnosis codes for amblyopia in electronic health records. This yielded ancestry-stratified subsets of 269 cases and 71 585 controls of AMR ancestry and 366 cases and 79 460 controls of AFR ancestry. METHODS: Stratified logistic regression models were adjusted for age, biological sex, and the top 10 principal components of genomic ancestry. GWAS was limited to common variants (minor allele frequency &#x2265;1%), and RVAS was limited to rare variants with coding sequence-altering effects (minor allele frequency >1%, exonic only, excluding synonymous variants) aggregated at the gene level using the SKAT algorithm. Downstream analyses of the significant variants were performed using KEGG and GO pathway analysis and STRING database queries for protein-protein interactions and gene-gene interactions. MAIN OUTCOME MEASURES: Single-nucleotide polymorphisms were determined to have genome-wide significance if P < 5e-8 in the GWAS, and genes were determined to have significant association with amblyopia in the RVAS if P < 8.0 &#xd7; 10-4. RESULTS: In the AMR GWAS, 245 unique single-nucleotide polymorphisms mapping to 97 distinct loci were identified, notably within neurodevelopmental and axonal guidance genes, including ROBO1, SEMA4B, PTPRD, NRXN1, and CAMK2D. The AFR GWAS identified 11 significant variants corresponding to 6 loci mapping primarily to long noncoding RNAs and pseudogenes. The AMR RVAS identified 15 genes, including axonal transport genes (KIF1B and KIF7) and growth factor signaling genes (EGF, ERBIN, and AKAP17A). The AFR RVAS identified a single gene, DLG2, which encodes the postsynaptic protein PSD-93, which promotes the closure of the sensitive period of neuroplasticity for vision in early childhood. CONCLUSIONS: Genetic risk architectures for amblyopia differ across ancestries but fundamentally converge on neurodevelopmental signaling, cortical synapse assembly, and sensitive period plasticity rather than ocular structural dynamics. FINANCIAL DISCLOSURE(S): The authors have no proprietary or commercial interest in any materials discussed in this article.

Amblyopia

Polygenic prediction of body mass index and obesity through the life course and across ancestries.

Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5&#x2009;years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55&#x2009;kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.

Adolescent

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans

Genetic analysis in African ancestry populations reveals genetic contributors to lung cancer susceptibility.

Striking disparities in lung cancer exist, with Black/African American individuals disproportionately affected by lung cancer, yet the genetic architecture in African ancestry individuals is poorly understood. We aimed to address this by performing a comprehensive genetic association study of lung cancer, incorporating local ancestry, across 6,490 African ancestry individuals (2,390 individuals with lung cancer and 4,100 control subjects). We identified a single genome-wide significant (p < 5 &#xd7; 10-8) locus, 15q25.1 (lead SNP rs17486278, OR [95% CI] = 1.34 [1.23-1.45], p = 4.52 &#xd7; 10-12), that has consistently shown a strong association with lung cancer across populations. Additionally, we identified nine suggestive (p < 1 &#xd7; 10-6) loci. Four of these loci (3p12.1, 8q22.2, 14q11.2, and 18q22.3) have no prior reported associations with lung cancer. We performed a multi-ancestry lung cancer meta-analysis using prior large-scale summary statistics from European and Asian ancestry populations, incorporating our African ancestry results. The meta-analysis identified 17 genome-wide significant loci, including an association with locus 4q35.2 (p = 1.22 &#xd7; 10-8), a genomic region that has been previously linked to forced expiratory volume. Genome-wide SNP-based heritability for lung cancer was 16% among African ancestry individuals. Follow-up in silico functional analyses identified genetically regulated gene expression (GReX) of nine genes (AC012184.3, ADK, CCDC12, CHRNA3, EML4, PSMA4, SNRNP200, TMEM50A, and ZYG11A) associated with lung cancer risk and biological pathways relevant to cancer and lung function. Cumulatively, these findings further elucidate the genetic architecture of lung cancer in African ancestry individuals, confirming prior loci and revealing new loci.

Female

A multi-ancestry polygenic risk score for body mass index predicts longitudinal weight change.

BACKGROUND: Identifying individuals at risk for future weight gain is challenging, partly because associations with traditional clinical risk factors may be biased by confounding and reverse causation. Polygenic risk scores (PRS) provide a stable, lifelong measure of genetic predisposition to obesity. However, existing PRS have not been evaluated for their association with longitudinal weight change in adulthood and often lack generalizability across diverse genetic ancestry groups. METHODS: We conducted ancestry-specific genome-wide association study meta-analyses of body mass index (BMI) in populations of European, African or African American, Admixed American, East Asian, and South Asian ancestries and developed ancestry-specific PRS. A multi-ancestry polygenic risk score (MAPRS) was trained using ancestry-specific PRS in a model selection dataset (N&#x2009;=&#x2009;39,685) from the All of Us Research Program (AoU). We evaluated the MAPRS in an independent AoU model evaluation dataset (N&#x2009;=&#x2009;158,743) for BMI prediction and in a separate AoU test dataset (N&#x2009;=&#x2009;78,219) with repeated measurements over 1.5-2.5 years for weight change prediction. The outcomes included change in BMI and&#x2009;&#x2265;&#x2009;10% or&#x2009;&#x2265;&#x2009;5% total body weight (TBW) gain. We further examined the relationship between MAPRS and 12 clinical risk factors commonly comorbid with obesity in relation to weight change. RESULTS: The MAPRS captured 7.05% of the variance in measured BMI in the AoU model evaluation dataset and demonstrated improved generalizability across all non-European genetic ancestry groups. In the AoU test dataset, conditioned on baseline BMI at the second-to-last measurement, a one SD increase in MAPRS was associated with a 0.16 kg/m2 increase in future BMI (standard error&#x2009;=&#x2009;0.012 kg/m2; p-value&#x2009;=&#x2009;2.2&#x2009;&#xd7;&#x2009;10-39), 1.27-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;10% TBW gain (95% CI: 1.24-1.31; p-value&#x2009;=&#x2009;1.4&#x2009;&#xd7;&#x2009;10-55), and 1.15-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;5% TBW gain (95% CI: 1.13-1.18; p-value&#x2009;=&#x2009;2.8&#x2009;&#xd7;&#x2009;10-39). These associations were observed across all genetic ancestry groups and remained highly consistent after adjustment for any clinical risk factor. In contrast, most clinical risk factors demonstrated inconsistent or weaker associations with weight change outcomes. CONCLUSIONS: We developed an MAPRS for BMI that represents a robust and generalizable risk factor for longitudinal weight gain in adulthood, providing a foundation for genetically informed risk stratification and earlier, more targeted obesity prevention strategies.

Humans

Local ancestry inference identifies robust evidence of selection in Neolithic Europe.

During the European Neolithic, migrating Anatolian farmers admixed with local hunter-gatherers, coinciding with major shifts in diet, environment, and lifestyle that imposed strong selective pressures. Local ancestry inference is widely used to detect selection following admixture, but most methods were developed and validated on present-day populations. Their performance in ancient DNA - where reference panels are smaller, data are sparser, and admixture is more ancient - remains unresolved. We benchmark eight local ancestry inference methods on 176 imputed Neolithic genomes. While individual-level ancestry estimates are highly correlated across methods, inferred tract lengths and admixture time estimates vary by an order of magnitude. Overall, we recommend Gnomix or RFMix for general use. We also investigated our ability to detect natural selection using LAI. Integrating results across methods and replicating across methods and in two independent datasets (n=378 and 1,121) we identify a robust ancestry deviation at FADS1/2, consistent with adaptation on metabolism. We also identify IRAK4 (innate immunity) as a candidate locus, but with less consistent signal across methods. Finally, we replicate previous reports of excess hunter-gatherer ancestry at the HLA, but these results are inconsistent across methods and suggest that they may be affected by bias in local ancestry inference. Our findings demonstrate that while local ancestry inference recovers biologically meaningful signals in ancient genomes, results can be sensitive to the methods used for inference, particularly in complex regions like the HLA. Method choice critically influences inferred ancestry patterns and selection signals, underscoring the importance of multi-method validation.

Journal Article

DiscoDivas: Leveraging genetic ancestry continuum information to interpolate PRS for admixed populations.

The relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits polygenic risk score (PRS) development and equitable clinical translation for admixed populations. Under the assumption that the most informative PRS model for a genetically homogeneous sample varies linearly in an ancestry continuum space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, genetic ancestries, leveraging multiple PRS models fine-tuned within existing samples, which are mostly of single ancestry, and genetic distance. DiscoDivas treats genetic ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar genetic ancestry samples. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed individuals.

PRS harmonization

Tractor Workflow Pipeline: A Scalable Nextflow Framework for Local Ancestry-Aware Genome-Wide Association Studies.

The routine exclusion of admixed individuals from traditional Genome-Wide Association Studies (GWAS) due to concerns about spurious associations has hindered genetic analyses involving multiple ancestries. Tractor GWAS addresses this issue by incorporating local ancestry into its analysis, empowering identification of ancestry-enriched hits and generating ancestry-specific summary statistics. However, Tractor requires accurate genomic phasing and local ancestry inference as prerequisite steps, which requires additional bioinformatics expertise and decision points regarding reference panel setup. To streamline, harmonize, and automate this process, we present a scalable Nextflow workflow that integrates all necessary steps, minimizing the need for manual intervention while remaining modular and customizable. The workflow supports multiple commonly used tools and offers flexibility in how Tractor is implemented. To demonstrate its utility, we applied this pipeline to analyze 32 blood biomarkers in 6,245 two-way AFR-EUR admixed individuals from the UK Biobank. This pipeline ran efficiently at scale, replicated known associations, and identified novel ancestry-specific loci. These novel associations were largely driven by variants present on African ancestral tracts but absent from European tracts, underscoring the value of local ancestry-aware methods in uncovering previously missed genetic signals. By enabling the efficient analysis of admixed individuals, our workflow facilitates Tractor use, paving the way for more broader genetic discovery.

Journal Article

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+&#x2009;approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+&#x2009;led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+&#x2009;also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+&#x2009;approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+&#x2009;can be integrated into existing R pipelines for commercial methylation arrays, such as 450&#xa0;K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation

Quo vadis, BGA? A collaborative EDNAP exercise on the challenges and progress in forensic biogeographical ancestry inference.

There is a broad consensus that forensic tests for the prediction of externally visible characteristics (EVC) and analysis of biogeographic ancestry (BGA) of an individual are technically reliable. However, interpretation of the results and population-specific genotype distribution patterns remains challenging. EVC and BGA analyses provide valuable information for population genetics studies and as investigative leads for criminal cases, as well as for historical and contemporary identification tests. However, inaccurate or incorrect predictions, for example, from subjective bias in the interpretations made, have the potential to misdirect police investigations. The legal situation regarding EVC and BGA testing varies by country: ranging from countries where it is explicitly prohibited, to those without specific regulations on biogeographic ancestry prediction, and others that have already enacted laws governing its use. The reluctance to utilize these analyses is not only due to legal restrictions and data protection concerns, but also to initial limited sets of sufficiently comprehensive forensic DNA assays. Forensic BGA marker panels typically contain up to &#x223c;300 SNPs. This relatively small number of genetic markers, along with limited reference population data, complicates the interpretation of results from donors of unknown origin. This paper presents the results of a collaborative EDNAP study, which, for the first time, evaluated the approach to reporting EVC and BGA data between international laboratories. For the study, DNA from nine individuals with self-reported ancestry was collected and analysed using various forensic panels differing in the number and composition of ancestry-informative markers genotyped, comprising: the Precision ID mtDNA Whole Genome Panel, the VISAGE Basic Tool and the VISAGE Enhanced Tool for Appearance and Ancestry Prediction, and the Ion AmpliSeq&#x2122; PhenoTrivium Panel. To ensure full data protection, all SNP genotypes and uniparental marker haplotypes obtained were not shared with third parties. Instead, the genetic data were analysed using a range of commonly used population analysis software packages. These analysis outcomes were then distributed to twelve European forensic laboratories (both academic and law enforcement institutions), who were asked to prepare reports based on their interpretation of the phenotypes and ancestry they inferred from the analysis data. A questionnaire sent alongside the genetic information, aimed to evaluate which difficulties were encountered by the participants in processing the BGA analysis data they were given.

Humans