Search PubMedSearch

SEARCH · Search PubMed

Results for “Ancestry Continuum”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

9 recordsLinked to original sources

DiscoDivas: Leveraging genetic ancestry continuum information to interpolate PRS for admixed populations.

The relatively low representation of admixed populations in both discovery and fine-tuning individual-level datasets limits polygenic risk score (PRS) development and equitable clinical translation for admixed populations. Under the assumption that the most informative PRS model for a genetically homogeneous sample varies linearly in an ancestry continuum space, we introduce a Genetic Distance-assisted PRS Combination Pipeline for Diverse Genetic Ancestries (DiscoDivas) to interpolate a harmonized PRS for diverse, especially admixed, genetic ancestries, leveraging multiple PRS models fine-tuned within existing samples, which are mostly of single ancestry, and genetic distance. DiscoDivas treats genetic ancestry as a continuous variable and does not require shifting between different models when calculating PRS for different ancestries. We generated PRS with DiscoDivas and the current conventional method, i.e. fine-tuning multiple GWAS PRS using the matched or similar genetic ancestry samples. DiscoDivas generated a harmonized PRS of the accuracy comparable to or higher than the conventional approach, with the greatest advantage exhibited in admixed individuals.

PRS harmonization

Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Ancestry Continuum

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans

Divergent Biological Consequences of APOE Isoforms Across Industrialized and Non-Industrial Environments.

The apolipoprotein ε4 (APOE ε4) isoform directly alters cholesterol and immune biology and is associated with an increased risk of neurodegenerative and cardiometabolic disease in industrialized settings; nevertheless, APOE ε4-which is ancestral in humans-has persisted over evolutionary time. One potential explanation is that the costs and benefits of APOE ε4 were significantly different in the environments in which humans evolved compared to those we experience today. In support, previous work has suggested that living in a high pathogen environment, engaging in high levels of physical activity, or eating a low fat diet can dampen the detrimental effects of APOE ε4, and has revealed positive effects for fertility. However, direct tests of whether APOE isoforms are associated with different biological outcomes in non-industrial versus industrialized contexts are lacking. Working with the Turkana of Kenya and the Orang Asli of Peninsular Malaysia-two Indigenous groups in which individuals of shared ancestry span a continuum of subsistence, non-industrial to urban, industrialized lifestyles-we investigated how APOE genotypes impact cholesterol, immunological, and reproductive traits and tested for genotype x environment (GxE) interactions. First, we confirmed established genotype effects across lifestyles, showing that more APOE ε4 alleles are associated with higher total cholesterol, higher LDL cholesterol, and lower HDL cholesterol. Second, we tested for lifestyle interactions, finding lifestyle-dependent effects of genotype on innate immune biomarkers in the Orang Asli but not Turkana. Finally, we show that more APOE ε4 alleles are correlated with an extended reproductive lifespan, however this effect is relatively weak, is not consistent across populations, and does not correspond with a higher reproductive output. Together, our study provides evidence that industrialized environments can modify the biology of APOE ε4; however, we find that APOE ε4 is not universally beneficial in non-industrial contexts, highlighting the role of local environmental variation in determining its specific costs and benefits.

Journal Article

Genomic reconstruction of the Pakistani Roma reveals dual South Asian ancestry, medieval bottlenecks, and the early dispersal routes of the Romani people.

The Roma people represent one of the largest and most historically enigmatic diasporas in Eurasia, illuminating human migration patterns and cultural resilience across continents. Despite extensive research on European Roma as the diaspora endpoint, the genetic legacy of their putative South Asian source populations remains critically underexplored, leaving fundamental gaps in understanding the pre-diaspora demographic structure and early dispersal dynamics. This study uniquely positions Pakistani Roma as a potential ancestral reservoir, offering a rare window into the pre-migration phase distinct from derived European Roma populations shaped by centuries of post-dispersal admixture. We analyze 82 Pakistani Roma from Punjab using high-resolution genome-wide SNP data and comprehensive mitochondrial haplogroup profiling to reconstruct their genetic origins, population structure, and historical trajectory. Analyses reveal a dual ancestry profile comprising 50-82% Indus Valley related, 20-30% Onge related, and up to 26% Steppe derived components, with three distinct subgroups exhibiting varying affinities along a South Asian to Central Western Eurasian continuum reflecting jati-like endogamy. A severe demographic bottleneck ~800 years ago coincides with medieval socio-political upheavals, while major Eurasian admixture is dated to ~660 years ago. Mitochondrial haplogroups H (45.12%) and M (26.83%) underscore dual maternal influences from West and South Eurasia. Pakistani Roma retain substantially higher South Asian ancestry than their European counterparts, establishing them as a genetically distinct population preserving the ancestral pre-diaspora state. These findings redefine the Romani origin narrative and underscore the critical value of understudied South Asian minorities in reconstructing complex human migration pathways and diaspora formation mechanisms.

Humans

Absence of triradius d in a three-generation pedigree and other variations of main-line D.

An example is reported of a rare dermatoglyphic variant (absence of triradius d) in a woman of mixed European and Cherokee American Indian ancestry. This variant was not present in her parents, her five siblings, four nephews or one niece. Attention is drawn to the continuum from an absent triradius d to a triradius with an abbreviated main-line associated with either an open field in interdigital area IV, or a loop in interdigital area IV or a tented arch at d. This same continuum occurs at c. The absent triradius at d is extremely rare and the tented arch at d is very rare.

Dermatoglyphics

Repeatable Genomic Outcomes Along the Speciation Continuum: Insights From Pine Hybrid Zones (Genus Pinus).

Hybridization is a widespread evolutionary process and a key source of evolutionary novelty. Despite intensive study, the extent to which hybridization is deterministic and repeatable, particularly in recurrent contact events involving the same species under varying ecological conditions, remains unclear. Here, we investigated three replicated contact zones between Scots pine (Pinus sylvestris) and dwarf mountain pine (Pinus mugo) in Central Europe: two occurring in peatland habitats and one in a contrasting sandstone outcrop. Using genome-wide SNP genotyping of over 1300 individuals, we analysed genomic structure, diversity, and ancestry patterns across these zones. All sites revealed pervasive hybridization, dominated by later-generation hybrids and a notable scarcity of pure P. mugo. Across environments, hybrid populations exhibited strikingly consistent genomic compositions, with asymmetric introgression strongly biased toward P. mugo ancestry, suggesting that hybrid genome structure may follow predictable patterns under similar ecological conditions and could be shaped by cytonuclear incompatibilities. Nonetheless, we also detected site-specific differences in hybrid diversity and phenotype, highlighting the influence of local environmental selection on shared hybrid genomic backgrounds. We provide genomic evidence that Pinus uliginosa, a morphologically distinct peat bog pine traditionally regarded as a relict and endangered species is instead a partially stabilised hybrid lineage. Its genome reflects incomplete hybridization and ecological filtering, yet it lacks sufficient genetic divergence to be recognised as a distinct species. Together, these results provide evidence for the repeatability of hybridization processes, which result in the formation of phenotypes reflecting a species continuum subjected to strong environmental pressures. The findings support the simplification of taxonomic nomenclature within the Pinus mugo complex, informing adaptive conservation strategies and the genetic management of hybrid lineages.

Hybridization, Genetic

Dynamic clustering of genomics cohorts beyond race, ethnicity-and ancestry.

BACKGROUND: Recent decades have witnessed a steady decrease in the use of race categories in genomic studies. While studies that still include race categories vary in goal and type, these categories already build on a history during which racial color lines have been enforced and adjusted in the service of social and political systems of power and disenfranchisement. For early modern classification systems, data collection was also considerably arbitrary and limited. Fixed, discrete classifications have limited the study of human genomic variation and disrupted widely spread genetic and phenotypic continuums across geographic scales. Relatedly, the use of broad and predefined classification schemes-e.g. continent-based-across traits can risk missing important trait-specific genomic signals. METHODS: To address these issues, we introduce a dynamic approach to clustering human genomics cohorts based on genomic variation in trait-specific loci and without using a set of predefined categories. We tested the approach on whole-exome sequencing datasets in ten cancer types and partitioned them based on germline variants in cancer-relevant genes that could confer cancer type-specific disease predisposition. RESULTS: Results demonstrate clustering patterns that transcend discrete continent-based categories across cancer types. Functional analysis based on cancer type-specific clusterings also captures the fundamental biological processes underlying cancer, differentiates between dynamic clusters on a functional level, and identifies novel potential drivers overlooked by a predefined continent-based clustering. CONCLUSIONS: Through a trait-based lens, the dynamic clustering approach reveals genomic patterns that transcend predefined classification categories. We propose that coupled with diverse data collection, new clustering approaches have the potential to draw a more complete portrait of genomic variation and to address, in parallel, technical and social aspects of its study.

Humans

Genetic diversity and molecular mechanisms in hypertrophic cardiomyopathy: toward personalized therapy.

Hypertrophic cardiomyopathy (HCM) is the most common inherited cardiac muscle disorder, yet contemporary genomic and mechanistic research still lacks a cohesive model explaining how diverse genetic architectures give rise to heterogeneous phenotypes. This review synthesizes advances across sarcomeric and nonsarcomeric mutations, including intermediate-effect variants, polygenic modifiers, and ancestry-dependent sources of variant misclassification to elucidate how these factors govern disease penetrance and clinical expression. It critically evaluates how genetic diversity intersects with key molecular pathways, including sarcomeric hypercontractility, calcium dysregulation, mitochondrial energy deficiency, and transforming growth factor-β (TGF-β) and protein kinase B (AKT)/mammalian target of rapamycin (mTOR) signaling, to drive hypertrophic and fibrotic remodeling. Emerging mechanism-based therapies, such as myosin inhibition, allele-specific silencing, clustered regularly interspaced short palindromic repeats (CRISPR)-based correction, and metabolic modulation, are examined with respect to their capacity to modify upstream molecular drivers rather than downstream hemodynamic consequences. Persistent challenges, including variants of uncertain significance classification, ancestry-biased databases, inequitable access to genetic testing, and unresolved safety concerns for gene-based therapies, are critically assessed as major barriers to precision-medicine integration. By linking genetic architecture, molecular pathogenesis, and targeted interventions, this review advances a contemporary, mechanistically grounded framework that informs both individualized management and future research directions. Future research should prioritize pathway-specific therapeutics, functional and mechanistic validation of emerging variants, deeper physiologic phenotyping to refine disease modeling, and accelerate translation throughout the continuum of HCM pathophysiology.

Humans