Search PubMedSearch

SEARCH · Search PubMed

Results for “African populations”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Low and differential polygenic score generalizability among African populations due largely to genetic diversity.

African populations are vastly underrepresented in genetic studies but have the most genetic variation and face wide-ranging environmental exposures globally. Because systematic evaluations of genetic prediction had not yet been conducted in ancestries that span African diversity, we calculated polygenic risk scores (PRSs) in simulations across Africa and in empirical data from South Africa, Uganda, and the United Kingdom to better understand the generalizability of genetic studies. PRS accuracy improves with ancestry-matched discovery cohorts more than from ancestry-mismatched studies. Within ancestrally and ethnically diverse South African individuals, we find that PRS accuracy is low for all traits but varies across groups. Differences in African ancestries contribute more to variability in PRS accuracy than other large cohort differences considered between individuals in the United Kingdom versus Uganda. We computed PRS in African ancestry populations using existing European-only versus ancestrally diverse genetic studies; the increased diversity produced the largest accuracy gains for hemoglobin concentration and white blood cell count, reflecting large-effect ancestry-enriched variants in genes known to influence sickle cell anemia and the allergic response, respectively. Differences in PRS accuracy across African ancestries originating from diverse regions are as large as across out-of-Africa continental ancestries, requiring commensurate nuance.

Humans

Genetic Basis of Pancreatic Steatosis: A Systematic Review of Comparison between African and Non-African Populations.

This systematic review compared genetic evidence of pancreatic steatosis across African and non-African populations to illuminate ancestry-specific mechanisms and precision-prevention opportunities. Following PRISMA guidelines for reporting, a search was conducted across PubMed, Scopus, Web of Science, and NHGRI-EBI GWAS Catalog for studies spanning 2011 to 31st March 2026. Eligible studies included genome-wide association studies (GWAS), polygenic risk score (PRS), and Mendelian randomization (MR) analyses that reported genetic associations with pancreatic fat phenotypes and had explicit ancestry stratification or comparison. Narrative thematic synthesis was performed due to methodological heterogeneity. Six core genetic studies (N > 120,000 participants) were included. The only multi-ethnic GWAS found a strong protective variant of African ancestry, rs73449607 (near PDX1/PLUTO), which reduced pancreatic fat (&#x3b2; = -0.67, P = 4.50 &#xd7; 10&#x207b;&#x2078;) and explained 14.3% of the variance in African Americans (versus 5.3% overall). UK Biobank race-stratified PRS analyses confirmed the lowest pancreatic fat fraction in Black participants, with the strongest HbA1c PRS-fat association in this group (&#x3c1; = 0.23, P < 0.0001). European-dominant GWAS highlighted risk loci, including FUT2 rs601338 (higher fat and chronic pancreatitis risk, OR 1.26). MR studies demonstrated causal links between genetically predicted intra-pancreatic fat deposition (IPFD) and pancreatic ductal adenocarcinoma (PDAC) (OR 2.46 per SD) but not diabetes. African-ancestry genomes confer substantial protection against pancreatic steatosis, whereas non-African genomes are enriched for risk alleles that amplify the non-alcoholic fatty pancreas disease (NAFPD)-to-PDAC cascade. These ancestry-differentiated mechanisms position NAFPD as a precision medicine target.

Africa

Malaria driven mechanisms shaping cancer risk and aggressiveness in African populations.

Malaria and cancer represent intersecting public health challenges in sub-Saharan Africa, where malaria remains endemic and cancer incidence is rapidly increasing. Emerging evidence indicates that chronic or recurrent malaria infection may influence carcinogenesis and tumour aggressiveness through complex biological mechanisms. This narrative review critically synthesizes data from PubMed, Scopus, and Web of Science to elucidate the mechanistic intersections between malaria and cancer risk, progression, and therapeutic response. The review highlights five principal axes linking malaria to oncogenesis: malaria-induced oxidative stress and chronic inflammation driving genomic instability; gut microbiome dysbiosis altering systemic immunity and tumour microenvironment; exploitation of shared molecular targets such as the endothelial protein C receptor (EPCR) and oncofetal chondroitin sulfate by Plasmodium parasites and cancer cells; cooperative interactions between malaria and oncogenic viruses like Epstein-Barr virus in lymphomagenesis; and malaria-associated vitamin D deficiency impairing immune surveillance. Furthermore, pharmacological evidence reveals that several antimalarial agents, including artemisinin derivatives, chloroquine, and quinacrine, possess anticancer properties, while some anticancer drugs exhibit antimalarial activity, underscoring opportunities for dual-action or repurposed therapeutics. The convergence of malaria and cancer biology underscores the urgent need for integrative, multidisciplinary research spanning molecular epidemiology, immunology, and pharmacology. Unveiling these mechanisms may unveil novel biomarkers and therapeutic targets, guiding context-specific interventions to reduce the disproportionate cancer burden in malaria-endemic African populations.

Humans

Age and gender profiles of HIV infection burden and viraemia: novel metrics for HIV epidemic control in African populations with high antiretroviral therapy coverage.

INTRODUCTION: To prioritize and tailor interventions for ending AIDS by 2030 in Africa, it is important to characterize the population groups in which HIV viraemia is concentrating. METHODS: We analysed HIV testing and viral load data collected between 2013-2019 from the open, population-based Rakai Community Cohort Study (RCCS) in Uganda, to estimate HIV seroprevalence and population viral suppression over time by gender, one-year age bands and residence in inland and fishing communities. All estimates were standardized to the underlying source population using census data. We then assessed 95-95-95 targets in their ability to identify the populations in which viraemia concentrates. RESULTS: Following the implementation of Universal Test and Treat, the proportion of individuals with viraemia decreased from 4.9% (4.6%-5.3%) in 2013 to 1.9% (1.7%-2.2%) in 2019 in inland communities and from 19.1% (18.0%-20.4%) in 2013 to 4.7% (4.0%-5.5%) in 2019 in fishing communities. Viraemia did not concentrate in the age and gender groups furthest from achieving 95-95-95 targets. Instead, in both inland and fishing communities, women aged 25-29 and men aged 30-34 were the 5-year age groups that contributed most to population-level viraemia in 2019, despite these groups being close to or had already achieved 95-95-95 targets. CONCLUSIONS: The 95-95-95 targets provide a useful benchmark for monitoring progress towards HIV epidemic control, but do not contextualize underlying population structures and so may direct interventions towards groups that represent a marginal fraction of the population with viraemia.

Universal Test and Treat

Alterations in DNA Methylation, Proteomic, and Metabolomic Profiles in African Ancestry Populations with APOL1 Risk Alleles.

KEY POINTS: We aimed to elucidate potential methylation, proteomic, and metabolomic mechanisms by which APOL1 variants may be linked to kidney disease. We report distinct methylation profiling between APOL1 risk allele carriers and noncarriers, many near APOL gene family. We report higher APOL1 protein and lower C18:1 cholesteryl ester in two risk allele carriers. BACKGROUND: The APOL1 high-risk haplotype has been associated with CKD and the deterioration of kidney function, particularly in populations with West African ancestry. However, the mechanisms by which APOL1 risk variants increase the risk for kidney disease and its progression have not been fully elucidated. METHODS: We compared methylation (N=3191; 715 [22%] carriers), proteomic (N=1240; 169 [14%] carriers), and metabolomic (N=6309; 674 [11%] carriers) profiles in African and Hispanic/Latino carriers of two APOL1 high-risk alleles (G1/G1, G2/G2, G1/G2) and noncarriers (G0/G0), excluding heterozygotes (G0/G1, G0/G2), from the Population Architecture using Genomics and Epidemiology Consortium and UK Biobank. In each study, the associations between the APOL1 high-risk haplotype and up to 722,719 cytosine-phosphate-guanine (CpG) sites, 2923 proteins, or 836 metabolites were estimated using covariate-adjusted linear regression models, followed by fixed-effects sample size&#x2013;weighted meta-analyses. RESULTS: Significant associations were observed between APOL1 high-risk haplotype and methylation at 52 CpG sites, with 48 located on chromosome 22 and 18 in the vicinity of APOL1&#x2013;4 and MYH9. All significant CpG sites near APOL2 were hypomethylated, whereas those near APOL3 and APOL4 were hypermethylated. APOL1-associated CpG sites were also identified in genes involved in ion transport and mitochondrial stress pathways. Sensitivity analyses indicated consistent yet attenuated effects among heterozygotes, supporting an additive effect of APOL1 risk alleles. Further analyses of the 52 CpG sites identified two near APOL4 exhibiting G1-specific effects, eight associated with CKD but none with eGFR, and three showing heterogeneity by CKD status. In addition, carrying two APOL1 risk alleles was associated with higher plasma APOL1 protein (&#x3b2;=1.12, PFDR = 2.26e-70) and lower C18:1 cholesteryl ester metabolite (Z=&#x2212;4.50, PFDR = 4.83e-3). CONCLUSIONS: Our results demonstrate differential methylation, proteomic, and metabolomic profiles associated with APOL1 high-risk haplotypes.

APOL1

Genetic Research on Cardiac Channelopathies in African and African-Descent Populations: A Scoping Review.

Cardiac channelopathies are inherited arrhythmias that can lead to sudden cardiac death. Despite Africa's extensive genomic diversity, African and African-descent populations remain underrepresented in genetic research, creating gaps in variant interpretation and clinical care. This scoping review aims to map the extent, range, and nature of genetic research on cardiac channelopathies in these populations and to identify key geographic, thematic, and methodological gaps. Using the Joanna Briggs Institute scoping review methodology and the Population-Concept-Context framework, systematic searches in PubMed, Embase, and Web of Science identified original human studies on cardiac channelopathies with genetic data. Extracted variables included study characteristics, populations, types of channelopathies, and reported genes and variants. Forty-four studies met the inclusion criteria. Most studies originated from the United States and South Africa, while West, Central, and East Africa were largely underrepresented. US Black individuals and South African individuals of continental African or African-descended ancestry (excluding populations of European descent such as Cape Afrikaner people) were the most studied groups, with other continental African groups rarely included. Long QT syndrome was the predominant focus, and SCN5A, KCNQ1, and KCNH2 were the most frequently analyzed genes. Many of the genetic variants discussed remained of uncertain significance due to limited functional validation and the underrepresentation of African genomes in reference databases. Genetic research on cardiac channelopathies in populations of African ancestry is limited, restricting variant interpretation, counseling, and risk prediction. Broader African inclusion, expanded gene screening, and functional studies are essential to improve diagnostics and promote equity in genomic medicine.

Humans

Clinical Utility of Trio Exome Sequencing in Rwandan Children With Autism Spectrum Disorder.

INTRODUCTION: Autism spectrum disorder (ASD) is a neurodevelopmental condition with substantial genetic and phenotypic heterogeneity. However, populations of African ancestry remain underrepresented in genomic studies, limiting understanding of ASD genetic architecture. This study aimed to characterize rare, clinically relevant genetic variants in a Rwandan pediatric ASD cohort using trio-based whole-exome sequencing (WES). METHODS: Trio-based WES was performed in 31 Rwandan pediatric patients with ASD (aged 2-18&#x2009;years) and their parents. Variants were analyzed using a trio-based workflow and classified according to American College of Medical Genetics and Genomics/Association for Molecular Pathology (ACMG/AMP) guidelines. RESULTS: Eleven candidate variants were identified in 9 of 31 patients, including four likely pathogenic variants and seven variants of uncertain significance. This resulted in a diagnostic yield of 12.9% (4/31), expanded to 29.0% when phenotypically concordant variants of uncertain significance were considered. Most likely pathogenic variants were identified in individuals with syndromic ASD who presented with intellectual disability, epilepsy, and global developmental delay. Likely pathogenic findings included two single nucleotide variants in GABRB3, SYNGAP1, and two copy-number variants involving the GNAS locus and chromosome 1p35.3-p35.2. CONCLUSIONS: The diagnostic yield observed in this cohort is consistent with previous trio-based WES studies of ASD. The findings support the clinical utility of WES for the genetic evaluation of ASD and underscore the need for expanded genomic studies in African populations.

Humans

Insights on the pathogenesis of type 2 diabetes as revealed by signature genomic classifiers in an African American population in the Washington, DC area.

AIMS: African Americans (AA) in the United States have a high risk of type 2 diabetes mellitus (T2DM) and suffer from disparities in the prevalence, mortality, and comorbidities of the disease compared to other Americans. The present study aimed to shed light on the molecular mechanisms of disease pathogenesis of T2DM among AA in the Washington, DC region. METHODS: We performed TaqMan Low Density Arrays (TLDA) on 24 genes of interest that belong to three categories: metabolic disease and disorders, cancer-related genes, and neurobehavioural disorders genes. The 18 genes, viz. ARNT, CYP2D6, IL6, INSR, RRAD, SLC2A2 (metabolic disease and disorders), APC, BCL2, CSNK1D, MYC, SOD2, TP53 (Cancer-related), APBA1, APBB2, APOC1, APOE, GSK3B, and NAE1 (neurobehavioural disorders), were differentially expressed in T2DM participants compared to controls. RESULTS: Our results suggest that factors including gender, smoking habits, and the severity or lack of control of T2DM (as indicated by HbA1c levels) were significantly associated with differential gene expression. APBA1 was significantly (p-value <0.05) downregulated in all diabetes participants. Upregulation of APOE and CYP2D6 genes and downregulation of the INSR gene were observed in the majority of diabetes patients. CONCLUSIONS: Tobacco smoking and gender were significantly associated with case-control differences in expression of the APBA1 and APOE genes (connected with Alzheimer's disease) and the INSR and CYP2D6 (associated with metabolic disorders). The results highlight the need for more effective management of T2DM and for tobacco smoking cessation interventions in this community, and further research on the associations of T2DM with other disease processes, including cancer and neurobehavioral pathways.

Humans

Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa.

African populations are the most diverse in the world yet are sorely underrepresented in medical genetics research. Here, we examine the structure of African populations using genetic and comprehensive multi-generational ethnolinguistic data from the Neuropsychiatric Genetics of African Populations-Psychosis study (NeuroGAP-Psychosis) consisting of 900 individuals from Ethiopia, Kenya, South Africa, and Uganda. We find that self-reported language classifications meaningfully tag underlying genetic variation that would be missed with consideration of geography alone, highlighting the importance of culture in shaping genetic diversity. Leveraging our uniquely rich multi-generational ethnolinguistic metadata, we track language transmission through the pedigree, observing the disappearance of several languages in our cohort as well as notable shifts in frequency over three generations. We find suggestive evidence for the rate of language transmission in matrilineal groups having been higher than that for patrilineal ones. We highlight both the diversity of variation within Africa as well as how within-Africa variation can be informative for broader variant interpretation; many variants that are rare elsewhere are common in parts of Africa. The work presented here improves the understanding of the spectrum of genetic variation in African populations and highlights the enormous and complex genetic and ethnolinguistic diversity across Africa.

Africa, Southern

Genetic analysis in African ancestry populations reveals genetic contributors to lung cancer susceptibility.

Striking disparities in lung cancer exist, with Black/African American individuals disproportionately affected by lung cancer, yet the genetic architecture in African ancestry individuals is poorly understood. We aimed to address this by performing a comprehensive genetic association study of lung cancer, incorporating local ancestry, across 6,490 African ancestry individuals (2,390 individuals with lung cancer and 4,100 control subjects). We identified a single genome-wide significant (p < 5 &#xd7; 10-8) locus, 15q25.1 (lead SNP rs17486278, OR [95% CI] = 1.34 [1.23-1.45], p = 4.52 &#xd7; 10-12), that has consistently shown a strong association with lung cancer across populations. Additionally, we identified nine suggestive (p < 1 &#xd7; 10-6) loci. Four of these loci (3p12.1, 8q22.2, 14q11.2, and 18q22.3) have no prior reported associations with lung cancer. We performed a multi-ancestry lung cancer meta-analysis using prior large-scale summary statistics from European and Asian ancestry populations, incorporating our African ancestry results. The meta-analysis identified 17 genome-wide significant loci, including an association with locus 4q35.2 (p = 1.22 &#xd7; 10-8), a genomic region that has been previously linked to forced expiratory volume. Genome-wide SNP-based heritability for lung cancer was 16% among African ancestry individuals. Follow-up in silico functional analyses identified genetically regulated gene expression (GReX) of nine genes (AC012184.3, ADK, CCDC12, CHRNA3, EML4, PSMA4, SNRNP200, TMEM50A, and ZYG11A) associated with lung cancer risk and biological pathways relevant to cancer and lung function. Cumulatively, these findings further elucidate the genetic architecture of lung cancer in African ancestry individuals, confirming prior loci and revealing new loci.

Female

Q&A with Mich&#xe8;le Ramsay.

Mich&#xe8;le Ramsay, PhD, is Director of the Sydney Brenner Institute for Molecular Bioscience, Professor in Human Genetics, and South African Research Chair in Genomics and Bioinformatics of African Populations at the University of the Witwatersrand, Johannesburg. While promoting research excellence in Africa and contributing to research that accurately represents African populations in global science, she supports capacity strengthening in the fields of genomics and precision medicine. Mich&#xe8;le is a founding member of the Human Heredity and Health in Africa Consortium, co-chair of the International Health Cohorts Consortium, member of the WHO Technical Advisory Group for Genomics (TAG-G), and co-chair of the Lancet Commission on Precision Health. She contributes low- and middle-income countries' perspectives to global genomics, promoting ethical, equitable and fair principles and practices.

Humans

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common &#x223c;4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

Proteome-wide association study of prostate cancer risk across populations.

There is insufficient understanding of the molecular basis of prostate cancer (PCa) across different populations. We perform a large-scale proteome-wide association study&#xa0;(PWAS) to identify proteins with genetically regulated expression in plasma to be associated with PCa risk across populations. We develop genetic prediction models for expression of 1578, 1993, 1218, and 1390 proteins for African (n&#x2009;=&#x2009;450), European (n&#x2009;=&#x2009;758), Asian (n&#x2009;=&#x2009;289), and Hispanic/Latino (n&#x2009;=&#x2009;474) males, respectively, and evaluate associations of genetically regulated protein expression with PCa risk in 19,391 PCa cases and 61,608 controls of African population, 122,188 cases and 604,640 controls of European population, 10,809 cases and 95,790 controls of Asian population, and 3931 cases and 26,405 controls of Hispanic/Latino population. We identify three, four, 15, and 73 PCa-associated proteins in African, Hispanic/Latino, Asian, and European populations, respectively, and 83 in trans-population meta-analysis. There are both pan-population and population-specific associations. Our findings provide valuable insights into etiology of PCa.

Humans

Carrying APOL1 G1 allele is associated with cardiovascular complications during COVID-19 in an admixed population.

BACKGROUND: The APOL1 G1 and G2 alleles were selected in the Sub-Saharan African population by conferring resistance to trypanosome infection. However, these alleles are associated with kidney diseases, and their role in cardiovascular complications remains uncertain. A second hit mediated by an inflammatory state is necessary for APOL1-mediated phenotypes. Thus, this cross-sectional study investigates the association of APOL1 alleles with COVID-19 outcomes such as cardiovascular complications and kidney injury in an admixed population. Whole-genome sequencing was performed for 485 patients with different outcomes from a Biobank in Southern Brazil. RESULTS: COVID-19 individuals presented median age of 51 years, 281 were hospitalized, and 10.9% had CKD previous to the infection. Global ancestry inference revealed 12.8% of African ancestry. The G1 allele frequency was 2.7% and G2 allele was 1.2%. Local ancestry inference evidenced African ancestry in the locus of APOL1 alleles. The G1 allele frequency was higher among patients with severe outcomes. The presence of this allele was associated with kidney injury (OR = 2.78; 95% CI = 1.04-7.42; p = 0.041) using a minimally adjusted model and cardiovascular complications with a minimally (OR = 4.61; 95% CI = 1.61-13.19; p = 0.004) and fully adjusted model (OR = 4.59; 95% CI = 1.41-14.96; p = 0.011). Four individuals carried two alleles (three G1/G1 and one G1/G2) and three of them progressed to severe COVID-19 developing kidney injury. CONCLUSION: APOL1 risk alleles are present in the Brazilian population due to genetic admixture and the G1 allele was associated with COVID-19 outcomes.

Humans

Blood-based biomarkers of Alzheimer's disease and neurodegeneration in an indigenous African cohort using both Simoa and NULISA platforms.

In low- and middle-income countries, Alzheimer's disease (AD) constitutes a growing public health burden. However, AD biomarkers research remains underrepresented in African populations. This study assesses core biomarkers of AD and their relevance in the African context as potential aid in clinical diagnosis. Nigerian older adults from VALIANT cohort (n&#x2009;=&#x2009;967) underwent biomarker quantification in plasma (p-tau217, GFAP, NfL, A&#x3b2;42 and A&#x3b2;40) employing both the Single Molecule Assay (Simoa, Quanterix) and Nucleic acid-Linked Immuno-Sandwich Assay (NULISA, Alamar). Biomarkers were associated with disease severity in clinical-diagnostic and clinical-biological groups, with stepwise increases of p-tau217, NfL and GFAP from cognitively unimpaired to dementia (p&#x2009;<&#x2009;0.05). Results were consistent across platforms. Comparison between sexes showed higher biomarker levels in male participants across diagnostic groups. A significant effect of apoE-E4 proteotype on p-tau217 levels, after adjusting for age and sex was identified. These findings support the application of plasma AD biomarkers in the African context and the relevance of further AD biomarker research in diverse populations.

Biomarkers

Efficient Detection and Characterization of Targets of Natural Selection Using Transfer Learning.

Natural selection leaves detectable patterns of altered spatial diversity within genomes, and identifying affected regions is crucial for understanding species evolution. Recently, machine learning approaches applied to raw population genomic data have been developed to uncover these adaptive signatures. Convolutional neural networks (CNNs) are particularly effective for this task, as they handle large data arrays while maintaining element correlations. However, shallow CNNs may miss complex patterns due to their limited capacity, while deep CNNs can capture these patterns but require extensive data and computational power. Transfer learning addresses these challenges by utilizing a deep CNN pretrained on a large dataset as a feature extraction tool for downstream classification and evolutionary parameter prediction. This approach reduces extensive training data generation requirements and computational needs while maintaining high performance. In this study, we developed TrIdent, a tool that uses transfer learning to enhance detection of adaptive genomic regions from image representations of multilocus variation. We evaluated TrIdent across various genetic, demographic, and adaptive settings, in addition to unphased data and other confounding factors. TrIdent demonstrated improved detection of adaptive regions compared to recent methods using similar data representations. We further explored model interpretability through class activation maps and adapted TrIdent to infer selection parameters for identified adaptive candidates. Using whole-genome haplotype data from European and African populations, TrIdent effectively recapitulated known sweep candidates and identified novel cancer, and other disease-associated genes as potential sweeps.

Selection, Genetic

Efficient detection and characterization of targets of natural selection using transfer learning.

Natural selection leaves detectable patterns of altered spatial diversity within genomes, and identifying affected regions is crucial for understanding species evolution. Recently, machine learning approaches applied to raw population genomic data have been developed to uncover these adaptive signatures. Convolutional neural networks (CNNs) are particularly effective for this task, as they handle large data arrays while maintaining element correlations. However, shallow CNNs may miss complex patterns due to their limited capacity, while deep CNNs can capture these patterns but require extensive data and computational power. Transfer learning addresses these challenges by utilizing a deep CNN pre-trained on a large dataset as a feature extraction tool for downstream classification and evolutionary parameter prediction. This approach reduces extensive training data generation requirements and computational needs while maintaining high performance. In this study, we developed TrIdent, a tool that uses transfer learning to enhance detection of adaptive genomic regions from image representations of multilocus variation. We evaluated TrIdent across various genetic, demographic, and adaptive settings, in addition to unphased data and other confounding factors. TrIdent demonstrated improved detection of adaptive regions compared to recent methods using similar data representations. We further explored model interpretability through class activation maps and adapted TrIdent to infer selection parameters for identified adaptive candidates. Using whole-genome haplotype data from European and African populations, TrIdent effectively recapitulated known sweep candidates and identified novel cancer, and other disease-associated genes as potential sweeps.

Journal Article