Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10 250 participants from the HostSeq cohort.
Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10 250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score ± SD: 0.012 ± 0.052), Black (0.013 ± 0.042), White (0.055 ± 0.022), and South Asian (0.057 ± 0.098) participants, whereas higher scores among Hispanic (0.083 ± 0.060) and Middle Eastern or Central Asian (0.122 ± 0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.