Search PubMedSearch

Biomedical subjects

Inanc Birol

Publications and source records attributed to Inanc Birol.

3 recordsLinked to original sources

ntSynt-viz: Visualizing synteny patterns across multiple genomes.

With the explosion of chromosome-scale genome assemblies being generated in recent years, there is vast potential for comparative genomics analyses through detecting multi-genome synteny. While existing tools can detect synteny blocks between multiple genomes, their text-based outputs make it challenging to intuitively explore large-scale synteny patterns. Interpretable, information-rich and easy-to-use synteny visualization tools are imperative to enable important biological insights from the synteny block data output by the aforementioned utilities. Here, we present ntSynt-viz, a command-line tool for automated sorting, normalization and plotting of multi-genome synteny blocks. We show how ntSynt-viz provides clearer and more easily interpretable chromosome painting ribbon plots compared to the state-of-the-art tools NGenomeSyn and plotsr when evaluating synteny between 14 human genomes, and compared to NGenomeSyn when comparing 9 hoverfly genomes. As plotsr is limited to comparing genomes with equal chromosome numbers, it was not applicable to the hoverfly dataset. Furthermore, we demonstrate how ntSynt-viz can also be applied to visualize syntenic patterns encoded in pangenome graphs, using a Minigraph-Cactus graph built from 16 Drosophila genomes. We expect that ntSynt-viz will provide crucial insights into large-scale synteny patterns between divergent genomes, thereby advancing research into key evolutionary questions.

Synteny

Population-scale disease-associated tandem repeat analysis reveals locus and ancestry-specific insights.

Tandem repeat (TR) expansions, including short TRs (motifs ≤6 bp) and variable number TRs (motifs >6 bp), underlie many monogenic disorders, with variable length and sequence influencing pathogenicity, penetrance, severity, and onset. Accurate genotype-phenotype correlation and disease prevalence estimation require characterization beyond repeat length. Here we present a population-scale analysis of 66 disease-associated TR loci using long-read assemblies from 2530 diverse haplotypes from 1265 unaffected donors. Integrating repeat length, motif composition, local ancestry, linkage disequilibrium, and phylogenetic analyses, we reveal extensive locus-, population-, and allele-specific variation shaping disease risk. Up to 8.5% of individuals carry expansions above established pathogenic thresholds, many containing interrupting motifs or sequence structures that attenuate pathogenicity. After excluding alleles from loci with uncertain disease association, non-pathogenic interrupted expansions, and carrier states inconsistent with inheritance patterns, ~4% carried expansions predicted to confer disease risk, largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses uncover population-specific TR architectures contributing to epidemiological disparities in repeat expansion disorders. Phylogenetic analyses identify conserved ancestral alleles and loci with recent instability. We describe variable linkage disequilibrium patterns and recombination signatures around specific disease-associated TR loci. Our findings emphasize integrating sequence, ancestry, and evolutionary context to understand the complex landscape of disease-associated TRs.

Humans

Concordance and divergence between self-declared ancestry and genome-derived ancestry composition in 10 250 participants from the HostSeq cohort.

Accurate characterization of human genetic diversity is essential for robust genomic analyses. We compared self-declared and genome-derived ancestry composition in 10 250 participants from the pan-Canadian HostSeq cohort using whole-genome sequencing data. Global and local ancestry were inferred at the continental super-population level using the alignment-free ntRoot algorithm and evaluated through both hard-label concordance and multiclass Brier score analyses incorporating full ancestry fraction profiles. Strong agreement was observed among East Asian / Pacific Islander (mean Brier score ± SD: 0.012 ± 0.052), Black (0.013 ± 0.042), White (0.055 ± 0.022), and South Asian (0.057 ± 0.098) participants, whereas higher scores among Hispanic (0.083 ± 0.060) and Middle Eastern or Central Asian (0.122 ± 0.034) participants reflected broader and more admixed ancestry profiles. Principal component analysis of centered log-ratio-transformed ancestry fractions revealed overlapping ancestry gradients rather than discrete continental groupings. Entropy- and dominance margin-based analyses further indicated that many discordant cases reflected diffuse admixture rather than categorical mismatch. Together, these findings support representing ancestry as a continuous compositional spectrum rather than discrete categories. Genome-derived ancestry estimates describe patterns of genomic variation and should not be interpreted as proxies for race.

Humans