Search PubMedSearch

Biomedical subjects

Ivan A Alexandrov

Publications and source records attributed to Ivan A Alexandrov.

2 recordsLinked to original sources

A complete genome for the common marmoset.

The common marmoset is a New World monkey widely used to study primate evolution and human disease. We present a telomere-to-telomere (T2T) reference assembly for the species, plus three near-T2T haplotypes. These resolve previously inaccessible regions, including the centromeres, sex chromosomes, subterminal satellites, acrocentric chromosomes, and the major histocompatibility complex (MHC). We find marmoset centromeres carry dimeric alpha satellites with chromosomal specificity, flanked by inactive layers interpreted as ancestral centromere remnants. We assemble gene-poor, satellite-rich short arms of the acrocentrics and find that most can harbor rDNA and all share pseudo-homolog regions (PHRs). PHR-sharing chromosomes also share closely related centromeric satellites, consistent with a model of ongoing rDNA-facilitated recombinational exchange between heterologous chromosomes. We further identify over 500 marmoset-lineage-specific transcribed genes with previously unknown transcript models or expansions. These resources, along with a preliminary pangenome, improve the utility of the marmoset as a model organism and address gaps in primate genome evolution.

Animals

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article