Search PubMedSearch

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Pan-genome-based resequencing of 2,320 accessions reveals structural variations and accelerates breeding advances in cultivated peanut.

The cultivated peanut is a crucial global legume crop that is essential for food security and nutrition, particularly in developing regions. However, its limited genetic variation hampers breeding progress and yield improvement. Here we constructed a graph-based pan-genome for peanut, incorporating 14 genomes that represent all 6 peanut varieties. Using this pan-genome, we genotyped 2,320 accessions, covering 88.03% of ICRISAT and 59.21% of USDA core germplasm, enriching valuable resources for genomic studies and breeding. We cataloged genomic structural variations and investigated the role of homoeologous exchanges in population divergence. Through our pan-genome approach, we overcame the challenges of genotyping posed by homoeologous exchanges and identified key genes associated with flowering and dwarfism in peanut. By integrating superior haplotypes and germplasm resources guided by the pan-genome, we further developed high-yield dwarf lines. This work provides essential genomic resources to accelerate functional gene discovery and modern peanut breeding.

Journal Article

Chromosome-scale genome remodeling in tumor evolution: Copy number alterations and structural variants as two sides of the same coin.

Chromosome-scale genomic rearrangements are a dominant force in tumor evolution. Copy-number alterations (CNAs) and structural variants (SVs) constitute two complementary axes of this process. Although detection technologies now deliver near-comprehensive catalogs, technical resolution has outpaced conceptual integration. In this review, we frame CNAs and SVs as inextricable facets of chromosomal aberrations. They reshape cancer genomes through altered gene dosage and three-dimensional regulatory rewiring. CNAs quantify the gene-dosage imbalance, yet arise through mechanistically distinct routes. Segmental CNAs typically require chromosomal breakage, and therefore often coincide with SV junctions. By contrast, whole-chromosome aneuploidy and whole-genome doubling (WGD) primarily reflect mitotic or cytokinetic failure and can occur without local breakpoints, while nevertheless reshaping the karyotypic landscape and seeding subsequent structural complexity. SVs, in turn, range from unbalanced events that alter copy number to ostensibly balanced exchanges that predominantly rewire regulatory architecture. Despite their diverse and sometimes catastrophic architectures, SVs are ultimately rooted in double-strand break formation and error-prone resolution. By integrating CNAs and SVs within a unified mechanistic and functional framework, we aim to convert catalogs into concepts and distill the organizing principles that govern tumor genome evolution.

Humans

Comparative genomics reveals lineage-associated structural variation and diversification in a barley fungal pathogen.

Leaf rust, caused by Puccinia hordei, is a major barley disease worldwide. Despite repeated shifts in virulence, contrasting reproductive histories, and emerging fungicide insensitivity, the genomic basis of its diversification and adaptation remains poorly understood. In this study, we generated haplotype-resolved, chromosome-level genome assemblies for two isolates with contrasting virulence and analyzed 41 Australian isolates collected over 54 yr (1966-2020), integrating comparative and population genomics, mating-type gene phylogenies, chromosome-specific k-mer profiling, genome-wide copy-number variation (CNV) analysis, and gene-expression analysis. We identified a structurally dynamic chromosome characterized by repeat-associated rearrangements, structural variation, and lineage-associated CNV, representing the first evidence in a rust fungus of chromosome-scale structural diversification of this extent. Population analyses distinguished clonally expanded lineages from recombination-associated lineages, with mating-type gene phylogenies providing further support for lineage differentiation. More recently collected isolates showed increased duplication-associated variation, and CNV boundaries were associated with structural-variant breakpoints. We also identified lineage-associated amplification of Cyp51, with increased copy number associated with higher transcript abundance, supporting a potential role in fungicide adaptation. Overall, our findings highlight structural variation, contrasting reproductive histories, and lineage-associated CNV as important contributors to diversification in P. hordei, providing insights for future rust pathogen surveillance and management strategies.

Cyp51 gene

Pan-Genome Analysis Reveals Local Adaptation to Climate Driven by Introgression in Oak Species.

The genetic base of local adaptation has been extensively studied in natural populations. However, a comprehensive genome-wide perspective on the contribution of structural variants (SVs) and adaptive introgression to local adaptation remains limited. In this study, we performed de novo assembly and annotation of 22 representative accessions of Quercus variabilis, identifying a total of 543,372 SVs. These SVs play crucial roles in shaping genomic structure and influencing gene expression. By analyzing range-wide genomic data, we identified both SNPs and SVs associated with local adaptation in Q. variabilis and Quercus acutissima. Notably, SV-outliers exhibit selection signals that did not overlap with SNP-outliers, indicating that SNP-based analyses may not detect the same candidate genes associated with SV-outliers. Remarkably, 29%-37% of candidate SNPs were located in a 250 kb region on chromosome 9, referred to as Chr9-ERF. This region contains 8 duplicated ethylene-responsive factor (ERF) genes, which may have contributed to local adaptation of Q. variabilis and Q. acutissima. We also found that a considerable number of candidate SNPs were shared between Q. variabilis and Q. acutissima in the Chr9-ERF region, suggesting a pattern of repeated selection. We further demonstrated that advantageous variants in this region were introgressed from western populations of Q. acutissima into Q. variabilis, providing compelling evidence that introgression facilitates local adaptation. This study offers a valuable genomic resource for future studies on oak species and highlights the importance of pan-genome analysis in understating mechanism driving adaptation and evolution.

Quercus

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

Genome-wide variation analysis of two Salvia hispanica L. genotypes and implication for associations with metabolic and adaptive traits.

BACKGROUND: Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. METHODS: Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28×, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. RESULTS: A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS ≈ 1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (Hₒ ≈ 0.71) and nucleotide diversity (π ≈ 7 × 10-3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. CONCLUSION: This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.

Copy-number variation, structural variation

Structural variants linked to Alzheimer's disease and other common age-related clinical and neuropathologic traits.

BACKGROUND: Alzheimer's disease (AD) is a complex neurodegenerative disorder with substantial genetic influence. While genome-wide association studies (GWAS) have identified numerous risk loci for late-onset AD (LOAD), the functional mechanisms underlying most of these associations remain unresolved. Large genomic rearrangements, known as structural variants (SVs), represent a promising avenue for elucidating such mechanisms within some of these loci. METHODS: By leveraging data from two ongoing cohort studies of aging and dementia, the Religious Orders Study and Rush Memory and Aging Project (ROS/MAP), we performed genome-wide association analysis testing 20,205 common SVs from 1088 participants with whole genome sequencing (WGS) data. A range of Alzheimer's disease and other common age-related clinical and neuropathologic traits were examined. RESULTS: First, we mapped SVs across 81 AD risk loci and discovered 22 SVs in linkage disequilibrium (LD) with GWAS lead variants and directly associated with the phenotypes tested. The strongest association was a deletion of an Alu element in the 3'UTR of the TMEM106B gene, in high LD with the respective AD GWAS locus and associated with multiple AD and AD-related disorders (ADRD) phenotypes, including tangles density, TDP-43, and cognitive resilience. The deletion of this element was also linked to lower TMEM106B protein abundance. We also found a 22-kb deletion associated with depression in ROS/MAP and bearing similar association patterns as GWAS SNPs at the IQCK locus. In addition, we leveraged our catalog of SV-GWAS to replicate and characterize independent findings in SV-based GWAS for AD and five other neurodegenerative diseases. Among these findings, we highlight the replication of genome-wide significant SVs for progressive supranuclear palsy (PSP), including markers for the 17q21.31 MAPT locus inversion and a 1483-bp deletion at the CYP2A13 locus, along with other suggestive associations, such as a 994-bp duplication in the LMNTD1 locus, suggestively linked to AD and a 3958-bp deletion at the DOCK5 locus linked to Lewy body disease (LBD) (P = 3.36 × 10-4). CONCLUSIONS: While still limited in sample size, this study highlights the utility of including analysis of SVs for elucidating mechanisms underlying GWAS loci and provides a valuable resource for the characterization of the effects of SVs in neurodegenerative disease pathogenesis.

Humans

Structural variation detection and association analysis of whole-genome-sequence data from 16,543 Alzheimer's disease sequencing project subjects.

INTRODUCTION: The role of structural variations (SVs) in Alzheimer's disease (AD) remains understudied. METHODS: We analyzed whole-genome sequencing data from the Alzheimer's Disease Sequencing Project (N&#xa0;=&#xa0;16,543) and identified 400,234 (168,223 high-quality) SVs. Laboratory validation yielded a sensitivity of 82% (85% for high-quality). RESULTS: We found a burden of singletons (odds ratio [OR]&#xa0;=&#xa0;1.07, p&#xa0;=&#xa0;0.0017) and homozygous deletions (OR&#xa0;=&#xa0;1.14, p&#xa0;<&#xa0;0.0001) in cases. On AD genes, we observed the ultra-rare SVs associated with the disease, including protein-altering SVs in ABCA7, APP, PLCG2, and SORL1. Twenty-one SVs are in linkage disequilibrium (LD) with known AD-risk variants, exemplified by a 5k deletion in LD (R2&#xa0;=&#xa0;0.99) with rs143080277 in NCK2. We identified a rare deletion near RNA5SP293 associated with AD (OR&#xa0;=&#xa0;1.99, p&#xa0;=&#xa0;1.3&#xa0;&#xd7;&#xa0;10-5), which was replicated using an independent dataset. DISCUSSION: This study highlights the pivotal role of SVs in AD genetics. HIGHLIGHTS: Observed a significant burden of singletons and homozygous deletions in Alzheimer's disease (AD) patients. Identified rare protein-altering structural variations (SVs) in ABCA7, APP, PLCG2, and SORL1. Established linkages between SVs and AD risk-associated single nucleotide variants (SNVs). Discovered a novel deletion near RNA5SP293 linked to AD, replicated independently. Uncovered over-representation of SVs in neuronal function pathways.

Humans

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Complex de novo structural variants are an underestimated cause of rare disorders.

Complex de novo structural variants (dnSVs) are crucial genetic factors in rare disorders, yet their prevalence and characteristics in rare disorders remain poorly understood. Here, we conduct a comprehensive analysis of whole-genome sequencing data of 12,568 families, including 13,698 offspring with rare diseases, obtained as part of the UK 100,000 Genomes Project. We identify 1,870 dnSVs, constituting the largest dnSV dataset reported to date. Complex dnSVs (n&#x2009;=&#x2009;158; 8.4%) emerge as the third most common type of SV, following simple deletions and duplications. We classify 65% of these complex dnSVs into 11 subtypes. Among probands with dnSVs (n&#x2009;=&#x2009;1,696), 9% exhibit exon-disrupting pathogenic dnSVs associated with the probands' phenotype. Notably, 12% of exon-disrupting pathogenic dnSVs and 22% of de novo deletions or duplications previously identified by array-based or whole-exome sequencing methods are found to be complex dnSVs. We also find distinct genomic properties of de novo deletions depending on the parent of origin. This study highlights the importance of complex dnSVs in the cause of rare disorders and demonstrates the necessity of specific genomic analysis to avoid overlooking these variants.

Humans

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We develop a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implement it in the SuPreMo-Akita software. With this tool, we rank hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This reveals that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritize candidate variants that disrupt ASD CREints and validate our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study suggests that disrupted genome folding is a potential genetic mechanism in a subset of ASD cases and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Humans

Aplf/Dna2 variants drive chromosomal fission and accelerate speciation in zokors.

Chromosomal fissions and fusions are common, yet the molecular mechanisms and implications in speciation remain poorly understood. Here, we confirm a fission event in one zokor species through multiple-omics and functional analyses. We traced this event to a mutation in a splicing enhancer of the DNA repair gene Aplf in the fission-bearing species, which caused exon skipping and produced a truncated protein that disrupted DNA repair. An intronic deletion in Dna2, known to facilitate neo-telomere formation when knocked out, reduced gene activity. These variants collectively drove chromosomal fission in this zokor species. The newly formed chromosome became fixed due to carrying essential genes and strong selective pressure. While geographic isolation likely initiated the divergence of this species and the sister one, the fission event and associated decline at the chromosome level in gene flow probably exacerbated the speciation process. Our work elucidates the genetic basis of chromosomal fission and underscores its role in speciation dynamics.

Multiomics

Comparative Analysis of Mammalian Adaptive Immune Loci Revealed Spectacular Divergence and Common Genetic Patterns.

Adaptive immune responses are mediated by the production of adaptive immune receptors, antibodies, and T-cell receptors, which bind antigens, thus causing their neutralization. Unlike other proteins, adaptive immune receptors are not fully encoded in the germline genome and result from a complex of somatic processes collectively called V(D)J recombination affecting germline immunoglobulin (IG) and T-cell receptor (TR) loci consisting of template genes. While various existing studies report extreme diversity of antibodies and T-cell receptors, little is known about the diversity of germline IG and TR loci. To overcome this gap, the first comparative analysis of full-length sequences of IG/TR loci across 46 mammalian species from 13 taxonomic orders was performed. First, germline gene counts were shown to correlate in immunoglobulin heavy chain immunoglobulin heavy chain (IGH)/immunoglobulin lambda (IGL) loci and T-cell receptor alpha (TRA)/T-cell receptor beta (TRB) and anticorrelate in immunoglobulin kappa (IGK)/IGL, possibly indicating coevolution between corresponding chains. Second, structures of IG/TR loci were analyzed, and it was shown that IG/TR loci formed by long arrays of high multiplicity repeats are more common for species that have experienced population bottlenecks. Finally, haplotypes of IG/TR loci with little or no sequence similarity within a species were found, suggesting that they may have a limited potential for homologous recombination. These results demonstrate that IG/TR loci are rapidly evolving genomic regions whose structural variation is shaped by the population history of the species and open new perspectives for immunogenomics studies.

Animals

PARTAGE: Parallel analysis of replication timing and gene expression.

The human genome is partitioned into functional compartments that replicate at specific times during the S-phase. This temporal program, referred to as replication timing (RT), is co-regulated with the 3D genome organization, is cell type-specific, and changes during development in coordination with gene expression. Moreover, RT alterations are linked to abnormal gene expression, genome instability, and structural variation in multiple diseases, including cancer. However, mechanistic links between RT, large-scale 3D genome architecture, and transcriptional regulation remain poorly understood. A major limitation is that current approaches require the separate profiling of RT and transcriptomes from independent batches of samples, obscuring the complex co-regulation between the epigenome and transcriptome. Here, we developed PARTAGE, a multiomics approach that enables joint profiling of copy number variation (CNV), RT, and gene expression from the same sample, providing a more accurate integrative view of the complex relationships between RT and gene regulation.

Journal Article

Graph-based pan-genome reveals structural and functional diversity across oil palm domestication gradients.

BACKGROUND: Oil palm (Elaeis guineensis Jacq.), the world's most land-efficient oil crop, underpins global vegetable oil supply yet faces mounting constraints from limited expansion, climate stress, and disease pressure. These challenges highlight the urgent need for genomic resources that capture species-wide diversity to support sustainable improvement. While recent reference assemblies have advanced trait discovery, single linear genomes fail to represent the full spectrum of structural and gene-content variation, limiting resolution of agronomic alleles. RESULTS: Here, we constructed a graph-based pan-genome from 30 diverse oil palm assemblies representing wild, semi-domesticated, and commercial accessions. We characterized structural variants, gene presence-absence variation, and copy-number gains, with focusing on functional stratification and resistance gene dynamics. The graph-based pan-genome revealed extensive structural and gene-content variation, including a large conserved core, complemented by shell and unique fractions enriched or biased toward regulatory, stress-responsive, and defense-related functions. Structural variation and duplication-derived copy-number gains contributed substantially to gene-content diversity, with semi-domesticated accessions exhibiting the greatest variability. Resistance gene repertoires showed contrasting patterns: receptor-like kinases remained comparatively stable, whereas the CNL subclass of NLR genes contributed disproportionately to shell-genome variation and duplication-associated turnover. CONCLUSIONS: This graph-based pan-genome provides a curated multi-assembly reference and comparative framework for oil palm genomics. By capturing structural variants, gene-content variations, copy-number gains, and resistance gene dynamics across domestication gradients, it establishes a foundation for future pan-GWAS analysis, functional genomics, and molecular breeding strategies aimed at improving resilience and productivity in this globally important crop.

Arecaceae

Genome-wide SNP data reveal geographic structure and landscape-associated genomic differentiation in a widespread lizard in arid Eastern Central Asia.

Arid landscapes provide important systems for examining how geographic structure and environmental heterogeneity shape genomic differentiation. In topographically complex desert regions, however, it remains challenging to determine whether population structure primarily reflects landscape resistance, geographic distance, or contemporary environmental variation. Here, we use genome-wide SNP data to investigate population structure, phylogenetic relationships, historical gene flow, demographic history, and landscape correlates of genomic differentiation in the variegated racerunner (Eremias vermiculata), a widespread lacertid lizard across arid Eastern Central Asia. Analyses of 164 individuals recovered six geographically structured nuclear clusters associated with major desert basins and mountain-bounded regions. Nuclear phylogenies resolved two broad regional clades corresponding to northeastern and southwestern parts of the species' range, while PCA and ADMIXTURE analyses recovered six finer-scale genetic clusters. Mitochondrial phylogenies, based on combined NCBI-derived Cyt b and COI sequences from the same individuals, recovered four deeper maternal lineages. These patterns indicate overall phylogeographic agreement between nuclear and mitochondrial datasets, with genome-wide SNPs providing finer-scale resolution of population structure. Demographic reconstructions further uncovered regionally heterogeneous Late Pleistocene histories among clusters, including signals of expansion, stability, and decline. Landscape genomic analyses revealed that genomic differentiation is primarily associated with landscape resistance, particularly elevation and land cover, as well as geographic distance, whereas contemporary environmental variables explained comparatively little variation after controlling for spatial structure. Together, our results suggest that genomic differentiation in E. vermiculata reflects the interplay of persistent landscape configuration, historical connectivity, and region-specific demographic histories across arid Eastern Central Asia. More broadly, this study highlights the value of integrating phylogeographic and landscape genomic approaches for understanding population differentiation and evolutionary history in topographically heterogeneous desert ecosystems.

Arid Eastern Central Asia

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99&#xd7;) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including&#x2009;~&#x2009;17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8&#x2009;&#xb1;&#x2009;8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (&#x3c0;&#x2009;=&#x2009;0.00267), followed by lowland (&#x3c0;&#x2009;=&#x2009;0.00233), whereas highland chickens showed the lowest diversity (&#x3c0;&#x2009;=&#x2009;0.00203) and elevated genomic inbreeding (FROH and FHOM &#x2248; 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals