Search PubMedSearch

SEARCH · Search PubMed

Results for “gene copy number variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Pervasive positive selection on X-linked ampliconic genes in primates.

Mammalian sex chromosomes harbour ampliconic gene families, which are multi-copy genes with ≥97% sequence identity, predominantly expressed in testis tissue and essential for male fertility. The amplification of testis-specific genes is conserved across mammals, yet the specific gene families that expand show striking lineage-specific variation. Previous studies suggest a dynamic turnover with adaptive evolution for several of these families, but their analysis has been limited by the quality of reference genomes of repetitive regions. To characterise the molecular evolutionary processes of ampliconic gene families on both sex chromosomes, we analysed telomere-to-telomere genome assemblies from eight primate species spanning 25 million years of evolution. We identified 53 X-linked and 19 Y-linked ampliconic gene families with dynamic copy number variation. Gene conversion through palindromic pairing and tandem arrays maintained high sequence similarity despite accumulating mutations. X-linked families maintained conserved chromosomal positions despite copy number changes, whereas Y-linked families showed frequent positional turnover. Strikingly, multiple X-linked families (GAGE, SSX, CSAG, and VCX) showed pervasive positive selection across the primate phylogeny and multiple (MAGEB, CT45, HSFX) showed lineage specific positive selection. Y-linked families predominantly evolve under purifying selection. Examining intraspecific copy number variation of the X-linked ampliconic families in chimpanzees, humans, and gorillas, we found variation among individuals but clear differences between species, with the largest families varying the most. These patterns could suggest that sperm competition, meiotic drive, or dosage-dependent selection drive the rapid, lineage-specific evolution of testis-expressed ampliconic genes in primates.

Journal Article

Genomic insights into karyotype evolution and adaptive mechanisms in Polygonaceae species.

Polygonaceae, with ecological versatility and global distribution, is an ideal system for investigating plant adaptation. However, the genomic mechanisms underlying its karyotype evolution and environmental resilience remain unclear. We herein present chromosome-level genomes of 11 species from 10 Polygonaceae genera. Our analyses reveal that Gypsy retrotransposons are key drivers of genome size variations in Polygonaceae. We reconstructed a Polygonaceae ancestral karyotype comprising 28 proto-chromosomes and elucidated evolutionary trajectories via extensive chromosomal rearrangements. Furthermore, we constructed a cross-genus super pan-genome for Polygonaceae, identifying 80,055 gene families, of which 9,845 (12.30%) are core gene families. Private genes are found to contribute significantly to interspecific differences in adaptability. Notably, gene copy number variations are identified as a critical factor influencing adaptations to diverse niches involving species-specific increases in metabolic pathways. This study provides a genomic framework for Polygonaceae karyotype plasticity and adaptive innovation, offering insights into plant evolution under environmental challenges.

Karyotype

Anticancer drug response prediction integrating multi-omics pathway-based difference features and multiple deep learning techniques.

Individualized prediction of cancer drug sensitivity is of vital importance in precision medicine. While numerous predictive methodologies for cancer drug response have been proposed, the precise prediction of an individual patient's response to drug and a thorough understanding of differences in drug responses among individuals continue to pose significant challenges. This study introduced a deep learning model PASO, which integrated transformer encoder, multi-scale convolutional networks and attention mechanisms to predict the sensitivity of cell lines to anticancer drugs, based on the omics data of cell lines and the SMILES representations of drug molecules. First, we use statistical methods to compute the differences in gene expression, gene mutation, and gene copy number variations between within and outside biological pathways, and utilized these pathway difference values as cell line features, combined with the drugs' SMILES chemical structure information as inputs to the model. Then the model integrates various deep learning technologies multi-scale convolutional networks and transformer encoder to extract the properties of drug molecules from different perspectives, while an attention network is devoted to learning complex interactions between the omics features of cell lines and the aforementioned properties of drug molecules. Finally, a multilayer perceptron (MLP) outputs the final predictions of drug response. Our model exhibits higher accuracy in predicting the sensitivity to anticancer drugs comparing with other methods proposed recently. It is found that PARP inhibitors, and Topoisomerase I inhibitors were particularly sensitive to SCLC when analyzing the drug response predictions for lung cancer cell lines. Additionally, the model is capable of highlighting biological pathways related to cancer and accurately capturing critical parts of the drug's chemical structure. We also validated the model's clinical utility using clinical data from The Cancer Genome Atlas. In summary, the PASO model suggests potential as a robust support in individualized cancer treatment. Our methods are implemented in Python and are freely available from GitHub (https://github.com/queryang/PASO).

Deep Learning

A nonlinear multi-omics data integration and classification model based on pathway self-attention and graph convolutional networks.

The abundance of omics data has significantly advanced the development of multi-omics data integration techniques. Non-linear embedding approaches for data integration have gradually become the mainstream in multi-omics research, as these approaches can substantially improve cancer analysis by enhancing the quality of the embeddings. However, current multi-omics data integration methods are typically confined to omics measurements, neglecting domain-specific prior knowledge encompassing biological pathways. In this study, we proposed a multi-omics integrated classification model, PathTransGCN, based on pathway self-attention and graph convolutional networks (GCN). The model integrated biological pathway information into multi-omics data analysis with the aim of enhancing the accuracy of cancer classification. Multi-omics data for breast cancer (BRCA), non-small cell lung cancer (NSCLC), and low-grade glioma (LGG) were obtained from The Cancer Genome Atlas (TCGA) and UCSC Xena databases. These data included gene mutations, DNA methylation, copy number variations, and gene expression, and were used to assess the model's generalizability across different cancers. First, PathTransGCN employed a pathway self-attention module to learn latent representations of samples across different pathways, thereby obtaining multi-omics integration vectors. Concurrently, a patient similarity network (PSN) was constructed using the similarity network fusion (SNF) approach. Second, the integrated vectors and the PSN were jointly fed into a GCN for end-to-end training, enabling precise classification of cancer subtypes. Through multi-omics data analysis of the BRCA dataset, PathTransGCN outperformed several popular algorithms (such as MoGCN and DeePathNet) in the five-class classification of cancer subtypes, achieving an accuracy rate of 87.6% and an F1 score of 86.4%. Moreover, the model demonstrated robust generalization capabilities across both NSCLC and LGG datasets, while effectively identifying key disease-associated biomarkers at the pathway level. Experimental results demonstrate that PathTransGCN exhibits outstanding performance in integrating omics data and delivering interpretable classification outcomes, presenting significant potential for clinical applications.

Humans

Graph-based pan-genome reveals structural and functional diversity across oil palm domestication gradients.

BACKGROUND: Oil palm (Elaeis guineensis Jacq.), the world's most land-efficient oil crop, underpins global vegetable oil supply yet faces mounting constraints from limited expansion, climate stress, and disease pressure. These challenges highlight the urgent need for genomic resources that capture species-wide diversity to support sustainable improvement. While recent reference assemblies have advanced trait discovery, single linear genomes fail to represent the full spectrum of structural and gene-content variation, limiting resolution of agronomic alleles. RESULTS: Here, we constructed a graph-based pan-genome from 30 diverse oil palm assemblies representing wild, semi-domesticated, and commercial accessions. We characterized structural variants, gene presence-absence variation, and copy-number gains, with focusing on functional stratification and resistance gene dynamics. The graph-based pan-genome revealed extensive structural and gene-content variation, including a large conserved core, complemented by shell and unique fractions enriched or biased toward regulatory, stress-responsive, and defense-related functions. Structural variation and duplication-derived copy-number gains contributed substantially to gene-content diversity, with semi-domesticated accessions exhibiting the greatest variability. Resistance gene repertoires showed contrasting patterns: receptor-like kinases remained comparatively stable, whereas the CNL subclass of NLR genes contributed disproportionately to shell-genome variation and duplication-associated turnover. CONCLUSIONS: This graph-based pan-genome provides a curated multi-assembly reference and comparative framework for oil palm genomics. By capturing structural variants, gene-content variations, copy-number gains, and resistance gene dynamics across domestication gradients, it establishes a foundation for future pan-GWAS analysis, functional genomics, and molecular breeding strategies aimed at improving resilience and productivity in this globally important crop.

Arecaceae

Extrachromosomal DNA-Driven Oncogene Dosage Heterogeneity Promotes Rapid Adaptation to Therapy in MYCN-Amplified Cancers.

UNLABELLED: Extrachromosomal DNA (ecDNA) amplification enhances intercellular oncogene dosage variability and accelerates tumor evolution by violating foundational principles of genetic inheritance through its asymmetric mitotic segregation. Spotlighting high-risk neuroblastoma, we demonstrate how ecDNA amplification undermines the clinical efficacy of current therapies in cancers with extrachromosomal MYCN amplification. Integrating theoretical models of oncogene copy number-dependent fitness with single-cell ecDNA quantification and phenotype analyses, we reveal that ecDNA copy-number heterogeneity drives phenotypic diversity and determines treatment sensitivity through mechanisms unattainable by chromosomal oncogene amplification. We demonstrate that ecDNA copy number directly influences cell fate decisions in cancer cell lines, patient-derived xenografts, and primary neuroblastomas, illustrating how extrachromosomal oncogene dosage-driven phenotypic diversity offers a strong evolutionary advantage under therapeutic pressure. Furthermore, we identify senescent cells with reduced ecDNA copy numbers as a source of treatment resistance in neuroblastomas and outline a strategy for their targeted elimination to improve the treatment of MYCN-amplified cancers. SIGNIFICANCE: ecDNA-driven tumor genome evolution provides a major challenge to curative cancer therapies. We demonstrate that ecDNA copy-number dynamics drives treatment resistance by promoting oncogene dosage-dependent phenotypic heterogeneity in MYCN-amplified cancers. Exploiting phenotype-specific vulnerabilities of ecDNA cells, therefore, presents a powerful strategy to overcome treatment resistance. See related commentary by Korsah, p. 1979.

Humans

Integrating mutation, copy number, and gene expression data to identify driver genes of recurrent chromosome-arm losses.

Aneuploidy is a hallmark of cancer, yet the genes driving recurrent chromosome-arm losses remain largely unknown. We present a systematic framework integrating mutation, copy number, and gene expression data to identify candidate driver genes of cancer type-specific recurrent chromosome-arm losses across 20 cancer types, using ∼7,500 tumors from The Cancer Genome Atlas. By analyzing focal deletions and point mutations that co-occur, or are mutually exclusive, with chromosome-arm losses, we pinpoint 322 candidate drivers associated with 159 recurring events. Our approach identifies known aneuploidy drivers such as TP53 and PTEN, while revealing multiple additional candidates, including tumor suppressors not previously linked to aneuploidy. We leverage expression changes associated with chromosome-arm losses to propose cancer-promoting pathway-level alterations. Integrating these findings highlights key candidate drivers that underlie the observed expression alterations, reinforcing their biological relevance. We provide a comprehensive catalog of candidate driver genes for recurrently lost chromosome-arms in human cancer.

Humans

Comprehensive identification and evolutionary analysis of the Wnt gene family in bivalves: Insights into the larval development of the noble scallop Chlamys nobilis.

The Wnt gene family regulates fundamental developmental processes in metazoans, but its evolutionary composition and developmental deployment in bivalves remain largely unresolved. Here, we performed a comparative genomic analysis of Wnt genes in 19 bivalve species and examined developmental expression profiles in the noble scallop Chlamys nobilis, with Crassostrea gigas and Chlamys farreri used for cross-species comparison. A total of 235 Wnt genes were identified and assigned to 12 subfamilies. No reliable Wnt3 ortholog was detected in any analyzed bivalve, supporting the view that Wnt3 loss occurred early during lophotrochozoan evolution rather than representing a lineage-specific absence. Most Wnt proteins retained the conserved WNT domain, indicating strong structural conservation, whereas lineage-specific copy-number variation and gene loss were observed among species. C. farreri and C. gigas each retained 12 Wnt genes and lacked Wnt3, whereas C. nobilis lacked Wnt3, Wnt7, and Wnt16. Developmental transcriptome analysis and RT-qPCR revealed clear stage-specific expression patterns. In C. gigas, Wnt2/10/A were highly expressed during earlydevelopment and peaked around the D-shaped larval stage, while Wnt8 and Wnt11 showed distinct stage-specific peaks. By contrast, Wnt1/5/6/9 were more active during later larval development or juvenile formation. These results provide a comparative framework for bivalve Wnt evolution and identify candidate Wnt genes potentially involved in larval development and aquaculture-relevant developmental transitions.

Animals

ELViS: an R package for estimating copy number levels of viral genomic segments at base-resolution.

MOTIVATION: Tumor viruses account for ∼10% of cancer diagnoses. Virally induced tumorigenesis is understood as direct signaling through oncogenes such as E6 and E7 genes in the case of human papillomavirus. Furthermore, pathogen characteristics such as viral oncogene dose may impact the disease course. To our knowledge, no tool has been proposed to assess the intra-viral copy number alterations that define the gene dose of viral oncogenes and associated suppressive pathways native to the pathogen's normal life cycle. RESULTS: We propose an R package, "ELViS," that analyzes viral copy number changes from DNA sequencing of whole viral genomes. The method adjusts for viral load with 2D transformation and segmentation to offer the relative viral gene doses. AVAILABILITY AND IMPLEMENTATION: The ELViS R package is available from https://bioconductor.org/packages/ELViS. This article used controlled access data from dbGaP (phs001713.v1.p1).

Software

Evidence for variation in the number of functional gene copies at the AmaR locus in Chinese hamster cell lines.

The hypothesis of functional hemizygosity has been examined for the alpha-amanitin resistant (AmaR, a codominant marker) locus in a series of Chinese hamster cell lines. AmaR mutants were obtained from different cell lines, e.g., CHO, DHW, M3- 1 and CHO-Kl, at similar frequencies. After fractionation of different RNA polymerase activities in the extracts by chromatographic procedures, the sensitivity of the mutant RNA polymerase II towards alpha-amanitin was determined. While all of the RNA polymerase II activity in mutant CHO and CHO-Kl lines became resistant to alpha-amanitin inhibition, only about 50% of the activity is highly resistant in AmaR mutants of CHW and M3- 1 cell lines. The remaining activity in the latter cell lines shows alpha-amanitin sensitivity similar to that seen with the wild-type enzyme. This behaviour is similar to that observed with a 1:1 mixture of resistant and sensitive enzymes from CHO cells. These results, therefore, strongly indicate that while only one functional copy of the gene affected by alpha-amanitin is present in CHO and CHO-Kl cells, two copies of this gene are functional in the CHW and M3-1 cell lines.

Amanitins

Comprehensive molecular profiling of multiple myeloma identifies refined copy number and expression subtypes.

Multiple myeloma is a treatable, but currently incurable, hematological malignancy of plasma cells characterized by diverse and complex tumor genetics for which precision medicine approaches to treatment are lacking. The Multiple Myeloma Research Foundation's Relating Clinical Outcomes in Multiple Myeloma to Personal Assessment of Genetic Profile study ( NCT01454297 ) is a longitudinal, observational clinical study of newly diagnosed patients with multiple myeloma (n = 1,143) where tumor samples are characterized using whole-genome sequencing, whole-exome sequencing and RNA sequencing at diagnosis and progression, and clinical data are collected every 3 months. Analyses of the baseline cohort identified genes that are the target of recurrent gain-of-function and loss-of-function events. Consensus clustering identified 8 and 12 unique copy number and expression subtypes of myeloma, respectively, identifying high-risk genetic subtypes and elucidating many of the molecular underpinnings of these unique biological groups. Analysis of serial samples showed that 25.5% of patients transition to a high-risk expression subtype at progression. We observed robust expression of immunotherapy targets in this subtype, suggesting a potential therapeutic option.

Humans

SegMantX: A Novel Tool for Detecting DNA Duplications Uncovers Prevalent Duplications in Plasmids.

Segmental duplications play an important role in genome evolution via their contribution to copy-number variation, gene-family diversification, and the emergence of novel functions. The detection of segmental duplications is challenging due to heterogeneous amelioration of sequence similarity among duplicates, which hinders the reconstruction of continuous sequence alignment. Here we introduce SegMantX, a novel approach for the identification of diverged segmental duplications in prokaryote genomes using local alignment chaining. In this approach, local alignments resulting from a preliminary sequence similarity search (e.g. BLASTn) are chained into continuous segments. Evaluating the performance of SegMantX using simulated sequences shows that the tool can detect diverged duplications beyond the sensitivity limits of standard alignment-based methods. Applying SegMantX to 6,784 enterobacterial plasmids, we find that 65% plasmids contain duplicated regions and gene duplications, most of which correspond either to dispersed, noncoding regions or duplicated mobile genetic elements (MGEs; e.g. transposons and insertion sequences). Furthermore, we demonstrate the applicability of SegMantX for the identification of diverged gene transfers between replicons and plasmid hybridization events. Our findings highlight MGEs as drivers of segmental duplications in plasmid evolution, leading to the amplification of their cargo genes, including antibiotic resistance genes. SegMantX provides a powerful framework for reconstructing diverged segmental duplications and other alignment problems.

Plasmids

Genetic and phenotypic diversity of wine-associated Hanseniaspora species.

The genus Hanseniaspora includes apiculate yeasts commonly found in fruit- and fermentation-associated environments. Their genetic diversity and evolutionary adaptations remain largely unexplored despite their ecological and oenological significance. This study investigated the phylogenetic relationships, genome structure, selection patterns, and phenotypic diversity of Hanseniaspora species isolated primarily from Australian wine environments, focusing on Hanseniaspora uvarum, the most abundant non-Saccharomyces yeast in wine fermentation. A total of 151 isolates were sequenced, including long-read genomes for representatives of the main phylogenetic clades. Comparative genomics revealed ancestral chromosomal rearrangements between the slow-evolving lineage (SEL) and fast-evolving lineage (FEL) that could have contributed to their evolutionary split, as well as significant loss of genes associated with mRNA splicing, chromatid segregation and signal recognition particle protein targeting in the FEL. Pangenome analysis within H. uvarum identified extensive copy number variation, particularly in genes related to xenobiotic tolerance and nutrient transport. Investigation into the selective landscape following the FEL/SEL divergence identified diversifying selection in 229 genes in the FEL, with significant enrichment in genes within the lysine biosynthetic pathway. Furthermore, phenotypic screening of 116 isolates revealed substantial intraspecific diversity, with specific species exhibiting enhanced ethanol, osmotic, copper, SO₂, and cold tolerance.

Wine

Hurdles to horizontal gene transfer: species-specific effects of synonymous variation and plasmid copy number determine antibiotic resistance phenotype.

Could codon composition condition the immediate success and the orientation of horizontal gene transfer? Horizontal gene transfer represents a change in the genome of expression of the transferred gene, and experimental evidence has accumulated indicating that the codon composition of a sequence is an important determinant of its compatibility with the translation machinery of the genome in which it is expressed. This suggests that codon composition influences the phenotype and the fitness conferred by a transferred gene and thus the immediate success of the transfer. To directly test this hypothesis, we characterized the resistance conferred by synonymous variants of a gentamicin resistance gene in three bacterial species: Escherichia coli, Acinetobacter baylyi and Pseudomonas aeruginosa. The strongest determinant of the resistance level conferred was the species in which the resistance gene was transferred, very likely because of important differences in the copy number of the plasmid carrying the gene. Significant differences in resistance were also found between synonymous variants within each of the three species, but more importantly, there was a strong interaction between species and variant: variants conferring high resistance in one species confer low resistance in another. However, the similarity in codon usage between the synonymous variants and the host genome only explained part of the phenotypic differences between variants in one species, P. aeruginosa. Further investigation of alternative explanations did not reveal common universal mechanisms across our three bacterial species. We conclude that codon composition can be a determinant of post-horizontal gene transfer success. However, there are multiple paths leading from synonymous sequence to phenotype, and sensitivity to these different paths is species-specific.

Gene Transfer, Horizontal

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean&#x2009;=&#x2009;0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Comparative genomics reveals lineage-associated structural variation and diversification in a barley fungal pathogen.

Leaf rust, caused by Puccinia hordei, is a major barley disease worldwide. Despite repeated shifts in virulence, contrasting reproductive histories, and emerging fungicide insensitivity, the genomic basis of its diversification and adaptation remains poorly understood. In this study, we generated haplotype-resolved, chromosome-level genome assemblies for two isolates with contrasting virulence and analyzed 41 Australian isolates collected over 54&#x2009;yr (1966-2020), integrating comparative and population genomics, mating-type gene phylogenies, chromosome-specific k-mer profiling, genome-wide copy-number variation (CNV) analysis, and gene-expression analysis. We identified a structurally dynamic chromosome characterized by repeat-associated rearrangements, structural variation, and lineage-associated CNV, representing the first evidence in a rust fungus of chromosome-scale structural diversification of this extent. Population analyses distinguished clonally expanded lineages from recombination-associated lineages, with mating-type gene phylogenies providing further support for lineage differentiation. More recently collected isolates showed increased duplication-associated variation, and CNV boundaries were associated with structural-variant breakpoints. We also identified lineage-associated amplification of Cyp51, with increased copy number associated with higher transcript abundance, supporting a potential role in fungicide adaptation. Overall, our findings highlight structural variation, contrasting reproductive histories, and lineage-associated CNV as important contributors to diversification in P. hordei, providing insights for future rust pathogen surveillance and management strategies.

Cyp51 gene

Biochemical and Genomic Underpinnings of Carotenoid Colour Variation Across a Hybrid Zone Between South Asian Flameback Woodpeckers.

Colouration and patterning have been implicated in lineage diversification across various taxa, as colour traits are heavily influenced by sexual and natural selection. Investigating the biochemical and genomic foundations of these traits therefore provides deeper insights into the interplay between genetics, ecology and social interactions in shaping the diversity of life. In this study, we assessed the pigment chemistries and genomic underpinnings of carotenoid colour variation in naturally hybridising Dinopium flamebacks in tropical South Asia. We employed reflectance spectrometric analysis to quantify species-specific plumage colouration, High-Performance Liquid Chromatography (HPLC) to elucidate the feather carotenoids of flamebacks across the hybrid zone, and Genome-Wide Association Study (GWAS) using next-generation sequencing data to uncover the genetic factors underlying carotenoid colour variation in flamebacks. Our analysis revealed that the red mantle feathers of D. psarodes primarily contained astaxanthin, with small amounts of other 4-keto-carotenoids. In contrast, the yellow mantle feathers of D. benghalense predominantly contained lutein and 3'-dehydro-lutein, alongside minor amounts of zeaxanthin, &#x3b2;-cryptoxanthin and canary-xanthophylls A and B. Hybrids with an intermediate, orange colouration deposited all of these pigments in their mantle feathers, with notably higher concentrations of carotenoids with &#x3b5;-end rings. The GWAS analysis identified the CYP2J2 gene, which plays a role in carotenoid ketolation, as associated with the expression of carotenoid colouration. Read depth data suggested variation in copy number of this gene in flamebacks. These findings contribute to the growing knowledge of avian carotenoid metabolism and highlight how genomic architecture can influence phenotypic diversity.

Animals