Search PubMedSearch

SEARCH · Search PubMed

Results for “polyploid genomes”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Transposable Element Dynamics Drive the Genomic Evolution and Phenotypic Diversification of Allotetraploid Common Carp.

An important question in evolutionary biology is how polyploidization generates raw material for phenotypic diversification. Transposable elements (TEs) represent an underestimated source of genetic variation in eukaryotic genomes. By integrating 516 whole-genome resequencing datasets and 236 transcriptomes from common carp (Cyprinus carpio), a representative allotetraploid fish, we constructed the first population-scale landscape of TE insertions in teleosts. TE insertions are widespread in the carp genome and preferentially associated with stress-responsive genes, with DNA transposons as major contributors. Relaxed purifying selection and TE burst events coexist, generating abundant variation for subsequent subspecies differentiation. Compared with a closely related diploid species, carp exhibits more exonic TE insertions and shorter TE-gene distances, and multiple TE superfamilies expanded during tetraploidization. Genome-wide association analyses uncovered intragenic TE variants underlying domesticated traits missed by SNPs, including DNA transposon deletions associated with scale reduction and altered body shape. Notably, lighter-colored individuals harbor homozygous deletions of LTR and DNA transposons within mdfic2, whose knockout in zebrafish reduces pigmentation. Most trait-associated variants reflect lineage-specific loss of ancient TE insertions rather than recent transposition. Overall, these findings highlight the distinct role of TEs in polyploid genome evolution and phenotypic diversification, providing new insights into TE dynamics in vertebrates.

allotetraploidization

SpacerScope: binary-vectorized, genome-wide off-target profiling for RNA-guided nucleases without prior candidate-site bias.

The precision of CRISPR/Cas systems is fundamental to their application in plant and animal biotechnology. However, comprehensive sequence-based off-target candidate discovery remains a computational bottleneck, particularly in large and complex genomes. Here we developed SpacerScope, an off-target candidate discovery framework that enables unbiased, genome-wide discovery by leveraging binary vectorization, bitwise filtering, and right-end-anchored alignment. Benchmarking against human CIRCLE-seq data demonstrated that SpacerScope recovered 100% of validated off-target sites (6142/6142), matching the sensitivity of exhaustive algorithms. Crucially, SpacerScope achieved this maximum candidate recovery while substantially reducing computational overhead. In large-genome evaluations, SpacerScope maintained low peak memory usage of 2.20 GiB and achieved substantial runtime improvements over indel-aware comparator tools, including more than 50-fold speedup relative to Cas-OFFinder 3 (544 s versus 29 185 s). Furthermore, comparative analyses in polyploid species, such as the octoploid strawberry, revealed that SpacerScope identified larger sequence-compatible candidate burdens than standard web-based design platforms. Our results establish SpacerScope as a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes. The source code and program was publicly available at https://github.com/charlesqu666/SpacerScope. Short Abstract CRISPR/Cas sequence-based off-target candidate discovery remains computationally challenging in large, repetitive, and polyploid genomes. Existing tools either miss indel-containing candidate sites or incur prohibitive runtime and memory costs. We developed SpacerScope, a binary-vectorized framework that enables unbiased, genome-wide off-target candidate discovery without pre-selected candidate sites. By integrating bitwise filtering with right-end-anchored alignment, SpacerScope recovered 100% of validated off-target sites in human CIRCLE-seq data while using only 2.20 GiB of memory and achieving more than 10-fold speedup over indel-aware alternatives. Evaluation in plant genomes, including rice and octoploid strawberry, further demonstrated SpacerScope's capacity to identify larger sequence-compatible candidate burdens overlooked by standard tools. SpacerScope thus provides a high-speed framework for sequence-based genome-wide off-target candidate discovery across diverse and highly repetitive genomic landscapes, supporting downstream prioritization.

CRISPR-Cas Systems

Complexity and polyadenylic acid content of visna virus 60-70S RNA.

The genomic complexity of visna virus was measured by quantitative analysis of 18 RNase T1-resistant oligonucleotides from 60-70S RNA. T1-resistant oligonucleotides were separated by two-dimensional polyacrylamide gel electrophoresis. Visna virus had a genomic complexity of 3.6 X 10(6) daltons, very close to the size of a single 30-40S RNA subunit. It was therefore concluded that the visna virus genome is largely polyploid. Visna virus 60-70S RNA polyadenylic acid segment was purified by T1 RNase digestion followed by oligodeoxythymidylic acid-cellulose column chromatography. It contained over 99% AMP and had a size of about 200 nucleotides. The binding capacities on oligodeoxythymidylic acid-cellulose of native 60-70S RNA and purified 30-40S RNA subunits were examined. It was concluded that two out of three intact subunits contain a polyadenylic acid segment.

Avian Sarcoma Viruses

Genome-wide cyclin gene evolution in Arabidopsis and Brassica reveals polyploidization-driven duplication and flowering-time associations.

Cyclin genes are plant cell cycle regulators that play essential roles in growth, development, and reproduction. However, the evolutionary dynamics and genomic organization of cyclin genes across the Brassicaceae family remain poorly understood, particularly in the context of allotetraploid genome evolution. Here, we investigated the diversity, expansion mechanisms, and potential functional diversification of cyclin genes across ten Brassicaceae genomes, including four Arabidopsis and six Brassica species. A total of 1087 cyclin genes representing 23 cyclin types were identified. Comparative genomic analyses revealed that cyclin gene expansion was strongly influenced by polyploidization in Brassica species, with 1845 duplication events involving 1063 genes. Whole-genome duplication was the predominant mechanism driving expansion, while both inter- and intra-genomic duplications contributed to gene retention in tetraploid Brassica species, with the highest duplication frequency observed in Brassica juncea. Across genomes, 120 physical gene clusters were identified, including homogeneous and heterogeneous types. Ortholog analysis between progenitor and allotetraploid species identified 852 orthologous pairs involving 366 genes, indicating extensive conservation following allotetraploid formation. Phylogenetic analysis resolved cyclins into three major clades, while expression-based clustering in Brassica napus grouped genes into four major clusters, suggesting functional diversification. Integration of pan-genomic and flowering-time QTL analyses further identified two cyclin genes, Bna21cycA2 and Bna113cycD4, which contain amino acid polymorphisms and represent putative candidate variations potentially associated with flowering-time variation across multiple genomes. These findings provide new insights into the evolutionary expansion, retention, and potential functional divergence of cyclin genes in Brassicaceae and highlight candidate loci for future functional studies and crop improvement.

Evolution, Molecular

Distinct evolutionary trajectories of subgenomic centromeres in polyploid wheat.

BACKGROUND: Centromeres are crucial for precise chromosome segregation and maintaining genome stability during cell division. However, their evolutionary dynamics, particularly in polyploid organisms with complex genomic architectures, remain largely enigmatic. Allopolyploid wheat, with its well-defined hierarchical ploidy series and recent polyploidization history, serves as an excellent model to explore centromere evolution. RESULTS: In this study, we perform a systematic comparative analysis of centromeres in common wheat and its corresponding ancestral species, utilizing the latest comprehensive reference genome assembly available. Our findings reveal that wheat centromeres predominantly consist of five types of centromeric-specific retrotransposon elements (CRWs), with CRW1 and CRW2 being the most prevalent. We identify distinct evolutionary trajectories in the functional centromeres of each subgenome, characterized by variations in copy number, insertion age, and CRW composition. By utilizing CENH3-ChIP data across various ploidy levels, we uncover a series of CRW invasion events that have shaped the evolution of AA subgenome centromeres. Conversely, the evolutionary process of the DD subgenome centromeres involves their expansion from diploid to hexaploid wheat, facilitating adaptation to a larger genomic context. Integration of complete einkorn centromere assemblies and Aegilops tauschii pan-genomes further revealed subgenome-specific centromere evolutionary trajectories. By inclusion of synthetic hexaploid from S2-S3 generations, alongside 2x/6 × natural accessions, we demonstrate that DD subgenome centromere expansion represents a gradual evolutionary process rather than an immediate response to polyploidization. CONCLUSIONS: Our study provides a comprehensive landscape of centromere adaptation, evolution, and maturation, along with insights into how retrotransposon invasions drive centromere evolution in polyploid wheat.

Centromere

Genome diversity and evolution of the duckweed section Alatae comprising diploids, polyploids, and interspecific hybrids.

The section Alatae of genus Lemna of the monocotyledonous aquatic duckweed family (Lemnaceae) consists of rather diverse accessions with unknown phylogeny and unclear taxonomic assignment. In contrast to other duckweeds, some Alatae accessions, in addition to mainly vegetative propagation, produce readily flowers and viable seeds. We analyzed the genomic diversity and phylogenetic relationship of 52 Alatae accessions. For this purpose, we applied multiple molecular and cytogenetic approaches, including plastid and nuclear sequence polymorphisms, chromosome counting, genome size determination, and genomic in situ hybridization in combination with geographic distribution. We uncovered ploidy variation, recurrent hybridization, and backcrosses between species and their hybrids. The latter successfully spread over three continents. The results elucidate the evolution of Alatae accessions and explain the difficult taxonomic assignment of distinct accessions. Our study might be an example for analogous studies to resolve the hitherto unclear relationships among accessions of the duckweed genera Wolffiella and Wolffia.

Araceae

Forty new genomes shed light on sexual reproduction and the origin of tetraploidy in Microsporidia.

Microsporidia are single-celled, obligately intracellular parasites with growing public health, agricultural, and economic importance. Despite this, Microsporidia remain relatively enigmatic, with many aspects of their biology and evolution unexplored. Key questions include whether Microsporidia undergo sexual reproduction, and the nature of the relationship between tetraploid and diploid lineages. While few high-quality microsporidian genomes currently exist to help answer such questions, large-scale biodiversity genomics initiatives, such as the Darwin Tree of Life project, can generate high-quality genome assemblies for microsporidian parasites when sequencing infected host species. Here, we present 40 new microsporidian genome assemblies from infected arthropod hosts that were sequenced to create reference genomes. Out of the 40, 32 are complete genomes, eight of which are chromosome-level, and eight are partial microsporidian genomes. We characterized 14 of these as polyploid and five as diploid. We found that tetraploid genome haplotypes are consistent with autopolyploidy, in that they coalesce more recently than species, and that they likely recombine. Within some genomes, we found large-scale rearrangements between the homeologous genomes. We also observed a high rate of rearrangement between genomes from different microsporidian groups, and a striking tolerance for segmental duplications. Analysis of chromatin conformation capture (Hi-C) data indicated that tetraploid genomes are likely organized into two diploid units, similar to dikaryotic cells in fungi, with evidence of recombination within and between units. Together, our results provide evidence for the existence of a sexual cycle in Microsporidia, and suggest a model for the microsporidian lifecycle that mirrors fungal reproduction.

Genome, Fungal

From family trials to genomic mate allocation: statistical and genomic strategies to accelerate sugarcane genetic improvement.

Sugarcane (Saccharum spp.) underpins global sugar and bioenergy supply and is increasingly valued as a renewable biomass feedstock. Sustained improvement in commercial traits and resilience is constrained by long breeding cycles, clonal propagation, multi-stage testing, and a highly polyploid, heterozygous, and frequently aneuploid genome with substantial non-additive genetic variation. Genomic selection has demonstrated value for predicting elite-clone performance, yet its operational use remains limited at earlier decision points, including family selection, parent evaluation, and cross design. This review examines the biological, statistical, and genomic factors that shape these decisions, with emphasis on the Australian breeding context based on progeny assessment trials (PATs), clonal assessment trials (CATs), and final assessment trials (FATs). We evaluate challenges arising from family plot means, the use of different full-sib samples as nominal family replicates, spatial heterogeneity, competition, genotype-by-environment interaction, and the partitioning of additive and non-additive effects. We also assess the integration of pedigree and genomic relationship, genotype representation, allele-dosage estimation, aneuploidy, genomic prediction models, and training-population design. We then consider genomic prediction of cross performance and constrained mate allocation as approaches for improving expected family performance, accounting for cross-specific non-additive effects and managing relatedness. We propose a decision-centred framework that links family and clonal data across breeding stages, tracks the propagation of information and uncertainty, and supports parent recycling and cross allocation. We conclude with a practical research agenda for stage-integrated mixed-model and single-step analyses that connect early family evaluation with genomic prediction and cross-level decision support in sugarcane breeding.

Saccharum

A chromosome-level genome assembly of Coffea arabica L. var. 'Kona Typica'.

Coffea arabica L. var. 'Kona Typica' is renowned for its premium cup quality, but its vulnerability to pests and diseases limits production. To accelerate cultivar improvement, we generated a chromosome-level genome assembly of 'Kona Typica' using PacBio HiFi sequencing and Hi-C scaffolding technology. The final assembly spans 1.13 Gb, with a scaffold N50 of 50.50 Mb, organized into 22 chromosomes. BUSCO assessment indicated a high completeness at 99.1%. We annotated 65,458 protein-coding genes and identified 1,073,545 interspersed repeats, accounting for 65.16% of the genome. Analysis of transposon insertion ages revealed that most long terminal repeat retrotransposons proliferated after the polyploidization event. This high-quality genome assembly of 'Kona Typica' provides a valuable resource for exploring coffee genomic evolution and genetic mechanisms of complex traits, facilitating genomics studies and the development of improved coffee cultivars with enhanced disease resistance and quality traits.

Coffea

Adaptive evolution of polyploid crops.

Crop evolution represents a fundamental biological process through which plants respond to selection in different environments. This encompasses mechanisms operating at multiple scales of biological organization, including genetic and epigenetic regulation and higher-order interactions among molecular complexes. This Review synthesizes how polyploidy shapes crop evolution by generating duplicated genes, driving genome reorganization, altering dosage relationships and promoting regulatory divergence, which together influence crop metabolism, physiology, development and environmental responses. We focus mainly on the mechanisms underlying adaptation in polyploid crops, including the consequences of gene and genome duplication, genome reorganization and subfunctionalization. We also examine how hybridization, phenotypic plasticity and crop-microbiome interactions intersect with polyploidy to expand or constrain adaptive potential. Together, these processes affect crop survival, fitness and breeding value under changing environments. We suggest that future research connect polyploid genome architecture with experimentally validated signatures of selection and field performance to make better use of polyploidy-derived variation in crop improvement.

Polyploidy

Research Progress in the Cytogenetics of Sweetpotato and Its Wild Relatives.

Cultivated sweetpotato (Ipomoea batatas (L.) Lam.), a hexaploid (2n = 6x = 90) crop, is the most economically important species within the morning glory genus Ipomoea (Convolvulaceae). Fourteen diploid Ipomoea species and several polyploid accessions have been confirmed to be closely related to sweetpotato, often termed its wild relatives. These wild species harbor abundant elite genes beneficial to sweetpotato improvement and thereby serve as indispensable germplasm reservoirs for breeding programs. In addition, several wild taxa are proposed as potential ancestors of domesticated sweetpotato. Nevertheless, the evolutionary origin and genomic architecture of cultivated sweetpotato have not yet been fully resolved. Cytological investigations, particularly chromosome karyotyping and meiotic pairing analyses, have been pivotal in unravelling the genomic architecture and evolutionary trajectories of polyploid taxa. Herein, we systematically summarize advances in chromosome counting, genome size, karyotyping, and meiotic pairing research on sweetpotato and its wild relatives.

Ipomoea

Divergent trajectories of genome architecture and chromosome evolution in ferns and angiosperms.

Ferns and angiosperms represent the two largest vascular plant lineages but exhibit striking genomic and ecological contrasts. We investigated whether differences in genome size, chromosome architecture, GC content, and stomatal traits reveal divergent evolutionary trajectories between these lineages. We assembled the most comprehensive dataset to date, integrating genome size, chromosome number and size, GC content, and stomatal traits for over 1100 fern species and compared it with an extensive angiosperm dataset. Ferns exhibited markedly lower variability and c. 16-fold slower rates of chromosome size evolution than angiosperms. A persistent positive relationship between genome size and chromosome number in ferns suggests limited cytological post-polyploid diploidization. While ferns generally possess larger stomata, this difference disappears after accounting for genome size, indicating that nucleotypic constraints, rather than lineage-specific physiology, dictate stomatal dimensions. Both groups share a unimodal GC-genome size relationship peaking at c. 14 Gbp. Larger fern chromosomes imply lower genome-wide recombination rates, potentially limiting genetic reshuffling and adaptive potential. Our results highlight fundamentally divergent evolutionary trajectories, likely shaped by meiotic symmetry in ferns and meiotic asymmetry, possibly centromere drive, and post-polyploid diploidization in angiosperms, defining the functional and genomic landscapes of these lineages across deep evolutionary timescales.

Genome, Plant

Borrelia burgdorferi loses essential genetic elements and cell proliferative potential during stationary phase in culture but not in the tick vector.

The Lyme disease agent Borrelia burgdorferi is a polyploid bacterium with a segmented genome in which both the chromosome and over 20 distinct plasmids are present in multiple copies per cell. This pathogen can survive for at least 9 months in its tick vector in an apparent dormant state between blood meals, without losing cell proliferative capability when re-exposed to nutrients. Cultivated B. burgdorferi cells grown to stationary phase or resuspended in nutrient-limited media are often used to study the effects of nutrient deprivation. However, a thorough assessment of the spirochete's ability to recover from nutrient depletion has been lacking. Our study shows that starved B. burgdorferi cultures rapidly lose cell proliferative ability. Loss of genetic elements essential for cell proliferation contributes to the observed proliferative defect in stationary phase. The gradual decline in copies of genetic elements is not perfectly synchronized between chromosomes and plasmids, generating cells that harbor one or more copies of the essential chromosome but lack all copies of one or more non-essential plasmids. This phenomenon likely contributes to the well-documented issue of plasmid loss during in vitro cultivation of B. burgdorferi. In contrast, B. burgdorferi cells from ticks starved for 14 months showed no evidence of reduced cell proliferative ability or plasmid loss. Beyond their practical implications for studying B. burgdorferi, these findings suggest that the midgut of the tick vector offers a unique environment that supports the maintenance of B. burgdorferi's segmented genome and cell proliferative potential during periods of tick fasting.IMPORTANCEBorrelia burgdorferi causes Lyme disease, a prevalent tick-borne illness. B. burgdorferi must survive long periods (months to a year) of apparent dormancy in the midgut of the tick vector between blood meals. Resilience to starvation is a common trait among bacteria. However, this study reveals that, in laboratory cultures, B. burgdorferi poorly endures starvation and rapidly loses viability. This decline is linked to a gradual loss of genetic elements required for cell proliferation. These results suggest that the persistence of B. burgdorferi in nature is likely shaped more by unique environmental conditions in the midgut of the tick vector than by an innate ability of this bacterium to endure nutrient deprivation.

Borrelia burgdorferi

Genomic complexities of murine leukemia and sarcoma, reticuloendotheliosis, and visna viruses.

The genetic complexities of several ribodeoxyviruses were measured by quantitative analysis of unique RNase T1-resistant oligonucleotides from 60-70S viral RNAs. Moloney murine leukemia virus was found to have an RNA complexity of 3.5 x 10(6) daltons, whereas Moloney murine sarcoma virus had a significantly smaller genome size of 2.3 x 10(6). Reticuleondotheliosis and visna virus RNAs had complexities of 3.9 x 10(6), respectively. Analysis of RNase A-resistant oligonucleotides of Rous sarcoma virus RNA gave a complexity of 3.6 x 10(6), similar to that previously obtained with RNase T1-resistant oligonucleotides. Since each of these viruses was found to have a unique sequence genomic complexity near the molecular weight of a single 30-40S viral RNA subunit, it was concluded that ribodeoxyvirus genomes are at least largely polyploid.

Base Sequence

Unraveling evolutionary pathways: allopolyploidization and introgression in polyploid Prunus (Rosaceae).

Allopolyploidization, resulting from hybridization and subsequent whole-genome duplication (WGD), is a fundamental mechanism driving evolutionary diversification across various lineages within the Tree of Life. The polyploid Prunus (Rosaceae), significant for its economic and agricultural value, provides an ideal model for investigating the evolutionary dynamics associated with allopolyploidy. In this study, we utilized deep genome skimming (DGS) data to demonstrate a comprehensive analytical framework for elucidating the underlying allopolyploidy that includes a newly adapted tool (DGS-Tree2GD) tailored explicitly for accurately detecting WGD events. Additionally, we introduced two methods to evaluate the contribution of incomplete lineage sorting (ILS) to lineage diversification. Phylogenomic discordance analyses revealed that allopolyploidization, rather than ILS, played a dominant role in the origin and dynamics of polyploid Prunus. Moreover, we inferred that the uplift of the Himalayas from the Middle to Late Miocene was a key driver in the rapid diversification of the Maddenia clade, an endemic group in East Asia. This geological event facilitated extensive hybridization and allopolyploidization, particularly the introgression between the Himalayas-Hengduan and Central-Eastern China clades. This case study demonstrates the robustness and efficacy of our analytical approach in precisely identifying WGD events and elucidating the evolutionary mechanisms underlying allopolyploidization in polyploid Prunus.

Polyploidy

Genomic and evolutionary basis of parthenogenesis in a disease-vector tick species.

Haemaphysalis longicornis is an important tick species and pathogen vector characterized by the co-circulation of triploid parthenogenetic and diploid bisexual strains. However, the evolutionary basis of parthenogenesis in this species is unclear. Here we report reference-quality, haplotype-resolved genome assemblies of the parthenogenetic strain and two reference-quality genomes of the bisexual strains. Comparative genomic analysis revealed high collinearity between the parthenogenetic and bisexual genomes, with a stable chromosomal architecture maintained among the three haplotypes of the parthenogenetic strain. The parthenogenetic H. longicornis genome exhibited a major expansion in cell cycle-related gene families, including the inhibitor of apoptosis protein (IAP) family, but was characterized by a contraction in other gene families. Population resequencing of 179 individuals revealed two distinct subpopulations, with chromosome 7 harbouring high genetic differentiation and several candidate genes probably associated with parthenogenesis. Functional experiments showed that knockdown of the BIRC5 gene, a member of the IAP family, suppressed oviposition in both strains, with the parthenogenetic strain exhibiting milder adverse effects probably due to a stronger transcriptional response. Overall, our results reveal the genomic and evolutionary features associated with polyploid parthenogenesis in H. longicornis.

Animals

Cultivar-dependent regulation of cytokinin biosynthesis in wheat: developmental expression of TaIPT genes and hormonal crosstalk during reproductive development.

BACKGROUND: Cytokinins are key regulators of plant growth, reproductive development, and yield formation. In cereals, cytokinin biosynthesis is catalyzed by isopentenyltransferase (IPT) enzymes, yet the genomic organization and developmental regulation of IPT genes in polyploid wheat remain incompletely understood, especially at the cultivar level. RESULTS: Here, we present an integrated genomic, transcriptional, and hormonal analysis of the TaIPT gene family during vegetative and reproductive development in two wheat cultivars, awnless Kontesa and awned Ostka. Genome-wide analysis identified nine core TaIPT genes represented by 25 homoeologs distributed across the A, B, and D subgenomes, for which a unified nomenclature was established. Phylogenetic analysis resolved TaIPTs into conserved evolutionary clades corresponding to ATP/ADP-dependent and tRNA-dependent IPT groups. Expression profiling revealed distinct spatial and temporal patterns of TaIPT transcription across roots, leaves, inflorescences, and developing spikes. Several TaIPT genes showed enhanced expression during early reproductive stages, coinciding with dynamic changes in cytokinin concentrations. Comparative analyses revealed cultivar-specific expression and co-variation patterns, with Kontesa displaying more compartmentalized TaIPT expression and Ostka showing coordinated activation of multiple TaIPT genes during early grain development. Hormone profiling further indicated stage-dependent associations between TaIPT expression, cytokinin metabolism, and the balance between cytokinins and abscisic acid. These relationships are interpreted as correlative and provide a framework for future functional testing rather than direct evidence of causality. CONCLUSIONS: Together, these results provide a cultivar-focused framework for understanding the organization and regulation of cytokinin biosynthesis genes in wheat. The data highlight cultivar-dependent TaIPT expression patterns and their association with cytokinin dynamics during reproductive development, while also identifying the need for homoeolog-specific and functional validation. This study establishes a foundation for future research on cytokinin-mediated regulation of wheat growth and grain development.

Triticum

Contrasting regulation of protein-coding genes and lncRNA homeologs in allotetraploid Coffea arabica.

A chromosome-level Bourbon assembly revealed that protein-coding homeologs are predominantly co-regulated between subgenomes. In contrast, intergenic lncRNAs display a modest, but statistically consistent bias toward subgenome E across diverse developmental and stress contexts. Coffea arabica is an allotetraploid species derived from natural hybridization between C. canephora and C. eugenioides, which contributed the C and E subgenomes, respectively. This genomic origin poses major challenges for genome assembly, annotation, and the interpretation of gene regulation. In this study, a high-quality genome assembly of C. arabica was generated and annotated, with particular emphasis on identifying protein-coding genes and intergenic long non-coding RNAs (lincRNAs). Homeologous relationships between genes from the C and E subgenomes were established, providing a robust framework to investigate subgenomic conservation and regulatory divergence. Using an extensive collection of publicly available RNA-seq libraries spanning multiple developmental stages, tissues, and environmental conditions, the relative transcriptional contribution of each subgenome was evaluated. On a global scale, gene expression was largely balanced between subgenomes, with no consistent evidence of subgenome dominance. While protein-coding genes showed comparable regulatory behavior across subgenomes, lincRNAs exhibited a more asymmetric expression pattern, suggesting higher subgenome-specific expression that is interpreted here as a consistent directional tendency rather than as evidence of subgenome dominance. Together, these results provide new insights into the regulatory architecture of the C. arabica genome and establish a foundational genomic and transcriptomic resource for future functional studies and crop improvement efforts.

Coffea